← Search

Junhyeok Kim

9 accepted papers

2026

Anchoring and Rescaling Attention for Semantically Coherent Inbetweening

CVPR 2026

Generative inbetweening (GI) seeks to synthesize realistic intermediate frames between the first and last keyframes beyond mere interpolation. As sequences become sparser and motions larger, previous GI models struggle with inconsistent frames with unstable pacing and semantic misalignment. Since GI

Cited by 0SourcecodeScholar
2026

Real-Time Visual Attribution Streaming in Thinking Model

ICML 2026spotlight

We present an amortized framework for real-time visual attribution streaming in multimodal thinking models. When these models generate code from a screenshot or solve math problems from images, their long reasoning traces should be grounded in visual evidence. However, verifying this reliance is cha…

Cited by 0SourceScholar
2025

EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild

NAACL 2025findings

Predicting when to initiate speech in real-world environments remains a fundamental challenge for conversational agents. We introduce , a novel framework for real-time speech initiation prediction in egocentric streaming video. By modeling the conversation from the speaker’s first-person viewpoint,…

Cited by 0SourcePDFScholar
2025

Interpreting vision transformers via residual replacement model

NeurIPS 2025poster

How do vision transformers (ViTs) represent and process the world? This paper addresses this long-standing question through the first systematic analysis of 6.6K features across all layers, extracted via sparse autoencoders, and by introducing the residual replacement model, which replaces ViT compu…

Cited by 0SourceScholar
2025

See What You Are Told: Visual Attention Sink in Large Multimodal Models

ICLR 2025poster

Large multimodal models (LMMs) "see" images by leveraging the attention mechanism between text and visual tokens in the transformer decoder. Ideally, these models should focus on key visual information relevant to the text token. However, recent findings indicate that LMMs have an extraordinary tend…

Cited by 2SourcePDFScholar
2025

Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues

ACL 2025long

Nonverbal communication is integral to human interaction, with gestures, facial expressions, and body language conveying critical aspects of intent and emotion. However, existing large language models (LLMs) fail to effectively incorporate these nonverbal elements, limiting their capacity to create…

2025

Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding

CVPR 2025highlight

Visual grounding seeks to localize the image region corresponding to a free-form text description. Recently, the strong multimodal capabilities of Large Vision-Language Models (LVLMs) have driven substantial improvements in visual grounding, though they inevitably require fine-tuning and additional…

Cited by 2SourcePDFScholar
2023

Reading Books is Great, But Not if You Are Driving! Visually Grounded Reasoning about Defeasible Commonsense Norms

EMNLP 2023long main

Commonsense norms are defeasible by context: reading books is usually great, but not when driving a car. While contexts can be explicitly described in language, in embodied scenarios, contexts are often provided visually. This type of visually grounded reasoning about defeasible commonsense norms is…

Cited by 0SourcecodeScholar
2019

Vehicular Multi-Camera Sensor System for Automated Visual Inspection of Electric Power Distribution Equipment

IROS 2019poster

In this paper, we present a multi-camera sensor system along with its control algorithm for automated visual inspection from a moving vehicle. To accomplish this task, we propose a unique hardware configuration consisting of a frontal stereo vision system, six lateral cameras motorized to tilt, and…

Cited by 7SourceScholar