← Search

Jinyeong Kim

5 accepted papers

2026

FEAST: Fully Connected Expressive Attention for Spatial Transcriptomics

CVPR 2026

Spatial Transcriptomics (ST) provides spatially-resolved gene expression, offering crucial insights into tissue architecture and complex diseases. However, its prohibitive cost limits widespread adoption, leading to significant attention on inferring spatial gene expression from readily available wh

Cited by 0SourcecodeScholar
2026

Real-Time Visual Attribution Streaming in Thinking Model

ICML 2026spotlight

We present an amortized framework for real-time visual attribution streaming in multimodal thinking models. When these models generate code from a screenshot or solve math problems from images, their long reasoning traces should be grounded in visual evidence. However, verifying this reliance is cha…

Cited by 0SourceScholar
2025

Interpreting vision transformers via residual replacement model

NeurIPS 2025poster

How do vision transformers (ViTs) represent and process the world? This paper addresses this long-standing question through the first systematic analysis of 6.6K features across all layers, extracted via sparse autoencoders, and by introducing the residual replacement model, which replaces ViT compu…

Cited by 0SourceScholar
2025

See What You Are Told: Visual Attention Sink in Large Multimodal Models

ICLR 2025poster

Large multimodal models (LMMs) "see" images by leveraging the attention mechanism between text and visual tokens in the transformer decoder. Ideally, these models should focus on key visual information relevant to the text token. However, recent findings indicate that LMMs have an extraordinary tend…

Cited by 2SourcePDFScholar
2025

Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding

CVPR 2025highlight

Visual grounding seeks to localize the image region corresponding to a free-form text description. Recently, the strong multimodal capabilities of Large Vision-Language Models (LVLMs) have driven substantial improvements in visual grounding, though they inevitably require fine-tuning and additional…

Cited by 2SourcePDFScholar