← Search

Yejin Kim

16 accepted papers

2026

A More Word-like Image Tokenization for MLLMs

CVPR 2026

Modern multimodal large language models (MLLMs) typically keep the language model fixed and train a visual projector that maps the pixels into a sequence of tokens in its embedding space, so that images can be presented in essentially the same form as text. However, the language model has been optim

Cited by 0SourcecodeScholar
2026

Contact-Anchored Policies: Contact Conditioning Creates Strong Robot Utility Models

RSS 2026poster

The prevalent paradigm in robot learning attempts to generalize across environments, embodiments, and tasks with language prompts at runtime. A fundamental tension limits this approach: language is often too abstract to guide the concrete physical understanding required for robust manipulation. In t…

Cited by 0SourceScholar
2026

Enhancing Multi-Image Understanding through Delimiter Token Scaling

ICLR 2026poster

Large Vision-Language Models (LVLMs) achieve strong performance on single-image tasks, but their performance declines when multiple images are provided as input. One major reason is the cross-image information leakage, where the model struggles to distinguish information across different images. Exi…

Cited by 0SourcecodeScholar
2026

Fine-Grained Multi Image Object Hallucination Benchmark

CVPR 2026

Multimodal Large Language Models (MLLMs) are increasingly deployed in multi-image scenarios requiring complex reasoning across visual contexts. However, current MLLMs remain fundamentally limited by object hallucination--generating plausible yet factually inconsistent descriptions about objects. Exi

Cited by 0SourceScholar
2026

MolmoSpaces: Large-Scale Open Ecosystem for Robot Manipulation and Navigation

RSS 2026poster

Deploying robots at scale demands robustness to the long tail of everyday situations. The countless variations in scene layout, object geometry, and task specifications that characterize real environments are vast and underrepresented in existing robot benchmarks. Measuring this level of generalizat…

Cited by 0SourceScholar
2026

TripleSumm: Adaptive Triple-Modality Fusion for Video Summarization

ICLR 2026poster

The exponential growth of video content highlights the importance of video summarization, a task that efficiently extracts key information from long videos. However, existing video summarization studies face inherent limitations in understanding complex, multimodal videos. This limitation stems from…

Cited by 0SourcecodeScholar
2025

Measuring Representational Shifts in Continual Learning: A Linear Transformation Perspective

ICML 2025poster

In continual learning scenarios, catastrophic forgetting of previously learned tasks is a critical issue, making it essential to effectively measure such forgetting. Recently, there has been growing interest in focusing on representation forgetting, the forgetting measured at the hidden layer. In th…

Cited by 0SourcePDFScholar
2025

OVS Meets Continual Learning: Towards Sustainable Open-Vocabulary Segmentation

NeurIPS 2025poster

Open-Vocabulary Segmentation (OVS) aims to segment classes that are not present in the training dataset. However, most existing studies assume that the training data is fixed in advance, overlooking more practical scenarios where new datasets are continuously collected over time. To address this, we…

Cited by 0SourceScholar
2024

Improving Content Recommendation: Knowledge Graph-Based Semantic Contrastive Learning for Diversity and Cold-Start Users

COLING 2024main

Addressing the challenges related to data sparsity, cold-start problems, and diversity in recommendation systems is both crucial and demanding. Many current solutions leverage knowledge graphs to tackle these issues by combining both item-based and user-item collaborative signals. A common trend in…

2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration

ICRA 2024

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man

Cited by 910SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration0

ICRA 2024poster

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man…

Cited by 259SourcecodeScholar
2024

Prompts have evil twins

EMNLP 2024main

We discover that many natural-language prompts can be replaced by corresponding prompts that are unintelligible to humans but that provably elicit similar behavior in language models. We call these prompts “evil twins” because they are obfuscated and uninterpretable (evil), but at the same time mimi…

2024

SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World

CVPR 2024poster

Reinforcement learning (RL) with dense rewards and imitation learning (IL) with human-generated trajectories are the most widely used approaches for training modern embodied agents. RL requires extensive reward shaping and auxiliary losses and is often too slow and ineffective for long-horizon tasks…

2022

Don’t Judge a Language Model by Its Last Layer: Contrastive Learning with Layer-Wise Attention Pooling

COLING 2022main

Recent pre-trained language models (PLMs) achieved great success on many natural language processing tasks through learning linguistic features and contextualized sentence representation. Since attributes captured in stacked layers of PLMs are not clearly identified, straightforward approaches such…

2018

Self-Adaptive Machine Learning Operating Systems for Security Applications

ICASSP 2018accepted

This paper proposes a reliable and self-adaptive operating system management policy for CCTV-based security applications which controls arrival image compression rates. After receiving image sequences via CCTV cameras, the system enqueues the sequences of images and processes them for face recogniti…

Cited by 0SourceScholar