← Search

Sangho Lee

19 accepted papers

2026

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

CVPR 2026

Today's strongest video-language models (VLMs) remain proprietary, and the strongest open-weight models often rely on synthetic data from proprietary VLMs and do not disclose their training data or recipe. As a result, the open-source community lacks the foundations needed to improve on the state-of

Cited by 0SourcecodeScholar
2026

MolmoAct: Action Reasoning Models That Can Reason in Space

ICRA 2026poster

Reasoning is essential for purposeful action, yet most robotic foundation models map perception and instructions directly to control, limiting adaptability, generalization, and semantic grounding. We introduce Action Reasoning Models (ARMs), which integrate perception, planning, and control through …

2026

SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning

CVPR 2026

As humans, we are natural any-horizon reasoners, i.e., we can decide whether to iteratively skim long videos or watch short ones in full when necessary for a given task. With this in mind, one would expect video reasoning models to reason flexibly across different durations. However, SOTA models are

Cited by 0SourcecodeScholar
2025

Adaptive Time Encoding for Irregular Multivariate Time-Series Classification

NeurIPS 2025poster

Time series are often irregularly sampled with uneven time intervals. In multivariate cases, such irregularities may lead to misaligned observations across variables and varying observation counts, making it difficult to extract intrinsic patterns and degrading the classification performance of deep…

Cited by 0SourceScholar
2025

MAMS: Model-Agnostic Module Selection Framework for Video Captioning

AAAI 2025technical

Multi-modal transformers are rapidly gaining attention in video captioning tasks. Existing multi-modal video captioning methods extract a fixed number of frames, but this has critical challenges. If a limited number of frames are extracted, important frames with essential information for caption gen…

2025

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

CVPR 2025award

Today's most advanced vision-language models (VLMs) remain proprietary. The strongest open-weight models rely heavily on synthetic data from proprietary VLMs to achieve good performance, effectively distilling these closed VLMs into open ones. As a result, the community has been missing foundational…

2025

One Diffusion to Generate Them All

CVPR 2025poster

We introduce \texttt OneDiffusion - a single large-scale diffusion model designed to tackle a wide range of image synthesis and understanding tasks. It can generate images conditioned on text, depth, pose, layout, or semantic maps. It also handles super-resolution, multi-view generation, instant p…

2025

ReSpec: Relevance and Specificity Grounded Online Filtering for Learning on Video-Text Data Streams

CVPR 2025poster

The rapid growth of video-text data presents challenges in storage and computation during training. Online learning, which processes streaming data in real-time, offers a promising solution to these issues while also allowing swift adaptations in scenarios demanding real-time responsiveness. One str…

2024

Finding NeMo: Negative-mined Mosaic Augmentation for Referring Image Segmentation

ECCV 2024poster

"Referring Image Segmentation is a comprehensive task to segment an object referred by a textual query from an image. In nature, the level of difficulty in this task is affected by the existence of similar objects and the complexity of the referring expression. Recent RIS models still show a signifi…

Cited by 0SourcePDFScholar
2024

Proxyformer: Nyström-Based Linear Transformer with Trainable Proxy Tokens

AAAI 2024technical

Transformer-based models have demonstrated remarkable performance in various domains, including natural language processing, image processing and generative modeling. The most significant contributor to the successful performance of Transformer models is the self-attention mechanism, which allows fo…

Cited by 3SourcePDFScholar
2024

Towards a Complete Benchmark on Video Moment Localization

AISTATS 2024poster

In this paper, we propose and conduct a comprehensive benchmark on moment localization task, which aims to retrieve a segment that corresponds to a text query from a single untrimmed video. Our study starts from an observation that most moment localization papers report experimental results only on…

2024

Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision Language Audio and Action

CVPR 2024highlight

We present Unified-IO 2 a multimodal and multi-skill unified model capable of following novel instructions. Unified-IO 2 can use text images audio and/or videos as input and can generate text image or audio outputs which is accomplished in a unified way by tokenizing these different inputs and outpu…

2021

ACAV100M: Automatic Curation of Large-Scale Datasets for Audio-Visual Video Representation Learning

ICCV 2021poster

The natural association between visual observations and their corresponding sound provides powerful self-supervisory signals for learning video representations, which makes the ever-growing amount of online videos an attractive source of training data. However, large portions of online videos contai…

Cited by 55PDFScholar
2021

Parameter Efficient Multimodal Transformers for Video Representation Learning

ICLR 2021poster

The recent success of Transformers in the language domain has motivated adapting it to a multimodal setting, where a new visual model is trained in tandem with an already pretrained language model. However, due to the excessive memory requirements from Transformers, existing work typically fixes the…

Cited by 94SourcePDFScholar
2021

Unsupervised Representation Learning via Neural Activation Coding

ICML 2021oral

We present neural activation coding (NAC) as a novel approach for learning deep representations from unlabeled data for downstream applications. We argue that the deep encoder should maximize its nonlinear expressivity on the data for downstream predictors to take full advantage of its representatio…

2018

A Memory Network Approach for Story-Based Temporal Summarization of 360° Videos

CVPR 2018poster

We address the problem of story-based temporal summarization of long 360° videos. We propose a novel memory network model named Past-Future Memory Network (PFMN), in which we first compute the scores of 81 normal field of view (NFOV) region proposals cropped from the input 360° video, and then recov…

Cited by 85SourcePDFScholar