← Search

Reuben Tan

14 accepted papers

2026

Learning Sparse Visual Representations via Spatial-Semantic Factorization

ICML 2026poster

Self-supervised learning (SSL) faces a fundamental conflict between semantic understanding and image reconstruction. High-level semantic SSL (e.g., DINO) relies on global tokens that are forced to be location-invariant for augmentation alignment, a process that inherently discards the spatial coordi…

Cited by 0SourceScholar
2026

OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning

CVPR 2026

High-quality and carefully curated data is a cornerstone of training medical large language models, as it directly impacts both generalization and robustness to unseen clinical tasks. We investigate strategies for training and data curation to develop a robust multimodal reasoning model in the medic

Cited by 0SourceScholar
2025

GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents

NeurIPS 2025poster

One of the principal challenges in building VLM-powered GUI agents is visual grounding—localizing the appropriate screen region for action execution based on both the visual content and the textual plans. Most existing work formulates this as a text-based coordinate generation task. However, these a…

Cited by 0SourceScholar
2025

Latent Action Pretraining from Videos

ICLR 2025poster

We introduce Latent Action Pretraining for general Action models (LAPA), the first unsupervised method for pretraining Vision-Language-Action (VLA) models without ground-truth robot action labels. Existing Vision-Language-Action models require action labels typically collected by human teleoperators…

Cited by 20SourcePDFScholar
2025

Magma: A Foundation Model for Multimodal AI Agents

CVPR 2025poster

We present Magma, a foundation model that serves multimodal AI agentic tasks in both the digital and physical worlds. Magma is a significant extension of vision-language (VL) models in that it not only retains the VL understanding ability (verbal intelligence) of the latter, but is also equipped wi…

2025

MindJourney: Test-Time Scaling with World Models for Spatial Reasoning

NeurIPS 2025poster

Spatial reasoning in 3D space is central to human cognition and indispensable for embodied tasks such as navigation and manipulation. However, state-of-the-art vision–language models (VLMs) struggle frequently with tasks as simple as anticipating how a scene will look after an egocentric motion: the…

Cited by 0SourceScholar
2025

SITE: towards Spatial Intelligence Thorough Evaluation

ICCV 2025poster

Spatial intelligence (SI) represents a cognitive ability encompassing the visualization, manipulation, and reasoning about spatial relationships, underpinning disciplines from neuroscience to robotics. We introduce SITE, a benchmark dataset towards SI Thorough Evaluation in a standardized format of…

Cited by 0SourcePDFScholar
2024

Koala: Key Frame-Conditioned Long Video-LLM

CVPR 2024highlight

Long video question answering is a challenging task that involves recognizing short-term activities and reasoning about their fine-grained relationships. State-of-the-art video Large Language Models (vLLMs) hold promise as a viable solution due to their demonstrated emergent capabilities on new task…

Cited by 33SourcePDFScholar
2023

Language-Guided Audio-Visual Source Separation via Trimodal Consistency

CVPR 2023poster

We propose a self-supervised approach for learning to perform audio source separation in videos based on natural language queries, using only unlabeled video and audio pairs as training data. A key challenge in this task is learning to associate the linguistic description of a sound-emitting object…

Cited by 19SourcePDFScholar
2022

NewsStories: Illustrating Articles with Visual Summaries

ECCV 2022poster

"Recent self-supervised approaches have used large-scale image-text datasets to learn powerful representations that transfer to many tasks without finetuning. These methods often assume that there is one-to-one correspondence between its images and their (short) captions. However, many tasks require…

2021

Look at What I’m Doing: Self-Supervised Spatial Grounding of Narrations in Instructional Videos

NeurIPS 2021spotlight

We introduce the task of spatially localizing narrated interactions in videos. Key to our approach is the ability to learn to spatially localize interactions with self-supervision on a large corpus of videos with accompanying transcribed narrations. To achieve this goal, we propose a multilayer cro…

Cited by 30SourcePDFScholar
2019

Language Features Matter: Effective Language Representations for Vision-Language Tasks

ICCV 2019poster

Shouldn't language and vision features be treated equally in vision-language (VL) tasks? Many VL approaches treat the language component as an afterthought, using simple language models that are either built upon fixed word embeddings trained on text-only data or are learned from scratch. We conclud…

Cited by 39PDFScholar
2019

Learning Similarity Conditions Without Explicit Supervision

ICCV 2019poster

Many real-world tasks require models to compare images along multiple similarity conditions (e.g. similarity in color, category or shape). Existing methods often reason about these complex similarity relationships by learning condition-aware embeddings. While such embeddings aid models in learning d…

Cited by 113PDFcodeScholar