← Search

David Fan

17 accepted papers

2026

Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training

ICLR 2026oral

Large Language Models (LLMs), despite being trained on text alone, surprisingly develop rich visual priors. These priors allow latent visual capabilities to be unlocked for vision tasks with a relatively small amount of multimodal data, and to perform symbolic visual generation tasks without ever ha…

Cited by 0SourceScholar
2026

Towards Unified Multimodal Pretraining

ICML 2026spotlight

Unified multimodal models aim to input and output both vision and language data within a single system. In this work, we explore the design space of Unified Multimodal Pretraining through a controlled, from-scratch study. We find that leveraging a single high-dimensional semantic encoder (e.g. SigLI…

Cited by 0SourceScholar
2026

TravSUITE: Traversability via Self-Supervised, Uncertainty-Aware IRL and Terrain Estimation

RSS 2026poster

Traversability analysis in off-road settings remains a fundamental challenge for mobile robots. Key difficulties include constructing an accurate, expressive local map from multi-modal sensor data and using the map to design traversability rules that yield desirable navigation behavior. Importantly,…

Cited by 0SourceScholar
2025

Enter the Mind Palace: Reasoning and Planning for Long-term Active Embodied Question Answering

CoRL 2025poster

As robots become increasingly capable of operating over extended periods—spanning days, weeks, and even months—they are expected to accumulate knowledge of their environments and leverage this experience to assist humans more effectively. This paper studies the problem of Long-term Active Embodied Q…

Cited by 0SourceScholar
2025

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

ICCV 2025poster

In this work, we propose Visual-Predictive Instruction Tuning (VPiT) - a simple and effective extension to visual instruction tuning that enables a pretrained LLM to quickly morph into an unified autoregressive model capable of generating both text and visual tokens. VPiT teaches an LLM to predict d…

Cited by 0SourcePDFScholar
2025

Scaling Language-Free Visual Representation Learning

ICCV 2025poster

Visual Self-Supervised Learning (SSL) currently underperforms Contrastive Language-Image Pretraining (CLIP) in multimodal settings such as Visual Question Answering (VQA). This multimodal gap is often attributed to the semantics introduced by language supervision, even though visual SSL and CLIP mod…

2024

SEEK: Semantic Reasoning for Object Goal Navigation in Real World Inspection Tasks

RSS 2024poster

This paper addresses the problem of object-goal navigation in autonomous inspections in real-world environments. Object-goal navigation is crucial to enable effective inspections in various settings, often requiring the robot to identify the target object within a large search space. Current object…

Cited by 8SourcePDFScholar
2024

Velociraptor: Leveraging Visual Foundation Models for Label-Free, Risk-Aware Off-Road Navigation

CoRL 2024poster

Traversability analysis in off-road regimes is a challenging task that requires understanding of multi-modal inputs such as camera and LiDAR. These measurements are often sparse, noisy, and difficult to interpret, particularly in the off-road setting. Existing systems are very engineering-intensive,…

Cited by 2SourceScholar
2023

A Multi-step Dynamics Modeling Framework For Autonomous Driving In Multiple Environments

ICRA 2023poster

Modeling dynamics is often the first step to making a vehicle autonomous. While on-road autonomous vehicles have been extensively studied, off-road vehicles pose many challenging modeling problems. An off-road vehicle encounters highly complex and difficult-to-model terrain/vehicle interactions, as…

Cited by 15SourceScholar
2023

MEGA: Multimodal Alignment Aggregation and Distillation For Cinematic Video Segmentation

ICCV 2023poster

Previous research has studied the task of segmenting cinematic videos into scenes and into narrative acts. However, these studies have overlooked the essential task of multimodal alignment and fusion for effectively and efficiently processing long-form videos (>60min). In this paper, we introduce Mu…

Cited by 4PDFcodeScholar
2023

Motion-Guided Masking for Spatiotemporal Representation Learning

ICCV 2023poster

Several recent works have directly extended the image masked autoencoder (MAE) with random masking into video domain, achieving promising results. However, unlike images, both spatial and temporal information are important for video understanding. This suggests that the random masking strategy that…

Cited by 16PDFScholar
2022

Hybrid Imitative Planning with Geometric and Predictive Costs in Off-road Environments

ICRA 2022poster

Geometric methods for solving open-world off-road navigation tasks, by learning occupancy and metric maps, provide good generalization but can be brittle in outdoor environments that violate their assumptions (e.g., tall grass). Learning-based methods can directly learn collision-free behavior from…

Cited by 17SourceScholar
2022

PrePARE: Predictive Proprioception for Agile Failure Event Detection in Robotic Exploration of Extreme Terrains

IROS 2022poster

Legged robots can traverse a wide variety of terrains, some of which may be challenging for wheeled robots, such as stairs or highly uneven surfaces. However, quadruped robots face stability challenges on slippery surfaces. This can be resolved by adjusting the robot's locomotion by switching to mor…

Cited by 9SourceScholar
2021

Shot Contrastive Self-Supervised Learning for Scene Boundary Detection

CVPR 2021poster

Scenes play a crucial role in breaking the storyline of movies and TV episodes into semantically cohesive parts. However, given their complex temporal structure, finding scene boundaries can be a challenging task requiring large amounts of labeled training data. To address this challenge, we present…

Cited by 87PDFScholar
2020

OASIS: A Large-Scale Dataset for Single Image 3D in the Wild

CVPR 2020poster

Single-view 3D is the task of recovering 3D properties such as depth and surface normals from a single image. We hypothesize that a major obstacle to single-image 3D is data. We address this issue by presenting Open Annotations of Single Image Surfaces (OASIS), a dataset for single-image 3D in the w…

Cited by 83PDFScholar