← Search

Chen Sun

81 accepted papers

2026

EIMC: Efficient Instance-Aware Multi-Modal Collaborative Perception

ICRA 2026poster

Multi-modal collaborative perception calls for great attention to enhancing the safety of autonomous driving. However, current multi-modal approaches remain a ``local fusion to communication” sequence, which fuses multi-modal data locally and needs high bandwidth to transmit an individual's feature …

2026

Goal Force: Teaching Video Models To Accomplish Physics-Conditioned Goals

CVPR 2026

Recent advancements in video generation have enabled the development of "world models" capable of simulating potential futures for robotics and planning. However, specifying precise goals for these models remains a challenge; text instructions are often too abstract to capture physical nuances, whil

Cited by 0SourcecodeScholar
2026

Implicit Action Chunking for Smooth Continuous Control

ICML 2026poster

Reinforcement learning often produces high-frequency oscillatory control signals that undermine the safety and stability required for physical deployment. Explicit action chunking addresses this by predicting fixed-horizon trajectories but increases the policy output dimension to R^hd, leading to op…

Cited by 0SourceScholar
2026

Towards Robust Sequential Decomposition for Complex Image Editing

CVPR 2026

Recent advances in visual generative models have enabled high-fidelity image editing guided by human instructions. However, these models often struggle with complex instructions involving combinatorial editing operations or inter-step dependencies. This difficulty stems from the limitations of two c

Cited by 0SourceScholar
2026

WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving

ICML 2026poster

Deploying multiple models within shared GPU clusters is a key strategy to improve resource efficiency in large language model (LLM) serving. Existing multi-LLM serving systems improve GPU utilization at the cost of degraded inference performance, particularly time-to-first-token (TTFT). We attribute…

Cited by 0SourceScholar
2025

Automated Video Object Detection of Motile Cells Under Microscopy

ICRA 2025

Video object detection (VOD) of motile cells (e.g., bacteria and sperm) under microscopy is challenging due to motion blur, sporadic out-of-focus, and pose variations. Compared with VOD in generic scenes, the lower contrast and smaller color space of microscopy imaging further introduce feature over

Cited by 0SourceScholar
2025

Dense Video Object Captioning from Disjoint Supervision

ICLR 2025spotlight

We propose a new task and model for dense video object captioning -- detecting, tracking and captioning trajectories of objects in a video. This task unifies spatial and temporal localization in video, whilst also requiring fine-grained visual understanding that is best described by natural language…

2025

Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens

ICLR 2025poster

Scaling up autoregressive models in vision has not proven as beneficial as in large language models. In this work, we investigate this scaling problem in the context of text-to-image generation, focusing on two critical factors: whether models use discrete or continuous tokens, and whether tokens ar…

Cited by 35SourcePDFScholar
2025

Force Prompting: Video Generation Models Can Learn And Generalize Physics-based Control Signals

NeurIPS 2025poster

Recent advances in video generation models have sparked interest in world models capable of simulating realistic environments. While navigation has been well-explored, physically meaningful interactions that mimic real-world forces remain largely understudied. In this work, we investigate using phy…

Cited by 0SourceScholar
2025

Fourier Head: Helping Large Language Models Learn Complex Probability Distributions

ICLR 2025poster

As the quality of large language models has improved, there has been increased interest in using them to model non-linguistic tokens. For example, the Decision Transformer recasts agentic decision making as a sequence modeling problem, using a decoder-only LLM to model the distribution over the disc…

Cited by 13SourcePDFScholar
2025

How Can Objects Help Video-Language Understanding?

ICCV 2025poster

Do we still need to represent objects explicitly in multimodal large language models (MLLMs)? To one extreme, pre-trained encoders convert images into visual tokens, with which objects and spatiotemporal relationships may be implicitly modeled. To the other extreme, image captions by themselves prov…

2025

How new data permeates LLM knowledge and how to dilute it

ICLR 2025spotlight

Large language models continually learn through the accumulation of gradient-based updates, but how individual pieces of new information affect existing knowledge, leading to both beneficial generalization and problematic hallucination, remains poorly understood. We demonstrate that when learning ne…

2025

HyperFree: A Channel-adaptive and Tuning-free Foundation Model for Hyperspectral Remote Sensing Imagery

CVPR 2025poster

Advanced interpretation of hyperspectral remote sensing images benefits many precise Earth observation tasks. Recently, visual foundation models have promoted the remote sensing interpretation but concentrating on RGB and multispectral images. Due to the varied hyperspectral channels, existing found…

Cited by 0SourcePDFScholar
2025

Image-Based Compliance Control for Robotic Steering of a Ferromagnetic Guidewire

ICRA 2025

Robotic steering of magnetic guidewires has shown great potential in accelerating endovascular interventions, enhancing the success rate of time-sensitive surgeries such as stroke treatment. Incomplete state feedback of the guidewire from 2D perspective images and unknown interactions with the surro

Cited by 0SourceScholar
2025

MotiF: Making Text Count in Image Animation with Motion Focal Loss

CVPR 2025poster

Text-Image-to-Video (TI2V) generation aims to generate a video from an image following a text description, which is also referred to as text-guided image animation. Most existing methods struggle to generate videos that align well with the text prompts, particularly when motion is specified. To over…

2025

Motion Prompting: Controlling Video Generation with Motion Trajectories

CVPR 2025poster

Motion control is crucial for generating expressive and compelling video content; however, most existing video generation models rely mainly on text prompts for control, which struggle to capture the nuances of dynamic actions and temporal compositions. To this end, we train a video generation model…

Cited by 22SourcePDFScholar
2025

Multi-Timescale Hierarchical Reinforcement Learning for Unified Behavior and Control of Autonomous Driving

RA-L 2025

Reinforcement Learning (RL) is increasingly used in autonomous driving (AD) and shows clear advantages. However, most RL-based AD methods overlook policy structure design. An RL policy that only outputs short-timescale vehicle control commands results in fluctuating driving behavior due to fluctuati

Cited by 3SourceScholar
2025

SAGE: A Unified Framework for Generalizable Object State Recognition with State-Action Graph Embedding

NeurIPS 2025oral

Recognizing the physical states of objects and their transformations within videos is crucial for structured video understanding and enabling robust real-world applications, such as robotic manipulation. However, pretrained vision-language models often struggle to capture these nuanced dynamics and…

Cited by 0SourceScholar
2024

AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?

ICLR 2024poster

Can we better anticipate an actor’s future actions (e.g. mix eggs) by knowing what commonly happens after the current action (e.g. crack eggs)? What if the actor also shares the goal (e.g. make fried rice) with us? The long-term action anticipation (LTA) task aims to predict an actor’s future behavi…

2024

Delta-AI: Local objectives for amortized inference in sparse graphical models

ICLR 2024poster

We present a new algorithm for amortized inference in sparse probabilistic graphical models (PGMs), which we call $\Delta$-amortized inference ($\Delta$-AI). Our approach is based on the observation that when the sampling of variables in a PGM is seen as a sequence of actions taken by an agent, spar…

2024

EPO: Hierarchical LLM Agents with Environment Preference Optimization

EMNLP 2024main

Long-horizon decision-making tasks present significant challenges for LLM-based agents due to the need for extensive planning over multiple steps. In this paper, we propose a hierarchical framework that decomposes complex tasks into manageable subgoals, utilizing separate LLMs for subgoal prediction…

2024

End-to-End Spatio-Temporal Action Localisation with Video Transformers

CVPR 2024poster

The most performant spatio-temporal action localisation models use external person proposals and complex external memory banks. We propose a fully end-to-end transformer based model that directly ingests an input video and outputs tubelets -- a sequence of bounding boxes and the action classes at ea…

Cited by 21SourcePDFScholar
2024

Self-Correcting Self-Consuming Loops for Generative Model Training

ICML 2024poster

As synthetic data becomes higher quality and proliferates on the internet, machine learning models are increasingly trained on a mix of human- and machine-generated data. Despite the successful stories of using synthetic data for representation learning, using synthetic data for generative model tra…

2024

Vamos: Versatile Action Models for Video Understanding

ECCV 2024poster

"What makes good representations for video understanding, such as anticipating future activities, or answering video-conditioned questions? While earlier approaches focus on end-to-end learning directly from video pixels, we propose to revisit text-based representations, such as general-purpose vide…

2023

AVIS: Autonomous Visual Information Seeking with Large Language Model Agent

NeurIPS 2023poster

In this paper, we propose an autonomous information seeking visual question answering framework, AVIS. Our method leverages a Large Language Model (LLM) to dynamically strategize the utilization of external tools and to investigate their outputs via tree search, thereby acquiring the indispensable k…

Cited by 51SourcePDFScholar
2023

Contrastive Retrospection: honing in on critical steps for rapid learning and generalization in RL

NeurIPS 2023poster

In real life, success is often contingent upon multiple critical steps that are distant in time from each other and from the final reward. These critical steps are challenging to identify with traditional reinforcement learning (RL) methods that rely on the Bellman equation for credit assignment. He…

2023

Deja Vu: Continual Model Generalization for Unseen Domains

ICLR 2023poster

In real-world applications, deep learning models often run in non-stationary environments where the target data distribution continually shifts over time. There have been numerous domain adaptation (DA) methods in both online and offline modes to improve cross-domain adaptation ability. However, the…

Cited by 27SourcePDFScholar
2023

Does Visual Pretraining Help End-to-End Reasoning?

NeurIPS 2023poster

We aim to investigate whether end-to-end learning of visual reasoning can be achieved with general-purpose neural networks, with the help of visual pretraining. A positive result would refute the common belief that explicit visual abstraction (e.g. object detection) is essential for compositional ge…

Cited by 4SourcePDFScholar
2023

Emergence of Abstract State Representations in Embodied Sequence Modeling

EMNLP 2023long main

Decision making via sequence modeling aims to mimic the success of language models, where actions taken by an embodied agent are modeled as tokens to predict. Despite their promising performance, it remains unclear if embodied sequence modeling leads to the emergence of internal representations that…

Cited by 0SourceScholar
2023

Goal-Conditioned Predictive Coding for Offline Reinforcement Learning

NeurIPS 2023poster

Recent work has demonstrated the effectiveness of formulating decision making as supervised learning on offline-collected trajectories. Powerful sequence models, such as GPT or BERT, are often employed to encode the trajectories. However, the benefits of performing sequence modeling on trajectory da…

Cited by 12SourcePDFScholar
2023

REVEAL: Retrieval-Augmented Visual-Language Pre-Training With Multi-Source Multimodal Knowledge Memory

CVPR 2023highlight

In this paper, we propose an end-to-end Retrieval-Augmented Visual Language Model (REVEAL) that learns to encode world knowledge into a large-scale memory, and to retrieve from it to answer knowledge-intensive queries. REVEAL consists of four key components: the memory, the encoder, the retriever an…

2022

Learning Audio-Video Modalities from Image Captions

ECCV 2022poster

"There has been a recent explosion of large-scale image-text datasets, as images with alt-text captions can be easily obtained online. Obtaining large-scale, high quality data for video in the form of text-video and text-audio pairs however, is more challenging. To close this gap we propose a new vi…

Cited by 109SourcePDFScholar
2022

Multiview Transformers for Video Recognition

CVPR 2022poster

Video understanding requires reasoning at multiple spatiotemporal resolutions -- from short fine-grained motions to events taking place over longer durations. Although transformer architectures have recently advanced the state-of-the-art, they have not explicitly modelled different spatiotemporal re…

Cited by 348PDFcodeScholar
2022

TL;DW? Summarizing Instructional Videos with Task Relevance & Cross-Modal Saliency

ECCV 2022poster

"YouTube users looking for instructions for a specific task may spend a long time browsing content trying to find the right video that matches their needs. Creating a visual summary (abridged version of a video) provides viewers with a quick overview and massively reduces search time. In this work,…

Cited by 0SourcePDFScholar
2022

Trajectory balance: Improved credit assignment in GFlowNets

NeurIPS 2022accept

Generative flow networks (GFlowNets) are a method for learning a stochastic policy for generating compositional objects, such as graphs or strings, from a given unnormalized density by sequences of actions, where many possible action sequences may lead to the same object. We find previously proposed…

2021

Attention Bottlenecks for Multimodal Fusion

NeurIPS 2021poster

Humans perceive the world by concurrently processing and fusing high-dimensional inputs from multiple modalities such as vision and audio. Machine perception models, in stark contrast, are typically modality-specific and optimised for unimodal benchmarks. A common approach for building multimodal m…

2021

Composable Augmentation Encoding for Video Representation Learning

ICCV 2021poster

We focus on contrastive methods for self-supervised video representation learning. A common paradigm in contrastive learning is to construct positive pairs by sampling different data views for the same instance, with different data instances as negatives. These methods implicitly assume a set of rep…

Cited by 26PDFcodeScholar
2021

Discrete-Valued Neural Communication

NeurIPS 2021poster

Deep learning has advanced from fully connected architectures to structured models organized into components, e.g., the transformer composed of positional elements, modular architectures divided into slots, and graph neural nets made up of nodes. The nature of structured models is that communication…

Cited by 58SourcePDFScholar
2021

HDMapGen: A Hierarchical Graph Generative Model of High Definition Maps

CVPR 2021poster

High Definition (HD) maps are maps with precise definitions of road lanes with rich semantics of the traffic rules. They are critical for several key stages in an autonomous driving system, including motion forecasting and planning. However, there are only a small amount of real-world road topologie…

Cited by 71PDFScholar
2021

ViViT: A Video Vision Transformer

ICCV 2021poster

We present pure-transformer based models for video classification, drawing upon the recent success of such models in image classification. Our model extracts spatio-temporal tokens from the input video, which are then encoded by a series of transformer layers. In order to handle the long sequences o…

Cited by 2888PDFcodeScholar
2020

Imitative Reinforcement Learning Fusing Vision and Pure Pursuit for Self-driving

ICRA 2020poster

Autonomous urban driving navigation is still an open problem and has ample room for improvement in unknown complex environments and terrible weather conditions. In this paper, we propose a two-stage framework, called IPP-RL, to handle these problems. IPP means an Imitation learning method fusing vis…

Cited by 14SourceScholar
2020

Speech2Action: Cross-Modal Supervision for Action Recognition

CVPR 2020poster

Is it possible to guess human action from dialogue alone? In this work we investigate the link between spoken words and actions in movies. We note that movie screenplays describe actions, as well as contain the speech of characters and hence can be used to learn this correlation with no additional s…

Cited by 78PDFScholar
2020

Uncertainty-Aware Weakly Supervised Action Detection from Untrimmed Videos

ECCV 2020poster

Despite the recent advances in video classification, progress in spatio-temporal action recognition has lagged behind. A major contributing factor has been the prohibitive cost of annotating videos frame-by-frame. In this paper, we present a spatio-temporal action recognition model that is trained w…

2020

VectorNet: Encoding HD Maps and Agent Dynamics From Vectorized Representation

CVPR 2020poster

Behavior prediction in dynamic, multi-agent systems is an important problem in the context of self-driving cars, due to the complex representations and interactions of road components, including moving agents (e.g. pedestrians and vehicles) and road context information (e.g. lanes, traffic lights).…

Cited by 1022PDFScholar
2020

What Makes for Good Views for Contrastive Learning?

NeurIPS 2020poster

Contrastive learning between multiple views of the data has recently achieved state of the art performance in the field of self-supervised representation learning. Despite its success, the influence of different view choices has been less studied. In this paper, we use theoretical and empirical anal…

Cited by 1641SourcePDFScholar
2019

Composing Text and Image for Image Retrieval - an Empirical Odyssey

CVPR 2019oral

In this paper, we study the task of image retrieval, where the input query is specified in the form of an image plus some text that describes desired modifications to the input image. For example, we may present an image of the Eiffel tower, and ask the system to find images which are visually simil…

Cited by 442PDFScholar
2019

Hyperspectral Image Reconstruction Using a Deep Spatial-Spectral Prior

CVPR 2019poster

Regularization is a fundamental technique to solve an ill-posed optimization problem robustly and is essential to reconstruct compressive hyperspectral images. Various hand-crafted priors have been employed as a regularizer but are often insufficient to handle the wide variety of spectra of natural…

Cited by 218PDFScholar
2019

Relational Action Forecasting

CVPR 2019oral

This paper focuses on multi-person action forecasting in videos. More precisely, given a history of H previous frames, the goal is to detect actors and to predict their future actions for the next T frames. Our approach jointly models temporal and spatial interactions among different actors by const…

Cited by 100PDFScholar
2019

Stochastic Prediction of Multi-Agent Interactions from Partial Observations

ICLR 2019poster

We present a method which learns to integrate temporal information, from a learned dynamics model, with ambiguous visual information, from a learned vision model, in the context of interacting agents. Our method is based on a graph-structured variational recurrent neural network, which is trained en…

Cited by 118SourcePDFScholar
2019

Unsupervised Discovery of Parts, Structure, and Dynamics

ICLR 2019poster

Humans easily recognize object parts and their hierarchical structure by watching how they move; they can then predict how each part moves in the future. In this paper, we propose a novel formulation that simultaneously learns a hierarchical, disentangled object representation and a dynamics model f…

Cited by 84SourcePDFScholar
2019

Unsupervised learning of object structure and dynamics from videos

NeurIPS 2019poster

Extracting and predicting object structure and dynamics from videos without supervision is a major challenge in machine learning. To address this challenge, we adopt a keypoint-based image representation and learn a stochastic dynamics model of the keypoints. Future frames are reconstructed from the…

2019

VideoBERT: A Joint Model for Video and Language Representation Learning

ICCV 2019poster

Self-supervised learning has become increasingly important to leverage the abundance of unlabeled data available on platforms like YouTube. Whereas most existing approaches learn low-level representations, we propose a joint visual-linguistic model to learn high-level features without any explicit s…

Cited by 1568PDFScholar
2018

AVA: A Video Dataset of Spatio-Temporally Localized Atomic Visual Actions

CVPR 2018poster

This paper introduces a video dataset of spatio-temporally localized Atomic Visual Actions (AVA). The AVA dataset densely annotates 80 atomic visual actions in 437 15-minute video clips, where actions are localized in space and time, resulting in 1.59M action labels with multiple labels per person o…

Cited by 1319SourcePDFScholar
2018

Actor-centric Relation Network

ECCV 2018poster

Current state-of-the-art approaches for spatio-temporal action localization rely on detections at the frame level and model temporal context with 3D ConvNets. Here, we go one step further and model spatio-temporal relations to capture the interactions between human actors, relevant objects and scene…

Cited by 280SourcePDFScholar
2018

Large Scale Fine-Grained Categorization and Domain-Specific Transfer Learning

CVPR 2018poster

Transferring the knowledge learned from large scale datasets (e.g., ImageNet) via fine-tuning offers an effective solution for domain-specific fine-grained visual categorization (FGVC) tasks (e.g., recognizing bird species or car make & model). In such scenarios, data annotation often calls for spec…

Cited by 656SourcePDFScholar
2018

Rethinking Spatiotemporal Feature Learning: Speed-Accuracy Trade-offs in Video Classification

ECCV 2018poster

Despite the steady progress in video analysis led by the adoption of convolutional neural networks (CNNs), the relative improvement has been less drastic as that in 2D static image classification. Three main challenges exist including spatial (image) feature representation, temporal information repr…

Cited by 1765SourcePDFScholar
2018

The INaturalist Species Classification and Detection Dataset

CVPR 2018poster

Existing image classification datasets used in computer vision tend to have a uniform distribution of images across object categories. In contrast, the natural world is heavily imbalanced, as some species are more abundant and easier to photograph than others. To encourage further progress in challe…

2017

Revisiting Unreasonable Effectiveness of Data in Deep Learning Era

ICCV 2017spotlight

The success of deep learning in vision can be attributed to: (a) models with high capacity; (b) increased computational power; and (c) availability of large-scale labeled data. Since 2012, there have been significant advances in representation capabilities of the models and computational capabilitie…

Cited by 3433PDFScholar
2017

Speed/Accuracy Trade-Offs for Modern Convolutional Object Detectors

CVPR 2017spotlight

The goal of this paper is to serve as a guide for selecting a detection architecture that achieves the right speed/memory/accuracy balance for a given application and platform. To this end, we investigate various ways to trade accuracy for speed and memory usage in modern convolutional object detect…

Cited by 3693PDFcodeScholar
2017

TURN TAP: Temporal Unit Regression Network for Temporal Action Proposals

ICCV 2017poster

We address the problem of Temporal Action Proposal (TAP) generation. This is an important problem, as fast extraction of semantically important (e.g. human actions) segments from untrimmed videos is an important step for large-scale video analysis. To tackle this problem, we propose a novel Temporal…

Cited by 486PDFcodeScholar
2017

VQS: Linking Segmentations to Questions and Answers for Supervised Attention in VQA and Question-Focused Semantic Segmentation

ICCV 2017poster

Rich and dense human labeled datasets are the main enabling factor, among others, for the recent exciting work on vision-language understanding. Many seemingly distinct annotations (e.g., semantic segmentation and visual questions answering (VQA)) are inherently connected in that they reveal differe…

Cited by 145PDFcodeScholar
2016

ProNet: Learning to Propose Object-Specific Boxes for Cascaded Neural Networks

CVPR 2016poster

This paper aims to classify and locate objects accurately and efficiently, without using bounding box annotations. It is challenging as objects in the wild could appear at arbitrary locations and in different scales. In this paper, we propose a novel classification architecture ProNet based on convo…

Cited by 80PDFScholar