← Search

Jiachen Li

73 accepted papers

2026

AutoFocus-IL: VLM-Based Saliency Maps for Data-Efficient Visual Imitation Learning without Extra Human Annotations

ICRA 2026poster

We present AutoFocus-IL, a simple yet effective method to improve data efficiency and generalization in visual imitation learning by guiding policies to attend to task-relevant features rather than distractors and spurious correlations. Saliency regularization has emerged as a promising way to achie…

2026

CommCP: Efficient Multi-Agent Coordination Via LLM-Based Communication with Conformal Prediction

ICRA 2026poster

To complete assignments provided by humans in natural language, robots must interpret commands, generate and answer relevant questions for scene understanding, and manipulate target objects. Real-world deployments often require multiple heterogeneous robots with different manipulation capabilities t…

2026

Drive My Way: Preference Alignment of Vision-Language-Action Model for Personalized Driving

CVPR 2026

Human driving behavior is inherently personal, which is shaped by long-term habits and influenced by short-term intentions. Individuals differ in how they accelerate, brake, merge, yield, and overtake across diverse situations. However, existing end-to-end autonomous driving systems either optimize

Cited by 0SourcecodeScholar
2026

GUIDES: Guidance Using Instructor-Distilled Embeddings for Pre-Trained Robot Policy Enhancement

ICRA 2026poster

Pre-trained robot policies serve as the foundation of many validated robotic systems, which encapsulate extensive embodied knowledge. However, they often lack the semantic awareness characteristic of foundation models, and replacing them entirely is impractical in many situations due to high costs a…

2026

Modality-Aware Bias Mitigation and Invariance Learning for Unsupervised Visible-Infrared Person Re-Identification

AAAI 2026technical

Unsupervised visible-infrared person re-identification (USVI-ReID) aims to match individuals across visible and infrared cameras without relying on any annotation. Given the significant gap across visible and infrared modality, estimating reliable cross-modality association becomes a major challenge

Cited by 0SourcePDFScholar
2026

Reducing Oracle Feedback with Vision Language Embeddings for Preference Based RL

ICRA 2026poster

Preference-based reinforcement learning (RL) offers a promising approach for aligning policies with human intent but is often constrained by the high cost of human feedback. In this work, we introduce ROVED, a framework that integrates Vision-Language Models (VLMs) with selective human feedback to s…

2026

TrajEvo: Trajectory Prediction Heuristics Design via LLM-driven Evolution

AAAI 2026technical

Trajectory prediction is a crucial task in modeling human behavior, especially in safety-critical fields such as social robotics and autonomous vehicle navigation. Traditional heuristics based on handcrafted rules often lack accuracy, while recently proposed deep learning approaches suffer from comp

Cited by 0SourcePDFScholar
2026

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

CVPR 2026

The rapid advancement of Large Multimodal Models (LMMs) for 2D images and videos has sparked interest in extending these models to 3D scenes, with the goal of human-like visual-spatial intelligence. However, achieving deep spatial understanding comparable to human capabilities remains challenging fo

Cited by 0SourcecodeScholar
2025

A Visual Servo System for Robotic on-Orbit Servicing Based on 3D Perception of Non-Cooperative Satellite

ICRA 2025

The 3D perception of satellites, including both their shape and pose, is a key foundation for robotic on-orbit servicing. However, the demanding space environment-such as intense and dim illumination-presents significant challenges. Previous non-cooperative methods focus on specific geometric featur

Cited by 0SourceScholar
2025

Adaptive Prediction Ensemble: Improving Out-of-Distribution Generalization of Motion Forecasting

RA-L 2025

Deep learning-based trajectory prediction models for autonomous driving often struggle with generalization to out-of-distribution (OOD) scenarios, sometimes performing worse than simple rule-based models. To address this limitation, we propose a novel framework, Adaptive Prediction Ensemble (APE), w

Cited by 10SourceScholar
2025

CMP: Cooperative Motion Prediction With Multi-Agent Communication

RA-L 2025

The confluence of the advancement of Autonomous Vehicles (AVs) and the maturity of Vehicle-to-Everything (V2X) communication has enabled the capability of cooperative connected and automated vehicles (CAVs). Building on top of cooperative perception, this letter explores the feasibility and effectiv

Cited by 37SourceScholar
2025

CoMamba: Real-time Cooperative Perception Unlocked with State-Space Models

IROS 2025

Cooperative perception systems play a vital role in enhancing the safety and efficiency of vehicular autonomy. Although recent studies have highlighted the efficacy of vehicle-to-everything (V2X) communication techniques in autonomous driving, a significant challenge persists: how to efficiently int

Cited by 7SourceScholar
2025

Conformal Prediction and MLLM aided Uncertainty Quantification in Scene Graph Generation

CVPR 2025poster

Scene Graph Generation (SGG) aims to represent visual scenes by identifying objects and their pairwise relationships, providing a structured understanding of image content. However, inherent challenges like long-tailed class distributions and prediction variability necessitate uncertainty quantifica…

Cited by 7SourcePDFScholar
2025

HEAL: An Empirical Study on Hallucinations in Embodied Agents Driven by Large Language Models

EMNLP 2025

Large language models (LLMs) are increasingly being adopted as the cognitive core of embodied agents. However, inherited hallucinations, which stem from failures to ground user instructions in the observed physical environment, can lead to navigation errors, such as searching for a refrigerator that

Cited by 0SourcePDFScholar
2025

HeightAware-BEV: Height-Aware Feature Mapping for Efficient Bird's-Eye-View Perception

IROS 2025

Bird’s-Eye View (BEV) perception has gained significant attention in autonomous driving and robotics due to its advantages in simplifying modality alignment and feature fusion. Addressing the challenge of jointly optimizing performance and efficiency in 2D-3D view transformation, we identify that, c

Cited by 0SourcecodeScholar
2025

Human Implicit Preference-Based Policy Fine-tuning for Multi-Agent Reinforcement Learning in USV Swarm

IROS 2025

Multi-Agent Reinforcement Learning (MARL) has shown promise in solving complex problems involving cooperation and competition among agents, such as an Unmanned Surface Vehicle (USV) swarm used in search and rescue, surveillance, and vessel protection. However, aligning system behavior with user pref

Cited by 5SourceScholar
2025

Importance Sampling-Guided Meta-Training for Intelligent Agents in Highly Interactive Environments

RA-L 2025

Training intelligent agents to navigate highly interactive environments presents significant challenges. While guided meta reinforcement learning (RL) approach that first trains a guiding policy to train the ego agent has proven effective in improving generalizability across scenarios with various l

Cited by 4SourceScholar
2025

LaMMA-P: Generalizable Multi-Agent Long-Horizon Task Allocation and Planning with LM-Driven PDDL Planner

ICRA 2025

Language models (LMs) possess a strong capability to comprehend natural language, making them effective in translating human instructions into detailed plans for simple robot tasks. Nevertheless, it remains a significant challenge to handle long-horizon tasks, especially in subtask identification an

Cited by 41SourcecodeScholar
2025

MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

ICLR 2025poster

Multimodal Language Language Models (MLLMs) demonstrate the emerging abilities of "world models"---interpreting and reasoning about complex real-world dynamics. To assess these abilities, we posit videos are the ideal medium, as they encapsulate rich representations of real-world dynamics and causal…

2025

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

ICCV 2025poster

In this work, we propose Visual-Predictive Instruction Tuning (VPiT) - a simple and effective extension to visual instruction tuning that enables a pretrained LLM to quickly morph into an unified autoregressive model capable of generating both text and visual tokens. VPiT teaches an LLM to predict d…

Cited by 0SourcePDFScholar
2025

PRESS: Defending Privacy in Retrieval-Augmented Generation via Embedding Space Shifting

ICASSP 2025accepted

Retrieval-augmented generation (RAG) expands the capabilities of large language models (LLMs) in various applications by integrating relevant information retrieved from external data sources. However, the RAG systems are exposed to substantial privacy risks during the information retrieval process,…

Cited by 0SourceScholar
2025

Prior-free 3D Object Tracking

CVPR 2025highlight

In this paper, we introduce a novel, truly prior-free 3D object tracking method that operates without given any model or training priors. Unlike existing methods that typically require pre-defined 3D models or specific training datasets as priors, which limit their applicability, our method is free…

2025

RDD: Retrieval-Based Demonstration Decomposer for Planner Alignment in Long-Horizon Tasks

NeurIPS 2025poster

To tackle long-horizon tasks, recent hierarchical vision-language-action (VLAs) frameworks employ vision-language model (VLM)-based planners to decompose complex manipulation tasks into simpler sub-tasks that low-level visuomotor policies can easily handle. Typically, the VLM planner is finetuned to…

Cited by 0SourcecodeScholar
2025

STAD: Joint Spatial-Temporal Dimension and Channel Correlation for Time Series Anomaly Detection

ICASSP 2025accepted

Accurately identifying real anomalies and pseudo-anomalies in complex multi-dimensional time series data has been a difficult problem in time series anomaly detection. To solve this problem, this paper proposes a new framework, STAD, that joint temporal and spatial dimensions. This framework guides…

Cited by 0SourceScholar
2025

STAMP: Scalable Task- And Model-agnostic Collaborative Perception

ICLR 2025poster

Perception is a crucial component of autonomous driving systems. However, single-agent setups often face limitations due to sensor constraints, especially under challenging conditions like severe occlusion, adverse weather, and long-range object detection. Multi-agent collaborative perception (CP) o…

2025

Self-supervised Multi-future Occupancy Forecasting for Autonomous Driving

RSS 2025poster

Environment prediction frameworks are critical for the safe navigation of autonomous vehicles (AVs) in dynamic settings. LiDAR-generated occupancy grid maps (L-OGMs) offer a robust bird’s-eye view scene representation, enabling self-supervised joint scene predictions while exhibiting resilience to p…

Cited by 3PDFScholar
2025

T2V-Turbo-v2: Enhancing Video Model Post-Training through Data, Reward, and Conditional Guidance Design

ICLR 2025poster

In this paper, we focus on enhancing a diffusion-based text-to-video (T2V) model during the post-training phase by distilling a highly capable consistency model from a pretrained T2V model. Our proposed method, T2V-Turbo-v2, introduces a significant advancement by integrating various supervision sig…

Cited by 17SourcePDFScholar
2025

TC-Bench: Benchmarking Temporal Compositionality in Conditional Video Generation

ACL 2025finding

Video generation has many unique challenges beyond those of image generation. The temporal dimension introduces extensive possible variations across frames, over which consistency and continuity may be violated. In this work, we evaluate the emergence of new concepts and relation transitions as time…

Cited by 0SourcePDFScholar
2025

Towards Generalizable Safety in Crowd Navigation via Conformal Uncertainty Handling

CoRL 2025poster

Mobile robots navigating in crowds trained using reinforcement learning are known to suffer performance degradation when faced with out-of-distribution scenarios. We propose that by properly accounting for the uncertainties of pedestrians, a robot can learn safe navigation policies that are robust t…

Cited by 0SourceScholar
2025

Uncertainty-Aware Diffusion-Guided Refinement of 3D Scenes

ICCV 2025poster

Reconstructing 3D scenes from a single image is a fundamentally ill-posed task due to the severely under-constrained nature of the problem. Consequently, when the scene is rendered from novel camera views, particularly in unseen regions far away from the input camera, existing single image to 3D rec…

Cited by 0SourcePDFScholar
2025

UniOcc: A Unified Benchmark for Occupancy Forecasting and Prediction in Autonomous Driving

ICCV 2025poster

We introduce UniOcc, a comprehensive, unified benchmark and toolkit for occupancy forecasting (i.e., predicting future occupancies based on historical information) and occupancy prediction (i.e., predicting current-frame occupancy from camera images. UniOcc unifies the data from multiple real-world…

2025

VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search

EMNLP 2025

Vision-Language Models have made significant progress on many perception-focused tasks. However, their progress on reasoning-focused tasks remains limited due to the lack of high-quality and diverse training data. In this work, we aim to address the scarcity of reasoning-focused multimodal datasets.

Cited by 0SourcePDFScholar
2024

Adversarial Attacks on Parts of Speech: An Empirical Study in Text-to-Image Generation

EMNLP 2024finding

Recent studies show that text-to-image (T2I) models are vulnerable to adversarial attacks, especially with noun perturbations in text prompts. In this study, we investigate the impact of adversarial attacks on different POS tags within text prompts on the images generated by T2I models. We create a…

2024

BPO: Staying Close to the Behavior LLM Creates Better Online LLM Alignment

EMNLP 2024main

Direct alignment from preferences (DAP) has emerged as a promising paradigm for aligning large language models (LLMs) to human desiderata from pre-collected, offline preference datasets. While recent studies indicate that existing offline DAP methods can directly benefit from online training samples…

2024

CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts

NeurIPS 2024poster

Recent advancements in Multimodal Large Language Models (LLMs) have focused primarily on scaling by increasing text-image pair data and enhancing LLMs to improve performance on multimodal tasks. However, these scaling approaches are computationally expensive and overlook the significance of efficien…

2024

Disentangled Neural Relational Inference for Interpretable Motion Prediction

RA-L 2024

Effective interaction modeling and behavior prediction of dynamic agents play a significant role in interactive motion planning for autonomous robots. Although existing methods have improved prediction accuracy, few research efforts have been devoted to enhancing prediction model interpretability an

Cited by 9SourceScholar
2024

Language Model Adaption for Reinforcement Learning with Natural Language Action Space

ACL 2024long

Reinforcement learning with natural language action space often suffers from the curse of dimensionality due to the combinatorial nature of the natural language. Previous research leverages pretrained language models to capture action semantics and reduce the size of the action space. However, since…

2024

MATRIX: Multi-Agent Trajectory Generation with Diverse Contexts

ICRA 2024poster

Data-driven methods have great advantages in modeling complicated human behavioral dynamics and dealing with many human-robot interaction applications. However, collecting massive and annotated real-world human datasets has been a laborious task, especially for highly interactive scenarios. On the o…

Cited by 6SourceScholar
2024

Mastering Robot Manipulation with Multimodal Prompts through Pretraining and Multi-task Fine-tuning

ICML 2024poster

Prompt-based learning has been demonstrated as a compelling paradigm contributing to large language models' tremendous success (LLMs). Inspired by their success in language tasks, existing research has leveraged LLMs in embodied instruction following and task planning. In this work, we tackle the pr…

Cited by 11SourcePDFScholar
2024

More Samples or More Prompts? Exploring Effective Few-Shot In-Context Learning for LLMs with In-Context Sampling

NAACL 2024findings

While most existing works on LLM prompting techniques focus only on how to select a better set of data samples inside one single prompt input (In-Context Learning or ICL), why can not we design and leverage multiple prompts together to further improve the LLM’s performance? In this work, we propose…

Cited by 11SourcePDFScholar
2024

Scene Informer: Anchor-based Occlusion Inference and Trajectory Prediction in Partially Observable Environments

ICRA 2024poster

Navigating complex and dynamic environments requires autonomous vehicles (AVs) to reason about both visible and occluded regions. This involves predicting the future motion of observed agents, inferring occluded ones, and modeling their interactions based on vectorized scene representations of the p…

Cited by 10SourcecodeScholar
2024

T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback

NeurIPS 2024poster

Diffusion-based text-to-video (T2V) models have achieved significant success but continue to be hampered by the slow sampling speed of their iterative sampling processes. To address the challenge, consistency models have been proposed to facilitate fast inference, albeit at the cost of sample qualit…

2023

Causal Balancing for Domain Generalization

ICLR 2023poster

While machine learning models rapidly advance the state-of-the-art on various real-world tasks, out-of-domain (OOD) generalization remains a challenging problem given the vulnerability of these models to spurious correlations. We propose a balanced mini-batch sampling strategy to transform a biased…

2023

Offline Reinforcement Learning with Closed-Form Policy Improvement Operators

ICML 2023poster

Behavior constrained policy optimization has been demonstrated to be a successful paradigm for tackling Offline Reinforcement Learning. By exploiting historical transitions, a policy is trained to maximize a learned value function while constrained by the behavior policy to avoid a significant distr…

2023

OneFormer: One Transformer To Rule Universal Image Segmentation

CVPR 2023poster

Universal Image Segmentation is not a new concept.Past attempts to unify image segmentation include scene parsing, panoptic segmentation, and, more recently, new panoptic architectures. However, such panoptic architectures do not truly unify image segmentation because they need to be trained individ…

2023

Online Hand-Eye Calibration with Decoupling by 3D Textureless Object Tracking

ICRA 2023poster

Hand-eye calibration estimates the pose of a camera relative to a robot, which is a fundamental problem for visually guided robots, especially for dynamic object grasping. Most methods use 2D fiducial markers with distinctive visual features and require pre-calibration for accurate calibration, whic…

Cited by 2SourceScholar
2023

Pedestrian Crossing Action Recognition and Trajectory Prediction with 3D Human Keypoints

ICRA 2023poster

Accurate understanding and prediction of human behaviors are critical prerequisites for autonomous vehicles, especially in highly dynamic and interactive scenarios such as intersections in dense urban areas. In this work, we aim at identifying crossing pedestrians and predicting their future traject…

Cited by 19SourceScholar
2022

Alleviating Semantics Distortion in Unsupervised Low-Level Image-to-Image Translation via Structure Consistency Constraint

CVPR 2022poster

Unsupervised image-to-image (I2I) translation aims to learn a domain mapping function that can preserve the semantics of the input images without paired data. However, because the underlying semantics distributions in the source and target domains are often mismatched, current distribution matching-…

Cited by 33PDFcodeScholar
2022

BCOT: A Markerless High-Precision 3D Object Tracking Benchmark

CVPR 2022poster

Template-based 3D object tracking still lacks a high-precision benchmark of real scenes due to the difficulty of annotating the accurate 3D poses of real moving video objects without using markers. In this paper, we present a multi-view approach to estimate the accurate 3D poses of real moving objec…

Cited by 17PDFcodeScholar
2022

Dynamics-Aware Spatiotemporal Occupancy Prediction in Urban Environments

IROS 2022poster

Detection and segmentation of moving obstacles, along with prediction of the future occupancy states of the local environment, are essential for autonomous vehicles to proactively make safe and informed decisions. In this paper, we propose a framework that integrates the two capabilities together us…

Cited by 18SourceScholar
2022

Grouptron: Dynamic Multi-Scale Graph Convolutional Networks for Group-Aware Dense Crowd Trajectory Forecasting

ICRA 2022poster

Accurate, long-term forecasting of pedestrian trajectories in highly dynamic and interactive scenes is a longstanding challenge. Recent advances in using data-driven approaches have achieved significant improvements in terms of prediction accuracy. However, the lack of group-aware analysis has limit…

Cited by 33SourceScholar
2022

Important Object Identification with Semi-Supervised Learning for Autonomous Driving

ICRA 2022poster

Accurate identification of important objects in the scene is a prerequisite for safe and high-quality decision making and motion planning of intelligent agents (e.g., autonomous vehicles) that navigate in complex and dynamic environments. Most existing approaches attempt to employ attention mechanis…

Cited by 19SourceScholar
2022

Interaction Modeling with Multiplex Attention

NeurIPS 2022accept

Modeling multi-agent systems requires understanding how agents interact. Such systems are often difficult to model because they can involve a variety of types of interactions that layer together to drive rich social behavioral dynamics. Here we introduce a method for accurately modeling multi-agent…

Cited by 24SourcePDFScholar
2022

Learning Physical Dynamics with Subequivariant Graph Neural Networks

NeurIPS 2022accept

Graph Neural Networks (GNNs) have become a prevailing tool for learning physical dynamics. However, they still encounter several challenges: 1) Physical laws abide by symmetry, which is a vital inductive bias accounting for model generalization and should be incorporated into the model design. Exis…

Cited by 46SourcePDFScholar
2022

Multi-Objective Diverse Human Motion Prediction With Knowledge Distillation

CVPR 2022oral

Obtaining accurate and diverse human motion prediction is essential to many industrial applications, especially robotics and autonomous driving. Recent research has explored several techniques to enhance diversity and maintain the accuracy of human motion prediction at the same time. However, most o…

Cited by 48PDFScholar
2022

Point-to-Box Network for Accurate Object Detection via Single Point Supervision

ECCV 2022poster

"Object detection using single point supervision has received increasing attention over the years. However, the performance gap between point supervised object detection (PSOD) and bounding box supervised detection remains large. In this paper, we attribute such a large performance gap to the failur…

2021

Continual Multi-Agent Interaction Behavior Prediction With Conditional Generative Memory

RA-L 2021

Multi-agent trajectory prediction plays a crucial role in robotics and autonomous driving. The current mainstream research focuses on how to achieve accurate prediction on one large dataset. However, whether the multi-agent trajectory prediction model can be trained with a sequence of datasets, i.e.

Cited by 39SourceScholar
2021

LOKI: Long Term and Key Intentions for Trajectory Prediction

ICCV 2021poster

Recent advances in trajectory prediction have shown that explicit reasoning about agents' intent is important to accurately forecast their motion. However, the current research activities are not directly applicable to intelligent and safety critical systems. This is mainly because very few public d…

Cited by 113PDFScholar
2021

Orientation-Aware Planning for Parallel Task Execution of Omni-Directional Mobile Robot

IROS 2021poster

Omni-directional mobile robot (OMR) systems have been very popular in academia and industry for their superb maneuverability and flexibility. Yet their potential has not been fully exploited, where the extra degree of freedom in OMR can potentially enable the robot to carry out extra tasks. For inst…

Cited by 2SourceScholar
2021

RAIN: Reinforced Hybrid Attention Inference Network for Motion Forecasting

ICCV 2021poster

Motion forecasting plays a significant role in various domains (e.g., autonomous driving, human-robot interaction), which aims to predict future motion sequences given a set of historical observations. However, the observed elements may be of different levels of importance. Some information may be i…

Cited by 49PDFScholar
2021

Reinforcement Learning for Autonomous Driving with Latent State Inference and Spatial-Temporal Relationships

ICRA 2021poster

Deep reinforcement learning (DRL) provides a promising way for learning navigation in complex autonomous driving scenarios. However, identifying the subtle cues that can indicate drastically different outcomes remains an open problem with designing autonomous systems that operate in human environmen…

Cited by 80SourceScholar
2020

A Speech-to-Knowledge-Graph Construction System

IJCAI 2020poster

This paper presents a HAO-Graph system that generates and visualizes knowledge graphs from a speech in real-time. When a user speaks to the system, HAO-Graph transforms the voice into knowledge graphs with key phrases from the original speech as nodes and edges. Different from language-to-language s…

Cited by 0SourcePDFScholar
2020

EvolveGraph: Multi-Agent Trajectory Prediction with Dynamic Relational Reasoning

NeurIPS 2020poster

Multi-agent interacting systems are prevalent in the world, from purely physical systems to complicated social dynamic systems. In many applications, effective understanding of the situation and accurate trajectory prediction of interactive agents play a significant role in downstream tasks, such as…

Cited by 260SourcePDFScholar
2020

Multi-task Batch Reinforcement Learning with Metric Learning

NeurIPS 2020poster

We tackle the Multi-task Batch Reinforcement Learning problem. Given multiple datasets collected from different tasks, we train a multi-task policy to perform well in unseen tasks sampled from the same distribution. The task identities of the unseen tasks are not provided. To perform well, the polic…

Cited by 60SourcePDFScholar
2019

Conditional Generative Neural System for Probabilistic Trajectory Prediction

IROS 2019poster

Effective understanding of the environment and accurate trajectory prediction of surrounding dynamic obstacles are critical for intelligent systems such as autonomous vehicles and wheeled mobile robotics navigating in complex scenarios to achieve safe and high-quality decision making, motion plannin…

Cited by 239SourceScholar
2019

Interaction-aware Multi-agent Tracking and Probabilistic Behavior Prediction via Adversarial Learning

ICRA 2019poster

In order to enable high-quality decision making and motion planning of intelligent systems such as robotics and autonomous vehicles, accurate probabilistic predictions for surrounding interactive objects is a crucial prerequisite. Although many research studies have been devoted to making prediction…

Cited by 78SourceScholar