← Search

Li Chen

78 accepted papers

2026

Agility Meets Stability: Versatile Humanoid Control with Heterogeneous Data

ICRA 2026poster

Humanoid robots are envisioned to perform a wide range of tasks in human-centered environments, requiring controllers that combine agility with robust balance. Recent advances in locomotion and whole-body tracking have enabled impressive progress in either agile dynamic skills or stability-critical …

2026

Configurable Reward Model for Balanced Safety Alignment

ICML 2026poster

Aligning large language models (LLMs) to heterogeneous and rapidly evolving safety requirements remains a critical challenge. Existing instruction-tuned LLMs and standalone safety classifiers often fail to generalize to new safety configurations, motivating the need for Reward Models (RMs) that are …

Cited by 0SourceScholar
2026

DataGuard: A Non-intrusive Dataset Auditing Framework via Differential Information Forensics

ICML 2026poster

Concerns over dataset misuse in deep learning have highlighted the need for effective auditing. Unlike existing intrusive methods that require dataset modifications, which risk model performance and security, we present DataGuard, a non-intrusive framework for quantitative dataset auditing. Specific…

Cited by 0SourceScholar
2026

ENTROPY-GUIDED DATA-EFFICIENT TRAINING FOR MULTIMODAL REASONING REWARD MODELS

ICASSP 2026poster

Multimodal reward models are crucial for aligning multimodal large language models with human preferences. Recent works have incorporated reasoning capabilities into these models, achieving promising results. However, training these models suffers from two critical challenges: (1) the inherent noise…

Cited by 0SourcePDFScholar
2026

FreeTacMan: Robot-Free Visuo-Tactile Data Collection System for Contact-Rich Manipulation

ICRA 2026poster

Enabling robots with contact-rich manipulation remains a pivotal challenge in robot learning, which is substantially hindered by the data collection gap, including its inefficiency and limited sensor setup. While prior work has explored handheld paradigms, their rod-based mechanical structures remai…

2026

Ice Cream Doesn’t Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal Inference

ICLR 2026poster

Reliable causal inference is essential for making decisions in high-stakes areas like medicine, economics, and public policy. However, it remains unclear whether large language models (LLMs) can handle rigorous and trustworthy \textit{statistical causal inference}. Current benchmarks usually involve…

Cited by 0SourcecodeScholar
2026

Model Whisper: Steering Vectors Unlock Large Language Models’ Potential in Test-Time

AAAI 2026technical

It is a critical challenge to efficiently unlock the powerful reasoning potential of Large Language Models (LLMs) for specific tasks or new distributions. Existing test-time adaptation methods often require tuning model parameters, which is not only computationally expensive but also risks degrading

Cited by 0SourcePDFScholar
2026

Revisiting the Seasonal Trend Decomposition for Enhanced Time Series Forecasting

ICASSP 2026poster

Time series forecasting presents significant challenges in real-world applications across various domains. Building upon the decomposition of the time series, we enhance the architecture of machine learning models for better multivariate time series forecasting. To achieve this, we focus on the tren…

Cited by 0SourcePDFScholar
2026

Seeing Farther and Smarter: Value-Guided Multi-Path Reflection for VLM Policy Optimization

ICRA 2026poster

Solving complex, long-horizon robotic manipulation tasks requires a deep understanding of physical interactions, reasoning about their long-term consequences, and precise high-level planning. Vision-Language Models (VLMs) offer a general perceive-reason-act framework for this goal. However, previous…

2026

Self-Improving Robot Policy with Compositional World Model

RSS 2026poster

Despite the sustained scaling on model capacity and data acquisition, Vision–Language–Action (VLA) models remain brittle in contact-rich and dynamic manipulation tasks, where minor execution deviations can compound into failures. While reinforcement learning (RL) offers a principled path to robustne…

Cited by 0SourceScholar
2026

Task-and-Model-Aware Fractal-Consistency for Efficient LLM Reasoning

ICML 2026poster

While self-consistency methods have emerged as a promising approach to enhance the correctness of large language model (LLM) outputs by aggregating multiple stochastic samples, they suffer from two critical limitations, resulting in high computation cost. First, they evaluate output consistency mono…

Cited by 0SourceScholar
2026

Unlocking In-the-Wild Loco-Manipulation with Robot-Free Egocentric Demonstration

RSS 2026poster

Human demonstrations offer rich environmental diversity and scale naturally, making them an appealing alternative to robot teleoperation. While this paradigm has advanced robot-arm manipulation, its potential for the more challenging, data-hungry problem of humanoid loco-manipulation remains largely…

Cited by 0SourceScholar
2026

VGDM: Visual Localization-Guided 3D Dental Segmentation via Extrinsic–Intrinsic Bridging

IJCAI 2026

3D dental segmentation is a key task in digital dentistry. In real intraoral scans data (IOS), occlusion, scanning noise, and reconstruction artifacts often break down the geometric separation structure between teeth, resulting in adjacent teeth being incorrectly merged or a single tooth being over-

Cited by 0Scholar
2026

WholeBodyVLA: Towards Unified Latent VLA for Whole-body Loco-manipulation Control

ICLR 2026poster

Humanoid robots require precise locomotion and dexterous manipulation to per- form challenging locomanipulation tasks. Yet existing approaches, modular or end-to-end, are deficient in manipulation-aware locomotion. This confines the robot to a limited workspace, preventing it from performing large-s…

Cited by 0SourcecodeScholar
2026

WideSearch: Benchmarking Agentic Broad Info-Seeking

ICLR 2026poster

From professional research to everyday planning, many tasks are bottlenecked by wide-scale information seeking, which is more repetitive than cognitively complex. With the rapid development of Large Language Models (LLMs), automated search agents powered by LLMs offer a promising solution to liberat…

Cited by 0SourcecodeScholar
2025

From Pairwise to Ranking: Climbing the Ladder to Ideal Collaborative Filtering with Pseudo-Ranking

AAAI 2025technical

Intuitively, an ideal collaborative filtering (CF) model should learn from users' full rankings over all items to make optimal top-K recommendations. Due to the absence of such full rankings in practice, most CF models rely on pairwise loss functions to approximate full rankings, resulting in an imm…

Cited by 1SourcePDFScholar
2025

MLLM-as-a-Judge for Image Safety without Human Labeling

CVPR 2025highlight

Image content safety has become a significant challenge with the rise of visual media on online platforms. Meanwhile, in the age of AI-generated content (AIGC), many image generation models are capable of producing harmful content, such as images containing sexual or violent material. Thus, it becom…

Cited by 2SourcePDFScholar
2025

QCS:Feature Refining from Quadruplet Cross Similarity for Facial Expression Recognition

AAAI 2025technical

Facial expression recognition faces challenges where labeled significant features in datasets are mixed with unlabeled redundant ones. In this paper, we introduce Cross Similarity Attention (CSA) to mine richer intrinsic information from image pairs, overcoming a limitation when the Scaled Dot-Produ…

2025

ReSim: Reliable World Simulation for Autonomous Driving

NeurIPS 2025spotlight

How can we reliably simulate future driving scenarios under a wide range of ego driving behaviors? Recent driving world models, developed exclusively on real-world driving data composed mainly of safe expert trajectories, struggle to follow hazardous or non-expert behaviors, which are rare in such d…

Cited by 0SourceScholar
2025

Simplifying Control Mechanism in Text-to-Image Diffusion Models

AAAI 2025technical

ControlNet has significantly advanced controllable image generation by integrating dense conditions (such as depth and canny edges) with text-to-image diffusion models. However, ControlNet's integration requires an additional amount nearly equal to half of the base diffusion model's parameters, maki…

2025

Transcending Cost-Quality Tradeoff in Agent Serving via Session-Awareness

NeurIPS 2025poster

Large Language Model (LLM) agents are capable of task execution across various domains by autonomously interacting with environments and refining LLM responses based on feedback. However, existing model serving systems are not optimized for the unique demands of serving agents. Compared to classic m…

Cited by 0SourceScholar
2025

XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation

NeurIPS 2025poster

Achieving fine-grained control over subject identity and semantic attributes (pose, style, lighting) in text-to-image generation, particularly for multiple subjects, often undermines the editability and coherence of Diffusion Transformers (DiTs). Many approaches introduce artifacts or suffer from at…

Cited by 0SourcecodeScholar
2024

$\texttt{Model-GLUE}$: Democratized LLM Scaling for A Large Model Zoo in the Wild

NeurIPS 2024poster

As Large Language Models (LLMs) excel across tasks and specialized domains, scaling LLMs based on existing models has gained significant attention, which is challenged by potential performance drop when combining disparate models. Various techniques have been proposed to aggregate pre-trained LLMs,…

2024

Adaptive Hardness Negative Sampling for Collaborative Filtering

AAAI 2024technical

Negative sampling is essential for implicit collaborative filtering to provide proper negative training signals so as to achieve desirable performance. We experimentally unveil a common limitation of all existing negative sampling methods that they can only select negative samples of a fixed hardnes…

2024

AvatarVerse: High-Quality & Stable 3D Avatar Creation from Text and Pose

AAAI 2024technical

Creating expressive, diverse and high-quality 3D avatars from highly customized text descriptions and pose guidance is a challenging task, due to the intricacy of modeling and texturing in 3D that ensure details and various styles (realistic, fictional, etc). We present AvatarVerse, a stable pipelin…

2024

Closed-Loop Visuomotor Control with Generative Expectation for Robotic Manipulation

NeurIPS 2024poster

Despite significant progress in robotics and embodied AI in recent years, deploying robots for long-horizon tasks remains a great challenge. Majority of prior arts adhere to an open-loop philosophy and lack real-time feedback, leading to error accumulation and undesirable robustness. A handful of ap…

2024

DriveLM: Driving with Graph Visual Question Answering

ECCV 2024oral

"We study how vision-language models (VLMs) trained on web-scale data can be integrated into end-to-end driving systems to boost generalization and enable interactivity with human users. While recent approaches adapt VLMs to driving via single-round visual question answering (VQA), human drivers rea…

2024

Fairness-Aware Meta-Learning via Nash Bargaining

NeurIPS 2024poster

To address issues of group-level fairness in machine learning, it is natural to adjust model parameters based on specific fairness objectives over a sensitive-attributed validation set. Such an adjustment procedure can be cast within a meta-learning framework. However, naive integration of fairness…

Cited by 2SourcePDFScholar
2024

FontStudio: Shape-Adaptive Diffusion Model for Coherent and Consistent Font Effect Generation

ECCV 2024poster

"Recently, the application of modern diffusion-based text-to-image generation models for creating artistic fonts, traditionally the domain of professional designers, has garnered significant interest. Diverging from the majority of existing studies that concentrate on generating artistic typography,…

2024

Fully Sparse 3D Occupancy Prediction

ECCV 2024poster

"Occupancy prediction plays a pivotal role in autonomous driving. Previous methods typically construct dense 3D volumes, neglecting the inherent sparsity of the scene and suffering high computational costs. To bridge the gap, we introduce a novel fully sparse occupancy network, termed SparseOcc. Spa…

2024

Generalized Predictive Model for Autonomous Driving

CVPR 2024highlight

In this paper we introduce the first large-scale video prediction model in the autonomous driving discipline. To eliminate the restriction of high-cost data collection and empower the generalization ability of our model we acquire massive data from the web and pair it with diverse and high-quality t…

Cited by 61SourcePDFScholar
2024

LaneSegNet: Map Learning with Lane Segment Perception for Autonomous Driving

ICLR 2024poster

A map, as crucial information for downstream applications of an autonomous driving system, is usually represented in lanelines or centerlines. However, existing literature on map learning primarily focuses on either detecting geometry-based lanelines or perceiving topology relationships of centerlin…

2024

Large Language Models for Generative Recommendation: A Survey and Visionary Discussions

COLING 2024main

Large language models (LLM) not only have revolutionized the field of natural language processing (NLP) but also have the potential to reshape many other fields, e.g., recommender systems (RS). However, most of the related work treats an LLM as a component of the conventional recommendation pipeline…

Cited by 105SourcePDFScholar
2024

Learning Manipulation by Predicting Interaction

RSS 2024poster

Representation learning approaches for robotic manipulation have boomed in recent years. Due to the scarcity of in-domain robot data, prevailing methodologies tend to leverage large-scale human video datasets to extract generalizable features for visuomotor policy learning. Despite the progress achi…

2024

Named Entity Driven Zero-Shot Image Manipulation

CVPR 2024poster

We introduced StyleEntity a zero-shot image manipulation model that utilizes named entities as proxies during its training phase. This strategy enables our model to manipulate images using unseen textual descriptions during inference all within a single training phase. Additionally we proposed an in…

2024

Reasoning Multi-Agent Behavioral Topology for Interactive Autonomous Driving

NeurIPS 2024poster

Autonomous driving system aims for safe and social-consistent driving through the behavioral integration among interactive agents. However, challenges remain due to multi-agent scene uncertainty and heterogeneous interaction. Current dense and sparse behavioral representations struggle with ineffici…

2024

Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability

NeurIPS 2024poster

World models can foresee the outcomes of different actions, which is of paramount importance for autonomous driving. Nevertheless, existing driving world models still have limitations in generalization to unseen environments, prediction fidelity of critical details, and action controllability for fl…

2024

Visual Point Cloud Forecasting enables Scalable Autonomous Driving

CVPR 2024highlight

In contrast to extensive studies on general vision pre-training for scalable visual autonomous driving remains seldom explored. Visual autonomous driving applications require features encompassing semantics 3D geometry and temporal information simultaneously for joint perception prediction and plann…

2023

A Lightweight Fourier Convolutional Attention Encoder for Multi-Channel Speech Enhancement

ICASSP 2023accepted

Beamforming weights prediction via deep neural networks has been one of the main methods in multi-channel speech enhancement tasks. The spectral-spatial cues are crucial in beamforming weights estimation, however, many existing works fail to optimally predict the beamforming weights with an absence…

Cited by 0SourceScholar
2023

Distilling Focal Knowledge From Imperfect Expert for 3D Object Detection

CVPR 2023poster

Multi-camera 3D object detection blossoms in recent years and most of state-of-the-art methods are built up on the bird's-eye-view (BEV) representations. Albeit remarkable performance, these works suffer from low efficiency. Typically, knowledge distillation can be used for model compression. Howeve…

2023

DriveAdapter: Breaking the Coupling Barrier of Perception and Planning in End-to-End Autonomous Driving

ICCV 2023oral

End-to-end autonomous driving aims to build a fully differentiable system that takes raw sensor data as inputs and directly outputs the planned trajectory or control signals of the ego vehicle. State-of-the-art methods usually follow the `Teacher-Student' paradigm. The Teacher model uses privileged…

Cited by 59PDFcodeScholar
2023

ERNIE-ViLG 2.0: Improving Text-to-Image Diffusion Model With Knowledge-Enhanced Mixture-of-Denoising-Experts

CVPR 2023highlight

Recent progress in diffusion models has revolutionized the popular technology of text-to-image generation. While existing approaches could produce photorealistic high-resolution images with text conditions, there are still several open problems to be solved, which limits the further improvement of i…

Cited by 140SourcePDFScholar
2023

FABRIKv: A Fast, Iterative Inverse Kinematics Solver for Surgical Continuum Robot with Variable Curvature Model

IROS 2023poster

Due to the advantages of high flexibility, large workspace, and good human-body compatibility, flexible tendon-driven surgical continuum robots have attracted a lot of attention in robot-assisted minimally invasive surgery. However, due to the coupling of the position and angle of the continuum robo…

Cited by 1SourceScholar
2023

Local and Global Logit Adjustments for Long-Tailed Learning

ICCV 2023poster

Multi-expert ensemble models for long-tailed learning typically either learn diverse generalists from the whole dataset or aggregate specialists on different subsets. However, the former is insufficient for tail classes due to the high imbalance factor of the entire dataset, while the latter may bri…

Cited by 25PDFScholar
2023

MMST-ViT: Climate Change-aware Crop Yield Prediction via Multi-Modal Spatial-Temporal Vision Transformer

ICCV 2023poster

Precise crop yield prediction provides valuable information for agricultural planning and decision-making processes. However, timely predicting crop yields remains challenging as crop growth is sensitive to growing season weather variation and climate change. In this work, we develop a deep learning…

Cited by 45PDFcodeScholar
2023

OpenLane-V2: A Topology Reasoning Benchmark for Unified 3D HD Mapping

NeurIPS 2023poster

Accurately depicting the complex traffic scene is a vital component for autonomous vehicles to execute correct judgments. However, existing benchmarks tend to oversimplify the scene by solely focusing on lane perception tasks. Observing that human drivers rely on both lanes and traffic signals to op…

2023

Personalized Speech Enhancement Combining Band-Split RNN and Speaker Attentive Module

ICASSP 2023accepted

Target speaker information can be utilized in speech enhancement (SE) models to more effectively extract the desired speech. Previous works introduce the speaker embedding into speech enhancement models by means of concatenation or affine transformation. In this paper, we propose a speaker attentive…

Cited by 0SourceScholar
2023

Planning-Oriented Autonomous Driving

CVPR 2023poster

Modern autonomous driving system is characterized as modular tasks in sequential order, i.e., perception, prediction, and planning. In order to perform a wide diversity of tasks and achieve advanced-level intelligence, contemporary approaches either deploy standalone models for individual tasks, or…

2023

Policy Pre-training for Autonomous Driving via Self-supervised Geometric Modeling

ICLR 2023poster

Witnessing the impressive achievements of pre-training techniques on large-scale data in the field of computer vision and natural language processing, we wonder whether this idea could be adapted in a grab-and-go spirit, and mitigate the sample inefficiency problem for visuomotor driving. Given the…

2023

REASONER: An Explainable Recommendation Dataset with Comprehensive Labeling Ground Truths

NeurIPS 2023poster

Explainable recommendation has attracted much attention from the industry and academic communities. It has shown great potential to improve the recommendation persuasiveness, informativeness and user satisfaction. In the past few years, while a lot of promising explainable recommender models have be…

2023

Self-Paced Learning Based Graph Convolutional Neural Network for Mixed Integer Programming (Student Abstract)

AAAI 2023technical

Graph convolutional neural network (GCN) based methods have achieved noticeable performance in solving mixed integer programming problems (MIPs). However, the generalization of existing work is limited due to the problem structure. This paper proposes a self-paced learning (SPL) based GCN network (S…

Cited by 3SourcePDFScholar
2023

SwiftAvatar: Efficient Auto-Creation of Parameterized Stylized Character on Arbitrary Avatar Engines

AAAI 2023technical

The creation of a parameterized stylized character involves careful selection of numerous parameters, also known as the "avatar vectors" that can be interpreted by the avatar engine. Existing unsupervised avatar vector estimation methods that auto-create avatars for users, however, often fail to wor…

2023

Think Twice Before Driving: Towards Scalable Decoders for End-to-End Autonomous Driving

CVPR 2023poster

End-to-end autonomous driving has made impressive progress in recent years. Existing methods usually adopt the decoupled encoder-decoder paradigm, where the encoder extracts hidden features from raw sensor data, and the decoder outputs the ego-vehicle's future trajectories or actions. Under such a p…

2023

Two-Stage Neural Network for ICASSP 2023 Speech Signal Improvement Challenge

ICASSP 2023accepted

In ICASSP 2023 speech signal improvement challenge, we developed a dual-stage neural model which improves speech signal quality induced by different distortions in a stage-wise divide-and-conquer fashion. Specifically, in the first stage, the speech improvement network focuses on recovering the miss…

Cited by 0SourceScholar
2023

Two-Step Band-Split Neural Network Approach For Full-Band Residual Echo Suppression

ICASSP 2023accepted

This paper describes a Two-step Band-split Neural Network (TBNN) approach for full-band acoustic echo cancellation. Specifically, after linear filtering, we split the full-band signal into wideband (16KHz) and high-band (16-48KHz) for residual echo removal with lower modeling difficulty. The wide-ba…

Cited by 0SourceScholar
2022

Cloning One's Voice Using Very Limited Data in the Wild

ICASSP 2022accepted

With the increasing popularity of speech synthesis products, the industry has put forward more requirements for personalized speech synthesis: (1) How to use low-resource, easily accessible data to clone a person’s voice. (2) How to clone a person’s voice while controlling the style and prosody. To…

Cited by 0SourceScholar
2022

PersFormer: 3D Lane Detection via Perspective Transformer and the OpenLane Benchmark

ECCV 2022poster

"Methods for 3D lane detection have been recently proposed to address the issue of inaccurate lane layouts in many autonomous driving scenarios (uphill/downhill, bump, etc.). Previous work struggled in complex cases due to their simple designs of the spatial transformation between front view and bir…

2022

Robust Landmark-Based Stent Tracking in X-Ray Fluoroscopy

ECCV 2022poster

"In clinical procedures of angioplasty (i.e., open clogged coronary arteries), devices such as balloons and stents need to be placed and expanded in arteries under the guidance of X-ray fluoroscopy. Due to the limitation of X-ray dose, the resulting images are often noisy. To check the correct place…

Cited by 7SourcePDFScholar
2022

ST-P3: End-to-End Vision-Based Autonomous Driving via Spatial-Temporal Feature Learning

ECCV 2022poster

"Many existing autonomous driving paradigms involve a multi-stage discrete pipeline of tasks. To better predict the control signals and enhance user safety, an end-to-end approach that benefits from joint spatial-temporal feature learning is desirable. While there are some pioneering works on LiDAR-…

2022

SoftCollage: A Differentiable Probabilistic Tree Generator for Image Collage

CVPR 2022poster

Image collage task aims to create an informative and visual-aesthetic visual summarization for an image collection. While several recent works exploit tree-based algorithm to preserve image content better, all of them resort to hand-crafted adjustment rules to optimize the collage tree structure, le…

Cited by 2PDFcodeScholar
2022

Towards Capturing the Temporal Dynamics for Trajectory Prediction: a Coarse-to-Fine Approach

CoRL 2022poster

Trajectory prediction is one of the basic tasks in the autonomous driving field, which aims to predict the future position of other agents around the ego vehicle so that a safe yet efficient driving plan could be generated in the downstream module. Recently, deep learning based methods dominate the…

Cited by 37SourceScholar
2022

Trajectory-guided Control Prediction for End-to-end Autonomous Driving: A Simple yet Strong Baseline

NeurIPS 2022accept

Current end-to-end autonomous driving methods either run a controller based on a planned trajectory or perform control prediction directly, which have spanned two separately studied lines of research. Seeing their potential mutual benefits to each other, this paper takes the initiative to explore th…

2021

Finding Optimal Tangent Points for Reducing Distortions of Hard-label Attacks

NeurIPS 2021poster

One major problem in black-box adversarial attacks is the high query complexity in the hard-label attack setting, where only the top-1 predicted label is available. In this paper, we propose a novel geometric-based approach called Tangent Attack (TA), which identifies an optimal tangent point of a v…

2020

Content Adaptive and Error Propagation Aware Deep Video Compression

ECCV 2020poster

Recently, learning based video compression methods attract increasing attention. However, previous works suffer from error propagation, which stems from the accumulation of reconstructed error in inter predictive coding. Meanwhile, previous learning based video codecs are also not adaptive to differ…

Cited by 160SourcePDFScholar
2018

Post: Device Placement with Cross-Entropy Minimization and Proximal Policy Optimization

NeurIPS 2018poster

Training deep neural networks requires an exorbitant amount of computation resources, including a heterogeneous mix of GPU and CPU devices. It is critical to place operations in a neural network on these devices in an optimal way, so that the training process can complete within the shortest amount…

Cited by 36SourcePDFScholar
2018

Rcdfnn: Robust Change Detection Based on Convolutional Fusion Neural Network

ICASSP 2018accepted

Video change detection, which plays an important role in computer vision, is far from being well resolved due to the complexity of diverse scenes in real world. Most of the current methods are designed based on hand-crafted features and perform well in some certain scenes but may fail on others. Thi…

Cited by 0SourceScholar
2016

Principal components analysis-based visual saliency detection

ICASSP 2016accepted

In this paper, a novel patch-wise saliency detection algorithm is proposed based on Principal Component Analysis (PCA). As a powerful statistical procedure in data analysis, PCA are fully exploited to convert color space and produce compact patch representation. Specifically, images are first conver…

Cited by 0SourceScholar