← Search

yi chen

63 accepted papers

2026

ActAvatar: Temporally-Aware Precise Action Control for Talking Avatars

CVPR 2026

Despite significant advances in talking avatar generation, existing methods face critical challenges: insufficient text-following capability for diverse actions, lack of temporal alignment between actions and audio content, and dependency on additional control signals such as pose skeletons. We pres

Cited by 0SourceScholar
2026

An Open-Ended Benchmark and Formal Framework for Adjuvant Research with MLLM

ICLR 2026poster

Adjuvants play a critical role in modulating immune responses and are central to the development of vaccines and immunotherapies. Yet progress in this field is constrained by data scarcity and incomplete understanding of mechanisms of action, which limit the transition from experience-based design t…

Cited by 0SourceScholar
2026

CD-DPE: Dual-Prompt Expert Network Based on Convolutional Dictionary Feature Decoupling for Multi-Contrast MRI Super-Resolution

AAAI 2026technical

Multi-contrast magnetic resonance imaging (MRI) super-resolution intends to reconstruct high-resolution (HR) images from low-resolution (LR) scans by leveraging structural information present in HR reference images acquired with different contrasts. This technique enhances anatomical detail and soft

Cited by 0SourcePDFScholar
2026

FedMOP: Achieving Enhanced Privacy and Performance in Federated Learning via Momentum Orthogonal Projection

CVPR 2026

Federated Learning (FL) faces a fundamental dilemma: existing defenses against gradient leakage attacks (GLAs) invariably sacrifice model performance for privacy protection through noise injection or gradient clip. We introduce Federated Learning with Momentum-Based Orthogonal Projection (FedMOP), a

Cited by 0SourcecodeScholar
2026

Fine-Grained Post-Training Quantization for Large Vision Language Models with Quantization-Aware Integrated Gradients

CVPR 2026

Large Vision Language Models (LVLMs) have achieved remarkable success in a wide range of downstream tasks that require multimodal interaction, but their powerful capabilities come with substantial computational and memory overhead, which hinders practical deployment. Among numerous acceleration tech

Cited by 0SourcecodeScholar
2026

GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation

CVPR 2026

Driving World Models (DWMs) have been developing rapidly with the advances of generative models. However, existing DWMs lack 3D scene understanding capabilities and can only generate content conditioned on input data, without the ability to interpret or reason about the driving environment. Moreover

Cited by 0SourcecodeScholar
2026

Localizing, Structuring, and Rendering: Bridging 3D and 2D Vision-Language-Action Models for Robotic Manipulation

CVPR 2026

Robotic manipulation in complex 3D environments requires unifying spatial reasoning with intuitive visual perception, which is a capability that current Vision-Language-Action paradigms address separately. While 3D VLAs excel in geometric and physical reasoning, they lack intuitive, image-level unde

Cited by 0SourcecodeScholar
2026

MeteorPred: A Meteorological Multimodal Large Model and Dataset for Severe Weather Event Prediction

CVPR 2026

Timely and accurate forecasts of severe weather events are essential for early warning and for constraining downstream analysis and decision-making. Since severe weather events prediction still depends on subjective, time-consuming expert interpretation, end-to-end "AI weather station" systems are e

Cited by 0SourcecodeScholar
2026

MoCoDiff: A Controllable Autoregressive Diffusion Model for Expressive Motion Generation

CVPR 2026

Diffusion-based motion generation has advanced rapidly, but current methods still struggle with long-horizon consistency, style control, and multi-condition guidance. A major reason is the fused-conditioning design, where semantic, stylistic, and temporal signals share a single pathway, causing inte

Cited by 0SourceScholar
2026

Not Just What’s There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-Tuning

AAAI 2026technical

Vision-Language Models (VLMs) like CLIP struggle to understand negation, often embedding affirmatives and negatives similarly (e.g., matching "no dog" with dog images). Existing methods refine negation understanding via fine-tuning CLIP’s text encoder, risking overfitting. In this work, we propose C

Cited by 0SourcePDFScholar
2026

One Patch Doesn’t Fit All: Adaptive Patching for Native-Resolution Multimodal Large Language Models

ICLR 2026poster

Real-world visual signals are inherently variable in resolution, and it is natural to endow multimodal large language models (MLLMs) with such native-resolution perception capabilities. In principle, for general and straightforward multimodal understanding, low-resolution images are sufficient. Whil…

Cited by 0SourceScholar
2026

PCA-Seg: Revisiting Cost Aggregation for Open-Vocabulary Semantic and Part Segmentation

CVPR 2026

Recent advances in vision-language models (VLMs) have garnered substantial attention in open-vocabulary semantic and part segmentation (OSPS). However, existing methods extract image-text alignment cues from cost volumes through a serial structure of spatial and class aggregations, leading to knowle

Cited by 0SourcecodeScholar
2026

RANGER: A Monocular Zero-Shot Semantic Navigation Framework through Visual Contextual Adaptation

ICRA 2026poster

Efficient target localization and autonomous navigation in complex environments are fundamental to real-world embodied applications. While recent advances in multimodal foundation models have enabled zero-shot object goal navigation, allowing robots to search for arbitrary objects without fine-tunin…

2026

ROVER: Robust Generative Continual Identity Unlearning Against Relearning Attacks

AAAI 2026technical

Recent generative unlearning models synthesize high quality samples while protecting private information by unlearning the identity. However, existing generative identity unlearning methods face two challenges in multi-identity unlearning: 1) identity conflicts, which cause conflicts of model parame

Cited by 0SourcePDFScholar
2026

StreamAvatar: Streaming Diffusion Models for Real-Time Interactive Human Avatars

CVPR 2026

Real-time, streaming interactive avatars represent a critical yet challenging goal in digital human research. Although diffusion-based human avatar generation methods achieve remarkable success, their non-causal architecture and high computational costs make them unsuitable for streaming. Moreover,

Cited by 0SourcecodeScholar
2026

UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions

CVPR 2026

Due to the lack of effective cross-modal modeling, existing open-source audio-video generation methods often exhibit compromised lip synchronization and insufficient semantic consistency. To mitigate these drawbacks, we propose UniAVGen, a unified framework for human-centric joint audio and video ge

Cited by 0SourceScholar
2026

VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model

ICML 2026poster

Vision-Language-Action (VLA) models have demonstrated remarkable capabilities and generalization in embodied manipulation. However, their decision-making relies on a fast, instinctive process that lacks deliberation. This strategy often leads to suboptimal or catastrophic actions when facing complex…

Cited by 0SourceScholar
2025

Aerodynamic Coefficients Prediction via Cross-Attention Fusion and Physical-Informed Training

AAAI 2025technical

Aerodynamic coefficient prediction is pivotal in aircraft and vehicles' design, performance evaluation, and motion control. Integrating artificial neural networks into aerodynamic coefficient prediction offers a promising alternative to traditional numerical methods burdened by extensive computation…

Cited by 0SourcePDFScholar
2025

Learning Verified Safe Neural Network Controllers for Multi-Agent Path Finding

AAAI 2025technical

Multi-agent path finding (MAPF) is a safety-critical scenario where the goal is to secure collision-free trajectories from initial to desired locations. However, due to system complexity and uncertainty, integrating learning-based controllers with MAPF is challenging and cannot theoretically guarant…

Cited by 0SourcePDFScholar
2025

Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos

ICCV 2025poster

Recent developments in Large Language Models (LLMs) pre-trained on extensive corpora have shown significant success in various natural language processing (NLP) tasks with minimal fine-tuning. This success offers new promise for robotics, which has long been constrained by the high cost of action-la…

2025

PAL: Sample-Efficient Personalized Reward Modeling for Pluralistic Alignment

ICLR 2025poster

Foundation models trained on internet-scale data benefit from extensive alignment to human preferences before deployment. However, existing methods typically assume a homogeneous preference shared by all individuals, overlooking the diversity inherent in human values. In this work, we propose a gene…

Cited by 2SourcePDFScholar
2025

RecNet: Optimization for Dense Object Detection in Retail Scenarios Based on View Rectification

ICASSP 2025accepted

High-precision dense object detection in retail is crucial for automation, inventory management, and sales optimization. Our experiments revealed that detection models perform significantly better with frontal views than with oblique views, motivating the development of RecNet. RecNet utilizes a Rec…

Cited by 0SourceScholar
2025

Recoverable Compression: A Multimodal Vision Token Recovery Mechanism Guided by Text Information

AAAI 2025technical

With the advancement of large-scale language modeling techniques, large multimodal models combining visual encoders with large language models have demonstrated exceptional performance in various visual tasks. Most of the current large multimodal models achieve this by mapping visual features obtain…

2025

Rethinking Confidence Scores and Thresholds in Pseudolabeling-based SSL

ICML 2025poster

Modern semi-supervised learning (SSL) methods rely on pseudolabeling and consistency regularization. Pseudolabeling is typically performed by comparing the model's confidence scores and a predefined threshold. While several heuristics have been proposed to improve threshold selection, the underlyin…

Cited by 0SourcePDFScholar
2025

SN-LiDAR: Semantic Neural Fields for Novel Space-time View LiDAR Synthesis

IROS 2025

Recent research has begun exploring novel view synthesis (NVS) for LiDAR point clouds, aiming to generate realistic LiDAR scans from unseen viewpoints. However, most existing approaches do not reconstruct semantic labels, which are crucial for many downstream applications such as autonomous driving

Cited by 1SourcecodeScholar
2025

Sonic: Shifting Focus to Global Audio Perception in Portrait Animation

CVPR 2025poster

The study of talking face generation mainly explores the intricacies of synchronizing facial movements and crafting visually appealing, temporally-coherent animations. However, due to the limited exploration of global audio perception, current approaches predominantly employ auxiliary visual and sp…

Cited by 8SourcePDFScholar
2025

Supervised Optimism Correction: Be Confident When LLMs Are Sure

ACL 2025finding

In this work, we establish a novel theoretical connection between supervised fine-tuning and offline reinforcement learning under the token-level Markov decision process, revealing that large language models indeed learn an implicit Q-function for inference.Through this theoretical lens, we demonstr…

Cited by 0SourcePDFScholar
2025

TOPP-DWR: Time-Optimal Path Parameterization of Differential-Driven Wheeled Robots Considering Piecewise-Constant Angular Velocity Constraints

IROS 2025

Differential-driven wheeled robots (DWR) represent the quintessential type of mobile robots and find extensive applications across the robotic field. Most high-performance control approaches for DWR explicitly utilize the linear and angular velocities of the trajectory as control references. However

Cited by 1SourceScholar
2025

Towards Fast Correspondence-Free Odometry Using Multiple FMCW Lidars

RA-L 2025

3D FMCW lidars return relative velocity measurements via the Doppler effect, which provides a new form of information for motion estimation. In our prior work, we proposed an odometry method that avoids the conventional ICP-based approach and uses the Doppler velocity measurements in a correspondenc

Cited by 1SourceScholar
2025

Uncertainty-Participation Context Consistency Learning for Semi-supervised Semantic Segmentation

ICASSP 2025accepted

Semi-supervised semantic segmentation has attracted considerable attention for its ability to mitigate the reliance on extensive labeled data. However, existing consistency regularization methods only utilize high certain pixels with prediction confidence surpassing a fixed threshold for training, f…

Cited by 0SourceScholar
2025

Variance-Dependent Regret Bounds for Nonstationary Linear Bandits

AISTATS 2025poster

We investigate the non-stationary stochastic linear bandit problem where the reward distribution evolves each round. Existing algorithms characterize the non-stationarity by the total variation budget $B_K$, which is the summation of the change of the consecutive feature vectors of the linear bandit…

Cited by 0SourceScholar
2024

Beyond the Limit of Weight-Sharing: Pioneering Space-Evolving NAS with Large Language Models

ICASSP 2024accepted

Large language models (LLMs) offer impressive performance across diverse fields, but their increasing complexity raises both design costs and the need for specialized expertise. These challenges are intensified for Neural Architecture Search (NAS) methods reliant on weight-sharing techniques. This p…

Cited by 0SourceScholar
2024

Constrained Ensemble Exploration for Unsupervised Skill Discovery

ICML 2024poster

Unsupervised Reinforcement Learning (RL) provides a promising paradigm for learning useful behaviors via reward-free per-training. Existing methods for unsupervised RL mainly conduct empowerment-driven skill discovery or entropy-based exploration. However, empowerment often leads to static skills, a…

Cited by 6SourcePDFScholar
2024

Enhancing Large Language Models Against Inductive Instructions with Dual-critique Prompting

NAACL 2024long

Numerous works are proposed to align large language models (LLMs) with human intents to better fulfill instructions, ensuring they are trustful and helpful.Nevertheless, some human instructions are often malicious or misleading and following them will lead to untruthful and unsafe responses.Previous…

2024

Learning Populations of Preferences via Pairwise Comparison Queries

AISTATS 2024poster

Ideal point based preference learning using pairwise comparisons of type "Do you prefer a or b?" has emerged as a powerful tool for understanding how we make preferences. Existing preference learning approaches assume homogeneity and focus on learning preference on average over the population or req…

Cited by 5SourcePDFScholar
2024

MapLE: Matching Molecular Analogues Promptly with Low Computational Resources by Multi-Metrics Evaluation (Student Abstract)

AAAI 2024technical

Matching molecular analogues is a computational chemistry and bioinformatics research issue which is used to identify molecules that are structurally or functionally similar to a target molecule. Recent studies on matching analogous molecules have predominantly concentrated on enhancing effectivenes…

Cited by 0SourcePDFScholar
2024

Pearls from Pebbles: Improved Confidence Functions for Auto-labeling

NeurIPS 2024poster

Auto-labeling is an important family of techniques that produce labeled training sets with minimum manual annotation. A prominent variant, threshold-based auto-labeling (TBAL), works by finding thresholds on a model's confidence scores above which it can accurately automatically label unlabeled data…

Cited by 2SourcePDFScholar
2024

Role Prompting Guided Domain Adaptation with General Capability Preserve for Large Language Models

NAACL 2024findings

The growing interest in Large Language Models (LLMs) for specialized applications has revealed a significant challenge: when tailored to specific domains, LLMs tend to experience catastrophic forgetting, compromising their general capabilities and leading to a suboptimal user experience. Additionall…

2023

Enlightening the Student in Knowledge Distillation

ICASSP 2023accepted

Knowledge distillation is a common method of model compression, which uses large models (teacher networks) to guide the training of small models (student networks). However, the student may find a hard time absorbing the knowledge from a sophisticated teacher due to the capacity and confidence gaps…

Cited by 0SourceScholar
2023

Need for Speed: Fast Correspondence-Free Lidar-Inertial Odometry Using Doppler Velocity

IROS 2023poster

In this paper, we present a fast, lightweight odometry method that uses the Doppler velocity measurements from a Frequency-Modulated Continuous-Wave (FMCW) lidar without data association. FMCW lidar is a recently emerging technology that enables per-return relative radial velocity measurements via t…

Cited by 12SourceScholar
2023

Picking up Speed: Continuous-Time Lidar-Only Odometry Using Doppler Velocity Measurements

RA-L 2023

Frequency-Modulated Continuous-Wave (FMCW) lidar is a recently emerging technology that additionally enables per-return instantaneous relative radial velocity measurements via the Doppler effect. In this letter, we present the first continuous-time lidar-only odometry algorithm using these Doppler v

Cited by 40SourcecodeScholar
2023

Retrieval-free Knowledge Injection through Multi-Document Traversal for Dialogue Models

ACL 2023long

Dialogue models are often enriched with extensive external knowledge to provide informative responses through a retrieval-augmented pipeline. Nevertheless, retrieval-augmented approaches rely on finely annotated retrieval training data and knowledge-grounded response generation data, making it costl…

2023

SKD-NER: Continual Named Entity Recognition via Span-based Knowledge Distillation with Reinforcement Learning

EMNLP 2023long main

Continual learning for named entity recognition (CL-NER) aims to enable models to continuously learn new entity types while retaining the ability to recognize previously learned ones. However, the current strategies fall short of effectively addressing the catastrophic forgetting of previously learn…

Cited by 0SourceScholar
2022

Learning from Sibling Mentions with Scalable Graph Inference in Fine-Grained Entity Typing

ACL 2022long

In this paper, we firstly empirically find that existing models struggle to handle hard mentions due to their insufficient contexts, which consequently limits their overall typing performance. To this end, we propose to exploit sibling mentions for enhancing the mention representations. Specifically…

Cited by 10SourcePDFScholar
2022

MCPG: A Flexible Multi-Level Controllable Framework for Unsupervised Paraphrase Generation

EMNLP 2022finding

We present MCPG: a simple and effectiveapproach for controllable unsupervised paraphrase generation, which is also flexible toadapt to specific domains without extra training. MCPG is controllable in different levels: local lexicons, global semantics, and universal styles. The unsupervised paradigm…

Cited by 8SourcePDFScholar
2022

Signed Neuron with Memory: Towards Simple, Accurate and High-Efficient ANN-SNN Conversion

IJCAI 2022poster

Spiking Neural Networks (SNNs) are receiving increasing attention due to their biological plausibility and the potential for ultra-low-power event-driven neuromorphic hardware implementation. Due to the complex temporal dynamics and discontinuity of spikes, training SNNs directly usually suffers fro…

2021

An Empirical Study on Multiple Information Sources for Zero-Shot Fine-Grained Entity Typing

EMNLP 2021main

Auxiliary information from multiple sources has been demonstrated to be effective in zero-shot fine-grained entity typing (ZFET). However, there lacks a comprehensive understanding about how to make better use of the existing information sources and how they affect the performance of ZFET. In this p…

Cited by 16SourcePDFScholar
2021

Deep Spiking Neural Network with Neural Oscillation and Spike-Phase Information

AAAI 2021technical

Deep spiking neural network (DSNN) is a promising computational model towards artificial intelligence. It benefits from both the DNNs and SNNs through a hierarchy structure to extract multiple levels of abstraction and the event-driven computational manner to provide ultra-low-power neuromorphic imp…

Cited by 17SourcePDFScholar
2021

When Shall I Be Empathetic? The Utility of Empathetic Parameter Estimation in Multi-Agent Interactions

ICRA 2021poster

Human-robot interactions (HRI) can be modeled as differential games with incomplete information, where each agent holds private reward parameters. Due to the open challenge in finding perfect Bayesian equilibria of such games, existing studies often decouple the belief and physical dynamics by itera…

Cited by 11SourceScholar
2020

Enabling Robot to Assist Human in Collaborative Assembly using Convolutional Neural Networks

IROS 2020poster

Human-robot collaborative assembly consists of humans and automated robots, who cooperate with each other to accomplish complex assembly tasks, which are difficult for either humans or robots to accomplish alone. There has been some success in statistics-based and optimization-based approaches to re…

Cited by 7SourceScholar
2020

Intra Frame Rate Control for Versatile Video Coding with Quadratic Rate-Distortion Modelling

ICASSP 2020accepted

With numerous coding tools adopted in the forthcoming Versatile Video Coding (VVC) standard, much less work has been dedicated to study the corresponding Rate-Distortion (R-D) characteristics. This paper proposes a new quadratic R-D model for Versatile Video Coding. In particular, based on the propo…

Cited by 0SourceScholar
2020

Rapid Fabrication of Electro-Adhesive Devices With Inkjet Printed Electrodes

RA-L 2020

This letter proposes a procedure for the rapid prototyping and on-demand manufacturing of thin film flexible electro-adhesive devices (EADs) made with a commercial polyimide dielectric layer, inkjet printed interdigitated silver electrodes and blade coated silicone elastomer encapsulation backing. A

Cited by 14SourceScholar
2020

Real-Time Adaptive Assembly Scheduling in Human-Multi-Robot Collaboration According to Human Capability

ICRA 2020poster

Human-multi-robot collaboration is becoming more and more common in intelligent manufacturing. Optimal assembly scheduling of such systems plays a critical role in their production efficiency. Existing approaches mostly consider humans as agents with assumed or known capabilities, which leads to sub…

Cited by 37SourceScholar
2019

ACCELERATING NONCONVEX LEARNING VIA REPLICA EXCHANGE LANGEVIN DIFFUSION

ICLR 2019poster

Langevin diffusion is a powerful method for nonconvex optimization, which enables the escape from local minima by injecting noise into the gradient. In particular, the temperature parameter controlling the noise level gives rise to a tradeoff between ``global exploration'' and ``local exploitation''…

Cited by 44SourcePDFScholar