← Search

Cheng Chen

74 accepted papers

2026

Beyond Adapter Retrieval: Latent Geometry-Preserving Composition via Sparse Task Projection

AAAI 2026technical

Recent advances in parameter-efficient transfer learning have demonstrated the utility of composing LoRA adapters from libraries of pretrained modules. However, most existing approaches rely on simple retrieval heuristics or uniform averaging, which overlook the latent structure of task relationshi

Cited by 0SourcePDFScholar
2026

CodeDelegator: Mitigating Context Pollution via Role Separation in Code-as-Action Agents

IJCAI 2026

Recent advances in large language models (LLMs) allow agents to represent actions as executable code, offering greater expressivity than traditional tool-calling. However, real-world tasks often demand both strategic planning and detailed implementation. Using a single agent for both leads to contex

Cited by 0Scholar
2026

Compositional Perception and Generalizing Induction: Latent Compositional Manifold Assumption on Generalized Category Discovery

ICML 2026poster

Generalized Category Discovery (GCD) assigns unlabeled instances, mixed with labeled data, to known or novel categories, requiring human-like compositional reasoning: reusing primitives learned from known classes and deciding when new combinations imply new categories. Existing GCD methods operate o…

Cited by 0SourceScholar
2026

DGP: A Dual-Granularity Prompting Framework for Fraud Detection with Graph-Enhanced LLMs

AAAI 2026technical

Real-world fraud detection applications benefit from graph learning techniques that jointly exploit node features—often rich in textual data—and graph structural information. Recently, Graph-Enhanced LLMs have emerged as a promising graph learning approach that converts graph information into prompt

Cited by 0SourcePDFScholar
2026

HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives

CVPR 2026

State-of-the-art text-to-video models excel at generating isolated clips but fall short of creating the coherent, multi-shot narratives, which are the essence of storytelling. We bridge this "narrative gap" with HoloCine, a model that generates entire scenes holistically to ensure global consistency

Cited by 0SourcecodeScholar
2026

JANUS-LORA: A Balanced Low-Rank Adaptation for Continual Learning

ICML 2026poster

Low-Rank Adaptation (LoRA) has emerged as a promising paradigm for Continual Learning. It independently updates its low-rank factors ($A$ and $B$), creating a composite update to the full weight matrix through their interaction. To prevent catastrophic forgetting, this update should remain orthogona…

Cited by 0SourceScholar
2026

Multi-Subspace Multi-Modal Modeling for Diffusion Models: Estimation, Convergence and Mixture of Experts

ICLR 2026poster

Recently, diffusion models have achieved a great performance with a small dataset of size $n$ and a fast optimization process. Despite the impressive performance, the estimation error suffers from the curse of dimensionality $n^{-1/D}$, where $D$ is the data dimension. Since images are usually a un…

Cited by 0SourceScholar
2026

OneSparse: A Unified Framework for Sparse Activation Layers in Vision Models

CVPR 2026

Sparse activation layers, primarily Mixture-of-Experts (MoE) and memory-based modules, have become a central approach for scaling large models and are gaining traction in vision tasks. Despite conceptual similarities, these paradigms have evolved independently, hindering systematic comparison and th

Cited by 0SourcecodeScholar
2026

Robust Decentralized Multi-armed Bandits: From Corruption-Resilience to Byzantine-Resilience

AAAI 2026technical

Decentralized cooperative multi-agent multi-armed bandits (DeCMA2B) considers how multiple agents collaborate in a decentralized multi-armed bandit setting. Though this problem has been extensively studied in previous work, most existing methods remain susceptible to various adversarial attacks. In

Cited by 0SourcePDFScholar
2026

SAMTok: Representing Any Mask with Two Words

CVPR 2026

Pixel-wise capabilities are essential for building interactive intelligent systems. However, pixel-wise multi-modal LLMs (MLLMs) remain difficult to scale due to complex region-level encoders, specialized segmentation decoders, and incompatible training objectives. To address these challenges, we pr

Cited by 0SourcecodeScholar
2026

SCOPE: Safety-Constrained Online Preview Enforcement for Efficient Encirclement in Multi-UAV Pursuit-Evasion

IJCAI 2026

Unmanned aerial vehicle swarms in pursuit-evasion requires encirclement efficiency while maintaining safety constraints, facing a critical safety-efficiency trade-off. Existing safe multi-agent reinforcement learning (MARL) methods often yield either unsafe task policies or conservative policies. Th

Cited by 0Scholar
2026

ScDiVa: Masked Discrete Diffusion for Joint Modeling of Single-Cell Identity and Expression

ICML 2026poster

Single-cell RNA-seq profiles are high-dimensional, sparse, and unordered, causing autoregressive generation to impose an artificial ordering bias and suffer from error accumulation. To address this, we propose scDiVa, a masked discrete diffusion foundation model that aligns generation with the dropo…

Cited by 0SourceScholar
2026

SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead

CVPR 2026

Vision-Language-Action (VLA) models built on pretrained Vision-Language Models (VLMs) show strong potential but are limited in practicality due to their large parameter counts. To mitigate this issue, using a lightweight VLM has been explored, but it compromises spatiotemporal reasoning. Although so

Cited by 0SourcecodeScholar
2026

SynerMedGen: Synergizing Medical Multimodal Understanding with Generation via Task Alignment

ICML 2026poster

Unifying multimodal understanding and generation is a compelling frontier that is beginning to emerge in the medical field. However, the limited existing unified medical models typically treat understanding and generation as disjoint objectives, lacking a meaningful functional synergy. In this work,…

Cited by 0SourceScholar
2026

Tackling Dual-stage Missing Modalities in Brain Tumor Segmentation via Robust Modality Reconstruction and Prompt-guided Modality Adaptation

AAAI 2026technical

Addressing missing modalities is a critical challenge in multimodal brain tumor segmentation. Most existing approaches merely handle modality-incomplete inputs during inference, assuming a full set of modalities for all training samples. However, this unrealistic assumption limits the usage of abund

Cited by 0SourcePDFScholar
2026

The Accumulation of Score Estimation Error in Diffusion Models

ICML 2026poster

Diffusion models are widely used for high-quality generation, but their performance is sensitive to the accuracy of the estimated score. We first develop our main bounds in a Gaussian-mixture setting, where the score admits a closed-form structure and the score Hessian can be controlled explicitly, …

Cited by 0SourceScholar
2026

Unleashing the Power of Image-Tabular Self-Supervised Learning via Breaking Cross-Tabular Barriers

AAAI 2026technical

Multi-modal learning integrating medical images and tabular data has significantly advanced clinical decision-making in recent years. Self-Supervised Learning (SSL) has emerged as a powerful paradigm for pretraining these models on large-scale unlabeled image-tabular data, aiming to learn discrimina

Cited by 0SourcePDFScholar
2026

Verification and Co-Alignment via Heterogeneous Consistency for Preference-Aligned LLM Annotations

ICLR 2026poster

Large Language Models (LLMs) are increasingly expected to be culturally customizable and personally aligned for natural language understanding (NLU). However, existing methods, from supervised fine-tuning (SFT) to personalized RLHF and prompting, either require costly large-scale annotations or rema…

Cited by 0SourceScholar
2026

Vision-Only Gaussian Splatting for Collaborative Semantic Occupancy Prediction

AAAI 2026technical

Collaborative perception enables connected vehicles to share information, overcoming occlusions and extending the limited sensing range inherent in single-agent (non-collaborative) systems. Existing vision-only methods for 3D semantic occupancy prediction commonly rely on dense 3D voxels, which incu

Cited by 0SourcePDFScholar
2026

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation

CVPR 2026

Pre-trained video models learn powerful priors for generating high-quality, temporally coherent content. While these models excel at temporal coherence, their dynamics are often constrained by the continuous nature of their training data. We hypothesize that by injecting the rich and unconstrained c

Cited by 0SourcecodeScholar
2025

A Near-optimal, Scalable and Parallelizable Framework for Stochastic Bandits Robust to Adversarial Corruptions and Beyond

NeurIPS 2025poster

We investigate various stochastic bandit problems in the presence of adversarial corruptions. A seminal work for this problem is the BARBAR~\cite{gupta2019better} algorithm, which achieves both robustness and efficiency. However, it suffers from a regret of $O(KC)$, which does not match the lower bo…

Cited by 0SourceScholar
2025

ADHMR: Aligning Diffusion-based Human Mesh Recovery via Direct Preference Optimization

ICML 2025poster

Human mesh recovery (HMR) from a single image is inherently ill-posed due to depth ambiguity and occlusions. Probabilistic methods have tried to solve this by generating numerous plausible 3D human mesh predictions, but they often exhibit misalignment with 2D image observations and weak robustness t…

2025

BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices

CVPR 2025poster

The emergence and growing popularity of multimodal large language models (MLLMs) have significant potential to enhance various aspects of daily life, from improving communication to facilitating learning and problem-solving. Mobile phones, as essential daily companions, represent the most effective…

2025

CADCrafter: Generating Computer-Aided Design Models from Unconstrained Images

CVPR 2025poster

Creating CAD digital twins from the physical world is crucial for manufacturing, design, and simulation. However, current methods typically rely on costly 3D scanning with labor-intensive post-processing. To provide a user-friendly design process, we explore the problem of reverse engineering from u…

Cited by 3SourcePDFScholar
2025

Can Post-Training Quantization Benefit from an Additional QLoRA Integration?

NAACL 2025industry

Large language models (LLMs) have transformed natural language processing but pose significant challenges for real-world deployment. These models necessitate considerable computing resources, which can be costly and frequently unavailable. Model compression techniques such as quantization are often…

2025

Improved Discretization Complexity Analysis of Consistency Models: Variance Exploding Forward Process and Decay Discretization Scheme

ICML 2025poster

Consistency models, a new class of one-step generative models, have shown competitive performance with multi-step diffusion models. The most challenging part of consistency models is the training process, which discretizes the continuous diffusion process into $K$ steps and trains a one-step mapping…

Cited by 0SourcePDFScholar
2025

LLM Evaluate: An Industry-Focused Evaluation Tool for Large Language Models

COLING 2025industry

Large Language Models (LLMs) have demonstrated impressive capability to solve a wide range of tasks in recent years. This has inspired researchers and practitioners in the real-world industrial domain to build useful products via leveraging LLMs. However, extensive evaluations of LLMs, in terms of a…

2025

Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context Learning

CVPR 2025poster

In-context learning (ICL), a predominant trend in instruction learning, aims at enhancing the performance of large language models by providing clear task guidance and examples, improving their capability in task understanding and execution. This paper investigates ICL on Large Vision-Language Model…

Cited by 0SourcePDFScholar
2025

RRInf: Efficient Influence Function Estimation via Ridge Regression for Large Language Models and Text-to-Image Diffusion Models

EMNLP 2025

The quality of data plays a vital role in the development of Large-scale Generative Models. Understanding how important a data point is for a generative model is essential for explaining its behavior and improving the performance. The influence function provides a framework for quantifying the impac

Cited by 0SourcePDFScholar
2024

CoIN: A Benchmark of Continual Instruction Tuning for Multimodel Large Language Models

NeurIPS 2024poster

Instruction tuning demonstrates impressive performance in adapting Multimodal Large Language Models (MLLMs) to follow task instructions and improve generalization ability. By extending tuning across diverse tasks, MLLMs can further enhance their understanding of world knowledge and instruction inte…

Cited by 14SourcePDFScholar
2024

Enhancing Steganography of Generative Image Based on Image Retouching

ICASSP 2024accepted

Steganography, which hides messages within innocent-looking carriers, is an essential technique to protect data privacy. The rapid advancement of generative models makes AI-generated images a potential steganographic carrier. However, the distortion resulting from the embedding of messages makes it…

Cited by 0SourceScholar
2024

Fast Updating Truncated SVD for Representation Learning with Sparse Matrices

ICLR 2024poster

Updating truncated Singular Value Decomposition (SVD) has extensive applications in representation learning. The continuous evolution of massive-scaled data matrices in practical scenarios highlights the importance of aligning SVD-based models with fast-paced updates. Recent methods for updating tru…

Cited by 2SourcePDFScholar
2024

Few-Shot Diffusion Models Escape the Curse of Dimensionality

NeurIPS 2024poster

While diffusion models have demonstrated impressive performance, there is a growing need for generating samples tailored to specific user-defined concepts. The customized requirements promote the development of few-shot diffusion models, which use limited $n_{ta}$ target samples to fine-tune a pre-t…

Cited by 1SourcePDFScholar
2024

Learn to Optimize Denoising Scores: A Unified and Improved Diffusion Prior for 3D Generation

ECCV 2024poster

"In this paper, we propose a unified framework aimed at enhancing the diffusion priors for 3D generation tasks. Despite the critical importance of these tasks, existing methodologies often struggle to generate high-caliber results. We begin by examining the inherent limitations in previous diffusion…

2024

Query-OPT: Optimizing Inference of Large Language Models via Multi-Query Instructions in Meeting Summarization

EMNLP 2024industry

This work focuses on the task of query-based meeting summarization in which the summary of a context (meeting transcript) is generated in response to a specific query. When using Large Language Models (LLMs) for this task, a new call to the LLM inference endpoint/API is required for each new query e…

2024

Robustness Verification of Deep Reinforcement Learning Based Control Systems Using Reward Martingales

AAAI 2024technical

Deep Reinforcement Learning (DRL) has gained prominence as an effective approach for control systems. However, its practical deployment is impeded by state perturbations that can severely impact system performance. Addressing this critical challenge requires robustness verification about system perf…

Cited by 1SourcePDFScholar
2024

Sculpt3D: Multi-View Consistent Text-to-3D Generation with Sparse 3D Prior

CVPR 2024poster

Recent works on text-to-3d generation show that using only 2D diffusion supervision for 3D generation tends to produce results with inconsistent appearances (e.g. faces on the back view) and inaccurate shapes (e.g. animals with extra legs). Existing methods mainly address this issue by retraining di…

Cited by 18SourcePDFScholar
2024

Tiny Titans: Can Smaller Large Language Models Punch Above Their Weight in the Real World for Meeting Summarization?

NAACL 2024industry

Large Language Models (LLMs) have demonstrated impressive capabilities to solve a wide range of tasks without being explicitly fine-tuned on task-specific datasets. However, deploying LLMs in the real world is not trivial, as it requires substantial computing resources. In this paper, we investigate…

Cited by 23SourcePDFScholar
2024

Zeroth-Order Methods for Constrained Nonconvex Nonsmooth Stochastic Optimization

ICML 2024oral

This paper studies the problem of solving nonconvex nonsmooth optimization over a closed convex set. Most previous works tackle such problems by transforming the constrained problem into an unconstrained problem that can be solved by the techniques developed in the unconstrained setting. However, th…

Cited by 1SourcePDFScholar
2023

A Progressive Neural Network for Acoustic Echo Cancellation

ICASSP 2023accepted

Acoustic echo cancellation is a key issue in hand-free communication systems. In this paper, we proposed a hybrid signal processing and deep echo cancellation method, where a two-stage neural network is designed to remove residual echo progressively. For the personalized acoustic echo cancellation,…

Cited by 0SourceScholar
2023

AI Coach Assist: An Automated Approach for Call Recommendation in Contact Centers for Agent Coaching

ACL 2023industry

In recent years, the utilization of Artificial Intelligence (AI) in the contact center industry is on the rise. One area where AI can have a significant impact is in the coaching of contact center agents. By analyzing call transcripts, AI can quickly determine which calls are most relevant for coach…

Cited by 5SourcePDFScholar
2023

Boosting Verification of Deep Reinforcement Learning via Piece-Wise Linear Decision Neural Networks

NeurIPS 2023poster

Formally verifying deep reinforcement learning (DRL) systems suffers from both inaccurate verification results and limited scalability. The major obstacle lies in the large overestimation introduced inherently during training and then transforming the inexplicable decision-making models, i.e., deep…

Cited by 2SourcePDFScholar
2023

CauSSL: Causality-inspired Semi-supervised Learning for Medical Image Segmentation

ICCV 2023poster

Semi-supervised learning (SSL) has recently demonstrated great success in medical image segmentation, significantly enhancing data efficiency with limited annotations. However, despite its empirical benefits, there are still concerns in the literature about the theoretical foundation and explanation…

Cited by 57PDFcodeScholar
2023

HumanBench: Towards General Human-Centric Perception With Projector Assisted Pretraining

CVPR 2023poster

Human-centric perceptions include a variety of vision tasks, which have widespread industrial applications, including surveillance, autonomous driving, and the metaverse. It is desirable to have a general pretrain model for versatile human-centric downstream tasks. This paper forges ahead along this…

2023

Personalized Speech Enhancement Combining Band-Split RNN and Speaker Attentive Module

ICASSP 2023accepted

Target speaker information can be utilized in speech enhancement (SE) models to more effectively extract the desired speech. Previous works introduce the speaker embedding into speech enhancement models by means of concatenation or affine transformation. In this paper, we propose a speaker attentive…

Cited by 0SourceScholar
2023

Uncertainty Estimation for Safety-critical Scene Segmentation via Fine-grained Reward Maximization

NeurIPS 2023poster

Uncertainty estimation plays an important role for future reliable deployment of deep segmentation models in safety-critical scenarios such as medical applications. However, existing methods for uncertainty estimation have been limited by the lack of explicit guidance for calibrating the prediction…

2022

BLINK with Elasticsearch for Efficient Entity Linking in Business Conversations

NAACL 2022industry

An Entity Linking system aligns the textual mentions of entities in a text to their corresponding entries in a knowledge base. However, deploying a neural entity linking system for efficient real-time inference in production environments is a challenging task. In this work, we present a neural entit…

2022

Developing a Production System for Purpose of Call Detection in Business Phone Conversations

NAACL 2022industry

For agents at a contact centre receiving calls, the most important piece of information is the reason for a given call. An agent cannot provide support on a call if they do not know why a customer is calling. In this paper we describe our implementation of a commercial system to detect Purpose of Ca…

Cited by 6SourcePDFScholar
2022

Entity-level Sentiment Analysis in Contact Center Telephone Conversations

EMNLP 2022industry

Entity-level sentiment analysis predicts the sentiment about entities mentioned in a given text. It is very useful in a business context to understand user emotions towards certain entities, such as products or companies. In this paper, we demonstrate how we developed an entity-level sentiment analy…

Cited by 14SourcePDFScholar
2022

Finding Second-Order Stationary Points in Nonconvex-Strongly-Concave Minimax Optimization

NeurIPS 2022accept

We study the smooth minimax optimization problem $\min_{\bf x}\max_{\bf y} f({\bf x},{\bf y})$, where $f$ is $\ell$-smooth, strongly-concave in ${\bf y}$ but possibly nonconvex in ${\bf x}$. Most of existing works focus on finding the first-order stationary point of the function $f({\bf x},{\bf y})$…

Cited by 33SourcePDFScholar
2022

Inverse is Better! Fast and Accurate Prompt for Few-shot Slot Tagging

ACL 2022findings

Prompting methods recently achieve impressive success in few-shot learning. These methods modify input samples with prompt sentence pieces, and decode label tokens to map samples to corresponding labels. However, such a paradigm is very inefficient for the task of slot tagging. Since slot tagging sa…

2022

Online Active Regression

ICML 2022oral

Active regression considers a linear regression problem where the learner receives a large number of data points but can only observe a small number of labels. Since online algorithms can deal with incremental training data and take advantage of low computational cost, we consider an online extensio…

Cited by 10SourcePDFScholar
2022

Self-Supervised Graph Neural Networks via Diverse and Interactive Message Passing

AAAI 2022technical

By interpreting Graph Neural Networks (GNNs) as the message passing from the spatial perspective, their success is attributed to Laplacian smoothing. However, it also leads to serious over-smoothing issue by stacking many layers. Recently, many efforts have been paid to overcome this issue in semi-s…

Cited by 12SourcePDFScholar
2022

Simultaneously Learning Stochastic and Adversarial Bandits under the Position-Based Model

AAAI 2022technical

Online learning to rank (OLTR) interactively learns to choose lists of items from a large collection based on certain click models that describe users' click behaviors. Most recent works for this problem focus on the stochastic environment where the item attractiveness is assumed to be invariant dur…

Cited by 6SourcePDFScholar
2022

Single-Domain Generalization in Medical Image Segmentation via Test-Time Adaptation from Shape Dictionary

AAAI 2022technical

Domain generalization typically requires data from multiple source domains for model learning. However, such strong assumption may not always hold in practice, especially in medical field where the data sharing is highly concerned and sometimes prohibitive due to privacy issue. This paper studies th…

Cited by 47SourcePDFScholar
2022

UTC: A Unified Transformer With Inter-Task Contrastive Learning for Visual Dialog

CVPR 2022poster

Visual Dialog aims to answer multi-round, interactive questions based on the dialog history and image content. Existing methods either consider answer ranking and generating individually or only weakly capture the relation across the two tasks implicitly by two separate models. The research on a uni…

Cited by 57PDFScholar
2022

bert2BERT: Towards Reusable Pretrained Language Models

ACL 2022long

In recent years, researchers tend to pre-train ever-larger language models to explore the upper limit of deep models. However, large language model pre-training costs intensive computational resources, and most of the models are trained from scratch without reusing the existing pre-trained models, w…

Cited by 86SourcePDFScholar
2021

AutoTinyBERT: Automatic Hyper-parameter Optimization for Efficient Pre-trained Language Models

ACL 2021long

Pre-trained language models (PLMs) have achieved great success in natural language processing. Most of PLMs follow the default setting of architecture hyper-parameters (e.g., the hidden dimension is a quarter of the intermediate dimension in feed-forward sub-networks) in BERT. Few studies have been…

2021

C2C-GenDA: Cluster-to-Cluster Generation for Data Augmentation of Slot Filling

AAAI 2021technical

Slot filling, a fundamental module of spoken language understanding, often suffers from insufficient quantity and diversity of training data. To remedy this, we propose a novel Cluster-to-Cluster generation framework for Data Augmentation (DA), named C2C-GenDA. It enlarges the training set by recons…

2021

FedDG: Federated Domain Generalization on Medical Image Segmentation via Episodic Learning in Continuous Frequency Space

CVPR 2021poster

Federated learning allows distributed medical institutions to collaboratively learn a shared prediction model with privacy protection. While at clinical deployment, the models trained in federated learning can still suffer from performance drop when applied to completely unseen hospitals outside the…

Cited by 586PDFcodeScholar
2021

Revisiting Co-Occurring Directions: Sharper Analysis and Efficient Algorithm for Sparse Matrices

AAAI 2021technical

We study the streaming model for approximate matrix multiplication (AMM). We are interested in the scenario that the algorithm can only take one pass over the data with limited memory. The state-of-the-art deterministic sketching algorithm for streaming AMM is the co-occurring directions (COD), whic…

Cited by 4SourcePDFScholar
2020

Efficient and Robust High-Dimensional Linear Contextual Bandits

IJCAI 2020poster

The linear contextual bandits is a sequential decision-making problem where an agent decides among sequential actions given their corresponding contexts. Since large-scale data sets become more and more common, we study the linear contextual bandits in high-dimensional situations. Recent works focus…

Cited by 0SourcePDFScholar
2020

Hybrid Deep-Semantic Matrix Factorization for Tag-Aware Personalized Recommendation

ICASSP 2020accepted

Matrix factorization has now become a dominant solution for personalized recommendation on the Social Web. To alleviate the cold start problem, previous approaches have incorporated various additional sources of information into traditional matrix factorization models. These upgraded models, however…

Cited by 0SourceScholar
2020

Hybrid fluidic actuation for a foam-based soft actuator

IROS 2020poster

Actuation means for soft robotic structures are manifold: despite actuation mechanisms such as tendon-driven manipulators or shape memory alloys, the majority of soft robotic actuators are fluidically actuated - either purely by positive or negative air pressure or by hydraulic actuation only. This…

Cited by 18SourceScholar
2015

A Real-time relative probabilistic mapping algorithm for high-speed off-road autonomous driving

IROS 2015poster

Reliable mapping and hazard detection are prerequisites for autonomous navigation for unmanned ground vehicles. Because of the uncertainty and vibration induced by high-speed navigation and rugged terrain, the problem of mapping for high-speed off-road autonomous navigation has not been completely s…

Cited by 8SourceScholar