← Search

Xuxin Cheng

59 accepted papers

2026

Breaking the Illusion: When Positive Meets Negative in Multimodal Decoding

CVPR 2026

Vision-Language Models (VLMs) are frequently undermined by object hallucination--generating content that contradicts visual reality--due to an over-reliance on linguistic priors. We introduce Positive-and-Negative Decoding (PND), a training-free inference framework that intervenes directly in the de

Cited by 0SourcecodeScholar
2026

ExBody2: Advanced Expressive Humanoid Whole-Body Control

ICRA 2026poster

This paper tackles the challenge of enabling real-world humanoid robots to perform expressive and dynamic whole-body motions while maintaining stability. We propose ExBody2, a whole-body tracking framework trained in simulation with Reinforcement Learning and then transferred to the real world. The …

2026

Towards Self-Robust LLMs: Intrinsic Prompt Noise Resistance via CoIPO

ICLR 2026poster

Large language models (LLMs) have demonstrated remarkable and steadily improving performance across a wide range of tasks. However, LLM performance may be highly sensitive to prompt variations especially in scenarios with limited openness or strict output formatting requirements, indicating insuffic…

Cited by 0SourcecodeScholar
2025

AMO: Adaptive Motion Optimization for Hyper-Dexterous Humanoid Whole-Body Control

RSS 2025poster

Humanoid robots derive much of their dexterity from hyper-dexterous whole-body movements, enabling tasks that require a large operational workspace—such as picking objects off the ground. However, achieving these capabilities on real humanoids remains challenging due to their high degrees of freedom…

Cited by 0PDFScholar
2025

CountLLM: Towards Generalizable Repetitive Action Counting via Large Language Model

CVPR 2025poster

Repetitive action counting, which aims to count periodic movements in a video, is valuable for video analysis applications such as fitness monitoring. However, existing methods largely rely on regression networks with limited representational capacity, which hampers their ability to accurately captu…

Cited by 1SourcePDFScholar
2025

DisPose: Disentangling Pose Guidance for Controllable Human Image Animation

ICLR 2025poster

Controllable human image animation aims to generate videos from reference images using driving videos. Due to the limited control signals provided by sparse guidance (e.g., skeleton pose), recent works have attempted to introduce additional dense conditions (e.g., depth map) to ensure motion alignme…

2025

EXCGEC: A Benchmark for Edit-Wise Explainable Chinese Grammatical Error Correction

AAAI 2025technical

Existing studies explore the explainability of Grammatical Error Correction (GEC) in a limited scenario, where they ignore the interaction between corrections and explanations and have not established a corresponding comprehensive benchmark. To bridge the gap, this paper first introduces the task of…

2025

Helpful DoggyBot: Open-World Object Fetching using Legged Robots and Vision-Language Models

IROS 2025

Learning-Based methods have achieved strong performance for quadrupedal locomotion. However, several challenges prevent quadrupeds from learning helpful indoor skills that require interaction with environments and humans: lack of end-effectors for manipulation, limited semantic under-standing using

Cited by 15SourceScholar
2025

ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training

CoRL 2025poster

Generative models based on flow matching offer significant potential for learning robot policies, particularly in generating high-dimensional, dexterous behaviors that are conditioned on diverse observations. In this work, we introduce ManiFlow, an advanced flow matching model specifically designed…

Cited by 0SourceScholar
2025

Mobile-TeleVision: Predictive Motion Priors for Humanoid Whole-Body Control

ICRA 2025

Humanoid robots require both robust lower-body locomotion and precise upper-body manipulation. While recent Reinforcement Learning (RL) approaches provide whole-body loco-manipulation policies, they lack precise manipulation with high DoF arms. In this paper, we propose decoupling upper-body control

Cited by 85SourceScholar
2025

UniCoTT: A Unified Framework for Structural Chain-of-Thought Distillation

ICLR 2025poster

Chains of thought (CoTs) have achieved success in enhancing the reasoning capabilities of large language models (LLMs), while their effectiveness is predominantly observed in LLMs. Existing solutions methods adopt distillation to inject chain-of-thought capabilities into small models (SLMs). Howeve…

2024

ACE: A Cross-platform and visual-Exoskeletons System for Low-Cost Dexterous Teleoperation

CoRL 2024poster

Bimanual robotic manipulation with dexterous hands has a large potential workability and a wide workspace as it follows the most natural human workflow. Learning from human demonstrations has proven highly effective for learning a dexterous manipulation policy. To collect such data, teleoperation se…

Cited by 37SourceScholar
2024

Aligner²: Enhancing Joint Multiple Intent Detection and Slot Filling via Adjustive and Forced Cross-Task Alignment

AAAI 2024technical

Multi-intent spoken language understanding (SLU) has garnered growing attention due to its ability to handle multiple intent utterances, which closely mirrors practical scenarios. Unlike traditional SLU, each intent in multi-intent SLU corresponds to its designated scope for slots, which occurs in…

2024

Alignment before Awareness: Towards Visual Question Localized-Answering in Robotic Surgery via Optimal Transport and Answer Semantics

COLING 2024main

The visual question localized-answering (VQLA) system has garnered increasing attention due to its potential as a knowledgeable assistant in surgical education. Apart from providing text-based answers, VQLA can also pinpoint the specific region of interest for better surgical scene understanding. Al…

2024

Code-Switching Can be Better Aligners: Advancing Cross-Lingual SLU through Representation-Level and Prediction-Level Alignment

ACL 2024short

Zero-shot cross-lingual spoken language understanding (SLU) can promote the globalization application of dialog systems, which has attracted increasing attention. While current code-switching based cross-lingual SLU frameworks have shown promising results, they (i) predominantly utilize contrastive…

2024

Cyclical Contrastive Learning Based on Geodesic for Zero-shot Cross-lingual Spoken Language Understanding

ACL 2024findings

Owing to the scarcity of labeled training data, Spoken Language Understanding (SLU) is still a challenging task in low-resource languages. Therefore, zero-shot cross-lingual SLU attracts more and more attention. Contrastive learning is widely applied to explicitly align representations of similar se…

2024

Embracing Language Inclusivity and Diversity in CLIP through Continual Language Learning

AAAI 2024technical

While vision-language pre-trained models (VL-PTMs) have advanced multimodal research in recent years, their mastery in a few languages like English restricts their applicability in broader communities. To this end, there is an increasing interest in developing multilingual VL models via a joint-lear…

2024

Enhancing Dialogue State Tracking Models through LLM-backed User-Agents Simulation

ACL 2024long

Dialogue State Tracking (DST) is designed to monitor the evolving dialogue state in the conversations and plays a pivotal role in developing task-oriented dialogue systems. However, obtaining the annotated data for the DST task is usually a costly endeavor. In this paper, we focus on employing LLMs…

2024

Exploiting Auxiliary Caption for Video Grounding

AAAI 2024technical

Video grounding aims to locate a moment of interest matching the given query sentence from an untrimmed video. Previous works ignore the sparsity dilemma in video annotations, which fails to provide the context information between potential events and query sentences in the dataset. In this paper, w…

Cited by 17SourcePDFScholar
2024

Expressive Whole-Body Control for Humanoid Robots

RSS 2024poster

Can we enable humanoid robots to generate rich, diverse, and expressive motions in the real world? We propose to learn a whole-body control policy on a human-sized robot to mimic human motions as realistic as possible. To train such a policy, we leverage the large-scale human motion capture data fro…

Cited by 94SourcePDFScholar
2024

KC-Prompt: End-To-End Knowledge-Complementary Prompting for Rehearsal-Free Continual Learning

ICASSP 2024accepted

Continuous learning requires adapting quickly to incoming tasks while avoiding catastrophic forgetting. Typical solutions resort to a rehearsal buffer to replay old data, which is intractable to apply in real-world scenarios with limited memory and inaccessible privacy. Recently, with the emergence…

Cited by 0SourceScholar
2024

KDProR: A Knowledge-Decoupling Probabilistic Framework for Video-Text Retrieval

ECCV 2024poster

"Existing video-text retrieval methods predominantly focus on designing diverse cross-modal interaction mechanisms between captions and videos. However, those approaches diverge from human learning paradigms, where humans possess the capability to seek and associate knowledge from an open set, rathe…

Cited by 8SourcePDFScholar
2024

Knowledge-enhanced Prompt Tuning for Dialogue-based Relation Extraction with Trigger and Label Semantic

COLING 2024main

Dialogue-based relation extraction (DRE) aims to determine the semantic relation of a given pair of arguments from a piece of dialogue, which has received increasing attention. Due to the low information density of dialogue text, it is difficult for the model to focus on key information. To this end…

2024

Learning to Match Representations is Better for End-to-End Task-Oriented Dialog System

EMNLP 2024finding

Due to the rapid development with pre-trained language models, fully end-to-end Task-Oriented Dialogue (TOD) systems exhibit superior performance. How to achieve the ability to efficiently retrieve entities in cross-domain large-scale databases is a key issue. Most existing end-to-end Task-Oriented…

Cited by 0SourcePDFScholar
2024

MaCSC: Towards Multimodal-augmented Pre-trained Language Models via Conceptual Prototypes and Self-balancing Calibration

NAACL 2024long

Pre-trained language models (PLMs) that rely solely on textual data may exhibit limitations in multimodal semantics comprehension. Existing solutions attempt to alleviate this issue by incorporating explicit image retrieval or generation techniques.However, these methods: (1) focus exclusively on th…

2024

MoE-SLU: Towards ASR-Robust Spoken Language Understanding via Mixture-of-Experts

ACL 2024findings

As a crucial task in the task-oriented dialogue systems, spoken language understanding (SLU) has garnered increasing attention. However, errors from automatic speech recognition (ASR) often hinder the performance of understanding. To tackle this problem, we propose MoE-SLU, an ASR-Robust SLU framewo…

Cited by 2SourcePDFScholar
2024

Open-TeleVision: Teleoperation with Immersive Active Visual Feedback

CoRL 2024poster

Teleoperation serves as a powerful method for collecting on-robot data essential for robot learning from demonstrations. The intuitiveness and ease of use of the teleoperation system are crucial for ensuring high-quality, diverse, and scalable data. To achieve this, we propose an immersive teleopera…

Cited by 99SourceScholar
2024

PCAD: Towards ASR-Robust Spoken Language Understanding via Prototype Calibration and Asymmetric Decoupling

ACL 2024long

Spoken language understanding (SLU) inevitably suffers from error propagation from automatic speech recognition (ASR) in actual scenarios. Some recent works attempt to alleviate this issue through contrastive learning. However, they (1) sample negative pairs incorrectly in pre-training; (2) only foc…

2024

PolyVoice: Language Models for Speech to Speech Translation

ICLR 2024poster

With the huge success of GPT models in natural language processing, there is a growing interest in applying language modeling approaches to speech tasks. Currently, the dominant architecture in speech-to-speech translation (S2ST) remains the encoder-decoder paradigm, creating a need to investigate t…

2024

RAG-HAT: A Hallucination-Aware Tuning Pipeline for LLM in Retrieval-Augmented Generation

EMNLP 2024industry

Retrieval-augmented generation (RAG) has emerged as a significant advancement in the field of large language models (LLMs). By integrating up-to-date information not available during their initial training, RAG greatly enhances the practical utility of LLMs in real-world applications. However, even…

Cited by 5SourcePDFScholar
2024

Soul-Mix: Enhancing Multimodal Machine Translation with Manifold Mixup

ACL 2024long

Multimodal machine translation (MMT) aims to improve the performance of machine translation with the help of visual information, which has received widespread attention recently. It has been verified that visual information brings greater performance gains when the textual information is limited. Ho…

2024

Towards Explainable Joint Models via Information Theory for Multiple Intent Detection and Slot Filling

AAAI 2024technical

Recent joint models for multi-intent detection and slot filling have obtained promising results through modeling the unidirectional or bidirectional guidance between intent and slot. However, existing works design joint models heuristically and lack some theoretical exploration, including (1) theore…

2024

Towards Multi-Intent Spoken Language Understanding via Hierarchical Attention and Optimal Transport

AAAI 2024technical

Multi-Intent spoken language understanding (SLU) can handle complicated utterances expressing multiple intents, which has attracted increasing attention from researchers. Although existing models have achieved promising performance, most of them still suffer from two leading problems: (1) each inten…

2024

Towards Multi-modal Sarcasm Detection via Disentangled Multi-grained Multi-modal Distilling

COLING 2024main

Multi-modal sarcasm detection aims to identify whether a given sample with multi-modal information (i.e., text and image) is sarcastic, which has received increasing attention due to the rapid growth of multi-modal posts on modern social media. However, mainstream models process the input of each mo…

2024

Uncertainty-aware sign language video retrieval with probability distribution modeling

ECCV 2024poster

"Sign language video retrieval plays a key role in facilitating information access for the deaf community. Despite significant advances in video-text retrieval, the complexity and inherent uncertainty of sign language preclude direct applications of these techniques. Previous methods achieve mapping…

2024

Visual Whole-Body Control for Legged Loco-Manipulation

CoRL 2024poster

We study the problem of mobile manipulation using legged robots equipped with an arm, namely legged loco-manipulation. The robot legs, while usually utilized for mobility, offer an opportunity to amplify the manipulation capabilities by conducting whole-body control. That is, the robot can control t…

Cited by 47SourceScholar
2024

What are the Generator Preferences for End-to-end Task-Oriented Dialog System?

EMNLP 2024main

Fully end-to-end task-oriented dialogue (EToD) systems have shown excellent performance, which requires the ability to retrieve entities accurately for generation. Existing methods improve the accuracy of entity retrieval and construct data flows between retrieval results and response generator, ach…

Cited by 0SourcePDFScholar
2024

Zero-Shot Spoken Language Understanding via Large Language Models: A Preliminary Study

COLING 2024main

Zero-shot Spoken Language Understanding (SLU) aims to enable task-oriented dialogue systems to understand user needs without training data. Challenging but worthwhile, zero-shot SLU reduces the time and effort that data labeling takes. Recent advancements in large language models (LLMs), such as GPT…

Cited by 14SourcePDFScholar
2023

A Dynamic Graph Interactive Framework with Label-Semantic Injection for Spoken Language Understanding

ICASSP 2023accepted

Multi-intent detection and slot filling joint models are gaining increasing traction since they are closer to complicated real-world scenarios. However, existing approaches (1) focus on identifying implicit correlations between utterances and one-hot encoded labels in both tasks while ignoring expli…

Cited by 0SourceScholar
2023

Accelerating Multiple Intent Detection and Slot Filling via Targeted Knowledge Distillation

EMNLP 2023long findings

Recent non-autoregressive Spoken Language Understanding (SLU) models attracts increasing attention owing to the high inference speed. However, most of them still (1) suffer from the multi-modality problem since the prior knowledge about the reference is relatively poor during inference; (2) fail to…

Cited by 0SourceScholar
2023

Discover and Align Taxonomic Context Priors for Open-world Semi-Supervised Learning

NeurIPS 2023poster

Open-world Semi-Supervised Learning (OSSL) is a realistic and challenging task, aiming to classify unlabeled samples from both seen and novel classes using partially labeled samples from the seen classes. Previous works typically explore the relationship of samples as priors on the pre-defined sing…

2023

Enhancing Code-Switching for Cross-lingual SLU: A Unified View of Semantic and Grammatical Coherence

EMNLP 2023short main

Despite the success of spoken language understanding (SLU) in high-resource languages, achieving similar performance in low-resource settings, such as zero-shot scenarios, remains challenging due to limited labeled training data. To improve zero-shot cross-lingual SLU, recent studies have explored c…

Cited by 0SourceScholar
2023

G2L: Semantically Aligned and Uniform Video Grounding via Geodesic and Game Theory

ICCV 2023oral

The recent video grounding works attempt to introduce vanilla contrastive learning into video grounding. However, we claim that this naive solution is suboptimal. Contrastive learning requires two key properties: (1) alignment of features of similar samples, and (2) uniformity of the induced distrib…

Cited by 52PDFScholar
2023

M3ST: Mix at Three Levels for Speech Translation

ICASSP 2023accepted

How to solve the data scarcity problem for end-to-end speech-to-text translation (ST)? It’s well known that data augmentation is an efficient method to improve performance for many tasks by enlarging the dataset. In this paper, we propose Mix at three levels for Speech Translation (M <sup xmlns:mml=…

Cited by 0SourceScholar
2023

MCLF: A Multi-grained Contrastive Learning Framework for ASR-robust Spoken Language Understanding

EMNLP 2023long findings

Enhancing the robustness towards Automatic Speech Recognition (ASR) errors is of great importance for Spoken Language Understanding (SLU). Trending ASR-robust SLU systems have witnessed impressive improvements through global contrastive learning. However, although most ASR errors occur only at local…

Cited by 0SourceScholar
2023

ML-LMCL: Mutual Learning and Large-Margin Contrastive Learning for Improving ASR Robustness in Spoken Language Understanding

ACL 2023findings

Spoken language understanding (SLU) is a fundamental task in the task-oriented dialogue systems. However, the inevitable errors from automatic speech recognition (ASR) usually impair the understanding performance and lead to error propagation. Although there are some attempts to address this problem…

2023

MRRL: Modifying the Reference via Reinforcement Learning for Non-Autoregressive Joint Multiple Intent Detection and Slot Filling

EMNLP 2023long findings

With the rise of non-autoregressive approach, some non-autoregressive models for joint multiple intent detection and slot filling have obtained the promising inference speed. However, most existing SLU models (1) suffer from the multi-modality problem that leads to reference intents and slots may no…

Cited by 0SourceScholar
2023

SSVMR: Saliency-Based Self-Training for Video-Music Retrieval

ICASSP 2023accepted

With the rise of short videos, the demand for selecting appropriate background music (BGM) for a video has increased significantly, video-music retrieval (VMR) task gradually draws much attention by research community. As other cross-modal learning tasks, existing VMR approaches usually attempt to m…

Cited by 0SourceScholar
2023

Syntax Matters: Towards Spoken Language Understanding via Syntax-Aware Attention

EMNLP 2023short findings

Spoken Language Understanding (SLU), a crucial component of task-oriented dialogue systems, has consistently garnered attention from both academic and industrial communities. Although incorporating syntactic information into models has the potential to enhance the comprehension of user utterances an…

Cited by 0SourceScholar
2023

Towards Unified Spoken Language Understanding Decoding via Label-aware Compact Linguistics Representations

ACL 2023findings

Joint intent detection and slot filling models have shown promising success in recent years due to the high correlations between the two tasks. However, previous works independently decode the two tasks, which could result in misaligned predictions for both tasks. To address this shortcoming, we pro…

2023

Unify, Align and Refine: Multi-Level Semantic Alignment for Radiology Report Generation

ICCV 2023poster

Automatic radiology report generation has attracted enormous research interest due to its practical value in reducing the workload of radiologists. However, simultaneously establishing global correspondences between the image (e.g., Chest X-ray) and its related report and local alignments between im…

Cited by 42PDFScholar
2022

Deep Whole-Body Control: Learning a Unified Policy for Manipulation and Locomotion

CoRL 2022oral

An attached arm can significantly increase the applicability of legged robots to several mobile manipulation tasks that are not possible for the wheeled or tracked counterparts. The standard modular control pipeline for such legged manipulators is to decouple the controller into that of manipulation…

Cited by 165SourcecodeScholar
2021

Reinforcement Learning for Robust Parameterized Locomotion Control of Bipedal Robots

ICRA 2021poster

Developing robust walking controllers for bipedal robots is a challenging endeavor. Traditional model-based locomotion controllers require simplifying assumptions and careful modelling; any small errors can result in unstable control. To address these challenges for bipedal locomotion, we present a…

Cited by 287SourceScholar