← Search

Chengyuan Ma

11 accepted papers

2026

ReIn: Conversational Error Recovery with Reasoning Inception

ICLR 2026poster

Conversational agents powered by large language models (LLMs) with tool integration achieve strong performance on fixed task-oriented dialogue datasets but remain vulnerable to unanticipated, user-induced errors. Rather than focusing on error prevention, this work focuses on error recovery, which ne…

Cited by 0SourcecodeScholar
2026

TLDIFFGAN: A LATENT DIFFUSION-GAN FRAMEWORK WITH TEMPORAL INFORMATION FUSION FOR ANOMALOUS SOUND DETECTION

ICASSP 2026poster

Existing generative models for unsupervised anomalous sound detection are limited by their inability to fully capture the complex feature distribution of normal sounds, while the potential of powerful diffusion models in this domain remains largely unexplored. To address this challenge, we propose a…

Cited by 0SourcePDFScholar
2026

VIVIDVOICE: A UNIFIED FRAMEWORK FOR SCENE-AWARE VISUALLY-DRIVEN SPEECH SYNTHESIS

ICASSP 2026poster

We introduce and define a novel task-Scene-Aware Visually-Driven Speech Synthesis, aimed at addressing the limitations of existing speech generation models in creating immersive auditory experiences that align with the real physical world. To tackle the two core challenges of data scarcity and modal…

Cited by 0SourcePDFScholar
2025

SLIM: Subtrajectory-Level Elimination for More Effective Reasoning

EMNLP 2025

In recent months, substantial progress has been made in complex reasoning of Large Language Models (LLMs), particularly through the application of test-time scaling. Notable examples include, though are not limited to, OpenAI’s o1/o3/o4 series and DeepSeek-R1. When responding to a query, these model

Cited by 0SourcePDFScholar
2024

Demonstrating Language-Grounded Motion Controller

RSS 2024poster

Recent advancements have enabled human-robot collaboration through physical assistance and verbal guidance. However, limitations persist in coordinating robots' physical motions and speech in response to real-time changes in human behavior during collaborative contact tasks. We first derive principl…

Cited by 0SourcePDFScholar
2023

An Avatar Robot Overlaid with the 3D Human Model of a Remote Operator

IROS 2023poster

Although telepresence assistive robots have made significant progress, they still lack the sense of realism and physical presence of the remote operator. This results in a lack of trust and adoption of such robots. In this paper, we introduce an Avatar Robot System which is a mixed real/virtual robo…

Cited by 5SourceScholar
2023

Clicker: Attention-Based Cross-Lingual Commonsense Knowledge Transfer

ICASSP 2023accepted

Recent advances in cross-lingual commonsense reasoning (CSR) are facilitated by the development of multilingual pre-trained models (mPTMs). While mPTMs show the potential to encode commonsense knowledge for different languages, transferring commonsense knowledge learned in large-scale English corpus…

Cited by 0SourceScholar
2023

KEPLET: Knowledge-Enhanced Pretrained Language Model with Topic Entity Awareness

EMNLP 2023long findings

In recent years, Pre-trained Language Models (PLMs) have shown their superiority by pre-training on unstructured text corpus and then fine-tuning on downstream tasks. On entity-rich textual resources like Wikipedia, Knowledge-Enhanced PLMs (KEPLMs) incorporate the interactions between tokens and men…

Cited by 0SourceScholar
2022

Incremental User Embedding Modeling for Personalized Text Classification

ICASSP 2022accepted

Individual user profiles and interaction histories play a significant role in providing customized experiences in real-world applications such as chatbots, social media, retail, and education. Adaptive user representation learning by utilizing user personalized information has be-come increasingly c…

Cited by 0SourceScholar
2022

Self-Aware Feedback-Based Self-Learning in Large-Scale Conversational AI

NAACL 2022industry

Self-learning paradigms in large-scale conversational AI agents tend to leverage user feedback in bridging between what they say and what they mean. However, such learning, particularly in Markov-based query rewriting systems have far from addressed the impact of these models on future training wher…

Cited by 3SourcePDFScholar
2018

Combining Acoustic Embeddings and Decoding Features for End-of-Utterance Detection in Real-Time Far-Field Speech Recognition Systems

ICASSP 2018accepted

We present an end-of-utterance detector for real-time automatic speech recognition in far-field scenarios. The proposed system consists of three components: a long short-term memory (LSTM) neural network trained on acoustic features, an LSTM trained on l-best recognition hypotheses of the automatic…

Cited by 0SourceScholar