← Search

Qian Chen

69 accepted papers

2026

Beyond Entity Correlations: Disentangling Event Causal Puzzles in Temporal Knowledge Graphs

ICLR 2026poster

Existing Temporal Knowledge Graph (TKG) representation learning approaches focus on modeling entity correlations. However, since TKG datasets are constructed from events, which inherently contain heterogeneous causalities, focusing solely on entity or relation level correlations is inadequate for ev…

Cited by 0SourceScholar
2026

DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations

ICLR 2026poster

Recent studies on end-to-end (E2E) speech generation with large language models (LLMs) have attracted significant community attention, with multiple works extending text-based LLMs to generate discrete speech tokens. Existing E2E approaches primarily fall into two categories: (1) Methods that genera…

Cited by 0SourceScholar
2026

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs

ICML 2026poster

Although Multimodal Large Language Models (MLLMs) demonstrate strong omni-modal perception, their ability to forecast future events from audio-visual cues remains largely unexplored, as existing benchmarks focus mainly on retrospective understanding. To bridge this gap, we introduce FutureOmni, the …

Cited by 0SourceScholar
2026

Look on Demand: A Cognitive Scheduling Framework for Visual Evidence Acquisition in Multimodal Reasoning

ICML 2026poster

Existing multimodal reasoning approaches predominantly follow two paradigms: converting visual inputs into text prior to reasoning, or performing end-to-end reasoning within a unified vision–language representation space. Despite their empirical progress, both paradigms suffer from fundamental struc…

Cited by 0SourceScholar
2026

MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models

ICLR 2026poster

Recent advances in multimodal large language models (MLLMs) have catalyzed transformative progress in affective computing, enabling models to exhibit emergent emotional intelligence. Despite substantial methodological progress, current emotional benchmarks remain limited, as it is still unknown: (a)…

Cited by 0SourcecodeScholar
2026

PrismAudio: Decomposed Chain-of-Thought and Multi-dimensional Rewards for Video-to-Audio Generation

ICLR 2026poster

Video-to-Audio (V2A) generation requires balancing four critical perceptual dimensions: semantic consistency, audio-visual temporal synchrony, aesthetic quality, and spatial accuracy; yet existing methods suffer from objective entanglement that conflates competing goals in single loss functions and…

Cited by 0SourcecodeScholar
2026

STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation

ICML 2026poster

Continuous Variational Autoencoders (VAEs) serve as the fundamental continuous tokenizer for modern neural audio generation systems, enabling high-fidelity reconstruction while providing a compact, smooth latent space for downstream generative priors. However, continuous VAEs face a fundamental conf…

Cited by 0SourcecodeScholar
2026

Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration Coding

AAAI 2026technical

Existing speech tokenizers typically assign a fixed number of tokens per second, regardless of the varying information density or temporal fluctuations in the speech signal. This uniform token allocation mismatches the intrinsic structure of speech, where information is distributed unevenly over tim

Cited by 0SourcePDFScholar
2026

When Simple Problems Wear Complex Costumes: Improving Efficiency in LRM’s Adaptive Reasoning

ICML 2026poster

Recent Large Reasoning Models (LRMs) have demonstrated powerful multi-step problem-solving capabilities but often suffer from inefficiency due to an ``overthinking phenomenon", where they apply complex reasoning to simple tasks, resulting in unnecessary computational cost and latency. While adaptive…

Cited by 0SourceScholar
2025

3D-Speaker-Toolkit: An Open-Source Toolkit for Multimodal Speaker Verification and Diarization

ICASSP 2025accepted

We introduce 3D-Speaker-Toolkit, an open-source toolkit for multimodal speaker verification and diarization, designed for meeting the needs of academic researchers and industrial practitioners. The 3D-Speaker-Toolkit adeptly leverages the combined strengths of acoustic, semantic, and visual data, se…

Cited by 0SourceScholar
2025

ClinicalLab: Aligning Agents for Multi-Departmental Clinical Diagnostics in the Real World

NeurIPS 2025poster

Large language models (LLMs) have achieved significant performance progress in various natural language processing applications. However, LLMs still struggle to meet the strict requirements for accuracy and reliability in the medical field and face many challenges in clinical applications. Existing…

Cited by 0SourcecodeScholar
2025

CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification

AAAI 2025technical

Large Language Models (LLMs) have made significant progress in code generation, offering developers groundbreaking automated programming support. However, LLMs often generate code that is syntactically correct and even semantically plausible, but may not execute as expected or fulfill specified requ…

2025

ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control

ACL 2025long

In this paper, we present ControlSpeech, a text-to-speech (TTS) system capable of fully cloning the speaker’s voice and enabling arbitrary control and adjustment of speaking style. Prior zero-shot TTS models only mimic the speaker’s voice without further control and adjustment capabilities while pri…

2025

Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning

ACL 2025long

Our quality audit for three widely used public multilingual speech datasets Mozilla Common Voice 17.0, FLEURS, and VoxPopuli shows that in some languages, these datasets suffer from significant quality issues. We believe addressing these issues will make these datasets more useful as evaluation sets…

2025

Distilled Decoding 2: One-step Sampling of Image Auto-regressive Models with Conditional Score Distillation

NeurIPS 2025poster

Image Auto-regressive (AR) models have emerged as a powerful paradigm of visual generative models. Despite their promising performance, they suffer from slow generation speed due to the large number of sampling steps required. Although Distilled Decoding 1 (DD1) was recently proposed to enable few-s…

Cited by 0SourcecodeScholar
2025

Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization on Multi-party Conversation

ACL 2025long

Speaker diarization aims to segment an audio stream into homogeneous partitions based on speaker identity, playing a crucial role in speech comprehension and analysis. Mainstream speaker diarization systems rely only on acoustic information, making the task particularly challenging in complex acoust…

2025

LED-Merging: Mitigating Safety-Utility Conflicts in Model Merging with Location-Election-Disjoint

ACL 2025long

Fine-tuning pre-trained Large Language Models (LLMs) for specialized tasks incurs substantial computational and data costs. While model merging offers a training-free solution to integrate multiple task-specific models, existing methods suffer from safety-utility conflicts where enhanced general cap…

2025

Mitigating Pervasive Modality Absence Through Multimodal Generalization and Refinement

AAAI 2025technical

The performance of multimodal models often deteriorates when modality absence occurs. The absence disrupts the learned inter-modal correlations, resulting in biased multimodal representations. This challenge is especially pronounced when the absence is pervasive, affecting both the training and infe…

Cited by 0SourcePDFScholar
2025

Multimodal Fusion and Coherence Modeling for Video Topic Segmentation

ACL 2025finding

The video topic segmentation (VTS) task segments videos into intelligible, non-overlapping topics, facilitating efficient comprehension of video content and quick access to specific content. VTS is also critical to various downstream video understanding tasks. Traditional VTS methods using shallow f…

2025

OmniAudio: Generating Spatial Audio from 360-Degree Video

ICML 2025poster

Traditional video-to-audio generation techniques primarily focus on perspective video and non-spatial audio, often missing the spatial cues necessary for accurately representing sound sources in 3D environments. To address this limitation, we introduce a novel task, \textbf{360V2SA}, to generate spa…

2025

OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation

ACL 2025long

Full-duplex spoken dialogue systems significantly surpass traditional turn-based dialogue systems, as they allow simultaneous bidirectional communication, closely mirroring human-human interactions. However, achieving low latency and natural interactions in full-duplex dialogue systems remains a sig…

2025

QuadEnhancer: Leveraging Quadratic Transformations to Enhance Deep Neural Networks

NeurIPS 2025poster

The combination of linear transformations and nonlinear activation functions forms the foundation of most modern deep neural networks, enabling them to approximate highly complex functions. This paper explores the introduction of quadratic transformations to further increase the nonlinearity of the…

Cited by 0SourceScholar
2025

Recording for Eyes, Not Echoing to Ears: Contextualized Spoken-to-Written Conversion of ASR Transcripts

AAAI 2025technical

Automatic Speech Recognition (ASR) transcripts exhibit recognition errors and various spoken language phenomena such as disfluencies, ungrammatical sentences, and incomplete sentences, hence suffering from poor readability. To improve readability, we propose a Contextualized Spoken-to-Written conver…

2025

SURE: Mutually Visible Objects and Self-generated Candidate Labels For Relation Extraction

COLING 2025main

Joint relation extraction models effectively mitigate the error propagation problem inherently present in pipeline models. Nevertheless, joint models face challenges including high computational complexity, complex network architectures, difficult parameter tuning, and notably, limited interpretabil…

2025

Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision

ICASSP 2025accepted

Training speaker-discriminative and robust speaker verification systems without explicit speaker labels remains a persistent challenge. In this paper, we propose a novel self-supervised speaker verification approach, Self-Distillation Prototypes Network (SDPN), which effectively facilitates self-sup…

Cited by 0SourceScholar
2025

Speech Recognition Meets Large Language Model: Benchmarking, Models, and Exploration

AAAI 2025technical

In this paper, we focus on prompting one of the most important tasks in the field of speech processing, i.e., automatic speech recognition (ASR), with speech foundation encoders and large language models (LLM). Despite the growing body of research in this area, we find that many crucial design decis…

2025

ThinkSound: Chain-of-Thought Reasoning in Multimodal LLMs for Audio Generation and Editing

NeurIPS 2025poster

While end-to-end video-to-audio generation has greatly improved, producing high-fidelity audio that authentically captures the nuances of visual content remains challenging. Like professionals in the creative industries, this generation requires sophisticated reasoning about items such as visual dyn…

Cited by 0SourcecodeScholar
2025

UniCodec: Unified Audio Codec with Single Domain-Adaptive Codebook

ACL 2025long

The emergence of audio language models is empowered by neural audio codecs, which establish critical mappings between continuous waveforms and discrete tokens compatible with language model paradigms. The evolutionary trends from multi-layer residual vector quantizer to single-layer quantizer are be…

2025

V2C-CBM: Building Concept Bottlenecks with Vision-to-Concept Tokenizer

AAAI 2025technical

Concept Bottleneck Models (CBMs) offer inherent interpretability by initially translating images into human-comprehensible concepts, followed by a linear combination of these concepts for classification. However, the annotation of concepts for visual recognition tasks requires extensive expert knowl…

2025

WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

ICLR 2025poster

Language models have been effectively applied to modeling natural signals, such as images, video, speech, and audio. A crucial component of these models is the codec tokenizer, which compresses high-dimensional natural signals into lower-dimensional discrete tokens. In this paper, we introduce WavTo…

2025

When GNNs meet symmetry in ILPs: an orbit-based feature augmentation approach

ICLR 2025poster

A common characteristic in integer linear programs (ILPs) is symmetry, allowing variables to be permuted without altering the underlying problem structure. Recently, GNNs have emerged as a promising approach for solving ILPs. However, a significant challenge arises when applying GNNs to ILPs with s…

2024

CIDR: A Cooperative Integrated Dynamic Refining Method for Minimal Feature Removal Problem

AAAI 2024technical

The minimal feature removal problem in the post-hoc explanation area aims to identify the minimal feature set (MFS). Prior studies using the greedy algorithm to calculate the minimal feature set lack the exploration of feature interactions under a monotonic assumption which cannot be satisfied in ge…

2024

CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation

ACL 2024long

Large Language Models (LLMs) have demonstrated remarkable performance on assisting humans in programming and facilitating programming automation. However, existing benchmarks for evaluating the code understanding and generation capacities of LLMs suffer from severe limitations. First, most benchmark…

2024

DECRL: A Deep Evolutionary Clustering Jointed Temporal Knowledge Graph Representation Learning Approach

NeurIPS 2024poster

Temporal Knowledge Graph (TKG) representation learning aims to map temporal evolving entities and relations to embedded representations in a continuous low-dimensional vector space. However, existing approaches cannot capture the temporal evolution of high-order correlations in TKGs. To this end, we…

Cited by 0SourcePDFScholar
2024

Leveraging Speech PTM, Text LLM, And Emotional TTS For Speech Emotion Recognition

ICASSP 2024accepted

In this paper, we explored how to boost speech emotion recognition (SER) with the state-of-the-art speech pre-trained model (PTM), data2vec, text generation technique, GPT-4, and speech synthesis technique, Azure TTS. First, we investigated the representation ability of different speech self-supervi…

Cited by 0SourceScholar
2024

Loss Masking Is Not Needed In Decoder-Only Transformer For Discrete-Token-Based ASR

ICASSP 2024accepted

Recently, unified speech-text models, such as SpeechGPT, VioLA, and AudioPaLM, have achieved remarkable performance on various speech tasks. These models discretize speech signals into tokens (speech discretization) and use a shared vocabulary for both text and speech tokens. Then they train a singl…

Cited by 0SourceScholar
2024

PDHG-Unrolled Learning-to-Optimize Method for Large-Scale Linear Programming

ICML 2024poster

Solving large-scale linear programming (LP) problems is an important task in various areas such as communication networks, power systems, finance and logistics. Recently, two distinct approaches have emerged to expedite LP solving: (i) First-order methods (FOMs); (ii) Learning to optimize (L2O). In…

2024

PE: A Poincare Explanation Method for Fast Text Hierarchy Generation

EMNLP 2024finding

The black-box nature of deep learning models in NLP hinders their widespread application. The research focus has shifted to Hierarchical Attribution (HA) for its ability to model feature interactions. Recent works model non-contiguous combinations with a time-costly greedy search in Eculidean spaces…

2024

SymILO: A Symmetry-Aware Learning Framework for Integer Linear Optimization

NeurIPS 2024poster

Integer linear programs (ILPs) are commonly employed to model diverse practical problems such as scheduling and planning. Recently, machine learning techniques have been utilized to solve ILPs. A straightforward idea is to train a model via supervised learning, with an ILP as the input and an opti…

2024

Thermal3D-GS: Physics-induced 3D Gaussians for Thermal Infrared Novel-view Synthesis

ECCV 2024poster

"Novel-view synthesis based on visible light has been extensively studied. In comparison to visible light imaging, thermal infrared imaging offers the advantage of all-weather imaging and strong penetration, providing increased possibilities for reconstruction in nighttime and adverse weather scenar…

2024

TruthReader: Towards Trustworthy Document Assistant Chatbot with Reliable Attribution

EMNLP 2024system demonstrations

Document assistant chatbots are empowered with extensive capabilities by Large Language Models (LLMs) and have exhibited significant advancements. However, these systems may suffer from hallucinations that are difficult to verify in the context of given documents.Moreover, despite the emergence of p…

2023

A GNN-Guided Predict-and-Search Framework for Mixed-Integer Linear Programming

ICLR 2023poster

Mixed-integer linear programming (MILP) is widely employed for modeling combinatorial optimization problems. In practice, similar MILP instances with only coefficient variations are routinely solved, and machine learning (ML) algorithms are capable of capturing common patterns across these MILP inst…

2023

Adaptive Knowledge Distillation Between Text and Speech Pre-Trained Models

ICASSP 2023accepted

Learning on a massive amount of speech corpus leads to the recent success of many self-supervised speech models. With knowledge distillation, these models may also benefit from the knowledge encoded by language models that are pre-trained on rich sources of texts. The distillation process, however,…

Cited by 0SourceScholar
2023

Auxiliary Pooling Layer For Spoken Language Understanding

ICASSP 2023accepted

End-to-end spoken language understanding requires speech data annotated with semantic information and may suffer from the shortage of annotated data. Recent progresses leverage unlabelled speech data to pre-train a speech encoder. However, it remains a challenge for the pre-trained speech encoder to…

Cited by 0SourceScholar
2023

CodeTransOcean: A Comprehensive Multilingual Benchmark for Code Translation

EMNLP 2023long findings

Recent code translation techniques exploit neural machine translation models to translate source code from one programming language to another to satisfy production compatibility or to improve efficiency of codebase maintenance. Most existing code translation datasets only focus on a single pair of…

Cited by 0SourcecodeScholar
2023

Ditto: A Simple and Efficient Approach to Improve Sentence Embeddings

EMNLP 2023short main

Prior studies diagnose the anisotropy problem in sentence representations from pre-trained language models, e.g., BERT, without fine-tuning. Our analysis reveals that the sentence embeddings from BERT suffer from a bias towards uninformative words, limiting the performance in semantic textual simila…

Cited by 0SourcecodeScholar
2023

DopplerBAS: Binaural Audio Synthesis Addressing Doppler Effect

ACL 2023findings

Recently, binaural audio synthesis (BAS) has emerged as a promising research field for its applications in augmented and virtual realities. Binaural audio helps ususers orient themselves and establish immersion by providing the brain with interaural time differences reflecting spatial information. H…

Cited by 1SourcePDFScholar
2023

Exploring Speaker-Related Information in Spoken Language Understanding for Better Speaker Diarization

ACL 2023findings

Speaker diarization is a classic task in speech processing and is crucial in multi-party scenarios such as meetings and conversations. Current mainstream speaker diarization approaches consider acoustic information only, which result in performance degradation when encountering adverse acoustic envi…

Cited by 7SourcePDFScholar
2023

Improving Long Document Topic Segmentation Models With Enhanced Coherence Modeling

EMNLP 2023long main

Topic segmentation is critical for obtaining structured documents and improving down- stream tasks such as information retrieval. Due to its ability of automatically exploring clues of topic shift from abundant labeled data, recent supervised neural models have greatly promoted the development of lo…

Cited by 0SourcecodeScholar
2023

MIMO Is All You Need:A Strong Multi-in-Multi-Out Baseline for Video Prediction

AAAI 2023technical

The mainstream of the existing approaches for video prediction builds up their models based on a Single-In-Single-Out (SISO) architecture, which takes the current frame as input to predict the next frame in a recursive manner. This way often leads to severe performance degradation when they try to e…

2023

MUG: A General Meeting Understanding and Generation Benchmark

ICASSP 2023accepted

Listening to long video/audio recordings from video conferencing and online courses for acquiring information is extremely inefficient. Even after ASR systems transcribe recordings into long-form spoken language documents, reading ASR transcripts only partly speeds up seeking information. It has bee…

Cited by 0SourceScholar
2023

Meeting Action Item Detection with Regularized Context Modeling

ICASSP 2023accepted

Meetings are increasingly important for collaborations. Action items in meeting transcripts are crucial for managing post-meeting to-do tasks, which usually are summarized laboriously. The Action Item Detection task aims to automatically detect meeting content associated with action items. However,…

Cited by 0SourceScholar
2023

Overview of the ICASSP 2023 General Meeting Understanding and Generation Challenge (MUG)

ICASSP 2023accepted

ICASSP2023 General Meeting Understanding and Generation Challenge (MUG) focuses on prompting a wide range of spoken language processing (SLP) research on meeting transcripts, as SLP applications are critical to improve users’ efficiency in grasping important information in meetings. MUG includes fiv…

Cited by 0SourceScholar
2023

Pushing the Limits of Self-Supervised Speaker Verification using Regularized Distillation Framework

ICASSP 2023accepted

Training robust speaker verification systems without speaker labels has long been a challenging task. Previous studies observed a large performance gap between self-supervised and fully supervised methods. In this paper, we apply a non-contrastive self-supervised learning framework called DIstillati…

Cited by 0SourceScholar
2023

RECESS Vaccine for Federated Learning: Proactive Defense Against Model Poisoning Attacks

NeurIPS 2023poster

Model poisoning attacks greatly jeopardize the application of federated learning (FL). The effectiveness of existing defenses is susceptible to the latest model poisoning attacks, leading to a decrease in prediction accuracy. Besides, these defenses are intractable to distinguish benign outliers fro…

Cited by 13SourcePDFScholar
2023

Weighted Sampling for Masked Language Modeling

ICASSP 2023accepted

Masked Language Modeling (MLM) is widely used to pretrain language models. The standard random masking strategy in MLM causes the pre-trained language models (PLMs) to be biased towards high-frequency tokens. Representation learning of rare tokens is poor and PLMs have limited performance on downstr…

Cited by 0SourceScholar
2022

Bagging Regional Classification Activation Maps for Weakly Supervised Object Localization

ECCV 2022poster

"Classification activation map (CAM), utilizing the classification structure to generate pixel-wise localization maps, is a crucial mechanism for weakly supervised object localization (WSOL). However, CAM directly uses the classifier trained on image-level features to locate objects, making it prefe…

2022

ICASSP-SPGC 2022: Root Cause Analysis for Wireless Network Fault Localization

ICASSP 2022accepted

Localizing the root cause of network faults is crucial to network operation and maintenance (O&M). Significant operational expenses will be saved if the root cause can be identified agilely and accurately. However, this is challenging for human beings due to the complicated wireless environments and…

Cited by 0SourceScholar
2022

MDERank: A Masked Document Embedding Rank Approach for Unsupervised Keyphrase Extraction

ACL 2022findings

Keyphrase extraction (KPE) automatically extracts phrases in a document that provide a concise summary of the core content, which benefits downstream information retrieval and NLP tasks. Previous state-of-the-art methods select candidate keyphrases based on the similarity between learned representat…

2022

PoNet: Pooling Network for Efficient Token Mixing in Long Sequences

ICLR 2022poster

Transformer-based models have achieved great success in various NLP, vision, and speech tasks. However, the core of Transformer, the self-attention mechanism, has a quadratic time and memory complexity with respect to the sequence length, which hinders applications of Transformer-based models to lon…

2022

Prosospeech: Enhancing Prosody with Quantized Vector Pre-Training in Text-To-Speech

ICASSP 2022accepted

Expressive text-to-speech (TTS) has become a hot research topic recently, mainly focusing on modeling prosody in speech. Prosody modeling has several challenges: 1) the extracted pitch used in previous prosody modeling works have inevitable errors, which hurts the prosody modeling; 2) different attr…

Cited by 0SourceScholar
2022

Weakly Supervised Object Localization As Domain Adaption

CVPR 2022poster

Weakly supervised object localization (WSOL) focuses on localizing objects only with the supervision of image-level classification masks. Most previous WSOL methods follow the classification activation map (CAM) that localizes objects based on the classification structure with the multi-instance lea…

Cited by 44PDFcodeScholar
2021

RGB-D Salient Object Detection via 3D Convolutional Neural Networks

AAAI 2021technical

RGB-D salient object detection (SOD) recently has attracted increasing research interest and many deep learning methods based on encoder-decoder architectures have emerged. However, most existing RGB-D SOD models conduct feature fusion either in the single encoder or the decoder stage, which hardly…

2021

TRS: Transferability Reduced Ensemble via Promoting Gradient Diversity and Model Smoothness

NeurIPS 2021poster

Adversarial Transferability is an intriguing property - adversarial perturbation crafted against one model is also effective against another model, while these models are from different model families or training processes. To better protect ML systems against adversarial attacks, several questions…

Cited by 77SourcePDFScholar
2020

Controllable Time-Delay Transformer for Real-Time Punctuation Prediction and Disfluency Detection

ICASSP 2020accepted

With the increased applications of automatic speech recognition (ASR) in recent years, it is essential to automatically insert punctuation marks and remove disfluencies in transcripts, to improve the readability of the transcripts as well as the performance of subsequent applications, such as machin…

Cited by 0SourceScholar