← Search

Xu LI

37 accepted papers

2026

A Logical Analysis of an Information Filtering Architecture Based on Epistemic Trust Inference

AAAI 2026technical

In agent theory, epistemic trust is used to infer beliefs, for example by filtering out the information the agent receives from untrustworthy agents. Moreover, trust itself can be inferred from other information. We introduce a simple information filtering architecture that clearly distinguishes the

Cited by 0SourcePDFScholar
2026

D&R: Recovery-based AI-Generated Text Detection via a Single Black-box LLM Call

ICLR 2026poster

Large language models (LLMs) generate increasingly human-like text, raising concerns about misinformation and authenticity. Detecting AI-generated text remains challenging: existing methods often underperform, especially on short texts, require probability access unavailable in real-world black-box…

Cited by 0SourcecodeScholar
2026

Decomposition of Concept-Level Rules in Visual Scenes

ICLR 2026poster

Human cognition is compositional, and one can parse a visual scene into independent concepts and the corresponding concept-changing rules. By contrast, many vision-language systems process images holistically, with limited support for explicit decomposition. And previous methods of decomposing conce…

Cited by 0SourceScholar
2026

Dual-stream Relation-modeling Disentanglement for Cloth-Changing Person Re-Identification

AAAI 2026technical

Cloth-changing person re-identification (CC-ReID) aims to identify individuals across non-overlapping cameras despite clothing variations. Existing methods are often constrained by two primary limitations: approaches using auxiliary modalities typically rely on a single specific cue, limiting their

Cited by 0SourcePDFScholar
2026

Envision, Attend, Then Respond: Counterfactual Hallucination Mitigation in Large Vision-Language Models

CVPR 2026

Large Vision-Language Models (LVLMs) often hallucinate when visual evidence conflicts with world knowledge, i.e., in counterfactual scenarios. We propose Envision-Attend-Respond (EnAR), a training-free framework that leverages visual priors to steer the model's attention toward counterfactual elemen

Cited by 0SourcecodeScholar
2026

Explore to Learn: Latent Exploration Through Disentangled Synergy Patterns for Reinforcement Learning in Overactuated Control

AAAI 2026technical

Control in high-dimensional action spaces remains a fundamental challenge in reinforcement learning (RL), primarily due to inefficient exploration of the action space. While recent methods attempt to guide exploration, they often fall short of achieving the agility and coordination exhibited in biol

Cited by 0SourcePDFScholar
2026

LiteFT-PR: Lightweight and Fault-Tolerant LiDAR-Camera Fusion Network for Robust Place Recognition via Model Distillation

RA-L 2026

Place recognition (PR) is a key component of simultaneous localization and mapping (SLAM) in autonomous vehicles and robotics. By efficiently matching descriptors generated from the current scene with a prebuilt reference database, existing PR methods enable accurate vehicle re-localization. However

Cited by 0SourceScholar
2026

Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents

ICML 2026poster

LLM-based agents are becoming increasingly capable, yet their safety lags behind. This creates a gap between what agents can do and should do. This gap widens as agents engage in multi-turn interactions and employ diverse tools, introducing new risks overlooked by existing benchmarks. To systematica…

Cited by 0SourceScholar
2026

YuE: Scaling Open Foundation Models for Long-Form Music Generation

ICLR 2026poster

We tackle the task of long-form music generation, particularly the challenging \textbf{lyrics-to-song} problem, by introducing \textbf{YuE (乐)}, a family of open-source music generation foundation models. Specifically, YuE scales to trillions of tokens and generates up to five minutes of music while…

Cited by 0SourcecodeScholar
2025

An Inflatable Deployable Origami Grasper for Adaptive and High-Load Grasping

IROS 2025

Robotic graspers are essential for enhancing the efficiency and versatility of robots in grasping tasks. In this paper, we propose a novel inflatable deployable origami grasper with a rigid-flexible coupling structure. The proposed grasper can achieve multiple deployment configurations under a singl

Cited by 0SourceScholar
2025

Audio-Faces Intra-Frame Alignment with Graph Attention Networks for Active Speaker Detection

ICASSP 2025accepted

Audio-Visual Active Speaker Detection(ASD) is the task of identifying, at any given moment, who is actively speaking in a multi-person scene by using audio and visual cues. Current main stream ASD methods separately encode audio and facial features, then adopt post-feature fusion approach where the…

Cited by 0SourceScholar
2025

Capturing Rich Behavior Representations: A Dynamic Action Semantic-Aware Graph Transformer for Video Captioning

ICASSP 2025accepted

Existing video captioning methods merely provide shallow or simplistic representations of object behaviors, resulting in superficial and ambiguous descriptions. However, object behavior is dynamic and complex. To comprehensively capture the essence of object behavior, we propose a dynamic action sem…

Cited by 0SourceScholar
2025

Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration With Improved Intelligibility

ICASSP 2025accepted

Speech restoration aims at restoring full-band speech with high quality and intelligibility, considering a diverse set of distortions. MaskSR is a recently proposed generative model for this task. As other models of its kind, MaskSR attains high quality but, as we show, intelligibility can be substa…

Cited by 0SourceScholar
2025

Language-Queried Target Sound Extraction Without Parallel Training Data

ICASSP 2025accepted

Language-queried target sound extraction (TSE) aims to extract specific sounds from mixtures based on language queries. Traditional fully-supervised training schemes require extensively annotated parallel audio-text data, which are labor-intensive. We introduce a parallel-data-free training scheme,…

Cited by 0SourceScholar
2025

NoT: Federated Unlearning via Weight Negation

CVPR 2025poster

Federated unlearning (FU) aims to remove a participant's data contributions from a trained federated learning (FL) model, ensuring privacy and regulatory compliance. Traditional FU methods often depend on auxiliary storage on either the client or server side or require direct access to the data targ…

Cited by 1SourcePDFScholar
2025

Zero-shot Stance Detection with Logically Consistent Data Augmentation

ICASSP 2025accepted

Zero-shot stance detection (ZSSD) is a challenging task that requires classifying stances towards unseen targets without large, well-curated training datasets. Existing data augmentation methods for ZSSD often suffer from semantic inconsistencies, hindering their effectiveness. To address these limi…

Cited by 0SourceScholar
2024

EA-VTR: Event-Aware Video-Text Retrieval

ECCV 2024poster

"Understanding the content of events occurring in the video and their inherent temporal logic is crucial for video-text retrieval. However, web-crawled pre-training datasets often lack sufficient event information, and the widely adopted video-level cross-modal contrastive learning also struggles to…

Cited by 3SourcePDFScholar
2024

Humtrans: A Novel Open-Source Dataset for Humming Melody Transcription and Beyond

ICASSP 2024accepted

This paper introduces the HumTrans dataset, which is publicly available and primarily designed for humming melody transcription. The dataset can also serve as a foundation for downstream tasks such as humming melody based music generation. It consists of 500 musical compositions of different genres…

Cited by 0SourceScholar
2024

Neural Concatenative Singing Voice Conversion: Rethinking Concatenation-Based Approach for One-Shot Singing Voice Conversion

ICASSP 2024accepted

Any-to-any singing voice conversion (SVC) is confronted with the challenge of "timbre leakage" issue caused by inadequate disentanglement between the content and the speaker timbre. To address this issue, this study introduces NeuCoSVC, a novel neural concatenative SVC framework. It consists of a se…

Cited by 0SourceScholar
2023

Covariance Regularization for Probabilistic Linear Discriminant Analysis

ICASSP 2023accepted

Probabilistic linear discriminant analysis (PLDA) is commonly used in speaker verification systems to score the similarity of speaker embeddings. Recent studies improved the performance of PLDA in domain-matched conditions by diagonalizing its covariance. We suspect such a brutal pruning approach co…

Cited by 0SourceScholar
2023

Enhancing the Vocal Range of Single-Speaker Singing Voice Synthesis with Melody-Unsupervised Pre-Training

ICASSP 2023accepted

The single-speaker singing voice synthesis (SVS) usually underperforms at pitch values that are out of the singer's vocal range or associated with limited training samples. Based on our previous work, this work proposes a melody-unsupervised multi-speaker pretraining method conducted on a multi-sing…

Cited by 0SourceScholar
2023

Quantitative Evidence on Overlooked Aspects of Enrollment Speaker Embeddings for Target Speaker Separation

ICASSP 2023accepted

Single channel target speaker separation (TSS) aims at extracting a speaker’s voice from a mixture of multiple talkers given an enrollment utterance of that speaker. A typical deep learning TSS framework consists of an upstream model that obtains enrollment speaker embeddings and a downstream model…

Cited by 0SourceScholar
2022

Characterizing the Adversarial Vulnerability of Speech self-Supervised Learning

ICASSP 2022accepted

A leaderboard named Speech processing Universal PERformance Benchmark (SUPERB), which aims at benchmarking the performance of a shared self-supervised learning (SSL) speech model across various downstream speech tasks with minimal modification of architectures and a small amount of data, has fueled…

Cited by 0SourceScholar
2021

Adversarial Defense for Automatic Speaker Verification by Cascaded Self-Supervised Learning Models

ICASSP 2021accepted

Automatic speaker verification (ASV) is one of the core technologies in biometric identification. With the ubiquitous usage of ASV systems in safety-critical applications, more and more malicious attackers attempt to launch adversarial attacks at ASV systems. In the midst of the arms race between at…

Cited by 0SourceScholar
2021

Configuration Transformation of the Wheel-Legged Robot Using Inverse Dynamics Control

ICRA 2021poster

In this paper, the configuration transformation of Wheel-Legged Robot (WLR) is studied, which can enable the robot to change its multilinks configuration on Inverted Equilibrium Manifold (IEM), while keeping balance with a small location drift on the floor. First of all, the general form of dynamics…

Cited by 13SourceScholar
2021

Replay and Synthetic Speech Detection with Res2Net Architecture

ICASSP 2021accepted

Existing approaches for replay and synthetic speech detection still lack generalizability to unseen spoofing attacks. This work proposes to leverage a novel model structure, so-called Res2Net, to improve the anti-spoofing countermeasure’s generalizability. Res2Net mainly modifies the ResNet block to…

Cited by 0SourceScholar
2020

Adversarial Attacks on GMM I-Vector Based Speaker Verification Systems

ICASSP 2020accepted

This work investigates the vulnerability of Gaussian Mixture Model (GMM) i-vector based speaker verification systems to adversarial attacks, and the transferability of adversarial samples crafted from GMM i-vector based systems to x-vector based systems. In detail, we formulate the GMM i-vector syst…

Cited by 0SourceScholar
2019

End-to-end Code-switched TTS with Mix of Monolingual Recordings

ICASSP 2019accepted

State-of-the-art text-to-speech (TTS) synthesis models can produce monolingual speech with high intelligibility and naturalness. However, when the models are applied to synthesize code-switched (CS) speech, the performance declines seriously. Conventionally, developing a CS TTS system requires multi…

Cited by 0SourceScholar
2019

Speech Emotion Recognition Using Capsule Networks

ICASSP 2019accepted

Speech emotion recognition (SER) is a fundamental step towards fluent human-machine interaction. One challenging problem in SER is obtaining utterance-level feature representation for classification. Recent works on SER have made significant progress by using spectrogram features and introducing neu…

Cited by 0SourceScholar
2018

Applying Multitask Learning to Acoustic-Phonemic Model for Mispronunciation Detection and Diagnosis in L2 English Speech

ICASSP 2018accepted

For mispronunciation detection and diagnosis (MDD), nowadays approaches generally treat the phonemes in correct and mispronunciations as the same despite the fact they may actually carry different characteristics. Furthermore, serious data imbalance issue between correct and mispronunciation in data…

Cited by 0SourceScholar
2018

Design and Experiments of a Novel Hydraulic Wheel-Legged Robot (WLR)

IROS 2018poster

Wheel-legged hybrid robot with multi-modal locomotion can efficiently adapt to different terrain environments, as well as realize rapid maneuver on flat ground. We have developed a novel hydraulic wheel-legged robot (WLR) combined with a humanoid structural design. This robot can assist to emergency…

Cited by 89SourceScholar
2018

Unsupervised Discovery of an Extended Phoneme Set in L2 English Speech for Mispronunciation Detection and Diagnosis

ICASSP 2018accepted

Second language (L2) speech is often labelled with the native, phoneme categories. Hence, we often observe segments for which it is difficult, if not impossible, to decide on a categorical phoneme label. We refer to these segments as “non-categorical” phoneme units. Existing approaches to mispronunc…

Cited by 0SourceScholar
2016

Low level descriptors based DBLSTM bottleneck feature for speech driven talking avatar

ICASSP 2016accepted

Speech is bimodal in nature. There are close correlations between the acoustic speech signals and the visual gestures such as lip movements, facial expressions and head motions. For speech driven talking avatar, how to derive more representative acoustic features from which to predict more accurate…

Cited by 0SourceScholar