← Search

Xiaodong Li

34 accepted papers

2026

GOMPSNR: Reflourish the Signal-to-Noise Ratio Metric for Audio Generation Tasks

AAAI 2026technical

In the field of audio generation, signal-to-noise ratio (SNR) has long served as an objective metric for evaluating audio quality. Nevertheless, recent studies have shown that SNR and its variants are not always highly correlated with human perception, prompting us to raise the questions: Why does

Cited by 0SourcePDFScholar
2026

SLD-L2S: Hierarchical Subspace Latent Diffusion for High-Fidelity Lip to Speech Synthesis

AAAI 2026technical

Although lip-to-speech synthesis (L2S) has achieved significant progress in recent years, current state-of-the-art methods typically rely on intermediate representations such as mel-spectrograms or discrete self-supervised learning (SSL) tokens. The potential of latent diffusion models (LDMs) in thi

Cited by 0SourcePDFScholar
2026

Safe and Efficient Control: A Subgraph-Augmented Hierarchical Reinforcement Learning Framework for Dynamically Reconfigurable Battery Systems

IJCAI 2026

Dynamically Reconfigurable Battery (DRB) systems employ power electronic switches to create dynamic topologies. They enable effective management of cell difference through real-time adjustment of cell connections. However, existing DRB control methods struggle to learn effective strategies due to sp

Cited by 0Scholar
2026

Taming Noise-Induced Prototype Degradation for Privacy-Preserving Personalized Federated Fine-Tuning

CVPR 2026

Prototype-based Personalized Federated Learning (ProtoPFL) enables efficient multi-domain adaptation by communicating compact class prototypes, but directly sharing them poses privacy risks. A common defense involves per-example l_2 clipping before prototype computation to bound sensitivity, followe

Cited by 0SourcecodeScholar
2025

Are Large Language Models Ready for Multi-Turn Tabular Data Analysis?

ICML 2025poster

Conversational Tabular Data Analysis, a collaboration between humans and machines, enables real-time data exploration for informed decision-making. The challenges and costs of collecting realistic conversational logs for tabular data analysis hinder comprehensive quantitative evaluation of Large Lan…

Cited by 0SourcePDFScholar
2025

Audiogram-Informed End-to-End Noise Reduction and Wide Dynamic Range Compression for Hearing Aids

ICASSP 2025accepted

Wide dynamic range compression (WDRC) provides level-dependent amplification, intended to make the output of a hearing aid fall between the hearing threshold and the highest comfortable level of the listener. Hearing aids often combine noise reduction with WDRC, applied sequentially. Unfortunately,…

Cited by 0SourceScholar
2025

Beyond the Answer: Advancing Multi-Hop QA with Fine-Grained Graph Reasoning and Evaluation

ACL 2025long

Recent advancements in large language models (LLMs) have significantly improved the performance of multi-hop question answering (MHQA) systems. Despite the success of MHQA systems, the evaluation of MHQA is not deeply investigated. Existing evaluations mainly focus on comparing the final answers of…

2025

DSINet: Towards Real-Time Target Speaker Extraction with Dynamic Speaker Information Fusion

ICASSP 2025accepted

Target speaker extraction (TSE) aims to directly extract the desired speech given enrollment utterances of the target speaker. Despite significant progress in recent years, most existing methods remain non-causal and computationally intensive. This paper introduces DSINet, a real-time time-frequency…

Cited by 0SourceScholar
2025

DeepPEM-AFC: An Improved Prediction-Error-Method-based Adaptive Feedback Cancellation with Deep Learning for Hearing Aids

ICASSP 2025accepted

Hearing assistive devices aim to compensate hearing loss for hearing-impaired listeners, and their maximum stable gain (MSG) is constrained because of the existence of the acoustic feedback between the receiver and microphone, resulting in their inefficiency for individuals with severe or profound h…

Cited by 0SourceScholar
2025

Factor Graph-based Interpretable Neural Networks

ICLR 2025poster

Comprehensible neural network explanations are foundations for a better understanding of decisions, especially when the input data are infused with malicious perturbations. Existing solutions generally mitigate the impact of perturbations through adversarial training, yet they fail to generate compr…

2025

Hyperbolic-PDE GNN: Spectral Graph Neural Networks in the Perspective of A System of Hyperbolic Partial Differential Equations

ICML 2025poster

Graph neural networks (GNNs) leverage message passing mechanisms to learn the topological features of graph data. Traditional GNNs learns node features in a spatial domain unrelated to the topology, which can hardly ensure topological features. In this paper, we formulates message passing as a syste…

2025

Learning Neural Vocoder from Range-Null Space Decomposition

IJCAI 2025

Despite the rapid development of neural vocoders in recent years, they usually suffer from some intrinsic challenges like opaque modeling, and parameter-performance trade-off. In this study, we propose an innovative time-frequency (T-F) domain-based neural vocoder to resolve the above-mentioned chal

2025

LiteFat: Lightweight Spatio-Temporal Graph Learning for Real-Time Driver Fatigue Detection

IROS 2025

Detecting driver fatigue is critical for road safety, as drowsy driving remains a leading cause of traffic accidents. Many existing solutions rely on computationally demanding deep learning models, which result in high latency and are unsuitable for embedded robotic devices with limited resources (s

Cited by 3SourceScholar
2025

M2PA: A Multi-Memory Planning Agent for Open Worlds Inspired by Cognitive Theory

ACL 2025finding

Open-world planning poses a significant challenge for general artificial intelligence due to environmental complexity and task diversity, especially in long-term tasks and lifelong learning. Inspired by cognitive theories, we propose M2PA, an open-world multi-memory planning agent. M2PA innovates by…

Cited by 0SourcePDFScholar
2025

Metagent-P: A Neuro-Symbolic Planning Agent with Metacognition for Open Worlds

ACL 2025finding

The challenge of developing agents capable of open-world planning remains fundamental to artificial general intelligence (AGI). While large language models (LLMs) have made progress with their vast world knowledge, their limitations in perception, memory, and reliable reasoning still hinder LLM-base…

Cited by 0SourcePDFScholar
2025

Micro-Act: Mitigate Knowledge Conflict in Question Answering via Actionable Self-Reasoning

ACL 2025long

Retrieval-Augmented Generation (RAG) systems commonly suffer from **Knowledge Conflicts**, where retrieved external knowledge contradicts the inherent, parametric knowledge of large language models (LLMs). It adversely affects performance on downstream tasks such as question answering (QA). Existing…

2025

Optimize Battery Control: A Multi-Objective Evolutionary Ensemble Reinforcement Learning Approach

IJCAI 2025

The Dynamically Reconfigurable Battery (DRB) systems, which use high-speed power electronic switches to dynamically adjust battery interconnections in real-time, are critical to the performance of the battery pack. Traditional battery management strategies often fail to address multi-objective optim

Cited by 0SourcePDFScholar
2025

SOTOPIA-Ω: Dynamic Strategy Injection Learning and Social Instruction Following Evaluation for Social Agents

ACL 2025long

Despite the abundance of prior social strategies possessed by humans, there remains a paucity of research dedicated to their transfer and integration into social agents. Our proposed SOTOPIA-Ω framework aims to address and bridge this gap, with a particular focus on enhancing the social capabilities…

2025

Sharper Error Bounds in Late Fusion Multi-view Clustering with Eigenvalue Proportion Optimization

AAAI 2025technical

Multi-view clustering (MVC) aims to integrate complementary information from multiple views to enhance clustering performance. Late Fusion Multi-View Clustering (LFMVC) has shown promise by synthesizing diverse clustering results into a unified consensus. However, current LFMVC methods struggle with…

2024

Adaptive Stabilization Based on Machine Learning for Column Generation

ICML 2024poster

Column generation (CG) is a well-established method for solving large-scale linear programs. It involves iteratively optimizing a subproblem containing a subset of columns and using its dual solution to generate new columns with negative reduced costs. This process continues until the dual values co…

2024

All Neural Kronecker Product Beamforming for Speech Extraction with Large-Scale Microphone Arrays

ICASSP 2024accepted

Existing frame-wise neural beamformers for speech extraction can obtain promising performance in relatively high signal-to-noise ratio (SNR) scenarios using small microphone arrays, while they still suffer from performance degradation in relatively low SNR environments, e.g., SNR<-5 dB. As an attemp…

Cited by 0SourceScholar
2024

Fine-Grained Task Planning for Service Robots Based on Object Ontology Knowledge via Large Language Models

RA-L 2024

In domestic environment, the successful execution of service tasks heavily relies on the robot's capability to identify and understand objects within its surrounding. This crucial process predominantly takes place during task planning, prior to the actual performance of service tasks. Therefore, it

Cited by 10SourceScholar
2024

Higher Order Multiple Graph Filtering for Structured Graph Learning

ICASSP 2024accepted

In the field of machine learning, multi-view clustering aims to reveal hidden clustering patterns across different data perspectives. However, traditional methods often struggle due to their reliance on low-order similarity data. To overcome this, we propose a new approach that integrates the learni…

Cited by 0SourceScholar
2024

Transformer-Based Relationship Inference Model for Household Object Organization by Integrating Graph Topology and Ontology

IROS 2024poster

In domestic environments, the conventional organization of objects by service robots often relies on the inherent properties of each object, such as placing fragile bowls in enclosed cupboards. However, this approach tends to overlook the importance of the orderly arrangement of objects, neglecting…

Cited by 0SourcecodeScholar
2023

Learning to Generate Columns with Application to Vertex Coloring

ICLR 2023poster

We present a new column generation approach based on Machine Learning (ML) for solving combinatorial optimization problems. The aim of our method is to generate high-quality columns that belong to an optimal integer solution, in contrast to the traditional approach that aims at solving linear progra…

Cited by 2SourcePDFScholar
2022

Embedding and Beamforming: All-Neural Causal Beamformer for Multichannel Speech Enhancement

ICASSP 2022accepted

Standing upon the intersection of traditional beamformers and deep neural networks, we propose a causal neural beamformer paradigm called Embedding and Beamforming, and two core modules are devised accordingly, namely EM and BM. For EM, instead of estimating spatial covariance matrix explicitly, the…

Cited by 0SourceScholar
2022

Enhancing Column Generation by a Machine-Learning-Based Pricing Heuristic for Graph Coloring

AAAI 2022technical

Column Generation (CG) is an effective method for solving large-scale optimization problems. CG starts by solving a subproblem with a subset of columns (i.e., variables) and gradually includes new columns that can improve the solution of the current subproblem. The new columns are generated as neede…

2022

PANDORA: A Panoramic Detection Dataset for Object with Orientation

ECCV 2022poster

"Panoramic images have become increasingly popular as omnidirectional panoramic technology has advanced. Many datasets and works resort to object detection to better understand the content of the panoramic image. These datasets and detectors use a Bounding Field of View (BFoV) as a bounding box in p…

2022

Taylor, Can You Hear Me Now? A Taylor-Unfolding Framework for Monaural Speech Enhancement

IJCAI 2022poster

While the deep learning techniques promote the rapid development of the speech enhancement (SE) community, most schemes only pursue the performance in a black-box manner and lack adequate model interpretability. Inspired by Taylor's approximation theory, we propose an interpretable decoupling-style…

2022

Unbiased IoU for Spherical Image Object Detection

AAAI 2022technical

As one of the fundamental components of object detection, intersection-over-union (IoU) calculations between two bounding boxes play an important role in samples selection, NMS operation and evaluation of object detection algorithms. This procedure is well-defined and solved for planar images, while…

Cited by 12SourcePDFScholar
2021

ICASSP 2021 Acoustic Echo Cancellation Challenge: Integrated Adaptive Echo Cancellation with Time Alignment and Deep Learning-Based Residual Echo Plus Noise Suppression

ICASSP 2021accepted

This paper describes a three-stage acoustic echo cancellation (AEC) and suppression framework for the ICASSP 2021 AEC Challenge. In the first stage, a partitioned block frequency domain adaptive filtering is implemented to cancel the linear echo components without introducing the near-end speech dis…

Cited by 0SourceScholar
2021

ICASSP 2021 Deep Noise Suppression Challenge: Decoupling Magnitude and Phase Optimization with a Two-Stage Deep Network

ICASSP 2021accepted

It remains a tough challenge to recover the speech signals contaminated by various noises under real acoustic environments. To this end, we propose a novel system for denoising in the complicated applications, which is mainly comprised of two pipelines, namely a two-stage network and a post-processi…

Cited by 0SourceScholar