← Search

Xiaofei Wang

45 accepted papers

2026

EdiVal-Agent: An Object-Centric Framework for Automated, Fine-Grained Evaluation of Multi-Turn Editing

ICLR 2026poster

Instruction-based image editing has advanced rapidly, yet reliable and interpretable evaluation remains a bottleneck. Current protocols either (i) depend on paired reference images—resulting in limited coverage and inheriting biases from prior generative models—or (ii) rely *solely* on zero-shot vis…

Cited by 0SourcecodeScholar
2026

EpiTwin: Spatiotemporal Graph Transformers for Epileptic sEEG Signal Reconstruction

ICML 2026poster

Stereotactic electroencephalography (sEEG) provides temporally precise intracranial recordings but is inherently constrained by sparse and irregular spatial sampling due to clinical limitations on electrode implantation. Signal reconstruction under this setting aims to infer neural activity at unmon…

Cited by 0SourceScholar
2026

FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates

ICLR 2026poster

Neural audio codecs are foundational to speech language models. It is expected to have a low frame rate and decoupled semantic and acoustic information. A lower frame rate codec can reduce the computational cost of speech language models by shortening the sequence length. Recent studies have develop…

Cited by 0SourcecodeScholar
2026

Independence Test for Linear Non-Gaussian Data and Applications in Causal Discovery

ICLR 2026poster

Independence testing involves determining whether two variables are independent based on observed samples, which is a fundamental problem in statistics and machine learning. Existing testing methods, such as HSIC, can theoretically detect broad forms of dependence, but may sacrifice statistical powe…

Cited by 0SourceScholar
2026

Less Is More: Sparse and Cooperative Perturbation for Point Cloud Attacks

AAAI 2026technical

Most adversarial attacks on point clouds perturb a large number of points, causing widespread geometric changes and limiting applicability in real-world scenarios. While recent works explore sparse attacks by modifying only a few points, such approaches often struggle to maintain effectiveness due t

Cited by 0SourcePDFScholar
2026

MAPS: Memory-Aware Predictive Scheduling Framework for Large Language Models Serving

ICML 2026poster

The surge of large language model (LLM) applications on personal devices imposes massive, bursty workloads on cloud serving infrastructure. While prefill-decode disaggregation improves throughput and scalability, memory-bound decode instances often suffer from persistent load imbalance, as output le…

Cited by 0SourceScholar
2026

Noise Suppression Method of Passive Electric Sense System for Robotic Shark

RA-L 2026

Passive electric sense, a unique underwater perception ability in aquatic species, enables fish to detect external electrical signals and perform essential life activities in dark or turbid environments. Inspired by the passive electrosensing capability of natural shark, this letter proposes a low-c

Cited by 0SourceScholar
2026

Optimal Transport-Induced Samples against Out-of-Distribution Overconfidence

ICLR 2026poster

Deep neural networks (DNNs) often produce overconfident predictions on out-of-distribution (OOD) inputs, undermining their reliability in open-world environments. Singularities in semi-discrete optimal transport (OT) mark regions of semantic ambiguity, where classifiers are particularly prone to unw…

Cited by 0SourceScholar
2026

STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language Models

ICLR 2026poster

Spoken Language Models (SLMs) are designed to take speech inputs and produce spoken responses. However, current SLMs lack the ability to perform an internal, unspoken thinking process before responding. In contrast, humans typically engage in complex mental reasoning internally, enabling them to com…

Cited by 0SourcecodeScholar
2026

Stratos: An End-to-End Distillation Pipeline for Customized LLMs Under Distributed Cloud Environments

AAAI 2026technical

The growing industrial demand for customized and cost-efficient large language models (LLMs) is fueled by the rise of vertical, domain-specific tasks and the need to optimize performance under constraints such as latency and budget. Knowledge distillation, as an efficient model compression and trans

Cited by 0SourcePDFScholar
2025

AdvGrasp: Adversarial Attacks on Robotic Grasping from a Physical Perspective

IJCAI 2025

Adversarial attacks on robotic grasping provide valuable insights into evaluating and improving the robustness of these systems. Unlike studies that focus solely on neural network predictions while overlooking the physical principles of grasping, this paper introduces AdvGrasp, a framework for adver

Cited by 0SourcePDFScholar
2025

Audio-Aware Large Language Models as Judges for Speaking Styles

EMNLP 2025

Audio-aware large language models (ALLMs) can understand the textual and non-textual information in the audio input. In this paper, we explore using ALLMs as an automatic judge to assess the speaking styles of speeches. We use ALLM judges to evaluate the speeches generated by SLMs on two tasks: voic

2025

CoVoMix2: Advancing Zero-Shot Dialogue Generation with Fully Non-Autoregressive Flow Matching

NeurIPS 2025poster

Generating natural-sounding, multi-speaker dialogue is crucial for applications such as podcast creation, virtual agents, and multimedia content generation. However, existing systems struggle to maintain speaker consistency, model overlapping speech, and synthesize coherent conversations efficiently…

Cited by 0SourceScholar
2025

ELLA-V: Stable Neural Codec Language Modeling with Alignment-Guided Sequence Reordering

AAAI 2025technical

The language model (LM) approach based on acoustic and linguistic prompts, such as VALL-E, has achieved remarkable progress in the field of zero-shot audio generation. However, existing methods still have some limitations: 1) repetitions, transpositions, and omissions in the output synthesized speec…

2025

Enhancing Statistical Validity and Power in Hybrid Controlled Trials: A Randomization Inference Approach with Conformal Selective Borrowing

ICML 2025poster

External controls from historical trials or observational data can augment randomized controlled trials when large-scale randomization is impractical or unethical, such as in drug evaluation for rare diseases. However, non-randomized external controls can introduce biases, and existing Bayesian and…

2025

Extracting Rare Dependence Patterns via Adaptive Sample Reweighting

ICML 2025poster

Discovering dependence patterns between variables from observational data is a fundamental issue in data analysis. However, existing testing methods often fail to detect subtle yet critical patterns that occur within small regions of the data distribution--patterns we term rare dependence. These rar…

Cited by 0SourcePDFScholar
2025

Imperceptible 3D Point Cloud Attacks on Lattice-based Barycentric Coordinates

AAAI 2025technical

Imperceptible adversarial attacks on 3D point clouds rely on effective constraints. While manifold constraints have notable advantages over Euclidean ones, the global parameterization used in current methods often fails to fully preserve manifold properties. In this paper, we propose to constrain la…

Cited by 1SourcePDFScholar
2025

Imperceptible Adversarial Attacks on Point Clouds Guided by Point-to-Surface Field

ICASSP 2025accepted

Adversarial attacks on point clouds are crucial for assessing and improving the adversarial robustness of 3D deep learning models. Traditional solutions strictly limit point displacement during attacks, making it challenging to balance imperceptibility with adversarial effectiveness. In this paper,…

Cited by 0SourceScholar
2024

CoVoMix: Advancing Zero-Shot Speech Generation for Human-like Multi-talker Conversations

NeurIPS 2024poster

Recent advancements in zero-shot text-to-speech (TTS) modeling have led to significant strides in generating high-fidelity and diverse speech. However, dialogue generation, along with achieving human-like naturalness in speech, continues to be a challenge. In this paper, we introduce CoVoMix: Conver…

2024

Diarist: Streaming Speech Translation with Speaker Diarization

ICASSP 2024accepted

End-to-end speech translation (ST) for conversation recordings involves several under-explored challenges such as speaker diarization (SD) without accurate word time stamps and handling of overlapping speech in a streaming fashion. In this work, we propose DiariST, the first streaming ST and SD solu…

Cited by 0SourceScholar
2024

EasyTS: The Express Lane to Long Time Series Forecasting

AAAI 2024technical

Responding to the escalating interest in long-term forecasting within the industry, we introduce EasyTS, a comprehensive toolkit engineered to streamline data collection, analysis, and model creation procedures. EasyTS acts as a unified solution, driving progress in long-term time series forecasting…

2024

Eliminating Cross-modal Conflicts in BEV Space for LiDAR-Camera 3D Object Detection

ICRA 2024poster

Recent 3D object detectors typically utilize multi-sensor data and unify multi-modal features in the shared bird’s-eye view (BEV) representation space. However, our empirical findings indicate that previous methods have limitations in generating fusion BEV features free from cross-modal conflicts. T…

Cited by 12SourcecodeScholar
2024

FLAT: Flux-aware Imperceptible Adversarial Attacks on 3D Point Clouds

ECCV 2024poster

"Adversarial attacks on point clouds play a vital role in assessing and enhancing the adversarial robustness of 3D deep learning models. While employing a variety of geometric constraints, existing adversarial attack solutions often display unsatisfactory imperceptibility due to inadequate considera…

Cited by 5SourcePDFScholar
2024

Global-Local Collaborative Inference with LLM for Lidar-Based Open-Vocabulary Detection

ECCV 2024poster

"Open-Vocabulary Detection (OVD) is the task of detecting all interesting objects in a given scene without predefined object classes. Extensive work has been done to deal with the OVD for 2D RGB images, but the exploration of 3D OVD is still limited. Intuitively, lidar point clouds provide 3D inform…

2024

TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation

NeurIPS 2024poster

There is a rising interest and trend in research towards directly translating speech from one language to another, known as end-to-end speech-to-speech translation. However, most end-to-end models struggle to outperform cascade models, i.e., a pipeline framework by concatenating speech recognition,…

2023

DATA2VEC-SG: Improving Self-Supervised Learning Representations for Speech Generation Tasks

ICASSP 2023accepted

Self-supervised learning has been successfully applied to various speech recognition and understanding tasks. However, for generative tasks such as speech enhancement and speech separation, most self-supervised speech representations did not show substantial improvements. To deal with this problem,…

Cited by 0SourceScholar
2023

Self-Supervised Learning with Bi-Label Masked Speech Prediction for Streaming Multi-Talker Speech Recognition

ICASSP 2023accepted

Self-supervised learning (SSL), which utilizes the input data itself for representation learning, has achieved state-of-the-art results for various downstream speech tasks. However, most of the previous studies focused on offline single-talker applications, with limited investigations in multi-talke…

Cited by 0SourceScholar
2023

Simulating Realistic Speech Overlaps Improves Multi-Talker ASR

ICASSP 2023accepted

Multi-talker automatic speech recognition (ASR) has been studied to generate transcriptions of natural conversation including over-lapping speech of multiple speakers. Due to the difficulty in acquiring real conversation data with high-quality human transcriptions, a naïve simulation of multi-talker…

Cited by 18SourceScholar
2023

Speech Separation with Large-Scale Self-Supervised Learning

ICASSP 2023accepted

Self-supervised learning (SSL) methods such as WavLM have shown promising speech separation (SS) results in small-scale simulation-based experiments. In this work, we extend the exploration of the SSL-based SS by massively scaling up both the pre-training data (more than 300K hours) and fine-tuning…

Cited by 0SourceScholar
2023

Vararray Meets T-Sot: Advancing the State of the Art of Streaming Distant Conversational Speech Recognition

ICASSP 2023accepted

This paper presents a novel streaming automatic speech recognition (ASR) framework for multi-talker overlapping speech captured by a distant microphone array with an arbitrary geometry. Our framework, named t-SOT-VA, capitalizes on independently developed two recent technologies; array-geometry-agno…

Cited by 0SourceScholar
2022

All-Neural Beamformer for Continuous Speech Separation

ICASSP 2022accepted

Continuous speech separation (CSS) aims to separate overlapping voices from a continuous influx of conversational audio containing an unknown number of utterances spoken by an unknown number of speakers. A common application scenario is transcribing a meeting conversation recorded by a microphone ar…

Cited by 0SourceScholar
2022

Improving Noise Robustness of Contrastive Speech Representation Learning with Speech Reconstruction

ICASSP 2022accepted

Noise robustness is essential for deploying automatic speech recognition (ASR) systems in real-world environments. One way to reduce the effect of noise interference is to employ a preprocessing module that conducts speech enhancement, and then feed the enhanced speech to an ASR backend. In this wor…

Cited by 0SourceScholar
2022

Personalized speech enhancement: new models and Comprehensive evaluation

ICASSP 2022accepted

Personalized speech enhancement (PSE) models utilize additional cues, such as speaker embeddings like d-vectors, to remove background noise and interfering speech in real-time and thus improve the speech quality of online video conferencing systems for various acoustic scenarios. In this work, we pr…

Cited by 0SourceScholar
2022

Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers Using End-to-End Speaker-Attributed ASR

ICASSP 2022accepted

This paper presents Transcribe-to-Diarize, a new approach for neural speaker diarization that uses an end-to-end (E2E) speaker-attributed automatic speech recognition (SA-ASR). The E2E SA-ASR is a joint model that was recently proposed for speaker counting, multi-talker speech recognition, and speak…

Cited by 0SourceScholar
2022

VarArray: Array-Geometry-Agnostic Continuous Speech Separation

ICASSP 2022accepted

Continuous speech separation using a microphone array was shown to be promising in dealing with the speech overlap problem in natural conversation transcription. This paper proposes VarArray, an array-geometry-agnostic speech separation neural network model. The proposed model is applicable to any n…

Cited by 47SourceScholar
2021

Deep Multi-Task Learning for Diabetic Retinopathy Grading in Fundus Images

AAAI 2021technical

Recent years have witnessed the growing interest in disease severity grading, especially for ocular diseases based on fundus images. The existing grading methods are usually trained with high resolution (HR) images. However, the grading performance decreases a lot given low resolution (LR) images, w…

Cited by 42SourcePDFScholar
2021

Hypothesis Stitcher for End-to-End Speaker-Attributed ASR on Long-Form Multi-Talker Recordings

ICASSP 2021accepted

An end-to-end (E2E) speaker-attributed automatic speech recognition (SA-ASR) model was proposed recently to jointly perform speaker counting, speech recognition and speaker identification. The model achieved a low speaker-attributed word error rate (SA-WER) for monaural overlapped speech comprising…

Cited by 0SourceScholar
2021

Minimum Bayes Risk Training for End-to-End Speaker-Attributed ASR

ICASSP 2021accepted

Recently, an end-to-end speaker-attributed automatic speech recognition (E2E SA-ASR) model was proposed as a joint model of speaker counting, speech recognition and speaker identification for monaural overlapped speech. In the previous study, the model parameters were trained based on the speaker-at…

Cited by 0SourceScholar
2021

Reinforcement Learning with Latent Flow

NeurIPS 2021poster

Temporal information is essential to learning effective policies with Reinforcement Learning (RL). However, current state-of-the-art RL algorithms either assume that such information is given as part of the state space or, when learning from pixels, use the simple heuristic of frame-stacking to imp…

Cited by 29SourcePDFScholar
2021

Skill Preferences: Learning to Extract and Execute Robotic Skills from Human Feedback

CoRL 2021poster

A promising approach to solving challenging long-horizon tasks has been to extract behavior priors (skills) by fitting generative models to large offline datasets of demonstrations. However, such generative models inherit the biases of the underlying data and result in poor and unusable skills when…

Cited by 48SourceScholar
2020

A Practical Two-Stage Training Strategy for Multi-Stream End-to-End Speech Recognition

ICASSP 2020accepted

The multi-stream paradigm of audio processing, in which several sources are simultaneously considered, has been an active research area for information fusion. Our previous study offered a promising direction within end-to-end automatic speech recognition, where parallel encoders aim to capture dive…

Cited by 0SourceScholar
2019

Attention Based Glaucoma Detection: A Large-Scale Database and CNN Model

CVPR 2019poster

Recently, the attention mechanism has been successfully applied in convolutional neural networks (CNNs), significantly boosting the performance of many computer vision tasks. Unfortunately, few medical image recognition approaches incorporate the attention mechanism in the CNNs. In particular, there…

Cited by 302PDFcodeScholar
2019

Stream Attention-based Multi-array End-to-end Speech Recognition

ICASSP 2019accepted

Automatic Speech Recognition (ASR) using multiple microphone arrays has achieved great success in the far-field robustness. Taking advantage of all the information that each array shares and contributes is crucial in this task. Motivated by the advances of joint Connectionist Temporal Classification…

Cited by 21SourceScholar