← Search

Xiaoyu Liu

49 accepted papers

2026

Are Deep Speech Denoising Models Robust to Adversarial Noise?

ICLR 2026poster

Deep noise suppression (DNS) models enjoy widespread use throughout a variety of high-stakes speech applications. However, we show that four recent DNS models can each be reduced to outputting unintelligible gibberish through the addition of psychoacoustically hidden adversarial noise, even in low-…

Cited by 0SourceScholar
2026

Efficient Plug-and-Play Weight Refinement for Sparse Large Models

AAAI 2026technical

One-shot pruning efficiently compresses Large Language Models but produces coarse sparse weights, causing significant performance degradation. Traditional fine-tuning approaches to refine these weights are prohibitively expensive for large models. This highlights the need for a training-free weight

Cited by 0SourcePDFScholar
2026

Multi-Aspect Cross-modal Quantization for Generative Recommendation

AAAI 2026technical

Generative Recommendation (GR) has emerged as a new paradigm in recommender systems. This approach relies on quantized representations to discretize item features, modeling users’ historical interactions as sequences of discrete tokens. Based on these tokenized sequences, GR predicts the next item b

Cited by 0SourcePDFScholar
2026

SHERPA: Fine-tuning Segment Anything Models with Task-relevant Guidance

ICML 2026poster

Segment Anything Models (SAMs) often struggle with certain specialized tasks. A common approach is to fine-tune models with specific task labels, but this often leads to overfitting, introduces model bias and significantly degrades their generalization ability. To overcome these challenges, we propo…

Cited by 0SourceScholar
2025

A Framework for Effective Invocation Methods of Various LLM Services

COLING 2025main

Large Language Models (LLMs) have shown impressive abilities in solving various natural language processing tasks and are now widely offered as services. LLM services enable users to accomplish tasks without requiring specialized knowledge, simply by paying service providers. However, numerous provi…

2025

CBQ: Cross-Block Quantization for Large Language Models

ICLR 2025spotlight

Post-training quantization (PTQ) has played a pivotal role in compressing large language models (LLMs) at ultra-low costs. Although current PTQ methods have achieved promising results by addressing outliers and employing layer- or block-wise loss optimization techniques, they still suffer from signi…

Cited by 13SourcePDFScholar
2025

ChatVLA-2: Vision-Language-Action Model with Open-World Reasoning

NeurIPS 2025poster

Vision-language-action (VLA) models have emerged as the next generation of models in robotics. However, despite leveraging powerful pre-trained Vision-Language Models (VLMs), existing end-to-end VLA systems often lose key capabilities during fine-tuning as the model adapts to specific robotic tasks.…

Cited by 0SourceScholar
2025

CoA-VLA: Improving Vision-Language-Action Models via Visual-Text Chain-of-Affordance

ICCV 2025poster

Robot foundation models, particularly Vision-Language-Action (VLA) models, have garnered significant attention for their ability to enhance robot policy learning, greatly improving robot's generalization and robustness. OpenAI's recent model, O1, showcased impressive capabilities in solving complex…

Cited by 0SourcePDFScholar
2025

DISCO Balances the Scales: Adaptive Domain- and Difficulty-Aware Reinforcement Learning on Imbalanced Data

EMNLP 2025

Large Language Models (LLMs) are increasingly aligned with human preferences through Reinforcement Learning from Human Feedback (RLHF). Among RLHF methods, Group Relative Policy Optimization (GRPO) has gained attention for its simplicity and strong performance, notably eliminating the need for a lea

Cited by 0SourcePDFScholar
2025

DiffusionVLA: Scaling Robot Foundation Models via Unified Diffusion and Autoregression

ICML 2025poster

In this paper, we present DiffusionVLA, a novel framework that integrates autoregressive reasoning with diffusion policies to address the limitations of existing methods: while autoregressive Vision-Language-Action (VLA) models lack precise and robust action generation, diffusion-based policies inhe…

Cited by 0SourcePDFScholar
2025

Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration With Improved Intelligibility

ICASSP 2025accepted

Speech restoration aims at restoring full-band speech with high quality and intelligibility, considering a diverse set of distortions. MaskSR is a recently proposed generative model for this task. As other models of its kind, MaskSR attains high quality but, as we show, intelligibility can be substa…

Cited by 0SourceScholar
2025

Large Language Models and Causal Inference in Collaboration: A Comprehensive Survey

NAACL 2025findings

Causal inference has demonstrated significant potential to enhance Natural Language Processing (NLP) models in areas such as predictive accuracy, fairness, robustness, and explainability by capturing causal relationships among variables. The rise of generative Large Language Models (LLMs) has greatl…

Cited by 0SourcePDFScholar
2025

MaskTwins: Dual-form Complementary Masking for Domain-Adaptive Image Segmentation

ICML 2025poster

Recent works have correlated Masked Image Modeling (MIM) with consistency regularization in Unsupervised Domain Adaptation (UDA). However, they merely treat masking as a special form of deformation on the input images and neglect the theoretical analysis, which leads to a superficial understanding o…

2025

Multi-level Relevance Document Identifier Learning for Generative Retrieval

ACL 2025long

Generative Retrieval (GR) introduces a new information retrieval paradigm that directly generates unique document identifiers (DocIDs). The key challenge of GR lies in creating effective yet discrete DocIDs that preserve semantic relevance for similar documents while differentiating dissimilar ones.…

2025

Prompting Fairness: Integrating Causality to Debias Large Language Models

ICLR 2025poster

Large language models (LLMs), despite their remarkable capabilities, are susceptible to generating biased and discriminatory responses. As LLMs increasingly influence high-stakes decision-making (e.g., hiring and healthcare), mitigating these biases becomes critical. In this work, we propose a causa…

Cited by 0SourcePDFScholar
2025

Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation

ICASSP 2025accepted

Audio token modeling has become a powerful framework for speech synthesis, with two-stage approaches employing semantic tokens remaining prevalent. In this paper, we aim to simplify this process by introducing a semantic knowledge distillation method that enables high-quality speech generation in a…

Cited by 0SourceScholar
2025

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

NeurIPS 2025spotlight

Recent Multimodal Large Language Models (MLLMs) have typically focused on integrating visual and textual modalities, with less emphasis placed on the role of speech in enhancing interaction. However, speech plays a crucial role in multimodal dialogue systems, and implementing high-performance in bot…

Cited by 0SourcecodeScholar
2025

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

NeurIPS 2025poster

Reinforcement learning (RL) has shown great effectiveness for fine-tuning large language models (LLMs) using tasks that are challenging yet easily verifiable, such as math reasoning or code generation. However, extending this success to visual perception in vision–language models (VLMs) has been imp…

Cited by 0SourcecodeScholar
2024

AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models

EMNLP 2024finding

Large vision-language models (LVLMs) are prone to hallucinations, where certain contextual cues in an image can trigger the language module to produce overconfident and incorrect reasoning about abnormal or hypothetical objects. While some benchmarks have been developed to investigate LVLM hallucina…

2024

Cross-Dimension Affinity Distillation for 3D EM Neuron Segmentation

CVPR 2024poster

Accurate 3D neuron segmentation from electron microscopy (EM) volumes is crucial for neuroscience research. However the complex neuron morphology often leads to over-merge and over-segmentation results. Recent advancements utilize 3D CNNs to predict a 3D affinity map with improved accuracy but suffe…

2024

DrM: Mastering Visual Reinforcement Learning through Dormant Ratio Minimization

ICLR 2024spotlight

Visual reinforcement learning (RL) has shown promise in continuous control tasks. Despite its progress, current algorithms are still unsatisfactory in virtually every aspect of the performance such as sample efficiency, asymptotic performance, and their robustness to the choice of random seeds. In t…

2024

Explore Spurious Correlations at the Concept Level in Language Models for Text Classification

ACL 2024long

Language models (LMs) have achieved notable success in numerous NLP tasks, employing both fine-tuning and in-context learning (ICL) methods. While language models demonstrate exceptional performance, they face robustness challenges due to spurious correlations arising from imbalanced label distribut…

2024

GASS: Generalizing Audio Source Separation with Large-Scale Data

ICASSP 2024accepted

Universal source separation targets at separating the audio sources of an arbitrary mix, removing the constraint to operate on a specific domain like speech or music. Yet, the potential of universal source separation is limited because most existing works focus on mixes with predominantly sound even…

Cited by 0SourceScholar
2024

HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

CVPR 2024poster

We introduce "HallusionBench" a comprehensive benchmark designed for the evaluation of image-context reasoning. This benchmark presents significant challenges to advanced large visual-language models (LVLMs) such as GPT-4V(ision) Gemini Pro Vision Claude 3 and LLaVA-1.5 by emphasizing nuanced unders…

2024

Learning Multiscale Consistency for Self-Supervised Electron Microscopy Instance Segmentation

ICASSP 2024accepted

Electron microscopy (EM) images are notoriously challenging to segment due to their complex structures and lack of effective annotations. Fortunately, large-scale self-supervised pretraining offers a promising solution by allowing us to acquire prior knowledge of cell and subcellular tissue structur…

Cited by 0SourceScholar
2024

M3T: A New Benchmark Dataset for Multi-Modal Document-Level Machine Translation

NAACL 2024short

Document translation poses a challenge for Neural Machine Translation (NMT) systems. Most document-level NMT systems rely on meticulously curated sentence-level parallel data, assuming flawless extraction of text from documents along with their precise reading order. These systems also tend to disre…

2024

Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

ACL 2024long

Multimodal Large Language Models (MLLMs) have demonstrated proficiency in handling a variety of visual-language tasks. However, current MLLM benchmarks are predominantly designed to evaluate reasoning based on static information about a single image, and the ability of modern MLLMs to extrapolate fr…

2024

Multi-Stage Balanced Distillation: Addressing Long-Tail Challenges in Sequence-Level Knowledge Distillation

EMNLP 2024finding

Large language models (LLMs) have significantly advanced various natural language processing tasks, but deploying them remains computationally expensive. Knowledge distillation (KD) is a promising solution, enabling the transfer of capabilities from larger teacher LLMs to more compact student models…

2024

SmartControl: Enhancing ControlNet for Handling Rough Visual Conditions

ECCV 2024poster

"Recent text-to-image generation methods such as ControlNet have achieved remarkable success in controlling image layouts, where the generated images by the default model are constrained to strictly follow the visual conditions (e.g., depth maps). However, in practice, the conditions usually provide…

2023

A Soma Segmentation Benchmark in Full Adult Fly Brain

CVPR 2023poster

Neuron reconstruction in a full adult fly brain from high-resolution electron microscopy (EM) data is regarded as a cornerstone for neuroscientists to explore how neurons inspire intelligence. As the central part of neurons, somas in the full brain indicate the origin of neurogenesis and neural func…

2023

Beyond Image Borders: Learning Feature Extrapolation for Unbounded Image Composition

ICCV 2023poster

For improving image composition and aesthetic quality, most existing methods modulate the captured images by striking out redundant content near the image borders. However, such image cropping methods are limited in the range of image views. Some methods have been suggested to extrapolate the images…

Cited by 2PDFcodeScholar
2023

C-Disentanglement: Discovering Causally-Independent Generative Factors under an Inductive Bias of Confounder

NeurIPS 2023poster

Representation learning assumes that real-world data is generated by a few semantically meaningful generative factors (i.e., sources of variation) and aims to discover them in the latent space. These factors are expected to be causally disentangled, meaning that distinct factors are encoded into sep…

2023

Learning Cross-Representation Affinity Consistency for Sparsely Supervised Biomedical Instance Segmentation

ICCV 2023poster

Sparse instance-level supervision has recently been explored to address insufficient annotation in biomedical instance segmentation, which is easier to annotate crowded instances and better preserves instance completeness for 3D volumetric datasets compared to common semi-supervision.In this paper,…

Cited by 8PDFcodeScholar
2023

Quantitative Evidence on Overlooked Aspects of Enrollment Speaker Embeddings for Target Speaker Separation

ICASSP 2023accepted

Single channel target speaker separation (TSS) aims at extracting a speaker’s voice from a mixture of multiple talkers given an enrollment utterance of that speaker. A typical deep learning TSS framework consists of an upstream model that obtains enrollment speaker embeddings and a downstream model…

Cited by 0SourceScholar
2023

Spatially Adaptive Self-Supervised Learning for Real-World Image Denoising

CVPR 2023poster

Significant progress has been made in self-supervised image denoising (SSID) in the recent few years. However, most methods focus on dealing with spatially independent noise, and they have little practicality on real-world sRGB images with spatially correlated noise. Although pixel-shuffle downsampl…

2022

Biological Instance Segmentation with a Superpixel-Guided Graph

IJCAI 2022poster

Recent advanced proposal-free instance segmentation methods have made significant progress in biological images. However, existing methods are vulnerable to local imaging artifacts and similar object appearances, resulting in over-merge and over-segmentation. To reduce these two kinds of errors, we…

2022

Tuformer: Data-driven Design of Transformers for Improved Generalization or Efficiency

ICLR 2022poster

Transformers are neural network architectures that achieve remarkable performance in many areas. However, the core component of Transformers, multi-head self-attention (MHSA), is mainly derived from heuristics, and the interactions across its components are not well understood. To address the proble…

Cited by 6SourcePDFScholar
2021

Long-Range Modeling of Source Code Files with eWASH: Extended Window Access by Syntax Hierarchy

EMNLP 2021main

Statistical language modeling and translation with transformers have found many successful applications in program understanding and generation tasks, setting high benchmarks for tools in modern software development environments. The finite context window of these neural models means, however, that…

2021

Occluded Video Instance Segmentation: Dataset and ICCV 2021 Challenge

NeurIPS 2021poster

Although deep learning methods have achieved advanced video object recognition performance in recent years, perceiving heavily occluded objects in a video is still a very challenging task. To promote the development of occlusion understanding, we collect a large-scale dataset called OVIS for video i…

Cited by 16SourceScholar
2020

Self-Supervised Learning for Alignment of Objects and Sound

ICRA 2020poster

The sound source separation problem has many useful applications in the field of robotics, such as human-robot interaction, scene understanding, etc. However, it remains a very challenging problem. In this paper, we utilize both visual and audio information of videos to perform the sound source sepa…

Cited by 5SourceScholar
2020

Speaker-Aware Target Speaker Enhancement by Jointly Learning with Speaker Embedding Extraction

ICASSP 2020accepted

Deep learning based speech separation approaches have received great interest, among which the recent speaker-aware speech enhancement methods are promising for solving difficulties such as arbitrary source permutation and unknown number of sources. In this paper, we propose a novel training framewo…

Cited by 0SourceScholar
2017

A modulation feature set for robust Automatic Speech Recognition in additive noise and reverberation

ICASSP 2017accepted

In this paper, a feature set referred to as Discrete Cosine Series (DCS) is proposed for noise robust Automatic Speech Recognition (ASR). Unlike many other robust algorithms which use various forms of “long term” processing, DCS uses a small frame spacing to facilitate separating speech from noise a…

Cited by 0SourceScholar
2015

A unified framework for filterbank and time-frequency basis vectors in ASR frontends

ICASSP 2015accepted

For many years, filterbank have been widely used as one step of frontend feature extraction for Automatic Speech Recognition (ASR). In this paper, we propose a unified framework for ASR frontends, by first moving the nonlinear amplitude scaling, and then combining the filterbank weights with the cos…

Cited by 0SourceScholar