← Search

Xiaobin Zhuang

5 accepted papers

2026

AudioTrust: Benchmarking The Multifaceted Trustworthiness of Audio Large Language Models

ICLR 2026poster

The rapid development and widespread adoption of Audio Large Language Models (ALLMs) require a rigorous assessment of their trustworthiness. However, existing evaluation frameworks, primarily designed for text, are not equipped to handle the unique vulnerabilities introduced by audio’s acoustic prop…

Cited by 0SourcecodeScholar
2025

DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation

ICML 2025poster

Several recent studies have attempted to autoregressively generate continuous speech representations without discrete speech tokens by combining diffusion and autoregressive models, yet they often face challenges with excessive computational loads or suboptimal outcomes. In this work, we propose Dif…

Cited by 1SourcePDFScholar
2025

Sound-VECaps: Improving Audio Generation with Visually Enhanced Captions

ICASSP 2025accepted

Generative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performance degradation. We hypothesize that this problem stems from the simplicity and scarcity of the training data. This work…

Cited by 0SourceScholar
2025

Sounding that Object: Interactive Object-Aware Image to Audio Generation

ICML 2025poster

Generating accurate sounds for complex audio-visual scenes is challenging, especially in the presence of multiple objects and sound sources. In this paper, we propose an interactive object-aware audio generation model that grounds sound generation in user-selected visual objects within images. Our m…

Cited by 0SourcePDFScholar
2021

Litesing: Towards Fast, Lightweight and Expressive Singing Voice Synthesis

ICASSP 2021accepted

LiteSing proposed in this paper is a high-quality singing voice synthesis (SVS) system, which is fast, lightweight and expressive. This model mainly stacks several non-autoregressive WaveNet blocks in the encoder and decoder under a generative adversarial architecture, predicts full conditions from…

Cited by 0SourceScholar