← Search

Huadai Liu

17 accepted papers

2026

MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models

ICLR 2026poster

Recent advances in multimodal large language models (MLLMs) have catalyzed transformative progress in affective computing, enabling models to exhibit emergent emotional intelligence. Despite substantial methodological progress, current emotional benchmarks remain limited, as it is still unknown: (a)…

Cited by 0SourcecodeScholar
2026

PrismAudio: Decomposed Chain-of-Thought and Multi-dimensional Rewards for Video-to-Audio Generation

ICLR 2026poster

Video-to-Audio (V2A) generation requires balancing four critical perceptual dimensions: semantic consistency, audio-visual temporal synchrony, aesthetic quality, and spatial accuracy; yet existing methods suffer from objective entanglement that conflates competing goals in single loss functions and…

Cited by 0SourcecodeScholar
2026

STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation

ICML 2026poster

Continuous Variational Autoencoders (VAEs) serve as the fundamental continuous tokenizer for modern neural audio generation systems, enabling high-fidelity reconstruction while providing a compact, smooth latent space for downstream generative priors. However, continuous VAEs face a fundamental conf…

Cited by 0SourcecodeScholar
2026

SpaceVista: All-Scale Visual Spatial Reasoning from mm to km

ICML 2026poster

With the current surge in spatial reasoning, researchers have made significant progress in understanding indoor scenes, but still struggle with more diverse applications. This paper aims to advance all-scale spatial reasoning by tackling two key challenges: 1) the heavy reliance on indoor 3D scans a…

Cited by 0SourcecodeScholar
2025

Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation

ICLR 2025spotlight

Recently, diffusion models have achieved great success in mono-channel audio generation. However, when it comes to stereo audio generation, the soundscapes often have a complex scene of multiple objects and directions. Controlling stereo audio with spatial contexts remains challenging due to high da…

Cited by 2SourcePDFScholar
2025

Build LLM-Based Zero-Shot Streaming TTS System with Cosyvoice

ICASSP 2025accepted

LLM-based text-to-speech(TTS) system has becoming the new trend and SOTA due to its high naturalness and zero-shot capability. However, it relies heavily on training data, usually requires at least thousands hours of labeled audio. In this report, we describe how to use pretrained CosyVoice model, t…

Cited by 0SourceScholar
2025

Fast Adaptation of Pretrained Speaker Verification System for Source Speaker Tracking

ICASSP 2025accepted

Traditional speaker verification system aims at distinguish speaker identity in real world audio, and has achieved satisfying performance in many scenarios. However, it is also very vulnerable, and can be easily attacked by voice anonymization system. In this report, we describe how to fast adapt a…

Cited by 0SourceScholar
2025

FlashAudio: Rectified Flow for Fast and High-Fidelity Text-to-Audio Generation

ACL 2025long

Recent advancements in latent diffusion models (LDMs) have markedly enhanced text-to-audio generation, yet their iterative sampling processes impose substantial computational demands, limiting practical deployment. While recent methods utilizing consistency-based distillation aim to achieve few-step…

2025

OmniAudio: Generating Spatial Audio from 360-Degree Video

ICML 2025poster

Traditional video-to-audio generation techniques primarily focus on perspective video and non-spatial audio, often missing the spatial cues necessary for accurately representing sound sources in 3D environments. To address this limitation, we introduce a novel task, \textbf{360V2SA}, to generate spa…

2025

ThinkSound: Chain-of-Thought Reasoning in Multimodal LLMs for Audio Generation and Editing

NeurIPS 2025poster

While end-to-end video-to-audio generation has greatly improved, producing high-fidelity audio that authentically captures the nuances of visual content remains challenging. Like professionals in the creative industries, this generation requires sophisticated reasoning about items such as visual dyn…

Cited by 0SourcecodeScholar
2024

AntCritic: Argument Mining for Free-Form and Visually-Rich Financial Comments

COLING 2024main

Argument mining aims to detect all possible argumentative components and identify their relationships automatically. As a thriving task in natural language processing, there has been a large amount of corpus for academic study and application development in this field. However, the research in this…

2024

Extending Multi-modal Contrastive Representations

NeurIPS 2024poster

Multi-modal contrastive representation (MCR) of more than three modalities is critical in multi-modal learning. Although recent methods showcase impressive achievements, the high dependence on large-scale, high-quality paired data and the expensive training costs limit their further development. Ins…

2024

Wav2SQL: Direct Generalizable Speech-To-SQL Parsing

ACL 2024findings

We release a multi-accent dataset and propose speech-programming and gradient reversal classifier to improve the generalization.Abstract: Speech-to-SQL (S2SQL) aims to convert spoken questions into SQL queries given relational databases, which has been traditionally implemented in a cascaded manner…

Cited by 3SourcePDFScholar
2023

AV-TranSpeech: Audio-Visual Robust Speech-to-Speech Translation

ACL 2023long

Direct speech-to-speech translation (S2ST) aims to convert speech from one language into another, and has demonstrated significant progress to date. Despite the recent success, current S2ST models still suffer from distinct degradation in noisy environments and fail to translate visual speech (i.e.,…

2023

MixSpeech: Cross-Modality Self-Learning with Audio-Visual Stream Mixup for Visual Speech Translation and Recognition

ICCV 2023poster

Multi-media communications facilitate global interaction among people. However, despite researchers exploring cross-lingual translation techniques such as machine translation and audio speech translation to overcome language barriers, there is still a shortage of cross-lingual studies on visual spee…

Cited by 25PDFcodeScholar
2023

RMSSinger: Realistic-Music-Score based Singing Voice Synthesis

ACL 2023findings

We are interested in a challenging task, Realistic-Music-Score based Singing Voice Synthesis (RMS-SVS). RMS-SVS aims to generate high-quality singing voices given realistic music scores with different note types (grace, slur, rest, etc.). Though significant progress has been achieved, recent singing…

2023

TranSpeech: Speech-to-Speech Translation With Bilateral Perturbation

ICLR 2023poster

Direct speech-to-speech translation (S2ST) with discrete units leverages recent progress in speech representation learning. Specifically, a sequence of discrete representations derived in a self-supervised manner are predicted from the model and passed to a vocoder for speech reconstruction, while s…