← Search

Hualei Wang

4 accepted papers

2026

Audio-Thinker: Guiding Large Audio Language Model When and How to Think via Reinforcement Learning

AAAI 2026technical

Recent advancements in large language models, multimodal large language models, and large audio language models (LALMs) have significantly improved their reasoning capabilities through reinforcement learning utilizing rule-based rewards. However, the explicit reasoning process has not yet yielded su

Cited by 0SourcePDFScholar
2026

Enhancing Stability and Fidelity for Zero-Shot TTS with a Multi-Level Evaluator

AAAI 2026technical

Recent advances in zero-shot text-to-speech (TTS), driven by language models, diffusion models and masked generation, have achieved impressive naturalness in speech synthesis. Nevertheless, stability and fidelity remain key challenges, manifesting as mispronunciations, audible noise, and quality deg

Cited by 0SourcePDFScholar
2026

Listening Between the Frames: Bridging Temporal Gaps in Large Audio-Language Models

AAAI 2026technical

Recent Large Audio-Language Models (LALMs) exhibit impressive capabilities in understanding audio content for conversational QA tasks. However, these models struggle to accurately understand timestamps for temporal localization (e.g., Temporal Audio Grounding) and are restricted to short audio perce

Cited by 0SourcePDFScholar
2025

SleepSMC: Ubiquitous Sleep Staging via Supervised Multimodal Coordination

ICLR 2025poster

Sleep staging is critical for assessing sleep quality and tracking health. Polysomnography (PSG) provides comprehensive multimodal sleep-related information, but its complexity and impracticality limit its practical use in daily and ubiquitous monitoring. Conversely, unimodal devices offer more conv…

Cited by 0SourcePDFScholar