← Search

Tianxin Xie

5 accepted papers

2026

AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching

ICML 2026poster

REPresentation Alignment (REPA) improves the training of generative flow models by aligning intermediate hidden states with pretrained teacher features, but its effectiveness in token-conditioned audio Flow Matching critically depends on the choice of supervised layers, which is typically made heuri…

Cited by 0SourceScholar
2026

Boosting ASR Robustness via Test-Time Reinforcement Learning with Audio-Text Semantic Rewards

AAAI 2026technical

Recently, Automatic Speech Recognition (ASR) systems (e.g., Whisper) have achieved remarkable accuracy improvements but remain highly sensitive to real-world unseen data (data with large distribution shifts), including noisy environments and diverse accents. To address this issue, test-time adaptati

Cited by 0SourcePDFScholar
2026

Resp-Agent: An Agent-Based System for Multimodal Respiratory Sound Generation and Disease Diagnosis

ICLR 2026poster

Deep learning-based respiratory auscultation is currently hindered by two fundamental challenges: (i) inherent information loss, as converting signals into spectrograms discards transient acoustic events and clinical context; (ii) limited data availability, exacerbated by severe class imbalance. To…

Cited by 0SourcecodeScholar
2025

Inter- and Intra-Sentence Cuer-Invariant Representation Learning for Generalizable Cued Speech Recognition

ICASSP 2025accepted

Cued Speech (CS) is a visual coding system that combines lip movements and hand gestures to represent spoken languages for hearing-impaired people. Automatic Cued Speech Recognition (ACSR) is an emerging research topic, but the cuer (i.e., people who perform CS) generalization problem of ACSR remain…

Cited by 0SourceScholar
2025

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey

EMNLP 2025

Text-to-speech (TTS) has advanced from generating natural-sounding speech to enabling fine-grained control over attributes like emotion, timbre, and style. Driven by rising industrial demand and breakthroughs in deep learning, e.g., diffusion and large language models (LLMs), controllable TTS has be