← Search

Yixuan Zhou

13 accepted papers

2026

Hierarchical Semantic-Acoustic Modeling via Semi-Discrete Residual Representations for Expressive End-to-End Speech Synthesis

ICLR 2026poster

Generative models for speech synthesis face a fundamental trade-off: discrete tokens ensure stability but sacrifice expressivity, while continuous signals retain acoustic richness but suffer from error accumulation due to task entanglement. This challenge has driven the field towards multi-stage pip…

Cited by 0SourcecodeScholar
2026

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment

ICML 2026poster

Neural speech codecs based on Vector-Quantized VAEs (VQ-VAEs) are core audio tokenizers for speech LLMs, yet their reconstruction fidelity is bottlenecked by quantization error. Instead of modifying the quantizer or increasing model capacity—common approaches that complicate downstream language mode…

Cited by 0SourceScholar
2026

UTTG: A Universal Teleoperation Framework Via Online Trajectory Generation

ICRA 2026poster

Teleoperation is crucial for hazardous environment operations and serves as a key tool for collecting expert demonstrations in robot learning. However, existing methods face robotic hardware dependency and control frequency mismatches between teleoperation devices and robotic platforms. Our approach…

Cited by 0codeScholar
2025

DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models

ICASSP 2025accepted

Conversational speech synthesis (CSS) aims to synthesize both contextually appropriate and expressive speech, and considerable efforts have been made to enhance the understanding of conversational context. However, existing CSS systems are limited to deterministic prediction, overlooking the diversi…

Cited by 0SourceScholar
2025

TAU-106K: A New Dataset for Comprehensive Understanding of Traffic Accident

ICLR 2025poster

Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in general visual understanding tasks. However, their potential for high-level, fine-grained comprehension, such as anomaly understanding, remains unexplored. Focusing on traffic accidents, a critical and practical sce…

2025

Zero-Shot Temporal Interaction Localization for Egocentric Videos

IROS 2025

Locating human-object interaction (HOI) actions within video serves as the foundation for multiple downstream tasks, such as human behavior analysis and human-robot skill transfer. Current temporal action localization methods typically rely on annotated action and object categories of interactions f

Cited by 2SourcecodeScholar
2024

Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts

ICASSP 2024accepted

Zero-shot text-to-speech (TTS) synthesis aims to clone any unseen speaker’s voice without adaptation parameters. By quantizing speech waveform into discrete acoustic tokens and modeling these tokens with the language model, recent language model-based TTS models show zero-shot speaker adaptation cap…

Cited by 0SourceScholar
2024

SongCreator: Lyrics-based Universal Song Generation

NeurIPS 2024poster

Music is an integral part of human culture, embodying human intelligence and creativity, of which songs compose an essential part. While various aspects of song generation have been explored by previous works, such as singing voice, vocal composition and instrumental arrangement, etc., generating so…

2023

Context-Aware Coherent Speaking Style Prediction with Hierarchical Transformers for Audiobook Speech Synthesis

ICASSP 2023accepted

Recent advances in text-to-speech have significantly improved the expressiveness of synthesized speech. However, it is still challenging to generate speech with contextually appropriate and coherent speaking style for multi-sentence text in audiobooks. In this paper, we propose a context-aware coher…

Cited by 0SourceScholar
2023

ImbSAM: A Closer Look at Sharpness-Aware Minimization in Class-Imbalanced Recognition

ICCV 2023poster

Class imbalance is a common challenge in real-world recognition tasks, where the majority of classes have few samples, also known as tail classes. We address this challenge with the perspective of generalization and empirically find that the promising Sharpness-Aware Minimization (SAM) fails to addr…

Cited by 17PDFcodeScholar
2022

A Character-Level Span-Based Model for Mandarin Prosodic Structure Prediction

ICASSP 2022accepted

The accuracy of prosodic structure prediction is crucial to the naturalness of synthesized speech in Mandarin text-to-speech system, but now is limited by widely-used sequence-to-sequence framework and error accumulation from previous word segmentation results. In this paper, we propose a span-based…

Cited by 0SourceScholar
2022

Towards Expressive Speaking Style Modelling with Hierarchical Context Information for Mandarin Speech Synthesis

ICASSP 2022accepted

Previous works on expressive speech synthesis mainly focus on current sentence. The context in adjacent sentences is neglected, resulting in inflexible speaking style for the same text, which lacks speech variations. In this paper, we propose a hierarchical framework to model speaking style from con…

Cited by 0SourceScholar
2021

Syntactic Representation Learning For Neural Network Based TTS with Syntactic Parse Tree Traversal

ICASSP 2021accepted

Syntactic structure of a sentence text is correlated with the prosodic structure of the speech that is crucial for improving the prosody and naturalness of a text-to-speech (TTS) system. Nowadays TTS systems usually try to incorporate syntactic structure information with manually designed features b…

Cited by 0SourceScholar