← Search

Fei Tao

7 accepted papers

2026

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis

AAAI 2026technical

Diffusion-based text-to-speech (TTS) systems have made remarkable progress in zero-shot speech synthesis, yet optimizing all components for perceptual metrics remains challenging. Prior work with DMOSpeech demonstrated direct metric optimization for speech generation components, but duration predict

Cited by 0SourcePDFScholar
2026

FrameOracle: Learning What to See and How Much to See in Videos

ICML 2026poster

Vision-language models (VLMs) advance video understanding but operate under tight computational budgets, making performance dependent on selecting a small, high-quality subset of frames. Existing frame sampling strategies, such as uniform or fixed-budget selection, fail to adapt to variations in con…

Cited by 3SourceScholar
2026

When Attributes Disagree: Gradient Conflict in Image Aesthetic Assessment

ICML 2026spotlight

Image Aesthetic Assessment (IAA) predicts an image’s overall aesthetic score, yet aesthetic is influenced by multiple attributes whose relative importance varies with image content and usage scenarios. Under end-to-end training with only overall-score supervision, attribute signals are blended, whic…

Cited by 0SourceScholar
2022

Progressive Teacher-Student Training Framework for Music Tagging

ICASSP 2022accepted

Music tagging is the task of predicting multiple tags of a music excerpt, and plays an important role in modern music recommendation systems. To obtain superior performance, recent approaches of music tagging focus on developing sophisticated models or exploiting additional multi-modal information.…

Cited by 0SourceScholar
2018

An Ensemble Framework of Voice-Based Emotion Recognition System for Films and TV Programs

ICASSP 2018accepted

Employing voice-based emotion recognition function in artificial intelligence (AI) product will improve the user experience. Most of researches that have been done only focus on the speech collected under controlled conditions. The scenarios evaluated in these research were well controlled. The conv…

Cited by 0SourceScholar