← Search

Yong Ren

15 accepted papers

2026

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability

ICML 2026poster

Recent advances in Large Audio Language Models (LALMs) have extended Text-to-Speech (TTS) to interactive role-play scenarios, which demand high expressiveness and strict adherence to role-play instructions. However, existing models struggle to maintain stylistic consistency with character profiles a…

Cited by 0SourcecodeScholar
2026

OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech

ICASSP 2026oral

Instruct Text-to-Speech (InstructTTS) leverages natural language descriptions as style prompts to guide speech synthesis. However, existing InstructTTS methods mainly rely on a direct combination of audio-related labels or their diverse rephrasings, making it difficult to handle flexible, high-level…

Cited by 0SourcePDFScholar
2025

AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models

ICML 2025oral

The emergence of multimodal large language models (MLLMs) advances multimodal emotion recognition (MER) to the next level—from naive discriminative tasks to complex emotion understanding with advanced video understanding abilities and natural language description. However, the current community suff…

2025

Enhancing Multimodal Continual Instruction Tuning with BranchLoRA

ACL 2025long

Multimodal Continual Instruction Tuning (MCIT) aims to finetune Multimodal Large Language Models (MLLMs) to continually align with human intent across sequential tasks. Existing approaches often rely on the Mixture-of-Experts (MoE) LoRA framework to preserve previous instruction alignments. However,…

Cited by 0SourcePDFScholar
2025

Progressive LoRA for Multimodal Continual Instruction Tuning

ACL 2025finding

Multimodal Continual Instruction Tuning (MCIT) empowers Multimodal Large Language Models (MLLMs) to adapt to ever-evolving requirements without continuous costly retraining. However, MCIT faces challenges in mitigating Catastrophic Forgetting (CF) and enhancing Knowledge Transfer (KT). Existing work…

Cited by 0SourcePDFScholar
2025

Region-Based Optimization in Continual Learning for Audio Deepfake Detection

AAAI 2025technical

Rapid advancements in speech synthesis and voice conversion bring convenience but also new security risks, creating an urgent need for effective audio deepfake detection. Although current models perform well, their effectiveness diminishes when confronted with the diverse and evolving nature of real…

2025

STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment

ICASSP 2025accepted

Visual and auditory perception are two crucial ways humans experience the world. Text-to-video generation has made remarkable progress over the past year, but the absence of harmonious audio in generated video limits its broader applications. In this paper, we propose Semantic and Temporal Aligned V…

Cited by 0SourceScholar
2025

WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification

ICASSP 2025accepted

Recent advances in speech spoofing necessitate stronger verification mechanisms in neural speech codecs to ensure authenticity. Current methods embed numerical watermarks before compression and extract them from reconstructed speech for verification, but face limitations such as separate training pr…

Cited by 0SourceScholar
2024

CReStyler: Text-Guided Single Image Style Transfer Method Based on CNN and Restormer

ICASSP 2024accepted

Text-guided image style transfer methods have gradually become a research hotspot. However, existing text-guided style transfer method suffers from content information missing and artifacts in the generated stylized images. Therefore, we propose CReStyler, a text-guided image style method based on t…

Cited by 0SourceScholar
2024

Fewer-Token Neural Speech Codec with Time-Invariant Codes

ICASSP 2024accepted

Language model based text-to-speech (TTS) models, like VALL-E, have gained attention for their outstanding in-context learning capability in zero-shot scenarios. Neural speech codec is a critical component of these models, which can convert speech into discrete token representations. However, excess…

Cited by 0SourceScholar
2024

SpikeVoice: High-Quality Text-to-Speech Via Efficient Spiking Neural Network

ACL 2024long

Brain-inspired Spiking Neural Network (SNN) has demonstrated its effectiveness and efficiency in vision, natural language, and speech understanding tasks, indicating their capacity to “see”, “listen”, and “read”. In this paper, we design SpikeVoice, which performs high-quality Text-To-Speech (TTS) v…

2021

Improving Generative Moment Matching Networks with Distribution Partition

AAAI 2021technical

Generative moment matching networks (GMMN) present a theoretically sound approach to learning deep generative mod-els. However, such methods are typically limited by the high sample complexity, thereby impractical in generating complex data. In this paper, we present a new strategy to train GMMN wit…

2018

Smooth Neighbors on Teacher Graphs for Semi-Supervised Learning

CVPR 2018poster

The recently proposed self-ensembling methods have achieved promising results in deep semi-supervised learning, which penalize inconsistent predictions of unlabeled data under different perturbations. However, they only consider adding perturbations to each single data point, while ignoring the conn…