← Search

Yi Ren

86 accepted papers

2026

Arm-Aware Guided Dexterous Grasp Generation With Arm-Agnostic Grasp Models

RA-L 2026

Dexterous grasp generation that considers armrelated constraints is crucial in real-world scenarios involving armenvironment collision avoidance, workspace boundary grasps, and consecutive grasping. Existing hand-centric grasp models, which primarily focus on the floating hand's pose, are insufficie

Cited by 0SourcecodeScholar
2026

CPQS-Tuning: A Model Self-Perception-Based Data Filtering Algorithm for Efficient Instruction Fine-Tuning

ICLR 2026poster

Instruction fine-tuning is a key technique for enhancing the performance of large language models (LLMs), but low-quality and redundant data often hinder its effectiveness. Recent studies suggest that filtering a small amount of high-quality data for instruction fine-tuning can achieve faster and mo…

Cited by 0SourcecodeScholar
2026

CoorGrasp: Coordinated Contact Control for Adaptive Dexterous Grasping under Uncertainty

ICRA 2026poster

While recent research has focused heavily on dexterous grasp pose generation, less attention has been devoted to the execution of planned grasps. Under shape and position uncertainty, open-loop execution often yields uncoordinated contacts, causing undesired in-hand object motion and even grasp fail…

Cited by 0codeScholar
2026

Data-Efficient Real-Time Control of an Artificial-Muscle-Driven Continuum Robot With Physics-Informed Koopman Operator

RA-L 2026

Data-driven techniques enable modeling of nonlinear continuum robots directly from input-output data without relying on physics-based modeling. Among these techniques, the Koopman method has been particularly successful as it enables the application of linear optimal control techniques such as model

Cited by 0SourceScholar
2026

InfinityHuman: Towards Long-Term Audio-Driven Human Animation

CVPR 2026

Audio-driven human animation has attracted wide attention thanks to its practical applications. However, critical challenges remain in generating high-resolution, long-duration videos with consistent appearance and natural hand motions. Existing methods extend videos using overlapping motion frames

Cited by 0SourcecodeScholar
2026

On Group Relative Policy Optimization Collapse in Agent Search: The Lazy Likelihood-Displacement

ICML 2026poster

Tool-integrated (TI) reinforcement learning (RL) enables large language models (LLMs) to perform multi-step reasoning by interacting with external tools such as search engines and retrievers. Group Relative Policy Optimization (GRPO), exemplified by the recent Search-R1, offers fast convergence and …

Cited by 0SourceScholar
2026

Solving Football by Exploiting Equilibrium Structure of 2p0s Differential Games with One-Sided Information

ICLR 2026poster

For a two-player imperfect-information extensive-form game (IIEFG) with $K$ time steps and a player action space of size $U$, the game tree complexity is $U^{2K}$, causing existing IIEFG solvers to struggle with large or infinite $(U,K)$, e.g., differential games with continuous action spaces. To pa…

Cited by 0SourcecodeScholar
2026

Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement Learning

ICLR 2026poster

Reinforcement learning with verifiable rewards has significantly advanced the reasoning capabilities of large language models, yet how to explicitly steer training toward exploration or exploitation remains an open problem. We introduce Token Hidden Reward (THR), a token-level metric that quantifies…

Cited by 0SourceScholar
2025

Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis

ACL 2025finding

Conversational Speech Synthesis (CSS) aims to align synthesized speech with the emotional and stylistic context of user-agent interactions to achieve empathy. Current generative CSS models face interpretability limitations due to insufficient emotional perception and redundant discrete speech coding…

2025

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization

NeurIPS 2025poster

Reinforcement learning (RL) has become popular in enhancing the reasoning capabilities of large language models (LLMs), with Group Relative Policy Optimization (GRPO) emerging as a widely used algorithm in recent systems. Despite GRPO's widespread adoption, we identify a previously unrecognized phen…

Cited by 0SourceScholar
2024

A Robotic Solution to Peg in/out Hole Tasks with Latching Requirements

RA-L 2024

Connectors with latches, such as LC fiber connectors, RJ45 network cable connectors, and certain electronic connectors, have significant automation requirements for connection and disconnection. To accomplish not only the peg-in-hole but also the peg-out-hole tasks, careful consideration must be giv

Cited by 7SourceScholar
2024

AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head

AAAI 2024technical

Large language models (LLMs) have exhibited remarkable capabilities across a variety of domains and tasks, challenging our understanding of learning and cognition. Despite the recent success, current LLMs are not capable of processing complex audio information or conducting spoken conversations (lik…

2024

Bias Amplification in Language Model Evolution: An Iterated Learning Perspective

NeurIPS 2024poster

With the widespread adoption of Large Language Models (LLMs), the prevalence of iterative interactions among these models is anticipated to increase. Notably, recent advancements in multi-round on-policy self-improving methods allow LLMs to generate new examples for training subsequent models. At th…

2024

Emotion Rendering for Conversational Speech Synthesis with Heterogeneous Graph-Based Context Modeling

AAAI 2024technical

Conversational Speech Synthesis (CSS) aims to accurately express an utterance with the appropriate prosody and emotional inflection within a conversational setting. While recognising the significance of CSS task, the prior studies have not thoroughly investigated the emotional expressiveness problem…

2024

Hacking Task Confounder in Meta-Learning

IJCAI 2024poster

Meta-learning enables rapid generalization to new tasks by learning knowledge from various tasks. It is intuitively assumed that as the training progresses, a model will acquire richer knowledge, leading to better generalization performance. However, our experiments reveal an unexpected result: ther…

2024

Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis

ICLR 2024poster

Zero-shot text-to-speech (TTS) aims to synthesize voices with unseen speech prompts, which significantly reduces the data and computation requirements for voice cloning by skipping the fine-tuning process. However, the prompting mechanisms of zero-shot TTS still face challenges in the following aspe…

2024

MimicTalk: Mimicking a personalized and expressive 3D talking face in minutes

NeurIPS 2024poster

Talking face generation (TFG) aims to animate a target identity's face to create realistic talking videos. Personalized TFG is a variant that emphasizes the perceptual identity similarity of the synthesized result (from the perspective of appearance and talking style). While previous works typically…

2024

ModaLink: Unifying Modalities for Efficient Image-to-PointCloud Place Recognition

IROS 2024poster

Place recognition is an important task for robots and autonomous cars to localize themselves and close loops in pre-built maps. While single-modal sensor-based methods have shown satisfactory performance, cross-modal place recognition that retrieving images from a point-cloud database remains a chal…

Cited by 3SourcecodeScholar
2024

Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis

ICLR 2024spotlight

One-shot 3D talking portrait generation aims to reconstruct a 3D avatar from an unseen image, and then animate it with a reference video or audio to generate a talking portrait video. The existing methods fail to simultaneously achieve the goals of accurate 3D avatar reconstruction and stable talkin…

2024

State-Constrained Zero-Sum Differential Games with One-Sided Information

ICML 2024poster

We study zero-sum differential games with state constraints and one-sided information, where the informed player (Player 1) has a categorical payoff type unknown to the uninformed player (Player 2). The goal of Player 1 is to minimize his payoff without violating the constraints, while that of Playe…

2024

lpNTK: Better Generalisation with Less Data via Sample Interaction During Learning

ICLR 2024poster

Although much research has been done on proposing new models or loss functions to improve the generalisation of artificial neural networks (ANNs), less attention has been directed to the impact of the training data on generalisation. In this work, we start from approximating the interaction between…

Cited by 2SourcePDFScholar
2023

AV-TranSpeech: Audio-Visual Robust Speech-to-Speech Translation

ACL 2023long

Direct speech-to-speech translation (S2ST) aims to convert speech from one language into another, and has demonstrated significant progress to date. Despite the recent success, current S2ST models still suffer from distinct degradation in noisy environments and fail to translate visual speech (i.e.,…

2023

Approximating Discontinuous Nash Equilibrial Values of Two-Player General-Sum Differential Games

ICRA 2023poster

Finding Nash equilibrial policies for two-player differential games requires solving Hamilton-Jacobi-Isaacs (HJI) PDEs. Self-supervised learning has been used to approximate solutions of such PDEs while circumventing the curse of dimensionality. However, this method fails to learn discontinuous PDE…

Cited by 8SourceScholar
2023

Attributing Image Generative Models using Latent Fingerprints

ICML 2023poster

Generative models have enabled the creation of contents that are indistinguishable from those taken from nature. Open-source development of such models raised concerns about the risks of their misuse for malicious purposes. One potential risk mitigation strategy is to attribute generative models via…

2023

CLAPSpeech: Learning Prosody from Text Context with Contrastive Language-Audio Pre-Training

ACL 2023long

Improving text representation has attracted much attention to achieve expressive text-to-speech (TTS). However, existing works only implicitly learn the prosody with masked token reconstruction tasks, which leads to low training efficiency and difficulty in prosody modeling. We propose CLAPSpeech, a…

2023

FREDIS: A Fusion Framework of Refinement and Disambiguation for Unreliable Partial Label Learning

ICML 2023poster

To reduce the difficulty of annotation, partial label learning (PLL) has been widely studied, where each example is ambiguously annotated with a set of candidate labels instead of the exact correct label. PLL assumes that the candidate label set contains the correct label, which induces disambiguati…

Cited by 7SourcePDFScholar
2023

FastDiff 2: Revisiting and Incorporating GANs and Diffusion Models in High-Fidelity Speech Synthesis

ACL 2023findings

Generative adversarial networks (GANs) and denoising diffusion probabilistic models (DDPMs) have recently achieved impressive performances in image and audio synthesis. After revisiting their success in conditional speech synthesis, we find that 1) GANs sacrifice sample diversity for quality and spe…

2023

FluentSpeech: Stutter-Oriented Automatic Speech Editing with Context-Aware Diffusion Models

ACL 2023findings

Stutter removal is an essential scenario in the field of speech editing. However, when the speech recording contains stutters, the existing text-based speech editing approaches still suffer from: 1) the over-smoothing problem in the edited speech; 2) lack of robustness due to the noise introduced by…

2023

GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis

ICLR 2023poster

Generating photo-realistic video portraits with arbitrary speech audio is a crucial problem in film-making and virtual reality. Recently, several works explore the usage of neural radiance field (NeRF) in this task to improve 3D realness and image fidelity. However, the generalizability of previous…

2023

Improving Compositional Generalization using Iterated Learning and Simplicial Embeddings

NeurIPS 2023poster

Compositional generalization, the ability of an agent to generalize to unseen combinations of latent factors, is easy for humans but hard for deep neural networks. A line of research in cognitive science has hypothesized a process, "iterated learning," to help explain how human language developed th…

Cited by 11SourcePDFScholar
2023

MUG: A General Meeting Understanding and Generation Benchmark

ICASSP 2023accepted

Listening to long video/audio recordings from video conferencing and online courses for acquiring information is extremely inefficient. Even after ASR systems transcribe recordings into long-form spoken language documents, reading ASR transcripts only partly speeds up seeking information. It has bee…

Cited by 0SourceScholar
2023

Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models

ICML 2023poster

Large-scale multimodal generative modeling has created milestones in text-to-image and text-to-video generation. Its application to audio still lags behind for two main reasons: the lack of large-scale datasets with high-quality text-audio pairs, and the complexity of modeling long continuous audio…

2023

Overview of the ICASSP 2023 General Meeting Understanding and Generation Challenge (MUG)

ICASSP 2023accepted

ICASSP2023 General Meeting Understanding and Generation Challenge (MUG) focuses on prompting a wide range of spoken language processing (SLP) research on meeting transcripts, as SLP applications are critical to improve users’ efficiency in grasping important information in meetings. MUG includes fiv…

Cited by 0SourceScholar
2023

Performance Comparison of Typical Physics Engines Using Robot Models With Multiple Joints

RA-L 2023

Physics engines are essential components in simulating complex robotic systems. The accuracy and computational speed of these engines are crucial for reliable real-time simulation. This letter comprehensively evaluates the performance of five common physics engines, i.e., ODE, Bullet, DART, MuJoCo,

Cited by 9SourceScholar
2023

Prosody-TTS: Improving Prosody with Masked Autoencoder and Conditional Diffusion Model For Expressive Text-to-Speech

ACL 2023findings

Expressive text-to-speech aims to generate high-quality samples with rich and diverse prosody, which is hampered by dual challenges: 1) prosodic attributes in highly dynamic voices are difficult to capture and model without intonation; and 2) highly multimodal prosodic representations cannot be well…

2023

TranSpeech: Speech-to-Speech Translation With Bilateral Perturbation

ICLR 2023poster

Direct speech-to-speech translation (S2ST) with discrete units leverages recent progress in speech representation learning. Specifically, a sequence of discrete representations derived in a self-supervised manner are predicted from the model and passed to a vocoder for speech reconstruction, while s…

2023

Unsupervised Video Domain Adaptation for Action Recognition: A Disentanglement Perspective

NeurIPS 2023poster

Unsupervised video domain adaptation is a practical yet challenging task. In this work, for the first time, we tackle it from a disentanglement view. Our key idea is to handle the spatial and temporal domain divergence separately through disentanglement. Specifically, we consider the generation of c…

2023

VarietySound: Timbre-Controllable Video to Sound Generation Via Unsupervised Information Disentanglement

ICASSP 2023accepted

Video-to-sound generation aims to generate realistic and natural sound given a video input. However, previous video-to-sound generation methods can only generate a random or average timbre without any controls of the generated sound timbre, leading to the problem that people cannot obtain the desire…

Cited by 0SourceScholar
2022

A Study of Syntactic Multi-Modality in Non-Autoregressive Machine Translation

NAACL 2022long

It is difficult for non-autoregressive translation (NAT) models to capture the multi-modal distribution of target translations due to their conditional independence assumption, which is known as the “multi-modality problem”, including the lexical multi-modality and the syntactic multi-modality. Whil…

2022

DA${2}$ Dataset: Toward Dexterity-Aware Dual-Arm Grasping

RA-L 2022

In this paper, we introduce DA <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$^{2}$</tex-math></inline-formula> , the first large-scale dual-arm dexterity-aware dataset for the generation of optimal bimanual grasp

Cited by 21SourceScholar
2022

Dict-TTS: Learning to Pronounce with Prior Dictionary Knowledge for Text-to-Speech

NeurIPS 2022accept

Polyphone disambiguation aims to capture accurate pronunciation knowledge from natural text sequences for reliable Text-to-speech (TTS) systems. However, previous approaches require substantial annotated training data and additional efforts from language experts, making it difficult to extend high-q…

2022

DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism

AAAI 2022technical

Singing voice synthesis (SVS) systems are built to synthesize high-quality and expressive singing voice, in which the acoustic model generates the acoustic features (e.g., mel-spectrogram) given a music score. Previous singing acoustic models adopt a simple loss (e.g., L1 and L2) or generative adver…

2022

EditSinger: Zero-Shot Text-Based Singing Voice Editing System with Diverse Prosody Modeling

IJCAI 2022poster

Zero-shot text-based singing editing enables singing voice modification based on the given edited lyrics without any additional data from the target singer. However, due to the different demands, challenges occur when applying existing speech editing methods to singing voice editing task, mainly inc…

2022

Expressivity of Emergent Languages is a Trade-off between Contextual Complexity and Unpredictability

ICLR 2022poster

Researchers are using deep learning models to explore the emergence of language in various language games, where agents interact and develop an emergent language to solve tasks. We focus on the factors that determine the expressivity of emergent languages, which reflects the amount of information ab…

Cited by 15SourcePDFScholar
2022

FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis

IJCAI 2022poster

Denoising diffusion probabilistic models (DDPMs) have recently achieved leading performances in many generative tasks. However, the inherited iterative sampling process costs hindered their applications to speech synthesis. This paper proposes FastDiff, a fast conditional diffusion model for high-qu…

2022

Flow-Based Unconstrained Lip to Speech Generation

AAAI 2022technical

Unconstrained lip-to-speech aims to generate corresponding speeches based on silent facial videos with no restriction to head pose or vocabulary. It is desirable to generate intelligible and natural speech with a fast speed in unconstrained settings. Currently, to handle the more complicated scena…

2022

GenerSpeech: Towards Style Transfer for Generalizable Out-Of-Domain Text-to-Speech

NeurIPS 2022accept

Style transfer for out-of-domain (OOD) speech synthesis aims to generate speech samples with unseen style (e.g., speaker identity, emotion, and prosody) derived from an acoustic reference, while facing the following challenges: 1) The highly dynamic style features in expressive voice are difficult t…

2022

HiFiDenoise: High-Fidelity Denoising Text to Speech with Adversarial Networks

ICASSP 2022accepted

Building a high-fidelity speech synthesis system with noisy speech data is a challenging but valuable task, which could significantly reduce the cost of data collection. Existing methods usually train speech synthesis systems based on the speech denoised with an enhancement model or feed noise infor…

Cited by 0SourceScholar
2022

Learning the Beauty in Songs: Neural Singing Voice Beautifier

ACL 2022long

We are interested in a novel task, singing voice beautification (SVB). Given the singing voice of an amateur singer, SVB aims to improve the intonation and vocal tone of the voice, while keeping the content and vocal timbre. Current automatic pitch correction techniques are immature, and most of the…

2022

M4Singer: A Multi-Style, Multi-Singer and Musical Score Provided Mandarin Singing Corpus

NeurIPS 2022accept

The lack of publicly available high-quality and accurately labeled datasets has long been a major bottleneck for singing voice synthesis (SVS). To tackle this problem, we present M4Singer, a free-to-use Multi-style, Multi-singer Mandarin singing collection with elaborately annotated Musical scores a…

Cited by 89SourcePDFScholar
2022

Parallel and High-Fidelity Text-to-Lip Generation

AAAI 2022technical

As a key component of talking face generation, lip movements generation determines the naturalness and coherence of the generated talking face video. Prior literature mainly focuses on speech-to-lip generation while there is a paucity in text-to-lip (T2L) generation. T2L is a challenging task and ex…

2022

Prosospeech: Enhancing Prosody with Quantized Vector Pre-Training in Text-To-Speech

ICASSP 2022accepted

Expressive text-to-speech (TTS) has become a hot research topic recently, mainly focusing on modeling prosody in speech. Prosody modeling has several challenges: 1) the extracted pitch used in previous prosody modeling works have inevitable errors, which hurts the prosody modeling; 2) different attr…

Cited by 0SourceScholar
2022

Spherical Convolutional Recurrent Neural Network for Real-Time Sound Source Tracking

ICASSP 2022accepted

Neural networks have been widely applied in direction-of-arrival (DOA) estimation and source tracking systems. In this paper, we introduce a spherical convolutional recurrent neural network that utilizes Deepsphere, a graph-based spherical convolutional neural network, employing the steered response…

Cited by 7SourceScholar
2022

Targeted Attack on Deep RL-based Autonomous Driving with Learned Visual Patterns

ICRA 2022poster

Recent studies demonstrated the vulnerability of control policies learned through deep reinforcement learning against adversarial attacks, raising concerns about the application of such models to risk-sensitive tasks such as autonomous driving. Threat models for these demonstrations are limited to (…

Cited by 15SourcecodeScholar
2022

Toward Global Sensing Quality Maximization: A Configuration Optimization Scheme for Camera Networks

IROS 2022poster

The performance of a camera network monitoring a set of targets depends crucially on the configuration of the cameras. In this paper, we investigate the reconfiguration strategy for the parameterized camera network model, with which the sensing qualities of the multiple targets can be optimized glob…

Cited by 0SourcecodeScholar
2021

Denoispeech: Denoising Text to Speech with Frame-Level Noise Modeling

ICASSP 2021accepted

While neural-based text to speech (TTS) models can synthesize natural and intelligible voice, they usually require high-quality speech data, which is costly to collect. In many scenarios, only noisy speech of a target speaker is available, which presents challenges for TTS model training for this sp…

Cited by 0SourceScholar
2021

FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

ICLR 2021poster

Non-autoregressive text to speech (TTS) models such as FastSpeech can synthesize speech significantly faster than previous autoregressive models with comparable quality. The training of FastSpeech model relies on an autoregressive teacher model for duration prediction (to provide more information as…

2021

SongMASS: Automatic Song Writing with Pre-training and Alignment Constraint

AAAI 2021technical

Automatic song writing aims to compose a song (lyric and/or melody) by machine, which is an interesting topic in both academia and industry. In automatic song writing, lyric-to-melody generation and melody-to-lyric generation are two important tasks, both of which usually suffer from the following c…

2021

UWSpeech: Speech to Speech Translation for Unwritten Languages

AAAI 2021technical

Existing speech to speech translation systems heavily rely on the text of target language: they usually translate source language either to target text and then synthesize target speech from text, or directly to target speech with target text for auxiliary training. However, those methods cannot be…

2021

When Shall I Be Empathetic? The Utility of Empathetic Parameter Estimation in Multi-Agent Interactions

ICRA 2021poster

Human-robot interactions (HRI) can be modeled as differential games with incomplete information, where each agent holds private reward parameters. Due to the open challenge in finding perfect Bayesian equilibria of such games, existing studies often decouple the belief and physical dynamics by itera…

Cited by 11SourceScholar
2020

Compositional languages emerge in a neural iterated learning model

ICLR 2020poster

The principle of compositionality, which enables natural language to represent complex concepts via a structured combination of simpler ones, allows us to convey an open-ended set of messages using a limited vocabulary. If compositionality is indeed a natural property of language, we may expect it t…

Cited by 115SourcecodeScholar
2020

Practical Quasi-Newton Methods for Training Deep Neural Networks

NeurIPS 2020spotlight

We consider the development of practical stochastic quasi-Newton, and in particular Kronecker-factored block diagonal BFGS and L-BFGS methods, for training deep neural networks (DNNs). In DNN training, the number of variables and components of the gradient n is often of the order of tens of million…

2020

Task-Level Curriculum Learning for Non-Autoregressive Neural Machine Translation

IJCAI 2020poster

Non-autoregressive translation (NAT) achieves faster inference speed but at the cost of worse accuracy compared with autoregressive translation (AT). Since AT and NAT can share model structure and AT is an easier task than NAT due to the explicit dependency on previous target-side tokens, a natural…

2019

Almost Unsupervised Text to Speech and Automatic Speech Recognition

ICML 2019oral

Text to speech (TTS) and automatic speech recognition (ASR) are two dual tasks in speech processing and both achieve impressive performance thanks to the recent advance in deep learning and large amount of aligned speech and text data. However, the lack of aligned data poses a major practical proble…

Cited by 131SourcePDFScholar
2019

FastSpeech: Fast, Robust and Controllable Text to Speech

NeurIPS 2019poster

Neural network based end-to-end text to speech (TTS) has significantly improved the quality of synthesized speech. Prominent methods (e.g., Tacotron 2) usually first generate mel-spectrogram from text, and then synthesize speech from the mel-spectrogram using vocoder such as WaveNet. Compared with t…

2019

How Shall I Drive? Interaction Modeling and Motion Planning towards Empathetic and Socially-Graceful Driving

ICRA 2019poster

While intelligence of autonomous vehicles (AVs) has significantly advanced in recent years, accidents involving AVs suggest that these autonomous systems lack gracefulness in driving when interacting with human drivers. In the setting of a two-player game, we propose model predictive control based o…

Cited by 20SourceScholar
2019

Low-cost Measurement of Industrial Shock Signals via Deep Learning Calibration

ICASSP 2019accepted

Special high-end sensors with expensive hardware are usually needed to measure shock signals with high accuracy. In this paper, we show that cheap low-end sensors calibrated by deep neural networks are also capable to measure high-g shocks accurately. Firstly we perform drop shock tests to collect a…

Cited by 0SourceScholar
2019

Multilingual Neural Machine Translation with Knowledge Distillation

ICLR 2019poster

Multilingual machine translation, which translates multiple languages with a single model, has attracted much attention due to its efficiency of offline training and online serving. However, traditional multilingual translation usually yields inferior accuracy compared with the counterpart using ind…