← Search

Rui Tao

5 accepted papers

2026

OVLR: Efficient, Scalable, and Robust Training via Output-Level Variance-Reduced Likelihood Ratio

ICML 2026poster

Gradient-based optimization is fundamental to deep learning, yet standard backpropagation (BP) is inherently limited by the requirement of differentiability, rendering it brittle when encountering piecewise-constant objectives with vanishing gradients (e.g., hard 0-1 loss) or black-box feedback. Whi…

Cited by 0SourceScholar
2026

RiskPO: Risk-based Policy Optimization with Verifiable Reward for LLM Post-Training

ICLR 2026poster

Reinforcement learning with verifiable reward has recently emerged as a central paradigm for post-training large language models (LLMs); however, prevailing mean-based methods, such as Group Relative Policy Optimization (GRPO), suffer from entropy collapse and limited reasoning gains. We argue that…

Cited by 0SourcecodeScholar
2024

Audio Generation with Multiple Conditional Diffusion Model

AAAI 2024technical

Text-based audio generation models have limitations as they cannot encompass all the information in audio, leading to restricted controllability when relying solely on text. To address this issue, we propose a novel model that enhances the controllability of existing pre-trained text-to-audio models…

2024

Semi-Supervised Sound Event Detection with Local and Global Consistency Regularization

ICASSP 2024accepted

Learning meaningful frame-wise features on a partially labeled dataset is crucial to semi-supervised sound event detection. Prior works either maintain consistency on frame-level predictions or seek feature-level similarity among neighboring frames, which cannot exploit the potential of unlabeled da…

Cited by 0SourceScholar
2022

A Robotic End-Effector for Screwing and Unscrewing Bolts From the Side

RA-L 2022

This letter presents a novel robotic end-effector for screwing and unscrewing bolts. In many industrial scenarios, it is required to manipulate bolts from the side using robots with end-effectors. Besides, the reaction torque during tightening needs to be balanced to the stability of the robot. To a

Cited by 7SourceScholar