← Search

Ruiqi Li

23 accepted papers

2026

Active Learning of 3D Gaussian Splatting with Consistent Region Partition and Robust Pose Estimation

ICLR 2026poster

Radiance fields have been successful in reconstructing 3D assets for scenes presented in Virtual Reality and Augmented Reality (VR/AR). The general workflow of scanning objects with radiance field representation involves a heavy workload of capturing images depicting the object empirically by the us…

Cited by 0SourceScholar
2026

S3Audio: Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer

ICML 2026poster

Real-time and accurate spatial audio generation is pivotal for delivering an immersive experience. However, existing spatial audio synthesis technologies are often encumbered by a tradeoff between generation quality and high inference latency, as well as difficulty in capturing precise spatial infor…

Cited by 0SourcecodeScholar
2026

TAPO: Dynamic Teacher and Perturbed Answer Injection for Policy Optimization

AAAI 2026technical

Reinforcement learning (RL) has emerged as a powerful framework to improve the reasoning performance of large language models (LLMs), with approaches such as Group Relative Policy Optimization (GRPO) showing promising results. However, GRPO and its variants struggle with collapsed groups (i.e., all-

Cited by 0SourcePDFScholar
2025

STARS: A Unified Framework for Singing Transcription, Alignment, and Refined Style Annotation

ACL 2025finding

Recent breakthroughs in singing voice synthesis (SVS) have heightened the demand for high-quality annotated datasets, yet manual annotation remains prohibitively labor-intensive and resource-intensive. Existing automatic singing annotation (ASA) methods, however, primarily tackle isolated aspects of…

2025

TechSinger: Technique Controllable Multilingual Singing Voice Synthesis via Flow Matching

AAAI 2025technical

Singing voice synthesis has made remarkable progress in generating natural and high-quality voices. However, existing methods rarely provide precise control over vocal techniques such as intensity, mixed voice, falsetto, bubble, and breathy tones, thus limiting the expressive potential of synthetic…

2025

Versatile Framework for Song Generation with Prompt-based Control

EMNLP 2025

Song generation focuses on producing controllable high-quality songs based on various prompts. However, existing methods struggle to generate vocals and accompaniments with prompt-based control and proper alignment. Additionally, they fall short in supporting various tasks. To address these challeng

2025

WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

ICLR 2025poster

Language models have been effectively applied to modeling natural signals, such as images, video, speech, and audio. A crucial component of these models is the codec tokenizer, which compresses high-dimensional natural signals into lower-dimensional discrete tokens. In this paper, we introduce WavTo…

2024

Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching

NeurIPS 2024poster

Video-to-audio (V2A) generation aims to synthesize content-matching audio from silent video, and it remains challenging to build V2A models with high generation quality, efficiency, and visual-audio temporal synchrony. We propose Frieren, a V2A model based on rectified flow matching. Frieren regres…

2024

GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks

NeurIPS 2024spotlight

The scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datasets suffer from low quality, limited diversity of languages and singers, absence of multi-technique information and real…

2024

GaussianMarker: Uncertainty-Aware Copyright Protection of 3D Gaussian Splatting

NeurIPS 2024poster

3D Gaussian Splatting (3DGS) has become a crucial method for acquiring 3D assets. To protect the copyright of these assets, digital watermarking techniques can be applied to embed ownership information discreetly within 3DGS mod- els. However, existing watermarking methods for meshes, point clouds,…

Cited by 4SourcePDFScholar
2024

NaRuto: Automatically Acquiring Planning Models from Narrative Texts

AAAI 2024technical

Domain model acquisition has been identified as a bottleneck in the application of planning technology, especially within narrative planning. Learning action models from narrative texts in an automated way is essential to overcome this barrier, but challenging because of the inherent complexities of…

Cited by 5SourcePDFScholar
2024

Prompt-Singer: Controllable Singing-Voice-Synthesis with Natural Language Prompt

NAACL 2024long

Recent singing-voice-synthesis (SVS) methods have achieved remarkable audio quality and naturalness, yet they lack the capability to control the style attributes of the synthesized singing explicitly. We propose Prompt-Singer, the first SVS method that enables attribute controlling on singer gender,…

2024

Robust Singing Voice Transcription Serves Synthesis

ACL 2024long

Note-level Automatic Singing Voice Transcription (AST) converts singing recordings into note sequences, facilitating the automatic annotation of singing datasets for Singing Voice Synthesis (SVS) applications. Current AST methods, however, struggle with accuracy and robustness when used for practica…

2024

Self-Supervised Singing Voice Pre-Training towards Speech-to-Singing Conversion

ACL 2024findings

Speech-to-singing voice conversion (STS) task always suffers from data scarcity, because it requires paired speech and singing data. Compounding this issue are the challenges of content-pitch alignment and the suboptimal quality of generated outputs, presenting significant hurdles in STS research. T…

2024

StyleSinger: Style Transfer for Out-of-Domain Singing Voice Synthesis

AAAI 2024technical

Style transfer for out-of-domain (OOD) singing voice synthesis (SVS) focuses on generating high-quality singing voices with unseen styles (such as timbre, emotion, pronunciation, and articulation skills) derived from reference singing voice samples. However, the endeavor to model the intricate nuanc…

2024

TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control

EMNLP 2024main

Zero-shot singing voice synthesis (SVS) with style transfer and style control aims to generate high-quality singing voices with unseen timbres and styles (including singing method, emotion, rhythm, technique, and pronunciation) from audio and text prompts. However, the multifaceted nature of singing…

2024

Text-to-Song: Towards Controllable Music Generation Incorporating Vocal and Accompaniment

ACL 2024long

A song is a combination of singing voice and accompaniment. However, existing works focus on singing voice synthesis and music generation independently. Little attention was paid to exploring song synthesis. In this work, we propose a novel task called Text-to-Song synthesis which incorporates both…

2024

Variational Multi-scale Representation for Estimating Uncertainty in 3D Gaussian Splatting

NeurIPS 2024poster

Recently, 3D Gaussian Splatting (3DGS) has become popular in reconstructing dense 3D representations of appearance and geometry. However, the learning pipeline in 3DGS inherently lacks the ability to quantify uncertainty, which is an important factor in applications like robotics mapping and navigat…

Cited by 1SourcePDFScholar
2023

AlignSTS: Speech-to-Singing Conversion via Cross-Modal Alignment

ACL 2023findings

The speech-to-singing (STS) voice conversion task aims to generate singing samples corresponding to speech recordings while facing a major challenge: the alignment between the target (singing) pitch contour and the source (speech) content is difficult to learn in a text-free situation. This paper pr…

2022

M4Singer: A Multi-Style, Multi-Singer and Musical Score Provided Mandarin Singing Corpus

NeurIPS 2022accept

The lack of publicly available high-quality and accurately labeled datasets has long been a major bottleneck for singing voice synthesis (SVS). To tackle this problem, we present M4Singer, a free-to-use Multi-style, Multi-singer Mandarin singing collection with elaborately annotated Musical scores a…

Cited by 123SourcePDFScholar
2022

Towards Explainable Action Recognition by Salient Qualitative Spatial Object Relation Chains

AAAI 2022technical

In order to be trusted by humans, Artificial Intelligence agents should be able to describe rationales behind their decisions. One such application is human action recognition in critical or sensitive scenarios, where trustworthy and explainable action recognizers are expected. For example, reliable…

Cited by 6SourcePDFScholar