← Search

Zhiyuan Zhu

12 accepted papers

2026

FineCycle: Towards a Full-Cycle Management Paradigm for Robotic Deployment and Development

ICRA 2026poster

Typical robotic workflows involve deploying applications from servers to robots for testing or distributing validated applications across a fleet to unify capabilities. Because these processes are often slowed by tedious environment configurations, we need a way to improve deployment and development…

Cited by 0Scholar
2026

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments

ICML 2026poster

Current audio-visual large language models (AV-LLMs) are predominantly restricted to 2D perception, relying on RGB video and monaural audio. This design choice introduces a fundamental dimensionality mismatch that precludes reliable source localization and spatial reasoning in complex 3D environment…

Cited by 0SourceScholar
2025

EvolveBench: A Comprehensive Benchmark for Assessing Temporal Awareness in LLMs on Evolving Knowledge

ACL 2025long

Large language models (LLMs) are trained on extensive historical corpora, but their ability to understand time and maintain temporal awareness of time-evolving factual knowledge remains limited. Previous studies often neglect the critical aspect of utilizing knowledge from various sources. To addres…

2025

MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations

NeurIPS 2025poster

Humans rely on multisensory integration to perceive spatial environments, where auditory cues enable sound source localization in three-dimensional space. Despite the critical role of spatial audio in immersive technologies such as VR/AR, most existing multimodal datasets provide only monaural audi…

Cited by 0SourcecodeScholar
2025

STARS: A Unified Framework for Singing Transcription, Alignment, and Refined Style Annotation

ACL 2025finding

Recent breakthroughs in singing voice synthesis (SVS) have heightened the demand for high-quality annotated datasets, yet manual annotation remains prohibitively labor-intensive and resource-intensive. Existing automatic singing annotation (ASA) methods, however, primarily tackle isolated aspects of…

2025

TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis

ACL 2025finding

Customizable multilingual zero-shot singing voice synthesis (SVS) has various potential applications in music composition and short video dubbing. However, existing SVS models overly depend on phoneme and note boundary annotations, limiting their robustness in zero-shot scenarios and producing poor…

2025

Versatile Framework for Song Generation with Prompt-based Control

EMNLP 2025

Song generation focuses on producing controllable high-quality songs based on various prompts. However, existing methods struggle to generate vocals and accompaniments with prompt-based control and proper alignment. Additionally, they fall short in supporting various tasks. To address these challeng

2024

CE-VDG: Counterfactual Entropy-based Bias Reduction for Video-grounded Dialogue Generation

COLING 2024main

The Video-Grounded Dialogue generation (VDG) is a challenging task requiring a comprehensive understanding of the multi-modal information to produce a pertinent response. However, VDG models may rely on dataset bias as a shortcut and fail to learn the multi-modal knowledge from both video and audio.…

Cited by 1SourcePDFScholar
2024

GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks

NeurIPS 2024spotlight

The scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datasets suffer from low quality, limited diversity of languages and singers, absence of multi-technique information and real…

2024

RA2FD: Distilling Faithfulness into Efficient Dialogue Systems

EMNLP 2024main

Generating faithful and fast responses is crucial in the knowledge-grounded dialogue. Retrieval Augmented Generation (RAG) strategies are effective but are inference inefficient, while previous Retrieval Free Generations (RFG) are more efficient but sacrifice faithfulness. To solve this faithfulness…

2020

Eeg Feature Selection Using Orthogonal Regression: Application to Emotion Recognition

ICASSP 2020accepted

A common drawback of the EEG applications is that the volume conduction of human head leads to lots of redundant information in EEG recordings. To reduce the redundancy and choose informative EEG features, in this paper, we propose an EEG feature selection technique, termed as Feature Selection with…

Cited by 0SourceScholar