← Search

Dongjie Fu

6 accepted papers

2026

MARS-Sep: Multimodal-Aligned Reinforced Sound Separation

ICLR 2026poster

Universal sound separation faces a fundamental misalignment: models optimized for low-level signal metrics often produce semantically contaminated outputs, failing to suppress perceptually salient interference from acoustically similar sources. We introduce a preference alignment perspective, analog…

Cited by 0SourcecodeScholar
2026

MOGO: Residual Quantized Hierarchical Causal Transformer for Real-Time and Infinite-Length 3D Human Motion Generation

AAAI 2026technical

Recent advances in transformer-based text-to-motion generation have significantly improved motion quality. However, achieving both real-time performance and long-horizon scalability remains an open challenge. In this paper, we present MOGO (Motion Generation with One-pass), a novel autoregressive fr

Cited by 0SourcePDFScholar
2025

AHa-Bench: Benchmarking Audio Hallucinations in Large Audio-Language Models

NeurIPS 2025poster

Hallucinations present a significant challenge in the development and evaluation of large language models (LLMs), directly affecting their reliability and accuracy. While notable advancements have been made in research on textual and visual hallucinations, there is still a lack of a comprehensive be…

Cited by 0SourceScholar
2025

PACHAT: Persona-Aware Speech Assistant for Multi-party Dialogue

EMNLP 2025

Extensive research on LLM-based spoken dialogue systems has significantly advanced the development of intelligent voice assistants. However, the integration of role information within speech remains an underexplored area, limiting its application in real-world scenarios, particularly in multi-party

Cited by 0SourcePDFScholar
2025

VoxDialogue: Can Spoken Dialogue Systems Understand Information Beyond Words?

ICLR 2025poster

With the rapid advancement of large models, voice assistants are gradually acquiring the ability to engage in open-ended daily conversations with humans. However, current spoken dialogue systems often overlook multi-modal information in audio beyond text, such as speech rate, volume, emphasis, and b…

2024

MPOD123: One Image to 3D Content Generation Using Mask-enhanced Progressive Outline-to-Detail Optimization

CVPR 2024poster

Recent advancements in single image driven 3D content generation have been propelled by leveraging prior knowledge from pretrained 2D diffusion models. However the 3D content generated by existing methods often exhibits distorted outline shapes and inadequate details. To solve this problem we propos…

Cited by 1SourcePDFScholar