← Search

Yifu Chen

7 accepted papers

2025

InteractSpeech: A Speech Dialogue Interaction Corpus for Spoken Dialogue Model

EMNLP 2025

Spoken Dialogue Models (SDMs) have achieved significant progress in recent years, yet they continue to face challenges in handling nuanced interactional phenomena. A significant bottleneck hindering further advancement is the scarcity of publicly available, high-quality datasets meticulously designe

2025

Speech Watermarking with Discrete Intermediate Representations

AAAI 2025technical

Speech watermarking techniques can proactively mitigate the potential harmful consequences of instant voice cloning techniques. These techniques involve the insertion of signals into speech that are imperceptible to humans but can be detected by algorithms. Previous approaches typically embed waterm…

Cited by 1SourcePDFScholar
2025

WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models

ACL 2025long

Retrieval Augmented Generation (RAG) has gained widespread adoption owing to its capacity to empower large language models (LLMs) to integrate external knowledge. However, existing RAG frameworks are primarily designed for text-based LLMs and rely on Automatic Speech Recognition to process speech in…

Cited by 0SourcePDFScholar
2025

WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

ICLR 2025poster

Language models have been effectively applied to modeling natural signals, such as images, video, speech, and audio. A crucial component of these models is the codec tokenizer, which compresses high-dimensional natural signals into lower-dimensional discrete tokens. In this paper, we introduce WavTo…

2024

Hetecooper: Feature Collaboration Graph for Heterogeneous Collaborative Perception

ECCV 2024poster

"Collaborative perception effectively expands the perception range of agents by sharing perceptual information, and it addresses the occlusion problem in single-vehicle perception. Most of the existing works are based on the assumption of perception model homogeneity. However, in actual collaboratio…

Cited by 1SourcePDFScholar
2024

Improving Text-guided Object Inpainting with Semantic Pre-inpainting

ECCV 2024poster

"Recent years have witnessed the success of large text-to-image diffusion models and their remarkable potential to generate high-quality images. The further pursuit of enhancing the editability of images has sparked significant interest in the downstream task of inpainting a novel object described b…

2021

PLOP: Learning Without Forgetting for Continual Semantic Segmentation

CVPR 2021poster

Deep learning approaches are nowadays ubiquitously used to tackle computer vision tasks such as semantic segmentation, requiring large datasets and substantial computational power. Continual learning for semantic segmentation (CSS) is an emerging trend that consists in updating an old model by seque…

Cited by 305PDFcodeScholar