← Search

Peike Li

7 accepted papers

2025

Bridging Sign and Spoken Languages: Pseudo Gloss Generation for Sign Language Translation

NeurIPS 2025poster

Sign Language Translation (SLT) aims to map sign language videos to spoken language text. A common approach relies on gloss annotations as an intermediate representation, decomposing SLT into two sub-tasks: video-to-gloss recognition and gloss-to-text translation. While effective, this paradigm depe…

Cited by 0SourceScholar
2025

Dynamic Derivation and Elimination: Audio Visual Segmentation with Enhanced Audio Semantics

CVPR 2025poster

Sound-guided object segmentation has drawn considerable attention for its potential to enhance multimodal perception. Previous methods primarily focus on developing advanced architectures to facilitate effective audio-visual interactions, without fully addressing the inherent challenges posed by aud…

2025

JEN-1 Composer: A Unified Framework for High-Fidelity Multi-Track Music Generation

AAAI 2025technical

With rapid advances in generative artificial intelligence, the text-to-music synthesis task has emerged as a promising direction for music generation. Nevertheless, achieving precise control over multi-track generation remains an open challenge. While existing models excel in directly generating mul…

Cited by 12SourcePDFScholar
2025

JEN-1 DreamStyler: Customized Musical Concept Learning via Pivotal Parameters Tuning

AAAI 2025technical

Large models for text-to-music generation have achieved significant progress, facilitating the creation of high-quality and varied musical compositions from provided text prompts. However, input text prompts may not precisely capture user requirements, particularly when the objective is to generate…

Cited by 2SourcePDFScholar
2025

Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent Alignment

CVPR 2025poster

Accurately localizing audible objects based on audio-visual cues is the core objective of audio-visual segmentation. Most previous methods emphasize spatial or temporal multi-modal modeling, yet overlook challenges from ambiguous audio-visual correspondences--such as nearby visually similar but acou…

Cited by 0SourcePDFScholar