← Search

Jianxuan Yang

2 accepted papers

2026

AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control

AAAI 2026technical

Sound effect editing—modifying audio by adding, removing, or replacing elements—remains constrained by existing approaches that rely solely on low-level signal processing or coarse text prompts, often resulting in limited flexibility and suboptimal audio quality. To address this, we propose AV-Edit,

Cited by 0SourcePDFScholar
2026

MEANFLOW-ACCELERATED MULTIMODAL VIDEO-TO-AUDIO SYNTHESIS VIA ONE-STEP GENERATION

ICASSP 2026poster

A key challenge in synthesizing audios from silent videos is the inherent trade-off between synthesis quality and inference efficiency in existing methods. For instance, flow matching based models rely on modeling instantaneous velocity, inherently require an iterative sampling process, leading to s…

Cited by 0SourcePDFScholar