← Search

Wanchao Su

3 accepted papers

2026

MIDI-LLaMA: An Instruction-Following Multimodal LLM for Symbolic Music Understanding

ICASSP 2026poster

Recent advances in multimodal large language models (MLLM) for audio music have demonstrated strong capabilities in music understanding, yet symbolic music, a fundamental representation of musical structure, remains unexplored. In this work, we introduce MIDI-LLaMA, the first instruction-following M…

Cited by 0SourcePDFScholar
2026

RecEdit-Drive: 3D Reconstruction-Guided Spatiotemporal Video Editing for Autonomous Driving Scenes

CVPR 2026

High-quality video editing and processing are crucial in domains such as filmmaking and autonomous driving, where accurate visual refinement and data preparation are essential. However, it is challenging to achieve precise control over dynamic objects while maintaining spatiotemporal consistency. Cu

Cited by 0SourcecodeScholar
2025

Chat2SVG: Vector Graphics Generation with Large Language Models and Image Diffusion Models

CVPR 2025poster

Scalable Vector Graphics (SVG) has become the de facto standard for vector graphics in digital design, offering resolution independence and precise control over individual elements. Despite their advantages, creating high-quality SVG content remains challenging, as it demands technical expertise wit…