← Search

Muyang Zhang

4 accepted papers

2026

LaplacianFormer:Rethinking Linear Attention with Laplacian Kernel

ICLR 2026poster

The quadratic complexity of softmax attention presents a major obstacle for scaling Transformers to high-resolution vision tasks. Existing linear attention variants often replace the softmax with Gaussian kernels to reduce complexity, but such approximations lack theoretical grounding and tend to ov…

Cited by 0SourceScholar
2025

AccidentX: A Large-Scale Multimodal BEV Dataset for Traffic Accident Analysis and Prevention

IROS 2025

With the rapid development and widespread application of autonomous driving technology, the accurate analysis and prevention of traffic accidents have become critical challenges. However, current traffic accident datasets are often constrained by limited scale and diversity, impeding progress in thi

Cited by 0SourceScholar
2025

HumanDreamer: Generating Controllable Human-Motion Videos via Decoupled Generation

CVPR 2025poster

Human-motion video generation has been a challenging task, primarily due to the difficulty inherent in learning human body movements. While some approaches have attempted to drive human-centric video generation explicitly through pose control, these methods typically rely on poses derived from exist…

Cited by 2SourcePDFScholar
2025

PanoDiT: Panoramic Videos Generation with Diffusion Transformer

AAAI 2025technical

As immersive experiences become increasingly popular, panoramic video has garnered significant attention in both research and applications. The high cost associated with capturing panoramic video underscores the need for efficient prompt-based generation methods. Although recent text-to-video (T2V)…

Cited by 0SourcePDFScholar