← Search

Haikuo Peng

2 accepted papers

2026

Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial Generation

AAAI 2026technical

While significant progress has been achieved in multimodal facial generation using semantic masks and textual descriptions, conventional feature fusion approaches often fail to enable effective cross-modal interactions, thereby leading to suboptimal generation outcomes. To address this challenge, we

Cited by 0SourcePDFScholar
2026

Zero-shot Active Mapping via Fused 360-BEV Representations and Vision–Language Models

ICML 2026poster

Active mapping enables embodied agents to understand and interact in previously unseen environments. However, most methods struggle to achieve zero-shot generalization to large-scale scenes and lack support for language instructions. We propose a VLM-based active mapping method that achieves zero-sh…

Cited by 0SourceScholar