← Search

Zhenghao Zhang

7 accepted papers

2026

Dynamic Deep Graph Learning for Incomplete Multi-View Clustering with Masked Graph Reconstruction Loss

AAAI 2026technical

The prevalence of real-world multi-view data makes incomplete multi-view clustering (IMVC) a crucial research. The rapid development of Graph Neural Networks (GNNs) has established them as one of the mainstream approaches for multi-view clustering. Despite significant progress in GNNs-based IMVC, so

Cited by 0SourcePDFScholar
2026

LaTo: Landmark-tokenized Diffusion Transformer for Fine-grained Human Face Editing

ICLR 2026poster

Recent multimodal models for instruction-based face editing enable semantic manipulation but still struggle with precise attribute control and identity preservation. Structural facial representations such as landmarks are effective for intermediate supervision, yet most existing methods treat them a…

Cited by 0SourcecodeScholar
2026

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving

CVPR 2026

End-to-end autonomous driving methods built on vision language models (VLMs) have undergone rapid development driven by their universal visual understanding and strong reasoning capabilities obtained from the large-scale pretraining. However, we find that current VLMs struggle to understand fine-gra

Cited by 0SourcecodeScholar
2025

Federated Incomplete Multi-view Clustering with Globally Fused Graph Guidance

ICML 2025poster

Federated multi-view clustering has been proposed to mine the valuable information within multi-view data distributed across different devices and has achieved impressive results while preserving the privacy. Despite great progress, most federated multi-view clustering methods only used global pseu…

2025

Tora: Trajectory-oriented Diffusion Transformer for Video Generation

CVPR 2025poster

Recent advancements in Diffusion Transformer (DiT) have demonstrated remarkable proficiency in producing high-quality video content. Nonetheless, the potential of transformer-based diffusion models for effectively generating videos with controllable motion remains an area of limited exploration. Thi…

2025

TransVDM: Motion-Constrained Video Diffusion Model for Transparent Video Synthesis

ICASSP 2025accepted

Recent developments in Video Diffusion Models (VDMs) have demonstrated remarkable capability to generate high-quality video content. Nonetheless, the potential of VDMs for creating transparent videos remains largely uncharted. In this paper, we introduce TransVDM, the first diffusion-based model spe…

Cited by 0SourceScholar
2024

MapLocNet: Coarse-to-Fine Feature Registration for Visual Re-Localization in Navigation Maps

IROS 2024

Robust localization is the cornerstone of autonomous driving, especially in challenging urban environments where GPS signals suffer from multipath errors. Traditional localization approaches rely on high-definition (HD) maps, which consist of precisely annotated landmarks. However, building HD map i

Cited by 29SourceScholar