← Search

Guanghao Zheng

4 accepted papers

2026

GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs

ICLR 2026poster

Vision encoders are indispensable for allowing impressive performance of Multimodal Large Language Models (MLLMs) in vision–language tasks such as visual question answering and reasoning. However, existing vision encoders focus on global image representations but overlook fine-grained regional analy…

Cited by 0SourceScholar
2024

MC-DiT: Contextual Enhancement via Clean-to-Clean Reconstruction for Masked Diffusion Models

NeurIPS 2024poster

Diffusion Transformer (DiT) is emerging as a cutting-edge trend in the landscape of generative diffusion models for image generation. Recently, masked-reconstruction strategies have been considered to improve the efficiency and semantic consistency in training DiT but suffer from deficiency in conte…

Cited by 0SourcePDFScholar
2024

Towards Unified Representation of Invariant-Specific Features in Missing Modality Face Anti-Spoofing

ECCV 2024poster

"The effectiveness of Vision Transformers (ViTs) diminishes considerably in multi-modal face anti-spoofing (FAS) under missing modality scenarios. Existing approaches rely on modality-invariant features to alleviate this issue but ignore modality-specific features. To solve this issue, we propose a…

Cited by 4SourcePDFScholar
2023

Learning Causal Representations for Generalizable Face Anti Spoofing

ICASSP 2023accepted

Generalization ability of face anti-spoofing has been widely concerned in recent years. Existing domain generalization methods use adversarial learning or metric learning to extract invariant features across domains but are proved to be flawed from causal views. The learned domain-invariant features…

Cited by 0SourceScholar