← Search

zhengyi zhao

7 accepted papers

2026

Large Depth Completion Model from Sparse Observations

ICLR 2026poster

This work presents the Large Depth Completion Model (LDCM), a simple, effective, and robust framework for single-view metric depth estimation with sparse observations. Without relying on complex architectural designs, LDCM generates metric-accurate dense depth maps in one large transformer. It outpe…

Cited by 0SourceScholar
2025

MemeReaCon: Probing Contextual Meme Understanding in Large Vision-Language Models

EMNLP 2025

Memes have emerged as a popular form of multimodal online communication, where their interpretation heavily depends on the specific context in which they appear. Current approaches predominantly focus on isolated meme analysis, either for harmful content detection or standalone interpretation, overl

Cited by 0SourcePDFScholar
2025

T2: An Adaptive Test-Time Scaling Strategy for Contextual Question Answering

EMNLP 2025

Recent advances in large language models have demonstrated remarkable performance on Contextual Question Answering (CQA). However, prior approaches typically employ elaborate reasoning strategies regardless of question complexity, leading to low adaptability. Recent efficient test-time scaling metho

2024

An Optimization Framework to Enforce Multi-View Consistency for Texturing 3D Meshes

ECCV 2024poster

"A fundamental problem in the texturing of 3D meshes using pre-trained text-to-image models is to ensure multi-view consistency. State-of-the-art approaches typically use diffusion models to aggregate multi-view inputs, where common issues are the blurriness caused by the averaging operation in the…

2024

GPLD3D: Latent Diffusion of 3D Shape Generative Models by Enforcing Geometric and Physical Priors

CVPR 2024poster

State-of-the-art man-made shape generative models usually adopt established generative models under a suitable implicit shape representation. A common theme is to perform distribution alignment which does not explicitly model important shape priors. As a result many synthetic shapes are not connecte…

Cited by 8SourcePDFScholar
2024

High-Fidelity 3D Textured Shapes Generation by Sparse Encoding and Adversarial Decoding

ECCV 2024poster

"3D vision is inherently characterized by sparse spatial structures, which propels the necessity for an efficient paradigm tailored to 3D generation. Another discrepancy is the amount of training data, which undeniably affects generalization if we only use limited 3D data. To solve these, we design…

2020

Hand-3d-Studio: A New Multi-View System for 3d Hand Reconstruction

ICASSP 2020accepted

This paper proposes a new system named as Hand-3D-Studio to capture the 3D hand pose and shape information. Our system includes 15 synchronized DSLR cameras, which can acquire high quality multi-view 4K resolution color images in a circular manner. We then introduce a 2D hand keypoints guided iterat…

Cited by 0SourceScholar