← Search

Guangkai Xu

8 accepted papers

2026

CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models

ICML 2026poster

Large Language Models(LLMs) have revolutionized text generation and multimodal perception, but their capabilities in 3D content generation remain underexplored. Existing methods compromise by producing either low-resolution meshes or coarse structural proxies, failing to capture fine-grained geometr…

Cited by 0SourceScholar
2026

Unlocking the Power of Critical Factors for 3D Visual Geometry Estimation

CVPR 2026

Feed-forward visual geometry estimation has recently made rapid progress. However, an important gap remains: multi-frame models usually produce better cross-frame consistency, yet they often underperform strong per-frame methods on single-frame accuracy. This observation motivates our systematic inv

Cited by 0SourcecodeScholar
2025

DiffCalib: Reformulating Monocular Camera Calibration as Diffusion-Based Dense Incident Map Generation

AAAI 2025technical

Monocular camera calibration is a key precondition for numerous 3D vision applications. Despite considerable advancements, existing methods often hinge on specific assumptions and struggle to generalize across varied real-world scenarios, and the performance is limited by insufficient training data.…

Cited by 4SourcePDFScholar
2025

POMATO: Marrying Pointmap Matching with Temporal Motions for Dynamic 3D Reconstruction

ICCV 2025poster

Recent approaches to 3D reconstruction in dynamic scenes primarily rely on the integration of separate geometry estimation and matching modules, where the latter plays a critical role in distinguishing dynamic regions and mitigating the interference caused by moving objects. Furthermore, the matchin…

2025

What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?

ICLR 2025poster

Extensive pre-training with large data is indispensable for downstream geometry and semantic visual perception tasks. Thanks to large-scale text-to-image (T2I) pretraining, recent works show promising results by simply fine-tuning T2I diffusion models for a few dense perception tasks. However, sever…

2024

Improving Neural Indoor Surface Reconstruction with Mask-Guided Adaptive Consistency Constraints

ICRA 2024poster

3D scene reconstruction from 2D images has been a long-standing task. Instead of estimating per-frame depth maps and fusing them in 3D, recent researches leverage the neural implicit surface as a global representation for 3D reconstruction. Equipped with data-driven pre-trained geometric cues, these…

Cited by 2SourceScholar
2024

Unleashing the Potential of the Diffusion Model in Few-shot Semantic Segmentation

NeurIPS 2024poster

The Diffusion Model has not only garnered noteworthy achievements in the realm of image generation but has also demonstrated its potential as an effective pretraining method utilizing unlabeled data. Drawing from the extensive potential unveiled by the Diffusion Model in both semantic corresponden…

2023

FrozenRecon: Pose-free 3D Scene Reconstruction with Frozen Depth Models

ICCV 2023poster

3D scene reconstruction is a long-standing vision task. Existing approaches can be categorized into geometry-based and learning-based methods. The former leverages multi-view geometry but may face catastrophic failures due to the reliance on accurate pixel correspondence across views, while the latt…

Cited by 17PDFcodeScholar