← Search

Ye Fang

4 accepted papers

2026

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

ICLR 2026poster

In recent years, 2D Vision-Language Models (VLMs) have made significant strides in image-text understanding tasks. However, their performance in 3D spatial comprehension, which is critical for embodied intelligence, remains limited. Recent advances have leveraged 3D point clouds and multi-view image…

Cited by 0SourcecodeScholar
2026

V-RGBX: Video Editing with Accurate Controls over Intrinsic Properties

CVPR 2026

Large-scale video generation models have shown remarkable potential in modeling photorealistic appearance and lighting interactions in real-world scenes. However, a closed-loop framework that jointly understands intrinsic scene properties (e.g., albedo, normal, material, and irradiance), leverages t

Cited by 0SourcecodeScholar
2024

GPT4Point: A Unified Framework for Point-Language Understanding and Generation

CVPR 2024highlight

Multimodal Large Language Models (MLLMs) have excelled in 2D image-text comprehension and image generation but their understanding of the 3D world is notably deficient limiting progress in 3D language understanding and generation. To solve this problem we introduce GPT4Point an innovative groundbrea…

Cited by 43SourcePDFScholar
2024

Make-it-Real: Unleashing Large Multimodal Model for Painting 3D Objects with Realistic Materials

NeurIPS 2024poster

Physically realistic materials are pivotal in augmenting the realism of 3D assets across various applications and lighting conditions. However, existing 3D assets and generative models often lack authentic material properties. Manual assignment of materials using graphic software is a tedious and ti…

Cited by 2SourcePDFScholar