← Search

Zhangyang Qi

3 accepted papers

2026

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

ICLR 2026poster

In recent years, 2D Vision-Language Models (VLMs) have made significant strides in image-text understanding tasks. However, their performance in 3D spatial comprehension, which is critical for embodied intelligence, remains limited. Recent advances have leveraged 3D point clouds and multi-view image…

Cited by 0SourcecodeScholar
2026

Game Ground Bench: Probing the Limits of LVLMs in Complex Semantic Grounding Across Game Universes

AAAI 2026technical

Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities, yet their ability to ground language in complex, interactive environments such as video games remains a critical frontier. Existing benchmarks are inadequate for this purpose: real-world datasets like RefCOCO introduce a

Cited by 0SourcePDFScholar
2024

GPT4Point: A Unified Framework for Point-Language Understanding and Generation

CVPR 2024highlight

Multimodal Large Language Models (MLLMs) have excelled in 2D image-text comprehension and image generation but their understanding of the 3D world is notably deficient limiting progress in 3D language understanding and generation. To solve this problem we introduce GPT4Point an innovative groundbrea…

Cited by 43SourcePDFScholar