← Search

Kaizhi Zheng

7 accepted papers

2025

EditRoom: LLM-parameterized Graph Diffusion for Composable 3D Room Layout Editing

ICLR 2025poster

Given the steep learning curve of professional 3D software and the time- consuming process of managing large 3D assets, language-guided 3D scene editing has significant potential in fields such as virtual reality, augmented reality, and gaming. However, recent approaches to language-guided 3D scene…

Cited by 0SourcePDFScholar
2025

GRIT: Teaching MLLMs to Think with Images

NeurIPS 2025poster

Recent studies have demonstrated the efficacy of using Reinforcement Learning (RL) in building reasoning models that articulate chains of thoughts prior to producing final answers. However, despite ongoing advances that aim at enabling reasoning for vision-language tasks, existing open-source visual…

Cited by 0SourceScholar
2025

MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

ICLR 2025poster

Multimodal Language Language Models (MLLMs) demonstrate the emerging abilities of "world models"---interpreting and reasoning about complex real-world dynamics. To assess these abilities, we posit videos are the ideal medium, as they encapsulate rich representations of real-world dynamics and causal…

2023

ESC: Exploration with Soft Commonsense Constraints for Zero-shot Object Navigation

ICML 2023poster

The ability to accurately locate and navigate to a specific object is a crucial capability for embodied agents that operate in the real world and interact with objects to complete tasks. Such object navigation tasks usually require large-scale training in visual environments with labeled objects, wh…

Cited by 115SourcePDFScholar
2023

R2H: Building Multimodal Navigation Helpers that Respond to Help Requests

EMNLP 2023long main

Intelligent navigation-helper agents are critical as they can navigate users in unknown areas through environmental awareness and conversational ability, serving as potential accessibility tools for individuals with disabilities. In this work, we first introduce a novel benchmark, Respond to Help Re…

Cited by 0SourceScholar
2022

Composable Causality in Semantic Robot Programming

ICRA 2022poster

Assembly tasks are challenging for robot manipulation because the robot must reason over the composed effects of actions and execute multi-objective behaviors. Robots typically use predefined priorities provided by users to determine how to compose controller behaviors, but we want the robot to auto…

Cited by 2SourceScholar
2022

VLMbench: A Compositional Benchmark for Vision-and-Language Manipulation

NeurIPS 2022accept

Benefiting from language flexibility and compositionality, humans naturally intend to use language to command an embodied agent for complex tasks such as navigation and object manipulation. In this work, we aim to fill the blank of the last mile of embodied agents---object manipulation by following…

Cited by 66SourcePDFScholar