← Search

Ziyi Bai

3 accepted papers

2025

R2C: Mapping Room to Chessboard to Unlock LLM As Low-Level Action Planner

CVPR 2025poster

This paper explores using large language models (LLMs) as low-level action planners for embodied tasks. While LLMs excel as the robot's "brain" for high-level planning, they face challenges in directly controlling the "body" by generating precise low-level actions. This limitation arises from LLMs'…

2023

Glance and Focus: Memory Prompting for Multi-Event Video Question Answering

NeurIPS 2023poster

Video Question Answering (VideoQA) has emerged as a vital tool to evaluate agents’ ability to understand human daily behaviors. Despite the recent success of large vision language models in many multi-modal tasks, complex situation reasoning over videos involving multiple human-object interaction ev…

2021

Env-QA: A Video Question Answering Benchmark for Comprehensive Understanding of Dynamic Environments

ICCV 2021poster

Visual understanding goes well beyond the study of images or videos on the web. To achieve complex tasks in volatile situations, the human can deeply understand the environment, quickly perceive events happening around, and continuously track objects' state changes, which are still challenging for c…

Cited by 37PDFScholar