← Search

Weiyu Liu

21 accepted papers

2026

ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction

ICLR 2026poster

Embodied cognition argues that intelligence arises from continuous sensorimotor interaction with the world. It raises an intriguing question: do modern vision-language models (VLMs), trained largely in a disembodied manner, exhibit signs of embodied cognition? To investigate this, we introduce **ENA…

Cited by 0SourcecodeScholar
2026

Learning Composable Skills by Discovering Spatial and Temporal Structure with Foundation Models

ICRA 2026poster

We present STACK, a framework for discovering and learning composable manipulation skills from unsegmented demonstrations by leveraging spatial and temporal structure extracted from foundation models. STACK automatically extracts temporal structure by segmenting raw demonstrations into short-horizon…

Cited by 0codeScholar
2026

MoMaGen: Generating Demonstrations under Soft and Hard Constraints for Multi-Step Bimanual Mobile Manipulation

ICLR 2026poster

Imitation learning from large-scale, diverse human demonstrations has been shown to be effective for training robots, but collecting such data is costly and time-consuming. This challenge intensifies for multi-step bimanual mobile manipulation, where humans must teleoperate both the mobile base and…

Cited by 0SourcecodeScholar
2025

LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models

CVPR 2025poster

Spatial reasoning is a fundamental aspect of human cognition, enabling intuitive understanding and manipulation of objects in three-dimensional space. While foundation models demonstrate remarkable performance on some benchmarks, they still struggle with 3D reasoning tasks like arranging objects in…

Cited by 10SourcePDFScholar
2024

Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making

NeurIPS 2024oral

We aim to evaluate Large Language Models (LLMs) for embodied decision making. While a significant body of work has been leveraging LLMs for decision making in embodied environments, we still lack a systematic understanding of their performance because they are usually applied in different domains, f…

Cited by 33SourcePDFScholar
2024

IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos

NeurIPS 2024poster

Shape assembly is a ubiquitous task in daily life, integral for constructing complex 3D structures like IKEA furniture. While significant progress has been made in developing autonomous agents for shape assembly, existing datasets have not yet tackled the 4D grounding of assembly instructions in vid…

2024

Learning Compositional Behaviors from Demonstration and Language

CoRL 2024poster

We introduce Behavior from Language and Demonstration (BLADE), a framework for long-horizon robotic manipulation by integrating imitation learning and model-based planning. BLADE leverages language-annotated demonstrations, extracts abstract action knowledge from large language models (LLMs), and co…

Cited by 3SourceScholar
2024

MARPLE: A Benchmark for Long-Horizon Inference

NeurIPS 2024poster

Reconstructing past events requires reasoning across long time horizons. To figure out what happened, humans draw on prior knowledge about the world and human behavior and integrate insights from various sources of evidence including visual, language, and auditory cues. We introduce MARPLE, a benchm…

2024

Naturally Supervised 3D Visual Grounding with Language-Regularized Concept Learners

CVPR 2024poster

3D visual grounding is a challenging task that often requires direct and dense supervision notably the semantic label for each object in the scene. In this paper we instead study the naturally supervised setting that learns from only 3D scene and QA pairs where prior works underperform. We propose t…

Cited by 7SourcePDFScholar
2023

GraspGPT: Leveraging Semantic Knowledge From a Large Language Model for Task-Oriented Grasping

RA-L 2023

Task-oriented grasping (TOG) refers to the problem of predicting grasps on an object that enable subsequent manipulation tasks. To model the complex relationships between objects, tasks, and grasps, existing methods incorporate semantic knowledge as priors into TOG pipelines. However, the existing s

Cited by 122SourceScholar
2023

StructDiffusion: Language-Guided Creation of Physically-Valid Structures using Unseen Objects

RSS 2023poster

Robots operating in human environments must be able to rearrange objects into semantically-meaningful configurations, even if these objects are previously unseen. In this work, we focus on the problem of building physically-valid structures without step-by-step instructions. We propose StructDiffusi…

Cited by 45SourcePDFScholar
2023

Task-Oriented Grasp Prediction with Visual-Language Inputs

IROS 2023poster

To perform household tasks, assistive robots receive commands in the form of user language instructions for tool manipulation. The initial stage involves selecting the intended tool (i.e., object grounding) and grasping it in a task-oriented manner (i.e., task grounding). Nevertheless, prior researc…

Cited by 40SourceScholar
2022

StructFormer: Learning Spatial Structure for Language-Guided Semantic Rearrangement of Novel Objects

ICRA 2022poster

Geometric organization of objects into semantically meaningful arrangements pervades the built world. As such, assistive robots operating in warehouses, offices, and homes would greatly benefit from the ability to recognize and rearrange objects into these semantically meaningful structures. To be u…

Cited by 104SourceScholar
2021

Learning Instance-Level N-Ary Semantic Knowledge At Scale For Robots Operating in Everyday Environments

RSS 2021poster

Robots operating in everyday environments need to effectively perceive; model; and infer semantic properties of objects. Existing knowledge reasoning frameworks only model binary relations between an object's class label and its semantic properties; unable to collectively reason about object propert…

2021

Towards Robust One-shot Task Execution using Knowledge Graph Embeddings

ICRA 2021poster

Requiring multiple demonstrations of a task plan presents a burden to end-users of robots. However, robustly executing tasks plans from a single end-user demonstration is an ongoing challenge in robotics. We address the problem of one-shot task execution, in which a robot must generalize a single de…

Cited by 21SourceScholar
2020

Same Object, Different Grasps: Data and Semantic Knowledge for Task-Oriented Grasping

CoRL 2020

Despite the enormous progress and generalization in robotic grasping in recent years, existing methods have yet to scale and generalize task-oriented grasping to the same extent. This is largely due to the scale of the datasets both in terms of the number of objects and tasks studied. We address the