← Search

Zhengyuan Li

5 accepted papers

2026

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation

CVPR 2026

Recent advancements in vision-language-action (VLA) models have shown promise in robotic manipulation, yet they continue to struggle with long-horizon, multi-step tasks. Existing methods lack internal reasoning mechanisms that can identify task-relevant interaction cues or track progress within a su

Cited by 0SourceScholar
2025

EfficientEQA: An Efficient Approach to Open-Vocabulary Embodied Question Answering

IROS 2025

Embodied Question Answering (EQA) is an essential yet challenging task for robot assistants. Large vision-language models (VLMs) have shown promise for EQA, but existing approaches either treat it as static video question answering without active exploration or restrict answers to a closed set of ch

Cited by 11SourcecodeScholar
2025

MDD: A Dataset for Text-and-Music Conditioned Duet Dance Generation

ICCV 2025poster

We introduce Multimodal DuetDance (MDD), a diverse multimodal benchmark dataset designed for text-controlled and music-conditioned 3D duet dance motion generation. Our dataset comprises 620 minutes of high-quality motion capture data performed by professional dancers, synchronized with music, and de…

Cited by 0SourcePDFScholar
2025

SimMotionEdit: Text-Based Human Motion Editing with Motion Similarity Prediction

CVPR 2025poster

Text-based 3D human motion editing is a critical yet challenging task in computer vision and graphics. While training-free approaches have been explored, the recent release of the MotionFix dataset, which includes source-text-motion triplets, has opened new avenues for training, yielding promising r…

2023

InterDiff: Generating 3D Human-Object Interactions with Physics-Informed Diffusion

ICCV 2023poster

This paper addresses a novel task of anticipating 3D human-object interactions (HOIs). Most existing research on HOI synthesis lacks comprehensive whole-body interactions with dynamic objects, e.g., often limited to manipulating small or static objects. Our task is significantly more challenging, as…

Cited by 116PDFcodeScholar