← Search

Yong-Lu Li

44 accepted papers

2026

Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models

ICLR 2026poster

While Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in image and video understanding, their ability to comprehend the physical world has become an increasingly important research focus. Despite their improvements, current MLLMs struggle significantly with high-le…

Cited by 0SourceScholar
2026

Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols

CVPR 2026

Vision-Language-Action (VLA) models have recently achieved remarkable progress in robotic manipulation, yet they remain limited in failure diagnosis and learning from failures. Additionally, existing failure datasets are mostly generated programmatically in simulation, which limits their generalizat

Cited by 0SourcecodeScholar
2026

Motion before Action: Diffusing Object Motion As Manipulation Condition

ICRA 2026poster

Inferring object motion representations from observations enhances the performance of robotic manipulation tasks. This paper introduces a new paradigm for robot imitation learning that generates action sequences by reasoning about object motion from visual observations.We propose MBA, a novel module…

2026

OmniXtreme: Breaking the Generality Barrier in High-Dynamic Humanoid Control

RSS 2026poster

High-fidelity motion tracking serves as the ultimate litmus test for generalizable, human-level motor skills. However, current policies often hit a “generality barrier”: as motion libraries scale in diversity, tracking fidelity inevitably collapses—especially for real-world deployment of high-dynami…

Cited by 0SourceScholar
2026

SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning

ICML 2026poster

Pruning is a typical acceleration technique for compute-bound models by removing computation on unimportant values. Recently, it has been applied to accelerate Vision-Language-Action (VLA) model inference. However, existing acceleration methods focus on local information from the current action step…

Cited by 0SourceScholar
2026

Verb Mirage: Unveiling and Assessing Verb Concept Hallucinations in Multimodal Large Language Models

AAAI 2026technical

Multimodal Large Language Models (MLLMs) have garnered significant attention recently and demonstrate outstanding capabilities in various tasks such as OCR, VQA, captioning, etc. However, hallucination remains a persistent issue. While numerous methods have been proposed to mitigate hallucinations,

Cited by 0SourcePDFScholar
2025

Dense Policy: Bidirectional Autoregressive Learning of Actions

ICCV 2025poster

Mainstream visuomotor policies predominantly rely on generative models for holistic action prediction, while current autoregressive policies, predicting the next token or chunk, have shown suboptimal results. This motivates a search for more effective learning methods to unleash the potential of aut…

Cited by 0SourcePDFScholar
2025

Design2GarmentCode: Turning Design Concepts to Tangible Garments Through Program Synthesis

CVPR 2025poster

Sewing patterns, the essential blueprints for fabric cutting and tailoring, act as a crucial bridge between design concepts and producible garments. However, existing uni-modal sewing pattern generation models struggle to effectively encode complex design concepts with a multi-modal nature and corre…

2025

Homogeneous Dynamics Space for Heterogeneous Humans

CVPR 2025poster

Analyses of human motion kinematics have achieved tremendous advances. However, the production mechanism, known as human dynamics, is still undercovered. In this paper, we aim to push data-driven human dynamics understanding forward. We identify a major obstacle to this as the heterogeneity of exist…

2025

Human-Agent Joint Learning for Efficient Robot Manipulation Skill Acquisition

ICRA 2025

Employing a teleoperation system for gathering demonstrations offers the potential for more efficient learning of robot manipulation. However, teleoperating a robot arm equipped with a dexterous hand or gripper, via a teleoperation system presents inherent challenges due to the task's high dimension

Cited by 18SourcecodeScholar
2025

ImDy: Human Inverse Dynamics from Imitated Observations

ICLR 2025poster

Inverse dynamics (ID), which aims at reproducing the driven torques from human kinematic observations, has been a critical tool for gait analysis. However, it is hindered from wider application to general motion due to its limited scalability. Conventional optimization-based ID requires expensive la…

2025

Interacted Object Grounding in Spatio-Temporal Human-Object Interactions

AAAI 2025technical

Spatio-temporal Human-Object Interaction (ST-HOI) understanding aims at detecting HOIs from videos, which is crucial for activity understanding. However, existing whole-body-object interaction video benchmarks overlook the truth that open-world objects are diverse, that is, they usually provide limi…

2025

M^3-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation

CVPR 2025poster

Intelligent robots need to interact with diverse objects across various environments. The appearance and state of objects frequently undergo complex transformations depending on the object properties, e.g., phase transitions. However, in the vision community, segmenting dynamic objects with phase tr…

2025

Motion Before Action: Diffusing Object Motion as Manipulation Condition

RA-L 2025

Inferring object motion representations from observations enhances the performance of robotic manipulation tasks. This paper introduces a new paradigm for robot imitation learning that generates action sequences by reasoning about object motion from visual observations. We propose MBA (Motion Before

Cited by 16SourceScholar
2025

Reconstructing In-the-Wild Open-Vocabulary Human-Object Interactions

CVPR 2025poster

Reconstructing human-object interactions (HOI) from single images is fundamental in computer vision. Existing methods are primarily trained and tested on indoor scenes due to the lack of 3D data, particularly constrained by the object variety, making it challenging to generalize to real-world scenes…

Cited by 0SourcePDFScholar
2025

The Labyrinth of Links: Navigating the Associative Maze of Multi-modal LLMs

ICLR 2025poster

Multi-modal Large Language Models (MLLMs) have exhibited impressive capability. However, recently many deficiencies of MLLMs have been found compared to human intelligence, $\textit{e.g.}$, hallucination. To drive the MLLMs study, the community dedicated efforts to building larger benchmarks with co…

Cited by 0SourcePDFScholar
2025

exUMI: Extensible Robot Teaching System with Action-aware Task-agnostic Tactile Representation

CoRL 2025poster

Tactile-aware robot learning faces critical challenges in data collection and representation due to data scarcity and sparsity, and the absence of force feedback in existing systems. To address these limitations, we introduce a tactile robot learning system with both hardware and algorithm innovatio…

Cited by 0SourceScholar
2024

Dancing with Still Images: Video Distillation via Static-Dynamic Disentanglement

CVPR 2024poster

Recently dataset distillation has paved the way towards efficient machine learning especially for image datasets. However the distillation for videos characterized by an exclusive temporal dimension remains an underexplored domain. In this work we provide the first systematic study of video distilla…

2024

From Isolated Islands to Pangea: Unifying Semantic Space for Human Action Understanding

CVPR 2024highlight

Action understanding matters for intelligent agents and has attracted long-term attention. It can be formed as the mapping from the action physical space to the semantic space. Typically researchers built action datasets according to idiosyncratic choices to define classes and push the envelope of b…

Cited by 14SourcePDFScholar
2024

General Articulated Objects Manipulation in Real Images via Part-Aware Diffusion Process

NeurIPS 2024poster

Articulated object manipulation in real images is a fundamental step in computer and robotic vision tasks. Recently, several image editing methods based on diffusion models have been proposed to manipulate articulated objects according to text prompts. However, these methods often generate weird art…

Cited by 0SourcePDFScholar
2024

HumanVLA: Towards Vision-Language Directed Object Rearrangement by Physical Humanoid

NeurIPS 2024poster

Physical Human-Scene Interaction (HSI) plays a crucial role in numerous applications. However, existing HSI techniques are limited to specific object dynamics and privileged information, which prevents the development of more comprehensive applications. To address this limitation, we introd…

2024

Low-Rank Similarity Mining for Multimodal Dataset Distillation

ICML 2024poster

Though dataset distillation has witnessed rapid development in recent years, the distillation of multimodal data, e.g., image-text pairs, poses unique and under-explored challenges. Unlike unimodal data, image-text contrastive learning (ITC) data lack inherent categorization and should instead place…

2024

Primitive-Based 3D Human-Object Interaction Modelling and Programming

AAAI 2024technical

Embedding Human and Articulated Object Interaction (HAOI) in 3D is an important direction for a deeper human activity understanding. Different from previous works that use parametric and CAD models to represent humans and objects, in this work, we propose a novel 3D geometric primitive-based languag…

Cited by 3SourcePDFScholar
2023

Beyond Object Recognition: A New Benchmark towards Object Concept Learning

ICCV 2023poster

Understanding objects is a central building block of AI, especially for embodied AI. Even though object recognition excels with deep learning, current machines struggle to learn higher-level knowledge, e.g., what attributes an object has, and what we can do with it. Here, we propose a challenging Ob…

Cited by 9PDFScholar
2023

EgoPCA: A New Framework for Egocentric Hand-Object Interaction Understanding

ICCV 2023poster

With the surge in attention to Egocentric Hand-Object Interaction (Ego-HOI), large-scale datasets such as Ego4D and EPIC-KITCHENS have been proposed. However, most current research is built on resources derived from third-person video action recognition. This inherent domain gap between first- and t…

Cited by 13PDFScholar
2023

Symbol-LLM: Leverage Language Models for Symbolic System in Visual Human Activity Reasoning

NeurIPS 2023poster

Human reasoning can be understood as a cooperation between the intuitive, associative "System-1'' and the deliberative, logical "System-2''. For existing System-1-like methods in visual activity understanding, it is crucial to integrate System-2 processing to improve explainability, generalization,…

Cited by 15SourcePDFScholar
2022

Canonical Voting: Towards Robust Oriented Bounding Box Detection in 3D Scenes

CVPR 2022poster

3D object detection has attracted much attention thanks to the advances in sensors and deep learning methods for point clouds. Current state-of-the-art methods like VoteNet regress direct offset towards object centers and box orientations with an additional Multi-Layer-Perceptron network. Both their…

Cited by 15PDFcodeScholar
2022

Constructing Balance from Imbalance for Long-Tailed Image Recognition

ECCV 2022poster

"Long-tailed image recognition presents massive challenges to deep learning systems since the imbalance between majority (head) classes and minority (tail) classes severely skews the data-driven deep neural networks. Previous methods tackle with data imbalance from the viewpoints of data distributio…

2022

Highlighting Object Category Immunity for the Generalization of Human-Object Interaction Detection

AAAI 2022technical

Human-Object Interaction (HOI) detection plays a core role in activity understanding. As a compositional learning problem (human-verb-object), studying its generalization matters. However, widely-used metric mean average precision (mAP) fails to model the compositional generalization well. Thus, we…

2022

Human Trajectory Prediction With Momentary Observation

CVPR 2022poster

Human trajectory prediction task aims to analyze human future movements given their past status, which is a crucial step for many autonomous systems such as self-driving cars and social robots. In real-world scenarios, it is unlikely to obtain sufficiently long observations at all times for predicti…

Cited by 39PDFScholar
2022

Interactiveness Field in Human-Object Interactions

CVPR 2022poster

Human-Object Interaction (HOI) detection plays a core role in activity understanding. Though recent two/one-stage methods have achieved impressive results, as an essential step, discovering interactive human-object pairs remains challenging. Both one/two-stage methods fail to effectively extract int…

Cited by 65PDFcodeScholar
2022

Mining Cross-Person Cues for Body-Part Interactiveness Learning in HOI Detection

ECCV 2022poster

"Human-Object Interaction (HOI) detection plays a crucial role in activity understanding. Though significant progress has been made, interactiveness learning remains a challenging problem in HOI detection: existing methods usually generate redundant negative H-O pair proposals and fail to effectivel…

2022

UKPGAN: A General Self-Supervised Keypoint Detector

CVPR 2022poster

Keypoint detection is an essential component for the object registration and alignment. In this work, we reckon keypoint detection as information compression, and force the model to distill out important points of an object. Based on this, we propose UKPGAN, a general self-supervised 3D keypoint det…

Cited by 31PDFcodeScholar
2020

Detailed 2D-3D Joint Representation for Human-Object Interaction

CVPR 2020poster

Human-Object Interaction (HOI) detection lies at the core of action understanding. Besides 2D information such as human/object appearance and locations, 3D pose is also usually utilized in HOI learning since its view-independence. However, rough 3D body joints just carry sparse body information and…

Cited by 175PDFcodeScholar
2020

HOI Analysis: Integrating and Decomposing Human-Object Interaction

NeurIPS 2020poster

Human-Object Interaction (HOI) consists of human, object and implicit interaction/verb. Different from previous methods that directly map pixels to HOI semantics, we propose a novel perspective for HOI learning in an analytical manner. In analogy to Harmonic Analysis, whose goal is to study how to r…

2019

InstaBoost: Boosting Instance Segmentation via Probability Map Guided Copy-Pasting

ICCV 2019poster

Instance segmentation requires a large number of training samples to achieve satisfactory performance and benefits from proper data augmentation. To enlarge the training set and increase the diversity, previous methods have investigated using data annotation from other domain (e.g. bbox, point) in a…

Cited by 253PDFcodeScholar
2019

Transferable Interactiveness Knowledge for Human-Object Interaction Detection

CVPR 2019poster

Human-Object Interaction (HOI) Detection is an important problem to understand how humans interact with objects. In this paper, we explore Interactiveness Knowledge which indicates whether human and object interact with each other or not. We found that interactiveness knowledge can be learned across…

Cited by 383PDFcodeScholar