← Search

Cewu Lu

190 accepted papers

2026

DICArt: Advancing Category-level Articulated Object Pose Estimation in Discrete State-Spaces

CVPR 2026

Articulated object pose estimation is a core task in embodied AI and computer vision. Existing methods typically regress poses in a continuous space, but often struggle with 1) navigating a large, complex search space and 2) failing to incorporate intrinsic kinematic constraints. In this paper, we i

Cited by 0SourceScholar
2026

Empowering Precise Embodied Agents with Executable Analytic Concepts as Semantic-Physical Blueprints

IJCAI 2026

A core challenge for embodied agents is the ``semantic-to-physical gap"—the difficulty of mapping symbolic reasoning to precise execution. While Vision-Language Models (VLMs) enhance agent task planning, they often fail in problem classes requiring accurate alignment between functional geometry and

Cited by 0Scholar
2026

Exploring Category-level Articulated Object Pose Tracking on SE(3) Manifolds

AAAI 2026technical

Articulated objects are prevalent in daily life and robotic manipulation tasks. However, compared to rigid objects, pose tracking for articulated objects remains an underexplored problem due to their inherent kinematic constraints. To address these challenges, this work proposes a novel point-pair-b

Cited by 0SourcePDFScholar
2026

Flow before Imitation: Learning Dexterous In-Hand Manipulation with Dynamic Visuotactile Shortcut Policy

ICRA 2026poster

Dexterous in-hand manipulation remains a long-standing challenge in robotics, primarily due to the complex contact dynamics and partial observability. While humans synergize vision and touch for such tasks, robotic approaches often prioritize one modality, therefore limiting adaptability. This paper…

Cited by 0Scholar
2026

Force Policy: Learning Hybrid Force-Position Control Policy under Interaction Frame for Contact-Rich Manipulation

RSS 2026poster

Contact-rich manipulation demands human-like integration of perception and force feedback: vision should guide task progress, while high-frequency interaction control must stabilize contact under uncertainty. Existing learning-based policies often entangle these roles in a monolithic network, tradin…

Cited by 0SourceScholar
2026

ForceVLA2: Unleashing Hybrid Force-Position Control with Force Awareness for Contact-Rich Manipulation

CVPR 2026

Embodied intelligence for contact-rich manipulation has predominantly relied on position control, while explicit awareness and regulation of interaction forces remain under-explored, limiting stability, precision, and robustness in real-world tasks. We propose ForceVLA2, an end-to-end vision-languag

Cited by 0SourceScholar
2026

HiWET: Hierarchical World-Frame End-Effector Tracking for Long-Horizon Humanoid Loco-Manipulation

RSS 2026poster

Humanoid loco-manipulation requires executing precise manipulation tasks while maintaining dynamic stability amid base motion and impacts. Existing approaches typically formulate commands in body-centric frames, fail to inherently correct cumulative world-frame drift induced by legged locomotion. We…

Cited by 0SourceScholar
2026

History-Aware Visuomotor Policy Learning Via Point Tracking

ICRA 2026poster

Many manipulation tasks require memory beyond the current observation, yet most visuomotor policies rely on the Markov assumption and thus struggle with repeated states or long-horizon dependencies. Existing methods attempt to extend observation horizons but remain insufficient for diverse memory re…

2026

Learning Dexterous Manipulation with Quantized Hand State

ICRA 2026poster

Dexterous robotic hands enable robots to perform complex manipulations that require fine-grained control and adaptability. Achieving such manipulation is challenging because the high degrees of freedom tightly couple hand and arm motions, making learning and control difficult. Successful dexterous m…

2026

Learning Realistic Depth via Physics-Grounded Noise Disentanglement with Semantic-Geometric Collaboration

ICML 2026poster

Real-world physical sensing exhibits complex, heterogeneous noise patterns that deviate significantly from idealized simulation, posing a fundamental bottleneck for sim-to-real transfer. Existing sensor modelings typically treat depth noise as a monolithic black-box process, overlooking the distinct…

Cited by 0SourceScholar
2026

Motion before Action: Diffusing Object Motion As Manipulation Condition

ICRA 2026poster

Inferring object motion representations from observations enhances the performance of robotic manipulation tasks. This paper introduces a new paradigm for robot imitation learning that generates action sequences by reasoning about object motion from visual observations.We propose MBA, a novel module…

2026

Physically Ground Commonsense Knowledge for Articulated Object Manipulation with Analytic Concepts

CVPR 2026

We humans rely on a wide range of commonsense knowledge to interact with an extensive number and categories of objects in the physical world. Likewise, such commonsense knowledge is also crucial for robots to successfully develop generalized object manipulation skills. While recent advancements in M

Cited by 0SourceScholar
2026

REArtGS++: Generalizable Articulation Reconstruction with Temporal Geometry Constraint via Planar Gaussian Splatting

CVPR 2026

Articulated objects are pervasive in daily environments, such as drawers and refrigerators. Towards their part-level surface reconstruction and joint parameter estimation, REArtGS [??] introduces a category-agnostic approach using multi-view RGB images at two different states. However, we observe th

Cited by 0SourceScholar
2026

Rethinking Camera Choice: An Empirical Study on Fisheye Camera Properties in Robotic Manipulation

CVPR 2026

The adoption of fisheye cameras in robotic manipulation, driven by their exceptionally wide Field of View (FoV), is rapidly outpacing a systematic understanding of their downstream effects on policy learning. This paper presents the first comprehensive empirical study to bridge this gap, rigorously

Cited by 0SourceScholar
2026

Right-Side-Out: Learning Zero-Shot Sim-To-Real Garment Reversal

ICRA 2026poster

Turning garments right-side out is a challenging manipulation task: it is highly dynamic, entails rapid contact changes, and is subject to severe visual occlusion. We introduce Right-Side-Out, a zero-shot sim-to-real framework that effectively solves this challenge by exploiting task structures. We …

2026

SOE: Sample-Efficient Robot Policy Self-Improvement Via On-Manifold Exploration

ICRA 2026poster

Intelligent agents progress by continually refining their capabilities through actively exploring environments. Yet robot policies often lack sufficient exploration capability due to action mode collapse. Existing methods that encourage exploration typically rely on random perturbations, which are u…

2026

Scaling by Diversified Experience for Vision-Language-Action Models

ICML 2026poster

Vision-Language-Action models face significant challenges in real-world deployment due to the entanglement of high-level reasoning with low-level control, and the instability of policy optimization. In this paper, we introduce SyVLA, a robust VLA model trained with diversified experiences. We propos…

Cited by 0SourceScholar
2026

Stereo-Inertial Poser: Towards Metric-Accurate Shape-Aware Motion Capture Using Sparse IMUs and a Single Stereo Camera

ICRA 2026poster

Recent advancements in visual-inertial motion capture systems have demonstrated the potential of combining monocular cameras with sparse inertial measurement units (IMUs) as cost-effective solutions, which effectively mitigate occlusion and drift issues inherent in single-modality systems. However, …

2026

TrajBooster: Boosting Humanoid Whole-Body Manipulation Via Trajectory-Centric Learning

ICRA 2026poster

Recent Vision-Language-Action (VLA) models show potential to generalize across embodiments but struggle to quickly align with a new robot’s action space when high-quality demonstrations are scarce, especially for bipedal humanoids. We present TrajBooster, a cross-embodiment framework that leverages …

2026

Verb Mirage: Unveiling and Assessing Verb Concept Hallucinations in Multimodal Large Language Models

AAAI 2026technical

Multimodal Large Language Models (MLLMs) have garnered significant attention recently and demonstrate outstanding capabilities in various tasks such as OCR, VQA, captioning, etc. However, hallucination remains a persistent issue. While numerous methods have been proposed to mitigate hallucinations,

Cited by 0SourcePDFScholar
2025

AgentWorld: An Interactive Simulation Platform for Scene Construction and Mobile Robotic Manipulation

CoRL 2025poster

We introduce AgentWorld, an interactive simulation platform for developing household mobile manipulation capabilities. Our platform combines automated scene construction that encompasses layout generation, semantic asset placement, visual material configuration, and physics simulation, with a dual-m…

Cited by 0SourceScholar
2025

AirExo-2: Scaling up Generalizable Robotic Imitation Learning with Low-Cost Exoskeletons

CoRL 2025oral

Scaling up robotic imitation learning for real-world applications requires efficient and scalable demonstration collection methods. While teleoperation is effective, it depends on costly and inflexible robot platforms. In-the-wild demonstrations offer a promising alternative, but existing collection…

Cited by 0SourceScholar
2025

ArtGS: 3D Gaussian Splatting for Interactive Visual-Physical Modeling and Manipulation of Articulated Objects

IROS 2025

Articulated object manipulation remains a critical challenge in robotics due to the complex kinematic constraints and the limited physical reasoning of existing methods. In this work, we introduce ArtGS, a novel framework that extends 3D Gaussian Splatting (3DGS) by integrating visual-physical model

Cited by 8SourceScholar
2025

Arti-PG: A Toolbox for Procedurally Synthesizing Large-Scale and Diverse Articulated Objects with Rich Annotations

ICCV 2025poster

The acquisition of substantial volumes of 3D articulated object data is expensive and time-consuming, and consequently the scarcity of 3D articulated object data becomes an obstacle for deep learning methods to achieve remarkable performance in various articulated object understanding tasks. Meanwhi…

2025

Cage: Causal Attention Enables Data-Efficient Generalizable Robotic Manipulation

ICRA 2025

Generalization in robotic manipulation remains a critical challenge, particularly when scaling to new environments with limited demonstrations. This paper introduces CAGE, a novel robotic manipulation policy designed to overcome these generalization barriers by integrating the pretrained visual repr

Cited by 20SourcecodeScholar
2025

ChatGarment: Garment Estimation, Generation and Editing via Large Language Models

CVPR 2025poster

We introduce ChatGarment, a novel approach that leverages large vision-language models (VLMs) to automate the estimation, generation, and editing of 3D garment sewing patterns from images or text descriptions. Unlike previous methods that often lack robustness and interactive editing capabilities, C…

Cited by 5SourcePDFScholar
2025

Deformpam: Data-Efficient Learning for Long-Horizon Deformable Object Manipulation Via Preference-Based Action Alignment

ICRA 2025

In recent years, imitation learning has made progress in the field of robotic manipulation. However, it still faces challenges when addressing complex long-horizon tasks with deformable objects, such as high-dimensional state spaces, complex dynamics, and multimodal action distributions. Traditional

Cited by 5SourcecodeScholar
2025

Dense Policy: Bidirectional Autoregressive Learning of Actions

ICCV 2025poster

Mainstream visuomotor policies predominantly rely on generative models for holistic action prediction, while current autoregressive policies, predicting the next token or chunk, have shown suboptimal results. This motivates a search for more effective learning methods to unleash the potential of aut…

Cited by 0SourcePDFScholar
2025

DexTOG: Learning Task-Oriented Dexterous Grasp With Language Condition

RA-L 2025

This study introduces a novel language-guided diffusion-based learning framework, DexTOG, aimed at advancing the field of task-oriented grasping (TOG) with dexterous hands. Unlike existing methods that mainly focus on 2-finger grippers, this research addresses the complexities of dexterous manipulat

Cited by 8SourceScholar
2025

DiffGen: Robot Demonstration Generation via Differentiable Physics Simulation, Differentiable Rendering, and Vision-Language Model

IROS 2025

Generating robot demonstrations through simulation is widely recognized as an effective way to scale up robot data. Previous work often trained reinforcement learning agents to generate expert policies, but this approach lacks sample efficiency. Recently, a line of work has attempted to generate rob

Cited by 3SourceScholar
2025

Discovering Conceptual Knowledge with Analytic Ontology Templates for Articulated Objects

AAAI 2025technical

Human cognition can leverage fundamental conceptual knowledge, like geometry and kinematic ones, to appropriately perceive, comprehend and interact with novel objects. Motivated by this finding, we aim to endow machine intelligence with an analogous capability through performing at the conceptual le…

2025

Dynamic Reconstruction of Hand-Object Interaction with Distributed Force-aware Contact Representation

ICCV 2025poster

We present ViTaM-D, a novel visual-tactile framework for reconstructing dynamic hand-object interaction with distributed tactile sensing to enhance contact modeling. Existing methods, relying solely on visual inputs, often fail to capture occluded interactions and object deformation. To address this…

Cited by 0SourcePDFScholar
2025

FSGlove: An Inertial-Based Hand Tracking System with Shape-Aware Calibration

IROS 2025

Accurate hand motion capture (MoCap) is vital for applications in robotics, virtual reality, and biomechanics, yet existing systems face limitations in capturing high-degree-of-freedom (DoF) joint kinematics and personalized hand shape. Commercial gloves offer up to 21 DoFs, which are insufficient f

Cited by 3SourceScholar
2025

FoAR: Force-Aware Reactive Policy for Contact-Rich Robotic Manipulation

RA-L 2025

Contact-rich tasks present significant challenges for robotic manipulation policies due to the complex dynamics of contact and the need for precise control. Vision-based policies often struggle with the skill required for such tasks, as they typically lack critical contact feedback modalities like f

Cited by 40SourceScholar
2025

ForceMimic: Force-Centric Imitation Learning with Force-Motion Capture System for Contact-Rich Manipulation

ICRA 2025

In most contact-rich manipulation tasks, humans apply time-varying forces to the target object, compensating for inaccuracies in the vision-guided hand trajectory. However, current robot learning algorithms primarily focus on trajectory-based policy, with limited attention given to learning force-re

Cited by 66SourcecodeScholar
2025

ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation

NeurIPS 2025poster

Vision-Language-Action (VLA) models have advanced general-purpose robotic manipulation by leveraging pretrained visual and linguistic representations. However, they struggle with contact-rich tasks that require fine-grained control involving force, especially under visual occlusion or dynamic uncert…

Cited by 0SourceScholar
2025

GaPT-DAR: Category-level Garments Pose Tracking via Integrated 2D Deformation and 3D Reconstruction

CVPR 2025poster

Garments are common in daily life and are important for embodied intelligence community. Current category-level garments pose tracking works focus on predicting point-wise canonical correspondence and learning a shape deformation in point cloud sequences. In this paper, motivated by the 2D warping s…

Cited by 0SourcePDFScholar
2025

Generalizable Articulated Object Perception with Superpoints

ICASSP 2025accepted

Manipulating articulated objects with robotic arms is challenging due to the complex kinematic structure, which requires precise part segmentation for efficient manipulation. In this work, we introduce a novel superpoint-based perception method designed to improve part segmentation in 3D point cloud…

Cited by 4SourceScholar
2025

Generalizable and Actionable Part Detection and Manipulation with SAM-rectified Segmentation and Iterative Pose Refinement

IROS 2025

The ability to perform cross-category object perception and manipulation is highly desirable in building intelligent robots. One promising approach is to define the concept of Generalizable and Actionable Parts (GAParts), such as buttons and handles, on both seen and unseen object categories. Howeve

Cited by 0SourceScholar
2025

Homogeneous Dynamics Space for Heterogeneous Humans

CVPR 2025poster

Analyses of human motion kinematics have achieved tremendous advances. However, the production mechanism, known as human dynamics, is still undercovered. In this paper, we aim to push data-driven human dynamics understanding forward. We identify a major obstacle to this as the heterogeneity of exist…

2025

Human-Agent Joint Learning for Efficient Robot Manipulation Skill Acquisition

ICRA 2025

Employing a teleoperation system for gathering demonstrations offers the potential for more efficient learning of robot manipulation. However, teleoperating a robot arm equipped with a dexterous hand or gripper, via a teleoperation system presents inherent challenges due to the task's high dimension

Cited by 18SourcecodeScholar
2025

ImDy: Human Inverse Dynamics from Imitated Observations

ICLR 2025poster

Inverse dynamics (ID), which aims at reproducing the driven torques from human kinematic observations, has been a critical tool for gait analysis. However, it is hindered from wider application to general motion due to its limited scalability. Conventional optimization-based ID requires expensive la…

2025

Interacted Object Grounding in Spatio-Temporal Human-Object Interactions

AAAI 2025technical

Spatio-temporal Human-Object Interaction (ST-HOI) understanding aims at detecting HOIs from videos, which is crucial for activity understanding. However, existing whole-body-object interaction video benchmarks overlook the truth that open-world objects are diverse, that is, they usually provide limi…

2025

Interactive Adjustment for Human Trajectory Prediction with Individual Feedback

ICLR 2025poster

Human trajectory prediction is fundamental for autonomous driving and service robot. The research community has studied various important aspects of this task and made remarkable progress recently. However, there is an essential perspective which is not well exploited in previous research all along,…

Cited by 0SourcePDFScholar
2025

Knowledge-Driven Imitation Learning: Enabling Generalization Across Diverse Conditions

IROS 2025

Imitation learning has emerged as a powerful paradigm in robot manipulation, yet its generalization capability remains constrained by object-specific dependencies in limited expert demonstrations. To address this challenge, we propose knowledge-driven imitation learning, a framework that leverages e

Cited by 1SourcecodeScholar
2025

LDexMM: Language-Guided Dexterous Multi-Task Manipulation with Reinforcement Learning

IROS 2025

Language plays a crucial role in robotic manipulation, particularly in facilitating complex tasks. Previous work primarily focused on two-finger manipulation. However, leveraging language to guide reinforcement learning for dexterous hands remains a challenge due to their high degrees of freedom. In

Cited by 0SourceScholar
2025

M^3-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation

CVPR 2025poster

Intelligent robots need to interact with diverse objects across various environments. The appearance and state of objects frequently undergo complex transformations depending on the object properties, e.g., phase transitions. However, in the vision community, segmenting dynamic objects with phase tr…

2025

Motion Before Action: Diffusing Object Motion as Manipulation Condition

RA-L 2025

Inferring object motion representations from observations enhances the performance of robotic manipulation tasks. This paper introduces a new paradigm for robot imitation learning that generates action sequences by reasoning about object motion from visual observations. We propose MBA (Motion Before

Cited by 16SourceScholar
2025

Novel Demonstration Generation with Gaussian Splatting Enables Robust One-Shot Manipulation

RSS 2025poster

Visuomotor policies learned through imitation learning methods often struggle to generalize to new visual domains due to the limited diversity of expert demonstrations, and collecting extensive real-world data is exhaustive. To address this challenge, we propose a novel demonstration generation app…

Cited by 1PDFScholar
2025

REArtGS: Reconstructing and Generating Articulated Objects via 3D Gaussian Splatting with Geometric and Motion Constraints

NeurIPS 2025poster

Articulated objects, as prevalent entities in human life, their 3D representations play crucial roles across various applications. However, achieving both high-fidelity textured surface reconstruction and dynamic generation for articulated objects remains challenging for existing methods. In this pa…

Cited by 0SourceScholar
2025

RH20T-P: A Primitive-Level Robotic Manipulation Dataset towards Composable Generalization Agents in Real-world Scenarios

IROS 2025

Achieving generalizability in solving out-of-distribution tasks is one of the ultimate goals of learning robotic manipulation. Recent progress of Vision-Language Models (VLMs) has shown that VLM-based task planners can alleviate the difficulty of solving novel tasks, by decomposing the compounded ta

Cited by 1SourceScholar
2025

Reactive Diffusion Policy: Slow-Fast Visual-Tactile Policy Learning for Contact-Rich Manipulation

RSS 2025poster

Humans can accomplish complex contact-rich tasks using vision and touch, with highly reactive capabilities such as quick adjustments to environmental changes and adaptive control of contact forces; however, this remains challenging for robots. Existing visual imitation learning (IL) approaches rely…

Cited by 8PDFcodeScholar
2025

SIME: Enhancing Policy Self-Improvement with Modal-level Exploration

IROS 2025

Self-improvement requires robotic systems to initially learn from human-provided data and then gradually enhance their capabilities through interaction with the environment. This is similar to how humans improve their skills through continuous practice. However, achieving effective self-improvement

Cited by 4SourcecodeScholar
2025

SKT: Integrating State-Aware Keypoint Trajectories with Vision-Language Models for Robotic Garment Manipulation

IROS 2025

Automating garment manipulation poses a significant challenge for assistive robotics due to the diverse and de-formable nature of garments. Traditional approaches typically require separate models for each garment type, which limits scalability and adaptability. In contrast, this paper presents a un

Cited by 3SourceScholar
2025

The Labyrinth of Links: Navigating the Associative Maze of Multi-modal LLMs

ICLR 2025poster

Multi-modal Large Language Models (MLLMs) have exhibited impressive capability. However, recently many deficiencies of MLLMs have been found compared to human intelligence, $\textit{e.g.}$, hallucination. To drive the MLLMs study, the community dedicated efforts to building larger benchmarks with co…

Cited by 0SourcePDFScholar
2025

Towards Effective Utilization of Mixed-Quality Demonstrations in Robotic Manipulation via Segment-Level Selection and Optimization

ICRA 2025

Data is crucial for robotic manipulation, as it underpins the development of robotic systems for complex tasks. While high-quality, diverse datasets enhance the performance and adaptability of robotic manipulation policies, collecting extensive expert-level data is resource-intensive. Consequently,

Cited by 6SourceScholar
2025

Tru-POMDP: Task Planning Under Uncertainty via Tree of Hypotheses and Open-Ended POMDPs

NeurIPS 2025poster

Task planning under uncertainty is essential for home-service robots operating in the real world. Tasks involve ambiguous human instructions, hidden or unknown object locations, and open-vocabulary object types, leading to significant open-ended uncertainty and a boundlessly large planning space. To…

Cited by 0SourceScholar
2025

UniAff: A Unified Representation of Affordances for Tool Usage and Articulation with Vision-Language Models

ICRA 2025

Previous studies on robotic manipulation are based on a limited understanding of the underlying 3D motion constraints and affordances. To address these challenges, we propose a comprehensive paradigm, termed UniAff, that integrates 3D object-centric manipulation and task understanding in a unified f

Cited by 10SourceScholar
2025

UniDomain: Pretraining a Unified PDDL Domain from Real-World Demonstrations for Generalizable Robot Task Planning

NeurIPS 2025poster

Robotic task planning in real-world environments requires reasoning over implicit constraints from language and vision. While LLMs and VLMs offer strong priors, they struggle with long-horizon structure and symbolic grounding. Existing meth- ods that combine LLMs with symbolic planning often rely on…

Cited by 0SourceScholar
2024

AirExo: Low-Cost Exoskeletons for Learning Whole-Arm Manipulation in the Wild

ICRA 2024poster

While humans can use parts of their arms other than the hands for manipulations like gathering and supporting, whether robots can effectively learn and perform the same type of operations remains relatively unexplored. As these manipulations require joint-level control to regulate the complete poses…

Cited by 40SourcecodeScholar
2024

COIN: Control-Inpainting Diffusion Prior for Human and Camera Motion Estimation

ECCV 2024poster

"Estimating global human motion from moving cameras is challenging due to the entanglement of human and camera motions. To mitigate the ambiguity, existing methods leverage learned human motion priors, which however often result in oversmoothed motions with misaligned 2D projections. To tackle this…

2024

ConceptFactory: Facilitate 3D Object Knowledge Annotation with Object Conceptualization

NeurIPS 2024poster

We present ConceptFactory, a novel scope to facilitate more efficient annotation of 3D object knowledge by recognizing 3D objects through generalized concepts (i.e. object conceptualization), aiming at promoting machine intelligence to learn comprehensive object knowledge from both vision and roboti…

2024

Dancing with Still Images: Video Distillation via Static-Dynamic Disentanglement

CVPR 2024poster

Recently dataset distillation has paved the way towards efficient machine learning especially for image datasets. However the distillation for videos characterized by an exclusive temporal dimension remains an underexplored domain. In this work we provide the first systematic study of video distilla…

2024

DiPGrasp: Parallel Local Searching for Efficient Differentiable Grasp Planning

RA-L 2024

Grasp planning is an important task for robotic manipulation. Though it is a richly studied area, a standalone, fast, and differentiable grasp planner that can work with robot grippers of different DOFs has not been reported. In this work, we present DiPGrasp, a grasp planner that satisfies all thes

Cited by 8SourceScholar
2024

Differentiable Cloth Parameter Identification and State Estimation in Manipulation

RA-L 2024

In the realm of robotic cloth manipulation, accurately estimating the cloth state during or post-execution is imperative. However, the inherent complexities in a cloth's dynamic behavior and its near-infinite degrees of freedom (DoF) pose significant challenges. Traditional methods have been restric

Cited by 14SourceScholar
2024

Differentiable Fluid Physics Parameter Identification By Stirring and For Stirring

IROS 2024poster

Fluid interactions are crucial in daily tasks, with properties like density and viscosity being key parameters. The property states can be used as control signals for robot operation. While density estimation is simple, assessing viscosity, especially for different fluid types, is complex. This stud…

Cited by 0SourceScholar
2024

Distill Gold from Massive Ores: Bi-level Data Pruning towards Efficient Dataset Distillation

ECCV 2024poster

"Data-efficient learning has garnered significant attention, especially given the current trend of large multi-modal models. Recently, dataset distillation has become an effective approach by synthesizing data samples that are essential for network training. However, it remains to be explored which…

2024

EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI

CVPR 2024poster

In the realm of computer vision and robotics embodied agents are expected to explore their environment and carry out human instructions. This necessitates the ability to fully understand 3D scenes given their first-person observations and contextualize them into language for interaction. However tra…

2024

FAVOR: Full-Body AR-Driven Virtual Object Rearrangement Guided by Instruction Text

AAAI 2024technical

Rearrangement operations form the crux of interactions between humans and their environment. The ability to generate natural, fluid sequences of this operation is of essential value in AR/VR and CG. Bridging a gap in the field, our study introduces FAVOR: a novel dataset for Full-body AR-driven Virt…

2024

From Isolated Islands to Pangea: Unifying Semantic Space for Human Action Understanding

CVPR 2024highlight

Action understanding matters for intelligent agents and has attracted long-term attention. It can be formed as the mapping from the action physical space to the semantic space. Typically researchers built action datasets according to idiosyncratic choices to define classes and push the envelope of b…

Cited by 14SourcePDFScholar
2024

GAMMA: Generalizable Articulation Modeling and Manipulation for Articulated Objects

ICRA 2024poster

Articulated objects like cabinets and doors are widespread in daily life. However, directly manipulating 3D articulated objects is challenging because they have diverse geometrical shapes, semantic categories, and kinetic constraints. Prior works mostly focused on recognizing and manipulating articu…

Cited by 15SourcecodeScholar
2024

General Articulated Objects Manipulation in Real Images via Part-Aware Diffusion Process

NeurIPS 2024poster

Articulated object manipulation in real images is a fundamental step in computer and robotic vision tasks. Recently, several image editing methods based on diffusion models have been proposed to manipulate articulated objects according to text prompts. However, these methods often generate weird art…

Cited by 0SourcePDFScholar
2024

HumanVLA: Towards Vision-Language Directed Object Rearrangement by Physical Humanoid

NeurIPS 2024poster

Physical Human-Scene Interaction (HSI) plays a crucial role in numerous applications. However, existing HSI techniques are limited to specific object dynamics and privileged information, which prevents the development of more comprehensive applications. To address this limitation, we introd…

2024

Intersection-Free Robot Manipulation With Soft-Rigid Coupled Incremental Potential Contact

RA-L 2024

This paper presents a novel simulation platform, ZeMa, designed for robotic manipulation tasks concerning soft objects. Such simulation ideally requires three properties: two-way soft-rigid coupling, intersection-free guarantee, and frictional contact modeling, with acceptable runtime suitable for l

Cited by 9SourceScholar
2024

Low-Rank Similarity Mining for Multimodal Dataset Distillation

ICML 2024poster

Though dataset distillation has witnessed rapid development in recent years, the distillation of multimodal data, e.g., image-text pairs, poses unique and under-explored challenges. Unlike unimodal data, image-text contrastive learning (ITC) data lack inherent categorization and should instead place…

2024

MS-MANO: Enabling Hand Pose Tracking with Biomechanical Constraints

CVPR 2024poster

This work proposes a novel learning framework for visual hand dynamics analysis that takes into account the physiological aspects of hand motion. The existing models which are simplified joint-actuated systems often produce unnatural motions. To address this we integrate a musculoskeletal system wit…

Cited by 7SourcePDFScholar
2024

OAKINK2: A Dataset of Bimanual Hands-Object Manipulation in Complex Task Completion

CVPR 2024poster

We present OAKINK2 a dataset of bimanual object manipulation tasks for complex daily activities. In pursuit of constructing the complex tasks into a structured representation OAKINK2 introduces three level of abstraction to organize the manipulation tasks: Affordance Primitive Task and Complex Task.…

Cited by 18SourcePDFScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration

ICRA 2024

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man

Cited by 910SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration0

ICRA 2024poster

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man…

Cited by 259SourcecodeScholar
2024

Primitive-Based 3D Human-Object Interaction Modelling and Programming

AAAI 2024technical

Embedding Human and Articulated Object Interaction (HAOI) in 3D is an important direction for a deeper human activity understanding. Different from previous works that use parametric and CAD models to represent humans and objects, in this work, we propose a novel 3D geometric primitive-based languag…

Cited by 3SourcePDFScholar
2024

RFTrans: Leveraging Refractive Flow of Transparent Objects for Surface Normal Estimation and Manipulation

RA-L 2024

Transparent objects are widely used in our daily lives, making it important to teach robots to interact with them. However, it's not easy because the reflective and refractive effects can make depth cameras fail to give accurate geometry measurements. To solve this problem, this paper introduces RFT

Cited by 11SourceScholar
2024

RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot

ICRA 2024poster

A key challenge for robotic manipulation in open domains is how to acquire diverse and generalizable skills for robots. Recent progress in one-shot imitation learning and robotic foundation models have shown promise in transferring trained policies to new tasks based on demonstrations. This feature…

Cited by 86SourcecodeScholar
2024

RISE: 3D Perception Makes Real-World Robot Imitation Simple and Effective

IROS 2024poster

Precise robot manipulations require rich spatial information in imitation learning. Image-based policies model object positions from fixed cameras, which are sensitive to camera view changes. Policies utilizing 3D point clouds usually predict keyframes rather than continuous actions, posing difficul…

Cited by 24SourcecodeScholar
2024

RPMArt: Towards Robust Perception and Manipulation for Articulated Objects

IROS 2024poster

Articulated objects are commonly found in daily life. It is essential that robots can exhibit robust perception and manipulation skills for articulated objects in real-world robotic applications. However, existing methods for articulated objects insufficiently address noise in point clouds and strug…

Cited by 4SourcecodeScholar
2024

Revisit Human-Scene Interaction via Space Occupancy

ECCV 2024poster

"Human-scene Interaction (HSI) generation is a challenging task and crucial for various downstream tasks. However, one of the major obstacles is its limited data scale. High-quality data with simultaneously captured human and 3D environments is hard to acquire, resulting in limited data diversity an…

2024

ShapeBoost: Boosting Human Shape Estimation with Part-Based Parameterization and Clothing-Preserving Augmentation

AAAI 2024technical

Accurate human shape recovery from a monocular RGB image is a challenging task because humans come in different shapes and sizes and wear different clothes. In this paper, we propose ShapeBoost, a new human shape recovery framework that achieves pixel-level alignment even for rare body shapes and hi…

Cited by 2SourcePDFScholar
2024

TacIPC: Intersection- and Inversion-Free FEM-Based Elastomer Simulation for Optical Tactile Sensors

RA-L 2024

Tactile perception stands as a critical sensory modality for human interaction with the environment. Among various tactile sensor techniques, optical sensor-based approaches have gained traction, notably for producing high-resolution tactile images. This letter explores gel elastomer deformation sim

Cited by 17SourceScholar
2024

Take A Step Back: Rethinking the Two Stages in Visual Reasoning

ECCV 2024poster

"As a prominent research area, visual reasoning plays a crucial role in AI by facilitating concept formation and interaction with the world. However, current works are usually carried out separately on small datasets thus lacking generalization ability. Through rigorous evaluation of diverse benchma…

2024

TieBot: Learning to Knot a Tie from Visual Demonstration through a Real-to-Sim-to-Real Approach

CoRL 2024poster

The tie-knotting task is highly challenging due to the tie's high deformation and long-horizon manipulation actions. This work presents TieBot, a Real-to-Sim-to-Real learning from visual demonstration system for the robots to learn to knot a tie. We introduce the Hierarchical Feature Matching approa…

Cited by 2SourcecodeScholar
2023

Beyond Object Recognition: A New Benchmark towards Object Concept Learning

ICCV 2023poster

Understanding objects is a central building block of AI, especially for embodied AI. Even though object recognition excels with deep learning, current machines struggle to learn higher-level knowledge, e.g., what attributes an object has, and what we can do with it. Here, we propose a challenging Ob…

Cited by 9PDFScholar
2023

CHORD: Category-level Hand-held Object Reconstruction via Shape Deformation

ICCV 2023poster

In daily life, humans utilize hands to manipulate objects. Modeling the shape of objects that are manipulated by the hand is essential for AI to comprehend daily tasks and to learn manipulation skills. However, previous approaches have encountered difficulties in reconstructing the precise shapes of…

Cited by 15PDFcodeScholar
2023

CRIN: Rotation-Invariant Point Cloud Analysis and Rotation Estimation via Centrifugal Reference Frame

AAAI 2023technical

Various recent methods attempt to implement rotation-invariant 3D deep learning by replacing the input coordinates of points with relative distances and angles. Due to the incompleteness of these low-level features, they have to undertake the expense of losing global information. In this paper, we p…

2023

ClothPose: A Real-world Benchmark for Visual Analysis of Garment Pose via An Indirect Recording Solution

ICCV 2023oral

Garments are important and pervasive in daily life. However, visual analysis on them for pose estimation is challenging because it requires recovering the complete configurations of garments, which is difficult, if not impossible, to annotate in the real world. In this work, we propose a recording s…

Cited by 6PDFScholar
2023

ClothesNet: An Information-Rich 3D Garment Model Repository with Simulated Clothes Environment

ICCV 2023poster

We present ClothesNet: a large-scale dataset of 3D clothes objects with information-rich annotations. Our dataset consists of around 4000 models covering 11 categories annotated with clothes features, boundary lines, and keypoints. ClothesNet can be used to facilitate a variety of computer vision an…

Cited by 16PDFScholar
2023

Demonstrating RFUniverse: A Multiphysics Simulation Platform for Embodied AI

RSS 2023poster

Multiphysics phenomena, the coupling effects involving different aspects of physics laws, are pervasive in the real world and can often be encountered when performing everyday household tasks. Intelligent agents which seek to assist or replace human laborers will need to learn to cope with such phe…

2023

Diff-LfD: Contact-aware Model-based Learning from Visual Demonstration for Robotic Manipulation via Differentiable Physics-based Simulation and Rendering

CoRL 2023oral

Learning from Demonstration (LfD) is an efficient technique for robots to acquire new skills through expert observation, significantly mitigating the need for laborious manual reward function design. This paper introduces a novel framework for model-based LfD in the context of robotic manipulation.…

Cited by 19SourceScholar
2023

EgoPCA: A New Framework for Egocentric Hand-Object Interaction Understanding

ICCV 2023poster

With the surge in attention to Egocentric Hand-Object Interaction (Ego-HOI), large-scale datasets such as Ego4D and EPIC-KITCHENS have been proposed. However, most current research is built on resources derived from third-person video action recognition. This inherent domain gap between first- and t…

Cited by 13PDFScholar
2023

Flexible Handover with Real-Time Robust Dynamic Grasp Trajectory Generation

IROS 2023poster

In recent years, there has been a significant effort dedicated to developing efficient, robust, and general human-to-robot handover systems. However, the area of flexible handover in the context of complex and continuous objects' motion remains relatively unexplored. In this work, we propose an appr…

Cited by 8SourceScholar
2023

GarmentTracking: Category-Level Garment Pose Tracking

CVPR 2023poster

Garments are important to humans. A visual system that can estimate and track the complete garment pose can be useful for many downstream tasks and real-world applications. In this work, we present a complete package to address the category-level garment pose tracking task: (1) A recording system VR…

2023

NIKI: Neural Inverse Kinematics With Invertible Neural Networks for 3D Human Pose and Shape Estimation

CVPR 2023poster

With the progress of 3D human pose and shape estimation, state-of-the-art methods can either be robust to occlusions or obtain pixel-aligned accuracy in non-occlusion cases. However, they cannot obtain robustness and mesh-image alignment at the same time. In this work, we present NIKI (Neural Invers…

2023

POEM: Reconstructing Hand in a Point Embedded Multi-View Stereo

CVPR 2023poster

Enable neural networks to capture 3D geometrical-aware features is essential in multi-view based vision tasks. Previous methods usually encode the 3D information of multi-view stereo into the 2D features. In contrast, we present a novel method, named POEM, that directly operates on the 3D POints Emb…

2023

Precise Robotic Needle-Threading with Tactile Perception and Reinforcement Learning

CoRL 2023poster

This work presents a novel tactile perception-based method, named T-NT, for performing the needle-threading task, an application of deformable linear object (DLO) manipulation. This task is divided into two main stages: \textit{Tail-end Finding} and \textit{Tail-end Insertion}. In the first stage, t…

Cited by 8SourceScholar
2023

SAM-RL: Sensing-Aware Model-Based Reinforcement Learning via Differentiable Physics-Based Simulation and Rendering

RSS 2023poster

Model-based reinforcement learning (MBRL) is recognized with the potential to be significantly more sample efficient than model-free RL. How an accurate model can be developed automatically and efficiently from raw sensory inputs (such as images), especially for complex environments and tasks, is a…

Cited by 28SourcePDFScholar
2023

Stimulus Verification Is a Universal and Effective Sampler in Multi-Modal Human Trajectory Prediction

CVPR 2023poster

To comprehensively cover the uncertainty of the future, the common practice of multi-modal human trajectory prediction is to first generate a set/distribution of candidate future trajectories and then sample required numbers of trajectories from them as final predictions. Even though a large number…

Cited by 17SourcePDFScholar
2023

Symbol-LLM: Leverage Language Models for Symbolic System in Visual Human Activity Reasoning

NeurIPS 2023poster

Human reasoning can be understood as a cooperation between the intuitive, associative "System-1'' and the deliberative, logical "System-2''. For existing System-1-like methods in visual activity understanding, it is crucial to integrate System-2 processing to improve explainability, generalization,…

Cited by 15SourcePDFScholar
2023

Target-Referenced Reactive Grasping for Dynamic Objects

CVPR 2023poster

Reactive grasping, which enables the robot to successfully grasp dynamic moving objects, is of great interest in robotics. Current methods mainly focus on the temporal smoothness of the predicted grasp poses but few consider their semantic consistency. Consequently, the predicted grasps are not guar…

Cited by 14SourcePDFScholar
2023

UniFolding: Towards Sample-efficient, Scalable, and Generalizable Robotic Garment Folding

CoRL 2023poster

This paper explores the development of UniFolding, a sample-efficient, scalable, and generalizable robotic system for unfolding and folding various garments. UniFolding employs the proposed UFONet neural network to integrate unfolding and folding decisions into a single policy model that is adaptab…

Cited by 15SourceScholar
2023

Unsupervised 3D Point Cloud Representation Learning by Triangle Constrained Contrast for Autonomous Driving

CVPR 2023poster

Due to the difficulty of annotating the 3D LiDAR data of autonomous driving, an efficient unsupervised 3D representation learning method is important. In this paper, we design the Triangle Constrained Contrast (TriCC) framework tailored for autonomous driving scenes which learns 3D unsupervised repr…

2023

Upcycling Models Under Domain and Category Shift

CVPR 2023poster

Deep neural networks (DNNs) often perform poorly in the presence of domain shift and category shift. How to upcycle DNNs and adapt them to the target task remains an important open problem. Unsupervised Domain Adaptation (UDA), especially recently proposed Source-free Domain Adaptation (SFDA), has b…

2023

Visual-Tactile Sensing for In-Hand Object Reconstruction

CVPR 2023poster

Tactile sensing is one of the modalities human rely on heavily to perceive the world. Working with vision, this modality refines local geometry structure, measures deformation at contact area, and indicates hand-object contact state. With the availability of open-source tactile sensors such as DIGIT…

Cited by 24SourcePDFScholar
2022

AKB-48: A Real-World Articulated Object Knowledge Base

CVPR 2022poster

Human life is populated with articulated objects. A comprehensive understanding of articulated objects, namely appearance, structure, physics property, and semantics, will benefit many research communities. As current articulated object understanding solutions are usually based on synthetic object d…

Cited by 91PDFcodeScholar
2022

ArtiBoost: Boosting Articulated 3D Hand-Object Pose Estimation via Online Exploration and Synthesis

CVPR 2022oral

Estimating the articulated 3D hand-object pose from a single RGB image is a highly ambiguous and challenging problem, requiring large-scale datasets that contain diverse hand poses, object types, and camera viewpoints. Most real-world datasets lack these diversities. In contrast, data synthesis can…

Cited by 98PDFcodeScholar
2022

Canonical Voting: Towards Robust Oriented Bounding Box Detection in 3D Scenes

CVPR 2022poster

3D object detection has attracted much attention thanks to the advances in sensors and deep learning methods for point clouds. Current state-of-the-art methods like VoteNet regress direct offset towards object centers and box orientations with an additional Multi-Layer-Perceptron network. Both their…

Cited by 15PDFcodeScholar
2022

Constructing Balance from Imbalance for Long-Tailed Image Recognition

ECCV 2022poster

"Long-tailed image recognition presents massive challenges to deep learning systems since the imbalance between majority (head) classes and minority (tail) classes severely skews the data-driven deep neural networks. Previous methods tackle with data imbalance from the viewpoints of data distributio…

2022

Correlation Field for Boosting 3D Object Detection in Structured Scenes

AAAI 2022technical

Data augmentation is an efficient way to elevate 3D object detection performance. In this paper, we propose a simple but effective online crop-and-paste data augmentation pipeline for structured 3D point cloud scenes, named CorrelaBoost. Observing that 3D objects should have reasonable relative posi…

Cited by 10SourcePDFScholar
2022

D&D: Learning Human Dynamics from Dynamic Camera

ECCV 2022poster

"3D human pose estimation from a monocular video has recently seen significant improvements. However, most state-of-the-art methods are kinematics-based, which are prone to physically implausible motions with pronounced artifacts. Current dynamics-based methods can predict physically plausible motio…

2022

DART: Articulated Hand Model with Diverse Accessories and Rich Textures

NeurIPS 2022accept

Hand, the bearer of human productivity and intelligence, is receiving much attention due to the recent fever of digital twins. Among different hand morphable models, MANO has been widely used in vision and graphics community. However, MANO disregards textures and accessories, which largely limits it…

2022

Highlighting Object Category Immunity for the Generalization of Human-Object Interaction Detection

AAAI 2022technical

Human-Object Interaction (HOI) detection plays a core role in activity understanding. As a compositional learning problem (human-verb-object), studying its generalization matters. However, widely-used metric mean average precision (mAP) fails to model the compositional generalization well. Thus, we…

2022

Human Trajectory Prediction With Momentary Observation

CVPR 2022poster

Human trajectory prediction task aims to analyze human future movements given their past status, which is a crucial step for many autonomous systems such as self-driving cars and social robots. In real-world scenarios, it is unlikely to obtain sufficiently long observations at all times for predicti…

Cited by 39PDFScholar
2022

Interactiveness Field in Human-Object Interactions

CVPR 2022poster

Human-Object Interaction (HOI) detection plays a core role in activity understanding. Though recent two/one-stage methods have achieved impressive results, as an essential step, discovering interactive human-object pairs remains challenging. Both one/two-stage methods fail to effectively extract int…

Cited by 65PDFcodeScholar
2022

Mining Cross-Person Cues for Body-Part Interactiveness Learning in HOI Detection

ECCV 2022poster

"Human-Object Interaction (HOI) detection plays a crucial role in activity understanding. Though significant progress has been made, interactiveness learning remains a challenging problem in HOI detection: existing methods usually generate redundant negative H-O pair proposals and fail to effectivel…

2022

OakInk: A Large-Scale Knowledge Repository for Understanding Hand-Object Interaction

CVPR 2022poster

Learning how humans manipulate objects requires machines to acquire knowledge from two perspectives: one for understanding object affordances and the other for learning human's interactions based on the affordances. Even though these two knowledge bases are crucial, we find that current databases la…

Cited by 97PDFcodeScholar
2022

RCare World: A Human-centric Simulation World for Caregiving Robots

IROS 2022poster

We present RCareWorld, a human-centric simulation world for physical and social robotic caregiving designed with inputs from stakeholders. RCareWorld has realistic human models of care recipients with mobility limitations and caregivers, home environments with multiple levels of accessibility and as…

Cited by 43SourceScholar
2022

RoboTube: Learning Household Manipulation from Human Videos with Simulated Twin Environments

CoRL 2022oral

We aim to build a useful, reproducible, democratized benchmark for learning household robotic manipulation from human videos. To realize this goal, a diverse, high-quality human video dataset curated specifically for robots is desired. To evaluate the learning progress, a simulated twin environment…

Cited by 12SourceScholar
2022

SAGCI-System: Towards Sample-Efficient, Generalizable, Compositional, and Incremental Robot Learning

ICRA 2022poster

Building general-purpose robots to perform a diverse range of tasks in a large variety of environments in the physical world at the human level is extremely challenging. According to [1], it requires the robot learning to be sample-efficient, generalizable, compositional, and incremental. In this wo…

Cited by 28SourceScholar
2022

TransCG: A Large-Scale Real-World Dataset for Transparent Object Depth Completion and a Grasping Baseline

RA-L 2022

Transparent objects are common in our daily life and frequently handled in the automated production line. Robust vision-based robotic grasping and manipulation for these objects would be beneficial for automation. However, the majority of current grasping algorithms would fail in this case since the

Cited by 120SourcecodeScholar
2022

UKPGAN: A General Self-Supervised Keypoint Detector

CVPR 2022poster

Keypoint detection is an essential component for the object registration and alignment. In this work, we reckon keypoint detection as information compression, and force the model to distill out important points of an object. Based on this, we propose UKPGAN, a general self-supervised 3D keypoint det…

Cited by 31PDFcodeScholar
2022

Unified and Fast Human Trajectory Prediction Via Conditionally Parameterized Normalizing Flow

RA-L 2022

Human trajectory prediction is crucial for service robots, autonomous driving and advanced driver assistant systems. Current top-performing methods mainly rely on intractable generative models to learn a distribution of future trajectories, and sample multiple plausible ones as prediction results. I

Cited by 14SourceScholar
2022

Unsupervised Representation for Semantic Segmentation by Implicit Cycle-Attention Contrastive Learning

AAAI 2022technical

We study the unsupervised representation learning for the semantic segmentation task. Different from previous works that aim at providing unsupervised pre-trained backbones for segmentation models which need further supervised fine-tune, here, we focus on providing representation that is only traine…

Cited by 11SourcePDFScholar
2022

Unsupervised Visual Representation Learning by Synchronous Momentum Grouping

ECCV 2022poster

"In this paper, we propose a genuine group-level contrastive visual representation learning method whose linear evaluation performance on ImageNet surpasses the vanilla supervised learning. Two mainstream unsupervised learning schemes are the instance-level contrastive framework and clustering-based…

Cited by 36SourcePDFScholar
2021

CPF: Learning a Contact Potential Field To Model the Hand-Object Interaction

ICCV 2021poster

Modeling the hand-object (HO) interaction not only requires estimation of the HO pose, but also pays attention to the contact due to their interaction. Significant progress has been made in estimating hand and object separately with deep learning methods, simultaneous HO pose estimation and contact…

Cited by 141PDFcodeScholar
2021

DIRV: Dense Interaction Region Voting for End-to-End Human-Object Interaction Detection

AAAI 2021technical

Recent years, human-object interaction (HOI) detection has achieved impressive advances. However, conventional two-stage methods are usually slow in inference. On the other hand, existing one-stage methods mainly focus on the union regions of interactions, which introduce unnecessary visual informat…

2021

Graspness Discovery in Clutters for Fast and Accurate Grasp Detection

ICCV 2021poster

Efficient and robust grasp pose detection is vital for robotic manipulation. For general 6 DoF grasping, conventional methods treat all points in a scene equally and usually adopt uniform sampling to select grasp candidates. However, we discover that ignoring where to grasp greatly harms the speed a…

Cited by 120PDFcodeScholar
2021

H2O: A Benchmark for Visual Human-Human Object Handover Analysis

ICCV 2021poster

Object handover is a common human collaboration behavior that attracts attention from researchers in Robotics and Cognitive Science. Though visual perception plays an important role in the object handover task, the whole handover process has been specifically explored. In this work, we propose a nov…

Cited by 31PDFScholar
2021

Human Pose Regression With Residual Log-Likelihood Estimation

ICCV 2021poster

Heatmap-based methods dominate in the field of human pose estimation by modelling the output distribution through likelihood heatmaps. In contrast, regression-based methods are more efficient but suffer from inferior performance. In this work, we explore maximum likelihood estimation (MLE) to develo…

Cited by 277PDFcodeScholar
2021

HybrIK: A Hybrid Analytical-Neural Inverse Kinematics Solution for 3D Human Pose and Shape Estimation

CVPR 2021poster

Model-based 3D pose and shape estimation methods reconstruct a full 3D mesh for the human body by estimating several parameters. However, learning the abstract parameters is a highly non-linear process and suffers from image-model misalignment, leading to mediocre model performance. In contrast, 3D…

Cited by 469PDFcodeScholar
2021

RGB Matters: Learning 7-DoF Grasp Poses on Monocular RGBD Images

ICRA 2021poster

General object grasping is an important yet unsolved problem in the field of robotics. Most of the current methods either generate grasp poses with few DoF that fail to cover most of the success grasps, or only take the unstable depth image or point cloud as input which may lead to poor results in s…

Cited by 132SourcecodeScholar
2021

TDAF: Top-Down Attention Framework for Vision Tasks

AAAI 2021technical

Human attention mechanisms often work in a top-down manner, yet it is not well explored in vision research. Here, we propose the Top-Down Attention Framework (TDAF) to capture top-down attentions, which can be easily adopted in most existing models. The designed Recursive Dual-Directional Nested Str…

Cited by 13SourcePDFScholar
2021

Three Steps to Multimodal Trajectory Prediction: Modality Clustering, Classification and Synthesis

ICCV 2021poster

Multimodal prediction results are essential for trajectory prediction task as there is no single correct answer for the future. Previous frameworks can be divided into three categories: regression, generation and classification frameworks. However, these frameworks have weaknesses in different aspec…

Cited by 89PDFScholar
2020

6-PACK: Category-level 6D Pose Tracker with Anchor-Based Keypoints

ICRA 2020poster

We present 6-PACK, a deep learning approach to category-level 6D object pose tracking on RGB-D data. Our method tracks in real time novel object instances of known object categories such as bowls, laptops, and mugs. 6-PACK learns to compactly represent an object by a handful of 3D keypoints, based o…

Cited by 190SourcecodeScholar
2020

Asynchronous Interaction Aggregation for Action Detection

ECCV 2020poster

Understanding interaction is an essential part of video action detection. We propose the Asynchronous Interaction Aggregation network (AIA) that leverages different interactions to boost action detection. There are two key designs in it: one is the Interaction Aggregation structure (IA) adopting a u…

2020

Detailed 2D-3D Joint Representation for Human-Object Interaction

CVPR 2020poster

Human-Object Interaction (HOI) detection lies at the core of action understanding. Besides 2D information such as human/object appearance and locations, 3D pose is also usually utilized in HOI learning since its view-independence. However, rough 3D body joints just carry sparse body information and…

Cited by 175PDFcodeScholar
2020

GraspNet-1Billion: A Large-Scale Benchmark for General Object Grasping

CVPR 2020poster

Object grasping is critical for many applications, which is also a challenging computer vision problem. However, for cluttered scene, current researches suffer from the problems of insufficient training data and the lacking of evaluation benchmarks. In this work, we contribute a large-scale grasp po…

Cited by 650PDFcodeScholar
2020

HMOR: Hierarchical Multi-Person Ordinal Relations for Monocular Multi-Person 3D Pose Estimation

ECCV 2020poster

Remarkable progress has been made in 3D human pose estimation from a monocular RGB camera. However, only a few studies explored 3D multi-person cases. In this paper, we attempt to address the lack of a global perspective of the top-down approaches by introducing a novel form of supervision - Hierarc…

Cited by 75SourcePDFScholar
2020

HOI Analysis: Integrating and Decomposing Human-Object Interaction

NeurIPS 2020poster

Human-Object Interaction (HOI) consists of human, object and implicit interaction/verb. Different from previous methods that directly map pixels to HOI semantics, we propose a novel perspective for HOI learning in an analytical manner. In analogy to Harmonic Analysis, whose goal is to study how to r…

2020

Human Correspondence Consensus for 3D Object Semantic Understanding

ECCV 2020poster

Semantic understanding of 3D objects is crucial in many applications such as object manipulation. However, it is hard to give a universal definition of point-level semantics that everyone would agree on. We observe that people have a consensus on semantic correspondences between two areas from diffe…

2020

KeypointNet: A Large-Scale 3D Keypoint Dataset Aggregated From Numerous Human Annotations

CVPR 2020poster

Detecting 3D objects keypoints is ofgreat interest to the areas of both graphics and computer vision. There have been several 2D and 3D keypoint datasets aiming to address this problem in a data-driven way. These datasets, however, either lack scalability or bring ambiguity to the definition of keyp…

Cited by 89PDFcodeScholar
2020

Transferable Active Grasping and Real Embodied Dataset

ICRA 2020poster

Grasping in cluttered scenes is challenging for robot vision systems, as detection accuracy can be hindered by partial occlusion of objects. We adopt a reinforcement learning (RL) framework and 3D vision architectures to search for feasible viewpoints for grasping by the use of hand-mounted RGB-D ca…

Cited by 26SourcecodeScholar
2020

TubeTK: Adopting Tubes to Track Multi-Object in a One-Step Training Model

CVPR 2020oral

Multi-object tracking is a fundamental vision problem that has been studied for a long time. As deep learning brings excellent performances to object detection algorithms, Tracking by Detection (TBD) has become the mainstream tracking framework. Despite the success of TBD, this two-step method is to…

Cited by 344PDFcodeScholar
2019

Cross-Domain Adaptation for Animal Pose Estimation

ICCV 2019oral

In this paper, we are interested in pose estimation of animals. Animals usually exhibit a wide range of variations on poses and there is no available animal pose dataset for training and testing. To address this problem, we build an animal pose dataset to facilitate training and evaluation. Consider…

Cited by 223PDFcodeScholar
2019

CrowdPose: Efficient Crowded Scenes Pose Estimation and a New Benchmark

CVPR 2019oral

Multi-person pose estimation is fundamental to many computer vision tasks and has made significant progress in recent years. However, few previous methods explored the problem of pose estimation in crowded scenes while it remains challenging and inevitable in many scenarios. Moreover, current benchm…

Cited by 699PDFcodeScholar
2019

DenseFusion: 6D Object Pose Estimation by Iterative Dense Fusion

CVPR 2019poster

A key technical challenge in performing 6D object pose estimation from RGB-D image is to fully leverage the two complementary data sources. Prior works either extract information from the RGB image and depth separately or use costly post-processing steps, limiting their performances in highly clutte…

Cited by 1292PDFScholar
2019

InstaBoost: Boosting Instance Segmentation via Probability Map Guided Copy-Pasting

ICCV 2019poster

Instance segmentation requires a large number of training samples to achieve satisfactory performance and benefits from proper data augmentation. To enlarge the training set and increase the diversity, previous methods have investigated using data annotation from other domain (e.g. bbox, point) in a…

Cited by 253PDFcodeScholar
2019

TendencyRL: Multi-stage Discriminative Hints for Efficient Goal-Oriented Reverse Curriculum Learning

IROS 2019poster

Deep reinforcement learning algorithms have been proven successful in a variety of simulation tasks with dense reward feedback. However, real-world RL applications, e.g. robotic manipulation, remain challenging as most of them are multi-stage and a positive reward can only be received when the final…

Cited by 4SourceScholar
2019

Transferable Interactiveness Knowledge for Human-Object Interaction Detection

CVPR 2019poster

Human-Object Interaction (HOI) Detection is an important problem to understand how humans interact with objects. In this paper, we explore Interactiveness Knowledge which indicates whether human and object interact with each other or not. We found that interactiveness knowledge can be learned across…

Cited by 383PDFcodeScholar
2018

Beyond Holistic Object Recognition: Enriching Image Understanding With Part States

CVPR 2018poster

Important high-level vision tasks require rich semantic descriptions of objects at part level. Based upon previous work on part localization, in this paper, we address the problem of inferring rich semantics imparted by an object part in still images. Specifically, we propose to tokenize the semanti…

Cited by 35SourcePDFScholar
2018

Environment Upgrade Reinforcement Learning for Non-Differentiable Multi-Stage Pipelines

CVPR 2018poster

Recent advances in multi-stage algorithms have shown great promise, but two important problems still remain. First of all, at inference time, information can't feed back from downstream to upstream. Second, at training time, end-to-end training is not possible if the overall pipeline involves non-di…

Cited by 8SourcePDFScholar
2018

LiDAR-Video Driving Dataset: Learning Driving Policies Effectively

CVPR 2018poster

Learning autonomous-driving policies is one of the most challenging but promising tasks for computer vision. Most researchers believe that future research and applications should combine cameras, video recorders and laser scanners to obtain comprehensive semantic understanding of real traffic. Howev…

2018

Pairwise Body-Part Attention for Recognizing Human-Object Interactions

ECCV 2018poster

In human-object interactions (HOI) recognition, conventional methods consider the human body as a whole and pay a uniform attention to the entire body region. They ignore the fact that normally, human interacts with an object by using some parts of the body. In this paper, we argue that different bo…

Cited by 169SourcePDFScholar
2018

Recurrent Residual Module for Fast Inference in Videos

CVPR 2018poster

Deep convolutional neural networks (CNNs) have made impressive progress in many video recognition tasks such as video pose estimation and video object detection. However, running CNN inference on video requires numerous computation and is usually slow. In this work, we propose a framework called Rec…

Cited by 46SourcePDFScholar
2018

SRDA: Generating Instance Segmentation Annotation via Scanning, Reasoning and Domain Adaptation

ECCV 2018poster

Instance segmentation is a problem of significance in computer vision. However, preparing annotated data for this task is extremely time-consuming and costly. By combining the advantages of 3D scanning, reasoning, and GAN-based domain adaptation techniques, we introduce a novel pipeline named SRDA t…

2018

Weakly and Semi Supervised Human Body Part Parsing via Pose-Guided Knowledge Transfer

CVPR 2018poster

Human body part parsing, or human semantic part segmentation, is fundamental to many computer vision tasks. In conventional semantic segmentation methods, the ground truth segmentations are provided, and fully convolutional networks (FCN) are trained in an end-to-end scheme. Although these methods h…

2015

Complexity-Adaptive Distance Metric for Object Proposals Generation

CVPR 2015poster

Distance metric plays a key role in grouping superpixels to produce object proposals for object detection. We observe that existing distance metrics work primarily for low complexity cases. In this paper, we develop a novel distance metric for grouping two superpixels in high-complexity scenarios. C…

Cited by 47SourcePDFScholar
2015

Deep LAC: Deep Localization, Alignment and Classification for Fine-Grained Recognition

CVPR 2015poster

We propose a fine-grained recognition system that incorporates part localization, alignment, and classification in one deep neural network. This is a nontrivial process, as the input to the classification module should be functions that enable back-propagation in constructing the solver. Our major c…

Cited by 439SourcePDFScholar