← Search

Hao Dong

123 accepted papers

2026

A Supervised Multi-task Framework for Joint cryo-ET Restoration Enabled by Generative Physical Simulation

CVPR 2026

Cryo-electron tomography (cryo-ET) enables in-situ visualization of cellular ultrastructure, but reconstructions are severely degraded by extremely low SNR and missing-wedge artifacts due to dose limits and restricted tilt angles. Existing learning-based approaches are further constrained by inaccur

Cited by 0SourceScholar
2026

A3D: Adaptive Affordance Assembly with Dual-Arm Manipulation

AAAI 2026technical

Furniture assembly is a crucial yet challenging task for robots, requiring precise dual-arm coordination where one arm manipulates parts while the other provides collaborative support and stabilization. To accomplish this task more effectively, robots need to actively adapt support strategies throu

Cited by 5SourcePDFScholar
2026

AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Models

CVPR 2026

Vision-Language-Action (VLA) models have significantly advanced robotic agents capable of executing diverse tasks; however, they remain limited in contact-rich manipulation scenarios that require precise physical interactions. To address this limitation, recent studies have attempted to incorporate

Cited by 0SourceScholar
2026

BiPreManip: Learning Affordance-Based Bimanual Preparatory Manipulation through Anticipatory Collaboration

CVPR 2026

Many everyday objects are difficult to directly grasp (e.g., a flat iPad) or manipulate functionally (e.g., opening the cap of a pen lying on a desk). Such tasks require sequential, asymmetric coordination between two arms, where one arm performs preparatory manipulation that enables the other's goa

Cited by 0SourceScholar
2026

CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model

AAAI 2026technical

Existing vision-and-language navigation models often deviate from the correct trajectory when executing instructions. However, these models lack effective error correction capability, hindering their recovery from errors. To address this challenge, we propose Self-correction Flywheel, a novel post-t

Cited by 25SourcePDFScholar
2026

FlatLab: A Unified Methodology Framework and Simulation-Based Benchmark for Robotic Manipulation of Flat Objects

ICML 2026poster

Robotic manipulation of flat objects is challenging due to the ungraspable configurations and strong variations in object geometry and material. Existing methods rely on heuristic pre-manipulation and are often evaluated in closed settings with limited generalization. We propose a unified framework …

Cited by 0SourcecodeScholar
2026

FreeArtGS: Articulated Gaussian Splatting Under Free-moving Scenario

CVPR 2026

The increasing demand for augmented reality and robotics is driving the need for articulated object reconstruction with high scalability. However, existing settings for reconstructing from discrete articulation states or casual monocular videos require non-trivial axis alignment or suffer from insuf

Cited by 0SourcecodeScholar
2026

GarmentPile++: Affordance-Driven Cluttered Garments Retrieval with Vision-Language Reasoning

ICRA 2026poster

Garment manipulation has attracted increasing attention due to its critical role in home-assistant robotics. However, the majority of existing garment manipulation works assume an initial state consisting of only one garment, while piled garments are far more common in real-world settings. To bridge…

2026

GraspALL: Adaptive Structural Compensation from Illumination Variation for Robotic Garment Grasping in Any Low-Light Conditions

CVPR 2026

Achieving accurate garment grasping under dynamically changing illumination is crucial for all-day operation of service robots. However, the reduced illumination in low-light scenes severely degrades garment structural features, leading to a significant drop in grasping robustness. Existing methods

Cited by 0SourcecodeScholar
2026

Imagine2Act: Leveraging Object-Action Motion Consistency from Imagined Goals for Robotic Manipulation

ICRA 2026poster

Relational object rearrangement (ROR) tasks require a robot to manipulate objects with precise semantic and geometric reasoning. Existing approaches either rely on pre-collected demonstrations that struggle to capture complex geometric constraints, or generate goal-state observations to capture sema…

2026

InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training Generalist Policy

CVPR 2026

Recent work explores how real and synthetic data contribute to VLA model generalization. While the \pi-series model has shown the strong effectiveness of large-scale real-robot pre-training, synthetic data has not previously demonstrated comparable capability at scale.This paper provides the first e

Cited by 0SourceScholar
2026

Learning Part-Aware Dense 3D Feature Field For Generalizable Articulated Object Manipulation

ICLR 2026poster

Articulated object manipulation is essential for various real-world robotic tasks, yet generalizing across diverse objects remains a major challenge. A key to generalization lies in understanding functional parts (e.g., door handles and knobs), which indicate where and how to manipulate across diver…

Cited by 0SourceScholar
2026

NaturalVLM: Leveraging Fine-Grained Natural Language for Affordance-Guided Visual Manipulation

ICRA 2026poster

Enabling home-assistant robots to perceive and manipulate a diverse range of 3D objects based on human language instructions is a pivotal challenge. Prior research has predominantly focused on simplistic and task-oriented instructions, i.e., "Slide the top drawer open". However, many real-world task…

2026

NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions

ICRA 2026poster

Instruction-following navigation is a key step toward embodied intelligence. Prior benchmarks mainly focus on semantic understanding but overlook systematically evaluating navigation agents' spatial perception and reasoning capabilities. In this work, we introduce the NavSpace benchmark, which conta…

2026

Real2Edit2Real: Generating Robotic Demonstrations via a 3D Control Interface

CVPR 2026

Recent progress in robot learning has been driven by large-scale datasets and powerful visuomotor policy architectures, yet policy robustness remains limited by the substantial cost of collecting diverse demonstrations, particularly for spatial generalization in manipulation tasks. To reduce repetit

Cited by 0SourceScholar
2026

RealAppiance: Let High-fidelity Appliance Assets Controllable and Workable as Aligned Real Manauls

CVPR 2026

Existing appliance assets suffer from poor rendering, incomplete mechanisms, and misalignment with manuals, leading to simulation-reality gaps that hinder appliance manipulation development. In this work, we introduce the RealAppliance dataset, comprising 100 high-fidelity appliances with complete p

Cited by 0SourceScholar
2026

Sparse Meets Dense: Correspondence Guided Robotic Manipulation with Rigid-Deformable Interactions

ICRA 2026poster

Manipulation involving rigid-deformable interactions, such as hanging clothes or dressing humans, is essential for household robots. Compared to single-object manipulation or interactions between rigid bodies, these tasks are particularly challenging due to the rich multi-point contacts and the comp…

Cited by 0Scholar
2026

SpikeStereoNet: A Brain-Inspired Framework for Stereo Depth Estimation from Spike Streams

ICLR 2026poster

Conventional frame-based cameras often struggle with stereo depth estimation in rapidly changing scenes. In contrast, bio-inspired spike cameras emit asynchronous events at microsecond-level resolution, providing an alternative sensing modality. However, existing methods lack specialized stereo algo…

Cited by 0SourcecodeScholar
2026

Towards Multimodal Domain Generalization with Few Labels

CVPR 2026

Multimodal models ideally should generalize to unseen domains while remaining data-efficient to reduce annotation costs. To this end, we introduce and study a new problem, Semi-Supervised Multimodal Domain Generalization (SSMDG), which aims to learn robust multimodal models from multi-source data wi

Cited by 0SourcecodeScholar
2026

Tracking through Severe Occlusion via Event-Derived Transient Cues

CVPR 2026

Tracking targets with high-speed and nonlinear motion under occlusion remains challenging due to spatial appearance deprivation and temporal trajectory fragmentation caused by missing visual cues. Existing methods typically either dynamically update templates to maintain appearance similarity or emp

Cited by 0SourceScholar
2026

UniDoorManip: Learning Universal Door Manipulation Policy Over Large-Scale and Diverse Door Manipulation Environments

ICRA 2026poster

Learning a universal manipulation policy encompassing doors with diverse categories, geometries and mechanisms, is crucial for future embodied agents to effectively work in complex and broad real-world scenarios. Due to the limited datasets and unrealistic simulation environments, previous studies f…

2026

User-Centric Object Navigation: A Benchmark with Integrated User Habits for Personalized Embodied Object Search

ICRA 2026poster

In the evolving field of robotics, the challenge of Object Navigation (ON) in household environments has attracted significant interest. Existing ON benchmarks typically place objects in locations guided by general scene priors, without accounting for the specific placement habits of individual user…

2026

Why Tree-Style Branching Matters for Thought Advantage Estimation in GRPO

ICML 2026poster

Group Relative Policy Optimization (GRPO) trains Chain-of-Thought reasoning with verifiable rewards, but estimating thought-level advantages without value functions often suffers from high variance. Although tree-style branching is used in practice to reduce the variance, it lacks a theoretical expl…

Cited by 0SourceScholar
2025

3DS-VLA: A 3D Spatial-Aware Vision Language Action Model for Robust Multi-Task Manipulation

CoRL 2025poster

Recently, 2D vision-language-action (VLA) models have made significant strides in multi-task manipulation. However, these models struggle to reason about 3D spatial relationships from 2D image inputs. Although an increasing number of 3D approaches explicitly integrate 3D information, they encounter…

Cited by 0SourceScholar
2025

3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment

ICRA 2025

The 3D weakly-supervised visual grounding task aims to localize oriented 3D boxes in point clouds based on natural language descriptions without requiring annotations to guide model learning. This setting presents two primary challenges: category-level ambiguity and instance-level complexity. Catego

Cited by 2SourceScholar
2025

AdaManip: Adaptive Articulated Object Manipulation Environments and Policy Learning

ICLR 2025poster

Articulated object manipulation is a critical capability for robots to perform various tasks in real-world scenarios. Composed of multiple parts connected by joints, articulated objects are endowed with diverse functional mechanisms through complex relative motions. For example, a safe consists of a…

Cited by 4SourcePDFScholar
2025

Adaptive Articulated Object Manipulation On The Fly with Foundation Model Reasoning and Part Grounding

ICCV 2025poster

Articulated objects pose diverse manipulation challenges for robots. Since their internal structures are not directly observable, robots must adaptively explore and refine actions to generate successful manipulation trajectories. While existing works have attempted cross-category generalization in a…

Cited by 0SourcePDFScholar
2025

Adaptive Visuo-Tactile Fusion with Predictive Force Attention for Dexterous Manipulation

IROS 2025

Effectively utilizing multi-sensory data is important for robots to generalize across diverse tasks. However, the heterogeneous nature of these modalities makes fusion challenging. Existing methods propose strategies to obtain comprehensively fused features but often ignore the fact that each modali

Cited by 12SourcecodeScholar
2025

An Automatic Sound and Complete Abstraction Method for Generalized Planning with Baggable Types

AAAI 2025technical

Generalized planning is concerned with how to find a single plan to solve multiple similar planning instances. Abstractions are widely used for solving generalized planning, and QNP (qualitative numeric planning) is a popular abstract model. Recently, Cui et al. showed that a plan solves a sound and…

2025

BiAssemble: Learning Collaborative Affordance for Bimanual Geometric Assembly

ICML 2025poster

Shape assembly, the process of combining parts into a complete whole, is a crucial skill for robots with broad real-world applications. Among the various assembly tasks, geometric assembly—where broken parts are reassembled into their original form (e.g., reconstructing a shattered bowl)—is particul…

Cited by 0SourcePDFScholar
2025

CADGrasp: Learning Contact and Collision Aware General Dexterous Grasping in Cluttered Scenes

NeurIPS 2025poster

Dexterous grasping in cluttered environments presents substantial challenges due to the high degrees of freedom of dexterous hands, occlusion, and potential collisions arising from diverse object geometries and complex layouts. To address these challenges, we propose CADGrasp, a two-stage algorithm…

Cited by 0SourceScholar
2025

Canonical Representation and Force-Based Pretraining of 3D Tactile for Dexterous Visuo-Tactile Policy Learning

ICRA 2025

Tactile sensing plays a vital role in enabling robots to perform fine-grained, contact-rich tasks. However, the high dimensionality of tactile data, due to the large coverage on dexterous hands, poses significant challenges for effective tactile feature learning, especially for 3D tactile data, as t

Cited by 15SourcecodeScholar
2025

CheckManual: A New Challenge and Benchmark for Manual-based Appliance Manipulation

CVPR 2025highlight

Correct use of electrical appliances has significantly improved human life quality. Unlike simple tools that can be manipulated with common sense, different parts of electrical appliances have specific functions defined by manufacturers. If we want the robot to heat bread by microwave, we should ena…

Cited by 0SourcePDFScholar
2025

ClutterDexGrasp: A Sim-to-Real System for General Dexterous Grasping in Cluttered Scenes

CoRL 2025oral

Dexterous grasping in cluttered scenes presents significant challenges due to diverse object geometries, occlusions, and potential collisions. Existing methods primarily focus on single-object grasping or grasp-pose prediction without interaction, which are insufficient for complex, cluttered scenes…

Cited by 0SourceScholar
2025

CordViP: Correspondence-based Visuomotor Policy for Dexterous Manipulation in Real-World

RSS 2025poster

Achieving human-level dexterity in robots is a key objective in the field of robotic manipulation. Recent advancements in 3D-based imitation learning have shown promising results, providing an effective pathway to achieve this goal. However, obtaining high-quality 3D representations presents two key…

Cited by 2PDFScholar
2025

DPU: Dynamic Prototype Updating for Multimodal Out-of-Distribution Detection

CVPR 2025highlight

Out-of-distribution (OOD) detection is crucial for ensuring the robustness of machine learning models by identifying samples that deviate from the training distribution. While traditional OOD detection has predominantly focused on single-modality inputs, such as images, recent advancements in multim…

2025

DexFlyWheel: A Scalable and Self-improving Data Generation Framework for Dexterous Manipulation

NeurIPS 2025spotlight

Dexterous manipulation is critical for advancing robot capabilities in real-world applications, yet diverse and high-quality datasets remain scarce. Existing data collection methods either rely on human teleoperation or require significant human engineering, or generate data with limited diversity,…

Cited by 0SourceScholar
2025

DexGarmentLab: Dexterous Garment Manipulation Environment with Generalizable Policy

NeurIPS 2025spotlight

Garment manipulation is a critical challenge due to the diversity in garment categories, geometries, and deformations. Despite this, humans can effortlessly handle garments, thanks to the dexterity of our hands. However, existing research in the field has struggled to replicate this level of dexteri…

Cited by 0SourcecodeScholar
2025

Disentangled Multi-span Evolutionary Network against Temporal Knowledge Graph Reasoning

ACL 2025finding

Temporal Knowledge Graphs (TKGs) incorporate the temporal feature to express the transience of knowledge by describing when facts occur. TKG extrapolation aims to infer possible future facts based on known history, which has garnered significant attention in recent years. Some existing methods treat…

2025

ET-SEED: EFFICIENT TRAJECTORY-LEVEL SE(3) EQUIVARIANT DIFFUSION POLICY

ICLR 2025poster

Imitation learning, e.g., diffusion policy, has been proven effective in various robotic manipulation tasks. However, extensive demonstrations are required for policy robustness and generalization. To reduce the demonstration reliance, we leverage spatial symmetry and propose ET-SEED, an efficient t…

2025

Extremely Simple Multimodal Outlier Synthesis for Out-of-Distribution Detection and Segmentation

NeurIPS 2025poster

Out-of-distribution (OOD) detection and segmentation are crucial for deploying machine learning models in safety-critical applications such as autonomous driving and robot-assisted surgery. While prior research has primarily focused on unimodal image data, real-world applications are inherently mult…

Cited by 0SourcecodeScholar
2025

Foundation Feature-Driven Online End-Effector Pose Estimation: A Marker-Free and Learning-Free Approach

ICRA 2025

Accurate transformation estimation between camera space and robot space is essential. Traditional methods using markers for hand-eye calibration require offline image collection, limiting their suitability for online self-calibration. Recent learning-based robot pose estimation methods, while advanc

Cited by 2SourcecodeScholar
2025

GCAL: Adapting Graph Models to Evolving Domain Shifts

ICML 2025poster

This paper addresses the challenge of graph domain adaptation on evolving, multiple out-of-distribution (OOD) graphs. Conventional graph domain adaptation methods are confined to single-step adaptation, making them ineffective in handling continuous domain shifts and prone to catastrophic forgetting…

2025

GFPack++: Attention-Driven Gradient Fields for Optimizing 2D Irregular Packing

ICCV 2025poster

2D irregular packing is a classic combinatorial optimization problem with various applications, such as material utilization and texture atlas generation. Due to its NP-hard nature, conventional numerical approaches typically encounter slow convergence and high computational costs. Previous research…

2025

GarmentPile: Point-Level Visual Affordance Guided Retrieval and Adaptation for Cluttered Garments Manipulation

CVPR 2025poster

Cluttered garments manipulation poses significant challenges in robotics due to the complex, deformable nature of garments and intricate garment relations. Unlike single-garment manipulation, cluttered scenarios require managing complex garment entanglements and interactions, while maintaining garme…

2025

ManipGPT: Is Affordance Segmentation by Large Vision Models Enough for Articulated Object Manipulation?

IROS 2025

Visual actionable affordance has emerged as a transformative approach in robotics, focusing on perceiving interaction areas prior to manipulation. Traditional methods rely on pixel sampling to identify successful interaction samples or processing pointclouds for affordance mapping. However, these ap

Cited by 0SourcecodeScholar
2025

Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation

CVPR 2025poster

In robotic manipulation, task goals can be conveyed through various modalities, such as language, goal images, and goal videos. However, natural language can be ambiguous, while images or videos may offer overly detailed specifications. To address these challenges, we propose a novel approach using…

Cited by 0SourcePDFScholar
2025

OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints

CVPR 2025highlight

The development of general robotic systems capable of manipulating in unstructured environments is a significant challenge. While Vision-Language Models(VLM) excel in high-level commonsense reasoning, they lack the fine-grained 3D spatial understanding required for precise manipulation tasks. Fine-…

Cited by 8SourcePDFScholar
2025

PartRM: Modeling Part-Level Dynamics with Large Cross-State Reconstruction Model

CVPR 2025poster

As interest grows in world models that predict future states from current observations and actions, accurately modeling part-level dynamics has become increasingly relevant for various applications. Existing approaches, such as Puppet-Master, rely on fine-tuning large-scale pre-trained video diffusi…

Cited by 0SourcePDFScholar
2025

Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation

ICLR 2025oral

Current efforts to learn scalable policies in robotic manipulation primarily fall into two categories: one focuses on "action," which involves behavior cloning from extensive collections of robotic data, while the other emphasizes "vision," enhancing model generalization by pre-training representati…

2025

RoboVerse: A Unified Platform, Benchmark and Dataset for Scalable and Generalizable Robot Learning

RSS 2025poster

Data scaling and standardized evaluation benchmarks have driven remarkable advances in natural language processing and computer vision. However, in robotics, scaling up data and establishing evaluation protocols pose significant challenges. Directly collecting real-world data is inefficient and reso…

Cited by 0PDFScholar
2025

RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot

IROS 2025

Recent advancements in imitation learning have shown promising results in robotic manipulation, driven by the availability of high-quality training data. To improve data collection efficiency, some approaches focus on developing specialized teleoperation devices for robot control, while others direc

Cited by 2SourcecodeScholar
2025

SR3D: Unleashing Single-view 3D Reconstruction for Transparent and Specular Object Grasping

IROS 2025

Recent advancements in 3D robotic manipulation have improved grasping of everyday objects, but transparent and specular materials remain challenging due to depth sensing limitations. While several 3D reconstruction and depth completion approaches address these challenges, they suffer from setup comp

Cited by 0SourceScholar
2025

Shared-AE: Automatic Identification of Shared Subspaces in High-dimensional Neural and Behavioral Activity

ICLR 2025poster

Understanding the relationship between behavior and neural activity is crucial for understanding brain function. An effective method is to learn embeddings for interconnected modalities. For simple behavioral tasks, neural features can be learned based on labels. However, complex behaviors, such as…

2025

SimLauncher: Launching Sample-Efficient Real-World Robotic Reinforcement Learning via Simulation Pre-Training

IROS 2025

Autonomous learning of dexterous, long-horizon robotic skills has been a longstanding pursuit of embodied AI. Recent advances in robotic reinforcement learning (RL) have demonstrated remarkable performance and robustness in real-world visuomotor control tasks. However, applying RL in the real world

Cited by 3SourceScholar
2025

SpatialBot: Precise Spatial Understanding with Vision Language Models

ICRA 2025

Vision Language Models (VLMs) have achieved impressive performance in 2D image understanding; however, they still struggle with spatial understanding, which is fundamental to embodied AI. In this paper, we propose SpatialBot, a model designed to enhance spatial understanding by utilizing both RGB an

Cited by 167SourcecodeScholar
2025

Towards Robust Multimodal Open-set Test-time Adaptation via Adaptive Entropy-aware Optimization

ICLR 2025poster

Test-time adaptation (TTA) has demonstrated significant potential in addressing distribution shifts between training and testing data. Open-set test-time adaptation (OSTTA) aims to adapt a source pre-trained model online to an unlabeled target domain that contains unknown classes. This task becomes…

2025

TransDiff: Diffusion-Based Method for Manipulating Transparent Objects Using a Single RGB-D Image

ICRA 2025

Manipulating transparent objects presents significant challenges due to the complexities introduced by their reflection and refraction properties, which considerably hinder the accurate estimation of their 3D shapes. To address these challenges, we propose a single-view RGB-D-based depth completion

Cited by 3SourcecodeScholar
2025

UniTac2Pose: A Unified Approach Learned in Simulation for Category-level Visuotactile In-hand Pose Estimation

CoRL 2025poster

Accurate estimation of the in-hand pose of an object based on its CAD model is crucial in both industrial applications and everyday tasks—ranging from positioning workpieces and assembling components to seamlessly inserting devices like USB connectors. While existing methods often rely on regression…

Cited by 0SourceScholar
2024

A3VLM: Actionable Articulation-Aware Vision Language Model

CoRL 2024poster

Vision Language Models (VLMs) for robotics have received significant attention in recent years. As a VLM can understand robot observations and perform complex visual reasoning, it is regarded as a potential universal solution for general robotics challenges such as manipulation and navigation. Howev…

Cited by 12SourcecodeScholar
2024

Articulated Object Manipulation with Coarse-to-fine Affordance for Mitigating the Effect of Point Cloud Noise

ICRA 2024poster

3D articulated objects are inherently challenging for manipulation due to the varied geometries and intricate functionalities associated with articulated objects. Point-level affordance, which predicts the per-point actionable score and thus proposes the best point to interact with, has demonstrated…

Cited by 16SourceScholar
2024

Autonomous Interactive Correction MLLM for Robust Robotic Manipulation

CoRL 2024poster

The ability to reflect on and correct failures is crucial for robotic systems to interact stably with real-life objects. Observing the generalization and reasoning capabilities of Multimodal Large Language Models (MLLMs), previous approaches have aimed to utilize these models to enhance robotic syst…

Cited by 4SourceScholar
2024

Bridging Zero-shot Object Navigation and Foundation Models through Pixel-Guided Navigation Skill

ICRA 2024poster

Zero-shot object navigation is a challenging task for home-assistance robots. This task emphasizes visual grounding, commonsense inference and locomotion abilities, where the first two are inherent in foundation models. But for the locomotion part, most works still depend on map-based planning appro…

Cited by 39SourcecodeScholar
2024

Broadcasting Support Relations Recursively from Local Dynamics for Object Retrieval in Clutters

RSS 2024poster

In our daily life, cluttered objects are everywhere, from scattered stationery and books cluttering the table to bowls and plates filling the kitchen sink. Retrieving a target object from clutters is an essential while challenging skill for robots, for the difficulty of safely manipulating an object…

Cited by 5SourcePDFScholar
2024

Discuss Before Moving: Visual Language Navigation via Multi-expert Discussions

ICRA 2024poster

Visual language navigation (VLN) is an embodied task demanding a wide range of skills encompassing understanding, perception, and planning. For such a multifaceted challenge, previous VLN methods totally rely on one model’s own thinking to make predictions within one round. However, existing models,…

Cited by 51SourceScholar
2024

GarmentLab: A Unified Simulation and Benchmark for Garment Manipulation

NeurIPS 2024poster

Manipulating garments and fabrics has long been a critical endeavor in the development of home-assistant robots. However, due to complex dynamics and topological structures, garment manipulations pose significant challenges. Recent successes in reinforcement learning and vision-based methods offer p…

2024

InstructNav: Zero-shot System for Generic Instruction Navigation in Unexplored Environment

CoRL 2024poster

Enabling robots to navigate following diverse language instructions in unexplored environments is an attractive goal for human-robot interaction. However, this goal is challenging because different navigation tasks require different strategies. The scarcity of instruction navigation data hinders tra…

Cited by 34SourceScholar
2024

JSTR: Joint Spatio-Temporal Reasoning for Event-based Moving Object Detection

ICRA 2024poster

Event-based moving object detection is a challenging task, where static background and moving object are mixed together. Typically, existing methods mainly align the background events to the same spatial coordinate system via motion compensation to distinguish the moving object. However, they neglec…

Cited by 4SourceScholar
2024

LVDiffusor: Distilling Functional Rearrangement Priors From Large Models Into Diffusor

RA-L 2024

Object rearrangement, a fundamental challenge in robotics, demands versatile strategies to handle diverse objects, configurations, and functional needs. To achieve this, the AI robot needs to learn functional rearrangement priors to specify precise goals that meet the functional requirements. Previo

Cited by 12SourceScholar
2024

Learning Manipulation by Predicting Interaction

RSS 2024poster

Representation learning approaches for robotic manipulation have boomed in recent years. Due to the scarcity of in-domain robot data, prevailing methodologies tend to leverage large-scale human video datasets to extract generalizable features for visuomotor policy learning. Despite the progress achi…

2024

MO-DDN: A Coarse-to-Fine Attribute-based Exploration Agent for Multi-Object Demand-driven Navigation

NeurIPS 2024poster

The process of satisfying daily demands is a fundamental aspect of humans' daily lives. With the advancement of embodied AI, robots are increasingly capable of satisfying human demands. Demand-driven navigation (DDN) is a task in which an agent must locate an object to satisfy a specified demand ins…

Cited by 0SourcePDFScholar
2024

Make Graph Neural Networks Great Again: A Generic Integration Paradigm of Topology-Free Patterns for Traffic Speed Prediction

IJCAI 2024poster

Urban traffic speed prediction aims to estimate the future traffic speed for improving urban transportation services. Enormous efforts have been made to exploit Graph Neural Networks (GNNs) for modeling spatial correlations and temporal dependencies of traffic speed evolving patterns, regularized by…

2024

ManipLLM: Embodied Multimodal Large Language Model for Object-Centric Robotic Manipulation

CVPR 2024poster

Robot manipulation relies on accurately predicting contact points and end-effector directions to ensure successful operation. However learning-based robot manipulation trained on a limited category within a simulator often struggles to achieve generalizability especially when confronted with extensi…

Cited by 54SourcePDFScholar
2024

ManipVQA: Injecting Robotic Affordance and Physically Grounded Information into Multi-Modal Large Language Models

IROS 2024poster

While the integration of Multi-modal Large Language Models (MLLMs) with robotic systems has significantly improved robots’ ability to understand and execute natural language instructions, their performance in manipulation tasks remains limited due to a lack of robotics-specific knowledge. Convention…

Cited by 27SourcecodeScholar
2024

MultiOOD: Scaling Out-of-Distribution Detection for Multiple Modalities

NeurIPS 2024spotlight

Detecting out-of-distribution (OOD) samples is important for deploying machine learning models in safety-critical applications such as autonomous driving and robot-assisted surgery. Existing research has mainly focused on unimodal scenarios on image data. However, real-world applications are inheren…

2024

NaturalVLM: Leveraging Fine-Grained Natural Language for Affordance-Guided Visual Manipulation

RA-L 2024

Enabling home-assistant robots to perceive and manipulate a diverse range of 3D objects based on human language instructions is a pivotal challenge. Prior research has predominantly focused on simplistic and task-oriented instructions, i.e., “Slide the top drawer open”. However, many real-world task

Cited by 18SourceScholar
2024

No Time to Train: Empowering Non-Parametric Networks for Few-shot 3D Scene Segmentation

CVPR 2024highlight

To reduce the reliance on large-scale datasets recent works in 3D segmentation resort to few-shot learning. Current 3D few-shot segmentation methods first pre-train models on 'seen' classes and then evaluate their generalization performance on 'unseen' classes. However the prior pre-training stage n…

2024

Personalize Segment Anything Model with One Shot

ICLR 2024poster

Driven by large-data pre-training, Segment Anything Model (SAM) has been demonstrated as a powerful promptable framework, revolutionizing the segmentation field. Despite the generality, customizing SAM for specific visual concepts without man-powered prompting is under-explored, e.g., automatically…

2024

PreAfford: Universal Affordance-Based Pre-Grasping for Diverse Objects and Environments

IROS 2024poster

Robotic manipulation with two-finger grippers is challenged by objects lacking distinct graspable features. Traditional pre-grasping methods, which typically involve repositioning objects or utilizing external aids like table edges, are limited in their adaptability across different object categorie…

Cited by 4SourceScholar
2024

RGBGrasp: Image-Based Object Grasping by Capturing Multiple Views During Robot arm Movement With Neural Radiance Fields

RA-L 2024

Robotic research encounters a significant hurdle when it comes to the intricate task of grasping objects that come in various shapes, materials, and textures. Unlike many prior investigations that heavily leaned on specialized point-cloud cameras or abundant RGB visual data to gather 3D insights for

Cited by 25SourceScholar
2024

RGBManip: Monocular Image-based Robotic Manipulation through Active Object Pose Estimation

ICRA 2024poster

Robotic manipulation requires accurate perception of the environment, which poses a significant challenge due to its inherent complexity and constantly changing nature. In this context, RGB image and point-cloud observations are two commonly used modalities in visual-based robotic manipulation, but…

Cited by 17SourcecodeScholar
2024

Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation

AAAI 2024technical

Recently, video object segmentation (VOS) referred by multi-modal signals, e.g., language and audio, has evoked increasing attention in both industry and academia. It is challenging for exploring the semantic alignment within modalities and the visual correspondence across frames. However, existing…

2024

RoboKeyGen: Robot Pose and Joint Angles Estimation via Diffusion-based 3D Keypoint Generation

ICRA 2024poster

Estimating robot pose and joint angles is significant in advanced robotics, enabling applications like robot collaboration and online hand-eye calibration. However, the introduction of unknown joint angles makes prediction more complex than simple robot pose estimation, due to its higher dimensional…

Cited by 7SourcecodeScholar
2024

SCANet: Correcting LEGO Assembly Errors with Self-Correct Assembly Network

IROS 2024poster

Autonomous assembly in robotics and 3D vision presents significant challenges, particularly in ensuring assembly correctness. Presently, predominant methods such as MEPNet focus on assembling components based on manually provided images. However, these approaches often fall short in achieving satisf…

Cited by 3SourcecodeScholar
2024

Scalable Geometric Fracture Assembly via Co-creation Space among Assemblers

AAAI 2024technical

Geometric fracture assembly presents a challenging practical task in archaeology and 3D computer vision. Previous methods have focused solely on assembling fragments based on semantic information, which has limited the quantity of objects that can be effectively assembled. Therefore, there is a need…

2024

SparseDFF: Sparse-View Feature Distillation for One-Shot Dexterous Manipulation

ICLR 2024poster

Humans demonstrate remarkable skill in transferring manipulation abilities across objects of varying shapes, poses, and appearances, a capability rooted in their understanding of semantic correspondences between different instances. To equip robots with a similar high-level comprehension, we present…

Cited by 16SourcePDFScholar
2024

SuperFusion: Multilevel LiDAR-Camera Fusion for Long-Range HD Map Generation

ICRA 2024poster

High-definition (HD) semantic map generation of the environment is an essential component of autonomous driving. Existing methods have achieved good performance in this task by fusing different sensor modalities, such as LiDAR and camera. However, current works are based on raw data or network featu…

Cited by 54SourcecodeScholar
2024

UniGarmentManip: A Unified Framework for Category-Level Garment Manipulation via Dense Visual Correspondence

CVPR 2024poster

Garment manipulation (e.g. unfolding folding and hanging clothes) is essential for future robots to accomplish home-assistant tasks while highly challenging due to the diversity of garment configurations geometries and deformations. Although able to manipulate similar shaped garments in a certain ta…

2023

Adaptive Path-Memory Network for Temporal Knowledge Graph Reasoning

IJCAI 2023poster

Temporal knowledge graph (TKG) reasoning aims to predict the future missing facts based on historical information and has gained increasing research interest recently. Lots of works have been made to model the historical structural and temporal characteristics for the reasoning task. Most existing w…

2023

DualAfford: Learning Collaborative Visual Affordance for Dual-gripper Manipulation

ICLR 2023poster

It is essential yet challenging for future home-assistant robots to understand and manipulate diverse 3D objects in daily human environments. Towards building scalable systems that can perform diverse manipulation tasks over various 3D shapes, recent works have advocated and demonstrated promising r…

Cited by 16SourcePDFScholar
2023

Find What You Want: Learning Demand-conditioned Object Attribute Space for Demand-driven Navigation

NeurIPS 2023poster

The task of Visual Object Navigation (VON) involves an agent's ability to locate a particular object within a given scene. To successfully accomplish the VON task, two essential conditions must be fulfiled: 1) the user knows the name of the desired object; and 2) the user-specified object actually…

2023

GFPose: Learning 3D Human Pose Prior With Gradient Fields

CVPR 2023poster

Learning 3D human pose prior is essential to human-centered AI. Here, we present GFPose, a versatile framework to model plausible 3D human poses for various applications. At the core of GFPose is a time-dependent score network, which estimates the gradient on each body joint and progressively denois…

2023

Learning Environment-Aware Affordance for 3D Articulated Object Manipulation under Occlusions

NeurIPS 2023poster

Perceiving and manipulating 3D articulated objects in diverse environments is essential for home-assistant robots. Recent studies have shown that point-level affordance provides actionable priors for downstream manipulation tasks. However, existing works primarily focus on single-object scenarios wi…

Cited by 25SourcePDFScholar
2023

Learning Score-based Grasping Primitive for Human-assisting Dexterous Grasping

NeurIPS 2023poster

The use of anthropomorphic robotic hands for assisting individuals in situations where human hands may be unavailable or unsuitable has gained significant importance. In this paper, we propose a novel task called human-assisting dexterous grasping that aims to train a policy for controlling a roboti…

Cited by 16SourcePDFScholar
2023

Learning Semantic-Agnostic and Spatial-Aware Representation for Generalizable Visual-Audio Navigation

RA-L 2023

Visual-audio navigation (VAN) is attracting more and more attention from the robotic community due to its broad applications, e.g., household robots and rescue robots. In this task, an embodied agent must search for and navigate to the sound source with egocentric visual and audio observations. Howe

Cited by 12SourcecodeScholar
2023

Learning-Based Dimensionality Reduction for Computing Compact and Effective Local Feature Descriptors

ICRA 2023poster

A distinctive representation of image patches in form of features is a key component of many computer vision and robotics tasks, such as image matching, image retrieval, and visual localization. State-of-the-art descriptors, from hand-crafted descriptors such as SIFT to learned ones such as HardNet,…

Cited by 11SourcecodeScholar
2023

Leveraging SE(3) Equivariance for Learning 3D Geometric Shape Assembly

ICCV 2023poster

Shape assembly aims to reassemble parts (or fragments) into a complete object, which is a common task in our daily life. Different from the semantic part assembly (e.g., assembling a chair's semantic parts like legs into a whole chair), geometric part assembly (e.g., assembling bowl fragments into a…

Cited by 21PDFcodeScholar
2023

PartManip: Learning Cross-Category Generalizable Part Manipulation Policy From Point Cloud Observations

CVPR 2023poster

Learning a generalizable object manipulation policy is vital for an embodied agent to work in complex real-world scenes. Parts, as the shared components in different object categories, have the potential to increase the generalization ability of the manipulation policy and achieve cross-category obj…

Cited by 40SourcePDFScholar
2023

RLAfford: End-to-End Affordance Learning for Robotic Manipulation

ICRA 2023poster

Learning to manipulate 3D objects in an interactive environment has been a challenging problem in Reinforcement Learning (RL). In particular, it is hard to train a policy that can generalize over objects with different semantic categories, diverse shape geometry and versatile functionality. In this…

Cited by 73SourceScholar
2023

Resilient Binary Neural Network

AAAI 2023technical

Binary neural networks (BNNs) have received ever-increasing popularity for their great capability of reducing storage burden as well as quickening inference time. However, there is a severe performance drop compared with {real-valued} networks, due to its intrinsic frequent weight oscillation during…

2023

Robot Structure Prior Guided Temporal Attention for Camera-to-Robot Pose Estimation From Image Sequence

CVPR 2023poster

In this work, we tackle the problem of online camera-to-robot pose estimation from single-view successive frames of an image sequence, a crucial task for robots to interact with the world. The primary obstacles of this task are the robot's self-occlusions and the ambiguity of single-view images. Thi…

2023

Semi-supervised Domain Adaptation in Graph Transfer Learning

IJCAI 2023poster

As a specific case of graph transfer learning, unsupervised domain adaptation on graphs aims for knowledge transfer from label-rich source graphs to unlabeled target graphs. However, graphs with topology and attributes usually have considerable cross-domain disparity and there are numerous real-worl…

Cited by 30SourcePDFScholar
2023

SimMMDG: A Simple and Effective Framework for Multi-modal Domain Generalization

NeurIPS 2023poster

In real-world scenarios, achieving domain generalization (DG) presents significant challenges as models are required to generalize to unknown target distributions. Generalizing to unseen multi-modal distributions poses even greater difficulties due to the distinct properties exhibited by different m…

2023

Where2Explore: Few-shot Affordance Learning for Unseen Novel Categories of Articulated Objects

NeurIPS 2023poster

Articulated object manipulation is a fundamental yet challenging task in robotics. Due to significant geometric and semantic variations across object categories, previous manipulation models struggle to generalize to novel categories. Few-shot learning is a promising solution for alleviating this is…

Cited by 39SourcePDFScholar
2022

AdaAfford: Learning to Adapt Manipulation Affordance for 3D Articulated Objects via Few-Shot Interactions

ECCV 2022poster

"Perceiving and interacting with 3D articulated objects, such as cabinets, doors, and faucets, pose particular challenges for future home-assistant robots performing daily tasks in human environments. Besides parsing the articulated parts and joint parameters, researchers recently advocate learning…

Cited by 68SourcePDFScholar
2022

Domain Randomization-Enhanced Depth Simulation and Restoration for Perceiving and Grasping Specular and Transparent Objects

ECCV 2022poster

"Commercial depth sensors usually generate noisy and missing depths, especially on specular and transparent objects, which poses critical issues to downstream depth or point cloud-based tasks. To mitigate this problem, we propose a powerful RGBD fusion network, SwinDRNet, for depth restoration. We f…

2022

Scalable Model-based Policy Optimization for Decentralized Networked Systems

IROS 2022poster

Reinforcement learning algorithms require a large amount of samples; this often limits their real-world applications on even simple tasks. Such a challenge is more outstanding in multi-agent tasks, as each step of operation is more costly, requiring communications or shifting or resources. This work…

Cited by 10SourcecodeScholar
2022

TarGF: Learning Target Gradient Field to Rearrange Objects without Explicit Goal Specification

NeurIPS 2022accept

Object Rearrangement is to move objects from an initial state to a goal state. Here, we focus on a more practical setting in object rearrangement, i.e., rearranging objects from shuffled layouts to a normative target distribution without explicit goal specification. However, it remains challenging f…

Cited by 36SourcePDFScholar
2022

Towards Human-Level Bimanual Dexterous Manipulation with Reinforcement Learning

NeurIPS 2022accept

Achieving human-level dexterity is an important open problem in robotics. However, tasks of dexterous hand manipulation even at the baby level are challenging to solve through reinforcement learning (RL). The difficulty lies in the high degrees of freedom and the required cooperation among heterogen…

2022

VAT-Mart: Learning Visual Action Trajectory Proposals for Manipulating 3D ARTiculated Objects

ICLR 2022poster

Perceiving and manipulating 3D articulated objects (e.g., cabinets, doors) in human environments is an important yet challenging task for future home-assistant robots. The space of 3D articulated objects is exceptionally rich in their myriad semantic categories, diverse shape geometry, and complicat…

Cited by 104SourcePDFScholar
2021

Contrastive Multimodal Fusion With TupleInfoNCE

ICCV 2021poster

This paper proposes a method for representation learning of multimodal data using contrastive losses. A traditional approach is to contrast different modalities to learn the information shared between them. However, that approach could fail to learn the complementary synergies between modalities tha…

Cited by 83PDFcodeScholar
2021

DMotion: Robotic Visuomotor Control with Unsupervised Forward Model Learned from Videos

IROS 2021poster

Learning an accurate model of the environment is essential for model-based control tasks. Existing methods in robotic visuomotor control usually learn from data with heavily labelled actions, object entities or locations, which can be demanding in many cases. To cope with this limitation, we propose…

Cited by 2SourcecodeScholar
2021

Fast Online Planning for Bipedal Locomotion via Centroidal Model Predictive Gait Synthesis

RA-L 2021

The planning of whole-body motion and step time for bipedal locomotion is constructed as a model predictive control (MPC) problem, in which a sequence of optimization problems needs to be solved online. While directly solving these problems is extremely time-consuming, we propose a predictive gait s

Cited by 13SourceScholar
2020

Generative 3D Part Assembly via Dynamic Graph Learning

NeurIPS 2020poster

Autonomous part assembly is a challenging yet crucial task in 3D computer vision and robotics. Analogous to buying an IKEA furniture, given a set of 3D parts that can assemble a single shape, an intelligent agent needs to perceive the 3D part geometry, reason to propose pose estimations for the inpu…

Cited by 100SourcePDFScholar