← Search

Hongtao Wu

23 accepted papers

2026

Hermes: An Evidence-Driven Agentic Framework for Trustworthy and Explainable AI-Generated Video Detection

ICML 2026poster

Recent advances in generative video models have blurred the boundary between real and synthetic content, raising urgent concerns about digital authenticity. Multimodal large language models (MLLMs) are appealing for AI-generated video (AIGV) forensics due to their broad perceptual and reasoning capa…

Cited by 0SourceScholar
2026

RoboOmni: Actions Are Just Another Modality for Your Vision-Language Models

ICML 2026poster

Integrating Vision-Language Models (VLMs) into robotics has facilitated the development of generalizable Vision-Language Action (VLA) policies. However, unified discrete frameworks lag behind decoupled continuous designs due to limitations in action chunking and temporal modeling. To address this, w…

Cited by 0SourceScholar
2025

BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models

NeurIPS 2025poster

Recently, leveraging pre-trained vision-language models (VLMs) for building vision-language-action (VLA) models has emerged as a promising approach to effective robot manipulation learning. However, only few methods incorporate 3D signals into VLMs for action prediction, and they do not fully levera…

Cited by 0SourcecodeScholar
2025

Design and Experimental Verification of a Quasi-Passive Variable Stiffness Ankle Exoskeleton for Human Walking Assistance

RA-L 2025

Exoskeleton robots are an effective method for enhancing human walking ability. This paper introduces a quasi-passive variable stiffness ankle exoskeleton, which absorbs negative work produced during ankle dorsiflexion in the stance phase of the gait cycle and releases energy to assist plantar flexi

Cited by 7SourceScholar
2025

GR-MG: Leveraging Partially-Annotated Data via Multi-Modal Goal-Conditioned Policy

RA-L 2025

The robotics community has consistently aimed to achieve generalizable robot manipulation with flexible natural language instructions. One primary challenge is that obtaining robot trajectories fully annotated with both actions and texts is time-consuming and labor-intensive. However, partially-anno

Cited by 39SourcecodeScholar
2025

Human-assisted Robotic Policy Refinement via Action Preference Optimization

NeurIPS 2025poster

Establishing a reliable and iteratively refined robotic system is essential for deploying real-world applications. While Vision-Language-Action (VLA) models are widely recognized as the foundation model for such robotic deployment, their reliance on offline expert demonstrations critically limi…

Cited by 0SourcecodeScholar
2025

IRASim: A Fine-Grained World Model for Robot Manipulation

ICCV 2025poster

World models allow autonomous agents to plan and explore by predicting the visual outcomes of different actions. However, for robot manipulation, it is challenging to accurately model the fine-grained robot-object interaction within the visual space using existing methods which overlook precise alig…

2025

SnowMaster: Comprehensive Real-world Image Desnowing via MLLM with Multi-Model Feedback Optimization

CVPR 2025poster

Snowfall presents significant challenges for visual data processing, necessitating specialized desnowing algorithms. However, existing models often fail to generalize effectively due to their heavy reliance on synthetic datasets. Furthermore, current real-world snowfall datasets are limited in scale…

Cited by 0SourcePDFScholar
2025

World Model-Based Perception for Visual Legged Locomotion

ICRA 2025

Legged locomotion over various terrains is challenging and requires precise perception of the robot and its surroundings from both proprioception and vision. However, learning directly from high-dimensional visual input is often data-inefficient and intricate. To address this issue, traditional meth

Cited by 23SourcecodeScholar
2024

Genuine Knowledge from Practice: Diffusion Test-Time Adaptation for Video Adverse Weather Removal

CVPR 2024poster

Real-world vision tasks frequently suffer from the appearance of unexpected adverse weather conditions including rain haze snow and raindrops. In the last decade convolutional neural networks and vision transformers have yielded outstanding results in single-weather video removal. However due to the…

2024

Semi-Supervised Video Desnowing Network via Temporal Decoupling Experts and Distribution-Driven Contrastive Regularization

ECCV 2024poster

"Snow degradations present formidable challenges to the advancement of computer vision tasks by the undesirable corruption in outdoor scenarios. While current deep learning-based desnowing approaches achieve success on synthetic benchmark datasets, they struggle to restore out-of-distribution real-w…

2024

Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

ICLR 2024poster

Generative pre-trained models have demonstrated remarkable effectiveness in language and vision domains by learning useful representations. In this paper, we extend the scope of this effectiveness by showing that visual robot manipulation can significantly benefit from large-scale video generative p…

2024

Vision-Language Foundation Models as Effective Robot Imitators

ICLR 2024spotlight

Recent progress in vision language foundation models has shown their ability to understand multimodal data and resolve complicated vision language tasks, including robotics manipulation. We seek a straightforward way of making use of existing vision-language models (VLMs) with simple fine-tuning on…

Cited by 133SourcePDFScholar
2023

MOMA-Force: Visual-Force Imitation for Real-World Mobile Manipulation

IROS 2023poster

In this paper, we present a novel method for mobile manipulators to perform multiple contact-rich manipulation tasks. While learning-based methods have the potential to generate actions in an end-to-end manner, they often suffer from insufficient action accuracy and robustness against noise. On the…

Cited by 12SourcecodeScholar
2023

Prepare the Chair for the Bear! Robot Imagination of Sitting Affordance to Reorient Previously Unseen Chairs

RA-L 2023

In this letter, a paradigm for the classification and manipulation of novel objects is established and demonstrated with the example of chairs. Our approach leverages the robot's understanding of object stability, perceptibility, and affordance to prepare previously unseen and randomly oriented chai

Cited by 4SourceScholar
2023

Snow Removal in Video: A New Dataset and A Novel Method

ICCV 2023poster

Snowfall is a common weather phenomenon that can severely affect computer vision tasks by obscuring objects and scenes. However, existing deep learning-based snow removal methods are designed for single images only. In this paper, we target a more complex task -- video snow removal, which aims to re…

Cited by 22PDFcodeScholar
2022

Kinematic Compatible Design and Analysis of a Back Exoskeleton via a Hyper Redundant Hybrid Mechanism

RA-L 2022

Back exoskeletons can reduce low-back pain and injury for workers engaged in manual handling operations. However, conventional back exoskeletons are incompatible with human trunk kinematics, arousing uncomfortable human-exoskeleton interaction and limiting natural human movement. This article propos

Cited by 4SourceScholar
2022

Put the Bear on the Chair! Intelligent Robot Interaction with Previously Unseen Chairs via Robot Imagination

ICRA 2022poster

In this paper, we study the problem of autonomously seating a teddy bear on a previously unseen chair. To achieve this goal, we present a novel method for robots to imagine the sitting pose of the bear by physically simulating a virtual humanoid agent sitting on the chair. We also develop a robotic…

Cited by 8SourcecodeScholar
2022

Transporters with Visual Foresight for Solving Unseen Rearrangement Tasks

IROS 2022poster

Rearrangement tasks have been identified as a crucial challenge for intelligent robotic manipulation, but few methods allow for precise construction of unseen structures. We propose a visual foresight model for pick-and-place rearrangement manipulation which is able to learn efficiently. In addition…

Cited by 16SourcecodeScholar
2021

Can I Pour Into It? Robot Imagining Open Containability Affordance of Previously Unseen Objects via Physical Simulations

RA-L 2021

Open containers, i.e., containers without covers, are an important and ubiquitous class of objects in human life. In this letter, we propose a novel method for robots to “imagine” the open containability affordance of a previously unseen object via physical simulations. The robot autonomously scans

Cited by 21SourceScholar
2021

LSG-CPD: Coherent Point Drift With Local Surface Geometry for Point Cloud Registration

ICCV 2021poster

Probabilistic point cloud registration methods are becoming more popular because of their robustness. However, unlike point-to-plane variants of iterative closest point (ICP) which incorporate local surface geometric information such as surface normals, most probabilistic methods (e.g., coherent poi…

Cited by 41PDFcodeScholar
2020

"Good Robot!": Efficient Reinforcement Learning for Multi-Step Visual Tasks with Sim to Real Transfer

RA-L 2020

Current Reinforcement Learning (RL) algorithms struggle with long-horizon tasks where time can be wasted exploring dead ends and task progress may be easily reversed. We develop the SPOT framework, which explores within action safety zones, learns about unsafe regions without exploring them, and pri

Cited by 73SourcecodeScholar
2020

Is That a Chair? Imagining Affordances Using Simulations of an Articulated Human Body

ICRA 2020poster

For robots to exhibit a high level of intelligence in the real world, they must be able to assess objects for which they have no prior knowledge. Therefore, it is crucial for robots to perceive object affordances by reasoning about physical interactions with the object. In this paper, we propose a n…

Cited by 16SourceScholar