← Search

Tingfan Wu

12 accepted papers

2026

Cross-Embodiment Robot Foundation World Models with Latent Actions

ICML 2026poster

The diversity of robot embodiments and action spaces makes it challenging to build robot world models that generalize across different embodiments. We introduce a Latent Action Conditioned Robot World Model (LAC-WM), which operates within a learned unified latent action space shared across diverse e…

Cited by 0SourceScholar
2026

Dexterity from Smart Lenses: Multi-Fingered Robot Manipulation with In-The-Wild Human Demonstrations

ICRA 2026poster

Learning multi-fingered robot policies from humans performing daily tasks in natural environments has long been a grand goal in the robotics community. Achieving this would mark significant progress toward generalizable robot manipulation in human environments, as it would reduce the reliance on lab…

2025

DexterityGen: Foundation Controller for Unprecedented Dexterity

RSS 2025poster

Teaching robots dexterous manipulation skills, such as tool use, presents a significant challenge. Current approaches can be broadly categorized into two strategies: human teleoperation (for imitation learning) and sim-to-real reinforcement learning. The first approach is difficult as it is hard fo…

Cited by 9PDFScholar
2025

Geometric Retargeting: A Principled, Ultrafast Neural Hand Retargeting Algorithm

IROS 2025

We introduce Geometric Retargeting (GeoRT), an ultrafast, and principled neural hand retargeting algorithm for teleoperation, developed as part of our recent Dexterity Gen (DexGen) system [1]. GeoRT converts human finger keypoints to robot hand keypoints at 1KHz, achieving state-of-the-art speed and

Cited by 17SourceScholar
2025

OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction

ICML 2025poster

Vision-Language-Action (VLA) models aim to predict robotic actions based on visual observations and language instructions. Existing approaches require fine-tuning pre-trained vision-language models (VLMs) as visual and language features are independently fed into downstream policies, degrading the p…

2025

Self-supervised perception for tactile skin covered dexterous hands

CoRL 2025poster

We present PercepSkin, a pre-trained encoder for magnetic skin sensors distributed across the fingertips, phalanges, and palm of a dexterous robot hand. Magnetic tactile skins offer a flexible form factor for hand-wide coverage with fast response times, in contrast to vision-based tactile sensors t…

Cited by 0SourceScholar
2025

Tactile Beyond Pixels: Multisensory Touch Representations for Robot Manipulation

CoRL 2025oral

We present TacX, the first multisensory touch representations across four tactile modalities: image, audio, motion, and pressure. Trained on ~1M contact-rich interactions collected with the Digit 360 sensor, TacX captures complementary touch signals at diverse temporal and spatial scales. By leverag…

Cited by 0SourceScholar
2024

Sparsh: Self-supervised touch representations for vision-based tactile sensing

CoRL 2024poster

In this work, we introduce general purpose touch representations for the increasingly accessible class of vision-based tactile sensors. Such sensors have led to many recent advances in robot manipulation as they markedly complement vision, yet solutions today often rely on task and sensor specific h…

Cited by 10SourcecodeScholar
2024

What Do We Learn from a Large-Scale Study of Pre-Trained Visual Representations in Sim and Real Environments?

ICRA 2024poster

We present a large empirical investigation on the use of pre-trained visual representations (PVRs) for training downstream policies that execute real-world tasks. Our study involves five different PVRs, each trained for five distinct manipulation or indoor navigation tasks. We performed this evaluat…

Cited by 6SourceScholar
2023

Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?

NeurIPS 2023poster

We present the largest and most comprehensive empirical study of pre-trained visual representations (PVRs) or visual ‘foundation models’ for Embodied AI. First, we curate CortexBench, consisting of 17 different tasks spanning locomotion, navigation, dexterous, and mobile manipulation. Next, we syste…

Cited by 161SourcePDFScholar
2018

Leveraging Motion Priors in Videos for Improving Human Segmentation

ECCV 2018poster

Despite many advances in deep-learning based semantic segmentation, performance drop due to distribution mismatch is often encountered in the real world. Recently, a few domain adaptation and active learning approaches have been proposed to mitigate the performance drop. However, very little attenti…

Cited by 1SourcePDFScholar
2015

Using parallel stiffness to achieve improved locomotive efficiency with the Sandia STEPPR robot

ICRA 2015poster

In this paper we introduce STEPPR (Sandia Transmission-Efficient Prototype Promoting Research), a bipedal robot designed to explore efficient bipedal walking. The initial iteration of this robot achieves efficient motions through powerful electromagnetic actuators and highly back-drivable synthetic…

Cited by 0SourceScholar