← Search

Ruoxuan Feng

7 accepted papers

2026

AnyTouch 2: General Optical Tactile Representation Learning For Dynamic Tactile Perception

ICLR 2026poster

Real-world contact-rich manipulation demands robots to perceive temporal tactile feedback, capture subtle surface deformations, and reason about object properties and force dynamics. Although optical tactile sensors are uniquely capable of providing such rich information, existing tactile datasets a…

Cited by 0SourcecodeScholar
2025

AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-tactile Sensors

ICLR 2025poster

Visuo-tactile sensors aim to emulate human tactile perception, enabling robots to precisely understand and manipulate objects. Over time, numerous meticulously designed visuo-tactile sensors have been integrated into robotic systems, aiding in completing various tasks. However, the distinct data cha…

2025

Phoenix: A Motion-based Self-Reflection Framework for Fine-grained Robotic Action Correction

CVPR 2025poster

Building a generalizable self-correction system is crucial for robots to recover from failures. Despite advancements in Multimodal Large Language Models (MLLMs) that empower robots with semantic reflection ability for failure, translating semantic reflection into how to correct fine-grained robotic…

2024

Enhancing Multimodal Cooperation via Sample-level Modality Valuation

CVPR 2024poster

One primary topic of multimodal learning is to jointly incorporate heterogeneous information from different modalities. However most models often suffer from unsatisfactory multimodal cooperation which cannot jointly utilize all modalities well. Some methods are proposed to identify and enhance the…

2024

Play to the Score: Stage-Guided Dynamic Multi-Sensory Fusion for Robotic Manipulation

CoRL 2024poster

Humans possess a remarkable talent for flexibly alternating to different senses when interacting with the environment. Picture a chef skillfully gauging the timing of ingredient additions and controlling the heat according to the colors, sounds, and aromas, seamlessly navigating through every stage…

Cited by 7SourceScholar
2023

MMCosine: Multi-Modal Cosine Loss Towards Balanced Audio-Visual Fine-Grained Learning

ICASSP 2023accepted

Audio-visual learning helps to comprehensively under-stand the world by fusing practical information from multiple modalities. However, recent studies show that the imbalanced optimization of uni-modal encoders in a joint-learning model is a bottleneck to enhancing the model’s performance. We furthe…

Cited by 0SourceScholar