← Search

Renhao Wang

10 accepted papers

2025

The Sound of Simulation: Learning Multimodal Sim-to-Real Robot Policies with Generative Audio

CoRL 2025oral

Robots must integrate multiple sensory modalities to act effectively in the real world. Yet, learning such multimodal policies at scale remains challenging. Simulation offers a viable solution, but while vision has benefited from high-fidelity simulators, other modalities (e.g. sound) can be notorio…

Cited by 0SourceScholar
2024

Improving Distant 3D Object Detection Using 2D Box Supervision

CVPR 2024poster

Improving the detection of distant 3d objects is an important yet challenging task. For camera-based 3D perception the annotation of 3d bounding relies heavily on LiDAR for accurate depth information. As such the distance of annotation is often limited due to the sparsity of LiDAR points on distant…

Cited by 3SourcePDFScholar
2024

Self-Supervised Audio-Visual Soundscape Stylization

ECCV 2024poster

"Speech sounds convey a great deal of information about the scenes, resulting in a variety of effects ranging from reverberation to additional ambient sounds. In this paper, we manipulate input speech to sound as though it was recorded within a different scene, given an audio-visual conditional exam…

Cited by 4SourcePDFScholar
2023

For Pre-Trained Vision Models in Motor Control, Not All Policy Learning Methods are Created Equal

ICML 2023poster

In recent years, increasing attention has been directed to leveraging pre-trained vision models for motor control. While existing works mainly emphasize the importance of this pre-training phase, the arguably equally important role played by downstream policy learning during control-specific fine-tu…

Cited by 26SourcePDFScholar
2023

Programmatically Grounded, Compositionally Generalizable Robotic Manipulation

ICLR 2023top-25%

Robots operating in the real world require both rich manipulation skills as well as the ability to semantically reason about when to apply those skills. Towards this goal, recent works have integrated semantic representations from large-scale pretrained vision-language (VL) models into manipulation…

2023

Robust and Controllable Object-Centric Learning through Energy-based Models

ICLR 2023poster

Humans are remarkably good at understanding and reasoning about complex visual scenes. The capability of decomposing low-level observations into discrete objects allows us to build a grounded abstract representation and identify the compositional structure of the world. Thus it is a crucial step for…

Cited by 12SourcePDFScholar
2022

CYBORGS: Contrastively Bootstrapping Object Representations by Grounding in Segmentation

ECCV 2022poster

"Many recent approaches in contrastive learning have worked to close the gap between pretraining on iconic images like ImageNet and pretraining on complex scenes like COCO. This gap exists largely because commonly used random crop augmentations obtain semantically inconsistent content in crowded sce…