← Search

Jinghuan Shang

11 accepted papers

2026

VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic Routing

ICLR 2026poster

Pretrained vision foundation models (VFMs) advance robotic learning via rich visual representations, yet individual VFMs typically excel only in specific domains, limiting generality across tasks. Distilling multiple VFMs into a unified representation can mitigate this limitation but often yields in…

Cited by 0SourceScholar
2025

LLaRA: Supercharging Robot Learning Data for Vision-Language Policy

ICLR 2025poster

Vision Language Models (VLMs) have recently been leveraged to generate robotic actions, forming Vision-Language-Action (VLA) models. However, directly adapting a pretrained VLM for robotic control remains challenging, particularly when constrained by a limited number of robot demonstrations. In this…

2024

Crossway Diffusion: Improving Diffusion-based Visuomotor Policy via Self-supervised Learning

ICRA 2024poster

Diffusion models have been adopted for behavioral cloning in a sequence modeling fashion, benefiting from their exceptional capabilities in modeling complex data distributions. The standard diffusion-based policy iteratively denoises action sequences from random noise conditioned on the input states…

Cited by 29SourcecodeScholar
2024

Theia: Distilling Diverse Vision Foundation Models for Robot Learning

CoRL 2024poster

Vision-based robot policy learning, which maps visual inputs to actions, necessitates a holistic understanding of diverse visual tasks beyond single-task needs like classification or segmentation. Inspired by this, we introduce Theia, a vision foundation model for robot learning that distills multip…

Cited by 18SourcecodeScholar
2022

Does Self-supervised Learning Really Improve Reinforcement Learning from Pixels?

NeurIPS 2022accept

We investigate whether self-supervised learning (SSL) can improve online reinforcement learning (RL) from pixels. We extend the contrastive reinforcement learning framework (e.g., CURL) that jointly optimizes SSL and RL losses and conduct an extensive amount of experiments with various self-supervis…

Cited by 35SourcePDFScholar
2022

Learning Viewpoint-Agnostic Visual Representations by Recovering Tokens in 3D Space

NeurIPS 2022accept

Humans are remarkably flexible in understanding viewpoint changes due to visual cortex supporting the perception of 3D structure. In contrast, most of the computer vision models that learn visual representation from a pool of 2D images often fail to generalize over novel camera viewpoints. Recently,…

2022

StARformer: Transformer with State-Action-Reward Representations for Visual Reinforcement Learning

ECCV 2022poster

"Reinforcement Learning (RL) can be considered as a sequence modeling task: given a sequence of past state-action-reward experiences, an agent predicts a sequence of next actions. In this work, we propose State-Action-Reward Transformer (StARformer) for visual RL, which explicitly models short-term…

2021

Self-Supervised Disentangled Representation Learning for Third-Person Imitation Learning

IROS 2021poster

Humans learn to imitate by observing others. However, robot imitation learning generally requires expert demonstrations in the first-person view (FPV). Collecting such FPV videos for every robot could be very expensive.Third-person imitation learning (TPIL) is the concept of learning action policies…

Cited by 25SourceScholar