← Search

Ya Jing

9 accepted papers

2024

Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

ICLR 2024poster

Generative pre-trained models have demonstrated remarkable effectiveness in language and vision domains by learning useful representations. In this paper, we extend the scope of this effectiveness by showing that visual robot manipulation can significantly benefit from large-scale video generative p…

2024

Vision-Language Foundation Models as Effective Robot Imitators

ICLR 2024spotlight

Recent progress in vision language foundation models has shown their ability to understand multimodal data and resolve complicated vision language tasks, including robotics manipulation. We seek a straightforward way of making use of existing vision-language models (VLMs) with simple fine-tuning on…

Cited by 133SourcePDFScholar
2023

Exploring Visual Pre-training for Robot Manipulation: Datasets, Models and Methods

IROS 2023poster

Visual pre-training with large-scale real-world data has made great progress in recent years, showing great potential in robot learning with pixel observations. However, the recipes of visual pre-training for robot manipulation tasks are yet to be built. In this paper, we thoroughly investigate the…

Cited by 16SourcecodeScholar
2023

MOMA-Force: Visual-Force Imitation for Real-World Mobile Manipulation

IROS 2023poster

In this paper, we present a novel method for mobile manipulators to perform multiple contact-rich manipulation tasks. While learning-based methods have the potential to generate actions in an end-to-end manner, they often suffer from insufficient action accuracy and robustness against noise. On the…

Cited by 12SourcecodeScholar
2022

Towards Unifying Reference Expression Generation and Comprehension

EMNLP 2022main

Reference Expression Generation (REG) and Comprehension (REC) are two highly correlated tasks. Modeling REG and REC simultaneously for utilizing the relation between them is a promising way to improve both. However, the problem of distinct inputs, as well as building connections between them in a si…

2021

Locate Then Segment: A Strong Pipeline for Referring Image Segmentation

CVPR 2021poster

Referring image segmentation aims to segment the objects referred by a natural language expression. Previous methods usually focus on designing an implicit and recurrent feature interaction mechanism to fuse the visual-linguistic features to directly generate the final segmentation mask without expl…

Cited by 162PDFScholar
2018

Skeleton-Based Action Recognition with Spatial Reasoning and Temporal Stack Learning

ECCV 2018poster

Skeleton-based action recognition has made great progress recently, but many problems still remain unsolved. For example, the representations of skeleton sequences captured by most of the previous methods lack spatial structure information and detailed temporal dynamics features. In this paper, we p…

Cited by 439SourcePDFScholar