← Search

Tao Kong

42 accepted papers

2026

Manipulation as in Simulation: Enabling Accurate Geometry Perception in Robots

ICLR 2026poster

Modern robotic manipulation primarily relies on visual observations in a 2D color space for skill learning but suffers from poor generalization. In contrast, humans, living in a 3D world, depend more on physical properties-such as distance, size, and shape-than on texture when interacting with objec…

Cited by 0SourcecodeScholar
2026

RoboOmni: Actions Are Just Another Modality for Your Vision-Language Models

ICML 2026poster

Integrating Vision-Language Models (VLMs) into robotics has facilitated the development of generalizable Vision-Language Action (VLA) policies. However, unified discrete frameworks lag behind decoupled continuous designs due to limitations in action chunking and temporal modeling. To address this, w…

Cited by 0SourceScholar
2025

BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models

NeurIPS 2025poster

Recently, leveraging pre-trained vision-language models (VLMs) for building vision-language-action (VLA) models has emerged as a promising approach to effective robot manipulation learning. However, only few methods incorporate 3D signals into VLMs for action prediction, and they do not fully levera…

Cited by 0SourcecodeScholar
2025

Chain-of-Action: Trajectory Autoregressive Modeling for Robotic Manipulation

NeurIPS 2025poster

We present Chain-of-Action (CoA), a novel visuomotor policy paradigm built upon Trajectory Autoregressive Modeling. Unlike conventional approaches that predict next step action(s) forward, CoA generates an entire trajectory by explicit backward reasoning with task-specific goals through an action-le…

Cited by 0SourceScholar
2025

GR-MG: Leveraging Partially-Annotated Data via Multi-Modal Goal-Conditioned Policy

RA-L 2025

The robotics community has consistently aimed to achieve generalizable robot manipulation with flexible natural language instructions. One primary challenge is that obtaining robot trajectories fully annotated with both actions and texts is time-consuming and labor-intensive. However, partially-anno

Cited by 39SourcecodeScholar
2025

Human-assisted Robotic Policy Refinement via Action Preference Optimization

NeurIPS 2025poster

Establishing a reliable and iteratively refined robotic system is essential for deploying real-world applications. While Vision-Language-Action (VLA) models are widely recognized as the foundation model for such robotic deployment, their reliance on offline expert demonstrations critically limi…

Cited by 0SourcecodeScholar
2025

IRASim: A Fine-Grained World Model for Robot Manipulation

ICCV 2025poster

World models allow autonomous agents to plan and explore by predicting the visual outcomes of different actions. However, for robot manipulation, it is challenging to accurately model the fine-grained robot-object interaction within the visual space using existing methods which overlook precise alig…

2025

ProcWorld: Benchmarking Large Model Planning in Reachability-Constrained Environments

EMNLP 2025

We introduce ProcWorld, a large-scale benchmark for partially observable embodied spatial reasoning and long-term planning with large language models (LLM) and vision language models (VLM). ProcWorld features a wide range of challenging embodied navigation and object manipulation tasks, covering 16

Cited by 0SourcePDFScholar
2025

World Model-Based Perception for Visual Legged Locomotion

ICRA 2025

Legged locomotion over various terrains is challenging and requires precise perception of the robot and its surroundings from both proprioception and vision. However, learning directly from high-dimensional visual input is often data-inefficient and intricate. To address this issue, traditional meth

Cited by 23SourcecodeScholar
2024

Exploring Target Representations for Masked Autoencoders

ICLR 2024poster

Masked autoencoders have become popular training paradigms for self-supervised visual representation learning. These models randomly mask a portion of the input and reconstruct the masked portion according to assigned target representations. In this paper, we show that a careful choice of the target…

2024

Towards Unified Interactive Visual Grounding in The Wild

ICRA 2024poster

Interactive visual grounding in Human-Robot Interaction (HRI) is challenging yet practical due to the inevitable ambiguity in natural languages. It requires robots to disambiguate the user’s input by active information gathering. Previous approaches often rely on predefined templates to ask disambig…

Cited by 3SourcecodeScholar
2024

Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

ICLR 2024poster

Generative pre-trained models have demonstrated remarkable effectiveness in language and vision domains by learning useful representations. In this paper, we extend the scope of this effectiveness by showing that visual robot manipulation can significantly benefit from large-scale video generative p…

2024

Vision-Language Foundation Models as Effective Robot Imitators

ICLR 2024spotlight

Recent progress in vision language foundation models has shown their ability to understand multimodal data and resolve complicated vision language tasks, including robotics manipulation. We seek a straightforward way of making use of existing vision-language models (VLMs) with simple fine-tuning on…

Cited by 133SourcePDFScholar
2024

What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?

NAACL 2024long

Recent advancements in GPT-4V have displayed remarkable multi-modal capabilities in processing image inputs and following open-ended instructions. Despite these advancements, there is considerable scope for enhancing open-source multi-modal LLMs, especially in terms of multi-modal understanding accu…

2023

Exploring Visual Pre-training for Robot Manipulation: Datasets, Models and Methods

IROS 2023poster

Visual pre-training with large-scale real-world data has made great progress in recent years, showing great potential in robot learning with pixel observations. However, the recipes of visual pre-training for robot manipulation tasks are yet to be built. In this paper, we thoroughly investigate the…

Cited by 16SourcecodeScholar
2023

MOMA-Force: Visual-Force Imitation for Real-World Mobile Manipulation

IROS 2023poster

In this paper, we present a novel method for mobile manipulators to perform multiple contact-rich manipulation tasks. While learning-based methods have the potential to generate actions in an end-to-end manner, they often suffer from insufficient action accuracy and robustness against noise. On the…

Cited by 12SourcecodeScholar
2022

Generative Category-Level Shape and Pose Estimation with Semantic Primitives

CoRL 2022poster

Empowering autonomous agents with 3D understanding for daily objects is a grand challenge in robotics applications. When exploring in an unknown environment, existing methods for object pose estimation are still not satisfactory due to the diversity of object shapes. In this paper, we propose a nove…

Cited by 28SourcecodeScholar
2022

Image BERT Pre-training with Online Tokenizer

ICLR 2022poster

The success of language Transformers is primarily attributed to the pretext task of masked language modeling (MLM), where texts are first tokenized into semantically meaningful pieces. In this work, we study masked image modeling (MIM) and indicate the necessity and challenges of using a semanticall…

Cited by 1044SourcePDFScholar
2022

Learning Design and Construction with Varying-Sized Materials via Prioritized Memory Resets

ICRA 2022poster

Can a robot autonomously learn to design and construct a bridge from varying-sized blocks without a blueprint? It is a challenging task with long horizon and sparse reward - the robot has to figure out physically stable design schemes and feasible actions to manipulate and transport blocks. Due to d…

Cited by 4SourcecodeScholar
2022

Towards Unifying Reference Expression Generation and Comprehension

EMNLP 2022main

Reference Expression Generation (REG) and Comprehension (REC) are two highly correlated tasks. Modeling REG and REC simultaneously for utilizing the relation between them is a promising way to improve both. However, the problem of distinct inputs, as well as building connections between them in a si…

2021

Adversarial Option-Aware Hierarchical Imitation Learning

ICML 2021spotlight

It has been a challenge to learning skills for an agent from long-horizon unannotated demonstrations. Existing approaches like Hierarchical Imitation Learning(HIL) are prone to compounding errors or suboptimal solutions. In this paper, we propose Option-GAIL, a novel method to learn skills at long h…

2021

Dense Contrastive Learning for Self-Supervised Visual Pre-Training

CVPR 2021poster

To date, most existing self-supervised learning methods are designed and optimized for image classification. These pre-trained models can be sub-optimal for dense prediction tasks due to the discrepancy between image-level prediction and pixel-level prediction. To fill this gap, we aim to design an…

Cited by 856PDFScholar
2021

Locate Then Segment: A Strong Pipeline for Referring Image Segmentation

CVPR 2021poster

Referring image segmentation aims to segment the objects referred by a natural language expression. Previous methods usually focus on designing an implicit and recurrent feature interaction mechanism to fuse the visual-linguistic features to directly generate the final segmentation mask without expl…

Cited by 162PDFScholar
2021

Scale-Aware Automatic Augmentation for Object Detection

CVPR 2021poster

We propose Scale-aware AutoAug to learn data augmentation policies for object detection. We define a new scale-aware search space, where both image- and box-level augmentations are designed for maintaining scale invariance. Upon this search space, we propose a new search metric, termed Pareto Scale…

Cited by 56PDFcodeScholar
2021

Simultaneous Semantic and Collision Learning for 6-DoF Grasp Pose Estimation

IROS 2021poster

Grasping in cluttered scenes has always been a great challenge for robots, due to the requirement of the ability to well understand the scene and object information. Previous works usually assume that the geometry information of the objects is available, or utilize a step-wise, multi-stage strategy…

Cited by 64SourcecodeScholar
2021

Sparse R-CNN: End-to-End Object Detection With Learnable Proposals

CVPR 2021poster

We present Sparse R-CNN, a purely sparse method for object detection in images. Existing works on object detection heavily rely on dense object candidates, such as k anchor boxes pre-defined on all grids of image feature map of size HxW. In our method, however, a fixed sparse set of learned object p…

Cited by 1491PDFcodeScholar
2020

SOLOv2: Dynamic and Fast Instance Segmentation

NeurIPS 2020poster

In this work, we design a simple, direct, and fast framework for instance segmentation with strong performance. To this end, we propose a novel and effective approach, termed SOLOv2, following the principle of the SOLO method [32]. First, our new framework is empowered by an efficient and holistic i…

2019

Attention-based Transfer Learning for Brain-computer Interface

ICASSP 2019accepted

Different functional areas of the human brain play different roles in brain activity, which has not been paid sufficient research attention in the brain-computer interface (BCI) field. This paper presents a new approach for electroencephalography (EEG) classification that applies attention-based tra…

Cited by 0SourceScholar
2019

Zoom-In-To-Check: Boosting Video Interpolation via Instance-Level Discrimination

CVPR 2019poster

We propose a light-weight video frame interpolation algorithm. Our key innovation is an instance-level supervision that allows information to be learned from the high-resolution version of similar objects. Our experiment shows that the proposed method can generate state-of-the-art results across di…

Cited by 31PDFScholar
2018

Deep Feature Pyramid Reconfiguration for Object Detection

ECCV 2018poster

State-of-the-art object detectors usually learn multi-scale representations to get better results by employing feature pyramids. However, the current designs for feature pyramids are still inefficient to integrate the semantic information over different scales. In this paper, we begin by investigati…

2017

RON: Reverse Connection With Objectness Prior Networks for Object Detection

CVPR 2017poster

We present RON, an efficient and effective framework for generic object detection. Our motivation is to smartly associate the best of the region-based (e.g., Faster R-CNN) and region-free (e.g., SSD) methodologies. Under fully convolutional architecture, RON mainly focuses on two fundamental problem…

Cited by 539PDFScholar
2016

HyperNet: Towards Accurate Region Proposal Generation and Joint Object Detection

CVPR 2016spotlight

Almost all of the current top-performing object detection networks employ region proposals to guide the search for object instances. State-of-the-art region proposal methods usually need several thousand proposals to get high recall, thus hurting the detection efficiency. Although the latest Region…

Cited by 1150PDFScholar