← Search

Tiancai Wang

35 accepted papers

2026

BFA: Best-Feature-Aware Fusion for Multi-View Fine-Grained Manipulation

ICRA 2026poster

In real-world scenarios, multi-view cameras are typically employed for fine-grained manipulation tasks. Existing approaches (e.g., ACT ) tend to treat multi-view features equally and directly concatenate them for policy learning. How ever, it will introduce redundant visual information and bring hig…

2026

GS-Playground: A High-Throughput Photorealistic Simulator for Vision-Informed Robot Learning

RSS 2026poster

Embodied AI research is undergoing a shift toward vision-centric perceptual paradigms. While massively parallel simulators have catalyzed breakthroughs in proprioception-based locomotion, their potential remains largely untapped for vision-centric tasks due to the prohibitive computational overhead …

Cited by 0SourceScholar
2026

IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human–Robot Interaction

ICRA 2026poster

Vision-Language-Action (VLA) models leverage pretrained vision-language models (VLMs) to couple perception with robotic control, offering a promising path toward general purpose embodied intelligence. However, current SOTA VLAs are primarily pretrained on multimodal tasks with limited relevance to e…

2026

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation

ICLR 2026poster

Temporal context is essential for robotic manipulation because such tasks are inherently non-Markovian, yet mainstream VLA models typically overlook it and struggle with long-horizon, temporally dependent tasks. Cognitive science suggests that humans rely on working memory to buffer short-lived repr…

Cited by 0SourcecodeScholar
2026

SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic Manipulation

AAAI 2026technical

Robotic manipulation requires precise spatial understanding to interact with objects in the real world. Point-based methods suffer from sparse sampling, leading to the loss of fine-grained semantics. Image-based methods typically feed RGB and depth into 2D backbones pre-trained on 3D auxiliary tasks

Cited by 0SourcePDFScholar
2025

BFA: Best-Feature-Aware Fusion for Multi-View Fine-Grained Manipulation

RA-L 2025

In real-world scenarios, multi-view cameras are typically employed for fine-grained manipulation tasks. Existing approaches (e.g., ACT [1]) tend to treat multi-view features equally and directly concatenate them for policy learning. However, it will introduce redundant visual information and bring h

Cited by 8SourceScholar
2025

Glad: A Streaming Scene Generator for Autonomous Driving

ICLR 2025poster

The generation and simulation of diverse real-world scenes have significant application value in the field of autonomous driving, especially for the corner cases. Recently, researchers have explored employing neural radiance fields or diffusion models to generate novel views or synthetic data under…

Cited by 1SourcePDFScholar
2025

Holistic Tokenizer for Autoregressive Image Generation

ICCV 2025poster

Vanilla autoregressive image generation models generate visual tokens step-by-step, limiting their ability to capture holistic relationships among token sequences. Moreover, because most visual tokenizers map local image patches into latent tokens, global information is limited. To address this, we…

2025

Language Prompt for Autonomous Driving

AAAI 2025technical

A new trend in the computer vision community is to capture objects of interest following flexible human command represented by a natural language prompt. However, the progress of using language prompts in driving scenarios is stuck in a bottleneck due to the scarcity of paired prompt-instance data.…

2025

RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping

ICCV 2025poster

General robotic grasping systems require accurate object affordance perception in diverse open-world scenarios following human instructions. However, current studies suffer from the problem of lacking reasoning-based large-scale affordance prediction data, leading to considerable concern about open-…

Cited by 0SourcePDFScholar
2025

Reconstructive Visual Instruction Tuning

ICLR 2025poster

This paper introduces reconstructive visual instruction tuning (ROSS), a family of Large Multimodal Models (LMMs) that exploit vision-centric supervision signals. In contrast to conventional visual instruction tuning approaches that exclusively supervise text outputs, ROSS prompts LMMs to supervise…

Cited by 65SourcePDFScholar
2025

Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness

ICCV 2025poster

The rapid development of Large Multimodal Models (LMMs) for 2D images and videos has spurred efforts to adapt these models for interpreting 3D scenes. However, the absence of large-scale 3D vision-language datasets has posed a significant obstacle. To address this issue, typical approaches focus on…

Cited by 0SourcePDFScholar
2025

SubjectDrive: Scaling Generative Data in Autonomous Driving via Subject Control

AAAI 2025technical

Autonomous driving progress relies on large-scale annotated datasets. In this work, we explore the potential of generative models to produce vast quantities of freely-labeled data for autonomous driving applications and present SubjectDrive, the first model proven to scale generative data production…

Cited by 11SourcePDFScholar
2025

UniScene: Unified Occupancy-centric Driving Scene Generation

CVPR 2025poster

Generating high-fidelity, controllable, and annotated training data is critical for autonomous driving. Existing methods typically generate a single data form directly from a coarse scene layout, which not only fails to output rich data forms required for diverse downstream tasks but also struggles…

2025

Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Generation

NeurIPS 2025poster

In this work, we present a novel direction to build an image tokenizer directly on top of a frozen vision foundation model, which is a largely underexplored area. Specifically, we employ a frozen vision foundation model as the encoder of our tokenizer. To enhance its effectiveness, we introduce two…

Cited by 0SourcecodeScholar
2024

Far3D: Expanding the Horizon for Surround-View 3D Object Detection

AAAI 2024technical

Recently 3D object detection from surround-view images has made notable advancements with its low deployment cost. However, most works have primarily focused on close perception range while leaving long-range detection less explored. Expanding existing methods directly to cover long distances poses…

2024

Merlin: Empowering Multimodal LLMs with Foresight Minds

ECCV 2024poster

"Humans can foresee the future based on present observations, a skill we term as foresight minds. However, this capability remains under-explored within existing MLLMs, hindering their capacity to understand intentions behind subjects. To address this, we integrate the future modeling into MLLMs. By…

2024

Panacea: Panoramic and Controllable Video Generation for Autonomous Driving

CVPR 2024poster

The field of autonomous driving increasingly demands high-quality annotated training data. In this paper we propose Panacea an innovative approach to generate panoramic and controllable videos in driving scenarios capable of yielding an unlimited numbers of diverse annotated samples pivotal for auto…

Cited by 50SourcePDFScholar
2024

Stream Query Denoising for Vectorized HD-Map Construction

ECCV 2024poster

"This paper introduces the Stream Query Denoising (SQD) strategy, a novel and general approach for high-definition map (HD-map) construction. SQD is designed to improve the modeling capability of map elements by learning temporal consistency. Specifically, SQD involves the process of denoising the q…

Cited by 24SourcePDFScholar
2024

TopoMLP: A Simple yet Strong Pipeline for Driving Topology Reasoning

ICLR 2024poster

Topology reasoning aims to comprehensively understand road scenes and present drivable routes in autonomous driving. It requires detecting road centerlines (lane) and traffic elements, further reasoning their topology relationship, \textit{i.e.}, lane-lane topology, and lane-traffic topology. In thi…

2023

Cross Modal Transformer: Towards Fast and Robust 3D Object Detection

ICCV 2023poster

In this paper, we propose a robust 3D detector, named Cross Modal Transformer (CMT), for end-to-end 3D multi-modal detection. Without explicit view transformation, CMT takes the image and point clouds tokens as inputs and directly outputs accurate 3D bounding boxes. The spatial alignment of multi-mo…

Cited by 118PDFcodeScholar
2023

Exploring Object-Centric Temporal Modeling for Efficient Multi-View 3D Object Detection

ICCV 2023poster

In this paper, we propose a long-sequence modeling framework, named StreamPETR, for multi-view 3D object detection. Built upon the sparse query design in the PETR series, we systematically develop an object-centric temporal mechanism. The model is performed in an online manner and the long-term hist…

Cited by 235PDFcodeScholar
2023

MOTRv2: Bootstrapping End-to-End Multi-Object Tracking by Pretrained Object Detectors

CVPR 2023poster

In this paper, we propose MOTRv2, a simple yet effective pipeline to bootstrap end-to-end multi-object tracking with a pretrained object detector. Existing end-to-end methods, e.g. MOTR and TrackFormer are inferior to their tracking-by-detection counterparts mainly due to their poor detection perfor…

2023

OnlineRefer: A Simple Online Baseline for Referring Video Object Segmentation

ICCV 2023poster

Referring video object segmentation (RVOS) aims at segmenting an object in a video following human instruction. Current state-of-the-art methods fall into an offline pattern, in which each clip independently interacts with text embedding for cross-modal understanding. They usually present that the o…

Cited by 55PDFcodeScholar
2023

PETRv2: A Unified Framework for 3D Perception from Multi-Camera Images

ICCV 2023poster

In this paper, we propose PETRv2, a unified framework for 3D perception from multi-view images. Based on PETR, PETRv2 explores the effectiveness of temporal modeling, which utilizes the temporal information of previous frames to boost 3D object detection. More specifically, we extend the 3D position…

Cited by 397PDFcodeScholar
2023

Referring Multi-Object Tracking

CVPR 2023poster

Existing referring understanding tasks tend to involve the detection of a single text-referred object. In this paper, we propose a new and general referring understanding task, termed referring multi-object tracking (RMOT). Its core idea is to employ a language expression as a semantic cue to guide…

2022

MOTR: End-to-End Multiple-Object Tracking with TRansformer

ECCV 2022poster

"Temporal modeling of objects is a key challenge in multiple-object tracking (MOT). Existing methods track by associating detections through motion-based and appearance-based similarity heuristics. The post-processing nature of association prevents end-to-end exploitation of temporal variations in v…

2022

PETR: Position Embedding Transformation for Multi-View 3D Object Detection

ECCV 2022poster

"In this paper, we develop position embedding transformation (PETR) for multi-view 3D object detection. PETR encodes the position information of 3D coordinates into image features, producing the 3D position-aware features. Object query can perceive the 3D position-aware features and perform end-to-e…

2022

Tree Energy Loss: Towards Sparsely Annotated Semantic Segmentation

CVPR 2022poster

Sparsely annotated semantic segmentation (SASS) aims to train a segmentation network with coarse-grained (i.e.,point-, scribble-, and block-wise) supervisions, where only a small proportion of pixels are labeled in each image. In this paper, we propose a novel tree energy loss for SASS by providing…

Cited by 80PDFcodeScholar
2021

Co-mining: Self-Supervised Learning for Sparsely Annotated Object Detection

AAAI 2021technical

Object detectors usually achieve promising results with the supervision of complete instance annotations. However, their performance is far from satisfactory with sparse instance annotations. Most existing methods for sparsely annotated object detection either re-weight the loss of hard negative sam…

2021

SOLQ: Segmenting Objects by Learning Queries

NeurIPS 2021poster

In this paper, we propose an end-to-end framework for instance segmentation. Based on the recently introduced DETR, our method, termed SOLQ, segments objects by learning unified queries. In SOLQ, each query represents one object and has multiple representations: class, location and mask. The object…

2020

Learning Human-Object Interaction Detection Using Interaction Points

CVPR 2020poster

Understanding interactions between humans and objects is one of the fundamental problems in visual classification and an essential step towards detailed scene understanding. Human-object interaction (HOI) detection strives to localize both the human and an object as well as the identification of com…

Cited by 297PDFcodeScholar
2019

Deep Contextual Attention for Human-Object Interaction Detection

ICCV 2019poster

Human-object interaction detection is an important and relatively new class of visual relationship detection tasks, essential for deeper scene understanding. Most existing approaches decompose the problem into object localization and interaction recognition. Despite showing progress, these approache…

Cited by 130PDFScholar
2019

Efficient Featurized Image Pyramid Network for Single Shot Detector

CVPR 2019poster

Single-stage object detectors have recently gained popularity due to their combined advantage of high detection accuracy and real-time speed. However, while promising results have been achieved by these detectors on standard-sized objects, their performance on small objects is far from satisfactory.…

Cited by 134PDFScholar
2019

Learning Rich Features at High-Speed for Single-Shot Object Detection

ICCV 2019poster

Single-stage object detection methods have received significant attention recently due to their characteristic realtime capabilities and high detection accuracies. Generally, most existing single-stage detectors follow two common practices: they employ a network backbone that is pretrained on ImageN…

Cited by 143PDFcodeScholar