← Search

Xiao Tan

52 accepted papers

2026

Cooperative Informed Tree (CoIT*): Cooperative Bi-Directional Multi-Resolution Motion Planning with Adaptive Edge Screening

ICRA 2026poster

In informed search-based path planning, heuristic functions that incorporate problem knowledge are essential for guiding the search and improving efficiency. The accuracy and computational cost of these heuristics are therefore critical to performance. However, accuracy and computational efficiency …

Cited by 0Scholar
2026

From Intuition to Investigation: A Tool-Augmented Reasoning MLLM Framework for Generalizable Face Anti-Spoofing

CVPR 2026

Face recognition remains vulnerable to presentation attacks, calling for robust Face Anti-Spoofing (FAS) solutions. Recent MLLM-based FAS methods reformulate the binary classification task as the generation of brief textual descriptions to improve cross-domain generalization. However, their generali

Cited by 0SourceScholar
2026

Hugging Visual Prompt and Segmentation Tokens: Consistency Learning for Fine-Grained Visual Understanding in MLLMs

CVPR 2026

Recently, multimodal large language models (MLLMs) have achieved remarkable success in general multimodal tasks. Increasing attention has been given to leveraging MLLMs for fine-grained visual understanding, such as region-level captioning and pixel-level grounding. However, most existing approaches

Cited by 0SourceScholar
2026

OptiMVMap: Offline Vectorized Map Construction via Optimal Multi-vehicle Perspectives

CVPR 2026

Offline vectorized maps constitute critical infrastructure for high-precision autonomous driving and mapping services. Existing approaches rely predominantly on single ego-vehicle trajectories, which fundamentally suffer from viewpoint insufficiency: while memory-based methods extend observation tim

Cited by 0SourcecodeScholar
2026

Safe Navigation Under State Uncertainty: Online Adaptation for Robust Control Barrier Functions

RA-L 2026

Measurements and state estimates are often imperfect in control practice, posing challenges for safety-critical applications, where safety guarantees rely on accurate state information. In the presence of estimation errors, several prior robust control barrier function (R-CBF) formulations have impo

Cited by 5SourcecodeScholar
2026

Safe Navigation under State Uncertainty: Online Adaptation for Robust Control Barrier Functions

ICRA 2026poster

Measurements and state estimates are often imperfect in control practice, posing challenges for safety-critical applications, where safety guarantees rely on accurate state information. In the presence of estimation errors, several prior robust control barrier function (R-CBF) formulations have impo…

2026

UniFuture: A 4D Driving World Model for Future Generation and Perception

ICRA 2026poster

We present UniFuture, a unified 4D Driving World Model designed to simulate the dynamic evolution of the 3D physical world. Unlike existing driving world models that focus solely on 2D pixel-level video generation (lacking geometry) or static perception (lacking temporal dynamics), our approach brid…

2026

ViLoMem: Agentic Learner with Grow-and-Refine Multimodal Semantic Memory

CVPR 2026

MLLMs exhibit strong reasoning on isolated queries, yet they operate de novo--solving each problem independently and often repeating the same mistakes. Existing memory-augmented agents mainly store past trajectories for reuse. However, trajectory-based memory suffers from brevity bias, gradually los

Cited by 0SourcecodeScholar
2025

AdaDrive: Self-Adaptive Slow-Fast System for Language-Grounded Autonomous Driving

ICCV 2025poster

Effectively integrating Large Language Models (LLMs) into autonomous driving requires a balance between leveraging high-level reasoning and maintaining real-time efficiency. Existing approaches either activate LLMs too frequently, causing excessive computational overhead, or use fixed schedules, fai…

2025

Bi-perspective Splitting Defense: Achieving Clean-Seed-Free Backdoor Security

ICML 2025poster

Backdoor attacks have seriously threatened deep neural networks (DNNs) by embedding concealed vulnerabilities through data poisoning. To counteract these attacks, training benign models from poisoned data garnered considerable interest from researchers. High-performing defenses often rely on additio…

Cited by 0SourcePDFScholar
2025

Explore the LiDAR-Camera Dynamic Adjustment Fusion for 3D Object Detection

ICRA 2025

Camera and LiDAR serve as informative sensors for accurate and robust autonomous driving systems. However, these sensors often exhibit heterogeneous natures, resulting in distributional modality gaps that present significant challenges for fusion. To address this, a robust fusion technique is crucia

Cited by 0SourcecodeScholar
2025

Fin3R: Fine-tuning Feed-forward 3D Reconstruction Models via Monocular Knowledge Distillation

NeurIPS 2025poster

We present Fin3R, a simple, effective, and general fine-tuning method for feed-forward 3D reconstruction models. The family of feed-forward reconstruction model regresses pointmap of all input images to a reference frame coordinate system, along with other auxiliary outputs, in a single forward pass…

Cited by 0SourcecodeScholar
2025

LIRA: Reasoning Reconstruction via Multimodal Large Language Models

ICCV 2025poster

Existing language instruction-guided online 3D reconstruction systems mainly rely on explicit instructions or queryable maps, showing inadequate capability to handle implicit and complex instructions. In this paper, we first introduce a reasoning reconstruction task. This task inputs an implicit ins…

2025

LaneDiffusion: Improving Centerline Graph Learning via Prior Injected BEV Feature Generation

ICCV 2025poster

Centerline graphs, crucial for path planning in autonomous driving, are traditionally learned using deterministic methods. However, these methods often lack spatial reasoning and struggle with occluded or invisible centerlines. Generative approaches, despite their potential, remain underexplored in…

2025

MGMapNet: Multi-Granularity Representation Learning for End-to-End Vectorized HD Map Construction

ICLR 2025poster

The construction of vectorized high-definition map typically requires capturing both category and geometry information of map elements. Current state-of-the-art methods often adopt solely either point-level or instance-level representation, overlooking the strong intrinsic relationship between point…

Cited by 3SourcePDFScholar
2025

Secure Safety Filter: Towards Safe Flight Control under Sensor Attacks

IROS 2025

Modern autopilot systems are prone to sensor attacks that can jeopardize flight safety. To mitigate this risk, we proposed a modular solution: the secure safety filter, which extends the well-established control barrier function (CBF)-based safety filter to account for, and mitigate, sensor attacks.

Cited by 1SourcecodeScholar
2025

Uni$^2$Det: Unified and Universal Framework for Prompt-Guided Multi-dataset 3D Detection

ICLR 2025poster

We present Uni$^2$Det, a brand new framework for unified and universal multi-dataset training on 3D detection, enabling robust performance across diverse domains and generalization to unseen domains. Due to substantial disparities in data distribution and variations in taxonomy across diverse domain…

2025

VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving

ICCV 2025poster

Recent advancements in language-grounded autonomous driving have been significantly promoted by the sophisticated cognition and reasoning capabilities of large language models (LLMs). However, current LLM-based approaches encounter critical challenges: (1) Failure analysis reveals that frequent coll…

2024

Decoupled Pseudo-labeling for Semi-Supervised Monocular 3D Object Detection

CVPR 2024poster

We delve into pseudo-labeling for semi-supervised monocular 3D object detection (SSM3OD) and discover two primary issues: a misalignment between the prediction quality of 3D and 2D attributes and the tendency of depth supervision derived from pseudo-labels to be noisy leading to significant optimiza…

Cited by 7SourcePDFScholar
2024

FasMe: Fast and Sample-efficient Meta Estimator for Precision Matrix Learning in Small Sample Settings

NeurIPS 2024poster

Precision matrix estimation is a ubiquitous task featuring numerous applications such as rare disease diagnosis and neural connectivity exploration. However, this task becomes challenging in small sample settings, where the number of samples is significantly less than the number of dimensions, leadi…

Cited by 0SourcePDFScholar
2024

Interactive 3D Object Detection with Prompts

ECCV 2024poster

"The evolution of 3D object detection hinges not only on advanced models but also on effective and efficient annotation strategies. Despite this progress, the labor-intensive nature of 3D object annotation remains a bottleneck, hindering further development in the field. This paper introduces a nove…

Cited by 0SourcePDFScholar
2024

OPEN: Object-wise Position Embedding for Multi-view 3D Object Detection

ECCV 2024poster

"Accurate depth information is crucial for enhancing the performance of multi-view 3D object detection. Despite the success of some existing multi-view 3D detectors utilizing pixel-wise depth supervision, they overlook two significant phenomena: 1) the depth supervision obtained from LiDAR points is…

2024

PointMamba: A Simple State Space Model for Point Cloud Analysis

NeurIPS 2024poster

Transformers have become one of the foundational architectures in point cloud analysis tasks due to their excellent global modeling ability. However, the attention mechanism has quadratic complexity, making the design of a linear complexity method with global modeling appealing. In this paper, we pr…

2023

A Simple Vision Transformer for Weakly Semi-supervised 3D Object Detection

ICCV 2023poster

Advanced 3D object detection methods usually rely on large-scale, elaborately labeled datasets to achieve good performance. However, labeling the bounding boxes for the 3D objects is difficult and expensive. Although semi-supervised (SS3D) and weakly-supervised 3D object detection (WS3D) methods can…

Cited by 29PDFScholar
2023

Ambiguity-Resistant Semi-Supervised Learning for Dense Object Detection

CVPR 2023poster

With basic Semi-Supervised Object Detection (SSOD) techniques, one-stage detectors generally obtain limited promotions compared with two-stage clusters. We experimentally find that the root lies in two kinds of ambiguities: (1) Selection ambiguity that selected pseudo labels are less accurate, since…

2023

CAPE: Camera View Position Embedding for Multi-View 3D Object Detection

CVPR 2023poster

In this paper, we address the problem of detecting 3D objects from multi-view images. Current query-based methods rely on global 3D position embeddings (PE) to learn the geometric correspondence between images and 3D space. We claim that directly interacting 2D image features with global 3D PE could…

Cited by 51SourcePDFScholar
2023

CFCG: Semi-Supervised Semantic Segmentation via Cross-Fusion and Contour Guidance Supervision

ICCV 2023poster

Current state-of-the-art semi-supervised semantic segmentation (SSSS) methods typically adopt pseudo labeling and consistency regularization between multiple learners with different perturbations. Although the performance is desirable, many issues remain: (1) supervisions from a single learner tend…

Cited by 16PDFScholar
2023

Command-Driven Articulated Object Understanding and Manipulation

CVPR 2023poster

We present Cart, a new approach towards articulated-object manipulations by human commands. Beyond the existing work that focuses on inferring articulation structures, we further support manipulating articulated shapes to align them subject to simple command templates. The key of Cart is to utilize…

2023

Distributed barrier function-enabled human-in-the-loop control for multi-robot systems

ICRA 2023poster

In this work, we propose a distributed control scheme for multi-robot systems in the presence of multiple constraints using control barrier functions. The proposed scheme expands previous work where only one single constraint can be handled. Here we show how to transform multiple constraints to a co…

Cited by 12SourceScholar
2023

Forward Flow for Novel View Synthesis of Dynamic Scenes

ICCV 2023oral

This paper proposes a neural radiance field (NeRF) approach for novel view synthesis of dynamic scenes using forward warping. Existing methods often adopt a static NeRF to represent the canonical space, and render dynamic images at other time steps by mapping the sampled 3D points back to the canoni…

Cited by 48PDFcodeScholar
2023

Gradient-based Sampling for Class Imbalanced Semi-supervised Object Detection

ICCV 2023poster

Current semi-supervised object detection (SSOD) algorithms typically assume class balanced datasets (PASCAL VOC etc.) or slightly class imbalanced datasets (MSCOCO, etc). This assumption can be easily violated since real world datasets can be extremely class imbalanced in nature, thus making the per…

Cited by 13PDFcodeScholar
2023

Semi-DETR: Semi-Supervised Object Detection With Detection Transformers

CVPR 2023poster

We analyze the DETR-based framework on semi-supervised object detection (SSOD) and observe that (1) the one-to-one assignment strategy generates incorrect matching when the pseudo ground-truth bounding box is inaccurate, leading to training inefficiency; (2) DETR-based detectors lack deterministic c…

Cited by 61SourcePDFScholar
2023

StereoDistill: Pick the Cream from LiDAR for Distilling Stereo-Based 3D Object Detection

AAAI 2023technical

In this paper, we propose a cross-modal distillation method named StereoDistill to narrow the gap between the stereo and LiDAR-based approaches via distilling the stereo detectors from the superior LiDAR model at the response level, which is usually overlooked in 3D object detection distillation. Th…

Cited by 10SourcePDFScholar
2022

Detaching and Boosting: Dual Engine for Scale-Invariant Self-Supervised Monocular Depth Estimation

RA-L 2022

Monocular depth estimation (MDE) in the self-supervised scenario has emerged as a promising method as it refrains from the requirement of ground truth depth. Despite continuous efforts, MDE is still sensitive to scale changes especially when all the training samples are from one single camera. Meanw

Cited by 1SourcecodeScholar
2022

Diverse Learner: Exploring Diverse Supervision for Semi-Supervised Object Detection

ECCV 2022poster

"Current state-of-the-art semi-supervised object detection methods (SSOD) typically adopt the teacher-student framework featured with pseudo labeling and Exponential Moving Average (EMA). Although the performance is desirable, many remaining issues still need to be resolved, for example: (1) the tea…

Cited by 5SourcePDFScholar
2022

GitNet: Geometric Prior-Based Transformation for Birds-Eye-View Segmentation

ECCV 2022poster

"Birds-eye-view (BEV) semantic segmentation is critical for autonomous driving for its powerful spatial representation ability. It is challenging to estimate the BEV semantic maps from monocular images due to the spatial gap, since it is implicitly required to realize both the perspective-to-BEV tra…

Cited by 35SourcePDFScholar
2022

Reusing the Task-Specific Classifier as a Discriminator: Discriminator-Free Adversarial Domain Adaptation

CVPR 2022poster

Adversarial learning has achieved remarkable performances for unsupervised domain adaptation (UDA). Existing adversarial UDA methods typically adopt an additional discriminator to play the min-max game with a feature extractor. However, most of these methods failed to effectively leverage the predic…

Cited by 201PDFcodeScholar
2022

Rope3D: The Roadside Perception Dataset for Autonomous Driving and Monocular 3D Object Detection Task

CVPR 2022poster

Concurrent perception datasets for autonomous driving are mainly limited to frontal view with sensors mounted on the vehicle. None of them is designed for the overlooked roadside perception tasks. On the other hand, the data captured from roadside cameras have strengths over frontal-view data, which…

Cited by 135PDFScholar
2022

SGM3D: Stereo Guided Monocular 3D Object Detection

RA-L 2022

Monocular 3D object detection aims to predict the object location, dimension and orientation in 3D space alongside the object category given only a monocular image. It poses a great challenge due to its ill-posed property, which is a critical lack of depth information in the 2D image plane. While ex

Cited by 39SourcecodeScholar
2022

Spatial Pruned Sparse Convolution for Efficient 3D Object Detection

NeurIPS 2022accept

3D scenes are dominated by a large number of background points, which is redundant for the detection task that mainly needs to focus on foreground objects. In this paper, we analyze major components of existing sparse 3D CNNs and find that 3D CNNs ignores the redundancy of data and further amplifies…

Cited by 45SourcePDFScholar
2022

TWIST: Two-Way Inter-Label Self-Training for Semi-Supervised 3D Instance Segmentation

CVPR 2022poster

We explore the way to alleviate the label-hungry problem in a semi-supervised setting for 3D instance segmentation. To leverage the unlabeled data to boost model performance, we present a novel Two-Way Inter-label Self-Training framework named TWIST. It exploits inherent correlations between semanti…

Cited by 29PDFcodeScholar
2021

Revealing the Reciprocal Relations Between Self-Supervised Stereo and Monocular Depth Estimation

ICCV 2021poster

Current self-supervised depth estimation algorithms mainly focus on either stereo or monocular only, neglecting the reciprocal relations between them. In this paper, we propose a simple yet effective framework to improve both stereo and monocular depth estimation by leveraging the underlying complem…

Cited by 34PDFScholar
2021

The Devil Is in the Task: Exploiting Reciprocal Appearance-Localization Features for Monocular 3D Object Detection

ICCV 2021poster

Low-cost monocular 3D object detection plays a fundamental role in autonomous driving, whereas its accuracy is still far from satisfactory. Our objective is to dig into the 3D object detection task and reformulate it as the sub-tasks of object localization and appearance perception, which benefits t…

Cited by 58PDFScholar
2021

Weakly-Supervised Spatio-Temporal Anomaly Detection in Surveillance Video

IJCAI 2021poster

In this paper, we introduce a novel task, referred to as Weakly-Supervised Spatio-Temporal Anomaly Detection (WSSTAD) in surveillance video. Specifically, given an untrimmed video, WSSTAD aims to localize a spatio-temporal tube (i.e., a sequence of bounding boxes at consecutive times) that encloses…

Cited by 75SourcePDFScholar
2020

Associate-3Ddet: Perceptual-to-Conceptual Association for 3D Point Cloud Object Detection

CVPR 2020poster

Object detection from 3D point clouds remains a challenging task, though recent studies pushed the envelope with the deep learning techniques. Owing to the severe spatial occlusion and inherent variance of point density with the distance to sensors, appearance of a same object varies a lot in point…

Cited by 115PDFScholar
2020

Discriminative Sounding Objects Localization via Self-supervised Audiovisual Matching

NeurIPS 2020poster

Discriminatively localizing sounding objects in cocktail-party, i.e., mixed sound scenes, is commonplace for humans, but still challenging for machines. In this paper, we propose a two-stage learning framework to perform self-supervised class-aware sounding object localization. First, we propose to…

2020

Monocular 3D Object Detection via Feature Domain Adaptation

ECCV 2020poster

Monocular 3D object detection is a challenging task due to unreliable depth, resulting in a distinct performance gap between monocular and LiDAR-based approaches. In this paper, we propose a novel domain adaptation based monocular 3D object detection framework named DA-3Ddet, which adapts the featur…

Cited by 58SourcePDFScholar
2020

Segment as Points for Efficient Online Multi-Object Tracking and Segmentation

ECCV 2020poster

Current multi-object tracking and segmentation (MOTS) methods follow the tracking-by-detection paradigm and adopt convolutions for feature extraction. However, as affected by the inherent receptive field, convolution based feature extraction inevitably mixes up the foreground features and the backgr…

2019

Multi-Agent Reinforcement Learning Based Frame Sampling for Effective Untrimmed Video Recognition

ICCV 2019oral

Video Recognition has drawn great research interest and great progress has been made. A suitable frame sampling strategy can improve the accuracy and efficiency of recognition. However, mainstream solutions generally adopt hand-crafted frame sampling strategies for recognition. It could degrade the…

Cited by 165PDFScholar
2019

Perspective-Guided Convolution Networks for Crowd Counting

ICCV 2019poster

In this paper, we propose a novel perspective-guided convolution (PGC) for convolutional neural network (CNN) based crowd counting (i.e. PGCNet), which aims to overcome the dramatic intra-scene scale variations of people due to the perspective effect. While most state-of-the-arts adopt multi-scale o…

Cited by 245PDFcodeScholar
2019

Recognizing Part Attributes With Insufficient Data

ICCV 2019poster

Recognizing the attributes of objects and their parts is central to many computer vision applications. Although great progress has been made to apply object-level recognition, recognizing the attributes of parts remains less applicable since the training data for part attributes recognition is usual…

Cited by 22PDFcodeScholar
2018

Fine-grained Video Categorization with Redundancy Reduction Attention

ECCV 2018poster

For fine-grained categorization tasks, videos could serve as a better source than static images as videos have a higher chance of containing discriminative patterns. Nevertheless, a video sequence could also contain a lot of redundant and irrelevant frames. How to locate critical information of inte…

Cited by 60SourcePDFScholar