← Search

Tao Tang

23 accepted papers

2026

AtomicVLA: Unlocking the Potential of Atomic Skill Learning in Robots

CVPR 2026

Recent advances in Visual-Language-Action (VLA) models have shown promising potential for robotic manipulation tasks.However, real-world robotic tasks often involve long-horizon, multi-step problem-solving and require generalization for continual skill acquisition, extending beyond single actions or

Cited by 0SourceScholar
2026

CATNet: Collaborative Alignment and Transformation Network for Cooperative Perception

CVPR 2026

Cooperative perception significantly enhances scene understanding by integrating complementary information from diverse agents. However, existing research often overlooks critical challenges inherent in real-world multi-source data integration, specifically high temporal latency and multi-source noi

Cited by 0SourceScholar
2026

CAUSAL-SAM-LLM: LARGE LANGUAGE MODELS AS CAUSAL REASONERS FOR ROBUST MEDICAL SEGMENTATION

ICASSP 2026oral

The clinical utility of deep learning models for medical image segmentation is severely constrained by their inability to generalize to unseen domains. This failure is often rooted in the models learning spurious correlations between anatomical content and domain-specific imaging styles. To overcome…

Cited by 0SourcePDFScholar
2026

CorrectAD: A Self-Correcting Agentic System to Improve End-to-end Planning in Autonomous Driving

AAAI 2026technical

End-to-end planning methods are the de-facto standard of the current autonomous driving system, while the robustness of the data-driven approaches suffers due to the notorious long-tail problem (i.e., rare but safety-critical failure cases). In this work, we explore whether recent diffusion-based vi

Cited by 0SourcePDFScholar
2026

DriveCombo: Benchmarking Compositional Traffic Rule Reasoning in Autonomous Driving

CVPR 2026

Multimodal Large Language Models (MLLMs) are rapidly becoming the intelligence brain of end-to-end autonomous driving systems. A key challenge is to assess whether MLLMs can truly understand and follow complex real-world traffic rules. However, existing benchmarks mainly focus on single-rule scenari

Cited by 0SourceScholar
2026

LiDAR-GS++: Improving LiDAR Gaussian Reconstruction via Diffusion Priors

AAAI 2026technical

Recent GS-based rendering has made significant progress for LiDAR, surpassing Neural Radiance Fields (NeRF) in both quality and speed. However, these methods exhibit artifacts in extrapolated novel view synthesis due to the incomplete reconstruction from single traversal scans. To address this limit

Cited by 0SourcePDFScholar
2026

MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation

ICLR 2026poster

Precise recognition, editing, and generation of molecules are essential prerequisites for both chemists and AI systems tackling various chemical tasks. We present MolLangBench, a comprehensive benchmark designed to evaluate fundamental molecule-language interface tasks: language-prompted molecular s…

Cited by 0SourcecodeScholar
2026

Thinking with Geometry: Active Geometry Integration for Spatial Reasoning

ICML 2026poster

Recent progress in spatial reasoning with Multimodal Large Language Models (MLLMs) increasingly leverages geometric priors from 3D encoders. However, most existing integration strategies remain passive: geometry is exposed as a global stream and fused in an indiscriminate manner, which often induces…

Cited by 0SourceScholar
2025

Automatic Annotation Augmentation Boosts Translation between Molecules and Natural Language

NAACL 2025findings

Recent advancements in AI for biological research focus on integrating molecular data with natural language to accelerate drug discovery. However, the scarcity of high-quality annotations limits progress in this area. This paper introduces LA3, a Language-based Automatic Annotation Augmentation fram…

2025

BEV-TSR: Text-Scene Retrieval in BEV Space for Autonomous Driving

AAAI 2025technical

The rapid development of the autonomous driving industry has led to a significant accumulation of autonomous driving data. Consequently, there comes a growing demand for retrieving data to provide specialized optimization. However, directly applying previous image retrieval methods faces several cha…

Cited by 2SourcePDFScholar
2025

LABEL-SAM: A Semi-Automatic Interactive Annotation Model for Aortic Dissection Segmentation in 3D CTA Image

ICASSP 2025accepted

Aortic Dissection (AD) is a life-threatening disease that can be rapidly screened by using deep learning methods. However, deep learning model training requires a large amount of manual annotation of data. To improve the annotation efficiency and accuracy, we propose LABEL-SAM, a semi-automatic inte…

Cited by 2SourceScholar
2025

S2-Track: A Simple yet Strong Approach for End-to-End 3D Multi-Object Tracking

ICML 2025poster

3D multiple object tracking (MOT) plays a crucial role in autonomous driving perception. Recent end-to-end query-based trackers simultaneously detect and track objects, which have shown promising potential for the 3D MOT task. However, existing methods are still in the early stages of development an…

Cited by 0SourcePDFScholar
2025

UniGS: Unified Language-Image-3D Pretraining with Gaussian Splatting

ICLR 2025poster

Recent advancements in multi-modal 3D pre-training methods have shown promising efficacy in learning joint representations of text, images, and point clouds. However, adopting point clouds as 3D representation fails to fully capture the intricacies of the 3D world and exhibits a noticeable gap betwe…

Cited by 0SourcePDFScholar
2024

Dual-Branch Graph Transformer Network for 3D Human Mesh Reconstruction from Video

IROS 2024poster

Human Mesh Reconstruction (HMR) from monocular video plays an important role in human-robot interaction and collaboration. However, existing video-based human mesh reconstruction methods face a trade-off between accurate reconstruction and smooth motion. These methods design networks based on either…

Cited by 0SourcecodeScholar
2024

LiT: Unifying LiDAR "Languages" with LiDAR Translator

NeurIPS 2024poster

LiDAR data exhibits significant domain gaps due to variations in sensors, vehicles, and driving environments, creating “language barriers” that limit the effective use of data across domains and the scalability of LiDAR perception models. To address these challenges, we introduce the LiDAR Translato…

2024

MLP Can Be A Good Transformer Learner

CVPR 2024poster

Self-attention mechanism is the key of the Transformer but often criticized for its computation demands. Previous token pruning works motivate their methods from the view of computation redundancy but still need to load the full network and require same memory costs. This paper introduces a novel st…

2024

Making Large Language Models Better Planners with Reasoning-Decision Alignment

ECCV 2024oral

"Data-driven approaches for autonomous driving (AD) have been widely adopted in the past decade but are confronted with dataset bias and uninterpretability. Inspired by the knowledge-driven nature of human driving, recent approaches explore the potential of large language models (LLMs) to improve un…

Cited by 12SourcePDFScholar
2024

OpenSight: A Simple Open-Vocabulary Framework for LiDAR-Based Object Detection

ECCV 2024poster

"Traditional LiDAR-based object detection research primarily focuses on closed-set scenarios, which falls short in complex real-world applications. Directly transferring existing 2D open-vocabulary models with some known LiDAR classes for open-vocabulary ability, however, tends to suffer from over-f…

2023

BEVHeight: A Robust Framework for Vision-Based Roadside 3D Object Detection

CVPR 2023poster

While most recent autonomous driving system focuses on developing perception methods on ego-vehicle sensors, people tend to overlook an alternative approach to leverage intelligent roadside cameras to extend the perception ability beyond the visual range. We discover that the state-of-the-art vision…

2022

BEVFusion: A Simple and Robust LiDAR-Camera Fusion Framework

NeurIPS 2022accept

Fusing the camera and LiDAR information has become a de-facto standard for 3D object detection tasks. Current methods rely on point clouds from the LiDAR sensor as queries to leverage the feature from the image space. However, people discovered that this underlying assumption makes the current fusio…

2021

BossNAS: Exploring Hybrid CNN-Transformers With Block-Wisely Self-Supervised Neural Architecture Search

ICCV 2021poster

A myriad of recent breakthroughs in hand-crafted neural architectures for visual recognition have highlighted the urgent need to explore hybrid architectures consisting of diversified building blocks. Meanwhile, neural architecture search methods are surging with an expectation to reduce human effor…

Cited by 142PDFcodeScholar
2020

On-chip integration of ultra-thin glass cantilever for physical property measurement activated by femtosecond laser impulse

IROS 2020poster

Under the excitation of acoustic radiation, the amount of energy absorbed and rebounded by cells have the relationship with mechanical properties, e.g. stiffness, shape, weight and so on. In this paper, a femtosecond laser-activated micro-detector is designed to convert this relationship into an ele…

Cited by 3SourceScholar