← Search

Xu Cao

30 accepted papers

2026

HOLO: Homography-Guided Pose Estimator Network for Fine-Grained Visual Localization on SD Maps

CVPR 2026

Visual localization on standard-definition (SD) maps has emerged as a promising low-cost and scalable solution for autonomous driving. However, existing regression-based approaches often overlook inherent geometric priors, resulting in suboptimal training efficiency and limited localization accuracy

Cited by 0SourcecodeScholar
2026

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation

CVPR 2026

Recent advancements in vision-language-action (VLA) models have shown promise in robotic manipulation, yet they continue to struggle with long-horizon, multi-step tasks. Existing methods lack internal reasoning mechanisms that can identify task-relevant interaction cues or track progress within a su

Cited by 0SourceScholar
2026

Robust LLM Unlearning via Post Judgment and Multi-round Thinking

ICLR 2026poster

The unlearning capability of LLMs is vital for ensuring compliance and safety, especially when removing sensitive knowledge from deployed models. Pre-filtering methods, enabling rapid deployment without parameter changes, are a prominent unlearning approach. However, they exhibit significant robustn…

Cited by 0SourcecodeScholar
2025

EAROL: Environmental Augmented Perception-Aware Planning and Robust Odometry via Downward-Mounted Tilted LiDAR

IROS 2025

To address the challenges of localization drift and perception-planning coupling in unmanned aerial vehicles (UAVs) operating in open-top scenarios (e.g., collapsed buildings, roofless mazes), this paper proposes EAROL, a novel framework with a downward-mounted tilted LiDAR configuration (20° inclin

Cited by 0SourceScholar
2025

Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs

NeurIPS 2025poster

Current Vision-Language Models (VLMs) struggle with fine-grained spatial reasoning, particularly when multi-step logic and precise spatial alignment are required. In this work, we introduce SpatialReasoner-R1, a vision-language reasoning model designed to address these limitations. To construct high…

Cited by 0SourceScholar
2025

IMVPR: Implicit BEV-Enhanced Multi-View Aggregation for Visual Place Recognition

IROS 2025

Visual Place Recognition (VPR) is essential for robotics and autonomous driving, enabling localization by matching current observations with a database of known places. While monocular VPR methods rely on visual features, they are sensitive to environmental changes, and multimodal approaches using L

Cited by 0SourceScholar
2025

Neural Multi-View Self-Calibrated Photometric Stereo without Photometric Stereo Cues

ICCV 2025poster

We propose a neural inverse rendering approach that jointly reconstructs geometry, spatially varying reflectance, and lighting conditions from multi-view images captured under varying directional lighting. Unlike prior multi-view photometric stereo methods that require light calibration or intermedi…

2025

On-Board Vision-Language Models (VLMs) for Personalized Motion Control of Autonomous Vehicles

IROS 2025

Personalized driving refers to an autonomous vehicle’s ability to adapt its driving behavior or control strategies to match individual users’ preferences and driving styles while maintaining safety and comfort standards. However, existing works either fail to capture every individual’s preference pr

Cited by 1SourceScholar
2025

PMNI: Pose-free Multi-view Normal Integration for Reflective and Textureless Surface Reconstruction

CVPR 2025poster

Reflective and textureless surfaces remain a challenge in multi-view 3D reconstruction. Both camera pose calibration and shape reconstruction often fail due to insufficient or unreliable cross-view visual features. To address these issues, we present PMNI (Pose-free Multi-view Normal Integration), a…

2025

Planning and Control for Active Morphing Tensegrity Aerial Vehicles in Confined Spaces

IROS 2025

Morphing quadrotors are capable of adapting to constrained environments through geometric reconfiguration. However, existing systems are limited by mechanical complexity and rigid links, which affect both safety and performance in such environments. In this paper, we propose a strut-actuated tensegr

Cited by 0SourceScholar
2025

Proxy-Bridged Game Transformer for Interactive Extreme Motion Prediction

ICCV 2025poster

Multi-person motion prediction becomes particularly challenging when handling highly interactive scenarios involving extreme motions. Previous works focused more on the case of `moderate' motions (e.g., walking together), where predicting each pose in isolation often yields reasonable results. Howev…

2025

SocialGesture: Delving into Multi-person Gesture Understanding

CVPR 2025poster

Previous research in human gesture recognition has largely overlooked multi-person interactions, which are crucial for understanding the social context of naturally occurring gestures. This limitation in existing datasets presents a significant challenge in aligning human gestures with other modalit…

Cited by 0SourcePDFScholar
2025

Toward Human Deictic Gesture Target Estimation

NeurIPS 2025poster

Humans have a remarkable ability to use co-speech deictic gestures, such as pointing and showing, to enrich verbal communication and support social interaction. These gestures are so fundamental that infants begin to use them even before they acquire spoken language, which highlights their central r…

Cited by 0SourcecodeScholar
2024

LaMPilot: An Open Benchmark Dataset for Autonomous Driving with Language Model Programs

CVPR 2024poster

Autonomous driving (AD) has made significant strides in recent years. However existing frameworks struggle to interpret and execute spontaneous user instructions such as "overtake the car ahead." Large Language Models (LLMs) have demonstrated impressive reasoning capabilities showing potential to br…

2024

Learning Autonomous Driving Tasks via Human Feedbacks with Large Language Models

EMNLP 2024finding

Traditional autonomous driving systems have mainly focused on making driving decisions without human interaction, overlooking human-like decision-making and human preference required in complex traffic scenarios. To bridge this gap, we introduce a novel framework leveraging Large Language Models (LL…

Cited by 2SourcePDFScholar
2024

MAPLM: A Real-World Large-Scale Vision-Language Benchmark for Map and Traffic Scene Understanding

CVPR 2024poster

Vision-language generative AI has demonstrated remarkable promise for empowering cross-modal scene understanding of autonomous driving and high-definition (HD) map systems. However current benchmark datasets lack multi-modal point cloud image and language data pairs. Recent approaches utilize visual…

2023

ECON: Explicit Clothed Humans Optimized via Normal Integration

CVPR 2023highlight

The combination of deep learning, artist-curated scans, and Implicit Functions (IF), is enabling the creation of detailed, clothed, 3D humans from images. However, existing methods are far from perfect. IF-based methods recover free-form geometry, but produce disembodied limbs or degenerate shapes f…

2023

Mitigating Transformer Overconfidence via Lipschitz Regularization

UAI 2023poster

Though Transformers have achieved promising results in many computer vision tasks, they tend to be over-confident in predictions, as the standard Dot Product Self-Attention (DPSA) can barely preserve distance for the unbounded input domain. In this work, we fill this gap by proposing a novel Lipschi…

2023

Multi-View Azimuth Stereo via Tangent Space Consistency

CVPR 2023poster

We present a method for 3D reconstruction only using calibrated multi-view surface azimuth maps. Our method, multi-view azimuth stereo, is effective for textureless or specular surfaces, which are difficult for conventional multi-view stereo methods. We introduce the concept of tangent space consist…

2023

THMA: Tencent HD Map AI System for Creating HD Map Annotations

AAAI 2023technical

Nowadays, autonomous vehicle technology is becoming more and more mature. Critical to progress and safety, high-definition (HD) maps, a type of centimeter-level map collected using a laser sensor, provide accurate descriptions of the surrounding environment. The key challenge of HD map production is…

Cited by 13SourcePDFScholar
2023

Vitasd: Robust Vision Transformer Baselines for Autism Spectrum Disorder Facial Diagnosis

ICASSP 2023accepted

Autism spectrum disorder (ASD) is a lifelong neurodevelopmental disorder with very high prevalence around the world. Research progress in the field of ASD facial analysis in pediatric patients has been hindered due to a lack of well-established baselines. In this paper, we propose the use of the Vis…

Cited by 0SourceScholar
2022

AggPose: Deep Aggregation Vision Transformer for Infant Pose Estimation

IJCAI 2022poster

Movement and pose assessment of newborns lets experienced pediatricians predict neurodevelopmental disorders, allowing early intervention for related diseases. However, most of the newest AI approaches for human pose estimation methods focus on adults, lacking publicly benchmark for infant pose esti…

2021

Normal Integration via Inverse Plane Fitting With Minimum Point-to-Plane Distance

CVPR 2021poster

This paper presents a surface normal integration method that solves an inverse problem of local plane fitting. Surface reconstruction from normal maps is essential in photometric shape reconstruction. To this end, we formulate normal integration in the camera coordinates and jointly solve for 3D poi…

Cited by 23PDFcodeScholar
2021

Semi-supervised Vein Segmentation of Ultrasound Images for Autonomous Venipuncture

IROS 2021poster

Venipuncture is an indispensable procedure for both diagnosis and treatment. In this paper, unlike existing solutions that fully or partially rely on professional assistance, a compact robotic system integrating both novel hardware and software developments is introduced. The hardware consists of a…

Cited by 7SourceScholar
2020

Balanced Joint Adversarial Training for Robust Intent Detection and Slot Filling

COLING 2020main

Joint intent detection and slot filling has recently achieved tremendous success in advancing the performance of utterance understanding. However, many joint models still suffer from the robustness problem, especially on noisy inputs or rare/unseen events. To address this issue, we propose a Joint A…

Cited by 10SourcePDFScholar
2020

Stereoscopic Flash and No-Flash Photography for Shape and Albedo Recovery

CVPR 2020poster

We present a minimal imaging setup that harnesses both geometric and photometric approaches for shape and albedo recovery. We adopt a stereo camera and a flashlight to capture a stereo image pair and a flash/no-flash pair. From the stereo image pair, we recover a rough shape that captures low-freque…

Cited by 12PDFScholar