← Search

Xiao Zhao

10 accepted papers

2026

CoT-VLNBench: A Benchmark for Visual Chain-of-Thought Reasoning in Vision-Language-Navigation Robots

AAAI 2026technical

Recent advances in vision language models (VLMs) have demonstrated remarkable potential in embodied navigation tasks. However, existing robot-centric datasets primarily focus on traditional 3D tasks such as perception and prediction, lacking adequate support for vision-language tasks. Vision-languag

Cited by 0SourcePDFScholar
2026

GaussianFusion: Unified 3D Gaussian Representation for Multi-Modal Fusion Perception

ICLR 2026poster

The bird’s-eye view (BEV) representation enables multi-sensor features to be fused within a unified space, serving as the primary approach for achieving comprehensive multi-task perception. However, the discrete grid representation of BEV leads to significant detail loss and limits feature alignment…

Cited by 0SourceScholar
2025

BloomScene: Lightweight Structured 3D Gaussian Splatting for Crossmodal Scene Generation

AAAI 2025technical

With the widespread use of virtual reality applications, 3D scene generation has become a new challenging research frontier. 3D scenes have highly complex structures and need to ensure that the output is dense, coherent, and contains all necessary structures. Many current 3D scene generation methods…

2025

TAD-E2E: A Large-scale End-to-end Autonomous Driving Dataset

ICCV 2025poster

End-to-end autonomous driving technology has recently become a focal point of research and application in autonomous driving. State-of-the-art (SOTA) methods are often trained and evaluated on the NuScenes dataset. However, the NuScenes dataset, introduced in 2019 for 3D perception tasks, faces seve…

Cited by 0SourcePDFScholar
2025

V-Fusion: 2D Detection-enhanced Multimodal 3D BEV Object Detection

ICASSP 2025accepted

Integrating information from multiple sensors enhances the performance of autonomous vehicle perception systems. However, current multimodal 3D object detection methods focus on unifying modalities into a bird’s-eye view (BEV) representation, which overlooks the inherent characteristics of camera pe…

Cited by 0SourceScholar
2024

CPR-Coach: Recognizing Composite Error Actions based on Single-class Training

CVPR 2024poster

Fine-grained medical action analysis plays a vital role in improving medical skill training efficiency but it faces the problems of data and algorithm shortage. Cardiopulmonary Resuscitation (CPR) is an essential skill in emergency treatment. Currently the assessment of CPR skills mainly depends on…

2024

Correlation-Decoupled Knowledge Distillation for Multimodal Sentiment Analysis with Incomplete Modalities

CVPR 2024poster

Multimodal sentiment analysis (MSA) aims to understand human sentiment through multimodal data. Most MSA efforts are based on the assumption of modality completeness. However in real-world applications some practical factors cause uncertain modality missingness which drastically degrades the model's…

Cited by 15SourcePDFScholar
2024

HybridOcc: NeRF Enhanced Transformer-Based Multi-Camera 3D Occupancy Prediction

RA-L 2024

Vision-based 3D semantic scene completion (SSC) describes autonomous driving scenes through 3D volume representations. However, the occlusion of invisible voxels by scene surfaces poses challenges to current SSC methods in hallucinating refined 3D geometry. This letter proposes HybridOcc, a hybrid 3

Cited by 17SourceScholar
2023

Context De-Confounded Emotion Recognition

CVPR 2023poster

Context-Aware Emotion Recognition (CAER) is a crucial and challenging task that aims to perceive the emotional states of the target person with contextual information. Recent approaches invariably focus on designing sophisticated architectures or mechanisms to extract seemingly meaningful representa…

2023

D-CONFORMER: Deformable Sparse Transformer Augmented Convolution for Voxel-Based 3D Object Detection

ICASSP 2023accepted

Although CNN-based and Transformer-based detectors have made impressive improvements in 3D object detection, these two network paradigms suffer from the interference of insufficient receptive field and local detail weakening, which significantly limits the feature extraction performance of the backb…

Cited by 0SourceScholar