← Search

Xuesong Chen

16 accepted papers

2026

ColaVLA: Leveraging Cognitive Latent Reasoning for Hierarchical Parallel Trajectory Planning in Autonomous Driving

CVPR 2026

Autonomous driving requires generating safe and reliable trajectories from complex multimodal inputs. Traditional modular pipelines separate perception, prediction, and planning, while recent end-to-end (E2E) systems learn them jointly. Vision-language models (VLMs) further enrich this paradigm by i

Cited by 0SourcecodeScholar
2026

sleep2vec: Unified Cross-Modal Alignment for Heterogeneous Nocturnal Biosignals

ICLR 2026poster

Tasks ranging from sleep staging to clinical diagnosis traditionally rely on standard polysomnography (PSG) devices, bedside monitors and wearable devices, which capture diverse nocturnal biosignals (e.g., EEG, EOG, ECG, SpO$_2$). However, heterogeneity across devices and frequent sensor dropout pos…

Cited by 2SourceScholar
2025

Foresight in Motion: Reinforcing Trajectory Prediction with Reward Heuristics

ICCV 2025poster

Motion forecasting for on-road traffic agents presents both a significant challenge and a critical necessity for ensuring safety in autonomous driving systems. In contrast to most existing data-driven approaches that directly predict future trajectories, we rethink this task from a planning perspect…

Cited by 0SourcePDFScholar
2025

GaussianPainter: Painting Point Cloud into 3D Gaussians with Normal Guidance

AAAI 2025technical

In this paper, we present GaussianPainter, the first method to paint a point cloud into 3D Gaussians given a reference image. GaussianPainter introduces an innovative feed-forward approach to overcome the limitations of time-consuming test-time optimization in 3D Gaussian splatting. Our method addre…

Cited by 0SourcePDFScholar
2025

M3Net: Multimodal Multi-task Learning for 3D Detection, Segmentation, and Occupancy Prediction in Autonomous Driving

AAAI 2025technical

The perception system for autonomous driving generally requires to handle multiple diverse sub-tasks. However, current algorithms typically tackle individual sub-tasks separately, which leads to low efficiency when aiming at obtaining full-perception results. Some multi-task learning methods try to…

2025

SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous Driving

CVPR 2025poster

The integration of Vision-Language Models (VLMs) into autonomous driving systems has shown promise in addressing key challenges such as learning complexity, interpretability, and common-sense reasoning. However, existing approaches often struggle with efficient integration and real-time decision-mak…

Cited by 0SourcePDFScholar
2023

TrajectoryFormer: 3D Object Tracking Transformer with Predictive Trajectory Hypotheses

ICCV 2023poster

3D multi-object tracking (MOT) is vital for many applications including autonomous driving vehicles and service robots. With the commonly used tracking-by-detection paradigm, 3D MOT has made important progress in recent years. However, these methods only use the detection boxes of the current frame…

Cited by 15PDFcodeScholar
2022

MPPNet: Multi-Frame Feature Intertwining with Proxy Points for 3D Temporal Object Detection

ECCV 2022poster

"Accurate and reliable 3D detection is vital for many applications including autonomous driving vehicles and service robots. In this paper, we present a flexible and high-performance 3D detection frame-work, named MPPNet, for 3D temporal object detection with point cloud sequences. We propose a nove…

2021

A Unified Multi-Scenario Attacking Network for Visual Object Tracking

AAAI 2021technical

Existing methods of adversarial attacks successfully generate adversarial examples to confuse Deep Neural Networks (DNNs) of image classification and object detection, resulting in wrong predictions. However, these methods are difficult to attack models of video object tracking, because the tracking…

Cited by 19SourcePDFScholar
2021

Semantic Scene Completion via Integrating Instances and Scene In-the-Loop

CVPR 2021poster

Semantic Scene Completion aims at reconstructing a complete 3D scene with precise voxel-wise semantics from a single-view depth or RGBD image. It is a crucial but challenging problem for indoor scene understanding. In this work, we present a novel framework named Scene-Instance-Scene Network (SISNet…

Cited by 84PDFcodeScholar
2020

Hijacking Tracker: A Powerful Adversarial Attack on Visual Tracking

ICASSP 2020accepted

Visual object tracking has made important breakthroughs with the assistance of deep learning models. Unfortunately, recent research has clearly proved that deep learning models are vulnerable to malicious adversarial attacks, which mislead the models making wrong decisions by perturbing the input im…

Cited by 0SourceScholar
2020

One-Shot Adversarial Attacks on Visual Tracking With Dual Attention

CVPR 2020poster

Almost all adversarial attacks in computer vision are aimed at pre-known object categories, which could be offline trained for generating perturbations. But as for visual object tracking, the tracked target categories are normally unknown in advance. However, the tracking algorithms also have potent…

Cited by 100PDFScholar
2020

Salience-Guided Cascaded Suppression Network for Person Re-Identification

CVPR 2020poster

Employing attention mechanisms to model both global and local features as a final pedestrian representation has become a trend for person re-identification (Re-ID) algorithms. A potential limitation of these methods is that they focus on the most salient features, but the re-identification of a pers…

Cited by 307PDFScholar
2019

Discriminative Features Reconstruction Network for Semantic Segmentation

ICASSP 2019accepted

Thanks to the development of convolutional neural networks (CNNs), researchers have proposed lots of effective semantic segmentation models. However, there are still two problems disturbing researchers, one of which is objects misidentification on the image level and another one is poor performance…

Cited by 0SourceScholar
2019

Multi-attention Network for Thoracic Disease Classification and Localization

ICASSP 2019accepted

The chest X-ray is one of the most commonly available radiological examinations for diagnosing lung diseases. This task remains a major challenge due to 1) the shortage of accurate annotations for chest X-ray examinations, 2) the diversity of lesion areas on X-rays from different thoracic disease an…

Cited by 0SourceScholar
2019

Scanet: Spatial-channel Attention Network for 3D Object Detection

ICASSP 2019accepted

This paper aims to achieve high-accuracy 3D object detection, in which we propose a novel Spatial-Channel Attention Network (SCANet), a two-stage detector that takes both LIDAR point clouds and RGB images as input to generate 3D object estimates. The first stage is a 3D region proposal network (RPN)…

Cited by 0SourceScholar