← Search

Baojie Fan

20 accepted papers

2026

DuRP: Dual-Stage Physics-Embedded Learning for Joint Radiance and Polarization Restoration

ICML 2026poster

Polarization information is valuable for many computer vision applications. However, in hazy environments, polarization information is severely attenuated due to the degradation of captured polarized images. Existing dehazing methods struggle to effectively restore polarization information, as singl…

Cited by 0SourceScholar
2026

GMT: Effective Global Framework for Multi-Camera Multi-Target Tracking

CVPR 2026

Existing Multi-Camera Multi-Target (MCMT) tracking models typically adopt a two-stage framework, involving single-camera tracking followed by inter-camera tracking. However, in this paradigm, the use of multiple views is confined to recovering missed matches in the first stage, providing a limited c

Cited by 0SourcecodeScholar
2026

Point-Voxel Guidance Fusion With Unidirectional Interaction for 3D Single Object Tracking

RA-L 2026

Sparse point-based trackers struggle with textureless and incomplete point clouds. Conversely, dense voxel-based trackers have richer spatial and semantic information, but how to filter out interference from complex backgrounds remains a challenge. Besides, there is still a gap between point and vox

Cited by 0SourceScholar
2026

STUR3D: Spatio-Temporal Unified Representation Learning for 3D Object Detection

CVPR 2026

Existing surrounding-view 3D object detectors initialize high-confidence queries using current 2D information, while leveraging historical 3D features as priors. However, such heavy reliance on 2D cues introduces spatio-temporal inconsistencies between 2D and 3D representations. Specifically, 2D cue

Cited by 0SourcecodeScholar
2025

All-Day Multi-Camera Multi-Target Tracking

CVPR 2025poster

The capability of tracking objects in low-light environments like nighttime is crucial for numerous real-world applications such as crowd behavior analysis and traffic scene understanding. However, previous Multi-Camera Multi-Target(MCMT) tracking methods are primarily focused on tracking during day…

2025

MAFF-Net: Enhancing 3D Object Detection With 4D Radar via Multi-Assist Feature Fusion

RA-L 2025

Perception systems are crucial for the safe operation of autonomous vehicles, particularly for 3D object detection. While LiDAR-based methods are limited by adverse weather conditions, 4D radars offer promising all-weather capabilities. However, 4D radars introduce challenges such as extreme sparsit

Cited by 9SourceScholar
2025

RIOcc: Efficient Cross-Modal Fusion Transformer with Collaborative Feature Refinement for 3D Semantic Occupancy Prediction

ICCV 2025poster

The multi-modal 3D semantic occupancy task provides a comprehensive understanding of the scene and has received considerable attention in the field of autonomous driving. However, existing methods mainly focus on processing large-scale voxels, which bring high computational costs and degrade details…

Cited by 0SourcePDFScholar
2025

Unidirectional Point-Voxel Fusion for Enhanced 3D Single Object Tracking

IROS 2025

Sparse point-based trackers struggle with texture-less and incomplete point clouds. Conversely, dense voxel-based trackers have richer spatial and semantic information, but filtering out interference from complex backgrounds remains a challenge. Additionally, there is still a gap between point and v

Cited by 0SourceScholar
2024

Enhancing 3D Single Object Tracking with Efficient Point Cloud Segmentation

IROS 2024poster

3D single object tracking (SOT) based on point cloud has attracted much attention due to its important role in machine vision and autonomous driving. Recently, M2-Track proposes a two-stage tracking structure centered on motion, but they ignore the effect of segmentation errors in sparse point cloud…

Cited by 0SourceScholar
2024

GAFusion: Adaptive Fusing LiDAR and Camera with Multiple Guidance for 3D Object Detection

CVPR 2024poster

Recent years have witnessed the remarkable progress of 3D multi-modality object detection methods based on the Bird's-Eye-View (BEV) perspective. However most of them overlook the complementary interaction and guidance between LiDAR and camera. In this work we propose a novel multi-modality 3D objec…

Cited by 7SourcePDFScholar
2024

Integrating Scaling Strategy and Central Guided Voting for 3D Point Cloud Object Tracking

RA-L 2024

LiDAR-based 3D single object tracking has received remarkable attention due to its crucial role in robotics and autonomous driving. Most of them are based on hierarchical feature structures from PointNet++. However, existing based-stratified structure trackers ignore the fact that non-linearities in

Cited by 4SourceScholar
2024

Unbiased Faster R-CNN for Single-source Domain Generalized Object Detection

CVPR 2024highlight

Single-source domain generalization (SDG) for object detection is a challenging yet essential task as the distribution bias of the unseen domain degrades the algorithm performance significantly. However existing methods attempt to extract domain-invariant features neglecting that the biased data lea…

Cited by 9SourcePDFScholar