← Search

Huijie Fan

12 accepted papers

2026

GMT: Effective Global Framework for Multi-Camera Multi-Target Tracking

CVPR 2026

Existing Multi-Camera Multi-Target (MCMT) tracking models typically adopt a two-stage framework, involving single-camera tracking followed by inter-camera tracking. However, in this paradigm, the use of multiple views is confined to recovering missed matches in the first stage, providing a limited c

Cited by 0SourcecodeScholar
2026

STUR3D: Spatio-Temporal Unified Representation Learning for 3D Object Detection

CVPR 2026

Existing surrounding-view 3D object detectors initialize high-confidence queries using current 2D information, while leveraging historical 3D features as priors. However, such heavy reliance on 2D cues introduces spatio-temporal inconsistencies between 2D and 3D representations. Specifically, 2D cue

Cited by 0SourcecodeScholar
2026

Unleashing the Potential of Large Language Models for Text-to-Image Generation Through Autoregressive Representation Alignment

AAAI 2026technical

We present Autoregressive Representation Alignment (ARRA), a new training framework that unlocks global-coherent text-to-image generation in autoregressive LLMs without architectural modifications. Different from prior works that require complex architectural redesigns, ARRA aligns LLM

Cited by 0SourcePDFScholar
2025

All-Day Multi-Camera Multi-Target Tracking

CVPR 2025poster

The capability of tracking objects in low-light environments like nighttime is crucial for numerous real-world applications such as crowd behavior analysis and traffic scene understanding. However, previous Multi-Camera Multi-Target(MCMT) tracking methods are primarily focused on tracking during day…

2025

Learning Generalizable 3D Manipulation With 10 Demonstrations

IROS 2025

Learning robust and generalizable manipulation skills from few demonstrations remains a key challenge in robotics, with broad applications in industrial automation and service robotics. Although recent imitation learning methods have achieved impressive results, they often require a large amount of

Cited by 2SourcecodeScholar
2025

RIOcc: Efficient Cross-Modal Fusion Transformer with Collaborative Feature Refinement for 3D Semantic Occupancy Prediction

ICCV 2025poster

The multi-modal 3D semantic occupancy task provides a comprehensive understanding of the scene and has received considerable attention in the field of autonomous driving. However, existing methods mainly focus on processing large-scale voxels, which bring high computational costs and degrade details…

Cited by 0SourcePDFScholar
2025

Selective Aggregation for Low-Rank Adaptation in Federated Learning

ICLR 2025poster

We investigate LoRA in federated learning through the lens of the asymmetry analysis of the learned $A$ and $B$ matrices. In doing so, we uncover that $A$ matrices are responsible for learning general knowledge, while $B$ matrices focus on capturing client-specific knowledge. Based on this finding,…

2024

Exploiting Multi-Modal Synergies for Enhancing 3D Multi-Object Tracking

RA-L 2024

3D Multi-Object Tracking (MOT) aims to establish and maintain consistent object trajectories in continuously dynamic environments. At present, the tracking-by-detection has emerged as a dominant paradigm for 3D MOT, due to its simplicity and efficiency. However, this paradigm depends heavily on the

Cited by 5SourceScholar
2024

GAFusion: Adaptive Fusing LiDAR and Camera with Multiple Guidance for 3D Object Detection

CVPR 2024poster

Recent years have witnessed the remarkable progress of 3D multi-modality object detection methods based on the Bird's-Eye-View (BEV) perspective. However most of them overlook the complementary interaction and guidance between LiDAR and camera. In this work we propose a novel multi-modality 3D objec…

Cited by 7SourcePDFScholar
2024

Residual Denoising Diffusion Models

CVPR 2024poster

We propose residual denoising diffusion models (RDDM) a novel dual diffusion process that decouples the traditional single denoising diffusion process into residual diffusion and noise diffusion. This dual diffusion framework expands the denoising-based diffusion models initially uninterpretable for…

2022

Adaptive Learning Attention Network for Underwater Image Enhancement

RA-L 2022

Underwater images suffer from color casts and low illumination due to the scattering and absorption of light as it propagates in water. These problems can interfere with underwater vision tasks, such as recognition and detection. We propose an adaptive learning attention network for underwater image

Cited by 112SourcecodeScholar
2017

Deep learning of directional truncated signed distance function for robust 3D object recognition

IROS 2017poster

In this paper, we develop a novel 3D object recognition algorithm to perform detection and pose estimation jointly. We focus on analyzing the advantages of the 3D point cloud relative to the RGB-D image and try to eliminate the unpredictability of output values that inevitably occurs in regression t…

Cited by 14SourceScholar