← Search

Zehua Fu

6 accepted papers

2026

ResWorld: Temporal Residual World Model for End-to-End Autonomous Driving

ICLR 2026poster

The comprehensive understanding capabilities of world models for driving scenarios have significantly improved the planning accuracy of end-to-end autonomous driving frameworks. However, the redundant modeling of static regions and the lack of deep interaction with trajectories hinder world models f…

Cited by 0SourcecodeScholar
2026

Visual Prototype Conditioned Focal Region Generation for UAV-Based Object Detection

CVPR 2026

Unmanned aerial vehicle (UAV) based object detection is a critical but challenging task, when applied in dynamically changing scenarios with limited annotated training data. Layout-to-image generation approaches have proved effective in promoting detection accuracy by synthesizing labeled images bas

Cited by 0SourcecodeScholar
2025

GeoBEV: Learning Geometric BEV Representation for Multi-view 3D Object Detection

AAAI 2025technical

Bird's-Eye-View (BEV) representation has emerged as a mainstream paradigm for multi-view 3D object detection, demonstrating impressive perceptual capabilities. However, existing methods overlook the geometric quality of BEV representation, leaving it in a low-resolution state and failing to restore…

2022

Learning from Future: A Novel Self-Training Framework for Semantic Segmentation

NeurIPS 2022accept

Self-training has shown great potential in semi-supervised learning. Its core idea is to use the model learned on labeled data to generate pseudo-labels for unlabeled samples, and in turn teach itself. To obtain valid supervision, active attempts typically employ a momentum teacher for pseudo-label…

2022

SparseTT: Visual Tracking with Sparse Transformers

IJCAI 2022poster

Transformers have been successfully applied to the visual tracking task and significantly promote tracking performance. The self-attention mechanism designed to model long-range dependencies is the key to the success of Transformers. However, self-attention lacks focusing on the most relevant inform…

2021

STMTrack: Template-Free Visual Tracking With Space-Time Memory Networks

CVPR 2021poster

Boosting performance of the offline trained siamese trackers is getting harder nowadays since the fixed information of the template cropped from the first frame has been almost thoroughly mined, but they are poorly capable of resisting target appearance changes. Existing trackers with template updat…

Cited by 364PDFcodeScholar