← Search

Wenqi Shang

2 accepted papers

2026

D2FANet: Enhancing Video Object Detection with Dual-Domain Feature Aggregation Network

CVPR 2026

Accurately capturing and aggregating spatiotemporal information has become crucial for video object detection. Previous methods mainly perform feature aggregation in the spatiotemporal domain, treating all regions indiscriminately and overlooking both their relative importance and the frequency char

Cited by 0SourceScholar
2026

MSTDiff: Multiscale-Aware Transformer Diffusion Network for Video Object Detection

AAAI 2026technical

Video object detection is a fundamental yet challenging task in computer vision. Recently, DETR-based methods have gained prominence in this domain owing to their powerful global modeling capabilities. However, these methods are still confronted with two key limitations: frame-agnostic initializatio

Cited by 0SourcePDFScholar