MSTDiff: Multiscale-Aware Transformer Diffusion Network for Video Object Detection
Video object detection is a fundamental yet challenging task in computer vision. Recently, DETR-based methods have gained prominence in this domain owing to their powerful global modeling capabilities. However, these methods are still confronted with two key limitations: frame-agnostic initializatio