← Search

Qiang Qi

7 accepted papers

2026

D2FANet: Enhancing Video Object Detection with Dual-Domain Feature Aggregation Network

CVPR 2026

Accurately capturing and aggregating spatiotemporal information has become crucial for video object detection. Previous methods mainly perform feature aggregation in the spatiotemporal domain, treating all regions indiscriminately and overlooking both their relative importance and the frequency char

Cited by 0SourceScholar
2026

MCI-Net: A Robust Multi-Domain Context Integration Network for Point Cloud Registration

AAAI 2026technical

Robust and discriminative feature learning is critical for high-quality point cloud registration. However, existing deep learning–based methods typically rely on Euclidean neighborhood-based strategies for feature extraction, which struggle to effectively capture the implicit semantics and structura

Cited by 0SourcePDFScholar
2026

MSTDiff: Multiscale-Aware Transformer Diffusion Network for Video Object Detection

AAAI 2026technical

Video object detection is a fundamental yet challenging task in computer vision. Recently, DETR-based methods have gained prominence in this domain owing to their powerful global modeling capabilities. However, these methods are still confronted with two key limitations: frame-agnostic initializatio

Cited by 0SourcePDFScholar
2026

SC-Net: Robust Correspondence Learning via Spatial and Cross-Channel Context

AAAI 2026technical

Recent research has focused on using convolutional neural networks (CNNs) as the backbones in two-view correspondence learning, demonstrating significant superiority over methods based on multilayer perceptrons. However, CNN backbones that are not tailored to specific tasks may fail to effectively a

Cited by 0SourcePDFScholar
2026

When Transformers Meet Mamba: A Hybrid Transformer-Mamba Network for Video Object Detection

CVPR 2026

Video object detection has gained notable progress with the advent of transformers. While transformers excel at modeling long-range contextual dependencies, the quadratic complexity limits their efficiency in long-sequence processing. In contrast, Mamba offers greater efficiency in modeling long seq

Cited by 0SourceScholar
2024

Proposal Distillation of Multi-Modal Feature Aggregation Network for Video Object Detection

ICASSP 2024accepted

Video object detection is a challenging task due to deteriorated object appearances. In order to bolster per-frame feature representations, one way is to aggregate features from relevant frames. However, relying exclusively on RGB modal for feature aggregation may limit the detection performance for…

Cited by 0SourceScholar