← Search

Xi Qiu

10 accepted papers

2026

OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding

AAAI 2026technical

In this paper, we propose a novel framework for controllable video diffusion, OmniVDiff , aiming to synthesize and comprehend multiple video visual content in a single diffusion model. To achieve this, OmniVDiff treats all video visual modalities in the color space to learn a joint distribution, whi

Cited by 0SourcePDFScholar
2025

NFIG: Multi-Scale Autoregressive Image Generation via Frequency Ordering

NeurIPS 2025poster

Autoregressive models have achieved significant success in image generation. However, unlike the inherent hierarchical structure of image information in the spectral domain, standard autoregressive methods typically generate pixels sequentially in a fixed spatial order. To better leverage this spect…

Cited by 0SourceScholar
2024

M^2Depth: Self-supervised Two-Frame Multi-camera Metric Depth Estimation

ECCV 2024oral

"This paper presents a novel self-supervised two-frame multi-camera metric depth estimation network, termed M2 Depth, which is designed to predict reliable scale-aware surrounding depth in autonomous driving. Unlike the previous works that use multi-view images from a single time-step or multiple ti…

Cited by 3SourcePDFScholar
2023

End-to-End Vectorized HD-Map Construction With Piecewise Bezier Curve

CVPR 2023poster

Vectorized high-definition map (HD-map) construction, which focuses on the perception of centimeter-level environmental information, has attracted significant research interest in the autonomous driving community. Most existing approaches first obtain rasterized map with the segmentation-based pipel…

2023

PivotNet: Vectorized Pivot Learning for End-to-end HD Map Construction

ICCV 2023poster

Vectorized high-definition map online construction has garnered considerable attention in the field of autonomous driving research. Most existing approaches model changeable map elements using a fixed number of points, or predict local maps in a two-stage autoregressive manner, which may miss essent…

Cited by 84PDFcodeScholar
2021

Binocular Mutual Learning for Improving Few-Shot Classification

ICCV 2021poster

Most of the few-shot learning methods learn to transfer knowledge from datasets with abundant labeled data (i.e., the base set). From the perspective of class space on base set, existing methods either focus on utilizing all classes under a global view by normal pretraining, or pay more attention to…

Cited by 113PDFcodeScholar
2021

DeFRCN: Decoupled Faster R-CNN for Few-Shot Object Detection

ICCV 2021poster

Few-shot object detection, which aims at detecting novel objects rapidly from extremely few annotated examples of previously unseen classes, has attracted significant research interest in the community. Most existing approaches employ the Faster R-CNN as basic detection framework, yet, due to the la…

Cited by 347PDFcodeScholar
2021

Dynamic Metric Learning: Towards a Scalable Metric Space To Accommodate Multiple Semantic Scales

CVPR 2021poster

This paper introduces a new fundamental characteristics, i.e., the dynamic range, from real-world metric tools to deep visual recognition. In metrology, the dynamic range is a basic quality of a metric tool, indicating its flexibility to accommodate various scales. Larger dynamic range offers higher…

Cited by 20PDFcodeScholar
2020

SQE: a Self Quality Evaluation Metric for Parameters Optimization in Multi-Object Tracking

CVPR 2020poster

We present a novel self quality evaluation metric SQE for parameters optimization in the challenging yet critical multi-object tracking task. Current evaluation metrics all require annotated ground truth, thus will fail in the test environment and realistic circumstances prohibiting further optimiza…

Cited by 9PDFScholar