← Search

Xi Shen

17 accepted papers

2026

A Closer Look at Cross-Domain Few-Shot Object Detection: Fine-Tuning Matters and Parallel Decoder Helps

CVPR 2026

Few-shot object detection (FSOD) is challenging due to unstable optimization and limited generalization arising from the scarcity of training samples. To address these issues, we propose a hybrid ensemble decoder that enhances generalization during fine-tuning. Inspired by ensemble learning, the dec

Cited by 1SourcecodeScholar
2026

FSOD-VFM: Few-Shot Object Detection with Vision Foundation Models and Graph Diffusion

ICLR 2026poster

In this paper, we present FSOD-VFM: Few-Shot Object Detectors with Vision Foundation Models, a framework that leverages vision foundation models to tackle the challenge of few-shot object detection. FSOD-VFM integrates three key components: a universal proposal network (UPN) for category-agnostic bo…

Cited by 0SourcecodeScholar
2026

MoVie: Broaden Your Views with Human Motion for Action Detection

CVPR 2026

Human action detection in videos requires both semantic recognition and accurate modeling of motion. While recent video foundation models have advanced visual semantics, they still struggle to capture complex and compositional actions due to the limited representation ability of motion. Human skelet

Cited by 0SourceScholar
2026

SimROD: A Simple Baseline for Raw Object Detection with Global and Local Enhancements

AAAI 2026technical

Most visual models are designed for sRGB images, yet RAW data offers significant advantages for object detection by preserving sensor information before ISP processing. This enables improved detection accuracy and more efficient hardware designs by bypassing the ISP. However, RAW object detection is

Cited by 0SourcePDFScholar
2025

DEIM: DETR with Improved Matching for Fast Convergence

CVPR 2025poster

We introduce DEIM, an innovative and efficient training framework designed to accelerate convergence in real-time object detection with Transformer-based architectures (DETR). To mitigate the sparse supervision inherent in one-to-one (O2O) matching in DETR models, DEIM employs a Dense O2O matching s…

2024

SURE: SUrvey REcipes for building reliable and robust deep networks

CVPR 2024poster

In this paper we revisit techniques for uncertainty estimation within deep neural networks and consolidate a suite of techniques to enhance their reliability. Our investigation reveals that an integrated application of diverse techniques--spanning model regularization classifier and optimization--su…

2023

Generating Human Motion From Textual Descriptions With Discrete Representations

CVPR 2023poster

In this work, we investigate a simple and must-known conditional generative framework based on Vector Quantised-Variational AutoEncoder (VQ-VAE) and Generative Pre-trained Transformer (GPT) for human motion generation from textural descriptions. We show that a simple CNN-based VQ-VAE with commonly u…

2023

LivelySpeaker: Towards Semantic-Aware Co-Speech Gesture Generation

ICCV 2023poster

Gestures are non-verbal but important behaviors accompanying people's speech. While previous methods are able to generate speech rhythm-synchronized gestures, the semantic context of the speech is generally lacking in the gesticulations. Although semantic gestures do not occur very regularly in huma…

Cited by 26PDFcodeScholar
2023

SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation

CVPR 2023poster

Generating talking head videos through a face image and a piece of speech audio still contains many challenges. i.e., unnatural head movement, distorted expression, and identity modification. We argue that these issues are mainly caused by learning from the coupled 2D motion fields. On the other han…

2022

Self-Supervised Transformers for Unsupervised Object Discovery Using Normalized Cut

CVPR 2022poster

Transformers trained with self-supervision using self-distillation loss (DINO) have been shown to produce attention maps that highlight salient foreground objects. In this paper, we show a graph-based method that uses the self-supervised transformer features to discover an object from an image. Visu…

Cited by 194PDFScholar
2021

Re-ranking for image retrieval and transductive few-shot classification

NeurIPS 2021poster

In the problems of image retrieval and few-shot classification, the mainstream approaches focus on learning a better feature representation. However, directly tackling the distance or similarity measure between images could also be efficient. To this end, we revisit the idea of re-ranking the top-k…

Cited by 51SourcePDFScholar
2020

A Differential Approach for Rain Field Tomographic Reconstruction Using Microwave Signals from Leo Satellites

ICASSP 2020accepted

A differential approach is proposed for tomographic rain field reconstruction using the estimated signal-to-noise ratio of microwave signals from low earth orbit satellites at the ground receivers, with the unknown baseline values eliminated before using least squares to reconstruct the attenuation…

Cited by 0SourceScholar
2020

Empirical Bayes Transductive Meta-Learning with Synthetic Gradients

ICLR 2020poster

We propose a meta-learning approach that learns from multiple tasks in a transductive setting, by leveraging the unlabeled query set in addition to the support set to generate a more powerful model for each task. To develop our framework, we revisit the empirical Bayes formulation for multi-task le…

Cited by 186SourceScholar
2020

Performance Analysis for Path Attenuation Estimation of Microwave Signals Due to Rainfall and Beyond

ICASSP 2020accepted

The attenuation of microwave signals can be used for meteorological observations. For example, the received signal level (RSL) of backhaul links of cellular systems, which usually has a quantization error of 0.1 dB or more for commercial systems, has been used to measure rainfall. In this work, thro…

Cited by 0SourceScholar
2019

Discovering Visual Patterns in Art Collections With Spatially-Consistent Feature Learning

CVPR 2019poster

Our goal in this paper is to discover near duplicate patterns in large collections of artworks. This is harder than standard instance mining due to differences in the artistic media (oil, pastel, drawing, etc), and imperfections inherent in the copying process. Our key technical insight is to adapt…

Cited by 118PDFScholar
2019

MARGINALIZED AVERAGE ATTENTIONAL NETWORK FOR WEAKLY-SUPERVISED LEARNING

ICLR 2019poster

In weakly-supervised temporal action localization, previous works have failed to locate dense and integral regions for each entire action due to the overestimation of the most salient regions. To alleviate this issue, we propose a marginalized average attentional network (MAAN) to suppress the domin…

Cited by 107SourcePDFScholar