← Search

Qilong Wang

37 accepted papers

2026

Fine-Grained Generalization via Structuralizing Concept and Feature Space into Commonality, Specificity and Confounding

AAAI 2026technical

Fine-Grained Domain Generalization (FGDG) presents greater challenges than conventional domain generalization due to the subtle inter-class differences and relatively pronounced intra-class variations inherent in fine-grained recognition tasks. Under domain shifts, the model becomes overly sensitive

Cited by 0SourcePDFScholar
2026

Test-Time Multi-Prompt Adaptation for Open-Vocabulary Remote Sensing Image Segmentation

CVPR 2026

The rise of vision-language models (VLMs) has driven the initial exploration of open-vocabulary remote sensing image semantic segmentation (OVRSIS), enabling recognition of unseen categories in complex Earth observation scenes. However, existing methods primarily focus on enhancing visual representa

Cited by 0SourcecodeScholar
2025

Asymmetric Factorized Bilinear Operation for Vision Transformer

ICLR 2025poster

As a core component of Transformer-like deep architectures, a feed-forward network (FFN) for channel mixing is responsible for learning features of each token. Recent works show channel mixing can be enhanced by increasing computational burden or can be slimmed at the sacrifice of performance. Altho…

Cited by 0SourcePDFScholar
2025

CoE: Chain-of-Explanation via Automatic Visual Concept Circuit Description and Polysemanticity Quantification

CVPR 2025poster

Explainability is a critical factor influencing the wide deployment of deep vision models (DVMs). Concept-based post-hoc explanation methods can provide both global and local insights into model decisions. However, current methods in this field face challenges in that they are inflexible to automati…

2025

DALIP: Distribution Alignment-based Language-Image Pre-Training for Domain-Specific Data

ICCV 2025poster

Recently, Contrastive Language-Image Pre-training (CLIP) has shown promising performance in domain-specific data (e.g., biology), and has attracted increasing research attention. Existing works generally focus on collecting extensive domain-specific data and directly tuning the original CLIP models.…

2025

Decoupled Multi-Predictor Optimization for Inference-Efficient Model Tuning

ICCV 2025poster

Recently, remarkable progress has been made in large-scale pre-trained model tuning, and inference efficiency is becoming more crucial for practical deployment. Early exiting in conjunction with multi-stage predictors, when cooperated with a parameter-efficient fine-tuning strategy, offers a straigh…

2025

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark

ICLR 2025poster

The dynamic imbalance of the fore-background is a major challenge in video object counting, which is usually caused by the sparsity of target objects. This remains understudied in existing works and often leads to severe under-/over-prediction errors. To tackle this issue in video object counting, w…

Cited by 1SourcePDFScholar
2025

Generative Inbetweening through Frame-wise Conditions-Driven Video Generation

CVPR 2025poster

Generative inbetweening aims to generate intermediate frame sequences by utilizing two key frames as input. Although remarkable progress has been made in video generation models, generative inbetweening still faces challenges in maintaining temporal stability due to the ambiguous interpolation path…

2025

ImagineFSL: Self-Supervised Pretraining Matters on Imagined Base Set for VLM-based Few-shot Learning

CVPR 2025highlight

Adapting CLIP models for few-shot recognition has recently attracted significant attention. Despite considerable progress, these adaptations remain hindered by the pervasive challenge of data scarcity. Text-to-image models, capable of generating abundant photorealistic labeled images, offer a promis…

Cited by 0SourcePDFScholar
2025

Integrating Visual Interpretation and Linguistic Reasoning for Geometric Problem Solving

ICCV 2025poster

Current large vision-language models (LVLMs) typically employ a connector module to link visual features with text embeddings of large language models (LLMs) and use end-to-end training to achieve multi-modal understanding in a unified process. Well alignment needs high-quality pre-training data and…

2025

RoomEditor: High-Fidelity Furniture Synthesis with Parameter-Sharing U-Net

NeurIPS 2025poster

Virtual furniture synthesis, a critical task in image composition, aims to seamlessly integrate reference objects into indoor scenes while preserving geometric coherence and visual realism. Despite its significant potential in home design applications, this field remains underexplored due to two maj…

Cited by 0SourcecodeScholar
2025

TAMT: Temporal-Aware Model Tuning for Cross-Domain Few-Shot Action Recognition

CVPR 2025poster

Going beyond few-shot action recognition (FSAR), cross-domain FSAR (CDFSAR) has attracted recent research interests by solving the domain gap lying in source-to-target transfer learning. Existing CDFSAR methods mainly focus on joint training of source and target data to mitigate the side effect of d…

2025

Unknown Text Learning for CLIP-based Few-Shot Open-set Recognition

ICCV 2025poster

Recently, vision-language models (e.g., CLIP) with prompt learning have shown great potential in few-shot learning. However, an open issue remains for the effective extension of CLIP-based models to few-shot open-set recognition (FSOR), which requires classifying known classes and detecting unknown…

2024

AMU-Tuning: Effective Logit Bias for CLIP-based Few-shot Learning

CVPR 2024poster

Recently pre-trained vision-language models (e.g. CLIP) have shown great potential in few-shot learning and attracted a lot of research interest. Although efforts have been made to improve few-shot ability of CLIP key factors on the effectiveness of existing methods have not been well studied limiti…

2023

Reliable and Interpretable Personalized Federated Learning

CVPR 2023poster

Federated learning can coordinate multiple users to participate in data training while ensuring data privacy. The collaboration of multiple agents allows for a natural connection between federated learning and collective intelligence. When there are large differences in data distribution among clien…

Cited by 27SourcePDFScholar
2023

Tuning Pre-trained Model via Moment Probing

ICCV 2023poster

Recently, efficient fine-tuning of large-scale pre-trained models has attracted increasing research interests, where linear probing (LP) as a fundamental module is involved in exploiting the final representations for task-dependent classification. However, most of the existing methods focus on how t…

Cited by 8PDFcodeScholar
2022

DropCov: A Simple yet Effective Method for Improving Deep Architectures

NeurIPS 2022accept

Previous works show global covariance pooling (GCP) has great potential to improve deep architectures especially on visual recognition tasks, where post-normalization of GCP plays a very important role in final performance. Although several post-normalization strategies have been studied, these meth…

2022

Joint Distribution Matters: Deep Brownian Distance Covariance for Few-Shot Classification

CVPR 2022oral

Few-shot classification is a challenging problem as only very few training examples are given for each new task. One of the effective research lines to address this challenge focuses on learning deep representations driven by a similarity measure between a query image and few support images of some…

Cited by 264PDFcodeScholar
2021

Boosting Weakly Supervised Object Detection via Learning Bounding Box Adjusters

ICCV 2021poster

Weakly-supervised object detection (WSOD) has emerged as an inspiring recent topic to avoid expensive instance-level object annotations. However, the bounding boxes of most existing WSOD methods are mainly determined by precomputed proposals, thereby being limited in precise object localization. In…

Cited by 60PDFcodeScholar
2021

Detection, Tracking, and Counting Meets Drones in Crowds: A Benchmark

CVPR 2021poster

To promote the developments of object detection, tracking and counting algorithms in drone-captured videos, we construct a benchmark with a new drone-captured large-scale dataset, named as DroneCrowd, formed by 112 video clips with 33,600 HD frames in various scenarios. Notably, we annotate 20,800 p…

Cited by 133PDFcodeScholar
2021

Temporal-attentive Covariance Pooling Networks for Video Recognition

NeurIPS 2021poster

For video recognition task, a global representation summarizing the whole contents of the video snippets plays an important role for the final performance. However, existing video architectures usually generate it by using a simple, global average pooling (GAP) method, which has limited ability to c…

2020

ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks

CVPR 2020poster

Recently, channel attention mechanism has demonstrated to offer great potential in improving the performance of deep convolutional neural networks (CNNs). However, most existing methods dedicate to developing more sophisticated attention modules for achieving better performance, which inevitably inc…

Cited by 7872PDFcodeScholar
2020

What Deep CNNs Benefit From Global Covariance Pooling: An Optimization Perspective

CVPR 2020poster

Recent works have demonstrated that global covariance pooling (GCP) has the ability to improve performance of deep convolutional neural networks (CNNs) on visual classification task. Despite considerable advance, the reasons on effectiveness of GCP on deep CNNs have not been well studied. In this pa…

Cited by 30PDFcodeScholar
2018

Global Gated Mixture of Second-order Pooling for Improving Deep Convolutional Neural Networks

NeurIPS 2018poster

In most of existing deep convolutional neural networks (CNNs) for classification, global average (first-order) pooling (GAP) has become a standard module to summarize activations of the last convolution layer as final representation for prediction. Recent researches show integration of higher-order…

2018

Multi-Scale Location-Aware Kernel Representation for Object Detection

CVPR 2018poster

Although Faster R-CNN and its variants have shown promising performance in object detection, they only exploit simple first order representation of object proposals for final classification and regression. Recent classification methods demonstrate that the integration of high order statistics into d…

2018

Towards Faster Training of Global Covariance Pooling Networks by Iterative Matrix Square Root Normalization

CVPR 2018poster

Global covariance pooling in convolutional neural networks has achieved impressive improvement over the classical first-order pooling. Recent works have shown matrix square root normalization plays a central role in achieving state-of-the-art performance. However, existing methods depend heavily on…

2017

G2DeNet: Global Gaussian Distribution Embedding Network and Its Application to Visual Recognition

CVPR 2017oral

Recently, plugging trainable structural layers into deep convolutional neural networks (CNNs) as image representations has made promising progress. However, there has been little work on inserting parametric probability distributions, which can effectively model feature statistics, into deep CNNs in…

Cited by 140PDFScholar
2017

Is Second-Order Information Helpful for Large-Scale Visual Recognition?

ICCV 2017poster

By stacking layers of convolution and nonlinearity, convolutional networks (ConvNets) effectively learn from low-level to high-level features and discriminative representations. Since the end goal of large-scale recognition is to delineate complex boundaries of thousands of classes, adequate explora…

Cited by 353PDFcodeScholar
2017

Mind the Class Weight Bias: Weighted Maximum Mean Discrepancy for Unsupervised Domain Adaptation

CVPR 2017poster

In domain adaptation, maximum mean discrepancy (MMD) has been widely adopted as a discrepancy metric between the distributions of source and target domains. However, existing MMD-based domain adaptation methods generally ignore the changes of class prior distributions, i.e., class weight bias across…

Cited by 777PDFcodeScholar
2016

RAID-G: Robust Estimation of Approximate Infinite Dimensional Gaussian With Application to Material Recognition

CVPR 2016poster

Infinite dimensional covariance descriptors can provide richer and more discriminative information than their low dimensional counterparts. In this paper, we propose a novel image descriptor, namely, robust approximate infinite dimensional Gaussian (RAID-G). The challenges of RAID-G mainly lie on tw…

Cited by 75PDFScholar
2015

From Dictionary of Visual Words to Subspaces: Locality-Constrained Affine Subspace Coding

CVPR 2015poster

The locality-constrained linear coding (LLC) is a very successful feature coding method in image classification. It makes known the importance of locality constraint which brings high efficiency and local smoothness of the codes. However, in the LLC method the geometry of feature space is described…

Cited by 54SourcePDFScholar