← Search

Peihua Li

21 accepted papers

2026

Task-Specific Distance Correlation Matching for Few-Shot Action Recognition

AAAI 2026technical

Few-shot action recognition (FSAR) has recently made notable progress through set matching and efficient adaptation of large-scale pre-trained models. However, two key limitations persist. First, existing set matching metrics typically rely on cosine similarity to measure inter-frame linear dependen

Cited by 0SourcePDFScholar
2025

BDC-CLIP: Brownian Distance Covariance for Adapting CLIP to Action Recognition

ICML 2025poster

Bridging contrastive language-image pre-training (CLIP) to video action recognition has attracted growing interest. Human actions are inherently rich in spatial and temporal contexts, involving dynamic interactions among people, objects, and the environment. Accurately recognizing actions requires e…

Cited by 0SourcePDFScholar
2025

DALIP: Distribution Alignment-based Language-Image Pre-Training for Domain-Specific Data

ICCV 2025poster

Recently, Contrastive Language-Image Pre-training (CLIP) has shown promising performance in domain-specific data (e.g., biology), and has attracted increasing research attention. Existing works generally focus on collecting extensive domain-specific data and directly tuning the original CLIP models.…

2025

ImagineFSL: Self-Supervised Pretraining Matters on Imagined Base Set for VLM-based Few-shot Learning

CVPR 2025highlight

Adapting CLIP models for few-shot recognition has recently attracted significant attention. Despite considerable progress, these adaptations remain hindered by the pervasive challenge of data scarcity. Text-to-image models, capable of generating abundant photorealistic labeled images, offer a promis…

Cited by 0SourcePDFScholar
2025

TAMT: Temporal-Aware Model Tuning for Cross-Domain Few-Shot Action Recognition

CVPR 2025poster

Going beyond few-shot action recognition (FSAR), cross-domain FSAR (CDFSAR) has attracted recent research interests by solving the domain gap lying in source-to-target transfer learning. Existing CDFSAR methods mainly focus on joint training of source and target data to mitigate the side effect of d…

2024

Wasserstein Distance Rivals Kullback-Leibler Divergence for Knowledge Distillation

NeurIPS 2024poster

Since pioneering work of Hinton et al., knowledge distillation based on Kullback-Leibler Divergence (KL-Div) has been predominant, and recently its variants have achieved compelling performance. However, KL-Div only compares probabilities of the corresponding category between the teacher and stud…

Cited by 1SourcePDFScholar
2022

DropCov: A Simple yet Effective Method for Improving Deep Architectures

NeurIPS 2022accept

Previous works show global covariance pooling (GCP) has great potential to improve deep architectures especially on visual recognition tasks, where post-normalization of GCP plays a very important role in final performance. Although several post-normalization strategies have been studied, these meth…

2022

Joint Distribution Matters: Deep Brownian Distance Covariance for Few-Shot Classification

CVPR 2022oral

Few-shot classification is a challenging problem as only very few training examples are given for each new task. One of the effective research lines to address this challenge focuses on learning deep representations driven by a similarity measure between a query image and few support images of some…

Cited by 264PDFcodeScholar
2021

Temporal-attentive Covariance Pooling Networks for Video Recognition

NeurIPS 2021poster

For video recognition task, a global representation summarizing the whole contents of the video snippets plays an important role for the final performance. However, existing video architectures usually generate it by using a simple, global average pooling (GAP) method, which has limited ability to c…

2020

ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks

CVPR 2020poster

Recently, channel attention mechanism has demonstrated to offer great potential in improving the performance of deep convolutional neural networks (CNNs). However, most existing methods dedicate to developing more sophisticated attention modules for achieving better performance, which inevitably inc…

Cited by 7872PDFcodeScholar
2020

What Deep CNNs Benefit From Global Covariance Pooling: An Optimization Perspective

CVPR 2020poster

Recent works have demonstrated that global covariance pooling (GCP) has the ability to improve performance of deep convolutional neural networks (CNNs) on visual classification task. Despite considerable advance, the reasons on effectiveness of GCP on deep CNNs have not been well studied. In this pa…

Cited by 30PDFcodeScholar
2018

Global Gated Mixture of Second-order Pooling for Improving Deep Convolutional Neural Networks

NeurIPS 2018poster

In most of existing deep convolutional neural networks (CNNs) for classification, global average (first-order) pooling (GAP) has become a standard module to summarize activations of the last convolution layer as final representation for prediction. Recent researches show integration of higher-order…

2018

Multi-Scale Location-Aware Kernel Representation for Object Detection

CVPR 2018poster

Although Faster R-CNN and its variants have shown promising performance in object detection, they only exploit simple first order representation of object proposals for final classification and regression. Recent classification methods demonstrate that the integration of high order statistics into d…

2018

Towards Faster Training of Global Covariance Pooling Networks by Iterative Matrix Square Root Normalization

CVPR 2018poster

Global covariance pooling in convolutional neural networks has achieved impressive improvement over the classical first-order pooling. Recent works have shown matrix square root normalization plays a central role in achieving state-of-the-art performance. However, existing methods depend heavily on…

2017

G2DeNet: Global Gaussian Distribution Embedding Network and Its Application to Visual Recognition

CVPR 2017oral

Recently, plugging trainable structural layers into deep convolutional neural networks (CNNs) as image representations has made promising progress. However, there has been little work on inserting parametric probability distributions, which can effectively model feature statistics, into deep CNNs in…

Cited by 140PDFScholar
2017

Is Second-Order Information Helpful for Large-Scale Visual Recognition?

ICCV 2017poster

By stacking layers of convolution and nonlinearity, convolutional networks (ConvNets) effectively learn from low-level to high-level features and discriminative representations. Since the end goal of large-scale recognition is to delineate complex boundaries of thousands of classes, adequate explora…

Cited by 353PDFcodeScholar
2017

Mind the Class Weight Bias: Weighted Maximum Mean Discrepancy for Unsupervised Domain Adaptation

CVPR 2017poster

In domain adaptation, maximum mean discrepancy (MMD) has been widely adopted as a discrepancy metric between the distributions of source and target domains. However, existing MMD-based domain adaptation methods generally ignore the changes of class prior distributions, i.e., class weight bias across…

Cited by 777PDFcodeScholar
2016

RAID-G: Robust Estimation of Approximate Infinite Dimensional Gaussian With Application to Material Recognition

CVPR 2016poster

Infinite dimensional covariance descriptors can provide richer and more discriminative information than their low dimensional counterparts. In this paper, we propose a novel image descriptor, namely, robust approximate infinite dimensional Gaussian (RAID-G). The challenges of RAID-G mainly lie on tw…

Cited by 75PDFScholar
2015

From Dictionary of Visual Words to Subspaces: Locality-Constrained Affine Subspace Coding

CVPR 2015poster

The locality-constrained linear coding (LLC) is a very successful feature coding method in image classification. It makes known the importance of locality constraint which brings high efficiency and local smoothness of the codes. However, in the LLC method the geometry of feature space is described…

Cited by 54SourcePDFScholar