← Search

Guangxing Han

17 accepted papers

2026

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

CVPR 2026

Recent progress in vision-language pretraining has enabled significant improvements to many downstream computer vision applications, such as classification, retrieval, segmentation and depth prediction. However, a fundamental capability that these models still struggle with is aligning dense patch r

Cited by 0SourcecodeScholar
2025

TIPS: Text-Image Pretraining with Spatial awareness

ICLR 2025poster

While image-text representation learning has become very popular in recent years, existing models tend to lack spatial awareness and have limited direct applicability for dense understanding tasks. For this reason, self-supervised image-only pretraining is still the go-to method for many dense visio…

2024

Jack of All Tasks Master of Many: Designing General-Purpose Coarse-to-Fine Vision-Language Model

CVPR 2024highlight

The ability of large language models (LLMs) to process visual inputs has given rise to general-purpose vision systems unifying various vision-language (VL) tasks by instruction tuning. However due to the enormous diversity in input-output formats in the vision domain existing general-purpose models…

2023

DiGeo: Discriminative Geometry-Aware Learning for Generalized Few-Shot Object Detection

CVPR 2023poster

Generalized few-shot object detection aims to achieve precise detection on both base classes with abundant annotations and novel classes with limited training data. Existing approaches enhance few-shot generalization with the sacrifice of base-class performance, or maintain high precision in base-cl…

2023

Supervised Masked Knowledge Distillation for Few-Shot Transformers

CVPR 2023poster

Vision Transformers (ViTs) emerge to achieve impressive performance on many data-abundant computer vision tasks by capturing long-range dependencies among local features. However, under few-shot learning (FSL) settings on small datasets with only a few labeled data, ViT tends to overfit and suffers…

2023

TempCLR: Temporal Alignment Representation with Contrastive Learning

ICLR 2023poster

Video representation learning has been successful in video-text pre-training for zero-shot transfer, where each sentence is trained to be close to the paired video clips in a common feature space. For long videos, given a paragraph of description where the sentences describe different segments of th…

2022

Few-Shot End-to-End Object Detection via Constantly Concentrated Encoding across Heads

ECCV 2022poster

"Few-shot object detection (FSOD) aims to detect objects of new classes and learn effective models without exhaustive annotation. The end-to-end detection framework has been proposed to generate sparse proposals and set a stack of detection heads to improve the performance. For each proposal, the pr…

Cited by 20SourcePDFScholar
2022

Few-Shot Object Detection With Fully Cross-Transformer

CVPR 2022oral

Few-shot object detection (FSOD), with the aim to detect novel objects using very few training examples, has recently attracted great research interest in the community. Metric-learning based methods have been demonstrated to be effective for this task using a two-branch based siamese network, and c…

Cited by 193PDFcodeScholar
2022

Meta Faster R-CNN: Towards Accurate Few-Shot Object Detection with Attentive Feature Alignment

AAAI 2022technical

Few-shot object detection (FSOD) aims to detect objects using only a few examples. How to adapt state-of-the-art object detectors to the few-shot domain remains challenging. Object proposal is a key ingredient in modern object detectors. However, the quality of proposals generated for few-shot class…

2022

Task-Adaptive Negative Envision for Few-Shot Open-Set Recognition

CVPR 2022poster

We study the problem of few-shot open-set recognition (FSOR), which learns a recognition system capable of both fast adaptation to new classes with limited labeled examples and rejection of unknown negative samples. Traditional large-scale open-set methods have been shown ineffective for FSOR proble…

Cited by 50PDFcodeScholar
2022

Weakly-Supervised Temporal Article Grounding

EMNLP 2022main

Given a long untrimmed video and natural language queries, video grounding (VG) aims to temporally localize the semantically-aligned video segments. Almost all existing VG work holds two simple but unrealistic assumptions: 1) All query sentences can be grounded in the corresponding video. 2) All que…

2021

COVID-19 Literature Knowledge Graph Construction and Drug Repurposing Report Generation

NAACL 2021system demonstrations

To combat COVID-19, both clinicians and scientists need to digest the vast amount of relevant biomedical knowledge in literature to understand the disease mechanism and the related biological functions. We have developed a novel and comprehensive knowledge discovery framework, COVID-KG to extract fi…

2021

Partner-Assisted Learning for Few-Shot Image Classification

ICCV 2021poster

Few-shot Learning has been studied to mimic human visual capabilities and learn effective models without the need of exhaustive human annotation. Even though the idea of meta-learning for adaptation has dominated the few-shot learning methods, how to train a feature extractor is still a challenge. I…

Cited by 94PDFScholar
2021

Query Adaptive Few-Shot Object Detection With Heterogeneous Graph Convolutional Networks

ICCV 2021poster

Few-shot object detection (FSOD) aims to detect never-seen objects using few examples. This field sees recent improvement owing to the meta-learning techniques by learning how to match between the query image and few-shot class examples, such that the learned model can generalize to few-shot novel c…

Cited by 149PDFcodeScholar
2021

The Met Dataset: Instance-level Recognition for Artworks

NeurIPS 2021poster

This work introduces a dataset for large-scale instance-level recognition in the domain of artworks. The proposed benchmark exhibits a number of different challenges such as large inter-class similarity, long tail distribution, and many classes. We rely on the open access collection of The Met museu…

Cited by 47SourceScholar