← Search

Yongchao Xu

19 accepted papers

2026

From Pairs to Sequences: Track-Aware Policy Gradients for Keypoint Detection

CVPR 2026

Keypoint-based matching is a fundamental component of modern 3D vision systems, such as Structure-from-Motion (SfM) and SLAM. Most existing learning-based methods are trained on image pairs, a paradigm that fails to explicitly optimize for the long-term trackability of keypoints across sequences und

Cited by 0SourcecodeScholar
2026

IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token Pruning

ICLR 2026poster

Large Vision-Language Models (LVLMs) achieve impressive performance across multiple tasks. A significant challenge, however, is their prohibitive inference cost when processing high-resolution visual inputs. While visual token pruning has emerged as a promising solution, existing methods that primar…

Cited by 0SourceScholar
2026

Learning to Diversify and Focus: A Reinforcement Framework for Open-Vocabulary HOI Detection

CVPR 2026

Open-Vocabulary Human-Object Interaction (OV-HOI) detection aims to recognize novel HOI categories beyond the training set. Existing OV-HOI detection approaches typically leverage CLIP to extract global visual representations and perform cross-attention between learnable queries and global features

Cited by 0SourceScholar
2026

PSP: Prompt-Guided Self-Training Sampling Policy for Active Prompt Learning

ICLR 2026poster

Active Prompt Learning (APL) using vision-language models (\textit{e.g.}, CLIP) has attracted considerable attention for mitigating the dependence on fully labeled dataset in downstream task adaptation. However, existing methods fail to explicitly leverage prompt to guide sample selection, resulting…

Cited by 0SourcecodeScholar
2026

Physics-Guided Multistep Deformation Reversal for Ancient Bamboo Slip Restoration

CVPR 2026

Bamboo slips are essential media for recording ancient East Asian civilizations, but excavated slips often suffer severe deformation due to dehydration and stress effects, creating substantial challenges for restoration. Traditional manual restoration is time-consuming and risks damage, while existi

Cited by 0SourcecodeScholar
2025

CQ-DINO: Mitigating Gradient Dilution via Category Queries for Vast Vocabulary Object Detection

NeurIPS 2025poster

With the exponential growth of data, traditional object detection methods are increasingly struggling to handle vast vocabulary object detection tasks effectively. We analyze two key limitations of classification-based detectors: positive gradient dilution, where rare positive categories receive ins…

Cited by 0SourcecodeScholar
2025

HOIMamba: Efficient Mamba-based Disentangled Progressive Learning for HOI Detection

AAAI 2025technical

Human-object interaction (HOI) detection aims to detect the spatial positions of human-object pairs and recognize their interactions. Existing single-branch, two-branch, and three-branch methods are challenging to make an appropriate trade-off on efficiency, multi-task decoupling, and collaborative…

Cited by 0SourcePDFScholar
2025

Hierarchical Knowledge Prompt Tuning for Multi-task Test-Time Adaptation

CVPR 2025poster

Test-time adaptation using vision-language models (such as CLIP) to quickly adjust to distributional shifts of downstream tasks has shown great potential. Despite significant progress, existing methods are still limited to single-task test-time adaptation scenarios and have not effectively explored…

Cited by 0SourcePDFScholar
2025

LiftFeat: 3D Geometry-Aware Local Feature Matching

ICRA 2025

Robust and efficient local feature matching plays a crucial role in applications such as SLAM and visual localization for robotics. Despite great progress, it is still very challenging to extract robust and discriminative visual features in scenarios with drastic lighting changes, low texture areas,

Cited by 9SourcecodeScholar
2025

Pathological Prior-Guided Multiple Instance Learning For Mitigating Catastrophic Forgetting in Breast Cancer Whole Slide Image Classification

ICASSP 2025accepted

In histopathology, intelligent diagnosis of Whole Slide Images (WSIs) is essential for automating and objectifying diagnoses, reducing the workload of pathologists. However, diagnostic models often face the challenge of forgetting previously learned data during incremental training on datasets from…

Cited by 0SourceScholar
2025

Pixel-wise Divide and Conquer for Federated Vessel Segmentation

IJCAI 2025

Accurate vessel segmentation is essential for diagnosing and managing vascular and ophthalmic diseases. Traditional learning-based vessel segmentation methods heavily rely on high-quality, pixel-level annotated datasets. However, segmentation performance suffers significantly when applied in federat

Cited by 0SourcePDFScholar
2023

Scratch Each Other's Back: Incomplete Multi-Modal Brain Tumor Segmentation via Category Aware Group Self-Support Learning

ICCV 2023poster

Although Magnetic Resonance Imaging (MRI) is very helpful for brain tumor segmentation and discovery, it often lacks some modalities in clinical practice. As a result, degradation of prediction performance is inevitable. According to current implementations, different modalities are considered to be…

Cited by 19PDFcodeScholar
2020

AutoSTR: Efficient Backbone Search for Scene Text Recognition

ECCV 2020poster

Scene text recognition (STR) is challenging due to the diversity of text instances and the complexity of scenes. However, no STR methods can adapt backbones to different diversities and complexities. In this work, inspired by the success of neural architecture search (NAS), we propose automated STR…

2020

Intra-class Feature Variation Distillation for Semantic Segmentation

ECCV 2020poster

Current state-of-the-art semantic segmentation methods usually require high computational resources for accurate segmentation. One promising way to achieve a good trade-off between segmentation accuracy and efficiency is knowledge distillation. In this paper, different from previous methods performi…

2020

Super-BPD: Super Boundary-to-Pixel Direction for Fast Image Segmentation

CVPR 2020poster

Image segmentation is a fundamental vision task and still remains a crucial step for many applications. In this paper, we propose a fast image segmentation method based on a novel super boundary-to-pixel direction (super-BPD) and a customized segmentation algorithm with super-BPD. Precisely, we defi…

Cited by 31PDFcodeScholar
2019

Learn to Scale: Generating Multipolar Normalized Density Maps for Crowd Counting

ICCV 2019poster

Dense crowd counting aims to predict thousands of human instances from an image, by calculating integrals of a density map over image pixels. Existing approaches mainly suffer from the extreme density variations. Such density pattern shift poses challenges even for multi-scale model ensembling. In t…

Cited by 144PDFScholar
2018

Hard-Aware Point-to-Set Deep Metric for Person Re-identification

ECCV 2018poster

Person re-identification (re-ID) is a highly challenging task due to large variations of pose, viewpoint, illumination, and occlusion. Deep metric learning provides a satisfactory solution to person re-ID by training a deep network under supervision of metric loss, e.g., triplet loss. However, the p…

Cited by 180SourcePDFScholar