← Search

Shuanglin Yan

4 accepted papers

2026

Cross-modal Proxy Evolving for OOD Detection with Vision-Language Models

AAAI 2026technical

Reliable zero-shot detection of out-of-distribution (OOD) inputs is critical for deploying vision-language models in open-world settings. However, the lack of labeled negatives in zero-shot OOD detection necessitates proxy signals that remain effective under distribution shift. Existing negative-lab

Cited by 0SourcePDFScholar
2026

Hyperbolic Hierarchical Alignment for Video-Based Visible-Infrared Person Re-Identification

ICML 2026poster

Video-based visible-infrared person re-identification (VVI-ReID) aims to learn robust video-level representations under modality discrepancy. However, existing methods typically rely on Euclidean geometry, which is suboptimal for modeling the complex temporal dynamics within visible and infrared tra…

Cited by 0SourceScholar
2025

Cross-modal Collaborative Representation Learning for Text-to-Image Person Retrieval

IJCAI 2025

Text-to-image person retrieval (TIPR) aims to find images of the same identity that match a given text description. Current TIPR methods mainly focus on mining the association between images and texts, ignoring their potential complementarity. Besides, existing matching losses treat all positive pai

Cited by 0SourcePDFScholar
2025

Richer Semantics, Better Alignment: Aligning Visual Features with Explicit and Enriched Semantics for Visible-Infrared Person Re-Identification

IJCAI 2025

Visible-infrared person re-identification (VIReID) retrieves pedestrian images with the same identity across different modalities. Existing methods learn visual features solely from images, failing to align them into the modality-invariant semantic space. In this paper, we propose a novel framework,

Cited by 0SourcePDFScholar