← Search

Jiannan Ge

8 accepted papers

2026

Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas

ICML 2026poster

Long-form TV dramas present a formidable challenge for comprehensive video understanding, where deciphering complex storyline often relies on **speaker recognition**, the task of accurately attributing each spoken utterance to its respective character. In this paper, we advance this field through tw…

Cited by 0SourceScholar
2026

Robustifying Vision-Language Models via Test-Time Prompt Adaptation

ICML 2026poster

Pre-trained Vision-Language Models (VLMs) such as CLIP achieve strong zero-shot generalization, but their performance degrades sharply under adversarial perturbations. Existing test-time adaptation methods typically rely on sample-level confidence heuristics, overlooking the intrinsic distributional…

Cited by 0SourceScholar
2026

SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability

AAAI 2026technical

Multimodal Large Language Models (MLLMs) have shown remarkable progress in temporal or spatial localization tasks, but struggle with joint spatio-temporal video grounding (STVG). We identify two key bottlenecks hindering this capability: (1) the sheer number of visual tokens makes long-range and fin

Cited by 0SourcePDFScholar
2025

CLIP-Adapted Region-to-Text Learning for Generative Open-Vocabulary Semantic Segmentation

ICCV 2025poster

In recent years, Open-Vocabulary Semantic Segmentation (OVSS) has been largely advanced. However, existing methods mostly rely on a pre-trained vision-language model (e.g., CLIP) and require a predefined set of classes to guide the semantic segmentation process during the inference. This not only na…

Cited by 0SourcePDFScholar
2024

Towards Balanced Alignment: Modal-Enhanced Semantic Modeling for Video Moment Retrieval

AAAI 2024technical

Video Moment Retrieval (VMR) aims to retrieve temporal segments in untrimmed videos corresponding to a given language query by constructing cross-modal alignment strategies. However, these existing strategies are often sub-optimal since they ignore the modality imbalance problem, i.e., the semantic…

2023

Progressive Spatio-Temporal Prototype Matching for Text-Video Retrieval

ICCV 2023oral

The performance of text-video retrieval has been significantly improved by vision-language cross-modal learning schemes. The typical solution is to directly align the global video-level and sentence-level features during learning, which would ignore the intrinsic video-text relations, i.e., a text…

Cited by 46PDFcodeScholar
2022

Dual-Stream Knowledge-Preserving Hashing for Unsupervised Video Retrieval

ECCV 2022poster

"Unsupervised video hashing usually optimizes binary codes by learning to reconstruct input videos. Such reconstruction constraint spends much effort on frame-level temporal context changes without focusing on video-level global semantics that are more useful for retrieval. Hence, we address this pr…

Cited by 24SourcePDFScholar
2021

Semantic-guided Reinforced Region Embedding for Generalized Zero-Shot Learning

AAAI 2021technical

Generalized zero-shot Learning (GZSL) aims to recognize images from either seen or unseen domain, mainly by learning a joint embedding space to associate image features with the corresponding category descriptions. Recent methods have proved that localizing important object regions can effectively b…

Cited by 37SourcePDFScholar