← Search

Jile Jiao

6 accepted papers

2026

IPFormer: Instance Prompt-guided Transformer for Multi-modal Multi-shot Video Understanding

AAAI 2026technical

Video Large Language Models (VideoLLMs), which adopt large language models for video understanding, have been demonstrated for single-shot videos. However, they usually struggle in multi-shot videos with frequent shot changes, varying camera angles, etc., which makes VideoLLMs hardly answer question

Cited by 0SourcePDFScholar
2026

Memento: Toward an All-Day Proactive Assistant for Ultra-Long Streaming Video

ICLR 2026poster

Multimodal large language models have demonstrated impressive capabilities in visual-language understanding, particularly in offline video tasks. More recently, the emergence of online video modeling has introduced early forms of active interaction. However, existing models, typically limited to ten…

Cited by 0SourceScholar
2025

Mamba-3VL: Taming State Space Model for 3D Vision Language Learning

ICCV 2025poster

3D vision-language (3D-VL) reasoning, connecting natural language with 3D physical world, represents a milestone in advancing spatial intelligence. While transformer-based methods dominate 3D-VL research, their quadratic complexity and simplistic positional embedding mechanisms severely limits effec…

2021

BV-Person: A Large-Scale Dataset for Bird-View Person Re-Identification

ICCV 2021poster

Person Re-IDentification (ReID) aims at re-identifying persons from non-overlapping cameras. Existing person ReID studies focus on horizontal-view ReID tasks, in which the person images are captured by the cameras from a (nearly) horizontal view. In this work we introduce a new ReID task, bird-view…

Cited by 24PDFScholar
2021

Occluded Person Re-Identification With Single-Scale Global Representations

ICCV 2021poster

Occluded person re-identification (ReID) aims at re-identifying occluded pedestrians from occluded or holistic images taken across multiple cameras. Current state-of-the-art (SOTA) occluded ReID models rely on some auxiliary modules, including pose estimation, feature pyramid and graph matching modu…

Cited by 64PDFScholar
2021

Person30K: A Dual-Meta Generalization Network for Person Re-Identification

CVPR 2021poster

Recently, person re-identification (ReID) has vastly benefited from the surging waves of data-driven methods. However, these methods are still not reliable enough for real-world deployments, due to the insufficient generalization capability of the models learned on existing benchmarks that have limi…

Cited by 78PDFScholar