← Search

Xuetao Feng

9 accepted papers

2026

IPFormer: Instance Prompt-guided Transformer for Multi-modal Multi-shot Video Understanding

AAAI 2026technical

Video Large Language Models (VideoLLMs), which adopt large language models for video understanding, have been demonstrated for single-shot videos. However, they usually struggle in multi-shot videos with frequent shot changes, varying camera angles, etc., which makes VideoLLMs hardly answer question

Cited by 0SourcePDFScholar
2026

Memento: Toward an All-Day Proactive Assistant for Ultra-Long Streaming Video

ICLR 2026poster

Multimodal large language models have demonstrated impressive capabilities in visual-language understanding, particularly in offline video tasks. More recently, the emergence of online video modeling has introduced early forms of active interaction. However, existing models, typically limited to ten…

Cited by 0SourceScholar
2025

Mamba-3VL: Taming State Space Model for 3D Vision Language Learning

ICCV 2025poster

3D vision-language (3D-VL) reasoning, connecting natural language with 3D physical world, represents a milestone in advancing spatial intelligence. While transformer-based methods dominate 3D-VL research, their quadratic complexity and simplistic positional embedding mechanisms severely limits effec…

2024

Visual-Augmented Dynamic Semantic Prototype for Generative Zero-Shot Learning

CVPR 2024poster

Generative Zero-shot learning (ZSL) learns a generator to synthesize visual samples for unseen classes which is an effective way to advance ZSL. However existing generative methods rely on the conditions of Gaussian noise and the predefined semantic prototype which limit the generator only optimized…

Cited by 19SourcePDFScholar
2022

Neural Surface Reconstruction of Dynamic Scenes with Monocular RGB-D Camera

NeurIPS 2022accept

We propose Neural-DynamicReconstruction (NDR), a template-free method to recover high-fidelity geometry and motions of a dynamic scene from a monocular RGB-D camera. In NDR, we adopt the neural implicit function for surface representation and rendering such that the captured color and depth can be f…

2021

BV-Person: A Large-Scale Dataset for Bird-View Person Re-Identification

ICCV 2021poster

Person Re-IDentification (ReID) aims at re-identifying persons from non-overlapping cameras. Existing person ReID studies focus on horizontal-view ReID tasks, in which the person images are captured by the cameras from a (nearly) horizontal view. In this work we introduce a new ReID task, bird-view…

Cited by 24PDFScholar
2021

High-Performance Discriminative Tracking With Transformers

ICCV 2021poster

End-to-end discriminative trackers improve the state of the art significantly, yet the improvement in robustness and efficiency is restricted by the conventional discriminative model, i.e., least-squares based regression. In this paper, we present DTT, a novel single-object discriminative tracker, b…

Cited by 141PDFScholar
2021

Occluded Person Re-Identification With Single-Scale Global Representations

ICCV 2021poster

Occluded person re-identification (ReID) aims at re-identifying occluded pedestrians from occluded or holistic images taken across multiple cameras. Current state-of-the-art (SOTA) occluded ReID models rely on some auxiliary modules, including pose estimation, feature pyramid and graph matching modu…

Cited by 64PDFScholar
2021

Person30K: A Dual-Meta Generalization Network for Person Re-Identification

CVPR 2021poster

Recently, person re-identification (ReID) has vastly benefited from the surging waves of data-driven methods. However, these methods are still not reliable enough for real-world deployments, due to the insufficient generalization capability of the models learned on existing benchmarks that have limi…

Cited by 78PDFScholar