← Search

Kaixin Peng

1 accepted papers

2026

OmniPT: Unleashing the Potential of Large Vision Language Models for Pedestrian Tracking and Understanding

AAAI 2026technical

LVLMs have been shown to perform excellently in image-level tasks such as VQA and caption. However, in many instance-level tasks, such as visual grounding and object detection, LVLMs still show performance gaps compared to previous expert models. Meanwhile, although pedestrian tracking is a classica

Cited by 0SourcePDFScholar