← Search

Ye Yongjie

1 accepted papers

2024

Elysium: Exploring Object-level Perception in Videos through Semantic Integration Using MLLMs

ECCV 2024poster

"Multi-modal Large Language Models (MLLMs) have demonstrated their ability to perceive objects in still images, but their application in video-related tasks, such as object tracking, remains understudied. This lack of exploration is primarily due to two key challenges. Firstly, extensive pretraining…