← Search

Soroush Seifi

4 accepted papers

2026

Ego: Embedding-Guided Personalization of Vision-Language Models

CVPR 2026

AI assistants that support humans in daily life are becoming increasingly feasible, driven by the rapid advancements in multimodal language models. A key challenge lies in overcoming the generic nature of these models to deliver personalized experiences. Existing approaches to personalizing large vi

Cited by 0SourceScholar
2025

Recurrent Attention-based Token Selection for Efficient Streaming Video-LLMs

NeurIPS 2025poster

Video Large Language Models (Video-LLMs) excel at understanding videos in-context, assuming full access to the video when answering queries. However, these models face challenges in streaming scenarios where hour-long videos must be processed online, and questions need timely responses. In this work…

Cited by 6SourceScholar
2021

Glimpse-Attend-and-Explore: Self-Attention for Active Visual Exploration

ICCV 2021poster

Active visual exploration aims to assist an agent with a limited field of view to understand its environment based on partial observations made by choosing the best viewing directions in the scene. Recent methods have tried to address this problem either by using reinforcement learning, which is dif…

Cited by 13PDFcodeScholar