← Search

Raffay Hamid

11 accepted papers

2025

CoLLM: A Large Language Model for Composed Image Retrieval

CVPR 2025poster

Composed Image Retrieval (CIR) is a complex task that aims to retrieve images based on a multimodal query. Typical training data consists of triplets containing a reference image, a textual description of desired modifications, and the target image, which are expensive and time-consuming to acquire.…

2025

M-LLM Based Video Frame Selection for Efficient Video Understanding

CVPR 2025poster

Recent advances in Multi-Modal Large Language Models (M-LLMs) show promising results in video reasoning. Popular Multi-Modal Large Language Model (M-LLM) frameworks usually apply naive uniform sampling to reduce the number of video frames that are fed into an M-LLM, particularly for long context vid…

Cited by 3SourcePDFScholar
2023

LEMaRT: Label-Efficient Masked Region Transform for Image Harmonization

CVPR 2023poster

We present a simple yet effective self-supervised pretraining method for image harmonization which can leverage large-scale unannotated image datasets. To achieve this goal, we first generate pre-training data online with our Label-Efficient Masked Region Transform (LEMaRT) pipeline. Given an image,…

Cited by 23SourcePDFScholar
2023

Movies2Scenes: Using Movie Metadata To Learn Scene Representation

CVPR 2023poster

Understanding scenes in movies is crucial for a variety of applications such as video moderation, search, and recommendation. However, labeling individual scenes is a time-consuming process. In contrast, movie level metadata (e.g., genre, synopsis, etc.) regularly gets produced as part of the film p…

Cited by 19SourcePDFScholar
2023

Selective Structured State-Spaces for Long-Form Video Understanding

CVPR 2023poster

Effective modeling of complex spatiotemporal dependencies in long-form videos remains an open problem. The recently proposed Structured State-Space Sequence (S4) model with its linear complexity offers a promising direction in this space. However, we demonstrate that treating all image-tokens equall…

Cited by 126SourcePDFScholar
2021

Shot Contrastive Self-Supervised Learning for Scene Boundary Detection

CVPR 2021poster

Scenes play a crucial role in breaking the storyline of movies and TV episodes into semantically cohesive parts. However, given their complex temporal structure, finding scene boundaries can be a challenging task requiring large amounts of labeled training data. To address this challenge, we present…

Cited by 87PDFScholar