← Search

Ashish Shah

6 accepted papers

2024

MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

CVPR 2024poster

With the success of large language models (LLMs) integrating the vision model into LLMs to build vision-language foundation models has gained much more interest recently. However existing LLM-based large multimodal models (e.g. Video-LLaMA VideoChat) can only take in a limited number of frames for s…

2024

Spherical Linear Interpolation and Text-Anchoring for Zero-shot Composed Image Retrieval

ECCV 2024poster

"Composed Image Retrieval (CIR) is a complex task that retrieves images using a query, which is configured with an image and a caption that describes desired modifications to that image. Supervised CIR approaches have shown strong performance, but their reliance on expensive manually-annotated datas…

2023

Open Vocabulary Semantic Segmentation With Patch Aligned Contrastive Learning

CVPR 2023highlight

We introduce Patch Aligned Contrastive Learning (PACL), a modified compatibility function for CLIP's contrastive loss, intending to train an alignment between the patch tokens of the vision encoder and the CLS token of the text encoder. With such an alignment, a model can identify regions of an imag…

2022

Few-Shot Fast-Adaptive Anomaly Detection

NeurIPS 2022accept

The ability to detect anomaly has long been recognized as an inherent human ability, yet to date, practical AI solutions to mimic such capability have been lacking. This lack of progress can be attributed to several factors. To begin with, the distribution of ``abnormalities'' is intractable. Anythi…

Cited by 30SourcePDFScholar
2022

Object-Centric Unsupervised Image Captioning

ECCV 2022poster

"Image captioning is a longstanding problem in the field of computer vision and natural language processing. To date, researchers have produced impressive state-of-the-art performance in the age of deep learning. Most of these state-of-the-art, however, requires large volume of annotated image-capti…