← Search

Sanjoy Chowdhury*

1 accepted papers

2024

Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time

ECCV 2024poster

"Leveraging Large Language Models’ remarkable proficiency in text-based tasks, recent works on Multi-modal LLMs (MLLMs) extend them to other modalities like vision and audio. However, the progress in these directions has been mostly focused on tasks that only require a coarse-grained understanding o…