← Search

Tanzila Rahman

5 accepted papers

2026

To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models

ICLR 2026poster

Large Vision Language Models (LVLMs) have recently emerged as powerful architectures capable of understanding and reasoning over both visual and textual information. These models typically rely on two key components: a Vision Transformer (ViT) and a Large Language Model (LLM). ViT encodes visual con…

Cited by 0SourceScholar
2024

Prompting Hard or Hardly Prompting: Prompt Inversion for Text-to-Image Diffusion Models

CVPR 2024poster

The quality of the prompts provided to text-to-image diffusion models determines how faithful the generated content is to the user's intent often requiring `prompt engineering'. To harness visual concepts from target images without prompt engineering current approaches largely rely on embedding inve…

Cited by 16SourcePDFScholar
2023

Make-a-Story: Visual Memory Conditioned Consistent Story Generation

CVPR 2023poster

There has been a recent explosion of impressive generative models that can produce high quality images (or videos) conditioned on text descriptions. However, all such approaches rely on conditional sentences that contain unambiguous descriptions of scenes and main actors in them. Therefore employing…

2021

TriBERT: Human-centric Audio-visual Representation Learning

NeurIPS 2021poster

The recent success of transformer models in language, such as BERT, has motivated the use of such architectures for multi-modal feature learning and tasks. However, most multi-modal variants (e.g., ViLBERT) have limited themselves to visual-linguistic data. Relatively few have explored its use in au…

2019

Watch, Listen and Tell: Multi-Modal Weakly Supervised Dense Event Captioning

ICCV 2019poster

Multi-modal learning, particularly among imaging and linguistic modalities, has made amazing strides in many high-level fundamental visual understanding problems, ranging from language grounding to dense event captioning. However, much of the research has been limited to approaches that either do no…

Cited by 115PDFcodeScholar