← Search

Nalin Gupta

2 accepted papers

2024

Generating Contextual Images for Long-Form Text

COLING 2024main

We investigate the problem of synthesizing relevant visual imagery from generic long-form text, leveraging Large Language Models (LLMs) and Text-to-Image Models (TIMs). Current Text-to-Image models require short prompts that describe the image content and style explicitly. Unlike image prompts, gene…

Cited by 0SourcePDFScholar
2022

Multimodal Context Carryover

EMNLP 2022industry

Multi-modality support has become an integral part of creating a seamless user experience with modern voice assistants with smart displays. Users refer to images, video thumbnails, or the accompanying text descriptions on the screen through voice communication with AI powered devices. This raises th…

Cited by 3SourcePDFScholar