← Search

Koustava Goswami

18 accepted papers

2026

Crossing Borders: A Multimodal Challenge for Indian Poetry Translation and Image Generation

AAAI 2026technical

Indian poetry, known for its linguistic complexity and deep cultural resonance, has a rich and varied heritage spanning thousands of years. However, its layered meanings, cultural allusions, and sophisticated grammatical constructions often pose challenges for comprehension, especially for non-nativ

Cited by 0SourcePDFScholar
2026

SmartChunk Retrieval: Query-Aware Chunk Compression with Planning for Efficient Document RAG

ICLR 2026poster

Retrieval-augmented generation (RAG) has strong potential for producing accurate and factual outputs by combining language models (LMs) with evidence retrieved from large text corpora. However, current pipelines are limited by static chunking and flat retrieval: documents are split into short, prede…

Cited by 1SourceScholar
2026

Step-by-step Layered Design Generation

AAAI 2026technical

Design generation, in its essence, is a step-by-step process where designers progressively refine and enhance their work through careful modifications. Despite this fundamental characteristic, existing approaches mainly treat design synthesis as a single-step generation problem, significantly undere

Cited by 0SourcePDFScholar
2025

Beyond Logit Lens: Contextual Embeddings for Robust Hallucination Detection & Grounding in VLMs

NAACL 2025long

The rapid development of Large Multimodal Models (LMMs) has significantly advanced multimodal understanding by harnessing the language abilities of Large Language Models (LLMs) and integrating modality-specific encoders. However, LMMs are plagued by hallucinations that limit their reliability and ad…

Cited by 1SourcePDFScholar
2025

Do It Yourself (DIY): Modifying Images for Poems in a Zero-Shot Setting Using Weighted Prompt Manipulation

EMNLP 2025

Poetry is an expressive form of art that invites multiple interpretations, as readers often bring their own emotions, experiences, and cultural backgrounds into their understanding of a poem. Recognizing this, we aim to generate images for poems and improve these images in a zero-shot setting, enabl

Cited by 0SourcePDFScholar
2025

Poetry in Pixels: Prompt Tuning for Poem Image Generation via Diffusion Models

COLING 2025main

The task of text-to-image generation has encountered significant challenges when applied to literary works, especially poetry. Poems are a distinct form of literature, with meanings that frequently transcend beyond the literal words. To address this shortcoming, we propose a PoemToPixel framework de…

2025

Subjective Behaviors and Preferences in LLM: Language of Browsing

EMNLP 2025

A Large Language Model (LLM) offers versatility across domains and tasks, purportedly benefiting users with a wide variety of behaviors and preferences. We question this perception about an LLM when users have inherently subjective behaviors and preferences, as seen in their ubiquitous and idiosyncr

Cited by 0SourcePDFScholar
2024

An Audit on the Perspectives and Challenges of Hallucinations in NLP

EMNLP 2024main

We audit how hallucination in large language models (LLMs) is characterized in peer-reviewed literature, using a critical examination of 103 publications across NLP research. Through the examination of the literature, we identify a lack of agreement with the term ‘hallucination’ in the field of NLP.…

2024

CoPL: Contextual Prompt Learning for Vision-Language Understanding

AAAI 2024technical

Recent advances in multimodal learning has resulted in powerful vision-language models, whose representations are generalizable across a variety of downstream tasks. Recently, their generalization ability has been further extended by incorporating trainable prompts, borrowed from the natural languag…

Cited by 7SourcePDFScholar
2024

Enhancing Post-Hoc Attributions in Long Document Comprehension via Coarse Grained Answer Decomposition

EMNLP 2024main

Accurately attributing answer text to its source document is crucial for developing a reliable question-answering system. However, attribution for long documents remains largely unexplored. Post-hoc attribution systems are designed to map answer text back to the source document, yet the granularity…

Cited by 3SourcePDFScholar
2024

Peering into the Mind of Language Models: An Approach for Attribution in Contextual Question Answering

ACL 2024findings

With the enhancement in the field of generative artificial intelligence (AI), contextual question answering has become extremely relevant. Attributing model generations to the input source document is essential to ensure trustworthiness and reliability. We observe that when large language models (LL…

2024

SAFARI: Adaptive Sequence Transformer for Weakly Supervised Referring Expression Segmentation

ECCV 2024poster

"Referring Expression Segmentation (RES) aims to provide a segmentation mask of the target object in an image referred to by the text (i.e., referring expression). Existing methods require large-scale mask annotations. Moreover, such approaches do not generalize well to unseen/zero-shot scenarios. T…

2023

A-STAR: Test-time Attention Segregation and Retention for Text-to-image Synthesis

ICCV 2023poster

While recent developments in text-to-image generative models have led to a suite of high-performing methods capable of producing creative imagery from free-form text, there are several limitations. By analyzing the cross-attention representations of these models, we notice two key issues. First, for…

Cited by 44PDFScholar
2023

Drilling Down into the Discourse Structure with LLMs for Long Document Question Answering

EMNLP 2023long findings

We address the task of evidence retrieval for long document question answering, which involves locating relevant paragraphs within a document to answer a question. We aim to assess the applicability of large language models (LLMs) in the task of zero-shot long document evidence retrieval, owing to t…

Cited by 0SourceScholar
2021

Cross-lingual Sentence Embedding using Multi-Task Learning

EMNLP 2021main

Multilingual sentence embeddings capture rich semantic information not only for measuring similarity between texts but also for catering to a broad range of downstream cross-lingual NLP tasks. State-of-the-art multilingual sentence embedding models require large parallel corpora to learn efficiently…

Cited by 26SourcePDFScholar
2020

Suggest me a movie for tonight: Leveraging Knowledge Graphs for Conversational Recommendation

COLING 2020main

Conversational recommender systems focus on the task of suggesting products to users based on the conversation flow. Recently, the use of external knowledge in the form of knowledge graphs has shown to improve the performance in recommendation and dialogue systems. Information from knowledge graphs…

2020

Unsupervised Deep Language and Dialect Identification for Short Texts

COLING 2020main

Automatic Language Identification (LI) or Dialect Identification (DI) of short texts of closely related languages or dialects, is one of the primary steps in many natural language processing pipelines. Language identification is considered a solved task in many cases; however, in the case of very cl…

Cited by 13SourcePDFScholar