← Search

Balaji Vasan Srinivasan

16 accepted papers

2026

Step-by-step Layered Design Generation

AAAI 2026technical

Design generation, in its essence, is a step-by-step process where designers progressively refine and enhance their work through careful modifications. Despite this fundamental characteristic, existing approaches mainly treat design synthesis as a single-step generation problem, significantly undere

Cited by 0SourcePDFScholar
2025

Imposter: Text and Frequency Guidance for Subject Driven Action Personalization using Diffusion Models

COLING 2025main

We present ImPoster, a novel algorithm for generating a target image of a ‘source’ subject performing a ‘driving’ action. The inputs to our algorithm are a single pair of a source image with the subject that we wish to edit and a driving image with a subject of an arbitrary class performing the driv…

2024

CoPL: Contextual Prompt Learning for Vision-Language Understanding

AAAI 2024technical

Recent advances in multimodal learning has resulted in powerful vision-language models, whose representations are generalizable across a variety of downstream tasks. Recently, their generalization ability has been further extended by incorporating trainable prompts, borrowed from the natural languag…

Cited by 7SourcePDFScholar
2024

Enhancing Post-Hoc Attributions in Long Document Comprehension via Coarse Grained Answer Decomposition

EMNLP 2024main

Accurately attributing answer text to its source document is crucial for developing a reliable question-answering system. However, attribution for long documents remains largely unexplored. Post-hoc attribution systems are designed to map answer text back to the source document, yet the granularity…

Cited by 3SourcePDFScholar
2024

MeLFusion: Synthesizing Music from Image and Language Cues using Diffusion Models

CVPR 2024highlight

Music is a universal language that can communicate emotions and feelings. It forms an essential part of the whole spectrum of creative media ranging from movies to social media posts. Machine learning models that can synthesize music are predominantly conditioned on textual descriptions of it. Inspi…

2024

Peering into the Mind of Language Models: An Approach for Attribution in Contextual Question Answering

ACL 2024findings

With the enhancement in the field of generative artificial intelligence (AI), contextual question answering has become extremely relevant. Attributing model generations to the input source document is essential to ensure trustworthiness and reliability. We observe that when large language models (LL…

2023

A-STAR: Test-time Attention Segregation and Retention for Text-to-image Synthesis

ICCV 2023poster

While recent developments in text-to-image generative models have led to a suite of high-performing methods capable of producing creative imagery from free-form text, there are several limitations. By analyzing the cross-attention representations of these models, we notice two key issues. First, for…

Cited by 44PDFScholar
2023

Sketch Recognition via Part-based Hierarchical Analogical Learning

IJCAI 2023poster

Sketch recognition has been studied for decades, but it is far from solved. Drawing styles are highly variable across people and adapting to idiosyncratic visual expressions requires data-efficient learning. Explainability also matters, so that users can see why a system got confused about something…

Cited by 4SourcePDFScholar
2023

What to Read in a Contract? Party-Specific Summarization of Legal Obligations, Entitlements, and Prohibitions

EMNLP 2023long main

Reviewing and comprehending key obligations, entitlements, and prohibitions in legal contracts can be a tedious task due to their length and domain-specificity. Furthermore, the key rights and duties requiring review vary for each contracting party. In this work, we propose a new task of \textit{par…

Cited by 0SourceScholar
2022

Agent-Specific Deontic Modality Detection in Legal Language

EMNLP 2022main

Legal documents are typically long and written in legalese, which makes it particularly difficult for laypeople to understand their rights and duties. While natural language understanding technologies can be valuable in supporting such understanding in the legal domain, the limited availability of d…

2021

ClauseRec: A Clause Recommendation Framework for AI-aided Contract Authoring

EMNLP 2021main

Contracts are a common type of legal document that frequent in several day-to-day business workflows. However, there has been very limited NLP research in processing such documents, and even lesser in generating them. These contracts are made up of clauses, and the unique nature of these clauses cal…

Cited by 10SourcePDFScholar
2021

IGA: An Intent-Guided Authoring Assistant

EMNLP 2021main

While large-scale pretrained language models have significantly improved writing assistance functionalities such as autocomplete, more complex and controllable writing assistants have yet to be explored. We leverage advances in language modeling to build an interactive writing assistant that generat…

2021

MIMOQA: Multimodal Input Multimodal Output Question Answering

NAACL 2021long

Multimodal research has picked up significantly in the space of question answering with the task being extended to visual question answering, charts question answering as well as multimodal input question answering. However, all these explorations produce a unimodal textual output as the answer. In…

Cited by 40SourcePDFScholar
2021

Multi-Style Transfer with Discriminative Feedback on Disjoint Corpus

NAACL 2021long

Style transfer has been widely explored in natural language generation with non-parallel corpus by directly or indirectly extracting a notion of style from source and target domain corpus. A common shortcoming of existing approaches is the prerequisite of joint annotations across all the stylistic d…