← Search

Balaji Krishnamurthy

33 accepted papers

2026

ALPHA: Action-Based Learning for Pluralistic Human Alignment in Large Language Models

AAAI 2026technical

Large language models are widely used, but aligning them with societal values remains challenging. Current approaches often rely on human annotations, which are hard to scale, or synthetic data produced by models that may themselves be misaligned, making it difficult to capture genuine public opinio

Cited by 0SourcePDFScholar
2026

Social Agents: Collective Intelligence Improves LLM Predictions

ICLR 2026poster

In human society, collective decision making has often outperformed the judgment of individuals. Classic examples range from estimating livestock weights to predicting elections and financial markets, where averaging many independent guesses often yields results more accurate than experts. These suc…

Cited by 0SourceScholar
2025

AesthetiQ: Enhancing Graphic Layout Design via Aesthetic-Aware Preference Alignment of Multi-modal Large Language Models

CVPR 2025poster

Visual layouts are essential in graphic design fields such as advertising, posters, and web interfaces. The application of generative models for content-aware layout generation has recently gained traction. However, these models fail to understand the contextual aesthetic requirements of layout desi…

Cited by 0SourcePDFScholar
2025

It Helps to Take a Second Opinion: Teaching Smaller LLMs To Deliberate Mutually via Selective Rationale Optimisation

ICLR 2025poster

Very large language models (LLMs) such as GPT-4 have shown the ability to handle complex tasks by generating and self-refining step-by-step rationales. Smaller language models (SLMs), typically with < 13B parameters, have been improved by using the data generated from very-large LMs through knowledg…

Cited by 0SourcePDFScholar
2025

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning

ACL 2025long

LLMs such as GPT-4 have shown a remarkable ability to solve complex questions by generating step-by-step rationales. Prior works have utilized this capability to improve smaller and cheaper LMs (say, with 7B parameters). However, various practical constraints, such as copyright and legal issues, owi…

2025

Measuring And Improving Engagement of Text-to-Image Generation Models

ICLR 2025poster

Recent advances in text-to-image generation have achieved impressive aesthetic quality, making these models usable for both personal and commercial purposes. However, in the fields of marketing and advertising, images are often created to be more engaging, as reflected in user behaviors such as incr…

2025

Measuring And Improving Persuasiveness Of Large Language Models

ICLR 2025poster

Large Language Models (LLMs) are increasingly being used in workflows involving generating content to be consumed by humans (*e.g.,* marketing) and also in directly interacting with humans (*e.g.,* through chatbots). The development of such systems that are capable of generating verifiably persuasiv…

2025

SPRO: Improving Image Generation via Self-Play

NeurIPS 2025poster

Recent advances in diffusion models have dramatically improved image fidelity and diversity. However, aligning these models with nuanced human preferences -such as aesthetics, engagement, and subjective appeal remains a key challenge due to the scarcity of large-scale human annotations. Collecting s…

Cited by 0SourceScholar
2025

Teaching Human Behavior Improves Content Understanding Abilities Of VLMs

ICLR 2025poster

Communication is defined as "*Who* says *what* to *whom* with *what* effect." A message from a communicator generates downstream receiver effects, also known as behavior. Receiver behavior, being a downstream effect of the message, carries rich signals about it. Even after carrying signals about the…

2024

All Should Be Equal in the Eyes of LMs: Counterfactually Aware Fair Text Generation

AAAI 2024technical

Fairness in Language Models (LMs) remains a long-standing challenge, given the inherent biases in training data that can be perpetuated by models and affect the downstream tasks. Recent methods employ expensive retraining or attempt debiasing during inference by constraining model outputs to contras…

Cited by 1SourcePDFScholar
2024

CABINET: Content Relevance-based Noise Reduction for Table Question Answering

ICLR 2024spotlight

Table understanding capability of Large Language Models (LLMs) has been extensively studied through the task of question-answering (QA) over tables. Typically, only a small part of the whole table is relevant to derive the answer for a given question. The irrelevant parts act as noise and are distra…

2024

Evaluating the Efficacy of Prompting Techniques for Debiasing Language Model Outputs (Student Abstract)

AAAI 2024technical

Achieving fairness in Large Language Models (LLMs) continues to pose a persistent challenge, as these models are prone to inheriting biases from their training data, which can subsequently impact their performance in various applications. There is a need to systematically explore whether structured…

Cited by 2SourcePDFScholar
2024

Large Content And Behavior Models To Understand, Simulate, And Optimize Content And Behavior

ICLR 2024spotlight

Shannon and Weaver's seminal information theory divides communication into three levels: technical, semantic, and effectiveness. While the technical level deals with the accurate reconstruction of transmitted symbols, the semantic and effectiveness levels deal with the inferred meaning and its effec…

2023

A Video Is Worth 4096 Tokens: Verbalize Story Videos To Understand Them In Zero Shot

EMNLP 2023long main

Multimedia content, such as advertisements and story videos, exhibit a rich blend of creativity and multiple modalities. They incorporate elements like text, visuals, audio, and storytelling techniques, employing devices like emotions, symbolism, and slogans to convey meaning. There is a dearth of l…

Cited by 0SourceScholar
2023

Explaining RL Decisions with Trajectories

ICLR 2023poster

Explanation is a key component for the adoption of reinforcement learning (RL) in many real-world decision-making problems. In the literature, the explanation is often provided by saliency attribution to the features of the RL agent's state. In this work, we propose a complementary approach to thes…

2023

HyHTM: Hyperbolic Geometry-based Hierarchical Topic Model

ACL 2023findings

Hierarchical Topic Models (HTMs) are useful for discovering topic hierarchies in a collection of documents. However, traditional HTMs often produce hierarchies where lower-level topics are unrelated and not specific enough to their higher-level topics. Additionally, these methods can be computationa…

2023

INGENIOUS: Using Informative Data Subsets for Efficient Pre-Training of Language Models

EMNLP 2023long findings

A salient characteristic of pre-trained language models (PTLMs) is a remarkable improvement in their generalization capability and emergence of new capabilities with increasing model capacity and pre-training dataset size. Consequently, we are witnessing the development of enormous models pushing th…

Cited by 0SourcecodeScholar
2023

Parameter Efficient Local Implicit Image Function Network for Face Segmentation

CVPR 2023poster

Face parsing is defined as the per-pixel labeling of images containing human faces. The labels are defined to identify key facial regions like eyes, lips, nose, hair, etc. In this work, we make use of the structural consistency of the human face to propose a lightweight face-parsing method using a L…

Cited by 10SourcePDFScholar
2023

Persuasion Strategies in Advertisements

AAAI 2023technical

Modeling what makes an advertisement persuasive, i.e., eliciting the desired response from consumer, is critical to the study of propaganda, social psychology, and marketing. Despite its importance, computational modeling of persuasion in computer vision is still in its infancy, primarily due to the…

2023

UMFuse: Unified Multi View Fusion for Human Editing Applications

ICCV 2023poster

Numerous pose-guided human editing methods have been explored by the vision community due to their extensive practical applications. However, most of these methods still use an image-to-image formulation in which a single image is given as input to produce an edited image as output. This objective b…

Cited by 1PDFScholar
2023

VGFlow: Visibility Guided Flow Network for Human Reposing

CVPR 2023poster

The task of human reposing involves generating a realistic image of a model standing in an arbitrary conceivable pose. There are multiple difficulties in generating perceptually accurate images and existing methods suffers from limitations in preserving texture, maintaining pattern coherence, respec…

Cited by 7SourcePDFScholar
2022

CoSe-Co: Text Conditioned Generative CommonSense Contextualizer

NAACL 2022long

Pre-trained Language Models (PTLMs) have been shown to perform well on natural language tasks. Many prior works have leveraged structured commonsense present in the form of entities linked through labeled relations in Knowledge Graphs (KGs) to assist PTLMs. Retrieval approaches use KG as a separate…

Cited by 5SourcePDFScholar
2022

Distilling the Undistillable: Learning from a Nasty Teacher

ECCV 2022poster

"The inadvertent stealing of private/sensitive information using Knowledge Distillation (KD) has been getting significant attention recently and has guided subsequent defense efforts considering its critical nature. Recent work \textit{Nasty Teacher} proposed to develop teachers which can not be dis…

2022

LM-CORE: Language Models with Contextually Relevant External Knowledge

NAACL 2022findings

Large transformer-based pre-trained language models have achieved impressive performance on a variety of knowledge-intensive tasks and can capture factual knowledge in their parameters. We argue that storing large amounts of knowledge in the model parameters is sub-optimal given the ever-growing amo…

2022

MINIMAL: Mining Models for Universal Adversarial Triggers

AAAI 2022technical

It is well known that natural language models are vulnerable to adversarial attacks, which are mostly input-specific in nature. Recently, it has been shown that there also exist input-agnostic attacks in NLP models, called universal adversarial triggers. However, existing methods to craft universal…

Cited by 5SourcePDFScholar
2021

TAN-NTM: Topic Attention Networks for Neural Topic Modeling

ACL 2021long

Topic models have been widely used to learn text representations and gain insight into document corpora. To perform topic discovery, most existing neural models either take document bag-of-words (BoW) or sequence of tokens as input followed by variational inference and BoW reconstruction to learn to…

2021

What Ails One-Shot Image Segmentation: A Data Perspective

NeurIPS 2021poster

One-shot image segmentation (OSS) methods enable semantic labeling of image pixels without supervised training with an extensive dataset. They require just one example (image, mask) pair per target class. Most neural-network-based methods train on a large subset of dataset classes and are evaluated…

Cited by 3SourcecodeScholar
2021

ZFlow: Gated Appearance Flow-Based Virtual Try-On With 3D Priors

ICCV 2021poster

Image-based virtual try-on involves synthesizing perceptually convincing images of a model wearing a particular garment and has garnered significant research interest due to its immense practical applicability. Recent methods involve a two-stage process: i) warping of the garment to align with the m…

Cited by 77PDFScholar
2020

Attributional Robustness Training using Input-Gradient Spatial Alignment

ECCV 2020poster

Interpretability is an emerging area of research in trustworthy machine learning. Safe deployment of machine learning system mandates that the prediction and its explanation be reliable and robust. Recently, it has been shown that the explanations could be manipulated easily by adding visually imper…

2020

Document Structure Extraction using Prior based High Resolution Hierarchical Semantic Segmentation

ECCV 2020poster

Structure extraction from document images has been a long-standing research topic due to its high impact on a wide range of practical applications. In this paper, we share our findings on employing a hierarchical semantic segmentation network for this task of structure extraction. We propose a prior…

Cited by 22SourcePDFScholar
2020

Explain Your Move: Understanding Agent Actions Using Specific and Relevant Feature Attribution

ICLR 2020poster

As deep reinforcement learning (RL) is applied to more tasks, there is a need to visualize and understand the behavior of learned agents. Saliency maps explain agent behavior by highlighting the features of the input state that are most relevant for the agent in taking an action. Existing perturbati…

Cited by 99SourcecodeScholar
2020

SimPropNet: Improved Similarity Propagation for Few-shot Image Segmentation

IJCAI 2020poster

Few-shot segmentation (FSS) methods perform image segmentation for a particular object class in a target (query) image, using a small set of (support) image-mask pairs. Recent deep neural network based FSS methods leverage high-dimensional feature similarity between the foreground features of the su…

Cited by 0SourcePDFScholar
2017

Introspection:Accelerating Neural Network Training By Learning Weight Evolution

ICLR 2017poster

Neural Networks are function approximators that have achieved state-of-the-art accuracy in numerous machine learning tasks. In spite of their great success in terms of accuracy, their large training time makes it difficult to use them for various tasks. In this paper, we explore the idea of learning…

Cited by 31SourceScholar