← Search

Anjan Dutta

13 accepted papers

2026

RAIGen: Rare Attribute Identification in Text-to-Image Generative Models

ICML 2026poster

Text-to-image diffusion models achieve impressive generation quality but inherit and amplify training-data biases, skewing coverage of semantic attributes. Prior work addresses this in two ways. Closed-set approaches mitigate biases in predefined fairness categories (e.g., gender, race), assuming so…

Cited by 0SourceScholar
2025

A Closer Look at Multimodal Representation Collapse

ICML 2025spotlight

We aim to develop a fundamental understanding of modality collapse, a recently observed empirical phenomenon wherein models trained for multimodal fusion tend to rely only on a subset of the modalities, ignoring the rest. We show that modality collapse happens when noisy features from one modality a…

2025

OmniCount: Multi-label Object Counting with Semantic-Geometric Priors

AAAI 2025technical

Object counting is pivotal for understanding the composition of scenes. Previously, this task was dominated by class-specific methods, which have gradually evolved into more adaptable class-agnostic strategies. However, these strategies come with their own set of limitations, such as the need for ma…

Cited by 2SourcePDFScholar
2025

RespoDiff: Dual-Module Bottleneck Transformation for Responsible & Faithful T2I Generation

NeurIPS 2025poster

The rapid advancement of diffusion models has enabled high-fidelity and semantically rich text-to-image generation; however, ensuring fairness and safety remains an open challenge. Existing methods typically improve fairness and safety at the expense of semantic fidelity and image quality. In this w…

Cited by 0SourcecodeScholar
2024

DeNetDM: Debiasing by Network Depth Modulation

NeurIPS 2024poster

Neural networks trained on biased datasets tend to inadvertently learn spurious correlations, hindering generalization. We formally prove that (1) samples that exhibit spurious correlations lie on a lower rank manifold relative to the ones that do not; and (2) the depth of a network acts as an impli…

2024

Learning Conditional Invariances through Non-Commutativity

ICLR 2024poster

Invariance learning algorithms that conditionally filter out domain-specific random variables as distractors, do so based only on the data semantics, and not the target domain under evaluation. We show that a provably optimal and sample-efficient way of learning conditional invariances is by relaxin…

2023

Transitivity Recovering Decompositions: Interpretable and Robust Fine-Grained Relationships

NeurIPS 2023poster

Recent advances in fine-grained representation learning leverage local-to-global (emergent) relationships for achieving state-of-the-art results. The relational representations relied upon by such methods, however, are abstract. We aim to deconstruct this abstraction by expressing them as interpreta…

2022

Abstracting Sketches through Simple Primitives

ECCV 2022poster

"Humans show high-level of abstraction capabilities in games that require quickly communicating object information. They decompose the message content into multiple parts and communicate them in an interpretable protocol. Toward equipping machines with such capabilities, we propose the Primitive-bas…

2022

Relational Proxies: Emergent Relationships as Fine-Grained Discriminators

NeurIPS 2022accept

Fine-grained categories that largely share the same set of parts cannot be discriminated based on part information alone, as they mostly differ in the way the local parts relate to the overall global structure of the object. We propose Relational Proxies, a novel approach that leverages the relation…

2020

Learning Robust Representations via Multi-View Information Bottleneck

ICLR 2020poster

The information bottleneck principle provides an information-theoretic method for representation learning, by training an encoder to retain all information which is relevant for predicting the label while minimizing the amount of other, excess information in the representation. The original formulat…

Cited by 323SourcecodeScholar
2019

Doodle to Search: Practical Zero-Shot Sketch-Based Image Retrieval

CVPR 2019oral

In this paper, we investigate the problem of zero-shot sketch-based image retrieval (ZS-SBIR), where human sketches are used as queries to conduct retrieval of photos from unseen categories. We importantly advance prior arts by proposing a novel ZS-SBIR scenario that represents a firm step forward i…

Cited by 224PDFScholar