← Search

Utkarsh Ojha

12 accepted papers

2025

Aligned Datasets Improve Detection of Latent Diffusion-Generated Images

ICLR 2025poster

As latent diffusion models (LDMs) democratize image generation capabilities, there is a growing need to detect fake images. A good detector should focus on the generative model’s fingerprints while ignoring image properties such as semantic content, resolution, file format, etc. Fake image detectors…

Cited by 0SourcePDFScholar
2024

Edit One for All: Interactive Batch Image Editing

CVPR 2024poster

In recent years image editing has advanced remarkably. With increased human control it is now possible to edit an image in a plethora of ways; from specifying in text what we want to change to straight up dragging the contents of the image in an interactive point-based manner. However most of the fo…

Cited by 4SourcePDFScholar
2024

Yo'LLaVA: Your Personalized Language and Vision Assistant

NeurIPS 2024poster

Large Multimodal Models (LMMs) have shown remarkable capabilities across a variety of tasks (e.g., image captioning, visual question answering). While broad, their knowledge remains generic (e.g., recognizing a dog), and they are unable to handle personalized subjects (e.g., recognizing a user's pet…

2023

Towards Universal Fake Image Detectors That Generalize Across Generative Models

CVPR 2023poster

With generative models proliferating at a rapid rate, there is a growing need for general purpose fake image detectors. In this work, we first show that the existing paradigm, which consists of training a deep network for real-vs-fake classification, fails to detect fake images from newer breeds of…

2023

Visual Instruction Inversion: Image Editing via Image Prompting

NeurIPS 2023poster

Text-conditioned image editing has emerged as a powerful tool for editing images. However, in many situations, language can be ambiguous and ineffective in describing specific image edits. When faced with such challenges, visual prompts can be a more informative and intuitive way to convey ideas. We…

Cited by 48SourcePDFScholar
2023

What Knowledge Gets Distilled in Knowledge Distillation?

NeurIPS 2023poster

Knowledge distillation aims to transfer useful information from a teacher network to a student network, with the primary goal of improving the student's performance for the task at hand. Over the years, there has a been a deluge of novel techniques and use cases of knowledge distillation. Yet, despi…

Cited by 30SourcePDFScholar
2021

Few-Shot Image Generation via Cross-Domain Correspondence

CVPR 2021poster

Training generative models, such as GANs, on a target domain containing limited examples (e.g., 10) can easily result in overfitting. In this work, we seek to utilize a large source domain for pretraining and transfer the diversity information from source to target. We propose to preserve the relati…

Cited by 294PDFcodeScholar
2021

Generating Furry Cars: Disentangling Object Shape and Appearance across Multiple Domains

ICLR 2021poster

We consider the novel task of learning disentangled representations of object shape and appearance across multiple domains (e.g., dogs and cars). The goal is to learn a generative model that learns an intermediate distribution, which borrows a subset of properties from each domain, enabling the gen…

Cited by 13SourcePDFScholar
2020

Elastic-InfoGAN: Unsupervised Disentangled Representation Learning in Class-Imbalanced Data

NeurIPS 2020poster

We propose a novel unsupervised generative model that learns to disentangle object identity from other low-level aspects in class-imbalanced data. We first investigate the issues surrounding the assumptions about uniformity made by InfoGAN, and demonstrate its ineffectiveness to properly disentangle…

2020

MixNMatch: Multifactor Disentanglement and Encoding for Conditional Image Generation

CVPR 2020poster

We present MixNMatch, a conditional generative model that learns to disentangle and encode background, object pose, shape, and texture from real images with minimal supervision, for mix-and-match image generation. We build upon FineGAN, an unconditional generative model, to learn the desired disenta…

Cited by 101PDFcodeScholar
2019

FineGAN: Unsupervised Hierarchical Disentanglement for Fine-Grained Object Generation and Discovery

CVPR 2019oral

We propose FineGAN, a novel unsupervised GAN framework, which disentangles the background, object shape, and object appearance to hierarchically generate images of fine-grained object categories. To disentangle the factors without supervision, our key idea is to use information theory to associate e…

Cited by 177PDFcodeScholar