← Search

Krishna Kumar Singh

38 accepted papers

2026

Beyond Simple Edits: X-Planner for Complex Instruction-Based Image Editing

AAAI 2026technical

Recent diffusion-based image editing methods have made great strides in text-guided tasks but often struggle with complex, indirect instructions. Additionally, current models frequently exhibit poor identity preservation, unintended edits, or rely on manual masks. To overcome these limitations, we i

Cited by 0SourcePDFScholar
2026

Group Diffusion: Enhancing Image Generation by Unlocking Cross-Sample Collaboration

CVPR 2026

In this work, we explore an untapped signal in diffusion model inference. While all previous methods generate images independently at inference, we instead ask if samples can be generated collaboratively. We propose Group Diffusion, unlocking the attention mechanism to be shared across images, rathe

Cited by 0SourcecodeScholar
2026

Learning an Image Editing Model without Image Editing Pairs

ICLR 2026poster

Recent image editing models have achieved impressive results while following natural language editing instructions, but they rely on supervised fine-tuning with large datasets of input-target pairs. This is a critical bottleneck, as such naturally occurring pairs are hard to curate at scale. Curren…

Cited by 0SourcecodeScholar
2026

Stepwise Credit Assignment for GRPO on Flow-Matching Models

CVPR 2026

Flow-GRPO successfully applies reinforcement learning to flow models, but uses uniform credit assignment across all steps. This ignores the temporal structure of diffusion generation: early steps determine composition and content (low-frequency structure), while late steps resolve details and textur

Cited by 0SourceScholar
2025

Can3Tok: Canonical 3D Tokenization and Latent Modeling of Scene-Level 3D Gaussians

ICCV 2025poster

3D generation has made significant progress, however, it still largely remains at the object-level. Feedforward 3D scene-level generation has been rarely explored due to the lack of models capable of scaling-up latent representation learning on 3D scene-level data. Unlike object-level generative mod…

2025

Comprehensive Relighting: Generalizable and Consistent Monocular Human Relighting and Harmonization

CVPR 2025poster

This paper introduces Comprehensive Relighting, the first all-in-one approach that can both control and harmonize the lighting from an image or video of humans with arbitrary body parts from any scene. Building such a generalizable model is extremely challenging due to the lack of dataset, restricti…

Cited by 0SourcePDFScholar
2025

DOLLAR: Few-Step Video Generation via Distillation and Latent Reward Optimization

ICCV 2025poster

Diffusion probabilistic models have shown significant progress in video generation; however, their computational efficiency is limited by the large number of sampling steps required. Reducing sampling steps often compromises video quality or generation diversity. In this work, we introduce a distill…

Cited by 0SourcePDFScholar
2025

Generating, Fast and Slow: Scalable Parallel Video Generation with Video Interface Networks

ICCV 2025poster

Diffusion Transformers (DiTs) can generate short photorealistic videos, yet directly training and sampling longer videos with full attention across the video remains computationally challenging. Alternative methods break long videos down into sequential generation of short video segments, requiring…

2025

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models

CVPR 2025poster

Current diffusion-based text-to-video methods are limited to producing short video clips of a single shot and lack the capability to generate multi-shot videos with discrete transitions where the same character performs distinct activities across the same or different backgrounds. To address this li…

2025

Text2Relight: Creative Portrait Relighting with Text Guidance

AAAI 2025technical

We present a lighting-aware image editing pipeline that, given a portrait image and a text prompt, performs single image relighting. Our model modifies the lighting and color of both the foreground and background to align with the provided text description. The unbounded nature in creativeness of a…

Cited by 1SourcePDFScholar
2025

X-Fusion: Introducing New Modality to Frozen Large Language Models

ICCV 2025poster

We propose X-Fusion, a framework that extends pretrained Large Language Models (LLMs) for multimodal tasks while preserving their language capabilities. X-Fusion employs a dual-tower design with modality-specific weights, keeping the LLM's parameters frozen while integrating vision-specific informat…

Cited by 0SourcePDFScholar
2025

Yo'Chameleon: Personalized Vision and Language Generation

CVPR 2025poster

Large Multimodal Models (e.g., GPT-4, Gemini, Chameleon) have evolved into powerful tools with millions of users. However, they remain generic models and lack personalized knowledge of specific user concepts. Previous work has explored personalization for text generation, yet it remains unclear how…

Cited by 1SourcePDFScholar
2024

ActAnywhere: Subject-Aware Video Background Generation

NeurIPS 2024poster

We study a novel problem to automatically generate video background that tailors to foreground subject motion. It is an important problem for the movie industry and visual effects community, which traditionally requires tedious manual efforts to solve. To this end, we propose ActAnywhere, a video di…

2024

GroupDiff: Diffusion-based Group Portrait Editing

ECCV 2024poster

"Group portrait editing is highly desirable since users constantly want to add a person, delete a person, or manipulate existing persons. It is also challenging due to the intricate dynamics of human interactions and the diverse gestures. In this work, we present GroupDiff, a pioneering effort to ta…

2024

Removing Distributional Discrepancies in Captions Improves Image-Text Alignment

ECCV 2024poster

"In this paper, we introduce a model designed to improve the prediction of image-text alignment, targeting the challenge of compositional understanding in current visual-language models. Our approach focuses on generating high-quality training datasets for the alignment task by producing mixed-type…

2024

Revisiting Feature Disentanglement Strategy in Diffusion Training and Breaking Conditional Independence Assumption in Sampling

ECCV 2024poster

"As Diffusion Models have shown promising performance, a lot of efforts have been made to improve the controllability of Diffusion Models. However, how to train Diffusion Models to have the disentangled latent spaces and how to naturally incorporate the disentangled conditions during the sampling pr…

Cited by 0SourcePDFScholar
2024

UniHuman: A Unified Model For Editing Human Images in the Wild

CVPR 2024poster

Human image editing includes tasks like changing a person's pose their clothing or editing the image according to a text prompt. However prior work often tackles these tasks separately overlooking the benefit of mutual reinforcement from learning them jointly. In this paper we propose UniHuman a uni…

2023

Complete 3D Human Reconstruction From a Single Incomplete Image

CVPR 2023poster

This paper presents a method to reconstruct a complete human geometry and texture from an image of a person with only partial body observed, e.g., a torso. The core challenge arises from the occlusion: there exists no pixel to reconstruct where many existing single-view human reconstruction methods…

Cited by 16SourcePDFScholar
2023

Putting People in Their Place: Affordance-Aware Human Insertion Into Scenes

CVPR 2023poster

We study the problem of inferring scene affordances by presenting a method for realistically inserting people into scenes. Given a scene image with a marked region and an image of a person, we insert the person into the scene while respecting the scene affordances. Our model can infer the set of rea…

2023

UMFuse: Unified Multi View Fusion for Human Editing Applications

ICCV 2023poster

Numerous pose-guided human editing methods have been explored by the vision community due to their extensive practical applications. However, most of these methods still use an image-to-image formulation in which a single image is given as input to produce an edited image as output. This objective b…

Cited by 1PDFScholar
2023

VGFlow: Visibility Guided Flow Network for Human Reposing

CVPR 2023poster

The task of human reposing involves generating a realistic image of a model standing in an arbitrary conceivable pose. There are multiple difficulties in generating perceptually accurate images and existing methods suffers from limitations in preserving texture, maintaining pattern coherence, respec…

Cited by 7SourcePDFScholar
2022

Contrastive Learning for Diverse Disentangled Foreground Generation

ECCV 2022poster

"We introduce a new method for diverse foreground generation with explicit control over various factors. Existing image inpainting based foreground generation methods often struggle to generate diverse results and rarely allow users to explicitly control specific factors of variation (e.g., varying…

2022

InsetGAN for Full-Body Image Generation

CVPR 2022poster

While GANs can produce photo-realistic images in ideal conditions for certain domains, the generation of full-body human images remains difficult due to the diversity of identities, hairstyles, clothing, and the variance in pose. Instead of modeling this complex domain with a single GAN, we propose…

Cited by 69PDFcodeScholar
2022

Spatially-Adaptive Multilayer Selection for GAN Inversion and Editing

CVPR 2022poster

Existing GAN inversion and editing methods work well for aligned objects with a clean background, such as portraits and animal faces, but often struggle for more difficult categories with complex scene layouts and object occlusions, such as cars, animals, and outdoor images. We propose a new method…

Cited by 49PDFcodeScholar
2021

Collaging Class-Specific GANs for Semantic Image Synthesis

ICCV 2021poster

We propose a new approach for high resolution semantic image synthesis. It consists of one base image generator and multiple class-specific generators. The base generator generates high quality images based on a segmentation map. To further improve the quality of different objects, we create a bank…

Cited by 43PDFScholar
2021

Generating Furry Cars: Disentangling Object Shape and Appearance across Multiple Domains

ICLR 2021poster

We consider the novel task of learning disentangled representations of object shape and appearance across multiple domains (e.g., dogs and cars). The goal is to learn a generative model that learns an intermediate distribution, which borrows a subset of properties from each domain, enabling the gen…

Cited by 13SourcePDFScholar
2021

IMAGINE: Image Synthesis by Image-Guided Model Inversion

CVPR 2021poster

Synthesizing variations of a specific reference image with semantically valid content is an important task in terms of personalized generation as well as for data augmentation. In this work, we propose an inversion based method, denoted as IMAge-Guided model INvErsion (IMAGINE), to generate high-qua…

Cited by 37PDFScholar
2020

Don't Judge an Object by Its Context: Learning to Overcome Contextual Bias

CVPR 2020oral

Existing models often leverage co-occurrences between objects and their context to improve recognition accuracy. However, strongly relying on context risks a model's generalizability, especially when typical co-occurrence patterns are absent. This work focuses on addressing such contextual biases to…

Cited by 141PDFScholar
2020

Elastic-InfoGAN: Unsupervised Disentangled Representation Learning in Class-Imbalanced Data

NeurIPS 2020poster

We propose a novel unsupervised generative model that learns to disentangle object identity from other low-level aspects in class-imbalanced data. We first investigate the issues surrounding the assumptions about uniformity made by InfoGAN, and demonstrate its ineffectiveness to properly disentangle…

2020

MixNMatch: Multifactor Disentanglement and Encoding for Conditional Image Generation

CVPR 2020poster

We present MixNMatch, a conditional generative model that learns to disentangle and encode background, object pose, shape, and texture from real images with minimal supervision, for mix-and-match image generation. We build upon FineGAN, an unconditional generative model, to learn the desired disenta…

Cited by 101PDFcodeScholar
2019

FineGAN: Unsupervised Hierarchical Disentanglement for Fine-Grained Object Generation and Discovery

CVPR 2019oral

We propose FineGAN, a novel unsupervised GAN framework, which disentangles the background, object shape, and object appearance to hierarchically generate images of fine-grained object categories. To disentangle the factors without supervision, our key idea is to use information theory to associate e…

Cited by 177PDFcodeScholar
2019

You Reap What You Sow: Using Videos to Generate High Precision Object Proposals for Weakly-Supervised Object Detection

CVPR 2019poster

We propose a novel way of using videos to obtain high precision object proposals for weakly-supervised object detection. Existing weakly-supervised detection approaches use off-the-shelf proposal methods like edge boxes or selective search to obtain candidate boxes. These methods provide high recal…

Cited by 47PDFcodeScholar
2018

DOCK: Detecting Objects by transferring Common-sense Knowledge

ECCV 2018poster

We present a scalable approach for Detecting Objects by transferring Common-sense Knowledge (DOCK) from source to target categories. In our setting, the training data for the source categories have bounding box annotations, while those for the target categories only have image-level annotations. Cur…

Cited by 42SourcePDFScholar
2017

Hide-And-Seek: Forcing a Network to Be Meticulous for Weakly-Supervised Object and Action Localization

ICCV 2017poster

We propose 'Hide-and-Seek', a weakly-supervised framework that aims to improve object localization in images and action localization in videos. Most existing weakly-supervised methods localize only the most discriminative parts of an object rather than all relevant parts, which leads to suboptimal p…

Cited by 739PDFScholar
2017

Identifying First-Person Camera Wearers in Third-Person Videos

CVPR 2017poster

We consider scenarios in which we wish to perform joint scene understanding, object tracking, activity recognition, and other tasks in scenarios in which multiple people are wearing body-worn cameras while a third-person static camera also captures the scene. To do this, we need to establ…

Cited by 77PDFScholar
2016

Track and Transfer: Watching Videos to Simulate Strong Human Supervision for Weakly-Supervised Object Detection

CVPR 2016poster

The status quo approach to training object detectors requires expensive bounding box annotations. Our framework takes a markedly different direction: we transfer tracked object boxes from weakly-labeled videos to weakly-labeled images to automatically generate pseudo ground-truth boxes, which repla…

Cited by 80PDFScholar