← Search

Jaskirat Singh

13 accepted papers

2026

What matters for Representation Alignment: Global Information or Spatial Structure?

ICLR 2026poster

Representation alignment helps generation by distilling representations from a pretrained vision encoder to intermediate diffusion features. We investigate a fundamental question - `what aspect of the target representation matters for generation, its global information (measured by Imagenet1K accura…

Cited by 0SourcecodeScholar
2025

OpenHands: An Open Platform for AI Software Developers as Generalist Agents

ICLR 2025poster

Software is one of the most powerful tools that we humans have at our disposal; it allows a skilled programmer to interact with the world in complex and profound ways. At the same time, thanks to improvements in large language models (LLMs), there has also been a rapid development in AI agents that…

Cited by 32SourcePDFScholar
2025

Probability Density Geodesics in Image Diffusion Latent Space

CVPR 2025poster

Diffusion models indirectly estimate the probability density over a data space, which can be used to study its structure. In this work, we show that geodesics can be computed in diffusion latent space, where the norm induced by the spatially-varying inner product is inversely proportional to the pro…

2025

REPA-E: Unlocking VAE for End-to-End Tuning of Latent Diffusion Transformers

ICCV 2025poster

In this paper we tackle a fundamental question: "Can we train latent diffusion models together with the variational auto-encoder (VAE) tokenizer in an end-to-end manner?" Traditional deep-learning wisdom dictates that end-to-end training is often preferable when possible. However, for latent diffusi…

2025

Storybooth: Training-Free Multi-Subject Consistency for Improved Visual Storytelling

ICLR 2025poster

Consistent text-to-image generation depicting the *same* subjects across different images has gained significant recent attention due to its widespread applications in the fields of visual-storytelling and multiple-shot video generation. While remarkable, existing methods often require costly finet…

Cited by 0SourcePDFScholar
2025

Vec2Face: Scaling Face Dataset Generation with Loosely Constrained Vectors

ICLR 2025poster

This paper studies how to synthesize face images of non-existent persons, to create a dataset that allows effective training of face recognition (FR) models. Besides generating realistic face images, two other important goals are: 1) the ability to generate a large number of distinct identities (int…

Cited by 4SourcePDFScholar
2024

IMPUS: Image Morphing with Perceptually-Uniform Sampling Using Diffusion Models

ICLR 2024poster

We present a diffusion-based image morphing approach with perceptually-uniform sampling (IMPUS) that produces smooth, direct and realistic interpolations given an image pair. The embeddings of two images may lie on distinct conditioned distributions of a latent diffusion model, especially when they…

2024

SmartMask: Context Aware High-Fidelity Mask Generation for Fine-grained Object Insertion and Layout Control

CVPR 2024poster

The field of generative image inpainting and object insertion has made significant progress with the recent advent of latent diffusion models. Utilizing a precise object mask can greatly enhance these applications. However due to the challenges users encounter in creating high-fidelity masks there i…

Cited by 9SourcePDFScholar
2023

Divide, Evaluate, and Refine: Evaluating and Improving Text-to-Image Alignment with Iterative VQA Feedback

NeurIPS 2023poster

The field of text-conditioned image generation has made unparalleled progress with the recent advent of latent diffusion models. While revolutionary, as the complexity of given text input increases, the current state of art diffusion models may still fail in generating images that accurately convey…

Cited by 23SourcePDFScholar
2023

High-Fidelity Guided Image Synthesis With Latent Diffusion Models

CVPR 2023poster

Controllable image synthesis with user scribbles has gained huge public interest with the recent advent of text-conditioned latent diffusion models. The user scribbles control the color composition while the text prompt provides control over the overall image semantics. However, we find that prior w…

2022

Intelli-Paint: Towards Developing More Human-Intelligible Painting Agents

ECCV 2022poster

"Stroke based rendering methods have recently become a popular solution for the generation of stylized paintings. However, the current research in this direction is focused mainly on the improvement of final canvas quality, and thus often fails to consider the intelligibility of the generated painti…

Cited by 18SourcePDFScholar
2022

Paint2Pix: Interactive Painting Based Progressive Image Synthesis and Editing

ECCV 2022poster

"Controllable image synthesis with user scribbles is a topic of keen interest in the computer vision community. In this paper, for the first time we study the problem of photorealistic image synthesis from incomplete and primitive human paintings. In particular, we propose a novel approach paint2pix…

2021

Combining Semantic Guidance and Deep Reinforcement Learning for Generating Human Level Paintings

CVPR 2021poster

Generation of stroke-based non-photorealistic imagery, is an important problem in the computer vision community. As an endeavor in this direction, substantial recent research efforts have been focused on teaching machines "how to paint", in a manner similar to a human painter. However, the applicabi…

Cited by 32PDFcodeScholar