← Search

Maitreya Patel

14 accepted papers

2026

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations

CVPR 2026

We introduce an efficient, resolution-agnostic autoregressive (AR) image synthesis approach that generalizes to arbitrary resolutions and aspect ratios, narrowing the gap to diffusion models at scale. At its core is VibeToken, a novel resolution-agnostic 1D Transformer-based image tokenizer that enc

Cited by 0SourcecodeScholar
2025

AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models

EMNLP 2025

Text-to-Image (T2I) models have recently achieved remarkable success in generating images from textual descriptions. However, challenges still persist in accurately rendering complex scenes where actions and interactions form the primary semantic focus. Our key observation in this work is that T2I m

2025

EraseFlow: Learning Concept Erasure Policies via GFlowNet-Driven Alignment

NeurIPS 2025spotlight

Erasing harmful or proprietary concepts from powerful text‑to‑image generators is an emerging safety requirement, yet current ``concept erasure'' techniques either collapse image quality, rely on brittle adversarial losses, or demand prohibitive retraining cycles. We trace these limitations to a myo…

Cited by 0SourceScholar
2025

FlowChef: Steering of Rectified Flow Models for Controlled Generations

ICCV 2025poster

Despite recent advances in Rectified Flow Models (RFMs), unlocking their full potential for controlled generation tasks--such as inverse problems and image editing--remains a significant hurdle. Although RFMs and Diffusion Models (DMs) represent state-of-the-art approaches in generative modeling, th…

2025

RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions

ICCV 2025poster

Despite recent advances in inversion and instruction-based image editing, existing approaches primarily excel at editing single, prominent objects but significantly struggle when applied to complex scenes containing multiple entities. To quantify this gap, we first introduce **`RefEdit-Bench`**, a r…

Cited by 0SourcePDFScholar
2025

VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning

ICLR 2025poster

Multimodal Large Language Models (MLLMs) have become a powerful tool for integrating visual and textual information. Despite their exceptional performance on visual understanding benchmarks, measuring their ability to reason abstractly across multiple images remains a significant challenge. To addre…

Cited by 0SourcePDFScholar
2024

ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models

AAAI 2024technical

The ability to understand visual concepts and replicate and compose these concepts from images is a central goal for computer vision. Recent advances in text-to-image (T2I) models have lead to high definition and realistic image quality generation by learning from large databases of images and their…

2024

ECLIPSE: A Resource-Efficient Text-to-Image Prior for Image Generations

CVPR 2024poster

Text-to-image (T2I) diffusion models notably the unCLIP models (e.g. DALL-E-2) achieve state-of-the-art (SOTA) performance on various compositional T2I benchmarks at the cost of significant computational resources. The unCLIP stack comprises T2I prior and diffusion image decoder. The T2I prior model…

Cited by 21SourcePDFScholar
2024

Precision or Recall? An Analysis of Image Captions for Training Text-to-Image Generation Model

EMNLP 2024finding

Despite advancements in text-to-image models, generating images that precisely align with textual descriptions remains challenging due to misalignment in training data. In this paper, we analyze the critical role of caption precision and recall in text-to-image model training. Our analysis of human-…

2024

TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives

NeurIPS 2024poster

Contrastive Language-Image Pretraining (CLIP) models maximize the mutual information between text and visual modalities to learn representations. This makes the nature of the training data a significant factor in the efficacy of CLIP for downstream tasks. However, the lack of compositional diversity…

2024

WOUAF: Weight Modulation for User Attribution and Fingerprinting in Text-to-Image Diffusion Models

CVPR 2024poster

The rapid advancement of generative models facilitating the creation of hyper-realistic images from textual descriptions has concurrently escalated critical societal concerns such as misinformation. Although providing some mitigation traditional fingerprinting mechanisms fall short in attributing re…

2022

CRIPP-VQA: Counterfactual Reasoning about Implicit Physical Properties via Video Question Answering

EMNLP 2022main

Videos often capture objects, their visible properties, their motion, and the interactions between different objects. Objects also have physical properties such as mass, which the imaging pipeline is unable to directly capture. However, these properties can be estimated by utilizing cues from relati…

2022

Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks

EMNLP 2022main

How well can NLP models generalize to a variety of unseen tasks when provided with task instructions? To address this question, we first introduce Super-NaturalInstructions, a benchmark of 1,616 diverse NLP tasks and their expert-written instructions. Our collection covers 76 distinct task types, in…

2020

Mspec-Net : Multi-Domain Speech Conversion Network

ICASSP 2020accepted

In this paper, we present a multi-domain speech conversion technique by proposing a Multi-domain Speech Conversion Network (MSpeC-Net) architecture for solving the less-explored area of Non-Audible Murmur-to-SPeeCH (NAM2-SPCH) conversion. The murmur produced by the speaker and captured by the NAM mi…

Cited by 0SourceScholar