← Search

Ruchika Chavhan

9 accepted papers

2026

NanoFLUX: Distillation-Driven Compression of Large Text-to-Image Generation Models for Mobile Devices

ICML 2026poster

While large-scale text-to-image diffusion models continue to improve in visual quality, their increasing scale has widened the gap between state-of-the-art models and on-device solutions. To address this gap, we introduce NanoFLUX, a **2.4B** text-to-image flow-matching model distilled from **17B** …

Cited by 0SourceScholar
2026

RFDM: Residual Flow Diffusion Models for Video Editing

CVPR 2026

Instructional video editing applies edits to an input video using only text prompts, enabling intuitive natural-language control. Despite the rapid progress, most methods still require fixed-length inputs and substantial compute. Meanwhile, autoregressive video generation enables efficient variable-

Cited by 0SourcecodeScholar
2025

ConceptPrune: Concept Editing in Diffusion Models via Skilled Neuron Pruning

ICLR 2025poster

While large-scale text-to-image diffusion models have demonstrated impressive image-generation capabilities, there are significant concerns about their potential misuse for generating unsafe content, violating copyright, and perpetuating societal biases. Recently, the text-to-image generation commun…

2025

EDiT: Efficient Diffusion Transformers with Linear Compressed Attention

ICCV 2025poster

Diffusion Transformers (DiTs) have emerged as a leading architecture for text-to-image synthesis, producing high-quality and photorealistic images. However, the quadratic scaling properties of the attention in DiTs hinder image generation with higher resolution or devices with limited resources. Thi…

Cited by 0SourcePDFScholar
2025

Upcycling Text-to-Image Diffusion Models for Multi-Task Capabilities

ICML 2025poster

Text-to-image synthesis has witnessed remarkable advancements in recent years. Many attempts have been made to adopt text-to-image models to support multiple tasks. However, existing approaches typically require resource-intensive re-training or additional parameters to accommodate for the new tasks…

Cited by 0SourcePDFScholar
2024

Fool Your (Vision and) Language Model with Embarrassingly Simple Permutations

ICML 2024poster

Large language and vision-language models are rapidly being deployed in practice thanks to their impressive capabilities in instruction following, in-context learning, and so on. This raises an urgent need to carefully analyse their robustness so that stakeholders can understand if and when such mod…

2023

Amortised Invariance Learning for Contrastive Self-Supervision

ICLR 2023poster

Contrastive self-supervised learning methods famously produce high quality transferable representations by learning invariances to different data augmentations. Invariances established during pre-training can be interpreted as strong inductive biases. However these may or may not be helpful, dependi…

2023

Meta Omnium: A Benchmark for General-Purpose Learning-To-Learn

CVPR 2023poster

Meta-learning and other approaches to few-shot learning are widely studied for image recognition, and are increasingly applied to other vision tasks such as pose estimation and dense prediction. This naturally raises the question of whether there is any few-shot meta-learning algorithm capable of ge…