← Search

Pinar Yanardag

22 accepted papers

2026

DTG-Restore: Training-Free Diffusion Refinement for Generative Video Super-Resolution

CVPR 2026

Recent progress in video diffusion models has enabled remarkable generative fidelity, yet leveraging these priors for restoration remains limited by the strong coupling between conditional and unconditional branches in standard classifier-free guidance. We introduce a training-free framework that en

Cited by 0SourceScholar
2026

Diverse Video Generation with Determinantal Point Process-Guided Policy Optimization

CVPR 2026

While recent text-to-video (T2V) diffusion models have achieved impressive quality and prompt alignment, they often produce low-diversity outputs when sampling multiple videos from a single text prompt. We tackle this challenge by formulating it as a set-level policy optimization problem, with the g

Cited by 0SourceScholar
2026

Infinity-RoPE: Action-Controllable Infinite Video Generation Emerges From Autoregressive Self-Rollout

CVPR 2026

Current autoregressive video diffusion models are constrained by three core bottlenecks: (i) the finite temporal horizon imposed by the base model's 3D Rotary Positional Embedding (3D-RoPE), (ii) slow prompt responsiveness in maintaining fine-grained action control during long-form rollouts, and (ii

Cited by 0SourcecodeScholar
2026

MotionFlow: Attention-Driven Motion Transfer in Video Diffusion Models

AAAI 2026technical

Text-to-video models have demonstrated impressive capabilities in producing diverse video content, yet often lack fine-grained control over motion. We address the problem of motion transfer: given a source video and a target text prompt, generate a new video that preserves the source motion while ma

Cited by 0SourcePDFScholar
2026

Plot’n Polish: Zero-Shot Story Visualization and Disentangled Editing with Text-to-Image Diffusion Models

AAAI 2026technical

Text-to-image diffusion models have demonstrated significant capabilities to generate diverse and detailed visuals in various domains, and story visualization is emerging as a particularly promising application. However, as their use in real-world creative domains increases, the need for providing e

Cited by 0SourcePDFScholar
2025

CREA: A Collaborative Multi-Agent Framework for Creative Image Editing and Generation

NeurIPS 2025poster

Creativity in AI imagery remains a fundamental challenge, requiring not only the generation of visually compelling content but also the capacity to add novel, expressive, and artistically rich transformations to images. Unlike conventional editing tasks that rely on direct prompt-based modifications…

Cited by 0SourceScholar
2025

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features

ICML 2025oral

Do the rich representations of multi-modal diffusion transformers (DiTs) exhibit unique properties that enhance their interpretability? We introduce ConceptAttention, a novel method that leverages the expressive power of DiT attention layers to generate high-quality saliency maps that precisely loca…

2025

Contrastive Test-Time Composition of Multiple LoRA Models for Image Generation

ICCV 2025poster

Low-Rank Adaptation (LoRA) has emerged as a powerful and popular technique for personalization, enabling efficient adaptation of pre-trained image generation models for specific tasks without comprehensive retraining. While employing individual pre-trained LoRA models excels at representing single c…

Cited by 0SourcePDFScholar
2025

Explaining in Diffusion: Explaining a Classifier with Diffusion Semantics

CVPR 2025poster

Classifiers are important components in many computer vision tasks, serving as the foundational backbone of a wide variety of models employed across diverse applications. However, understanding the decision-making process of classifiers remains a significant challenge. We propose DiffEx, a novel me…

Cited by 0SourcePDFScholar
2025

LoRACLR: Contrastive Adaptation for Customization of Diffusion Models

CVPR 2025poster

Recent advances in text-to-image customization have enabled high-fidelity, context-rich generation of personalized images, allowing specific concepts to appear in a variety of scenarios. However, current methods struggle with combining multiple personalized models, often leading to attribute entangl…

Cited by 0SourcePDFScholar
2025

LoRAShop: Training-Free Multi-Concept Image Generation and Editing with Rectified Flow Transformers

NeurIPS 2025spotlight

We introduce LoRAShop, the first framework for multi-concept image generation and editing with LoRA models. LoRAShop builds on a key observation about the feature interaction patterns inside Flux-style diffusion transformers: concept-specific transformer features activate spatially coherent regions…

Cited by 0SourceScholar
2025

LoRAverse: A Submodular Framework to Retrieve Diverse Adapters for Diffusion Models

ICCV 2025poster

Low-rank Adaptation (LoRA) models have revolutionized the personalization of pre-trained diffusion models by enabling fine-tuning through low-rank, factorized weight matrices specifically optimized for attention layers. These models facilitate the generation of highly customized content across a var…

Cited by 0SourcePDFScholar
2025

Personalized Image Editing in Text-to-Image Diffusion Models via Collaborative Direct Preference Optimization

NeurIPS 2025poster

Text-to-image (T2I) diffusion models have made remarkable strides in generating and editing high-fidelity images from text. Yet, these models remain fundamentally generic, failing to adapt to the nuanced aesthetic preferences of individual users. In this work, we present the first framework for pers…

Cited by 0SourceScholar
2024

CONFORM: Contrast is All You Need for High-Fidelity Text-to-Image Diffusion Models

CVPR 2024poster

Images produced by text-to-image diffusion models might not always faithfully represent the semantic intent of the provided text prompt where the model might overlook or entirely fail to produce certain objects. While recent studies propose various solutions they often require customly tailored func…

Cited by 21SourcePDFScholar
2024

NoiseCLR: A Contrastive Learning Approach for Unsupervised Discovery of Interpretable Directions in Diffusion Models

CVPR 2024poster

Generative models have been very popular in the recent years for their image generation capabilities. GAN-based models are highly regarded for their disentangled latent space which is a key feature contributing to their success in controlled image editing. On the other hand diffusion models have eme…

Cited by 17SourcePDFScholar
2024

RAVE: Randomized Noise Shuffling for Fast and Consistent Video Editing with Diffusion Models

CVPR 2024highlight

Recent advancements in diffusion-based models have demonstrated significant success in generating images from text. However video editing models have not yet reached the same level of visual quality and user control. To address this we introduce RAVE a zero-shot video editing method that leverages p…

2024

Stylebreeder: Exploring and Democratizing Artistic Styles through Text-to-Image Models

NeurIPS 2024poster

Text-to-image models are becoming increasingly popular, revolutionizing the landscape of digital art creation by enabling highly detailed and creative visual content generation. These models have been widely employed across various domains, particularly in art generation, where they facilitate a bro…

2022

FairStyle: Debiasing StyleGAN2 with Style Channel Manipulations

ECCV 2022poster

"Recent advances in generative adversarial networks have shown that it is possible to generate high-resolution and hyperrealistic images. However, the images produced by GANs are only as fair and representative as the datasets on which they are trained. In this paper, we propose a method for directl…

2021

LatentCLR: A Contrastive Learning Approach for Unsupervised Discovery of Interpretable Directions

ICCV 2021poster

Recent research has shown that it is possible to find interpretable directions in the latent spaces of pre-trained Generative Adversarial Networks (GANs). These directions enable controllable image generation and support a wide range of semantic editing operations, such as zoom or rotation. The disc…

Cited by 77PDFcodeScholar