← Search

Mehdi Noroozi

11 accepted papers

2026

NanoFLUX: Distillation-Driven Compression of Large Text-to-Image Generation Models for Mobile Devices

ICML 2026poster

While large-scale text-to-image diffusion models continue to improve in visual quality, their increasing scale has widened the gap between state-of-the-art models and on-device solutions. To address this gap, we introduce NanoFLUX, a **2.4B** text-to-image flow-matching model distilled from **17B** …

Cited by 0SourceScholar
2026

RFDM: Residual Flow Diffusion Models for Video Editing

CVPR 2026

Instructional video editing applies edits to an input video using only text prompts, enabling intuitive natural-language control. Despite the rapid progress, most methods still require fixed-length inputs and substantial compute. Meanwhile, autoregressive video generation enables efficient variable-

Cited by 0SourcecodeScholar
2025

EDiT: Efficient Diffusion Transformers with Linear Compressed Attention

ICCV 2025poster

Diffusion Transformers (DiTs) have emerged as a leading architecture for text-to-image synthesis, producing high-quality and photorealistic images. However, the quadratic scaling properties of the attention in DiTs hinder image generation with higher resolution or devices with limited resources. Thi…

Cited by 0SourcePDFScholar
2025

Edge-SD-SR: Low Latency and Parameter Efficient On-device Super-Resolution with Stable Diffusion via Bidirectional Conditioning

CVPR 2025poster

There has been immense progress recently in the visual quality of Stable Diffusion-based Super Resolution (SD-SR). However, deploying large diffusion models on computationally restricted devices such as mobile phones remains impractical due to the large model size and high latency. This is compounde…

Cited by 0SourcePDFScholar
2025

Upcycling Text-to-Image Diffusion Models for Multi-Task Capabilities

ICML 2025poster

Text-to-image synthesis has witnessed remarkable advancements in recent years. Many attempts have been made to adopt text-to-image models to support multiple tasks. However, existing approaches typically require resource-intensive re-training or additional parameters to accommodate for the new tasks…

Cited by 0SourcePDFScholar
2022

Ranking Info Noise Contrastive Estimation: Boosting Contrastive Learning via Ranked Positives

AAAI 2022technical

This paper introduces Ranking Info Noise Contrastive Estimation (RINCE), a new member in the family of InfoNCE losses that preserves a ranked ordering of positive samples. In contrast to the standard InfoNCE loss, which requires a strict binary separation of the training pairs into similar and dissi…

2022

Unified Fully and Timestamp Supervised Temporal Action Segmentation via Sequence to Sequence Translation

ECCV 2022poster

"This paper introduces a unified framework for video action segmentation via sequence to sequence (seq2seq) translation in a fully and timestamp supervised setup. In contrast to current state-of-the-art frame-level prediction methods, we view action segmentation as a seq2seq translation task, i.e.,…

2021

3D CNNs With Adaptive Temporal Feature Resolutions

CVPR 2021poster

While state-of-the-art 3D Convolutional Neural Networks (CNN) achieve very good results on action recognition datasets, they are computationally very expensive and require many GFLOPs. While the GFLOPs of a 3D CNN can be decreased by reducing the temporal feature resolution within the network, there…

Cited by 39PDFcodeScholar
2021

Long Short View Feature Decomposition via Contrastive Video Representation Learning

ICCV 2021poster

Self-supervised video representation methods typically focus on the representation of temporal attributes in videos. However, the role of stationary versus non-stationary attributes is less explored: Stationary features, which remain similar throughout the video, enable the prediction of video-level…

Cited by 45PDFScholar
2018

Boosting Self-Supervised Learning via Knowledge Transfer

CVPR 2018poster

In self-supervised learning one trains a model to solve a so-called pretext task on a dataset without the need for human annotation. The main objective, however, is to transfer this model to a target domain and task. Currently, the most effective transfer strategy is fine-tuning, which restricts one…

Cited by 391SourcePDFScholar