← Search

Isma Hadji

12 accepted papers

2026

Hierarchical Image Tokenization for Multi-Scale Image Super Resolution

ICML 2026poster

We introduce a multi-scale Image Super Resolution (ISR) method building on recent advances in Visual Auto-Regressive (VAR) modeling. Recently, VAR models challenged the dominance of diffusion-based models by adopting a next-scale prediction paradigm. Specifically, VAR models iteratively estimate the…

Cited by 0SourceScholar
2026

Restore, Assess, Repeat: A Unified Framework for Iterative Image Restoration

CVPR 2026

Image restoration aims to recover high quality images from inputs degraded by various factors, such as adverse weather, blur, or low light. While recent studies have shown remarkable progress across individual or unified restoration tasks, they still suffer from limited generalization and inefficien

Cited by 0SourceScholar
2025

Edge-SD-SR: Low Latency and Parameter Efficient On-device Super-Resolution with Stable Diffusion via Bidirectional Conditioning

CVPR 2025poster

There has been immense progress recently in the visual quality of Stable Diffusion-based Super Resolution (SD-SR). However, deploying large diffusion models on computationally restricted devices such as mobile phones remains impractical due to the large model size and high latency. This is compounde…

Cited by 0SourcePDFScholar
2025

FAM Diffusion: Frequency and Attention Modulation for High-Resolution Image Generation with Stable Diffusion

CVPR 2025poster

Diffusion models are proficient at generating high-quality images. They are however effective only when operating at the resolution used during training. Inference at a scaled resolution leads to repetitive patterns and structural distortions. Retraining at higher resolutions quickly becomes prohib…

Cited by 1SourcePDFScholar
2023

GePSAn: Generative Procedure Step Anticipation in Cooking Videos

ICCV 2023poster

We study the problem of future step anticipation in procedural videos. Given a video of an ongoing procedural activity, we predict a plausible next procedure step described in rich natural language. While most previous work focus on the problem of data scarcity in procedural video datasets, another…

Cited by 11PDFcodeScholar
2023

StepFormer: Self-Supervised Step Discovery and Localization in Instructional Videos

CVPR 2023poster

Instructional videos are an important resource to learn procedural tasks from human demonstrations. However, the instruction steps in such videos are typically short and sparse, with most of the video being irrelevant to the procedure. This motivates the need to temporally localize the instruction s…

Cited by 29SourcePDFScholar
2022

Flow Graph to Video Grounding for Weakly-Supervised Multi-step Localization

ECCV 2022poster

"In this work, we consider the problem of weakly-supervised multi-step localization in instructional videos. An established approach to this problem is to rely on a given list of steps. However, in reality, there is often more than one way to execute a procedure successfully, by following the set of…

2022

P3IV: Probabilistic Procedure Planning From Instructional Videos With Weak Supervision

CVPR 2022oral

In this paper, we study the problem of procedure planning in instructional videos. Here, an agent must produce a plausible sequence of actions that can transform the environment from a given start to a desired goal state. When learning procedure planning from instructional videos, most recent work l…

Cited by 52PDFcodeScholar
2021

Drop-DTW: Aligning Common Signal Between Sequences While Dropping Outliers

NeurIPS 2021poster

In this work, we consider the problem of sequence-to-sequence alignment for signals containing outliers. Assuming the absence of outliers, the standard Dynamic Time Warping (DTW) algorithm efficiently computes the optimal alignment between two (generally) variable-length sequences. While DTW is ro…

Cited by 59SourcePDFScholar
2021

Representation Learning via Global Temporal Alignment and Cycle-Consistency

CVPR 2021poster

We introduce a weakly supervised method for representation learning based on aligning temporal sequences (e.g., videos) of the same process (e.g., human action). The main idea is to use the global temporal ordering of latent correspondences across sequence pairs as a supervisory signal. In particula…

Cited by 71PDFcodeScholar
2018

A New Large Scale Dynamic Texture Dataset with Application to ConvNet Understanding

ECCV 2018poster

This paper introduces a new large scale dynamic texture dataset. The dataset is provided with two complementary organizations, one based on dynamics independent of spatial appearance and one based on spatial appearance independent of dynamics. With over 10,000 videos, the proposed Dynamic Texture Da…

Cited by 39SourcePDFScholar