← Search

Mohsen Ghafoorian

9 accepted papers

2026

Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer

CVPR 2026

Transformer-based video diffusion models (VDMs) deliver state-of-the-art video generation quality but are constrained by the quadratic cost of self-attention, making long sequences and high resolutions computationally expensive. While linear attention offers sub-quadratic complexity, previous approa

Cited by 11SourceScholar
2026

MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models

ICLR 2026poster

Text-to-video diffusion models have enabled high-quality video synthesis, yet often fail to generate temporally coherent and physically plausible motion. A key reason is the models' insufficient understanding of complex motions that natural videos often entail. Recent works tackle this problem by al…

Cited by 0SourceScholar
2026

Neodragon: Mobile Video Generation Using Diffusion Transformer

ICLR 2026poster

We propose Neogradon, a video DiT (Diffusion Transformer) designed to run on a low-power NPU present in devices such as phones and laptop computers. We demonstrate that, despite video transformers' huge memory and compute cost, mobile devices can run these models when carefully optimised for efficie…

Cited by 0SourcecodeScholar
2026

PyramidalWan: On Making Pretrained Video Model Pyramidal for Efficient Inference

CVPR 2026

Recently proposed pyramidal models decompose the conventional forward and backward diffusion processes into multiple stages operating at varying resolutions. These models handle inputs with higher noise levels at lower resolutions, while less noisy inputs are processed at higher resolutions. This hi

Cited by 0SourceScholar
2025

AnyMap: Learning a General Camera Model for Structure-from-Motion with Unknown Distortion in Dynamic Scenes

CVPR 2025poster

Current learning-based Structure-from-Motion (SfM) methods struggle with videos of dynamic scenes from wide-angle cameras. We present AnyMap, a differentiable SfM framework that jointly addresses image distortion and motion estimation. By learning a general implicit camera model without predefined p…

2024

FastCAD: Real-Time CAD Retrieval and Alignment from Scans and Videos

ECCV 2024poster

"Digitising the 3D world into a clean, CAD model-based representation has important applications for augmented reality and robotics. Current state-of-the-art methods are computationally intensive as they individually encode each detected object and optimise CAD alignments in a second stage. In this…

Cited by 2SourcePDFScholar
2023

3D Distillation: Improving Self-Supervised Monocular Depth Estimation on Reflective Surfaces

ICCV 2023poster

Self-supervised monocular depth estimation (SSMDE) aims at predicting the dense depth maps of monocular images, by learning to minimize a photometric loss using spatially neighboring image pairs during training. While SSMDE offers a significant scalability advantage over supervised approaches, it pe…

Cited by 11PDFScholar
2023

DG-Recon: Depth-Guided Neural 3D Scene Reconstruction

ICCV 2023poster

A key challenge in neural 3D scene reconstruction from monocular images is to fuse features back projected from various views without any depth or occlusion information. We address this by leveraging monocular depth priors, which effectively guide the fusion to improve surface prediction and skip ov…

Cited by 15PDFScholar