← Search

Amir Habibian

4 accepted papers

2026

Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs

ICML 2026poster

Vision-Language Models (VLMs) achieve strong multimodal performance but are costly to deploy, and post-training quantization often causes significant accuracy loss. Despite its potential, quantization-aware training for VLMs remains underexplored. We propose GRACE, a framework unifying knowledge dis…

Cited by 0SourceScholar
2026

MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models

ICLR 2026poster

Text-to-video diffusion models have enabled high-quality video synthesis, yet often fail to generate temporally coherent and physically plausible motion. A key reason is the models' insufficient understanding of complex motions that natural videos often entail. Recent works tackle this problem by al…

Cited by 0SourceScholar
2026

Neodragon: Mobile Video Generation Using Diffusion Transformer

ICLR 2026poster

We propose Neogradon, a video DiT (Diffusion Transformer) designed to run on a low-power NPU present in devices such as phones and laptop computers. We demonstrate that, despite video transformers' huge memory and compute cost, mobile devices can run these models when carefully optimised for efficie…

Cited by 0SourcecodeScholar
2024

Skip-Attention: Improving Vision Transformers by Paying Less Attention

ICLR 2024poster

This work aims to improve the efficiency of vision transformers (ViTs). While ViTs use computationally expensive self-attention operations in every layer, we identify that these operations are highly correlated across layers -- a key redundancy that causes unnecessary computations. Based on this o…

Cited by 34SourcePDFScholar