← Search

Mohammad Mahdi Derakhshani

7 accepted papers

2026

Purrception: Variational Flow Matching for Vector-Quantized Image Generation

ICLR 2026poster

We introduce Purrception, a variational flow matching approach for vector-quantized image generation that provides explicit categorical supervision while maintaining continuous transport dynamics. Our method adapts Variational Flow Matching to vector-quantized latents by learning categorical posteri…

Cited by 0SourceScholar
2025

TULIP: Token-length Upgraded CLIP

ICLR 2025poster

We address the challenge of representing long captions in vision-language models, such as CLIP. By design these models are limited by fixed, absolute positional encodings, restricting inputs to a maximum of 77 tokens and hindering performance on tasks requiring longer descriptions. Although recent w…

2025

TWIST & SCOUT: Grounding Multimodal LLM-Experts by Forget-Free Tuning

ICCV 2025poster

Spatial awareness is key to enable embodied multimodal AI systems. Yet, without vast amounts of spatial supervision, current Multimodal Large Language Models (MLLMs) struggle at this task. In this paper, we introduce TWIST & SCOUT, a framework that equips pre-trained MLLMs with visual grounding abil…

Cited by 0SourcePDFScholar
2024

Any-Shift Prompting for Generalization over Distributions

CVPR 2024poster

Image-language models with prompt learning have shown remarkable advances in numerous downstream vision tasks. Nevertheless conventional prompt learning methods overfit the training distribution and lose the generalization ability on the test distributions. To improve the generalization across vario…

Cited by 17SourcePDFScholar
2023

Bayesian Prompt Learning for Image-Language Model Generalization

ICCV 2023poster

Foundational image-language models have generated considerable interest due to their efficient adaptation to downstream tasks by prompt learning. Prompt learning treats part of the language model input as trainable while freezing the rest, and optimizes an Empirical Risk Minimization objective. Howe…

Cited by 42PDFcodeScholar
2019

Assisted Excitation of Activations: A Learning Technique to Improve Object Detectors

CVPR 2019poster

We present a simple yet effective learning technique that significantly improves mAP of YOLO object detectors without compromising their speed. During network training, we carefully feed in localization information. We excite certain activations in order to help the network learn to better localize…

Cited by 39PDFScholar