← Search

Mohammad Akbari

10 accepted papers

2026

Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes

ICLR 2026poster

Understanding 3D spatial relationships remains a major limitation of current Vision-Language Models (VLMs). Prior work has addressed this issue by creating spatial question-answering (QA) datasets based on single images or indoor videos. However, real-world embodied AI agents—such as robots and self…

Cited by 0SourcecodeScholar
2025

CASP: Compression of Large Multimodal Models Based on Attention Sparsity

CVPR 2025highlight

In this work, we propose an extreme compression technique for Large Multimodal Models (LMMs). While previous studies have explored quantization as an efficient post-training compression method for Large Language Models (LLMs), low-bit compression for multimodal models remains under-explored. The red…

2025

DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models

CVPR 2025poster

Large Multimodal Models (LMMs) have emerged as powerful models capable of understanding various data modalities, including text, images, and videos. LMMs encode both text and visual data into tokens that are then combined and processed by an integrated Large Language Model (LLM). Including visual to…

2025

Task-Agnostic Language Model Watermarking via High Entropy Passthrough Layers

AAAI 2025technical

In the era of costly pre-training of large language models, ensuring the intellectual property rights of model owners, and insuring that said models are responsibly deployed, is becoming increasingly important. To this end, we propose model watermarking via passthrough layers, which are added to exi…

Cited by 0SourcePDFScholar
2024

GOLD: Generalized Knowledge Distillation via Out-of-Distribution-Guided Language Data Generation

NAACL 2024findings

Knowledge distillation from LLMs is essential for the efficient deployment of language models. Prior works have proposed data generation using LLMs for preparing distilled models. We argue that generating data with LLMs is prone to sampling mainly from the center of original content distribution. Th…

Cited by 4SourcePDFScholar
2023

ETran: Energy-Based Transferability Estimation

ICCV 2023poster

This paper addresses the problem of ranking pre-trained models for object detection and image classification. Selecting the best pre-trained model by fine-tuning is an expensive and time-consuming task. Previous works have proposed transferability estimation based on features extracted by the pre-tr…

Cited by 17PDFcodeScholar
2022

E-LANG: Energy-Based Joint Inferencing of Super and Swift Language Models

ACL 2022long

Building huge and highly capable language models has been a trend in the past years. Despite their great performance, they incur high computational cost. A common solution is to apply model compression or choose light-weight architectures, which often need a separate fixed-size model for each desira…

Cited by 11SourcePDFScholar
2021

Learned Bi-Resolution Image Coding using Generalized Octave Convolutions

AAAI 2021technical

Learned image compression has recently shown the potential to outperform the standard codecs. State-of-the-art rate-distortion (R-D) performance has been achieved by context-adaptive entropy coding approaches in which hyperprior and autoregressive models are jointly utilized to effectively capture t…

Cited by 20SourcePDFScholar