← Search

Mohsen Gholami

6 accepted papers

2026

Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes

ICLR 2026poster

Understanding 3D spatial relationships remains a major limitation of current Vision-Language Models (VLMs). Prior work has addressed this issue by creating spatial question-answering (QA) datasets based on single images or indoor videos. However, real-world embodied AI agents—such as robots and self…

Cited by 0SourcecodeScholar
2025

CASP: Compression of Large Multimodal Models Based on Attention Sparsity

CVPR 2025highlight

In this work, we propose an extreme compression technique for Large Multimodal Models (LMMs). While previous studies have explored quantization as an efficient post-training compression method for Large Language Models (LLMs), low-bit compression for multimodal models remains under-explored. The red…

2024

GOLD: Generalized Knowledge Distillation via Out-of-Distribution-Guided Language Data Generation

NAACL 2024findings

Knowledge distillation from LLMs is essential for the efficient deployment of language models. Prior works have proposed data generation using LLMs for preparing distilled models. We argue that generating data with LLMs is prone to sampling mainly from the center of original content distribution. Th…

Cited by 4SourcePDFScholar
2023

ETran: Energy-Based Transferability Estimation

ICCV 2023poster

This paper addresses the problem of ranking pre-trained models for object detection and image classification. Selecting the best pre-trained model by fine-tuning is an expensive and time-consuming task. Previous works have proposed transferability estimation based on features extracted by the pre-tr…

Cited by 17PDFcodeScholar
2022

AdaptPose: Cross-Dataset Adaptation for 3D Human Pose Estimation by Learnable Motion Generation

CVPR 2022poster

This paper addresses the problem of cross-dataset generalization of 3D human pose estimation models. Testing a pre-trained 3D pose estimator on a new dataset results in a major performance drop. Previous methods have mainly addressed this problem by improving the diversity of the training data. We a…

Cited by 51PDFcodeScholar