← Search

Muhammad Akhtar Munir

10 accepted papers

2026

TerraFM: A Scalable Foundation Model for Unified Multisensor Earth Observation

ICLR 2026poster

Modern Earth observation (EO) increasingly leverages deep learning to harness the scale and diversity of satellite imagery across sensors and regions. While recent foundation models have demonstrated promising generalization across EO tasks, many remain limited by the scale, geographical coverage, a…

Cited by 0SourcecodeScholar
2026

Towards Calibrating Prompt Tuning of Vision- Language Models

CVPR 2026

Prompt tuning of large-scale vision-language models such as CLIP enables efficienttask adaptation without updating model weights. However, it often leads to poorconfidence calibration and unreliable predictive uncertainty. We address thisproblem by proposing a calibration framework that enhances pre

Cited by 0SourcecodeScholar
2025

EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues

CVPR 2025poster

Automated analysis of vast Earth observation data via interactive Vision-Language Models (VLMs) can unlock new opportunities for environmental monitoring, disaster response, and resource management. Existing generic VLMs do not perform well on Remote Sensing data, while the recent Geo-spatial VLMs r…

2025

GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks

ICCV 2025poster

While numerous recent benchmarks focus on evaluating generic Vision-Language Models (VLMs), they do not effectively address the specific challenges of geospatial applications.Generic VLM benchmarks are not designed to handle the complexities of geospatial data, an essential component for application…

2025

O-TPT: Orthogonality Constraints for Calibrating Test-time Prompt Tuning in Vision-Language Models

CVPR 2025highlight

Test-time prompt tuning for vision-language models (VLMs) is getting attention because of their ability to learn with unlabeled data without fine-tuning. Although test-time prompt tuning methods for VLMs can boost accuracy, the resulting models tend to demonstrate poor calibration, which casts doubt…

2024

Improving Single Domain-Generalized Object Detection: A Focus on Diversification and Alignment

CVPR 2024poster

In this work we tackle the problem of domain generalization for object detection specifically focusing on the scenario where only a single source domain is available. We propose an effective approach that involves two key steps: diversifying the source domain and aligning detections based on class p…

2023

Bridging Precision and Confidence: A Train-Time Loss for Calibrating Object Detection

CVPR 2023poster

Deep neural networks (DNNs) have enabled astounding progress in several vision-based problems. Despite showing high predictive accuracy, recently, several works have revealed that they tend to provide overconfident predictions and thus are poorly calibrated. The majority of the works addressing the…

2023

Cal-DETR: Calibrated Detection Transformer

NeurIPS 2023poster

Albeit revealing impressive predictive performance for several computer vision tasks, deep neural networks (DNNs) are prone to making overconfident predictions. This limits the adoption and wider utilization of DNNs in many safety-critical applications. There have been recent efforts toward calibrat…

2022

Towards Improving Calibration in Object Detection Under Domain Shift

NeurIPS 2022accept

With deep neural network based solution more readily being incorporated in real-world applications, it has been pressing requirement that predictions by such models, especially in safety-critical environments, be highly accurate and well-calibrated. Although some techniques addressing DNN calibrati…

Cited by 22SourcePDFScholar
2021

SSAL: Synergizing between Self-Training and Adversarial Learning for Domain Adaptive Object Detection

NeurIPS 2021poster

We study adapting trained object detectors to unseen domains manifesting significant variations of object appearance, viewpoints and backgrounds. Most current methods align domains by either using image or instance-level feature alignment in an adversarial fashion. This often suffers due to the pres…

Cited by 84SourcePDFScholar