← Search

Basim Azam

4 accepted papers

2026

GHOST: Hallucination-Inducing Image Generation for Multimodal LLMs

ICLR 2026poster

Object hallucination in Multimodal Large Language Models (MLLMs) is a persistent failure mode that causes the model to perceive objects absent in the image. This weakness of MLLMs is currently studied using static benchmarks with fixed visual scenarios, which preempts the possibility of uncovering m…

Cited by 0SourcecodeScholar
2025

DDB: Diffusion Driven Balancing to Address Spurious Correlations

ICCV 2025poster

Deep neural networks trained with Empirical Risk Minimization (ERM) perform well when both training and test data come from the same domain, but they often fail to generalize to out-of-distribution samples. In image classification, these models may rely on spurious correlations that often exist betw…

2025

GO-N3RDet: Geometry Optimized NeRF-enhanced 3D Object Detector

CVPR 2025poster

We propose GO-N3RDet, a scene-geometry optimized multi-view 3D object detector enhanced by neural radiance fields. The key to accurate 3D object detection is in effective voxel representation. However, due to occlusion and lack of 3D information, constructing 3D features from multi-view 2D images is…

2025

Plug-and-Play Interpretable Responsible Text-to-Image Generation via Dual-Space Multi-facet Concept Control

CVPR 2025poster

Ethical issues around text-to-image (T2I) models demand a comprehensive control over the generative content. Existing techniques addressing these issues for responsible T2I models aim for the generated content to be fair and safe (non-violent/explicit). However, these methods remain bounded to han…