MLLM-ITM: Multimodal Large Language Model Promotes Inverse Tone Mapping
Jingchao Peng, Thomas Bashford-Rogers, Haitao Zhao, Kurt Debattista
Abstract
High dynamic range (HDR) imaging is crucial for capturing real-world lighting conditions. HDR imaging is traditionally achieved either by fusing multiple exposure frames or via inverse tone mapping from a single SDR image. However, the multi-exposure HDR method is prone to motion-induced artefacts and imposes demanding hardware requirements, limiting its practical applicability. Traditional inverse tone-mapping techniques primarily rely on pixel-wise regression methods, which ignore semantic scene contexts and thus frequently introduce halo artefacts and structural distortions. To address this limitation, this paper proposes MLLM-ITM, a novel inverse tone-mapping framework incorporating multimodal large language models (MLLMs). MLLM-ITM utilizes cross-modal features extracted from a frozen MLLM to simultaneously encode visual features and semantic understanding. These features are integrated into a downstream HDR reconstruction backbone through lightweight adapters, enabling content-aware dynamic-range expansion. Furthermore, the decoupled design between MLLM and HDR backbone avoids costly fine-tuning of MLLMs and remains model-agnostic, allowing effortless substitution with emerging multimodal architectures. Extensive experiments on public benchmarks demonstrate that the proposed MLLM-ITM achieves state-of-the-art performance compared with existing inverse tone-mapping methods, highlighting the effectiveness of cross-modal semantic priors in enhancing HDR imaging performance.
BibTeX
@inproceedings{ijcai2026_mllmitmmultimoda,
title = {MLLM-ITM: Multimodal Large Language Model Promotes Inverse Tone Mapping},
author = {Jingchao Peng and Thomas Bashford-Rogers and Haitao Zhao and Kurt Debattista},
booktitle = {IJCAI 2026},
year = {2026}
}