ICASSP 2025accepted0 citations

UML: A Unified Multimodal Learning Framework for Cataract Postoperative Visual Acuity Prediction with Uncertain Missing Modalities

Tongyu Yang, Qian Zhou, Hua Zou, Haifeng Jiang, Yong Wang

Abstract

Cataracts are the leading cause of blindness worldwide, with surgery as the only effective treatment. Accurate prediction of Best Corrected Visual Acuity (BCVA) is crucial for surgical planning. In this paper, we propose a novel Unified Multimodal Learning (UML) framework for BCVA prediction with uncertain missing modalities. Unlike existing methods that apply generic encoders and overlook critical image variability, UML leverages medical priors to enhance feature extraction through three modules: central concave region enhancement, OCT re-weighting, and multi-scale attention. To manage missing modality uncertainty, we design a missing modality mask fusion network using an attentional mask for unified feature fusion. Additionally, an auxiliary diagnostic text-image contrastive learning task is introduced to further refine image features. UML achieves state-of-the-art performance with a mean absolute error (MAE) of 0.0457 and 96.25% predictions fall within an error of ± 0.10 LogMAR. Codes are available at https://github.com/yty9941/Eyer-BCVA

BibTeX
@inproceedings{icassp2025_umlaunifiedmulti,
  title = {UML: A Unified Multimodal Learning Framework for Cataract Postoperative Visual Acuity Prediction with Uncertain Missing Modalities},
  author = {Tongyu Yang and Qian Zhou and Hua Zou and Haifeng Jiang and Yong Wang},
  booktitle = {ICASSP 2025},
  year = {2025}
}