IJCAI 2023poster5 citations

Quality-agnostic Image Captioning to Safely Assist People with Vision Impairment

Lu Yu, Malvina Nikandrou, Jiali Jin, Verena Rieser

Abstract

Automated image captioning has the potential to be a useful tool for people with vision impairments. Images taken by this user group are often noisy, which leads to incorrect and even unsafe model predictions. In this paper, we propose a quality-agnostic framework to improve the performance and robustness of image captioning models for visually impaired people. We address this problem from three angles: data, model, and evaluation. First, we show how data augmentation techniques for generating synthetic noise can address data sparsity in this domain. Second, we enhance the robustness of the model by expanding a state-of-the-art model to a dual network architecture, using the augmented data and leveraging different consistency losses. Our results demonstrate increased performance, e.g. an absolute improvement of 2.15 on CIDEr, compared to state-of-the-art image captioning networks, as well as increased robustness to noise with up to 3 points improvement on CIDEr in more noisy settings. Finally, we evaluate the prediction reliability using confidence calibration on images with different difficulty / noise levels, showing that our models perform more reliably in safety-critical situations. The improved model is part of an assisted living application, which we develop in partnership with the Royal National Institute of Blind People.

AI for Good: Humans and AIAI for Good: Uncertainty in AI
BibTeX
@inproceedings{ijcai2023p697,
  title     = {Quality-agnostic Image Captioning to Safely Assist People with Vision Impairment},
  author    = {Yu, Lu and Nikandrou, Malvina and Jin, Jiali and Rieser, Verena},
  booktitle = {Proceedings of the Thirty-Second International Joint Conference on
               Artificial Intelligence, {IJCAI-23}},
  publisher = {International Joint Conferences on Artificial Intelligence Organization},
  editor    = {Edith Elkind},
  pages     = {6281--6289},
  year      = {2023},
  month     = {8},
  note      = {AI for Good},
  doi       = {10.24963/ijcai.2023/697},
  url       = {https://doi.org/10.24963/ijcai.2023/697},
}
Quality-agnostic Image Captioning to Safely Assist People with Vision Impairment · IJCAI 2023