ACL 2023industry2 citations

Referring to Screen Texts with Voice Assistants

Shruti Bhargava, Anand Dhoot, Ing-marie Jonsson, Hoang Long Nguyen, Alkesh Patel, Hong Yu, Vincent Renkens

Abstract

Voice assistants help users make phone calls, send messages, create events, navigate and do a lot more. However assistants have limited capacity to understand their users’ context. In this work, we aim to take a step in this direction. Our work dives into a new experience for users to refer to phone numbers, addresses, email addresses, urls, and dates on their phone screens. We focus on reference understanding, which is particularly interesting when, similar to visual grounding, there are multiple similar texts on screen. We collect a dataset and propose a lightweight general purpose model for this novel experience. Since consuming pixels directly is expensive, our system is designed to rely only on text extracted from the UI. Our model is modular, offering flexibility, better interpretability and efficient run time memory.

BibTeX
@inproceedings{bhargava-etal-2023-referring,
    title = "Referring to Screen Texts with Voice Assistants",
    author = "Bhargava, Shruti  and
      Dhoot, Anand  and
      Jonsson, Ing-marie  and
      Nguyen, Hoang Long  and
      Patel, Alkesh  and
      Yu, Hong  and
      Renkens, Vincent",
    editor = "Sitaram, Sunayana  and
      Beigman Klebanov, Beata  and
      Williams, Jason D",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 5: Industry Track)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-industry.72/",
    doi = "10.18653/v1/2023.acl-industry.72",
    pages = "752--762"
}
Referring to Screen Texts with Voice Assistants · ACL 2023