AAAI 2024technical2 citations

Tell Me What Is Good about This Property: Leveraging Reviews for Segment-Personalized Image Collection Summarization

Monika Wysoczanska, Moran Beladev, Karen Lastmann Assaraf, Fengjun Wang, Ofri Kleinfeld, Gil Amsalem, Hadas Harush Boke

Abstract

Image collection summarization techniques aim to present a compact representation of an image gallery through a carefully selected subset of images that captures its semantic content. When it comes to web content, however, the ideal selection can vary based on the user's specific intentions and preferences. This is particularly relevant at Booking.com, where presenting properties and their visual summaries that align with users' expectations is crucial. To address this challenge, in this work, we consider user intentions in the summarization of property visuals by analyzing property reviews and extracting the most significant aspects mentioned by users. By incorporating the insights from reviews in our visual summaries, we enhance the summaries by presenting the relevant content to a user. Moreover, we achieve it without the need for costly annotations. Our experiments, including human perceptual studies, demonstrate the superiority of our cross-modal approach, which we coin as CrossSummarizer over the no-personalization and image-based clustering baselines.

BibTeX
@article{Wysoczanska_Beladev_Lastmann Assaraf_Wang_Kleinfeld_Amsalem_Harush Boke_2024, title={Tell Me What Is Good about This Property: Leveraging Reviews for Segment-Personalized Image Collection Summarization}, volume={38}, url={https://ojs.aaai.org/index.php/AAAI/article/view/30339}, DOI={10.1609/aaai.v38i21.30339}, abstractNote={Image collection summarization techniques aim to present a compact representation of an image gallery through a carefully selected subset of images that captures its semantic content. When it comes to web content, however, the ideal selection can vary based on the user’s specific intentions and preferences. This is particularly relevant at Booking.com, where presenting properties and their visual summaries that align with users’ expectations is crucial. To address this challenge, in this work, we consider user intentions in the summarization of property visuals by analyzing property reviews and extracting the most significant aspects mentioned by users. By incorporating the insights from reviews in our visual summaries, we enhance the summaries by presenting the relevant content to a user. Moreover, we achieve it without the need for costly annotations. Our experiments, including human perceptual studies, demonstrate the superiority of our cross-modal approach, which we coin as CrossSummarizer over the no-personalization and image-based clustering baselines.}, number={21}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, author={Wysoczanska, Monika and Beladev, Moran and Lastmann Assaraf, Karen and Wang, Fengjun and Kleinfeld, Ofri and Amsalem, Gil and Harush Boke, Hadas}, year={2024}, month={Mar.}, pages={22983-22989} }