2021
RedCaps: Web-curated image-text data created by the people, for the people
NeurIPS 2021poster
Large datasets of paired images and text have become increasingly popular for learning generic representations for vision and vision-and-language tasks. Such datasets have been built by querying search engines or collecting HTML alt-text – since web data is noisy, they require complex filtering pipe…