← Search

Yasumasa Onoe*

1 accepted papers

2024

DOCCI: Descriptions of Connected and Contrasting Images

ECCV 2024poster

"Vision-language datasets are vital for both text-to-image (T2I) and image-to-text (I2T) research. However, current datasets lack descriptions with fine-grained detail that would allow for richer associations to be learned by models. To fill the gap, we introduce Descriptions of Connected and Contra…

Cited by 52SourcePDFScholar