← Search

Giuseppe Amato

1 accepted papers

2026

One Patch to Caption Them All: A Unified Zero-Shot Captioning Framework

CVPR 2026

Zero-shot captioners are recently proposed models that utilize common-space vision-language representations to caption images without relying on paired image-text data. To caption an image, they proceed by textually decoding a text-aligned image feature, but they limit their scope to global represen

Cited by 0SourcecodeScholar