← Search

Julia Demarest

1 accepted papers

2026

PoSh: Using Scene Graphs to Guide LLMs-as-a-Judge for Detailed Image Descriptions

ICLR 2026poster

While vision-language models (VLMs) have advanced into detailed image description, evaluation remains a challenge. Standard metrics (e.g. CIDEr, SPICE) were designed for short texts and tuned to recognize errors that are now uncommon, such as object misidentification. In contrast, long texts require…

Cited by 0SourcecodeScholar