2024
Right this way: Can VLMs Guide Us to See More to Answer Questions?
NeurIPS 2024poster
In question-answering scenarios, humans can assess whether the available information is sufficient and seek additional information if necessary, rather than providing a forced answer. In contrast, Vision Language Models (VLMs) typically generate direct, one-shot responses without evaluating the suff…