← Search

Yutaro Yamada

4 accepted papers

2025

L3GO: Language Agents with Chain-of-3D-Thoughts for Generating Unconventional Objects

NAACL 2025system demonstrations

Diffusion-based image generation models such as DALL-E 3 and Stable Diffusion-XL demonstrate remarkable capabilities in generating images with realistic and unique compositions. Yet, these models are not robust in precisely reasoning about physical and spatial configurations of objects, especially w…

2023

When are Lemons Purple? The Concept Association Bias of Vision-Language Models

EMNLP 2023long main

Large-scale vision-language models such as CLIP have shown impressive performance on zero-shot image classification and image-to-text retrieval. However, such performance does not realize in tasks that require a finer-grained correspondence between vision and language, such as Visual Question Answer…

Cited by 0SourceScholar