← Search

Ilker Yildirim

4 accepted papers

2025

L3GO: Language Agents with Chain-of-3D-Thoughts for Generating Unconventional Objects

NAACL 2025system demonstrations

Diffusion-based image generation models such as DALL-E 3 and Stable Diffusion-XL demonstrate remarkable capabilities in generating images with realistic and unique compositions. Yet, these models are not robust in precisely reasoning about physical and spatial configurations of objects, especially w…

2023

When are Lemons Purple? The Concept Association Bias of Vision-Language Models

EMNLP 2023long main

Large-scale vision-language models such as CLIP have shown impressive performance on zero-shot image classification and image-to-text retrieval. However, such performance does not realize in tasks that require a finer-grained correspondence between vision and language, such as Visual Question Answer…

Cited by 0SourceScholar
2017

Self-Supervised Intrinsic Image Decomposition

NeurIPS 2017poster

Intrinsic decomposition from a single image is a highly challenging task, due to its inherent ambiguity and the scarcity of training data. In contrast to traditional fully supervised learning approaches, in this paper we propose learning intrinsic image decomposition by explaining the input image. O…

Cited by 143SourcePDFScholar
2015

Galileo: Perceiving Physical Object Properties by Integrating a Physics Engine with Deep Learning

NeurIPS 2015poster

Humans demonstrate remarkable abilities to predict physical events in dynamic scenes, and to infer the physical properties of objects from static images. We propose a generative model for solving these problems of physical scene understanding from real-world videos and images. At the core of our gen…

Cited by 455SourcePDFScholar