← Search

Aida Nematzadeh

9 accepted papers

2026

Dynamic Classifier-Free Diffusion Guidance via Online Feedback

ICLR 2026poster

Classifier-free guidance (CFG) is a cornerstone of text-to-image diffusion models, yet its effectiveness is limited by the use of static guidance scales. This ``one-size-fits-all'' approach fails to adapt to the diverse requirements of different prompts; moreover, prior solutions like gradient-based…

Cited by 0SourceScholar
2025

Revisiting text-to-image evaluation with Gecko: on metrics, prompts, and human rating

ICLR 2025spotlight

While text-to-image (T2I) generative models have become ubiquitous, they do not necessarily generate images that align with a given prompt. While many metrics and benchmarks have been proposed to evaluate T2I models and alignment metrics, the impact of the evaluation components (prompt sets, human…

Cited by 12SourcePDFScholar
2024

Evaluating Numerical Reasoning in Text-to-Image Models

NeurIPS 2024poster

Text-to-image generative models are capable of producing high-quality images that often faithfully depict concepts described using natural language. In this work, we comprehensively evaluate a range of text-to-image models on numerical reasoning tasks of varying difficulty, and show that even the mo…

2023

Measuring Progress in Fine-grained Vision-and-Language Understanding

ACL 2023long

While pretraining on large-scale image–text data from the Web has facilitated rapid progress on many vision-and-language (V&L) tasks, recent work has demonstrated that pretrained models lack “fine-grained” understanding, such as the ability to recognise relationships, verbs, and numbers in images. T…

2023

Pragmatics in Language Grounding: Phenomena, Tasks, and Modeling Approaches

EMNLP 2023long findings

People rely heavily on context to enrich meaning beyond what is literally said, enabling concise but effective communication. To interact successfully and naturally with people, user-facing artificial intelligence systems will require similar skills in pragmatics: relying on various types of context…

Cited by 0SourceScholar
2023

Weakly-Supervised Learning of Visual Relations in Multimodal Pretraining

EMNLP 2023long main

Recent work in vision-and-language pretraining has investigated supervised signals from object detection data to learn better, fine-grained multimodal representations. In this work, we take a step further and explore how we can tap into supervision from small-scale visual relation data. In particula…

Cited by 0SourcecodeScholar
2022

A Systematic Investigation of Commonsense Knowledge in Large Language Models

EMNLP 2022main

Language models (LMs) trained on large amounts of data have shown impressive performance on many NLP tasks under the zero-shot and few-shot setup. Here we aim to better understand the extent to which such models learn commonsense knowledge — a critical component of many NLP applications. We conduct…

Cited by 72SourcePDFScholar
2022

Flamingo: a Visual Language Model for Few-Shot Learning

NeurIPS 2022accept

Building models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research. We introduce Flamingo, a family of Visual Language Models (VLM) with this ability. We propose key architectural innovations to: (i) bri…

Cited by 4376SourcePDFScholar
2020

Visual Grounding in Video for Unsupervised Word Translation

CVPR 2020poster

There are thousands of actively spoken languages on Earth, but a single visual world. Grounding in this visual world has the potential to bridge the gap between all these languages. Our goal is to use visual grounding to improve unsupervised word mapping between languages. The key idea is to establi…

Cited by 58PDFcodeScholar