← Search

Alexandru Coca

5 accepted papers

2026

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

ICML 2026poster

Large Multimodal Models (LMMs) exhibit shortfalls when interpreting images and, by some measures, have poorer spatial cognition than young children or animals. Despite this, they attain high scores on many popular visual benchmarks, with headroom rapidly eroded by surging model progress. To address …

Cited by 0SourceScholar
2025

ASPERA: A Simulated Environment to Evaluate Planning for Complex Action Execution

ACL 2025long

This work evaluates the potential of large language models (LLMs) to power digital assistants capable of complex action execution. Such assistants rely on pre-trained programming knowledge to execute multi-step goals by composing objects and functions defined in assistant libraries into action execu…

2024

Effective and Efficient Conversation Retrieval for Dialogue State Tracking with Implicit Text Summaries

NAACL 2024long

Few-shot dialogue state tracking (DST) with Large Language Models (LLM) relies on an effective and efficient conversation retriever to find similar in-context examples for prompt learning. Previous works use raw dialogue context as search keys and queries, and a retriever is fine-tuned with annotate…

Cited by 4SourcePDFScholar
2023

Fine-grained Late-interaction Multi-modal Retrieval for Retrieval Augmented Visual Question Answering

NeurIPS 2023poster

Knowledge-based Visual Question Answering (KB-VQA) requires VQA systems to utilize knowledge from external knowledge bases to answer visually-grounded questions. Retrieval-Augmented Visual Question Answering (RA-VQA), a strong framework to tackle KB-VQA, first retrieves related documents with Dense…

2022

uFACT: Unfaithful Alien-Corpora Training for Semantically Consistent Data-to-Text Generation

ACL 2022findings

We propose uFACT (Un-Faithful Alien Corpora Training), a training corpus construction method for data-to-text (d2t) generation models. We show that d2t models trained on uFACT datasets generate utterances which represent the semantic content of the data sources more accurately compared to models tra…

Cited by 1SourcePDFScholar