← Search

Konstantinos Voudouris

5 accepted papers

2026

Can vision language models learn intuitive physics from interaction?

ICML 2026poster

Pre-trained vision language models do not have good intuitions about the physical world. Recent work has shown that supervised fine-tuning can improve model performance on simple physical tasks. However, fine-tuned models do not appear to learn robust physical rules that can generalize to new contex…

Cited by 0SourceScholar
2025

PredictaBoard: Benchmarking LLM Score Predictability

ACL 2025finding

Despite possessing impressive skills, Large Language Models (LLMs) often fail unpre-dictably, demonstrating inconsistent success in even basic common sense reasoning tasks. This unpredictability poses a significant challenge to ensuring their safe deployment, as identifying and operating within a re…

2025

Testing the Limits of Fine-Tuning for Improving Visual Cognition in Vision Language Models

ICML 2025poster

Pre-trained vision language models still fall short of human visual cognition. In an effort to improve visual cognition and align models with human behavior, we introduce visual stimuli and human judgments on visual cognition tasks, allowing us to systematically evaluate performance across cognitive…

Cited by 0SourcePDFScholar
2025

metabench - A Sparse Benchmark of Reasoning and Knowledge in Large Language Models

ICLR 2025poster

Large Language Models (LLMs) vary in their abilities on a range of tasks. Initiatives such as the Open LLM Leaderboard aim to quantify these differences with several large benchmarks (sets of test items to which an LLM can respond either correctly or incorrectly). However, high correlations withi…

Cited by 0SourcePDFScholar
2022

Not a Number: Identifying Instance Features for Capability-Oriented Evaluation

IJCAI 2022poster

In AI evaluation, performance is often calculated by averaging across various instances. But to fully understand the capabilities of an AI system, we need to understand the factors that cause its pattern of success and failure. In this paper, we present a new methodology to identify and build inform…