← Search

José Hernández-Orallo

8 accepted papers

2025

Contamination Budget: Trade-offs Between Breadth, Depth and Difficulty

IJCAI 2025

Contamination in large language models (LLMs), and machine learning more broadly, refers to the inclusion of equal --or very similar-- examples in both training and test sets. This phenomenon usually translates into better test performance. Here we explore when this contamination is performed intent

2025

Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture

IJCAI 2025

Research in AI evaluation has grown increasingly complex and multidisciplinary, attracting researchers with diverse backgrounds and objectives. As a result, divergent evaluation paradigms have emerged, often developing in isolation, adopting conflicting terminologies, and overlooking each other's co

Cited by 0SourcePDFScholar
2024

Your Prompt Is My Command: On Assessing the Human-Centred Generality of Multimodal Models (Abstract Reprint)

AAAI 2024technical

Even with obvious deficiencies, large prompt-commanded multimodal models are proving to be flexible cognitive tools representing an unprecedented generality. But the directness, diversity, and degree of user interaction create a distinctive “human-centred generality” (HCG), rather than a fully auton…

Cited by 0SourcePDFScholar
2022

How General-Purpose Is a Language Model? Usefulness and Safety with Human Prompters in the Wild

AAAI 2022technical

The new generation of language models is reported to solve some extraordinary tasks the models were never trained for specifically, in few-shot or zero-shot settings. However, these reports usually cherry-pick the tasks, use the best prompts, and unwrap or extract the solutions leniently even if the…

2022

Measuring the Occupational Impact of AI: Tasks, Cognitive Abilities and AI Benchmarks (Extended Abstract)*

IJCAI 2022poster

We present a framework for analysing the impact of AI on occupations. This framework maps 59 generic tasks from different occupational datasets to 14 cognitive abilities and these to a comprehensive list of 328 AI benchmarks used to evaluate research intensity in AI. The use of cognitive abilities…

2022

Non-Cheating Teaching Revisited: A New Probabilistic Machine Teaching Model

IJCAI 2022poster

Over the past decades in the field of machine teaching, several restrictions have been introduced to avoid ‘cheating’, such as collusion-free or non-clashing teaching. However, these restrictions forbid several teaching situations that we intuitively consider natural and fair, especially those ‘…

Cited by 6SourcePDFScholar
2022

Not a Number: Identifying Instance Features for Capability-Oriented Evaluation

IJCAI 2022poster

In AI evaluation, performance is often calculated by averaging across various instances. But to fully understand the capabilities of an AI system, we need to understand the factors that cause its pattern of success and failure. In this paper, we present a new methodology to identify and build inform…

2022

When AI Difficulty Is Easy: The Explanatory Power of Predicting IRT Difficulty

AAAI 2022technical

One of challenges of artificial intelligence as a whole is robustness. Many issues such as adversarial examples, out of distribution performance, Clever Hans phenomena, and the wider areas of AI evaluation and explainable AI, have to do with the following question: Did the system fail because it is…

Cited by 15SourcePDFScholar