← Search

Jose Hernandez-Orallo

5 accepted papers

2025

Evaluating Generalization Capabilities of LLM-Based Agents in Mixed-Motive Scenarios Using Concordia

NeurIPS 2025poster

Large Language Model (LLM) agents have demonstrated impressive capabilities for social interaction and are increasingly being deployed in situations where they might engage with both human and artificial agents. These interactions represent a critical frontier for LLM-based agents, yet existing eval…

Cited by 0SourceScholar
2025

Personalized Safety in LLMs: A Benchmark and A Planning-Based Agent Approach

NeurIPS 2025poster

Large language models (LLMs) typically generate identical or similar responses for all users given the same prompt, posing serious safety risks in high-stakes applications where user vulnerabilities differ widely. Existing safety evaluations primarily rely on context-independent metrics—such as fact…

Cited by 0SourcecodeScholar
2025

PredictaBoard: Benchmarking LLM Score Predictability

ACL 2025finding

Despite possessing impressive skills, Large Language Models (LLMs) often fail unpre-dictably, demonstrating inconsistent success in even basic common sense reasoning tasks. This unpredictability poses a significant challenge to ensuring their safe deployment, as identifying and operating within a re…

2024

Melting Pot Contest: Charting the Future of Generalized Cooperative Intelligence

NeurIPS 2024poster

Multi-agent AI research promises a path to develop human-like and human-compatible intelligent technologies that complement the solipsistic view of other approaches, which mostly do not consider interactions between agents. Aiming to make progress in this direction, the Melting Pot contest 2023 focu…

Cited by 0SourcePDFScholar
2021

Think Big, Teach Small: Do Language Models Distil Occam’s Razor?

NeurIPS 2021poster

Large language models have recently shown a remarkable ability for few-shot learning, including patterns of algorithmic nature. However, it is still an open question to determine what kind of patterns these models can capture and how many examples they need in their prompts. We frame this question a…