← Search

Alessio Cocchieri

5 accepted papers

2025

Can Large Language Models Win the International Mathematical Games?

EMNLP 2025

Recent advances in large language models (LLMs) have demonstrated strong mathematical reasoning abilities, even in visual contexts, with some models surpassing human performance on existing benchmarks. However, these benchmarks lack structured age categorization, clearly defined skill requirements,

2025

OpenBioNER: Lightweight Open-Domain Biomedical Named Entity Recognition Through Entity Type Description

NAACL 2025findings

Biomedical Named Entity Recognition (BioNER) faces significant challenges in real-world applications due to limited annotated data and the constant emergence of new entity types, making zero-shot learning capabilities crucial. While Large Language Models (LLMs) possess extensive domain knowledge nec…

2025

ZeroNER: Fueling Zero-Shot Named Entity Recognition via Entity Type Descriptions

ACL 2025finding

What happens when a named entity recognition (NER) system encounters entities it has never seen before? In practical applications, models must generalize to unseen entity types where labeled training data is either unavailable or severely limited—a challenge that demands zero-shot learning capabilit…

2025

“What do you call a dog that is incontrovertibly true? Dogma”: Testing LLM Generalization through Humor

ACL 2025long

Humor, requiring creativity and contextual understanding, is a hallmark of human intelligence, showcasing adaptability across linguistic scenarios. While recent advances in large language models (LLMs) demonstrate strong reasoning on various benchmarks, it remains unclear whether they truly adapt to…

Cited by 0SourcePDFScholar
2024

To Generate or to Retrieve? On the Effectiveness of Artificial Contexts for Medical Open-Domain Question Answering

ACL 2024long

Medical open-domain question answering demands substantial access to specialized knowledge. Recent efforts have sought to decouple knowledge from model parameters, counteracting architectural scaling and allowing for training on common low-resource hardware. The retrieve-then-read paradigm has becom…