← Search

Mario Sanz-Guerrero

2 accepted papers

2025

Mind the Gap: A Closer Look at Tokenization for Multiple-Choice Question Answering with LLMs

EMNLP 2025

When evaluating large language models (LLMs) with multiple-choice question answering (MCQA), it is common to end the prompt with the string “Answer:” to facilitate automated answer extraction via next-token probabilities. However, there is no consensus on how to tokenize the space following the colo

Cited by 0SourcePDFScholar
2025

Molecular String Representation Preferences in Pretrained LLMs: A Comparative Study in Zero- & Few-Shot Molecular Property Prediction

EMNLP 2025

Large Language Models (LLMs) have demonstrated capabilities for natural language formulations of molecular property prediction tasks, but little is known about how performance depends on the representation of input molecules to the model; the status quo approach is to use SMILES strings, although al