← Search

Tomás Vergara Browne

4 accepted papers

2026

Operationalizing the Superficial Alignment Hypothesis via Task Complexity

ICML 2026poster

The superficial alignment hypothesis (SAH) posits that large language models learn most of their knowledge during pre-training, and that post-training merely surfaces this knowledge. The SAH, however, lacks a precise definition, which has led to (i) different and seemingly orthogonal arguments suppo…

Cited by 0SourceScholar
2025

Tracr-Injection: Distilling Algorithms into Pre-trained Language Models

ACL 2025finding

Motivated by the surge of large language models, there has been a push to formally characterize the symbolic abilities intrinsic to the transformer architecture. A programming language, called RASP, has been proposed, which can be directly compiled into transformer weights to implement these algorit…

2024

From Insights to Actions: The Impact of Interpretability and Analysis Research on NLP

EMNLP 2024main

Interpretability and analysis (IA) research is a growing subfield within NLP with the goal of developing a deeper understanding of the behavior or inner workings of NLP systems and methods. Despite growing interest in the subfield, a criticism of this work is that it lacks actionable insights and th…

2023

Large Language Models are biased to overestimate profoundness

EMNLP 2023short main

Recent advancements in natural language processing by large language models (LLMs), such as GPT-4, have been suggested to approach Artificial General Intelligence. And yet, it is still under dispute whether LLMs possess similar reasoning abilities to humans. This study evaluates GPT-4 and various ot…

Cited by 0SourcecodeScholar