← Search

Ferenc Huszár

10 accepted papers

2026

Out-of-Distribution Evaluation of Rule-Based and Strategic Reasoning in Chess Transformers

ICML 2026poster

Modern decision transformers, trained similarly to LLMs, can achieve strong in-distribution performance in complex sequential domains like chess, but it remains unclear to what extent they reason systematically about rules and strategy. We study the reasoning capabilities of a 270M-parameter chess t…

Cited by 0SourceScholar
2025

Identifiable Exchangeable Mechanisms for Causal Structure and Representation Learning

ICLR 2025spotlight

Identifying latent representations or causal structures is important for good generalization and downstream task performance. However, both fields developed rather independently. We observe that several structure and representation identifiability methods, particularly those that require multiple en…

Cited by 3SourcePDFScholar
2024

Do Finetti: On Causal Effects for Exchangeable Data

NeurIPS 2024oral

We study causal effect estimation in a setting where the data are not i.i.d.$\ $(independent and identically distributed). We focus on exchangeable data satisfying an assumption of independent causal mechanisms. Traditional causal effect estimation frameworks, e.g., relying on structural causal mode…

Cited by 1SourcePDFScholar
2024

Position: Understanding LLMs Requires More Than Statistical Generalization

ICML 2024spotlight

The last decade has seen blossoming research in deep learning theory attempting to answer, ``Why does deep learning generalize?" A powerful shift in perspective precipitated this progress: the study of overparametrized models in the interpolation regime. In this paper, we argue that another perspect…

2024

Recurrent Early Exits for Federated Learning with Heterogeneous Clients

ICML 2024poster

Federated learning (FL) has enabled distributed learning of a model across multiple clients in a privacy-preserving manner. One of the main challenges of FL is to accommodate clients with varying hardware capacities; clients have differing compute and memory requirements. To tackle this challenge, r…

2024

Rule Extrapolation in Language Modeling: A Study of Compositional Generalization on OOD Prompts

NeurIPS 2024spotlight

LLMs show remarkable emergent abilities, such as inferring concepts from presumably out-of-distribution prompts, known as in-context learning. Though this success is often attributed to the Transformer architecture, our systematic understanding is limited. In complex real-world data sets, even defin…

Cited by 2SourcePDFScholar
2023

Causal de Finetti: On the Identification of Invariant Causal Structure in Exchangeable Data

NeurIPS 2023poster

Constraint-based causal discovery methods leverage conditional independence tests to infer causal relationships in a wide variety of applications. Just as the majority of machine learning methods, existing work focuses on studying $\textit{independent and identically distributed}$ data. However, it…

2023

FedL2P: Federated Learning to Personalize

NeurIPS 2023poster

Federated learning (FL) research has made progress in developing algorithms for distributed learning of global models, as well as algorithms for local personalization of those common models to the specifics of each client’s local data distribution. However, different FL problems may require differen…

2017

Amortised MAP Inference for Image Super-resolution

ICLR 2017oral

Image super-resolution (SR) is an underdetermined inverse problem, where a large number of plausible high resolution images can explain the same downsampled image. Most current single image SR methods use empirical risk minimisation, often with a pixel-wise mean squared error (MSE) loss. However, th…

Cited by 538SourceScholar