← Search

Heiko Ludwig

3 accepted papers

2026

When Data is the Algorithm: A Systematic Study and Curation of Preference Optimization Datasets

ICLR 2026poster

Aligning large language models (LLMs) is a central objective of post-training, often achieved through reward modeling and reinforcement learning methods. Among these, direct preference optimization (DPO) has emerged as a widely adopted technique that fine-tunes LLMs on preferred completions over les…

Cited by 0SourceScholar
2025

Fixing It in Post: A Comparative Study of LLM Post-Training Data Quality and Model Performance

NeurIPS 2025spotlight

Recent work on large language models (LLMs) has increasingly focused on post-training and alignment with datasets curated to enhance instruction following, world knowledge, and specialized skills. However, most post-training datasets used in leading open- and closed-source LLMs remain inaccessible t…

Cited by 0SourceScholar
2023

Single-shot General Hyper-parameter Optimization for Federated Learning

ICLR 2023top-25%

We address the problem of hyper-parameter optimization (HPO) for federated learning (FL-HPO). We introduce Federated Loss SuRface Aggregation (FLoRA), a general FL-HPO solution framework that can address use cases of tabular data and any Machine Learning (ML) model including gradient boosting traini…

Cited by 16SourcePDFScholar