← Search

Herbert Woisetschlaeger

2 accepted papers

2026

Position: Let's Develop Data Probes to Fundamentally Understand How Data Affects LLM Performance

ICML 2026poster

Data is fundamental to large language models (LLMs). However, understanding of what makes certain data useful for different stages of an LLM workflow, including training, tuning, alignment, in-context learning, etc., and why, remains an open question. Current approaches rely heavily on extensive exp…

Cited by 0SourceScholar
2023

A Survey on Dataset Distillation: Approaches, Applications and Future Directions

IJCAI 2023poster

Dataset distillation is attracting more attention in machine learning as training sets continue to grow and the cost of training state-of-the-art models becomes increasingly high. By synthesizing datasets with high information density, dataset distillation offers a range of potential applications, i…