← Search

Daria Gitman

2 accepted papers

2025

NeMo-Inspector: A Visualization Tool for LLM Generation Analysis

NAACL 2025system demonstrations

Adapting Large Language Models (LLMs) to novel tasks and enhancing their overall capabilities often requires large, high-quality training datasets. Synthetic data, generated at scale, serves a valuable alternative when real-world data is scarce or difficult to obtain. However, ensuring the quality o…

2024

OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset

NeurIPS 2024oral

Recent work has shown the immense potential of synthetically generated datasets for training large language models (LLMs), especially for acquiring targeted skills. Current large-scale math instruction tuning datasets such as MetaMathQA (Yu et al., 2024) and MAmmoTH (Yue et al., 2024) are constructe…