← Search

Maeda Hanafi

1 accepted papers

2025

Identifying Noise in Human-Created Datasets using Training Dynamics from Generative Models

EMNLP 2025

Instruction fine-tuning enhances the alignment of autoregressive language models (ArLMs) with human intent but relies on large-scale annotated datasets prone to label and text noise. In this paper, we show that existing noise detection techniques designed for autoencoder models (AeLMs) do not direct

Cited by 0SourcePDFScholar