← Search

Simin Fan

4 accepted papers

2025

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining

NeurIPS 2025poster

The performance of large language models (LLMs) across diverse downstream applications is fundamentally governed by the quality and composition of their pretraining corpora. Existing domain reweighting algorithms primarily optimize data mixtures for a single target task, thereby resulting in models…

Cited by 4SourceScholar
2025

Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling

ICLR 2025poster

Specialist language models (LMs) focus on a specific task or domain on which they often outperform generalist LMs of the same size. However, the specialist data needed to pretrain these models is only available in limited amount for most tasks. In this work, we build specialist models from large gen…

Cited by 3SourcePDFScholar
2022

Towards Process-Oriented, Modular, and Versatile Question Generation that Meets Educational Needs

NAACL 2022long

NLP-powered automatic question generation (QG) techniques carry great pedagogical potential of saving educators’ time and benefiting student learning. Yet, QG systems have not been widely adopted in classrooms to date. In this work, we aim to pinpoint key impediments and investigate how to improve t…