← Search

Jay Piplodiya

1 accepted papers

2026

BhashaKritika: Building Synthetic Pretraining Data at Scale for Indic Languages

AAAI 2026technical

In the context of pretraining of Large Language Models (LLMs), synthetic data has emerged as an alternative for generating high-quality pretraining data at scale. This is particularly beneficial in low resource language settings where the benefits of the recent LLMs have been unevenly distributed ac

Cited by 0SourcePDFScholar