← Search

Shaharukh Khan

3 accepted papers

2026

BhashaKritika: Building Synthetic Pretraining Data at Scale for Indic Languages

AAAI 2026technical

In the context of pretraining of Large Language Models (LLMs), synthetic data has emerged as an alternative for generating high-quality pretraining data at scale. This is particularly beneficial in low resource language settings where the benefits of the recent LLMs have been unevenly distributed ac

Cited by 0SourcePDFScholar
2026

IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs

ICLR 2026poster

Vision-language models (VLMs) have demonstrated impressive generalization across multimodal tasks, yet most evaluation benchmarks remain Western-centric, leaving open questions about their performance in culturally diverse and multilingual settings. To address this gap, we introduce IndicVisionBench…

Cited by 0SourcecodeScholar
2025

Chitrarth: Bridging Vision and Language for a Billion People

ICASSP 2025accepted

Recent multimodal foundation models are primarily trained on English or high resource European language data, which limits their applicability to other medium and low-resource languages, such as the Indian languages. To address this limitation, we introduce Chitrarth (Chitra: Image; Artha: Meaning),…

Cited by 12SourceScholar