← Search

Ayush Maheshwari

7 accepted papers

2025

ARISE: Iterative Rule Induction and Synthetic Data Generation for Text Classification

NAACL 2025findings

We propose ARISE, a framework that iteratively induces rules and generates synthetic data for text classification. We combine synthetic data generation and automatic rule induction, via bootstrapping, to iteratively filter the generated rules and data. We induce rules via inductive generalisation of…

Cited by 0SourcePDFScholar
2025

LexGen: Domain-aware Multilingual Lexicon Generation

ACL 2025long

Lexicon or dictionary generation across domains has the potential for societal impact, as it can potentially enhance information accessibility for a diverse user base while preserving language identity. Prior work in the field primarily focuses on bilingual lexical induction, which deals with word a…

2024

DictDis: Dictionary Constrained Disambiguation for Improved NMT

EMNLP 2024finding

Domain-specific neural machine translation (NMT) systems (, in educational applications) are socially significant with the potential to help make information accessible to a diverse set of users in multilingual societies. Such NMT systems should be lexically constrained and draw from domain-specific…

2024

Samayik: A Benchmark and Dataset for English-Sanskrit Translation

COLING 2024main

We release Saamayik, a dataset of around 53,000 parallel English-Sanskrit sentences, written in contemporary prose. Sanskrit is a classical language still in sustenance and has a rich documented heritage. However, due to the limited availability of digitized content, it still remains a low-resource…

2023

Adaptive Mixing of Auxiliary Losses in Supervised Learning

AAAI 2023technical

In many supervised learning scenarios, auxiliary losses are used in order to introduce additional information or constraints into the supervised learning objective. For instance, knowledge distillation aims to mimic outputs of a powerful teacher model; similarly, in rule-based approaches, weak label…

2022

A Benchmark and Dataset for Post-OCR text correction in Sanskrit

EMNLP 2022finding

Sanskrit is a classical language with about 30 million extant manuscripts fit for digitisation, available in written, printed or scanned-image forms. However, it is still considered to be a low-resource language when it comes to available digital resources. In this work, we release a post-OCR text c…

2022

Learning to Robustly Aggregate Labeling Functions for Semi-supervised Data Programming

ACL 2022findings

A critical bottleneck in supervised machine learning is the need for large amounts of labeled data which is expensive and time-consuming to obtain. Although a small amount of labeled data cannot be used to train a model, it can be used effectively for the generation of humaninterpretable labeling fu…