← Search

Ankur P. Parikh

6 accepted papers

2023

Extrapolative Controlled Sequence Generation via Iterative Refinement

ICML 2023poster

We study the problem of extrapolative controlled generation, i.e., generating sequences with attribute values beyond the range seen in training. This task is of significant importance in automated design, especially drug discovery, where the goal is to design novel proteins that are better (e.g., mo…

2023

SEAHORSE: A Multilingual, Multifaceted Dataset for Summarization Evaluation

EMNLP 2023long main

Reliable automatic evaluation of summarization systems is challenging due to the multifaceted and subjective nature of the task. This is especially the case for languages other than English, where human evaluations are scarce. In this work, we introduce SEAHORSE, a dataset for multilingual, multifac…

Cited by 0SourceScholar
2023

SQuId: Measuring Speech Naturalness in Many Languages

ICASSP 2023accepted

Much of text-to-speech research relies on human evaluation. This incurs heavy costs and slows down the development process, especially in heavily multilingual applications where recruiting and polling annotators can take weeks. We introduce SQuId (Speech Quality Identification), a multilingual natur…

Cited by 0SourceScholar
2023

TaTA: A Multilingual Table-to-Text Dataset for African Languages

EMNLP 2023long findings

Existing data-to-text generation datasets are mostly limited to English. To address this lack of data, we create Table-to-Text in African languages (TaTA), the first large multilingual table-to-text dataset with a focus on African languages. We created TaTA by transcribing figures and accompanying t…

Cited by 0SourcecodeScholar
2020

Thieves on Sesame Street! Model Extraction of BERT-based APIs

ICLR 2020poster

We study the problem of model extraction in natural language processing, in which an adversary with only query access to a victim model attempts to reconstruct a local copy of that model. Assuming that both the adversary and victim model fine-tune a large pretrained language model such as BERT (Devl…

Cited by 235SourcecodeScholar
2018

Multi-Mention Learning for Reading Comprehension with Neural Cascades

ICLR 2018poster

Reading comprehension is a challenging task, especially when executed across longer or across multiple evidence documents, where the answer is likely to reoccur. Existing neural architectures typically do not scale to the entire evidence, and hence, resort to selecting a single passage in the docume…

Cited by 41SourcePDFScholar