← Search

Rahul Aralikatte

6 accepted papers

2024

Chart-based Reasoning: Transferring Capabilities from LLMs to VLMs

NAACL 2024findings

Vision-language models (VLMs) are achieving increasingly strong performance on multimodal tasks. However, reasoning capabilities remain limited particularly for smaller VLMs, while those of large-language models (LLMs) have seen numerous improvements. We pro-pose a technique to transfer capabilities…

2023

Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic Languages

ACL 2023long

Building Natural Language Understanding (NLU) capabilities for Indic languages, which have a collective speaker base of more than one billion speakers is absolutely crucial. In this work, we aim to improve the NLU capabilities of Indic languages by making contributions along 3 important axes (i) mon…

2023

Varta: A Large-Scale Headline-Generation Dataset for Indic Languages

ACL 2023findings

We present Varta, a large-scale multilingual dataset for headline generation in Indic languages. This dataset includes more than 41 million pairs of headlines and articles in 14 different Indic languages (and English), which come from a variety of high-quality news sources. To the best of our knowle…

Cited by 10SourcePDFScholar
2021

Focus Attention: Promoting Faithfulness and Diversity in Summarization

ACL 2021long

Professional summaries are written with document-level information, such as the theme of the document, in mind. This is in contrast with most seq2seq decoders which simultaneously learn to focus on salient content, while deciding what to generate, at each decoding step. With the motivation to narrow…

2021

Joint Semantic Analysis with Document-Level Cross-Task Coherence Rewards

AAAI 2021technical

Coreference resolution and semantic role labeling are NLP tasks that capture different aspects of semantics, indicating respectively, which expressions refer to the same entity, and what semantic roles expressions serve in the sentence. However, they are often closely interdependent, and both genera…

2018

DuoRC: Towards Complex Language Understanding with Paraphrased Reading Comprehension

ICLR 2018workshop

We propose DuoRC, a novel dataset for Reading Comprehension (RC) that motivates several new challenges for neural approaches in language understanding beyond those offered by existing RC datasets. DuoRC contains 186,089 unique question-answer pairs created from a collection of 7680 pairs of movie pl…

Cited by 161SourcecodeScholar