← Search

Vivek Iyer

4 accepted papers

2025

XL-Suite: Cross-Lingual Synthetic Training and Evaluation Data for Open-Ended Generation

EMNLP 2025

Cross-lingual open-ended generation – responding in a language different from that of the query – is an important yet understudied problem. This work proposes XL-Instruct, a novel technique for generating high-quality synthetic data, and introduces XL-AlpacaEval, a new benchmark for evaluating cross

Cited by 0SourcePDFScholar
2023

Code-Switching with Word Senses for Pretraining in Neural Machine Translation

EMNLP 2023long findings

Lexical ambiguity is a significant and pervasive challenge in Neural Machine Translation (NMT), with many state-of-the-art (SOTA) NMT systems struggling to handle polysemous words (Campolungo et al., 2022). The same holds for the NMT pretraining paradigm of denoising synthetic "code-switched" text (…

Cited by 0SourceScholar
2021

VeeAlign: Multifaceted Context Representation Using Dual Attention for Ontology Alignment

EMNLP 2021main

Ontology Alignment is an important research problem applied to various fields such as data integration, data transfer, data preparation, etc. State-of-the-art (SOTA) Ontology Alignment systems typically use naive domain-dependent approaches with handcrafted rules or domain-specific architectures, ma…