← Search

Neha Verma

4 accepted papers

2024

Exploring Geometric Representational Disparities between Multilingual and Bilingual Translation Models

COLING 2024main

Multilingual machine translation has proven immensely useful for both parameter efficiency and overall performance across many language pairs via complete multilingual parameter sharing. However, some language pairs in multilingual models can see worse performance than in bilingual models, especiall…

Cited by 0SourcePDFScholar
2023

Multilingual Pixel Representations for Translation and Effective Cross-lingual Transfer

EMNLP 2023long main

We introduce and demonstrate how to effectively train multilingual machine translation models with pixel representations. We experiment with two different data settings with a variety of language and script coverage, demonstrating improved performance compared to subword embeddings. We explore vario…

Cited by 0SourceScholar
2022

IsoVec: Controlling the Relative Isomorphism of Word Embedding Spaces

EMNLP 2022main

The ability to extract high-quality translation dictionaries from monolingual word embedding spaces depends critically on the geometric similarity of the spaces—their degree of “isomorphism.” We address the root-cause of faulty cross-lingual mapping: that word embedding training resulted in the unde…

2021

DART: Open-Domain Structured Data Record to Text Generation

NAACL 2021long

We present DART, an open domain structured DAta Record to Text generation dataset with over 82k instances (DARTs). Data-to-text annotations can be a costly process, especially when dealing with tables which are the major source of structured data and contain nontrivial structures. To this end, we pr…