← Search

Ahmed Alajrami

2 accepted papers

2023

Understanding the Role of Input Token Characters in Language Models: How Does Information Loss Affect Performance?

EMNLP 2023long main

Understanding how and what pre-trained language models (PLMs) learn about language is an open challenge in natural language processing. Previous work has focused on identifying whether they capture semantic and syntactic information, and how the data or the pre-training objective affects their perfo…

Cited by 0SourcecodeScholar
2022

How does the pre-training objective affect what large language models learn about linguistic properties?

ACL 2022short

Several pre-training objectives, such as masked language modeling (MLM), have been proposed to pre-train language models (e.g. BERT) with the aim of learning better language representations. However, to the best of our knowledge, no previous work so far has investigated how different pre-training ob…