2023
Solving Cosine Similarity Underestimation between High Frequency Words by ℓ2 Norm Discounting
ACL 2023findings
Cosine similarity between two words, computed using their contextualised token embeddings obtained from masked language models (MLMs) such as BERT has shown to underestimate the actual similarity between those words CITATION.This similarity underestimation problem is particularly severe for high fre…