2017
Automated Generation of Multilingual Clusters for the Evaluation of Distributed Representations
ICLR 2017workshop
We propose a language-agnostic way of automatically generating sets of semantically similar clusters of entities along with sets of "outlier" elements, which may then be used to perform an intrinsic evaluation of word embeddings in the outlier detection task. We used our methodology to create a gold…