2022
Signal in Noise: Exploring Meaning Encoded in Random Character Sequences with Character-Aware Language Models
Mark Chu, Bhargav Srinivasa Desikan, Ethan Nadler, Donald Ruggiero Lo Sardo, Elise Darragh-Ford, Douglas Guilbeault
ACL 2022long
Natural language processing models learn word representations based on the distributional hypothesis, which asserts that word context (e.g., co-occurrence) correlates with meaning. We propose that n-grams composed of random character sequences, or garble, provide a novel context for studying word me…