← Search

Sheridan Feucht

3 accepted papers

2024

Token Erasure as a Footprint of Implicit Vocabulary Items in LLMs

EMNLP 2024main

LLMs process text as sequences of tokens that roughly correspond to words, where less common words are represented by multiple tokens. However, individual tokens are often semantically unrelated to the meanings of the words/concepts they comprise. For example, Llama-2-7b’s tokenizer splits the word…

Cited by 3SourcePDFScholar
2022

NEWTS: A Corpus for News Topic-Focused Summarization

ACL 2022findings

Text summarization models are approaching human levels of fidelity. Existing benchmarking corpora provide concordant pairs of full and abridged versions of Web, news or professional content. To date, all summarization datasets operate under a one-size-fits-all paradigm that may not reflect the full…