2025
Training LLMs to be Better Text Embedders through Bidirectional Reconstruction
EMNLP 2025
Large language models (LLMs) have increasingly been explored as powerful text embedders. Existing LLM-based text embedding approaches often leverage the embedding of the final token, typically a reserved special token such as ‘[EOS]‘. However, these tokens have not been intentionally trained to capt