2025
EMO: Embedding Model Distillation via Intra-Model Relation and Optimal Transport Alignments
EMNLP 2025
Knowledge distillation (KD) is crucial for compressing large text embedding models, but faces challenges when teacher and student models use different tokenizers (Cross-Tokenizer KD - CTKD). Vocabulary mismatches impede the transfer of relational knowledge encoded in deep representations, such as hi