2025
MedOdyssey: A Medical Domain Benchmark for Long Context Evaluation Up to 200K Tokens
NAACL 2025findings
Numerous advanced Large Language Models (LLMs) now support context lengths up to 128K, and some extend to 200K. Some benchmarks in the generic domain have also followed up on evaluating long-context capabilities. In the medical domain, tasks are distinctive due to the unique contexts and need for do…