2026
Position: The Open Benchmark Paradox Must Be Resolved through Sovereign Medical Evaluation
ICML 2026poster
As medical large language models become increasingly involved in clinical actions, public benchmarks are often treated as proxies of deployment-readiness. However, this reliance creates a false sense of security because public scores are often based on data the models have already seen. We call this…