Libra-Leaderboard: Towards Responsible AI through a Balanced Leaderboard of Safety and Capability
Haonan Li, Xudong Han, Zenan Zhai, Honglin Mu, Hao Wang, Zhenxuan Zhang, Yilin Geng, Shom Lin
Abstract
As large language models (LLMs) continue to evolve, leaderboards play a significant role in steering their development. Existing leaderboards often prioritize model capabilities while overlooking safety concerns, leaving a significant gap in responsible AI development. To address this gap, we introduce Libra-Leaderboard, a comprehensive framework designed to rank LLMs through a balanced evaluation of performance and safety. Combining a dynamic leaderboard with an interactive LLM arena, Libra-Leaderboard encourages the joint optimization of capability and safety. Unlike traditional approaches that average performance and safety metrics, Libra-Leaderboard uses a distance-to-optimal-score method to calculate the overall rankings. This approach incentivizes models to achieve a balance rather than excelling in one dimension at the expense of some other ones. In the first release, Libra-Leaderboard evaluates 26 mainstream LLMs from 14 leading organizations, identifying critical safety challenges even in state-of-the-art models.
BibTeX
@inproceedings{li-etal-2025-libra,
title = "Libra-Leaderboard: Towards Responsible {AI} through a Balanced Leaderboard of Safety and Capability",
author = "Li, Haonan and
Han, Xudong and
Zhai, Zenan and
Mu, Honglin and
Wang, Hao and
Zhang, Zhenxuan and
Geng, Yilin and
Lin, Shom and
Wang, Renxi and
Shelmanov, Artem and
Qi, Xiangyu and
Wang, Yuxia and
Hong, Donghai and
Yuan, Youliang and
Chen, Meng and
Tu, Haoqin and
Koto, Fajri and
Zeng, Cong and
Kuribayashi, Tatsuki and
Bhardwaj, Rishabh and
Zhao, Bingchen and
Duan, Yawen and
Liu, Yi and
Alghamdi, Emad A. and
Yang, Yaodong and
Dong, Yinpeng and
Poria, Soujanya and
Liu, Pengfei and
Liu, Zhengzhong and
Ren, Hector Xuguang and
Hovy, Eduard and
Gurevych, Iryna and
Nakov, Preslav and
Choudhury, Monojit and
Baldwin, Timothy",
editor = "Dziri, Nouha and
Ren, Sean (Xiang) and
Diao, Shizhe",
booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System Demonstrations)",
month = apr,
year = "2025",
address = "Albuquerque, New Mexico",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.naacl-demo.23/",
pages = "268--286",
ISBN = "979-8-89176-191-9"
}