NAACL 2025system demonstrations4 citations

Libra-Leaderboard: Towards Responsible AI through a Balanced Leaderboard of Safety and Capability

Haonan Li, Xudong Han, Zenan Zhai, Honglin Mu, Hao Wang, Zhenxuan Zhang, Yilin Geng, Shom Lin

Abstract

As large language models (LLMs) continue to evolve, leaderboards play a significant role in steering their development. Existing leaderboards often prioritize model capabilities while overlooking safety concerns, leaving a significant gap in responsible AI development. To address this gap, we introduce Libra-Leaderboard, a comprehensive framework designed to rank LLMs through a balanced evaluation of performance and safety. Combining a dynamic leaderboard with an interactive LLM arena, Libra-Leaderboard encourages the joint optimization of capability and safety. Unlike traditional approaches that average performance and safety metrics, Libra-Leaderboard uses a distance-to-optimal-score method to calculate the overall rankings. This approach incentivizes models to achieve a balance rather than excelling in one dimension at the expense of some other ones. In the first release, Libra-Leaderboard evaluates 26 mainstream LLMs from 14 leading organizations, identifying critical safety challenges even in state-of-the-art models.

BibTeX
@inproceedings{li-etal-2025-libra,
    title = "Libra-Leaderboard: Towards Responsible {AI} through a Balanced Leaderboard of Safety and Capability",
    author = "Li, Haonan  and
      Han, Xudong  and
      Zhai, Zenan  and
      Mu, Honglin  and
      Wang, Hao  and
      Zhang, Zhenxuan  and
      Geng, Yilin  and
      Lin, Shom  and
      Wang, Renxi  and
      Shelmanov, Artem  and
      Qi, Xiangyu  and
      Wang, Yuxia  and
      Hong, Donghai  and
      Yuan, Youliang  and
      Chen, Meng  and
      Tu, Haoqin  and
      Koto, Fajri  and
      Zeng, Cong  and
      Kuribayashi, Tatsuki  and
      Bhardwaj, Rishabh  and
      Zhao, Bingchen  and
      Duan, Yawen  and
      Liu, Yi  and
      Alghamdi, Emad A.  and
      Yang, Yaodong  and
      Dong, Yinpeng  and
      Poria, Soujanya  and
      Liu, Pengfei  and
      Liu, Zhengzhong  and
      Ren, Hector Xuguang  and
      Hovy, Eduard  and
      Gurevych, Iryna  and
      Nakov, Preslav  and
      Choudhury, Monojit  and
      Baldwin, Timothy",
    editor = "Dziri, Nouha  and
      Ren, Sean (Xiang)  and
      Diao, Shizhe",
    booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System Demonstrations)",
    month = apr,
    year = "2025",
    address = "Albuquerque, New Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.naacl-demo.23/",
    pages = "268--286",
    ISBN = "979-8-89176-191-9"
}
Libra-Leaderboard: Towards Responsible AI through a Balanced Leaderboard of Safety and Capability · NAACL 2025