NAACL 2025industry0 citations

Granite Guardian: Comprehensive LLM Safeguarding

Inkit Padhi, Manish Nagireddy, Giandomenico Cornacchia, Subhajit Chaudhury, Tejaswini Pedapati, Pierre Dognin, Keerthiram Murugesan, Erik Miehling

Abstract

The deployment of language models in real-world applications exposes users to various risks, including hallucinations and harmful or unethical content. These challenges highlight the urgent need for robust safeguards to ensure safe and responsible AI. To address this, we introduce Granite Guardian, a suite of advanced models designed to detect and mitigate risks associated with prompts and responses, enabling seamless integration with any large language model (LLM). Unlike existing open-source solutions, our Granite Guardian models provide comprehensive coverage across a wide range of risk dimensions, including social bias, profanity, violence, sexual content, unethical behavior, jailbreaking, and hallucination-related issues such as context relevance, groundedness, and answer accuracy in retrieval-augmented generation (RAG) scenarios. Trained on a unique dataset combining diverse human annotations and synthetic data, Granite Guardian excels in identifying risks often overlooked by traditional detection systems, particularly jailbreak attempts and RAG-specific challenges. https://github.com/ibm-granite/granite-guardian

BibTeX
@inproceedings{padhi-etal-2025-granite,
    title = "Granite Guardian: Comprehensive {LLM} Safeguarding",
    author = "Padhi, Inkit  and
      Nagireddy, Manish  and
      Cornacchia, Giandomenico  and
      Chaudhury, Subhajit  and
      Pedapati, Tejaswini  and
      Dognin, Pierre  and
      Murugesan, Keerthiram  and
      Miehling, Erik  and
      Santill{\'a}n Cooper, Mart{\'i}n  and
      Fraser, Kieran  and
      Zizzo, Giulio  and
      Hameed, Muhammad Zaid  and
      Purcell, Mark  and
      Desmond, Michael  and
      Pan, Qian  and
      Vejsbjerg, Inge  and
      Daly, Elizabeth M.  and
      Hind, Michael  and
      Geyer, Werner  and
      Rawat, Ambrish  and
      Varshney, Kush R.  and
      Sattigeri, Prasanna",
    editor = "Chen, Weizhu  and
      Yang, Yi  and
      Kachuee, Mohammad  and
      Fu, Xue-Yong",
    booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 3: Industry Track)",
    month = apr,
    year = "2025",
    address = "Albuquerque, New Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.naacl-industry.49/",
    pages = "607--615",
    ISBN = "979-8-89176-194-0"
}
Granite Guardian: Comprehensive LLM Safeguarding · NAACL 2025