IJCAI 20260 citations

Semantic Similarity Is Not Legal Correctness: Evaluating RAG Systems in Brazilian Civil Procedure

Eryclis Silva, Madelyn Sanfilippo

Abstract

Evaluating retrieval-augmented generation (RAG) systems in legal domains is challenging due to the nuanced nature of legal reasoning and the scarcity of domain-specific benchmarks. We investigate semantic similarity and legal correctness in Brazilian legal question answering by developing a synthetic dataset of 3,012 evaluation instances derived from the Brazilian Civil Procedure Code, spanning seven query types and validated through human expert assessment. Our evaluation framework combines BERTScore with a domain-adapted LLM-as-Judge (GPT-4o-mini), validated against expert legal assessment. The analysis reveals a 50.8% disagreement between semantic similarity and legal correctness, with correlations varying substantially by query type (rho ranging from 0.464 to 0.732). Explicit article citation emerges as the strongest predictor of quality (Cohen’s d = 1.099), with cited responses achieving 145% higher legal correctness scores. Concrete examples demonstrate bidirectional divergence: responses may achieve high semantic similarity while citing incorrect legal articles, or provide legally correct answers with minimal lexical overlap. These findings demonstrate that semantic metrics alone are insufficient for evaluating legal RAG systems, with disagreement patterns systematically structured by query type.

Natural Language Processing: Information retrieval and text miningNatural Language Processing: ApplicationsNatural Language Processing: Question answeringAI: Natural Language Processing
BibTeX
@inproceedings{ijcai2026_semanticsimilari,
  title = {Semantic Similarity Is Not Legal Correctness: Evaluating RAG Systems in Brazilian Civil Procedure},
  author = {Eryclis Silva and Madelyn Sanfilippo},
  booktitle = {IJCAI 2026},
  year = {2026}
}
Semantic Similarity Is Not Legal Correctness: Evaluating RAG Systems in Brazilian Civil Procedure · IJCAI 2026