← Search

Santosh T.Y.S.S

28 accepted papers

2025

AQuAECHR: Attributed Question Answering for European Court of Human Rights

ACL 2025finding

LLMs have become prevalent tools for information seeking across various fields, including law. However, their generated responses often suffer from hallucinations, hindering their widespread adoption in high stakes domains such as law, which can potentially mislead experts and propagate societal har…

Cited by 0SourcePDFScholar
2025

CoCoLex: Confidence-guided Copy-based Decoding for Grounded Legal Text Generation

ACL 2025long

Due to their ability to process long and complex contexts, LLMs can offer key benefits to the Legal domain, but their adoption has been hindered by their tendency to generate unfaithful, ungrounded, or hallucinatory outputs. While Retrieval-Augmented Generation offers a promising solution by groundi…

Cited by 0SourcePDFScholar
2025

CoPERLex: Content Planning with Event-based Representations for Legal Case Summarization

NAACL 2025findings

Legal professionals often struggle with lengthy judgments and require efficient summarization for quick comprehension. To address this challenge, we investigate the need for structured planning in legal case summarization, particularly through event-centric representations that reflect the narrative…

Cited by 1SourcePDFScholar
2025

Fairness Beyond Performance: Revealing Reliability Disparities Across Groups in Legal NLP

ACL 2025long

Fairness in NLP must extend beyond performance parity to encompass equitable reliability across groups. This study exposes a criticalblind spot: models often make less reliable or overconfident predictions for marginalized groups, even when overall performance appearsfair. Using the FairLex benchmar…

Cited by 0SourcePDFScholar
2025

LeCoPCR: Legal Concept-guided Prior Case Retrieval for European Court of Human Rights cases

NAACL 2025findings

Prior case retrieval (PCR) is crucial for legal practitioners to find relevant precedent cases given the facts of a query case. Existing approaches often overlook the underlying semantic intent in determining relevance with respect to the query case. In this work, we propose LeCoPCR, a novel approac…

Cited by 1SourcePDFScholar
2025

LexKeyPlan: Planning with Keyphrases and Retrieval Augmentation for Legal Text Generation: A Case Study on European Court of Human Rights Cases

ACL 2025short

Large language models excel at legal text generation but often produce hallucinations due to their sole reliance on parametric knowledge. Retrieval-augmented models mitigate this by providing relevant external documents to the model but struggle when retrieval is based only on past context, which ma…

2025

LexTempus: Enhancing Temporal Generalizability of Legal Language Models Through Dynamic Mixture of Experts

ACL 2025long

The rapid evolution of legal concepts over time necessitates that legal language models adapt swiftly accounting for the temporal dynamics. However, prior works have largely neglected this crucial dimension, treating legal adaptation as a static problem rather than a continuous process. To address t…

Cited by 0SourcePDFScholar
2025

ProMALex: Progressive Modular Adapters for Multi-Jurisdictional Legal Language Modeling

ACL 2025long

This paper addresses the challenge of adapting language models to the jurisdiction-specific nature of legal corpora. Existing approaches—training separate models for each jurisdiction or using a single shared model—either fail to leverage common legal principles beneficial for low-resource settings…

Cited by 0SourcePDFScholar
2025

QABISAR: Query-Article Bipartite Interactions for Statutory Article Retrieval

COLING 2025main

In this paper, we introduce QABISAR, a novel framework for statutory article retrieval, to overcome the semantic mismatch problem when modeling each query-article pair in isolation, making it hard to learn representation that can effectively capture multi-faceted information. QABISAR leverages bipar…

Cited by 0SourcePDFScholar
2025

RELexED: Retrieval-Enhanced Legal Summarization with Exemplar Diversity

NAACL 2025findings

This paper addresses the task of legal summarization, which involves distilling complex legal documents into concise, coherent summaries. Current approaches often struggle with content theme deviation and inconsistent writing styles due to their reliance solely on source documents. We propose RELexE…

Cited by 1SourcePDFScholar
2024

A Tale of Two Revisions: Summarizing Changes Across Document Versions

ACL 2024findings

Document revision is a crucial aspect of the writing process, particularly in collaborative environments where multiple authors contribute simultaneously. However, current tools lack an efficient way to provide a comprehensive overview of changes between versions, leading to difficulties in understa…

2024

Beyond Borders: Investigating Cross-Jurisdiction Transfer in Legal Case Summarization

NAACL 2024long

Legal professionals face the challenge of managing an overwhelming volume of lengthy judgments, making automated legal case summarization crucial. However, prior approaches mainly focused on training and evaluating these models within the same jurisdiction. In this study, we explore the cross-jurisd…

2024

ChronosLex: Time-aware Incremental Training for Temporal Generalization of Legal Classification Tasks

ACL 2024long

This study investigates the challenges posed by the dynamic nature of legal multi-label text classification tasks, where legal concepts evolve over time. Existing models often overlook the temporal dimension in their training process, leading to suboptimal performance of those models over time, as t…

Cited by 4SourcePDFScholar
2024

CuSINeS: Curriculum-driven Structure Induced Negative Sampling for Statutory Article Retrieval

COLING 2024main

In this paper, we introduce CuSINeS, a negative sampling approach to enhance the performance of Statutory Article Retrieval (SAR). CuSINeS offers three key contributions. Firstly, it employs a curriculum-based negative sampling strategy guiding the model to focus on easier negatives initially and pr…

Cited by 2SourcePDFScholar
2024

ECtHR-PCR: A Dataset for Precedent Understanding and Prior Case Retrieval in the European Court of Human Rights

COLING 2024main

In common law jurisdictions, legal practitioners rely on precedents to construct arguments, in line with the doctrine of stare decisis. As the number of cases grow over the years, prior case retrieval (PCR) has garnered significant attention. Besides lacking real-world scale, existing PCR datasets d…

2024

HiCuLR: Hierarchical Curriculum Learning for Rhetorical Role Labeling of Legal Documents

EMNLP 2024finding

Rhetorical Role Labeling (RRL) of legal documents is pivotal for various downstream tasks such as summarization, semantic case search and argument mining. Existing approaches often overlook the varying difficulty levels inherent in legal document discourse styles and rhetorical roles. In this work,…

Cited by 2SourcePDFScholar
2024

Incorporating Precedents for Legal Judgement Prediction on European Court of Human Rights Cases

EMNLP 2024finding

Inspired by the legal doctrine of stare decisis, which leverages precedents (prior cases) for informed decision-making, we explore methods to integrate them into LJP models. To facilitate precedent retrieval, we train a retriever with a fine-grained relevance signal based on the overlap ratio of all…

Cited by 1SourcePDFScholar
2024

LexAbSumm: Aspect-based Summarization of Legal Decisions

COLING 2024main

Legal professionals frequently encounter long legal judgments that hold critical insights for their work. While recent advances have led to automated summarization solutions for legal documents, they typically provide generic summaries, which may not meet the diverse information needs of users. To a…

2024

Mind Your Neighbours: Leveraging Analogous Instances for Rhetorical Role Labeling for Legal Documents

COLING 2024main

Rhetorical Role Labeling (RRL) of legal judgments is essential for various tasks, such as case summarization, semantic search and argument mining. However, it presents challenges such as inferring sentence roles from context, interrelated roles, limited annotated data, and label imbalance. This stud…

Cited by 1SourcePDFScholar
2024

Query-driven Relevant Paragraph Extraction from Legal Judgments

COLING 2024main

Legal professionals often grapple with navigating lengthy legal judgements to pinpoint information that directly address their queries. This paper focus on this task of extracting relevant paragraphs from legal judgements based on the query. We construct a specialized dataset for this task from the…

2024

The Craft of Selective Prediction: Towards Reliable Case Outcome Classification - An Empirical Study on European Court of Human Rights Cases

EMNLP 2024finding

In high-stakes decision-making tasks within legal NLP, such as Case Outcome Classification (COC), quantifying a model’s predictive confidence is crucial. Confidence estimation enables humans to make more informed decisions, particularly when the model’s certainty is low, or where the consequences of…

Cited by 0SourcePDFScholar
2024

Through the Lens of Split Vote: Exploring Disagreement, Difficulty and Calibration in Legal Case Outcome Classification

ACL 2024long

In legal decisions, split votes (SV) occur when judges cannot reach a unanimous decision, posing a difficulty for lawyers who must navigate diverse legal arguments and opinions. In high-stakes domains, %as human-AI interaction systems become increasingly important, understanding the alignment of per…

Cited by 5SourcePDFScholar
2024

Towards Explainability and Fairness in Swiss Judgement Prediction: Benchmarking on a Multilingual Dataset

COLING 2024main

The assessment of explainability in Legal Judgement Prediction (LJP) systems is of paramount importance in building trustworthy and transparent systems, particularly considering the reliance of these systems on factors that may lack legal relevance or involve sensitive attributes. This study delves…

Cited by 7SourcePDFScholar
2023

From Dissonance to Insights: Dissecting Disagreements in Rationale Construction for Case Outcome Classification

EMNLP 2023long main

In legal NLP, Case Outcome Classification (COC) must not only be accurate but also trustworthy and explainable. Existing work in explainable COC has been limited to annotations by a single expert. However, it is well-known that lawyers may disagree in their assessment of case facts. We hence collect…

Cited by 0SourceScholar
2023

VECHR: A Dataset for Explainable and Robust Classification of Vulnerability Type in the European Court of Human Rights

EMNLP 2023short main

Recognizing vulnerability is crucial for understanding and implementing targeted support to empower individuals in need. This is especially important at the European Court of Human Rights (ECtHR), where the court adapts Convention standards to meet actual individual needs and thus to ensure effectiv…

Cited by 0SourcecodeScholar
2022

Deconfounding Legal Judgment Prediction for European Court of Human Rights Cases Towards Better Alignment with Experts

EMNLP 2022main

This work demonstrates that Legal Judgement Prediction systems without expert-informed adjustments can be vulnerable to shallow, distracting surface signals that arise from corpus construction, case distribution, and confounding factors. To mitigate this, we use domain expertise to strategically ide…

2020

SaSAKE: Syntax and Semantics Aware Keyphrase Extraction from Research Papers

COLING 2020main

Keyphrases in a research paper succinctly capture the primary content of the paper and also assist in indexing the paper at a concept level. Given the huge rate at which scientific papers are published today, it is important to have effective ways of automatically extracting keyphrases from a resear…

Cited by 18SourcePDFScholar