← Search

Xiaomo Liu

7 accepted papers

2026

Perturb Your Data: Paraphrase-Guided Training Data Watermarking

AAAI 2026technical

Training data detection is critical for enforcing copyright and data licensing, as Large Language Models (LLM) are trained on massive text corpora scraped from the internet. We present SPECTRA, a watermarking approach that makes training data reliably detectable even when it comprises less than 0.00

Cited by 0SourcePDFScholar
2025

CoCoLex: Confidence-guided Copy-based Decoding for Grounded Legal Text Generation

ACL 2025long

Due to their ability to process long and complex contexts, LLMs can offer key benefits to the Legal domain, but their adoption has been hindered by their tendency to generate unfaithful, ungrounded, or hallucinatory outputs. While Retrieval-Augmented Generation offers a promising solution by groundi…

Cited by 0SourcePDFScholar
2025

The Impact of Domain-Specific Terminology on Machine Translation for Finance in European Languages

NAACL 2025long

Domain-specific machine translation (MT) poses significant challenges due to specialized terminology, particularly when translating across multiple languages with scarce resources. In this study, we present the first impact analysis of domain-specific terminology on multilingual MT for finance, focu…

2024

DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding

ACL 2024long

Enterprise documents such as forms, receipts, reports, and other such records, often carry rich semantics at the intersection of textual and spatial modalities. The visual cues offered by their complex layouts play a crucial role in comprehending these documents effectively. In this paper, we presen…

2024

Fine-Tuning Language Models with Differential Privacy through Adaptive Noise Allocation

EMNLP 2024finding

Language models are capable of memorizing detailed patterns and information, leading to a double-edged effect: they achieve impressive modeling performance on downstream tasks with the stored knowledge but also raise significant privacy concerns. Traditional differential privacy based training appro…

Cited by 1SourcePDFScholar
2024

The State of the Art of Large Language Models on Chartered Financial Analyst Exams

EMNLP 2024industry

The Chartered Financial Analyst (CFA) program is one of the most widely recognized financial certifications globally. In this work, we test a variety of state-of-the-art large language models (LLMs) on mock CFA exams to provide an overview of their financial analysis capabilities using the same eval…

Cited by 2SourcePDFScholar
2024

“What is the value of templates?” Rethinking Document Information Extraction Datasets for LLMs

EMNLP 2024finding

The rise of large language models (LLMs) for visually rich document understanding (VRDU) has kindled a need for prompt-response, document-based datasets. As annotating new datasets from scratch is labor-intensive, the existing literature has generated prompt-response datasets from available resource…

Cited by 0SourcePDFScholar