← Search

Xiaojun Wan

74 accepted papers

2026

SCOPE: Intrinsic Semantic Space Control for Mitigating Copyright Infringement in LLMs

AAAI 2026technical

Large language models sometimes inadvertently reproduce passages that are copyrighted, exposing downstream applications to legal risk. Most existing studies for inference-time defences focus on surface-level token matching and rely on external blocklists or filters, which add deployment complexity a

Cited by 0SourcePDFScholar
2026

SCOUT: Active Information Foraging for Long-Text Understanding with Decoupled Epistemic States

ICML 2026poster

Long-Text Understanding (LTU) at million-token scale requires balancing reasoning fidelity with computational efficiency. Frontier long-context LLMs can process millions of token contexts end-to-end, but they suffer from high token consumption and attention dilution. In parallel, specialized LTU age…

Cited by 0SourceScholar
2025

A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better Interpretability

ACL 2025long

In NLG meta-evaluation, evaluation metrics are typically assessed based on their consistency with humans. However, we identify some limitations in traditional NLG meta-evaluation approaches, such as issues in handling human ratings and ambiguous selections of correlation measures, which undermine th…

2025

Analyzing and Evaluating Correlation Measures in NLG Meta-Evaluation

NAACL 2025long

The correlation between NLG automatic evaluation metrics and human evaluation is often regarded as a critical criterion for assessing the capability of an evaluation metric. However, different grouping methods and correlation coefficients result in various types of correlation measures used in meta-…

2025

DAMON: A Dialogue-Aware MCTS Framework for Jailbreaking Large Language Models

EMNLP 2025

While large language models (LLMs) demonstrate remarkable capabilities across a wide range of tasks, they remain vulnerable to generating outputs that are potentially harmful. Red teaming, which involves crafting adversarial inputs to expose vulnerabilities, is a widely adopted approach for evaluati

2025

DSGram: Dynamic Weighting Sub-Metrics for Grammatical Error Correction in the Era of Large Language Models

AAAI 2025technical

Evaluating the performance of Grammatical Error Correction (GEC) models has become increasingly challenging, as large language model (LLM)-based GEC systems often produce corrections that diverge from provided gold references. This discrepancy undermines the reliability of traditional reference-base…

2025

Enhancing LLM Watermark Resilience Against Both Scrubbing and Spoofing Attacks

NeurIPS 2025spotlight

Watermarking is a promising defense against the misuse of large language models (LLMs), yet it remains vulnerable to scrubbing and spoofing attacks. This vulnerability stems from an inherent trade-off governed by watermark window size: smaller windows resist scrubbing better but are easier to rev…

Cited by 0SourcecodeScholar
2025

Evaluating Self-Generated Documents for Enhancing Retrieval-Augmented Generation with Large Language Models

NAACL 2025findings

The integration of documents generated by LLMs themselves (Self-Docs) alongside retrieved documents has emerged as a promising strategy for retrieval-augmented generation systems. However, previous research primarily focuses on optimizing the use of Self-Docs, with their inherent properties remainin…

Cited by 0SourcePDFScholar
2025

Exploring and Evaluating Multimodal Knowledge Reasoning Consistency of Multimodal Large Language Models

EMNLP 2025

In recent years, multimodal large language models (MLLMs) have achieved significant breakthroughs, enhancing understanding across text and vision. However, current MLLMs still face challenges in effectively integrating knowledge across these modalities during multimodal knowledge reasoning, leading

2025

Gödel Agent: A Self-Referential Agent Framework for Recursively Self-Improvement

ACL 2025long

The rapid advancement of large language models (LLMs) has significantly enhanced the capabilities of agents across various tasks. However, existing agentic systems, whether based on fixed pipeline algorithms or pre-defined meta-learning frameworks, cannot search the whole agent design space due to t…

2025

ICR Probe: Tracking Hidden State Dynamics for Reliable Hallucination Detection in LLMs

ACL 2025long

Large language models (LLMs) excel at various natural language processing tasks, but their tendency to generate hallucinations undermines their reliability. Existing hallucination detection methods leveraging hidden states predominantly focus on static and isolated representations, overlooking their…

Cited by 0SourcePDFScholar
2025

MC-MKE: A Fine-Grained Multimodal Knowledge Editing Benchmark Emphasizing Modality Consistency

ACL 2025finding

Multimodal large language models (MLLMs) are prone to non-factual or outdated knowledge issues, highlighting the importance of knowledge editing. Many benchmark has been proposed for researching multimodal knowledge editing. However, previous benchmarks focus on limited scenarios due to the lack of…

Cited by 0SourcePDFScholar
2025

R-Bind: Unified Enhancement of Attribute and Relation Binding in Text-to-Image Diffusion Models

EMNLP 2025

Text-to-image models frequently fail to achieve perfect alignment with textual prompts, particularly in maintaining proper semantic binding between semantic elements in the given prompt. Existing approaches typically require costly retraining or focus on only correctly generating the attributes of e

Cited by 0SourcePDFScholar
2025

Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference

NAACL 2025findings

Evaluating and ranking the capabilities of different LLMs is crucial for understanding their performance and alignment with human preferences. Due to the high cost and time-consuming nature of human evaluations, an automatic LLM bencher (i.e., an automatic evaluation framework that aims to rank LLMs…

2025

Towards A “Novel” Benchmark: Evaluating Literary Fiction with Large Language Models

ACL 2025finding

Current exploration on creative generation focuses mainly on short stories, poetry, and scripts. With the expansion of Large Language Models (LLMs) context windows, “novel” avenues emerge. This study aims to extend the boundaries of Natural Language Generation (NLG) evaluation by exploring LLMs’ cap…

2025

Tracing Training Footprints: A Calibration Approach for Membership Inference Attacks Against Multimodal Large Language Models

EMNLP 2025

With the increasing scale of training data for Multimodal Large Language Models (MLLMs) and the lack of data details, there is growing concern about privacy breaches and data security issues. Under black-box access, exploring effective Membership Inference Attacks (MIA) has garnered increasing atten

Cited by 0SourcePDFScholar
2025

TriEmbed: Bridge the Gap between Text and Token Indices with Embedding Reparameterization

ACL 2025finding

The current paradigm of language modeling is a two-stage pipeline that first transforms raw text to token indices, where the distribution is then estimated. It inherently discards linguistic relations between tokens during tokenization, creating a fundamental gap. To address this, we propose TriEmbe…

Cited by 0SourcePDFScholar
2025

WaterPool: A Language Model Watermark Mitigating Trade-Offs among Imperceptibility, Efficacy and Robustness

NAACL 2025long

Watermarking is a prominent technique to trace the usage of specific large language models (LLMs) by injecting patterns into model-generated content. An ideal watermark should be imperceptible, easily detectable, and robust to text alterations, yet existing methods typically face trade-offs among th…

Cited by 0SourcePDFScholar
2024

Are LLM-based Evaluators Confusing NLG Quality Criteria?

ACL 2024long

Some prior work has shown that LLMs perform well in NLG evaluation for different tasks. However, we discover that LLMs seem to confuse different evaluation criteria, which reduces their reliability. For further verification, we first consider avoiding issues of inconsistent conceptualization and vag…

2024

Benchmarking Knowledge Boundary for Large Language Models: A Different Perspective on Model Evaluation

ACL 2024long

In recent years, substantial advancements have been made in the development of large language models, achieving remarkable performance across diverse tasks.To evaluate the knowledge ability of language models, previous studies have proposed lots of benchmarks based on question-answering pairs.We arg…

2024

Better than Random: Reliable NLG Human Evaluation with Constrained Active Sampling

AAAI 2024technical

Human evaluation is viewed as a reliable evaluation method for NLG which is expensive and time-consuming. To save labor and costs, researchers usually perform human evaluation on a small subset of data sampled from the whole dataset in practice. However, different selection subsets will lead to diff…

2024

Contextual Modeling for Document-level ASR Error Correction

COLING 2024main

Contextual information, including the sentences in the same document and in other documents of the dataset, plays a crucial role in improving the accuracy of document-level ASR Error Correction (AEC), while most previous works ignore this. In this paper, we propose a context-aware method that utiliz…

Cited by 0SourcePDFScholar
2024

Cross Modal Training for ASR Error Correction with Contrastive Learning

ICASSP 2024accepted

ASR Error Correction (AEC) aims to post-process the output of ASR systems and further reduce the word error rate. In this paper, we propose a cross-modal training framework with contrastive learning on the AEC task. This framework enables a shared encoder-decoder model to learn text, pinyin (phoneme…

Cited by 0SourceScholar
2024

Defining and Detecting Vulnerability in Human Evaluation Guidelines: A Preliminary Study Towards Reliable NLG Evaluation

NAACL 2024long

Human evaluation serves as the gold standard for assessing the quality of Natural Language Generation (NLG) systems. Nevertheless, the evaluation guideline, as a pivotal element ensuring reliable and reproducible human assessment, has received limited attention. Our investigation revealed that only…

2024

Enhancing Large Language Models in Coding Through Multi-Perspective Self-Consistency

ACL 2024long

Large language models (LLMs) have exhibited remarkable ability in code generation. However, generating the correct solution in a single attempt still remains a challenge. Prior works utilize verification properties in software engineering to verify and re-rank solutions in a majority voting manner.…

2024

History Matters: Temporal Knowledge Editing in Large Language Model

AAAI 2024technical

The imperative task of revising or updating the knowledge stored within large language models arises from two distinct sources: intrinsic errors inherent in the model which should be corrected and outdated knowledge due to external shifts in the real world which should be updated. Prevailing efforts…

2024

Image Matters: A New Dataset and Empirical Study for Multimodal Hyperbole Detection

COLING 2024main

Hyperbole, or exaggeration, is a common linguistic phenomenon. The detection of hyperbole is an important part of understanding human expression. There have been several studies on hyperbole detection, but most of which focus on text modality only. However, with the development of social media, peop…

2024

Is Summary Useful or Not? An Extrinsic Human Evaluation of Text Summaries on Downstream Tasks

COLING 2024main

Research on automated text summarization typically uses human and automatic evaluation methods. While most recent studies focus on intrinsic evaluation, which assesses the general quality of summaries, e.g. coherence and informativeness, we concentrate on task-based extrinsic evaluation to determine…

2024

PaCoST: Paired Confidence Significance Testing for Benchmark Contamination Detection in Large Language Models

EMNLP 2024finding

Large language models (LLMs) are known to be trained on vast amounts of data, which may unintentionally or intentionally include data from commonly used benchmarks. This inclusion can lead to cheatingly high scores on model leaderboards, yet result in disappointing performance in real-world applicat…

Cited by 3SourcePDFScholar
2024

Selecting Large Language Model to Fine-tune via Rectified Scaling Law

ICML 2024poster

The ever-growing ecosystem of LLMs has posed a challenge in selecting the most appropriate pre-trained model to fine-tune amidst a sea of options. Given constrained resources, fine-tuning all models and making selections afterward is unrealistic. In this work, we formulate this resource-constrained…

2024

Style-Compress: An LLM-Based Prompt Compression Framework Considering Task-Specific Styles

EMNLP 2024finding

Prompt compression condenses contexts while maintaining their informativeness for different usage scenarios. It not only shortens the inference time and reduces computational costs during the usage of large language models, but also lowers expenses when using closed-source models. In a preliminary s…

2024

Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability

EMNLP 2024main

The evaluation of natural language generation (NLG) tasks is a significant and longstanding research area. With the recent emergence of powerful large language models (LLMs), some studies have turned to LLM-based automatic evaluation methods, which demonstrate great potential to become a new evaluat…

2023

A New Benchmark and Reverse Validation Method for Passage-level Hallucination Detection

EMNLP 2023long findings

Large Language Models (LLMs) have shown their ability to collaborate effectively with humans in real-world scenarios. However, LLMs are apt to generate hallucinations, i.e., makeup incorrect text and unverified information, which can cause significant damage when deployed for mission-critical tasks.…

Cited by 0SourcecodeScholar
2023

A New Dataset and Empirical Study for Sentence Simplification in Chinese

ACL 2023long

Sentence Simplification is a valuable technique that can benefit language learners and children a lot. However, current research focuses more on English sentence simplification. The development of Chinese sentence simplification is relatively slow due to the lack of data. To alleviate this limitatio…

2023

MIL-Decoding: Detoxifying Language Models at Token-Level via Multiple Instance Learning

ACL 2023long

Despite advances in large pre-trained neural language models, they are prone to generating toxic language, which brings security risks to their applications. We introduce MIL-Decoding, which detoxifies language models at token-level by interpolating it with a trained multiple instance learning (MIL)…

2023

New Datasets and Controllable Iterative Data Augmentation Method for Code-switching ASR Error Correction

EMNLP 2023long findings

With the wide use of automatic speech recognition(ASR) systems, researchers pay more attention to the ASR error correction task to improve the quality of recognition results. In particular, ASR in bilingual or multilingual settings, namely code-switching ASR, has greater challenges and research valu…

Cited by 0SourceScholar
2023

Reference Matters: Benchmarking Factual Error Correction for Dialogue Summarization with Fine-grained Evaluation Framework

ACL 2023long

Factuality is important to dialogue summarization. Factual error correction (FEC) of model-generated summaries is one way to improve factuality. Current FEC evaluation that relies on factuality metrics is not reliable and detailed enough. To address this problem, we are the first to manually annotat…

2023

SituatedGen: Incorporating Geographical and Temporal Contexts into Generative Commonsense Reasoning

NeurIPS 2023poster

Recently, commonsense reasoning in text generation has attracted much attention. Generative commonsense reasoning is the task that requires machines, given a group of keywords, to compose a single coherent sentence with commonsense plausibility. While existing datasets targeting generative commonsen…

2023

Teaching the Pre-trained Model to Generate Simple Texts for Text Simplification

ACL 2023findings

Randomly masking text spans in ordinary texts in the pre-training stage hardly allows models to acquire the ability to generate simple texts. It can hurt the performance of pre-trained models on text simplification tasks. In this paper, we propose a new continued pre-training strategy to teach the p…

2022

Diversifying Neural Text Generation with Part-of-Speech Guided Softmax and Sampling

COLING 2022main

Neural text generation models are likely to suffer from the low-diversity problem. Various decoding strategies and training-based methods have been proposed to promote diversity only by exploiting contextual features, but rarely do they consider incorporating syntactic structure clues. In this work,…

2022

Nearest Neighbor Knowledge Distillation for Neural Machine Translation

NAACL 2022long

k-nearest-neighbor machine translation (kNN-MT), proposed by Khandelwal et al. (2021), has achieved many state-of-the-art results in machine translation tasks. Although effective, kNN-MT requires conducting kNN searches through the large datastore for each decoding step during inference, prohibitive…

2021

Bridging the Domain Gap: Improve Informal Language Translation via Counterfactual Domain Adaptation

AAAI 2021technical

Despite the near-human performances already achieved on formal texts such as news articles, neural machine translation still has difficulty in dealing with "user-generated" texts that have diverse linguistic phenomena but lack large-scale high-quality parallel corpora. To address this problem, we pr…

Cited by 6SourcePDFScholar
2021

Revisiting Pivot-Based Paraphrase Generation: Language Is Not the Only Optional Pivot

EMNLP 2021main

Paraphrases refer to texts that convey the same meaning with different expression forms. Pivot-based methods, also known as the round-trip translation, have shown promising results in generating high-quality paraphrases. However, existing pivot-based methods all rely on language as the pivot, where…

2021

Towards Document-Level Paraphrase Generation with Sentence Rewriting and Reordering

EMNLP 2021finding

Paraphrase generation is an important task in natural language processing. Previous works focus on sentence-level paraphrase generation, while ignoring document-level paraphrase generation, which is a more challenging and valuable task. In this paper, we explore the task of document-level paraphrase…

2020

Improving Grammatical Error Correction with Data Augmentation by Editing Latent Representation

COLING 2020main

The incorporation of data augmentation method in grammatical error correction task has attracted much attention. However, existing data augmentation methods mainly apply noise to tokens, which leads to the lack of diversity of generated errors. In view of this, we propose a new data augmentation met…

2019

Controllable Unsupervised Text Attribute Transfer via Editing Entangled Latent Representation

NeurIPS 2019poster

Unsupervised text attribute transfer automatically transforms a text to alter a specific attribute (e.g. sentiment) without using any parallel data, while simultaneously preserving its attribute-independent content. The dominant approaches are trying to model the content-independent attribute separa…

2019

Generating Diverse and Descriptive Image Captions Using Visual Paraphrases

ICCV 2019poster

Recently there has been significant progress in image captioning with the help of deep learning. However, captions generated by current state-of-the-art models are still far from satisfactory, despite high scores in terms of conventional metrics such as BLEU and CIDEr. Human-written captions are div…

Cited by 50PDFScholar