ACL 2025 Accepted Papers
The full list of 3,086 papers accepted at ACL 2025 (Association for Computational Linguistics). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
Long: 1,602finding: 1,387Short: 97
- What’s the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token PatternsLong
- When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated ExplanationsLong
- When Benchmarks Talk: Re-Evaluating Code LLMs with Interactive Feedbackfinding
- When Claims Evolve: Evaluating and Enhancing the Robustness of Embedding Models Against Misinformation Editsfinding
- When Detection Fails: The Power of Fine-Tuned Models to Generate Human-Like Social Media Textfinding
- When GPT Spills the Tea: Comprehensive Assessment of Knowledge File Leakage in GPTsLong
- When Harry Meets Superman: The Role of The Interlocutor in Persona-Based Dialogue GenerationLong
- When Large Language Models Meet Speech: A Survey on Integration Approachesfinding
- When People are Floods: Analyzing Dehumanizing Metaphors in Immigration Discourse with Large Language ModelsLong
- When Should Dense Retrievers Be Updated in Evolving Corpora? Detecting Out-of-Distribution Corpora Using GradNormIRfinding
- When the LM misunderstood the human chuckled: Analyzing garden path effects in humans and language modelsLong
- When to Speak, When to Abstain: Contrastive Decoding with AbstentionLong
- Where Are We? Evaluating LLM Performance on African LanguagesLong
- Whether LLMs Know If They Know: Identifying Knowledge Boundaries via Debiased Historical In-Context Learningfinding
- WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher LearningLong
- Which Demographics do LLMs Default to During Annotation?Long
- Which Retain Set Matters for LLM Unlearning? A Case Study on Entity Unlearningfinding
- Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the AboveLong
- White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMsLong
- Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Modelsfinding
- Who Taught You That? Tracing Teachers in Model Distillationfinding
- Who Writes What: Unveiling the Impact of Author Roles on AI-generated Text DetectionLong
- Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User PersonasLong
- Why Are Positional Encodings Nonessential for Deep Autoregressive Transformers? A Petroglyph Revisitedfinding
- Why Multi-Interest Fairness Matters: Hypergraph Contrastive Multi-Interest Learning for Fair Conversational Recommender Systemfinding
- Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancementfinding
- Why Prompt Design Matters and Works: A Complexity Analysis of Prompt Search Space in LLMsLong
- Why Safeguarded Ships Run Aground? Aligned Large Language Models’ Safety Mechanisms Tend to Be Anchored in The Template RegionLong
- Why Uncertainty Estimation Methods Fall Short in RAG: An Axiomatic Analysisfinding
- Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understandingfinding
- WiCkeD: A Simple Method to Make Multiple Choice Benchmarks More ChallengingShort
- WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Chartsfinding
- WinSpot: GUI Grounding Benchmark with Multimodal Large Language ModelsShort
- WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communicationsfinding
- Wizard of Shopping: Target-Oriented E-commerce Dialogue Generation with Decision Tree BranchingLong
- Word Form Matters: LLMs’ Semantic Reconstruction under Typoglycemiafinding
- Word-Level Detection of Code-Mixed Hate Speech with Multilingual Domain Transferfinding
- Word2Passage: Word-level Importance Re-weighting for Query Expansionfinding
- Words of Warmth: Trust and Sociability Norms for over 26k English WordsLong
- World Knowledge Resolves Some Aspectual Ambiguityfinding
- World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task PlanningLong
- Worse than Random? An Embarrassingly Simple Probing Evaluation of Large Multimodal Models in Medical VQAfinding
- Writing Like the Best: Exemplar-Based Expository Text GenerationLong
- X-TURING: Towards an Enhanced and Efficient Turing Test for Long-Term Dialogue AgentsLong
- X-WebAgentBench: A Multilingual Interactive Web Benchmark for Evaluating Global Agentic Systemfinding
- XDAC: XAI-Driven Detection and Attribution of LLM-Generated News Comments in KoreanLong
- XFinBench: Benchmarking LLMs in Complex Financial Problem Solving and Reasoningfinding
- YESciEval: Robust LLM-as-a-Judge for Scientific Question AnsweringLong
- YinYang-Align: A new Benchmark for Competing Objectives and Introducing Multi-Objective Preference based Text-to-Image Alignmentfinding
- You need to MIMIC to get FAME: Solving Meeting Transcript Scarcity with Multi-Agent Conversationsfinding
- Your Language Model May Think Too Rigidly: Achieving Reasoning Consistency with Symmetry-Enhanced Trainingfinding
- Your Model is Overconfident, and Other Lies We Tell OurselvesLong
- YuLan-Mini: Pushing the Limits of Open Data-efficient Language ModelLong
- ZIPA: A family of efficient models for multilingual phone recognitionLong
- Zero-Shot Conversational Stance Detection: Dataset and Approachesfinding
- Zero-Shot Text-to-Speech for VietnameseShort
- ZeroDL: Zero-shot Distribution Learning for Text Clustering via Large Language Modelsfinding
- ZeroNER: Fueling Zero-Shot Named Entity Recognition via Entity Type Descriptionsfinding
- daDPO: Distribution-Aware DPO for Distilling Conversational Abilitiesfinding
- from Benign import Toxic: Jailbreaking the Language Model via Adversarial MetaphorsLong
- gMBA: Expression Semantic Guided Mixed Boolean-Arithmetic Deobfuscation Using Transformer Architecturesfinding
- iAgent: LLM Agent as a Shield between User and Recommender Systemsfinding
- iMOVE : Instance-Motion-Aware Video Understandingfinding
- iNews: A Multimodal Dataset for Modeling Personalized Affective Responses to NewsLong
- iQUEST: An Iterative Question-Guided Framework for Knowledge Base Question AnsweringLong
- mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpusfinding
- mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document UnderstandingLong
- mRAKL: Multilingual Retrieval-Augmented Knowledge Graph Construction for Low-Resourced Languagesfinding
- mStyleDistance: Multilingual Style Embeddings and their Evaluationfinding
- mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Datafinding
- nvAgent: Automated Data Visualization from Natural Language via Collaborative Agent WorkflowLong
- scRAG: Hybrid Retrieval-Augmented Generation for LLM-based Cross-Tissue Single-Cell Annotationfinding
- skLEP: A Slovak General Language Understanding Benchmarkfinding
- taz2024full: Analysing German Newspapers for Gender Bias and Discrimination across Decadesfinding
- uMedSum: A Unified Framework for Clinical Abstractive SummarizationLong
- ‘No’ Matters: Out-of-Distribution Detection in Multimodality Multi-Turn Interactive Dialogue Download PDFfinding
- “Give Me BF16 or Give Me Death”? Accuracy-Performance Trade-Offs in LLM QuantizationLong
- “I understand your perspective”: LLM Persuasion through the Lens of Communicative Action Theoryfinding
- “My life is miserable, have to sign 500 autographs everyday”: Exposing Humblebragging, the Brags in Disguisefinding
- “Well, Keep Thinking”: Enhancing LLM Reasoning with Adaptive Injection Decodingfinding
- “What do you call a dog that is incontrovertibly true? Dogma”: Testing LLM Generalization through HumorLong
- “Yes, My LoRD.” Guiding Language Model Extraction with Locality Reinforced DistillationLong
- “You are Beautiful, Body Image Stereotypes are Ugly!” BIStereo: A Benchmark to Measure Body Image Stereotypes in Language Modelsfinding
- 𝒜3: Automatic Alignment Framework for Attributed Text GenerationLong
- 𝛿-Stance: A Large-Scale Real World Dataset of Stances in Legal ArgumentationLong
- 𝜙-Decoding: Adaptive Foresight Sampling for Balanced Inference-Time Exploration and ExploitationLong
ACL accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.