← All conferences

ACL 2025 Accepted Papers

The full list of 3,086 papers accepted at ACL 2025 (Association for Computational Linguistics). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

Long: 1,602finding: 1,387Short: 97
  1. What’s the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token PatternsLong
  2. When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated ExplanationsLong
  3. When Benchmarks Talk: Re-Evaluating Code LLMs with Interactive Feedbackfinding
  4. When Claims Evolve: Evaluating and Enhancing the Robustness of Embedding Models Against Misinformation Editsfinding
  5. When Detection Fails: The Power of Fine-Tuned Models to Generate Human-Like Social Media Textfinding
  6. When GPT Spills the Tea: Comprehensive Assessment of Knowledge File Leakage in GPTsLong
  7. When Harry Meets Superman: The Role of The Interlocutor in Persona-Based Dialogue GenerationLong
  8. When Large Language Models Meet Speech: A Survey on Integration Approachesfinding
  9. When People are Floods: Analyzing Dehumanizing Metaphors in Immigration Discourse with Large Language ModelsLong
  10. When Should Dense Retrievers Be Updated in Evolving Corpora? Detecting Out-of-Distribution Corpora Using GradNormIRfinding
  11. When the LM misunderstood the human chuckled: Analyzing garden path effects in humans and language modelsLong
  12. When to Speak, When to Abstain: Contrastive Decoding with AbstentionLong
  13. Where Are We? Evaluating LLM Performance on African LanguagesLong
  14. Whether LLMs Know If They Know: Identifying Knowledge Boundaries via Debiased Historical In-Context Learningfinding
  15. WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher LearningLong
  16. Which Demographics do LLMs Default to During Annotation?Long
  17. Which Retain Set Matters for LLM Unlearning? A Case Study on Entity Unlearningfinding
  18. Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the AboveLong
  19. White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMsLong
  20. Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Modelsfinding
  21. Who Taught You That? Tracing Teachers in Model Distillationfinding
  22. Who Writes What: Unveiling the Impact of Author Roles on AI-generated Text DetectionLong
  23. Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User PersonasLong
  24. Why Are Positional Encodings Nonessential for Deep Autoregressive Transformers? A Petroglyph Revisitedfinding
  25. Why Multi-Interest Fairness Matters: Hypergraph Contrastive Multi-Interest Learning for Fair Conversational Recommender Systemfinding
  26. Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancementfinding
  27. Why Prompt Design Matters and Works: A Complexity Analysis of Prompt Search Space in LLMsLong
  28. Why Safeguarded Ships Run Aground? Aligned Large Language Models’ Safety Mechanisms Tend to Be Anchored in The Template RegionLong
  29. Why Uncertainty Estimation Methods Fall Short in RAG: An Axiomatic Analysisfinding
  30. Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understandingfinding
  31. WiCkeD: A Simple Method to Make Multiple Choice Benchmarks More ChallengingShort
  32. WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Chartsfinding
  33. WinSpot: GUI Grounding Benchmark with Multimodal Large Language ModelsShort
  34. WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communicationsfinding
  35. Wizard of Shopping: Target-Oriented E-commerce Dialogue Generation with Decision Tree BranchingLong
  36. Word Form Matters: LLMs’ Semantic Reconstruction under Typoglycemiafinding
  37. Word-Level Detection of Code-Mixed Hate Speech with Multilingual Domain Transferfinding
  38. Word2Passage: Word-level Importance Re-weighting for Query Expansionfinding
  39. Words of Warmth: Trust and Sociability Norms for over 26k English WordsLong
  40. World Knowledge Resolves Some Aspectual Ambiguityfinding
  41. World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task PlanningLong
  42. Worse than Random? An Embarrassingly Simple Probing Evaluation of Large Multimodal Models in Medical VQAfinding
  43. Writing Like the Best: Exemplar-Based Expository Text GenerationLong
  44. X-TURING: Towards an Enhanced and Efficient Turing Test for Long-Term Dialogue AgentsLong
  45. X-WebAgentBench: A Multilingual Interactive Web Benchmark for Evaluating Global Agentic Systemfinding
  46. XDAC: XAI-Driven Detection and Attribution of LLM-Generated News Comments in KoreanLong
  47. XFinBench: Benchmarking LLMs in Complex Financial Problem Solving and Reasoningfinding
  48. YESciEval: Robust LLM-as-a-Judge for Scientific Question AnsweringLong
  49. YinYang-Align: A new Benchmark for Competing Objectives and Introducing Multi-Objective Preference based Text-to-Image Alignmentfinding
  50. You need to MIMIC to get FAME: Solving Meeting Transcript Scarcity with Multi-Agent Conversationsfinding
  51. Your Language Model May Think Too Rigidly: Achieving Reasoning Consistency with Symmetry-Enhanced Trainingfinding
  52. Your Model is Overconfident, and Other Lies We Tell OurselvesLong
  53. YuLan-Mini: Pushing the Limits of Open Data-efficient Language ModelLong
  54. ZIPA: A family of efficient models for multilingual phone recognitionLong
  55. Zero-Shot Conversational Stance Detection: Dataset and Approachesfinding
  56. Zero-Shot Text-to-Speech for VietnameseShort
  57. ZeroDL: Zero-shot Distribution Learning for Text Clustering via Large Language Modelsfinding
  58. ZeroNER: Fueling Zero-Shot Named Entity Recognition via Entity Type Descriptionsfinding
  59. daDPO: Distribution-Aware DPO for Distilling Conversational Abilitiesfinding
  60. from Benign import Toxic: Jailbreaking the Language Model via Adversarial MetaphorsLong
  61. gMBA: Expression Semantic Guided Mixed Boolean-Arithmetic Deobfuscation Using Transformer Architecturesfinding
  62. iAgent: LLM Agent as a Shield between User and Recommender Systemsfinding
  63. iMOVE : Instance-Motion-Aware Video Understandingfinding
  64. iNews: A Multimodal Dataset for Modeling Personalized Affective Responses to NewsLong
  65. iQUEST: An Iterative Question-Guided Framework for Knowledge Base Question AnsweringLong
  66. mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpusfinding
  67. mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document UnderstandingLong
  68. mRAKL: Multilingual Retrieval-Augmented Knowledge Graph Construction for Low-Resourced Languagesfinding
  69. mStyleDistance: Multilingual Style Embeddings and their Evaluationfinding
  70. mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Datafinding
  71. nvAgent: Automated Data Visualization from Natural Language via Collaborative Agent WorkflowLong
  72. scRAG: Hybrid Retrieval-Augmented Generation for LLM-based Cross-Tissue Single-Cell Annotationfinding
  73. skLEP: A Slovak General Language Understanding Benchmarkfinding
  74. taz2024full: Analysing German Newspapers for Gender Bias and Discrimination across Decadesfinding
  75. uMedSum: A Unified Framework for Clinical Abstractive SummarizationLong
  76. ‘No’ Matters: Out-of-Distribution Detection in Multimodality Multi-Turn Interactive Dialogue Download PDFfinding
  77. “Give Me BF16 or Give Me Death”? Accuracy-Performance Trade-Offs in LLM QuantizationLong
  78. “I understand your perspective”: LLM Persuasion through the Lens of Communicative Action Theoryfinding
  79. “My life is miserable, have to sign 500 autographs everyday”: Exposing Humblebragging, the Brags in Disguisefinding
  80. “Well, Keep Thinking”: Enhancing LLM Reasoning with Adaptive Injection Decodingfinding
  81. “What do you call a dog that is incontrovertibly true? Dogma”: Testing LLM Generalization through HumorLong
  82. “Yes, My LoRD.” Guiding Language Model Extraction with Locality Reinforced DistillationLong
  83. “You are Beautiful, Body Image Stereotypes are Ugly!” BIStereo: A Benchmark to Measure Body Image Stereotypes in Language Modelsfinding
  84. 𝒜3: Automatic Alignment Framework for Attributed Text GenerationLong
  85. 𝛿-Stance: A Large-Scale Real World Dataset of Stances in Legal ArgumentationLong
  86. 𝜙-Decoding: Adaptive Foresight Sampling for Balanced Inference-Time Exploration and ExploitationLong

Looking for submission deadlines instead? See the conference deadline calendar.