← All conferences

EMNLP 2025 Accepted Papers

The full list of 3,211 papers accepted at EMNLP 2025 (Conference on Empirical Methods in Natural Language Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

accepted: 3,211
  1. Universal Acoustic Adversarial Attacks for Flexible Control of Speech-LLMsaccepted
  2. UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessmentaccepted
  3. Unlearning vs. Obfuscation: Are We Truly Removing Knowledge?accepted
  4. Unleashing the Reasoning Potential of LLMs by Critique Fine-Tuning on One Problemaccepted
  5. Unlocking Legal Knowledge: A Multilingual Dataset for Judicial Summarization in Switzerlandaccepted
  6. Unlocking Smarter Device Control: Foresighted Planning with a World Model-Driven Code Execution Approachaccepted
  7. Unlocking the Effectiveness of LoRA-FP for Seamless Transfer Implantation of Fingerprints in Downstream Modelsaccepted
  8. Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question Answeringaccepted
  9. Unmasking Fake Careers: Detecting Machine-Generated Career Trajectories via Multi-layer Heterogeneous Graphsaccepted
  10. Unpacking Let Alone: Human-Scale Models Generalize to a Rare Construction in Form but not Meaningaccepted
  11. Unraveling Interwoven Roles of Large Language Models in Authorship Privacy: Obfuscation, Mimicking, and Verificationaccepted
  12. Unraveling Misinformation Propagation in LLM Reasoningaccepted
  13. Unstructured Evidence Attribution for Long Context Query Focused Summarizationaccepted
  14. Unsupervised Concept Vector Extraction for Bias Control in LLMsaccepted
  15. Unsupervised Hallucination Detection by Inspecting Reasoning Processesaccepted
  16. Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreementaccepted
  17. Unveiling Internal Reasoning Modes in LLMs: A Deep Dive into Latent Reasoning vs. Factual Shortcuts with Attribute Rate Ratioaccepted
  18. Unveiling Multimodal Processing: Exploring Activation Patterns in Multimodal LLMs for Interpretability and Efficiencyaccepted
  19. Unveiling the Response of Large Vision-Language Models to Visually Absent Tokensaccepted
  20. Uplift-RAG: Uplift-Driven Knowledge Preference Alignment for Retrieval-Augmented Generationaccepted
  21. UrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and Benchmarkingaccepted
  22. Use Random Selection for Now: Investigation of Few-Shot Selection Strategies in LLM-based Text Augmentationaccepted
  23. User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning Signalaccepted
  24. Using tournaments to calculate AUROC for zero-shot classification with LLMsaccepted
  25. Utility-Focused LLM Annotation for Retrieval and Retrieval-Augmented Generationaccepted
  26. V-SEAM: Visual Semantic Editing and Attention Modulating for Causal Interpretability of Vision-Language Modelsaccepted
  27. V-VAE: A Variational Auto Encoding Framework Towards Fine-Grained Control over Human-Like Chataccepted
  28. VC4VG: Optimizing Video Captions for Text-to-Video Generationaccepted
  29. VCSearch: Bridging the Gap Between Well-Defined and Ill-Defined Problems in Mathematical Reasoningaccepted
  30. VEHME: A Vision-Language Model For Evaluating Handwritten Mathematics Expressionsaccepted
  31. VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captionsaccepted
  32. VERITAS: Leveraging Vision Priors and Expert Fusion to Improve Multimodal Dataaccepted
  33. VIBE: Can a VLM Read the Room?accepted
  34. VISaGE: Understanding Visual Generics and Exceptionsaccepted
  35. VIVA+: Human-Centered Situational Decision-Makingaccepted
  36. VLA-Mark: A cross modal watermark for large vision-language alignment modelsaccepted
  37. VLASCD: A Visual Language Action Model for Simultaneous Chatting and Decision Makingaccepted
  38. VLM Is a Strong Reranker: Advancing Multimodal Retrieval-augmented Generation via Knowledge-enhanced Reranking and Noise-injected Trainingaccepted
  39. VLP: Vision-Language Preference Learning for Embodied Manipulationaccepted
  40. VQA-Augmented Machine Translation with Cross-Modal Contrastive Learningaccepted
  41. VRoPE: Rotary Position Embedding for Video Large Language Modelsaccepted
  42. Value Profiles for Encoding Human Variationaccepted
  43. Variance Sensitivity Induces Attention Entropy Collapse and Instability in Transformersaccepted
  44. VehicleWorld: A Highly Integrated Multi-Device Environment for Intelligent Vehicle Interactionaccepted
  45. VerIF: Verification Engineering for Reinforcement Learning in Instruction Followingaccepted
  46. VeriFact: Enhancing Long-Form Factuality Evaluation with Refined Fact Extraction and Reference Factsaccepted
  47. VeriFastScore: Speeding up long-form factuality evaluationaccepted
  48. VeriLocc: End-to-End Cross-Architecture Register Allocation via LLMaccepted
  49. VerifiAgent: a Unified Verification Agent in Language Model Reasoningaccepted
  50. VersaTune: An Efficient Data Composition Framework for Training Multi-Capability LLMsaccepted
  51. Versatile Framework for Song Generation with Prompt-based Controlaccepted
  52. ViClaim: A Multilingual Multilabel Dataset for Automatic Claim Detection in Videosaccepted
  53. ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agentsaccepted
  54. ViFT: Towards Visual Instruction-Free Fine-tuning for Large Vision-Language Modelsaccepted
  55. ViLBench: A Suite for Vision-Language Process Reward Modelingaccepted
  56. ViPE: Visual Perception in Parameter Space for Efficient Video-Language Understandingaccepted
  57. Viability of Machine Translation for Healthcare in Low-Resourced Languagesaccepted
  58. Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Modelsaccepted
  59. Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoningaccepted
  60. Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoningaccepted
  61. Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agentsaccepted
  62. VideoEraser: Concept Erasure in Text-to-Video Diffusion Modelsaccepted
  63. VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Formataccepted
  64. VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignmentaccepted
  65. VisBias: Measuring Explicit and Implicit Social Biases in Vision Language Modelsaccepted
  66. VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Modelsaccepted
  67. VisCoder: Fine-Tuning LLMs for Executable Python Visualization Code Generationaccepted
  68. VisEscape: A Benchmark for Evaluating Exploration-driven Decision-making in Virtual Escape Roomsaccepted
  69. VisFinEval: A Scenario-Driven Chinese Multimodal Benchmark for Holistic Financial Understandingaccepted
  70. VisiPruner: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMsaccepted
  71. Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptionsaccepted
  72. Vision-and-Language Navigation with Analogical Textual Descriptions in LLMsaccepted
  73. VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for Minecraftaccepted
  74. Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injectionaccepted
  75. Visual Program Distillation with Template-Based Augmentationaccepted
  76. Visual Self-Refinement for Autoregressive Modelsaccepted
  77. Visual-Aware Speech Recognition for Noisy Scenariosaccepted
  78. VisualEDU: A Benchmark for Assessing Coding and Visual Comprehension through Educational Problem-Solving Video Generationaccepted
  79. VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Searchaccepted
  80. VocalNet: Speech LLMs with Multi-Token Prediction for Faster and High-Quality Generationaccepted
  81. Voice of a Continent: Mapping Africa’s Speech Technology Frontieraccepted
  82. VoiceBBQ: Investigating Effect of Content and Acoustics in Social Bias of Spoken Language Modelaccepted
  83. VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editingaccepted
  84. WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classificationaccepted
  85. Wait, We Don’t Need to “Wait”! Removing Thinking Tokens Improves Reasoning Efficiencyaccepted
  86. Walk and Read Less: Improving the Efficiency of Vision-and-Language Navigation via Tuning-Free Multimodal Token Pruningaccepted
  87. WangchanThaiInstruct: An instruction-following Dataset for Culture-Aware, Multitask, and Multi-domain Evaluation in Thaiaccepted
  88. Warm Up Before You Train: Unlocking General Reasoning in Resource-Constrained Settingsaccepted
  89. Waste-Bench: A Comprehensive Benchmark for Evaluating VLLMs in Cluttered Environmentsaccepted
  90. Watermark Smoothing Attacks against Language Modelsaccepted
  91. Watermark under Fire: A Robustness Evaluation of LLM Watermarkingaccepted
  92. Watermarking for Factuality: Guiding Vision-Language Models Toward Truth via Tri-layer Contrastive Decodingaccepted
  93. Watermarking with Low-Entropy POS-Guided Token Partitioning and Z-Score-Driven Dynamic Bias for Large Language Modelsaccepted
  94. We Argue to Agree: Towards Personality-Driven Argumentation-Based Negotiation Dialogue Systems for Tourismaccepted
  95. We Need to Measure Data Diversity in NLP — Better and Broaderaccepted
  96. We Politely Insist: Your LLM Must Learn the Persian Art of Taarofaccepted
  97. Weak2Wise: An Automated, Lightweight Framework for Weak-LLM-Friendly Reasoning Synthesisaccepted
  98. Weaver: Interweaving SQL and LLM for Table Reasoningaccepted
  99. Web Intellectual Property at Risk: Preventing Unauthorized Real-Time Retrieval by Large Language Modelsaccepted
  100. WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learningaccepted
  101. WebCoT: Enhancing Web Agent Reasoning by Reconstructing Chain-of-Thought in Reflection, Branching, and Rollbackaccepted
  102. WebEvolver: Enhancing Web Agent Self-Improvement with Co-evolving World Modelaccepted
  103. WebInject: Prompt Injection Attack to Web Agentsaccepted
  104. WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generationaccepted
  105. Weight-Aware Activation Sparsity with Constrained Bayesian Optimization Scheduling for Large Language Modelsaccepted
  106. Weights-Rotated Preference Optimization for Large Language Modelsaccepted
  107. What Do Indonesians Really Need from Language Technology? A Nationwide Surveyaccepted
  108. What Has Been Lost with Synthetic Evaluation?accepted
  109. What Makes a Good Reasoning Chain? Uncovering Structural Patterns in Long Chain-of-Thought Reasoningaccepted
  110. What Makes for Good Image Captions?accepted
  111. What Media Frames Reveal About Stance: A Dataset and Study about Memes in Climate Change Discourseaccepted
  112. What You Read Isn’t What You Hear: Linguistic Sensitivity in Deepfake Speech Detectionaccepted
  113. What You See is What You Ask: Evaluating Audio Descriptionsaccepted
  114. What are Foundation Models Cooking in the Post-Soviet World?accepted
  115. What data should I include in my POS tagging training set?accepted
  116. What if Othello-Playing Language Models Could See?accepted
  117. What’s Not Said Still Hurts: A Description-Based Evaluation Framework for Measuring Social Bias in LLMsaccepted
  118. What’s in a prompt? Language models encode literary style in prompt embeddingsaccepted
  119. When Allies Turn Foes: Exploring Group Characteristics of LLM-Based Multi-Agent Collaborative Systems Under Adversarial Attacksaccepted
  120. When Annotators Disagree, Topology Explains: Mapper, a Topological Tool for Exploring Text Embedding Geometry and Ambiguityaccepted
  121. When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Modelsaccepted
  122. When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMsaccepted
  123. When Format Changes Meaning: Investigating Semantic Inconsistency of Large Language Modelsaccepted
  124. When Instructions Multiply: Measuring and Estimating LLM Capabilities of Multiple Instructions Followingaccepted
  125. When Inverse Data Outperforms: Exploring the Pitfalls of Mixed Data in Multi-Stage Fine-Tuningaccepted
  126. When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMsaccepted
  127. When Long Helps Short: How Context Length in Supervised Fine-tuning Affects Behavior of Large Language Modelsaccepted
  128. When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQAaccepted
  129. When Models Reason in Your Language: Controlling Thinking Language Comes at the Cost of Accuracyaccepted
  130. When Personalization Meets Reality: A Multi-Faceted Analysis of Personalized Preference Learningaccepted
  131. When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMsaccepted
  132. When Truthful Representations Flip Under Deceptive Instructions?accepted
  133. When Words Smile: Generating Diverse Emotional Facial Expressions from Textaccepted
  134. When to Continue Thinking: Adaptive Thinking Mode Switching for Efficient Reasoningaccepted
  135. Where Confabulation Lives: Latent Feature Discovery in LLMsaccepted
  136. Where Did That Come From? Sentence-Level Error-Tolerant Attributionaccepted
  137. Where Fact Ends and Fairness Begins: Redefining AI Bias Evaluation through Cognitive Biasesaccepted
  138. Where to show Demos in Your Prompt: A Positional Bias of In-Context Learningaccepted
  139. Which Word Orders Facilitate Length Generalization in LMs? An Investigation with GCG-Based Artificial Languagesaccepted
  140. Whisper-UT: A Unified Translation Framework for Speech and Textaccepted
  141. Who Holds the Pen? Caricature and Perspective in LLM Retellings of Historyaccepted
  142. Who Speaks Matters: Analysing the Influence of the Speaker’s Linguistic Identity on Hate Classificationaccepted
  143. Who is in the Spotlight: The Hidden Bias Undermining Multimodal Retrieval-Augmented Generationaccepted
  144. Who’s the Author? How Explanations Impact User Reliance in AI-Assisted Authorship Attributionaccepted
  145. Why Do Some Inputs Break Low-Bit LLM Quantization?accepted
  146. Why Stop at One Error? Benchmarking LLMs as Data Science Code Debuggers for Multi-Hop and Multi-Bug Errorsaccepted
  147. Why We Feel What We Feel: Joint Detection of Emotions and Their Opinion Triggers in E-commerceaccepted
  148. Why and How LLMs Benefit from Knowledge Introspection in Commonsense Reasoningaccepted
  149. WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild?accepted
  150. WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music Reasoningaccepted
  151. Will Annotators Disagree? Identifying Subjectivity in Value-Laden Argumentsaccepted
  152. Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QAaccepted
  153. WojoodRelations: Arabic Relation Extraction Corpus and Modelingaccepted
  154. Women, Infamous, and Exotic Beings: A Comparative Study of Honorific Usages in Wikipedia and LLMs for Bengali and Hindiaccepted
  155. Word Salad Chopper: Reasoning Models Waste A Ton Of Decoding Budget On Useless Repetitions, Self-Knowinglyaccepted
  156. Words Like Knives: Backstory-Personalized Modeling and Detection of Violent Communicationaccepted
  157. X-Boundary: Establishing Exact Safety Boundary to Shield LLMs from Jailbreak Attacks without Compromising Usabilityaccepted
  158. X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought Reasoningaccepted
  159. X-FLoRA: Cross-modal Federated Learning with Modality-expert LoRA for Medical VQAaccepted
  160. X-LeBench: A Benchmark for Extremely Long Egocentric Video Understandingaccepted
  161. XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoMLaccepted
  162. XL-Suite: Cross-Lingual Synthetic Training and Evaluation Data for Open-Ended Generationaccepted
  163. XLQA: A Benchmark for Locale-Aware Multilingual Open-Domain Question Answeringaccepted
  164. XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compressionaccepted
  165. XRAG: Cross-lingual Retrieval-Augmented Generationaccepted
  166. XTRA: Cross-Lingual Topic Modeling with Topic and Representation Alignmentsaccepted
  167. You Are What You Train: Effects of Data Composition on Training Context-aware Machine Translation Modelsaccepted
  168. You Only Use Reactive Attention Slice When Retrieving From Long Contextaccepted
  169. Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectorsaccepted
  170. Your Mileage May Vary: How Empathy and Demographics Shape Human Preferences in LLM Responsesaccepted
  171. Your RAG is Unfair: Exposing Fairness Vulnerabilities in Retrieval-Augmented Generation via Backdoor Attacksaccepted
  172. Your Semantic-Independent Watermark is Fragile: A Semantic Perturbation Attack against EaaS Watermarkaccepted
  173. ZEBRA: Leveraging Model-Behavioral Knowledge for Zero-Annotation Preference Dataset Constructionaccepted
  174. ZERA: Zero-init Instruction Evolving Refinement Agent – From Zero Instructions to Structured Prompts via Principle-based Optimizationaccepted
  175. ZOGRASCOPE: A New Benchmark for Semantic Parsing over Property Graphsaccepted
  176. Zero-Shot Contextual Embeddings via Offline Synthetic Corpus Generationaccepted
  177. Zero-Shot Cross-Domain Aspect-Based Sentiment Analysis via Domain-Contextualized Chain-of-Thought Reasoningaccepted
  178. Zero-Shot Defense Against Toxic Images via Inherent Multimodal Alignment in LVLMsaccepted
  179. Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Modelsaccepted
  180. Zero-Shot Privacy-Aware Text Rewriting via Iterative Tree Searchaccepted
  181. Zero-shot Cross-lingual NER via Mitigating Language Difference: An Entity-aligned Translation Perspectiveaccepted
  182. Zero-shot Graph Reasoning via Retrieval Augmented Framework with LLMsaccepted
  183. Zero-shot Multimodal Document Retrieval via Cross-modal Question Generationaccepted
  184. Zipf’s and Heaps’ Laws for Tokens and LLM-generated Textsaccepted
  185. ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Explorationaccepted
  186. [MASK]ED - Language Modeling for Explainable Classification and Disentangling of Socially Unacceptable Discourse.accepted
  187. cAST: Enhancing Code Retrieval-Augmented Generation with Structural Chunking via Abstract Syntax Treeaccepted
  188. fLSA: Learning Semantic Structures in Document Collections Using Foundation Modelsaccepted
  189. iKnow-audio: Integrating Knowledge Graphs with Audio-Language Modelsaccepted
  190. iTool: Reinforced Fine-Tuning with Dynamic Deficiency Calibration for Advanced Tool Useaccepted
  191. iVISPAR — An Interactive Visual-Spatial Reasoning Benchmark for VLMsaccepted
  192. mrCAD: Multimodal Communication to Refine Computer-aided Designsaccepted
  193. pFedGPT: Hierarchically Optimizing LoRA Aggregation Weights for Personalized Federated GPT Modelsaccepted
  194. pFedRAG: A Personalized Federated Retrieval-Augmented Generation System with Depth-Adaptive Tiered Embedding Tuningaccepted
  195. polyBART: A Chemical Linguist for Polymer Property Prediction and Generative Designaccepted
  196. reWordBench: Benchmarking and Improving the Robustness of Reward Models with Transformed Inputsaccepted
  197. s1: Simple test-time scalingaccepted
  198. s3: You Don’t Need That Much Data to Train a Search Agent via RLaccepted
  199. seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMsaccepted
  200. so much depends / upon / a whitespace: Why Whitespace Matters for Poets and LLMsaccepted
  201. sudoLLM: On Multi-role Alignment of Language Modelsaccepted
  202. xCoRe: Cross-context Coreference Resolutionaccepted
  203. zFLoRA: Zero-Latency Fused Low-Rank Adaptersaccepted
  204. ‘Hello, World!’: Making GNNs Talk with LLMsaccepted
  205. ‘Rich Dad, Poor Lad’: How do Large Language Models Contextualize Socioeconomic Factors in College Admission ?accepted
  206. “Feels Feminine to Me”: Understanding Perceived Gendered Style through Human Annotationsaccepted
  207. “Going to a trap house” conveys more fear than “Going to a mall”: Benchmarking Emotion Context Sensitivity for LLMsaccepted
  208. “I’ve Decided to Leak”: Probing Internals Behind Prompt Leakage Intentsaccepted
  209. “Mm, Wat?” Detecting Other-initiated Repair Requests in Dialogueaccepted
  210. “What’s Up, Doc?”: Analyzing How Users Seek Health Information in Large-Scale Conversational AI Datasetsaccepted
  211. “Where Does This Strange Smell Come from?”: Enabling Conversational Interfaces for Artificial Olfactionaccepted

Looking for submission deadlines instead? See the conference deadline calendar.