EMNLP 2025 Accepted Papers
The full list of 3,211 papers accepted at EMNLP 2025 (Conference on Empirical Methods in Natural Language Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
accepted: 3,211
- Universal Acoustic Adversarial Attacks for Flexible Control of Speech-LLMsaccepted
- UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessmentaccepted
- Unlearning vs. Obfuscation: Are We Truly Removing Knowledge?accepted
- Unleashing the Reasoning Potential of LLMs by Critique Fine-Tuning on One Problemaccepted
- Unlocking Legal Knowledge: A Multilingual Dataset for Judicial Summarization in Switzerlandaccepted
- Unlocking Smarter Device Control: Foresighted Planning with a World Model-Driven Code Execution Approachaccepted
- Unlocking the Effectiveness of LoRA-FP for Seamless Transfer Implantation of Fingerprints in Downstream Modelsaccepted
- Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question Answeringaccepted
- Unmasking Fake Careers: Detecting Machine-Generated Career Trajectories via Multi-layer Heterogeneous Graphsaccepted
- Unpacking Let Alone: Human-Scale Models Generalize to a Rare Construction in Form but not Meaningaccepted
- Unraveling Interwoven Roles of Large Language Models in Authorship Privacy: Obfuscation, Mimicking, and Verificationaccepted
- Unraveling Misinformation Propagation in LLM Reasoningaccepted
- Unstructured Evidence Attribution for Long Context Query Focused Summarizationaccepted
- Unsupervised Concept Vector Extraction for Bias Control in LLMsaccepted
- Unsupervised Hallucination Detection by Inspecting Reasoning Processesaccepted
- Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreementaccepted
- Unveiling Internal Reasoning Modes in LLMs: A Deep Dive into Latent Reasoning vs. Factual Shortcuts with Attribute Rate Ratioaccepted
- Unveiling Multimodal Processing: Exploring Activation Patterns in Multimodal LLMs for Interpretability and Efficiencyaccepted
- Unveiling the Response of Large Vision-Language Models to Visually Absent Tokensaccepted
- Uplift-RAG: Uplift-Driven Knowledge Preference Alignment for Retrieval-Augmented Generationaccepted
- UrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and Benchmarkingaccepted
- Use Random Selection for Now: Investigation of Few-Shot Selection Strategies in LLM-based Text Augmentationaccepted
- User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning Signalaccepted
- Using tournaments to calculate AUROC for zero-shot classification with LLMsaccepted
- Utility-Focused LLM Annotation for Retrieval and Retrieval-Augmented Generationaccepted
- V-SEAM: Visual Semantic Editing and Attention Modulating for Causal Interpretability of Vision-Language Modelsaccepted
- V-VAE: A Variational Auto Encoding Framework Towards Fine-Grained Control over Human-Like Chataccepted
- VC4VG: Optimizing Video Captions for Text-to-Video Generationaccepted
- VCSearch: Bridging the Gap Between Well-Defined and Ill-Defined Problems in Mathematical Reasoningaccepted
- VEHME: A Vision-Language Model For Evaluating Handwritten Mathematics Expressionsaccepted
- VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captionsaccepted
- VERITAS: Leveraging Vision Priors and Expert Fusion to Improve Multimodal Dataaccepted
- VIBE: Can a VLM Read the Room?accepted
- VISaGE: Understanding Visual Generics and Exceptionsaccepted
- VIVA+: Human-Centered Situational Decision-Makingaccepted
- VLA-Mark: A cross modal watermark for large vision-language alignment modelsaccepted
- VLASCD: A Visual Language Action Model for Simultaneous Chatting and Decision Makingaccepted
- VLM Is a Strong Reranker: Advancing Multimodal Retrieval-augmented Generation via Knowledge-enhanced Reranking and Noise-injected Trainingaccepted
- VLP: Vision-Language Preference Learning for Embodied Manipulationaccepted
- VQA-Augmented Machine Translation with Cross-Modal Contrastive Learningaccepted
- VRoPE: Rotary Position Embedding for Video Large Language Modelsaccepted
- Value Profiles for Encoding Human Variationaccepted
- Variance Sensitivity Induces Attention Entropy Collapse and Instability in Transformersaccepted
- VehicleWorld: A Highly Integrated Multi-Device Environment for Intelligent Vehicle Interactionaccepted
- VerIF: Verification Engineering for Reinforcement Learning in Instruction Followingaccepted
- VeriFact: Enhancing Long-Form Factuality Evaluation with Refined Fact Extraction and Reference Factsaccepted
- VeriFastScore: Speeding up long-form factuality evaluationaccepted
- VeriLocc: End-to-End Cross-Architecture Register Allocation via LLMaccepted
- VerifiAgent: a Unified Verification Agent in Language Model Reasoningaccepted
- VersaTune: An Efficient Data Composition Framework for Training Multi-Capability LLMsaccepted
- Versatile Framework for Song Generation with Prompt-based Controlaccepted
- ViClaim: A Multilingual Multilabel Dataset for Automatic Claim Detection in Videosaccepted
- ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agentsaccepted
- ViFT: Towards Visual Instruction-Free Fine-tuning for Large Vision-Language Modelsaccepted
- ViLBench: A Suite for Vision-Language Process Reward Modelingaccepted
- ViPE: Visual Perception in Parameter Space for Efficient Video-Language Understandingaccepted
- Viability of Machine Translation for Healthcare in Low-Resourced Languagesaccepted
- Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Modelsaccepted
- Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoningaccepted
- Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoningaccepted
- Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agentsaccepted
- VideoEraser: Concept Erasure in Text-to-Video Diffusion Modelsaccepted
- VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Formataccepted
- VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignmentaccepted
- VisBias: Measuring Explicit and Implicit Social Biases in Vision Language Modelsaccepted
- VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Modelsaccepted
- VisCoder: Fine-Tuning LLMs for Executable Python Visualization Code Generationaccepted
- VisEscape: A Benchmark for Evaluating Exploration-driven Decision-making in Virtual Escape Roomsaccepted
- VisFinEval: A Scenario-Driven Chinese Multimodal Benchmark for Holistic Financial Understandingaccepted
- VisiPruner: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMsaccepted
- Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptionsaccepted
- Vision-and-Language Navigation with Analogical Textual Descriptions in LLMsaccepted
- VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for Minecraftaccepted
- Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injectionaccepted
- Visual Program Distillation with Template-Based Augmentationaccepted
- Visual Self-Refinement for Autoregressive Modelsaccepted
- Visual-Aware Speech Recognition for Noisy Scenariosaccepted
- VisualEDU: A Benchmark for Assessing Coding and Visual Comprehension through Educational Problem-Solving Video Generationaccepted
- VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Searchaccepted
- VocalNet: Speech LLMs with Multi-Token Prediction for Faster and High-Quality Generationaccepted
- Voice of a Continent: Mapping Africa’s Speech Technology Frontieraccepted
- VoiceBBQ: Investigating Effect of Content and Acoustics in Social Bias of Spoken Language Modelaccepted
- VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editingaccepted
- WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classificationaccepted
- Wait, We Don’t Need to “Wait”! Removing Thinking Tokens Improves Reasoning Efficiencyaccepted
- Walk and Read Less: Improving the Efficiency of Vision-and-Language Navigation via Tuning-Free Multimodal Token Pruningaccepted
- WangchanThaiInstruct: An instruction-following Dataset for Culture-Aware, Multitask, and Multi-domain Evaluation in Thaiaccepted
- Warm Up Before You Train: Unlocking General Reasoning in Resource-Constrained Settingsaccepted
- Waste-Bench: A Comprehensive Benchmark for Evaluating VLLMs in Cluttered Environmentsaccepted
- Watermark Smoothing Attacks against Language Modelsaccepted
- Watermark under Fire: A Robustness Evaluation of LLM Watermarkingaccepted
- Watermarking for Factuality: Guiding Vision-Language Models Toward Truth via Tri-layer Contrastive Decodingaccepted
- Watermarking with Low-Entropy POS-Guided Token Partitioning and Z-Score-Driven Dynamic Bias for Large Language Modelsaccepted
- We Argue to Agree: Towards Personality-Driven Argumentation-Based Negotiation Dialogue Systems for Tourismaccepted
- We Need to Measure Data Diversity in NLP — Better and Broaderaccepted
- We Politely Insist: Your LLM Must Learn the Persian Art of Taarofaccepted
- Weak2Wise: An Automated, Lightweight Framework for Weak-LLM-Friendly Reasoning Synthesisaccepted
- Weaver: Interweaving SQL and LLM for Table Reasoningaccepted
- Web Intellectual Property at Risk: Preventing Unauthorized Real-Time Retrieval by Large Language Modelsaccepted
- WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learningaccepted
- WebCoT: Enhancing Web Agent Reasoning by Reconstructing Chain-of-Thought in Reflection, Branching, and Rollbackaccepted
- WebEvolver: Enhancing Web Agent Self-Improvement with Co-evolving World Modelaccepted
- WebInject: Prompt Injection Attack to Web Agentsaccepted
- WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generationaccepted
- Weight-Aware Activation Sparsity with Constrained Bayesian Optimization Scheduling for Large Language Modelsaccepted
- Weights-Rotated Preference Optimization for Large Language Modelsaccepted
- What Do Indonesians Really Need from Language Technology? A Nationwide Surveyaccepted
- What Has Been Lost with Synthetic Evaluation?accepted
- What Makes a Good Reasoning Chain? Uncovering Structural Patterns in Long Chain-of-Thought Reasoningaccepted
- What Makes for Good Image Captions?accepted
- What Media Frames Reveal About Stance: A Dataset and Study about Memes in Climate Change Discourseaccepted
- What You Read Isn’t What You Hear: Linguistic Sensitivity in Deepfake Speech Detectionaccepted
- What You See is What You Ask: Evaluating Audio Descriptionsaccepted
- What are Foundation Models Cooking in the Post-Soviet World?accepted
- What data should I include in my POS tagging training set?accepted
- What if Othello-Playing Language Models Could See?accepted
- What’s Not Said Still Hurts: A Description-Based Evaluation Framework for Measuring Social Bias in LLMsaccepted
- What’s in a prompt? Language models encode literary style in prompt embeddingsaccepted
- When Allies Turn Foes: Exploring Group Characteristics of LLM-Based Multi-Agent Collaborative Systems Under Adversarial Attacksaccepted
- When Annotators Disagree, Topology Explains: Mapper, a Topological Tool for Exploring Text Embedding Geometry and Ambiguityaccepted
- When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Modelsaccepted
- When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMsaccepted
- When Format Changes Meaning: Investigating Semantic Inconsistency of Large Language Modelsaccepted
- When Instructions Multiply: Measuring and Estimating LLM Capabilities of Multiple Instructions Followingaccepted
- When Inverse Data Outperforms: Exploring the Pitfalls of Mixed Data in Multi-Stage Fine-Tuningaccepted
- When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMsaccepted
- When Long Helps Short: How Context Length in Supervised Fine-tuning Affects Behavior of Large Language Modelsaccepted
- When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQAaccepted
- When Models Reason in Your Language: Controlling Thinking Language Comes at the Cost of Accuracyaccepted
- When Personalization Meets Reality: A Multi-Faceted Analysis of Personalized Preference Learningaccepted
- When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMsaccepted
- When Truthful Representations Flip Under Deceptive Instructions?accepted
- When Words Smile: Generating Diverse Emotional Facial Expressions from Textaccepted
- When to Continue Thinking: Adaptive Thinking Mode Switching for Efficient Reasoningaccepted
- Where Confabulation Lives: Latent Feature Discovery in LLMsaccepted
- Where Did That Come From? Sentence-Level Error-Tolerant Attributionaccepted
- Where Fact Ends and Fairness Begins: Redefining AI Bias Evaluation through Cognitive Biasesaccepted
- Where to show Demos in Your Prompt: A Positional Bias of In-Context Learningaccepted
- Which Word Orders Facilitate Length Generalization in LMs? An Investigation with GCG-Based Artificial Languagesaccepted
- Whisper-UT: A Unified Translation Framework for Speech and Textaccepted
- Who Holds the Pen? Caricature and Perspective in LLM Retellings of Historyaccepted
- Who Speaks Matters: Analysing the Influence of the Speaker’s Linguistic Identity on Hate Classificationaccepted
- Who is in the Spotlight: The Hidden Bias Undermining Multimodal Retrieval-Augmented Generationaccepted
- Who’s the Author? How Explanations Impact User Reliance in AI-Assisted Authorship Attributionaccepted
- Why Do Some Inputs Break Low-Bit LLM Quantization?accepted
- Why Stop at One Error? Benchmarking LLMs as Data Science Code Debuggers for Multi-Hop and Multi-Bug Errorsaccepted
- Why We Feel What We Feel: Joint Detection of Emotions and Their Opinion Triggers in E-commerceaccepted
- Why and How LLMs Benefit from Knowledge Introspection in Commonsense Reasoningaccepted
- WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild?accepted
- WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music Reasoningaccepted
- Will Annotators Disagree? Identifying Subjectivity in Value-Laden Argumentsaccepted
- Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QAaccepted
- WojoodRelations: Arabic Relation Extraction Corpus and Modelingaccepted
- Women, Infamous, and Exotic Beings: A Comparative Study of Honorific Usages in Wikipedia and LLMs for Bengali and Hindiaccepted
- Word Salad Chopper: Reasoning Models Waste A Ton Of Decoding Budget On Useless Repetitions, Self-Knowinglyaccepted
- Words Like Knives: Backstory-Personalized Modeling and Detection of Violent Communicationaccepted
- X-Boundary: Establishing Exact Safety Boundary to Shield LLMs from Jailbreak Attacks without Compromising Usabilityaccepted
- X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought Reasoningaccepted
- X-FLoRA: Cross-modal Federated Learning with Modality-expert LoRA for Medical VQAaccepted
- X-LeBench: A Benchmark for Extremely Long Egocentric Video Understandingaccepted
- XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoMLaccepted
- XL-Suite: Cross-Lingual Synthetic Training and Evaluation Data for Open-Ended Generationaccepted
- XLQA: A Benchmark for Locale-Aware Multilingual Open-Domain Question Answeringaccepted
- XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compressionaccepted
- XRAG: Cross-lingual Retrieval-Augmented Generationaccepted
- XTRA: Cross-Lingual Topic Modeling with Topic and Representation Alignmentsaccepted
- You Are What You Train: Effects of Data Composition on Training Context-aware Machine Translation Modelsaccepted
- You Only Use Reactive Attention Slice When Retrieving From Long Contextaccepted
- Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectorsaccepted
- Your Mileage May Vary: How Empathy and Demographics Shape Human Preferences in LLM Responsesaccepted
- Your RAG is Unfair: Exposing Fairness Vulnerabilities in Retrieval-Augmented Generation via Backdoor Attacksaccepted
- Your Semantic-Independent Watermark is Fragile: A Semantic Perturbation Attack against EaaS Watermarkaccepted
- ZEBRA: Leveraging Model-Behavioral Knowledge for Zero-Annotation Preference Dataset Constructionaccepted
- ZERA: Zero-init Instruction Evolving Refinement Agent – From Zero Instructions to Structured Prompts via Principle-based Optimizationaccepted
- ZOGRASCOPE: A New Benchmark for Semantic Parsing over Property Graphsaccepted
- Zero-Shot Contextual Embeddings via Offline Synthetic Corpus Generationaccepted
- Zero-Shot Cross-Domain Aspect-Based Sentiment Analysis via Domain-Contextualized Chain-of-Thought Reasoningaccepted
- Zero-Shot Defense Against Toxic Images via Inherent Multimodal Alignment in LVLMsaccepted
- Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Modelsaccepted
- Zero-Shot Privacy-Aware Text Rewriting via Iterative Tree Searchaccepted
- Zero-shot Cross-lingual NER via Mitigating Language Difference: An Entity-aligned Translation Perspectiveaccepted
- Zero-shot Graph Reasoning via Retrieval Augmented Framework with LLMsaccepted
- Zero-shot Multimodal Document Retrieval via Cross-modal Question Generationaccepted
- Zipf’s and Heaps’ Laws for Tokens and LLM-generated Textsaccepted
- ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Explorationaccepted
- [MASK]ED - Language Modeling for Explainable Classification and Disentangling of Socially Unacceptable Discourse.accepted
- cAST: Enhancing Code Retrieval-Augmented Generation with Structural Chunking via Abstract Syntax Treeaccepted
- fLSA: Learning Semantic Structures in Document Collections Using Foundation Modelsaccepted
- iKnow-audio: Integrating Knowledge Graphs with Audio-Language Modelsaccepted
- iTool: Reinforced Fine-Tuning with Dynamic Deficiency Calibration for Advanced Tool Useaccepted
- iVISPAR — An Interactive Visual-Spatial Reasoning Benchmark for VLMsaccepted
- mrCAD: Multimodal Communication to Refine Computer-aided Designsaccepted
- pFedGPT: Hierarchically Optimizing LoRA Aggregation Weights for Personalized Federated GPT Modelsaccepted
- pFedRAG: A Personalized Federated Retrieval-Augmented Generation System with Depth-Adaptive Tiered Embedding Tuningaccepted
- polyBART: A Chemical Linguist for Polymer Property Prediction and Generative Designaccepted
- reWordBench: Benchmarking and Improving the Robustness of Reward Models with Transformed Inputsaccepted
- s1: Simple test-time scalingaccepted
- s3: You Don’t Need That Much Data to Train a Search Agent via RLaccepted
- seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMsaccepted
- so much depends / upon / a whitespace: Why Whitespace Matters for Poets and LLMsaccepted
- sudoLLM: On Multi-role Alignment of Language Modelsaccepted
- xCoRe: Cross-context Coreference Resolutionaccepted
- zFLoRA: Zero-Latency Fused Low-Rank Adaptersaccepted
- ‘Hello, World!’: Making GNNs Talk with LLMsaccepted
- ‘Rich Dad, Poor Lad’: How do Large Language Models Contextualize Socioeconomic Factors in College Admission ?accepted
- “Feels Feminine to Me”: Understanding Perceived Gendered Style through Human Annotationsaccepted
- “Going to a trap house” conveys more fear than “Going to a mall”: Benchmarking Emotion Context Sensitivity for LLMsaccepted
- “I’ve Decided to Leak”: Probing Internals Behind Prompt Leakage Intentsaccepted
- “Mm, Wat?” Detecting Other-initiated Repair Requests in Dialogueaccepted
- “What’s Up, Doc?”: Analyzing How Users Seek Health Information in Large-Scale Conversational AI Datasetsaccepted
- “Where Does This Strange Smell Come from?”: Enabling Conversational Interfaces for Artificial Olfactionaccepted
EMNLP accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.