← All conferences

NAACL 2025 Accepted Papers

The full list of 1,274 papers accepted at NAACL 2025 (North American Chapter of the Association for Computational Linguistics). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

Long: 614Findings: 457Short: 80Industry: 79System Demonstrations: 44
  1. Making Language Models Robust Against NegationLong
  2. Marrying LLMs with Dynamic Forecasting: A Graph Mixture-of-expert PerspectiveFindings
  3. Mastering the Craft of Data Synthesis for CodeLLMsLong
  4. Matina: A Large-Scale 73B Token Persian Text CorpusLong
  5. MeKB-Sim: Personal Knowledge Base-Powered Multi-Agent SimulationSystem Demonstrations
  6. Mechanistic Unveiling of Transformer Circuits: Self-Influence as a Key to Model ReasoningFindings
  7. MedCodER: A Generative AI Assistant for Medical CodingIndustry
  8. MedEthicEval: Evaluating Large Language Models Based on Chinese Medical EthicsIndustry
  9. MedEureka: A Medical Domain Benchmark for Multi-Granularity and Multi-Data-Type Embedding-Based RetrievalFindings
  10. MedThink: A Rationale-Guided Framework for Explaining Medical Visual Question AnsweringFindings
  11. Media of Langue: Exploring Word Translation NetworkFindings
  12. MergeME: Model Merging Techniques for Homogeneous and Heterogeneous MoEsLong
  13. Meta-Cultural Competence: Climbing the Right Hill of Cultural AwarenessLong
  14. Meta-Reasoning Improves Tool Use in Large Language ModelsFindings
  15. MetaScientist: A Human-AI Synergistic Framework for Automated Mechanical Metamaterial DesignSystem Demonstrations
  16. Mitigating Bias in Item Retrieval for Enhancing Exam Assembly in Vocational Education ServicesIndustry
  17. Mitigating Biases of Large Language Models in Stance Detection with Counterfactual Augmented CalibrationLong
  18. Mitigating Hallucinations in Multi-modal Large Language Models via Image Token Attention-Guided DecodingLong
  19. Mitigating Heterogeneity among Factor Tensors via Lie Group Manifolds for Tensor Decomposition Based Temporal Knowledge Graph EmbeddingLong
  20. Mitigating Tail Narrowing in LLM Self-Improvement via Socratic-Guided SamplingLong
  21. MixRevDetect: Towards Detecting AI-Generated Content in Hybrid Peer Reviews.Short
  22. Mixture of Multimodal Adapters for Sentiment AnalysisLong
  23. MoCE: Adaptive Mixture of Contextualization Experts for Byte-based Neural Machine TranslationLong
  24. MoDS: Moderating a Mixture of Document Speakers to Summarize Debatable Queries in Document CollectionsLong
  25. MoEMoE: Question Guided Dense and Scalable Sparse Mixture-of-Expert for Multi-source Multi-modal AnsweringIndustry
  26. MoFE: Mixture of Frozen Experts ArchitectureIndustry
  27. MoLA: MoE LoRA with Layer-wise Expert AllocationFindings
  28. MobA: Multifaceted Memory-Enhanced Adaptive Planning for Efficient Mobile Task AutomationSystem Demonstrations
  29. Modeling the Differential Prevalence of Online Supportive Interactions in Private Instant Messages of AdolescentsFindings
  30. MonoTODia: Translating Monologue Requests to Task-Oriented DialoguesIndustry
  31. MorphNLI: A Stepwise Approach to Natural Language Inference Using Text MorphingFindings
  32. Multi-Agent Simulator Drives Language Models for Legal Intensive InteractionFindings
  33. Multi-Condition Guided Diffusion Network for Multimodal Emotion Recognition in ConversationFindings
  34. MultiCAT: Multimodal Communication Annotations for TeamsFindings
  35. Multilingual Blending: Large Language Model Safety Alignment Evaluation with Language MixtureFindings
  36. Multilingual Reasoning via Self-trainingLong
  37. Multimodal Generation with Consistency TransferringFindings
  38. Music for All: Representational Bias and Cross-Cultural Adaptability of Music Generation ModelsFindings
  39. Mutual Reinforcement of LLM Dialogue Synthesis and Summarization Capabilities for Few-Shot Dialogue SummarizationFindings
  40. Mutual-pairing Data Augmentation for Fewshot Continual Relation ExtractionLong
  41. My LLM might Mimic AAE - But When Should It?Long
  42. NAT: Enhancing Agent Tuning with Negative SamplesLong
  43. NLI under the Microscope: What Atomic Hypothesis Decomposition RevealsLong
  44. NOTA: Multimodal Music Notation Understanding for Visual Large Language ModelFindings
  45. Natural Language Processing for Human Resources: A SurveyIndustry
  46. NeMo-Inspector: A Visualization Tool for LLM Generation AnalysisSystem Demonstrations
  47. No Simple Answer to Data Complexity: An Examination of Instance-Level Complexity Metrics for Classification TasksLong
  48. Not All Adapters Matter: Selective Adapter Freezing for Memory-Efficient Fine-Tuning of Language ModelsLong
  49. Not all Hallucinations are Good to Throw Away When it Comes to Legal Abstractive SummarizationLong
  50. Omni-Chart-600K: A Comprehensive Dataset of Chart Types for Chart UnderstandingFindings
  51. On A Scale From 1 to 5: Quantifying Hallucination in Faithfulness EvaluationFindings
  52. On Localizing and Deleting Toxic Memories in Large Language ModelsFindings
  53. On Using Arabic Language Dialects in Recommendation SystemsFindings
  54. On the Analysis and Distillation of Emergent Outlier Properties in Pre-trained Language ModelsLong
  55. On the Feasibility of In-Context Probing for Data AttributionFindings
  56. On the Influence of Context Size and Model Choice in Retrieval-Augmented Generation SystemsFindings
  57. On the Role of Key Phrases in Argument MiningFindings
  58. On the Role of Speech Data in Reducing Toxicity Detection BiasLong
  59. On the Vulnerability of Text SanitizationLong
  60. One Unified Model for Diverse Tasks: Emotion Cause Analysis via Self-Promote Cognitive Structure ModelingLong
  61. Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMsIndustry
  62. OpenBioNER: Lightweight Open-Domain Biomedical Named Entity Recognition Through Entity Type DescriptionFindings
  63. Optimizing Hidden Markov Language Models: An Empirical Study of Reparameterization and Initialization TechniquesFindings
  64. Optimizing LLMs for Italian: Reducing Token Fertility and Enhancing Efficiency Through Vocabulary AdaptationFindings
  65. Option Symbol Matters: Investigating and Mitigating Multiple-Choice Option Symbol Bias of Large Language ModelsLong
  66. Overcoming both Domain Shift and Label Shift for Referring Video SegmentationFindings
  67. PA-RAG: RAG Alignment via Multi-Perspective Preference OptimizationLong
  68. PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio ClassificationLong
  69. PEMV: Improving Spatial Distribution for Emotion Recognition in Conversations Using Proximal Emotion Mean VectorsFindings
  70. PLEX: Adaptive Parameter-Efficient Fine-Tuning for Code LLMs using Lottery-TicketsIndustry
  71. PRACTIQ: A Practical Conversational Text-to-SQL dataset with Ambiguous and Unanswerable QueriesLong
  72. PRDetect: Perturbation-Robust LLM-generated Text Detection Based on Syntax TreeFindings
  73. PREMISE: Matching-based Prediction for Accurate Review RecommendationFindings
  74. PROM: Pivoted and Regulated Optimization for Multilingual Instruction LearningShort
  75. PROMPTEVALS: A Dataset of Assertions and Guardrails for Custom Production Large Language Model PipelinesLong
  76. PairScale: Analyzing Attitude Change with Pairwise ComparisonsFindings
  77. Pairwise Prompt-Based Tuning with Parameter Efficient Fast Adaptation for Generalized Zero-Shot Intent DetectionFindings
  78. Palette of Language Models: A Solver for Controlled Text GenerationLong
  79. ParaICL: Towards Parallel In-Context LearningLong
  80. Parameter-free and Accessible Prompt Learning to Enhance Adversarial Robustness for Pre-trained Vision-Language ModelsLong
  81. Pay More Attention to Images: Numerous Images-Oriented Multimodal SummarizationLong
  82. PeerQA: A Scientific Question Answering Dataset from Peer ReviewsLong
  83. PerCul: A Story-Driven Cultural Evaluation of LLMs in PersianLong
  84. Persona-SQ: A Personalized Suggested Question Generation Framework For Real-world DocumentsSystem Demonstrations
  85. Personalize Your LLM: Fake it then Align itFindings
  86. Personalized Help for Optimizing Low-Skilled Users’ StrategyShort
  87. PicPersona-TOD : A Dataset for Personalizing Utterance Style in Task-Oriented Dialogue with Image PersonaLong
  88. Pipeline Analysis for Developing Instruct LLMs in Low-Resource Languages: A Case Study on BasqueLong
  89. Pisets: A Robust Speech Recognition System for Lectures and InterviewsIndustry
  90. Playing with Voices: Tabletop Role-Playing Game Recordings as a Diarization ChallengeFindings
  91. Pointwise Mutual Information as a Performance Gauge for Retrieval-Augmented GenerationLong
  92. PolyJoin: Semantic Multi-key Joinable Table Search in Data LakesFindings
  93. Position Really Matters: Towards a Holistic Approach for Prompt TuningFindings
  94. Predicting ICU Length of Stay for Patients using Latent Categorization of Health ConditionsIndustry
  95. Preference Consistency Matters: Enhancing Preference Learning in Language Models with Automated Self-Curation of Training CorporaLong
  96. Prepending or Cross-Attention for Speech-to-Text? An Empirical ComparisonLong
  97. Preserving Zero-shot Capability in Supervised Fine-tuning for Multi-label Text ClassificationFindings
  98. Private Synthetic Text Generation with Diffusion ModelsLong
  99. ProMQA: Question Answering Dataset for Multimodal Procedural Activity UnderstandingLong
  100. ProSE: Diffusion Priors for Speech EnhancementLong
  101. Prompt-Guided Selective Masking Loss for Context-Aware Emotive Text-to-SpeechFindings
  102. PromptRefine: Enhancing Few-Shot Performance on Low-Resource Indic Languages with Example Selection from related Example BanksLong
  103. Prompting with Phonemes: Enhancing LLMs’ Multilinguality for Non-Latin Script LanguagesLong
  104. Prompto: An open source library for asynchronous querying of LLM endpointsSystem Demonstrations
  105. Protein2Text: Resampling Mechanism to Translate Protein Sequences into Human-Interpretable TextIndustry
  106. Prototype Conditioned Generative Replay for Continual Learning in NLPLong
  107. Prototype Tuning: A Meta-Learning Approach for Few-Shot Document-Level Relation Extraction with Large Language ModelsFindings
  108. Prototypical Extreme Multi-label Classification with a Dynamic Margin LossLong
  109. Pula: Training Large Language Models for SetswanaLong
  110. PuzzleGPT: Emulating Human Puzzle-Solving Ability for Time and Location PredictionFindings
  111. Q-FAKER: Query-free Hard Black-box Attack via Controlled GenerationFindings
  112. QAVA: Query-Agnostic Visual Attack to Large Vision-Language ModelsLong
  113. QSpell 250K: A Large-Scale, Practical Dataset for Chinese Search Query Spell CorrectionIndustry
  114. Query Variant Detection Using Retriever as EnvironmentIndustry
  115. Query-focused Referentiability Learning for Zero-shot RetrievalLong
  116. QueryShield: A Platform to Mitigate Enterprise Data Leakage in Queries to External LLMsIndustry
  117. RAMQA: A Unified Framework for Retrieval-Augmented Multi-Modal Question AnsweringFindings
  118. RAP: A Metric for Balancing Repetition and Performance in Open-Source Large Language ModelsLong
  119. REFFLY: Melody-Constrained Lyrics Editing ModelLong
  120. RTSM: Knowledge Distillation with Diverse Signals for Efficient Real-Time Semantic Matching in E-CommerceIndustry
  121. Racing Thoughts: Explaining Contextualization Errors in Large Language ModelsLong
  122. RankAdaptor: Hierarchical Rank Allocation for Efficient Fine-Tuning Pruned LLMs via Performance ModelFindings
  123. ReGLA: Refining Gated Linear AttentionLong
  124. Reading between the Lines: Can LLMs Identify Cross-Cultural Communication Gaps?Long
  125. Reasoning Aware Self-Consistency: Leveraging Reasoning Paths for Efficient LLM SamplingLong
  126. Representation-to-Creativity (R2C): Automated Holistic Scoring Model for Essay CreativityFindings
  127. ResoFilter: Fine-grained Synthetic Data Filtering for Large Language Models through Data-Parameter Resonance AnalysisFindings
  128. Rethinking Smoothness for Fast and Adaptable Entity Alignment DecodingFindings
  129. Rethinking Word Similarity: Semantic Similarity through Classification ConfusionLong
  130. Rethinking the Role of LLMs for Document-level Relation Extraction: a Refiner with Task Distribution and Probability FusionLong
  131. RetrieverGuard: Empowering Information Retrieval to Combat LLM-Generated MisinformationFindings
  132. Reverse Modeling in Large Language ModelsShort
  133. Reversed Attention: On The Gradient Descent Of Attention Layers In GPTLong
  134. RevieWeaver: Weaving Together Review Insights by Leveraging LLMs and Semantic SimilarityIndustry
  135. Revisiting Early Detection of Sexual Predators via Turn-level OptimizationLong
  136. Reward-Guided Tree Search for Inference Time Alignment of Large Language ModelsLong
  137. Richer Output for Richer Countries: Uncovering Geographical Disparities in Generated Stories and Travel RecommendationsFindings
  138. Robust Bias Detection in MLMs and its Application to Human Trait RatingsFindings
  139. Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-SpeechLong
  140. RxLens: Multi-Agent LLM-powered Scan and Order for PharmacyIndustry
  141. SALAD: Improving Robustness and Generalization through Contrastive Learning with Structure-Aware and LLM-Driven Augmented DataLong
  142. SANDWiCH: Semantical Analysis of Neighbours for Disambiguating Words in Context ad HocLong
  143. SAPIENT: Mastering Multi-turn Conversational Recommendation with Strategic Planning and Monte Carlo Tree SearchLong
  144. SCIURus: Shared Circuits for Interpretable Uncertainty Representations in Language ModelsLong
  145. SCORE: Systematic COnsistency and Robustness Evaluation for Large Language ModelsIndustry
  146. SEEval: Advancing LLM Text Evaluation Efficiency and Accuracy through Self-Explanation PromptingFindings
  147. SEP-MLDC: A Simple and Effective Paradigm for Multi-Label Document ClassificationFindings
  148. SFMSS: Service Flow aware Medical Scenario Simulation for Conversational Data GenerationFindings
  149. SMAB: MAB based word Sensitivity Estimation Framework and its Applications in Adversarial Text GenerationLong
  150. SOTOPIA-S4: a user-friendly system for flexible, customizable, and large-scale social simulationSystem Demonstrations
  151. SSH: Sparse Spectrum Adaptation via Discrete Hartley TransformationLong
  152. SSMLoRA: Enhancing Low-Rank Adaptation with State Space ModelLong
  153. STEP: Staged Parameter-Efficient Pre-training for Large Language ModelsShort
  154. SUNAR: Semantic Uncertainty based Neighborhood Aware Retrieval for Complex QALong
  155. SURF: A System to Unveil Explainable Risk Relations between FirmsSystem Demonstrations
  156. SVD-LLM V2: Optimizing Singular Value Truncation for Large Language Model CompressionLong
  157. SWITCH: Studying with Teacher for Knowledge Distillation of Large Language ModelsFindings
  158. SafeQuant: LLM Safety Analysis via Quantized Gradient InspectionLong
  159. SafeSpeech: A Comprehensive and Interactive Tool for Analysing Sexist and Abusive Language in ConversationsSystem Demonstrations
  160. SafetyQuizzer: Timely and Dynamic Evaluation on the Safety of LLMsLong
  161. Scaling Graph-Based Dependency Parsing with Arc Vectorization and Attention-Based RefinementShort
  162. Scaling LLM Inference Efficiently with Optimized Sample Compute AllocationLong
  163. Scaling Multi-Document Event Summarization: Evaluating Compression vs. Full-Text ApproachesShort
  164. Schema and Natural Language Aware In-Context Learning for Improved GraphQL Query GenerationIndustry
  165. ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming ChallengesShort
  166. Script-Agnosticism and its Impact on Language Identification for Dravidian LanguagesLong
  167. Search Query Embeddings via User-behavior-driven Contrastive LearningIndustry
  168. See-Saw Modality Balance: See Gradient, and Sew Impaired Vision-Language Balance to Mitigate Dominant Modality BiasLong
  169. Seeds of Discourse: A Multilingual Corpus of Direct Quotations from African Media on Agricultural BiotechnologiesFindings
  170. Selective Self-to-Supervised Fine-Tuning for Generalization in Large Language ModelsFindings
  171. Self-Training Large Language Models for Tool-Use Without DemonstrationsFindings
  172. Self-calibration for Language Model Quantization and PruningLong
  173. Semi-automatic Sequential Sentence Classification in the Discourse Analysis Tool SuiteSystem Demonstrations
  174. SeqAR: Jailbreak LLMs with Sequential Auto-Generated CharactersLong
  175. Sequence-level Large Language Model Training with Contrastive Preference OptimizationFindings
  176. Sharpness-Aware Minimization for Topic Models with High-Quality Document RepresentationsLong
  177. SimSMoE: Toward Efficient Training Mixture of Experts via Solving Representational CollapseFindings
  178. Single Ground Truth Is Not Enough: Adding Flexibility to Aspect-Based Sentiment Analysis EvaluationLong
  179. Smurfs: Multi-Agent System using Context-Efficient DFSDT for Tool PlanningLong
  180. Soft Language Prompts for Language TransferLong
  181. Sports and Women’s Sports: Gender Bias in Text Generation with Olympic DataShort
  182. Step-by-Step Fact Verification System for Medical Claims with Explainable ReasoningShort
  183. Storybranch - generating multimedia content from novelsSystem Demonstrations
  184. Stronger Universal and Transferable Attacks by Suppressing RefusalsLong
  185. SuperRAG: Beyond RAG with Layout-Aware Graph ModelingIndustry
  186. Superlatives in Context: Modeling the Implicit Semantics of SuperlativesLong
  187. SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise UseIndustry
  188. Synonym-unaware Fast Adversarial Training against Textual Adversarial AttacksFindings
  189. SynthDetoxM: Modern LLMs are Few-Shot Parallel Detoxification Data AnnotatorsLong
  190. TRACE: Real-Time Multimodal Common Ground Tracking in Situated Collaborative DialoguesSystem Demonstrations
  191. TRANSIENTTABLES: Evaluating LLMs’ Reasoning on Temporally Evolving Semi-structured TablesLong
  192. TabComp: A Dataset for Visual Table Reading ComprehensionFindings
  193. Tackling Social Bias against the Poor: a Dataset and a Taxonomy on AporophobiaFindings
  194. TaeBench: Improving Quality of Toxic Adversarial ExamplesIndustry
  195. Target-Augmented Shared Fusion-based Multimodal Sarcasm Explanation GenerationFindings
  196. Task-driven Layerwise Additive Activation InterventionShort
  197. Task-wrapped Continual Learning in Task-Oriented Dialogue SystemsFindings
  198. Taxi1500: A Dataset for Multilingual Text Classification in 1500 LanguagesShort
  199. Taxonomy and Analysis of Sensitive User Queries in Generative AI Search SystemFindings
  200. TeCoFeS: Text Column Featurization using Semantic AnalysisFindings
  201. Tell Me What You Know About Sexism: Expert-LLM Interaction Strategies and Co-Created Definitions for Zero-Shot Sexism DetectionFindings
  202. Temporal-Aware Soft Prompt Tuning for Automatic Text DatingLong
  203. Test-Time Code-Switching for Cross-lingual Aspect Sentiment Triplet ExtractionLong
  204. Tethering Broken Themes: Aligning Neural Topic Models with Labels and AuthorsFindings
  205. Text Annotation via Inductive Coding: Comparing Human Experts to LLMs in Qualitative Data AnalysisFindings
  206. Text2Sql: Pure Fine-Tuning and Pure Knowledge DistillationIndustry
  207. The American Sign Language Knowledge Graph: Infusing ASL Models with Linguistic KnowledgeFindings
  208. The Impact of Domain-Specific Terminology on Machine Translation for Finance in European LanguagesLong
  209. The Impact of Inference Acceleration on Bias of LLMsLong
  210. The Power of Bullet Lists: A Simple Yet Effective Prompting Approach to Enhancing Spatial Reasoning in Large Language ModelsFindings
  211. The Role of Prosody in Spoken Question AnsweringFindings
  212. The State and Fate of Summarization Datasets: A SurveyLong
  213. Through the Lens of History: Methods for Analyzing Temporal Variation in Content and Framing of State-run Chinese NewspapersLong
  214. Time-aware ReAct Agent for Temporal Knowledge Graph Question AnsweringFindings
  215. TinyThinker: Distilling Reasoning through Coarse-to-Fine Knowledge Internalization with Self-ReflectionLong
  216. ToVo: Toxicity Taxonomy via VotingFindings
  217. Token-Level Density-Based Uncertainty Quantification Methods for Eliciting Truthfulness of Large Language ModelsLong
  218. Tomato, Tomahto, Tomate: Do Multilingual Language Models Understand Based on Subword-Level Semantic Concepts?Findings
  219. Tonguescape: Exploring Language Models Understanding of Vowel ArticulationLong
  220. Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy LossLong
  221. Towards Lifelong Dialogue Agents via Timeline-based Memory ManagementLong
  222. Towards Operationalizing Right to Data ProtectionLong
  223. Towards Prompt Generalization: Grammar-aware Cross-Prompt Automated Essay ScoringFindings
  224. Towards Reliable Agents: Benchmarking Customized LLM-Based Retrieval-Augmented Generation Frameworks with Deployment ValidationIndustry
  225. Towards Reliable and Practical Phishing DetectionIndustry
  226. Towards Sustainable NLP: Insights from Benchmarking Inference Energy in Large Language ModelsLong
  227. Towards Unified, Dynamic and Annotation-based Visualisations and Exploration of Annotated Big Data Corpora with the Help of Unified Corpus ExplorerSystem Demonstrations
  228. Towards a Perspectivist Turn in Argument Quality AssessmentLong
  229. Track-SQL: Enhancing Generative Language Models with Dual-Extractive Modules for Schema and Context Tracking in Multi-turn Text-to-SQLLong
  230. Transferable Post-training via Inverse Value LearningLong
  231. Transform Retrieval for Textual Entailment in RAGShort
  232. TransformerRanker: A Tool for Efficiently Finding the Best-Suited Language Models for Downstream Classification TasksSystem Demonstrations
  233. Tricking Retrievers with Influential Tokens: An Efficient Black-Box Corpus Poisoning AttackLong
  234. TurboFuzzLLM: Turbocharging Mutation-based Fuzzing for Effectively Jailbreaking Large Language Models in PracticeIndustry
  235. UCL-Bench: A Chinese User-Centric Legal Benchmark for Large Language ModelsFindings
  236. UOREX: Towards Uncertainty-Aware Open Relation ExtractionLong
  237. Understanding the Role of Mental Models in User Interaction with an Adaptive Dialog AgentFindings
  238. UniRAG: Universal Retrieval Augmentation for Large Vision Language ModelsFindings
  239. Unleashing Multi-Hop Reasoning Potential in Large Language Models through Repetition of Misordered ContextFindings
  240. Unlocking Korean Verbs: A User-Friendly Exploration into the Verb LexiconSystem Demonstrations
  241. Unlocking the Planning Capabilities of Large Language Models with Maximum Diversity Fine-tuningFindings
  242. Unmasking Implicit Bias: Evaluating Persona-Prompted LLM Responses in Power-Disparate Social ScenariosLong
  243. Unsupervised Sentence Representation Learning with Syntactically Aligned Negative SamplesFindings
  244. Upsample or Upweight? Balanced Training on Heavily Imbalanced DatasetsLong
  245. Using Contextually Aligned Online Reviews to Measure LLMs’ Performance Disparities Across Language VarietiesShort
  246. Using Linguistic Entrainment to Evaluate Large Language Models for Use in Cognitive Behavioral TherapyFindings
  247. Using Review Combination and Pseudo-Tokens for Aspect Sentiment Quad PredictionFindings
  248. Using Text-Based Causal Inference to Disentangle Factors Influencing Online Review RatingsLong
  249. VIT-Pro: Visual Instruction Tuning for Product ImagesIndustry
  250. VTechAGP: An Academic-to-General-Audience Text Paraphrase Dataset and Benchmark ModelsLong
  251. Verifiable Format Control for Large Language Model GenerationsFindings
  252. Verify-in-the-Graph: Entity Disambiguation Enhancement for Complex Claim Verification with Interactive Graph RepresentationLong
  253. Vulnerability of Large Language Models to Output Prefix Jailbreaks: Impact of Positions on SafetyFindings
  254. Waste Not, Want Not; Recycled Gumbel Noise Improves Consistency in Natural Language GenerationLong
  255. WaterPool: A Language Model Watermark Mitigating Trade-Offs among Imperceptibility, Efficacy and RobustnessLong
  256. WaveFM: A High-Fidelity and Efficient Vocoder Based on Flow MatchingLong
  257. WebQuality: A Large-scale Multi-modal Web Page Quality Assessment Dataset with Multiple Scoring DimensionsLong
  258. What the #?*!: Disentangling Hate Across Target IdentitiesLong
  259. When and How to Augment Your Input: Question Routing Helps Balance the Accuracy and Efficiency of Large Language ModelsFindings
  260. When natural language is not enough: The limits of in-context learning demonstrations in multilingual reasoningFindings
  261. When2Call: When (not) to Call ToolsLong
  262. Where is the answer? An empirical study of positional bias for parametric knowledge extraction in language modelLong
  263. Where is this coming from? Making groundedness count in the evaluation of Document VQA modelsFindings
  264. WorkTeam: Constructing Workflows from Natural Language with Multi-AgentsIndustry
  265. You Only Read Once (YORO): Learning to Internalize Database Knowledge for Text-to-SQLLong
  266. Zero-Shot ATC Coding with Large Language Models for Clinical AssessmentsIndustry
  267. Zero-Shot Keyphrase Generation: Investigating Specialized Instructions and Multi-sample Aggregation on Large Language ModelsFindings
  268. eC-Tab2Text: Aspect-Based Text Generation from e-Commerce Product TablesIndustry
  269. kNN For Whisper And Its Effect On Bias And Speaker AdaptationFindings
  270. kNN Retrieval for Simple and Effective Zero-Shot Multi-speaker Text-to-SpeechShort
  271. tRAG: Term-level Retrieval-Augmented Generation for Domain-Adaptive RetrievalLong
  272. uDistil-Whisper: Label-Free Data Filtering for Knowledge Distillation in Low-Data RegimesLong
  273. “All that Glitters”: Techniques for Evaluations with Unreliable Model and Human AnnotationsFindings
  274. “Women do not have heart attacks!” Gender Biases in Automatically Generated Clinical Cases in FrenchFindings

NAACL accepted papers in other years

Looking for submission deadlines instead? See the conference deadline calendar.

NAACL 2025 Accepted Papers · Full List of 1,274 Papers