NAACL 2025 Accepted Papers
The full list of 1,274 papers accepted at NAACL 2025 (North American Chapter of the Association for Computational Linguistics). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
Long: 614Findings: 457Short: 80Industry: 79System Demonstrations: 44
- Making Language Models Robust Against NegationLong
- Marrying LLMs with Dynamic Forecasting: A Graph Mixture-of-expert PerspectiveFindings
- Mastering the Craft of Data Synthesis for CodeLLMsLong
- Matina: A Large-Scale 73B Token Persian Text CorpusLong
- MeKB-Sim: Personal Knowledge Base-Powered Multi-Agent SimulationSystem Demonstrations
- Mechanistic Unveiling of Transformer Circuits: Self-Influence as a Key to Model ReasoningFindings
- MedCodER: A Generative AI Assistant for Medical CodingIndustry
- MedEthicEval: Evaluating Large Language Models Based on Chinese Medical EthicsIndustry
- MedEureka: A Medical Domain Benchmark for Multi-Granularity and Multi-Data-Type Embedding-Based RetrievalFindings
- MedThink: A Rationale-Guided Framework for Explaining Medical Visual Question AnsweringFindings
- Media of Langue: Exploring Word Translation NetworkFindings
- MergeME: Model Merging Techniques for Homogeneous and Heterogeneous MoEsLong
- Meta-Cultural Competence: Climbing the Right Hill of Cultural AwarenessLong
- Meta-Reasoning Improves Tool Use in Large Language ModelsFindings
- MetaScientist: A Human-AI Synergistic Framework for Automated Mechanical Metamaterial DesignSystem Demonstrations
- Mitigating Bias in Item Retrieval for Enhancing Exam Assembly in Vocational Education ServicesIndustry
- Mitigating Biases of Large Language Models in Stance Detection with Counterfactual Augmented CalibrationLong
- Mitigating Hallucinations in Multi-modal Large Language Models via Image Token Attention-Guided DecodingLong
- Mitigating Heterogeneity among Factor Tensors via Lie Group Manifolds for Tensor Decomposition Based Temporal Knowledge Graph EmbeddingLong
- Mitigating Tail Narrowing in LLM Self-Improvement via Socratic-Guided SamplingLong
- MixRevDetect: Towards Detecting AI-Generated Content in Hybrid Peer Reviews.Short
- Mixture of Multimodal Adapters for Sentiment AnalysisLong
- MoCE: Adaptive Mixture of Contextualization Experts for Byte-based Neural Machine TranslationLong
- MoDS: Moderating a Mixture of Document Speakers to Summarize Debatable Queries in Document CollectionsLong
- MoEMoE: Question Guided Dense and Scalable Sparse Mixture-of-Expert for Multi-source Multi-modal AnsweringIndustry
- MoFE: Mixture of Frozen Experts ArchitectureIndustry
- MoLA: MoE LoRA with Layer-wise Expert AllocationFindings
- MobA: Multifaceted Memory-Enhanced Adaptive Planning for Efficient Mobile Task AutomationSystem Demonstrations
- Modeling the Differential Prevalence of Online Supportive Interactions in Private Instant Messages of AdolescentsFindings
- MonoTODia: Translating Monologue Requests to Task-Oriented DialoguesIndustry
- MorphNLI: A Stepwise Approach to Natural Language Inference Using Text MorphingFindings
- Multi-Agent Simulator Drives Language Models for Legal Intensive InteractionFindings
- Multi-Condition Guided Diffusion Network for Multimodal Emotion Recognition in ConversationFindings
- MultiCAT: Multimodal Communication Annotations for TeamsFindings
- Multilingual Blending: Large Language Model Safety Alignment Evaluation with Language MixtureFindings
- Multilingual Reasoning via Self-trainingLong
- Multimodal Generation with Consistency TransferringFindings
- Music for All: Representational Bias and Cross-Cultural Adaptability of Music Generation ModelsFindings
- Mutual Reinforcement of LLM Dialogue Synthesis and Summarization Capabilities for Few-Shot Dialogue SummarizationFindings
- Mutual-pairing Data Augmentation for Fewshot Continual Relation ExtractionLong
- My LLM might Mimic AAE - But When Should It?Long
- NAT: Enhancing Agent Tuning with Negative SamplesLong
- NLI under the Microscope: What Atomic Hypothesis Decomposition RevealsLong
- NOTA: Multimodal Music Notation Understanding for Visual Large Language ModelFindings
- Natural Language Processing for Human Resources: A SurveyIndustry
- NeMo-Inspector: A Visualization Tool for LLM Generation AnalysisSystem Demonstrations
- No Simple Answer to Data Complexity: An Examination of Instance-Level Complexity Metrics for Classification TasksLong
- Not All Adapters Matter: Selective Adapter Freezing for Memory-Efficient Fine-Tuning of Language ModelsLong
- Not all Hallucinations are Good to Throw Away When it Comes to Legal Abstractive SummarizationLong
- Omni-Chart-600K: A Comprehensive Dataset of Chart Types for Chart UnderstandingFindings
- On A Scale From 1 to 5: Quantifying Hallucination in Faithfulness EvaluationFindings
- On Localizing and Deleting Toxic Memories in Large Language ModelsFindings
- On Using Arabic Language Dialects in Recommendation SystemsFindings
- On the Analysis and Distillation of Emergent Outlier Properties in Pre-trained Language ModelsLong
- On the Feasibility of In-Context Probing for Data AttributionFindings
- On the Influence of Context Size and Model Choice in Retrieval-Augmented Generation SystemsFindings
- On the Role of Key Phrases in Argument MiningFindings
- On the Role of Speech Data in Reducing Toxicity Detection BiasLong
- On the Vulnerability of Text SanitizationLong
- One Unified Model for Diverse Tasks: Emotion Cause Analysis via Self-Promote Cognitive Structure ModelingLong
- Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMsIndustry
- OpenBioNER: Lightweight Open-Domain Biomedical Named Entity Recognition Through Entity Type DescriptionFindings
- Optimizing Hidden Markov Language Models: An Empirical Study of Reparameterization and Initialization TechniquesFindings
- Optimizing LLMs for Italian: Reducing Token Fertility and Enhancing Efficiency Through Vocabulary AdaptationFindings
- Option Symbol Matters: Investigating and Mitigating Multiple-Choice Option Symbol Bias of Large Language ModelsLong
- Overcoming both Domain Shift and Label Shift for Referring Video SegmentationFindings
- PA-RAG: RAG Alignment via Multi-Perspective Preference OptimizationLong
- PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio ClassificationLong
- PEMV: Improving Spatial Distribution for Emotion Recognition in Conversations Using Proximal Emotion Mean VectorsFindings
- PLEX: Adaptive Parameter-Efficient Fine-Tuning for Code LLMs using Lottery-TicketsIndustry
- PRACTIQ: A Practical Conversational Text-to-SQL dataset with Ambiguous and Unanswerable QueriesLong
- PRDetect: Perturbation-Robust LLM-generated Text Detection Based on Syntax TreeFindings
- PREMISE: Matching-based Prediction for Accurate Review RecommendationFindings
- PROM: Pivoted and Regulated Optimization for Multilingual Instruction LearningShort
- PROMPTEVALS: A Dataset of Assertions and Guardrails for Custom Production Large Language Model PipelinesLong
- PairScale: Analyzing Attitude Change with Pairwise ComparisonsFindings
- Pairwise Prompt-Based Tuning with Parameter Efficient Fast Adaptation for Generalized Zero-Shot Intent DetectionFindings
- Palette of Language Models: A Solver for Controlled Text GenerationLong
- ParaICL: Towards Parallel In-Context LearningLong
- Parameter-free and Accessible Prompt Learning to Enhance Adversarial Robustness for Pre-trained Vision-Language ModelsLong
- Pay More Attention to Images: Numerous Images-Oriented Multimodal SummarizationLong
- PeerQA: A Scientific Question Answering Dataset from Peer ReviewsLong
- PerCul: A Story-Driven Cultural Evaluation of LLMs in PersianLong
- Persona-SQ: A Personalized Suggested Question Generation Framework For Real-world DocumentsSystem Demonstrations
- Personalize Your LLM: Fake it then Align itFindings
- Personalized Help for Optimizing Low-Skilled Users’ StrategyShort
- PicPersona-TOD : A Dataset for Personalizing Utterance Style in Task-Oriented Dialogue with Image PersonaLong
- Pipeline Analysis for Developing Instruct LLMs in Low-Resource Languages: A Case Study on BasqueLong
- Pisets: A Robust Speech Recognition System for Lectures and InterviewsIndustry
- Playing with Voices: Tabletop Role-Playing Game Recordings as a Diarization ChallengeFindings
- Pointwise Mutual Information as a Performance Gauge for Retrieval-Augmented GenerationLong
- PolyJoin: Semantic Multi-key Joinable Table Search in Data LakesFindings
- Position Really Matters: Towards a Holistic Approach for Prompt TuningFindings
- Predicting ICU Length of Stay for Patients using Latent Categorization of Health ConditionsIndustry
- Preference Consistency Matters: Enhancing Preference Learning in Language Models with Automated Self-Curation of Training CorporaLong
- Prepending or Cross-Attention for Speech-to-Text? An Empirical ComparisonLong
- Preserving Zero-shot Capability in Supervised Fine-tuning for Multi-label Text ClassificationFindings
- Private Synthetic Text Generation with Diffusion ModelsLong
- ProMQA: Question Answering Dataset for Multimodal Procedural Activity UnderstandingLong
- ProSE: Diffusion Priors for Speech EnhancementLong
- Prompt-Guided Selective Masking Loss for Context-Aware Emotive Text-to-SpeechFindings
- PromptRefine: Enhancing Few-Shot Performance on Low-Resource Indic Languages with Example Selection from related Example BanksLong
- Prompting with Phonemes: Enhancing LLMs’ Multilinguality for Non-Latin Script LanguagesLong
- Prompto: An open source library for asynchronous querying of LLM endpointsSystem Demonstrations
- Protein2Text: Resampling Mechanism to Translate Protein Sequences into Human-Interpretable TextIndustry
- Prototype Conditioned Generative Replay for Continual Learning in NLPLong
- Prototype Tuning: A Meta-Learning Approach for Few-Shot Document-Level Relation Extraction with Large Language ModelsFindings
- Prototypical Extreme Multi-label Classification with a Dynamic Margin LossLong
- Pula: Training Large Language Models for SetswanaLong
- PuzzleGPT: Emulating Human Puzzle-Solving Ability for Time and Location PredictionFindings
- Q-FAKER: Query-free Hard Black-box Attack via Controlled GenerationFindings
- QAVA: Query-Agnostic Visual Attack to Large Vision-Language ModelsLong
- QSpell 250K: A Large-Scale, Practical Dataset for Chinese Search Query Spell CorrectionIndustry
- Query Variant Detection Using Retriever as EnvironmentIndustry
- Query-focused Referentiability Learning for Zero-shot RetrievalLong
- QueryShield: A Platform to Mitigate Enterprise Data Leakage in Queries to External LLMsIndustry
- RAMQA: A Unified Framework for Retrieval-Augmented Multi-Modal Question AnsweringFindings
- RAP: A Metric for Balancing Repetition and Performance in Open-Source Large Language ModelsLong
- REFFLY: Melody-Constrained Lyrics Editing ModelLong
- RTSM: Knowledge Distillation with Diverse Signals for Efficient Real-Time Semantic Matching in E-CommerceIndustry
- Racing Thoughts: Explaining Contextualization Errors in Large Language ModelsLong
- RankAdaptor: Hierarchical Rank Allocation for Efficient Fine-Tuning Pruned LLMs via Performance ModelFindings
- ReGLA: Refining Gated Linear AttentionLong
- Reading between the Lines: Can LLMs Identify Cross-Cultural Communication Gaps?Long
- Reasoning Aware Self-Consistency: Leveraging Reasoning Paths for Efficient LLM SamplingLong
- Representation-to-Creativity (R2C): Automated Holistic Scoring Model for Essay CreativityFindings
- ResoFilter: Fine-grained Synthetic Data Filtering for Large Language Models through Data-Parameter Resonance AnalysisFindings
- Rethinking Smoothness for Fast and Adaptable Entity Alignment DecodingFindings
- Rethinking Word Similarity: Semantic Similarity through Classification ConfusionLong
- Rethinking the Role of LLMs for Document-level Relation Extraction: a Refiner with Task Distribution and Probability FusionLong
- RetrieverGuard: Empowering Information Retrieval to Combat LLM-Generated MisinformationFindings
- Reverse Modeling in Large Language ModelsShort
- Reversed Attention: On The Gradient Descent Of Attention Layers In GPTLong
- RevieWeaver: Weaving Together Review Insights by Leveraging LLMs and Semantic SimilarityIndustry
- Revisiting Early Detection of Sexual Predators via Turn-level OptimizationLong
- Reward-Guided Tree Search for Inference Time Alignment of Large Language ModelsLong
- Richer Output for Richer Countries: Uncovering Geographical Disparities in Generated Stories and Travel RecommendationsFindings
- Robust Bias Detection in MLMs and its Application to Human Trait RatingsFindings
- Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-SpeechLong
- RxLens: Multi-Agent LLM-powered Scan and Order for PharmacyIndustry
- SALAD: Improving Robustness and Generalization through Contrastive Learning with Structure-Aware and LLM-Driven Augmented DataLong
- SANDWiCH: Semantical Analysis of Neighbours for Disambiguating Words in Context ad HocLong
- SAPIENT: Mastering Multi-turn Conversational Recommendation with Strategic Planning and Monte Carlo Tree SearchLong
- SCIURus: Shared Circuits for Interpretable Uncertainty Representations in Language ModelsLong
- SCORE: Systematic COnsistency and Robustness Evaluation for Large Language ModelsIndustry
- SEEval: Advancing LLM Text Evaluation Efficiency and Accuracy through Self-Explanation PromptingFindings
- SEP-MLDC: A Simple and Effective Paradigm for Multi-Label Document ClassificationFindings
- SFMSS: Service Flow aware Medical Scenario Simulation for Conversational Data GenerationFindings
- SMAB: MAB based word Sensitivity Estimation Framework and its Applications in Adversarial Text GenerationLong
- SOTOPIA-S4: a user-friendly system for flexible, customizable, and large-scale social simulationSystem Demonstrations
- SSH: Sparse Spectrum Adaptation via Discrete Hartley TransformationLong
- SSMLoRA: Enhancing Low-Rank Adaptation with State Space ModelLong
- STEP: Staged Parameter-Efficient Pre-training for Large Language ModelsShort
- SUNAR: Semantic Uncertainty based Neighborhood Aware Retrieval for Complex QALong
- SURF: A System to Unveil Explainable Risk Relations between FirmsSystem Demonstrations
- SVD-LLM V2: Optimizing Singular Value Truncation for Large Language Model CompressionLong
- SWITCH: Studying with Teacher for Knowledge Distillation of Large Language ModelsFindings
- SafeQuant: LLM Safety Analysis via Quantized Gradient InspectionLong
- SafeSpeech: A Comprehensive and Interactive Tool for Analysing Sexist and Abusive Language in ConversationsSystem Demonstrations
- SafetyQuizzer: Timely and Dynamic Evaluation on the Safety of LLMsLong
- Scaling Graph-Based Dependency Parsing with Arc Vectorization and Attention-Based RefinementShort
- Scaling LLM Inference Efficiently with Optimized Sample Compute AllocationLong
- Scaling Multi-Document Event Summarization: Evaluating Compression vs. Full-Text ApproachesShort
- Schema and Natural Language Aware In-Context Learning for Improved GraphQL Query GenerationIndustry
- ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming ChallengesShort
- Script-Agnosticism and its Impact on Language Identification for Dravidian LanguagesLong
- Search Query Embeddings via User-behavior-driven Contrastive LearningIndustry
- See-Saw Modality Balance: See Gradient, and Sew Impaired Vision-Language Balance to Mitigate Dominant Modality BiasLong
- Seeds of Discourse: A Multilingual Corpus of Direct Quotations from African Media on Agricultural BiotechnologiesFindings
- Selective Self-to-Supervised Fine-Tuning for Generalization in Large Language ModelsFindings
- Self-Training Large Language Models for Tool-Use Without DemonstrationsFindings
- Self-calibration for Language Model Quantization and PruningLong
- Semi-automatic Sequential Sentence Classification in the Discourse Analysis Tool SuiteSystem Demonstrations
- SeqAR: Jailbreak LLMs with Sequential Auto-Generated CharactersLong
- Sequence-level Large Language Model Training with Contrastive Preference OptimizationFindings
- Sharpness-Aware Minimization for Topic Models with High-Quality Document RepresentationsLong
- SimSMoE: Toward Efficient Training Mixture of Experts via Solving Representational CollapseFindings
- Single Ground Truth Is Not Enough: Adding Flexibility to Aspect-Based Sentiment Analysis EvaluationLong
- Smurfs: Multi-Agent System using Context-Efficient DFSDT for Tool PlanningLong
- Soft Language Prompts for Language TransferLong
- Sports and Women’s Sports: Gender Bias in Text Generation with Olympic DataShort
- Step-by-Step Fact Verification System for Medical Claims with Explainable ReasoningShort
- Storybranch - generating multimedia content from novelsSystem Demonstrations
- Stronger Universal and Transferable Attacks by Suppressing RefusalsLong
- SuperRAG: Beyond RAG with Layout-Aware Graph ModelingIndustry
- Superlatives in Context: Modeling the Implicit Semantics of SuperlativesLong
- SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise UseIndustry
- Synonym-unaware Fast Adversarial Training against Textual Adversarial AttacksFindings
- SynthDetoxM: Modern LLMs are Few-Shot Parallel Detoxification Data AnnotatorsLong
- TRACE: Real-Time Multimodal Common Ground Tracking in Situated Collaborative DialoguesSystem Demonstrations
- TRANSIENTTABLES: Evaluating LLMs’ Reasoning on Temporally Evolving Semi-structured TablesLong
- TabComp: A Dataset for Visual Table Reading ComprehensionFindings
- Tackling Social Bias against the Poor: a Dataset and a Taxonomy on AporophobiaFindings
- TaeBench: Improving Quality of Toxic Adversarial ExamplesIndustry
- Target-Augmented Shared Fusion-based Multimodal Sarcasm Explanation GenerationFindings
- Task-driven Layerwise Additive Activation InterventionShort
- Task-wrapped Continual Learning in Task-Oriented Dialogue SystemsFindings
- Taxi1500: A Dataset for Multilingual Text Classification in 1500 LanguagesShort
- Taxonomy and Analysis of Sensitive User Queries in Generative AI Search SystemFindings
- TeCoFeS: Text Column Featurization using Semantic AnalysisFindings
- Tell Me What You Know About Sexism: Expert-LLM Interaction Strategies and Co-Created Definitions for Zero-Shot Sexism DetectionFindings
- Temporal-Aware Soft Prompt Tuning for Automatic Text DatingLong
- Test-Time Code-Switching for Cross-lingual Aspect Sentiment Triplet ExtractionLong
- Tethering Broken Themes: Aligning Neural Topic Models with Labels and AuthorsFindings
- Text Annotation via Inductive Coding: Comparing Human Experts to LLMs in Qualitative Data AnalysisFindings
- Text2Sql: Pure Fine-Tuning and Pure Knowledge DistillationIndustry
- The American Sign Language Knowledge Graph: Infusing ASL Models with Linguistic KnowledgeFindings
- The Impact of Domain-Specific Terminology on Machine Translation for Finance in European LanguagesLong
- The Impact of Inference Acceleration on Bias of LLMsLong
- The Power of Bullet Lists: A Simple Yet Effective Prompting Approach to Enhancing Spatial Reasoning in Large Language ModelsFindings
- The Role of Prosody in Spoken Question AnsweringFindings
- The State and Fate of Summarization Datasets: A SurveyLong
- Through the Lens of History: Methods for Analyzing Temporal Variation in Content and Framing of State-run Chinese NewspapersLong
- Time-aware ReAct Agent for Temporal Knowledge Graph Question AnsweringFindings
- TinyThinker: Distilling Reasoning through Coarse-to-Fine Knowledge Internalization with Self-ReflectionLong
- ToVo: Toxicity Taxonomy via VotingFindings
- Token-Level Density-Based Uncertainty Quantification Methods for Eliciting Truthfulness of Large Language ModelsLong
- Tomato, Tomahto, Tomate: Do Multilingual Language Models Understand Based on Subword-Level Semantic Concepts?Findings
- Tonguescape: Exploring Language Models Understanding of Vowel ArticulationLong
- Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy LossLong
- Towards Lifelong Dialogue Agents via Timeline-based Memory ManagementLong
- Towards Operationalizing Right to Data ProtectionLong
- Towards Prompt Generalization: Grammar-aware Cross-Prompt Automated Essay ScoringFindings
- Towards Reliable Agents: Benchmarking Customized LLM-Based Retrieval-Augmented Generation Frameworks with Deployment ValidationIndustry
- Towards Reliable and Practical Phishing DetectionIndustry
- Towards Sustainable NLP: Insights from Benchmarking Inference Energy in Large Language ModelsLong
- Towards Unified, Dynamic and Annotation-based Visualisations and Exploration of Annotated Big Data Corpora with the Help of Unified Corpus ExplorerSystem Demonstrations
- Towards a Perspectivist Turn in Argument Quality AssessmentLong
- Track-SQL: Enhancing Generative Language Models with Dual-Extractive Modules for Schema and Context Tracking in Multi-turn Text-to-SQLLong
- Transferable Post-training via Inverse Value LearningLong
- Transform Retrieval for Textual Entailment in RAGShort
- TransformerRanker: A Tool for Efficiently Finding the Best-Suited Language Models for Downstream Classification TasksSystem Demonstrations
- Tricking Retrievers with Influential Tokens: An Efficient Black-Box Corpus Poisoning AttackLong
- TurboFuzzLLM: Turbocharging Mutation-based Fuzzing for Effectively Jailbreaking Large Language Models in PracticeIndustry
- UCL-Bench: A Chinese User-Centric Legal Benchmark for Large Language ModelsFindings
- UOREX: Towards Uncertainty-Aware Open Relation ExtractionLong
- Understanding the Role of Mental Models in User Interaction with an Adaptive Dialog AgentFindings
- UniRAG: Universal Retrieval Augmentation for Large Vision Language ModelsFindings
- Unleashing Multi-Hop Reasoning Potential in Large Language Models through Repetition of Misordered ContextFindings
- Unlocking Korean Verbs: A User-Friendly Exploration into the Verb LexiconSystem Demonstrations
- Unlocking the Planning Capabilities of Large Language Models with Maximum Diversity Fine-tuningFindings
- Unmasking Implicit Bias: Evaluating Persona-Prompted LLM Responses in Power-Disparate Social ScenariosLong
- Unsupervised Sentence Representation Learning with Syntactically Aligned Negative SamplesFindings
- Upsample or Upweight? Balanced Training on Heavily Imbalanced DatasetsLong
- Using Contextually Aligned Online Reviews to Measure LLMs’ Performance Disparities Across Language VarietiesShort
- Using Linguistic Entrainment to Evaluate Large Language Models for Use in Cognitive Behavioral TherapyFindings
- Using Review Combination and Pseudo-Tokens for Aspect Sentiment Quad PredictionFindings
- Using Text-Based Causal Inference to Disentangle Factors Influencing Online Review RatingsLong
- VIT-Pro: Visual Instruction Tuning for Product ImagesIndustry
- VTechAGP: An Academic-to-General-Audience Text Paraphrase Dataset and Benchmark ModelsLong
- Verifiable Format Control for Large Language Model GenerationsFindings
- Verify-in-the-Graph: Entity Disambiguation Enhancement for Complex Claim Verification with Interactive Graph RepresentationLong
- Vulnerability of Large Language Models to Output Prefix Jailbreaks: Impact of Positions on SafetyFindings
- Waste Not, Want Not; Recycled Gumbel Noise Improves Consistency in Natural Language GenerationLong
- WaterPool: A Language Model Watermark Mitigating Trade-Offs among Imperceptibility, Efficacy and RobustnessLong
- WaveFM: A High-Fidelity and Efficient Vocoder Based on Flow MatchingLong
- WebQuality: A Large-scale Multi-modal Web Page Quality Assessment Dataset with Multiple Scoring DimensionsLong
- What the #?*!: Disentangling Hate Across Target IdentitiesLong
- When and How to Augment Your Input: Question Routing Helps Balance the Accuracy and Efficiency of Large Language ModelsFindings
- When natural language is not enough: The limits of in-context learning demonstrations in multilingual reasoningFindings
- When2Call: When (not) to Call ToolsLong
- Where is the answer? An empirical study of positional bias for parametric knowledge extraction in language modelLong
- Where is this coming from? Making groundedness count in the evaluation of Document VQA modelsFindings
- WorkTeam: Constructing Workflows from Natural Language with Multi-AgentsIndustry
- You Only Read Once (YORO): Learning to Internalize Database Knowledge for Text-to-SQLLong
- Zero-Shot ATC Coding with Large Language Models for Clinical AssessmentsIndustry
- Zero-Shot Keyphrase Generation: Investigating Specialized Instructions and Multi-sample Aggregation on Large Language ModelsFindings
- eC-Tab2Text: Aspect-Based Text Generation from e-Commerce Product TablesIndustry
- kNN For Whisper And Its Effect On Bias And Speaker AdaptationFindings
- kNN Retrieval for Simple and Effective Zero-Shot Multi-speaker Text-to-SpeechShort
- tRAG: Term-level Retrieval-Augmented Generation for Domain-Adaptive RetrievalLong
- uDistil-Whisper: Label-Free Data Filtering for Knowledge Distillation in Low-Data RegimesLong
- “All that Glitters”: Techniques for Evaluations with Unreliable Model and Human AnnotationsFindings
- “Women do not have heart attacks!” Gender Biases in Automatically Generated Clinical Cases in FrenchFindings
NAACL accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.