← All conferences

EMNLP 2025 Accepted Papers

The full list of 3,211 papers accepted at EMNLP 2025 (Conference on Empirical Methods in Natural Language Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

accepted: 3,211
  1. Examining False Positives under Inference Scaling for Mathematical Reasoningaccepted
  2. Examining Multilingual Embedding Models Cross-Lingually Through LLM-Generated Adversarial Examplesaccepted
  3. ExeCoder: Empowering Large Language Models with Executability Representation for Code Translationaccepted
  4. ExeSQL: Self-Taught Text-to-SQL Models with Execution-Driven Bootstrapping for SQL Dialectsaccepted
  5. ExpandR: Teaching Dense Retrievers Beyond Queries with LLM Guidanceaccepted
  6. Expanding before Inferring: Enhancing Factuality in Large Language Models through Premature Layers Interpolationaccepted
  7. Expectation Preference Optimization: Reliable Preference Estimation for Improving the Reasoning Capability of Large Language Modelsaccepted
  8. ExpertGenQA: Open-ended QA generation in Specialized Domainsaccepted
  9. Explainability and Interpretability of Multilingual Large Language Models: A Surveyaccepted
  10. Explainable Chain-of-Thought Reasoning: An Empirical Analysis on State-Aware Reasoning Dynamicsaccepted
  11. Explainable Text Classification with LLMs: Enhancing Performance through Dialectical Prompting and Explanation-Guided Trainingaccepted
  12. Explaining Differences Between Model Pairs in Natural Language through Sample Learningaccepted
  13. Explaining Length Bias in LLM-Based Preference Evaluationsaccepted
  14. Explaining novel senses using definition generation with open language modelsaccepted
  15. Explicit Learning and the LLM in Machine Translationaccepted
  16. Exploiting Prompt-induced Confidence for Black-Box Attacks on LLMsaccepted
  17. Exploration-Driven Reinforcement Learning for Expert Routing Improvement in Mixture-of-Experts Language Modelsaccepted
  18. Exploring Artificial Image Generation for Stance Detectionaccepted
  19. Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignmentaccepted
  20. Exploring Changes in Nation Perception with Nationality-Assigned Personas in LLMsaccepted
  21. Exploring Context Strategies in LLMs for Discourse-Aware Machine Translationaccepted
  22. Exploring Deductive and Inductive Reasoning Capabilities of Large Language Models in Procedural Planningaccepted
  23. Exploring Hyperbolic Hierarchical Structure for Multimodal Rumor Detectionaccepted
  24. Exploring Large Language Models for Detecting Mental Disordersaccepted
  25. Exploring Model Kinship for Merging Large Language Modelsaccepted
  26. Exploring Paraphrasing Strategies for CEFR A1-Level Constraints in LLMsaccepted
  27. Exploring Quality and Diversity in Synthetic Data Generation for Argument Miningaccepted
  28. Exploring Response Uncertainty in MLLMs: An Empirical Evaluation under Misleading Scenariosaccepted
  29. Exploring and Controlling Diversity in LLM-Agent Conversationaccepted
  30. Exploring and Detecting Self-disclosure in Multi-modal posts on Chinese Social Mediaaccepted
  31. Exploring and Evaluating Multimodal Knowledge Reasoning Consistency of Multimodal Large Language Modelsaccepted
  32. Exploring morphology-aware tokenization: A case study on Spanish language modelingaccepted
  33. Exploring the Generalizability of Factual Hallucination Mitigation via Enhancing Precise Knowledge Utilizationaccepted
  34. Exploring the Hidden Capacity of LLMs for One-Step Text Generationaccepted
  35. Exploring the Hidden Reasoning Process of Large Language Models by Misleading Themaccepted
  36. Exploring the Impact of Personality Traits on LLM Bias and Toxicityaccepted
  37. Exploring the Limitations of Mamba in COPY and CoT Reasoningaccepted
  38. Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulationaccepted
  39. Extending Automatic Machine Translation Evaluation to Book-Length Documentsaccepted
  40. Extracting Conceptual Spaces from LLMs Using Prototype Embeddingsaccepted
  41. Extracting Linguistic Information from Large Language Models: Syntactic Relations and Derivational Knowledgeaccepted
  42. Extracting and Combining Abilities For Building Multi-lingual Ability-enhanced Large Language Modelsaccepted
  43. Extractive Fact Decomposition for Interpretable Natural Language Inference in one Forward Passaccepted
  44. ExtremeAIGC: Benchmarking LMM Vulnerability to AI-Generated Extremist Contentaccepted
  45. F2TEval: Human-Aligned Multi-Dimensional Evaluation for Figure-to-Text Taskaccepted
  46. FACTCHECKMATE: Preemptively Detecting and Mitigating Hallucinations in LMsaccepted
  47. FAEDKV: Infinite-Window Fourier Transform for Unbiased KV Cache Compressionaccepted
  48. FANS: Formal Answer Selection for LLM Natural Language Math Reasoning Using Lean4accepted
  49. FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs’ Responsiveness to Human Feedbackaccepted
  50. FC-Attack: Jailbreaking Multimodal Large Language Models via Auto-Generated Flowchartsaccepted
  51. FESTA: Functionally Equivalent Sampling for Trust Assessment of Multimodal LLMsaccepted
  52. FG-PRM: Fine-grained Hallucination Detection and Mitigation in Language Model Mathematical Reasoningaccepted
  53. FIER: Fine-Grained and Efficient KV Cache Retrieval for Long-context LLM Inferenceaccepted
  54. FIRE: Flexible Integration of Data Quality Ratings for Effective Pretrainingaccepted
  55. FISTAPruner: Layer-wise Post-training Pruning for Large Language Modelsaccepted
  56. FLAIRR-TS - Forecasting LLM-Agents with Iterative Refinement and Retrieval for Time Seriesaccepted
  57. FLAMES: Improving LLM Math Reasoning via a Fine-Grained Analysis of the Data Synthesis Pipelineaccepted
  58. FLARE: Faithful Logic-Aided Reasoning and Explorationaccepted
  59. FLRC: Fine-grained Low-Rank Compressor for Efficient LLM Inferenceaccepted
  60. FLUID QA: A Multilingual Benchmark for Figurative Language Usage in Dialogue across English, Chinese, and Koreanaccepted
  61. FM2DS: Few-Shot Multimodal Multihop Data Synthesis with Knowledge Distillation for Question Answeringaccepted
  62. FNSCC: Fuzzy Neighborhood-Aware Self-Supervised Contrastive Clustering for Short Textaccepted
  63. FOSSIL: Harnessing Feedback on Suboptimal Samples for Data-Efficient Generalisation with Imitation Learning for Embodied Vision-and-Language Tasksaccepted
  64. FSTs vs ICL: Generalisation in LLMs for an under-resourced languageaccepted
  65. FaMTEB: Massive Text Embedding Benchmark in Persian Languageaccepted
  66. FaST: Feature-aware Sampling and Tuning for Personalized Preference Alignment with Limited Dataaccepted
  67. FaStFact: Faster, Stronger Long-Form Factuality Evaluations in LLMsaccepted
  68. FacLens: Transferable Probe for Foreseeing Non-Factuality in Fact-Seeking Question Answering of Large Language Modelsaccepted
  69. Facilitating Cognitive Accessibility with LLMs: A Multi-Task Approach to Easy-to-Read Text Generationaccepted
  70. Facilitating Cross-lingual Transfer of Empathy through Language-independent Latent Diffusion: A Case Study in Chineseaccepted
  71. Facilitating Long Context Understanding via Supervised Chain-of-Thought Reasoningaccepted
  72. Fact Verification on Knowledge Graph via Programmatic Graph Reasoningaccepted
  73. FactReasoner: A Probabilistic Approach to Long-Form Factuality Assessment for Large Language Modelsaccepted
  74. Facts Fade Fast: Evaluating Memorization of Outdated Medical Knowledge in Large Language Modelsaccepted
  75. Factuality Beyond Coherence: Evaluating LLM Watermarking Methods for Medical Textsaccepted
  76. Fair Text-Attributed Graph Representation Learningaccepted
  77. Fair or Framed? Political Bias in News Articles Generated by LLMsaccepted
  78. FairCoT: Enhancing Fairness in Text-to-Image Generation via Chain of Thought Reasoning with Multimodal Large Language Modelsaccepted
  79. FairGen: Controlling Sensitive Attributes for Fair Generations in Diffusion Models via Adaptive Latent Guidanceaccepted
  80. Fairness in Automatic Speech Recognition Isn’t a One-Size-Fits-Allaccepted
  81. FaithUn: Toward Faithful Forgetting in Language Models by Investigating the Interconnectedness of Knowledgeaccepted
  82. FakeSV-VLM: Taming VLM for Detecting Fake Short-Video News via Progressive Mixture-Of-Experts Adapteraccepted
  83. False Friends Are Not Foes: Investigating Vocabulary Overlap in Multilingual Language Modelsaccepted
  84. Familiarity-Aware Evidence Compression for Retrieval-Augmented Generationaccepted
  85. Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMsaccepted
  86. Fast Quiet-STaR: Thinking Without Thought Tokensaccepted
  87. Fast, Not Fancy: Rethinking G2P with Rich Data and Statistical Modelsaccepted
  88. FastCuRL: Curriculum Reinforcement Learning with Stage-wise Context Scaling for Efficient Training R1-like Reasoning Modelsaccepted
  89. Faster In-Context Learning for LLMs via N-Gram Trie Speculative Decodingaccepted
  90. Faster and Better LLMs via Latency-Aware Test-Time Scalingaccepted
  91. Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Modelsaccepted
  92. FedCoT: Federated Chain-of-Thought Distillation for Large Language Modelsaccepted
  93. FedMABench: Benchmarking Mobile GUI Agents on Decentralized Heterogeneous User Dataaccepted
  94. Federated Retrieval-Augmented Generation: A Systematic Mapping Studyaccepted
  95. Feel the Difference? A Comparative Analysis of Emotional Arcs in Real and LLM-Generated CBT Sessionsaccepted
  96. Ferret: Faster and Effective Automated Red Teaming with Reward-Based Scoring Techniqueaccepted
  97. Few-Shot Learning Translation from New Languagesaccepted
  98. Few-Shot Open-Set Classification via Reasoning-Aware Decompositionaccepted
  99. FiRST: Finetuning Router-Selective Transformers for Input-Adaptive Latency Reductionaccepted
  100. FicSim: A Dataset for Multi-Faceted Semantic Similarity in Long-Form Fictionaccepted
  101. FigEx: Aligned Extraction of Scientific Figures and Captionsaccepted
  102. FilBench: Can LLMs Understand and Generate Filipino?accepted
  103. FillerSpeech: Towards Human-Like Text-to-Speech Synthesis with Filler Insertion and Filler Style Controlaccepted
  104. FinGEAR: Financial Mapping-Guided Enhanced Answer Retrievalaccepted
  105. FinGrAct: A Framework for FINe-GRrained Evaluation of ACTionability in Explainable Automatic Fact-Checkingaccepted
  106. FinHEAR: Human Expertise and Adaptive Risk-Aware Temporal Reasoning for Financial Decision-Makingaccepted
  107. FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answeringaccepted
  108. FinMTEB: Finance Massive Text Embedding Benchmarkaccepted
  109. FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial Domainaccepted
  110. FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domainaccepted
  111. Financial Risk Relation Identification through Dual-view Adaptationaccepted
  112. Finding your MUSE: Mining Unexpected Solutions Engineaccepted
  113. Fine-Tuned Thoughts: Leveraging Chain-of-Thought Reasoning for Industrial Asset Health Monitoringaccepted
  114. Fine-Tuning Encoder-Decoder Models with Contrastive Learning for In-Context Distractor Generationaccepted
  115. Fine-tuning LLMs with Cross-Attention-based Weight Decay for Bias Mitigationaccepted
  116. Finetuning LLMs for Human Behavior Prediction in Social Science Experimentsaccepted
  117. Fingerprinting LLMs through Survey Item Factor Correlation: A Case Study on Humor Style Questionnaireaccepted
  118. Firewall Routing: Blocking Leads to Better Hybrid Inference for LLMsaccepted
  119. FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Gamesaccepted
  120. Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMsaccepted
  121. FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantizationaccepted
  122. Flexible Thinking for Multimodal Emotional Support Conversation via Reinforcement Learningaccepted
  123. Flexible-length Text Infilling for Discrete Diffusion Modelsaccepted
  124. Flexibly Utilize Memory for Long-Term Conversation via a Fragment-then-Compose Frameworkaccepted
  125. FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Modelsaccepted
  126. FlowMalTrans: Unsupervised Binary Code Translation for Malware Detection Using Flow-Adapter Architectureaccepted
  127. FoREST: Frame of Reference Evaluation in Spatial Reasoning Tasksaccepted
  128. Follow the Flow: Fine-grained Flowchart Attribution with Neurosymbolic Agentsaccepted
  129. Following Length Constraints in Instructionsaccepted
  130. Following Occam’s Razor: Dynamic Combination of Structured Knowledge for Multi-Hop Question Answering using LLMsaccepted
  131. Following the Autoregressive Nature of LLM Embeddings via Compression and Alignmentaccepted
  132. FoodSafeSum: Enabling Natural Language Processing Applications for Food Safety Document Summarization and Analysisaccepted
  133. Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluationaccepted
  134. Foot-In-The-Door: A Multi-turn Jailbreak for LLMsaccepted
  135. For a Fistful of Puns: Evaluating a Puns in Multiword Expressions Identification Algorithm Without Dedicated Datasetaccepted
  136. ForestCast: Open-Ended Event Forecasting with Semantic News Forestaccepted
  137. Forewarned is Forearmed: Pre-Synthesizing Jailbreak-like Instructions to Enhance LLM Safety Guardrail to Potential Attacksaccepted
  138. Forget What You Know about LLMs Evaluations - LLMs are Like a Chameleonaccepted
  139. Forget for Get: A Lightweight Two-phase Gradient Method for Knowledge Editing in Large Language Modelsaccepted
  140. Forget the Unneeded: Backdooring Large Language Models via Contrastive-enhanced Machine Unlearningaccepted
  141. Formalizing Style in Personal Narrativesaccepted
  142. FormosanBench: Benchmarking Low-Resource Austronesian Languages in the Era of Large Language Modelsaccepted
  143. FractalLLM: Lossless Self-Speculative Decoding with Layer Embedded Self-Compressionaccepted
  144. Frame First, Then Extract: A Frame-Semantic Reasoning Pipeline for Zero-Shot Relation Triplet Extractionaccepted
  145. FrameEOL: Semantic Frame Induction using Causal Language Modelsaccepted
  146. Frequency & Compositionality in Emergent Communicationaccepted
  147. Friend or Foe? A Computational Investigation of Semantic False Friends across Romance Languagesaccepted
  148. FroM: Frobenius Norm-Based Data-Free Adaptive Model Mergingaccepted
  149. From A and B to A+B: Can Large Language Models Solve Compositional Math Problems?accepted
  150. From Automation to Autonomy: A Survey on Large Language Models in Scientific Discoveryaccepted
  151. From Benchmark to Better Embeddings: Leveraging Synonym Substitution to Enhance Multimodal Models in Ukrainianaccepted
  152. From Capabilities to Performance: Evaluating Key Functional Properties of LLM Architectures in Penetration Testingaccepted
  153. From Characters to Tokens: Dynamic Grouping with Hierarchical BPEaccepted
  154. From Charts to Fair Narratives: Uncovering and Mitigating Geo-Economic Biases in Chart-to-Textaccepted
  155. From Chat Logs to Collective Insights: Aggregative Question Answeringaccepted
  156. From Confidence to Collapse in LLM Factual Robustnessaccepted
  157. From Cross-Task Examples to In-Task Prompts: A Graph-Based Pseudo-Labeling Framework for In-context Learningaccepted
  158. From General Reward to Targeted Reward: Improving Open-ended Long-context Generation Modelsaccepted
  159. From Generation to Detection: A Multimodal Multi-Task Dataset for Benchmarking Health Misinformationaccepted
  160. From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeaccepted
  161. From Generic Empathy to Personalized Emotional Support: A Self-Evolution Framework for User Preference Alignmentaccepted
  162. From Ground Trust to Truth: Disparities in Offensive Language Judgments on Contemporary Korean Political Discourseaccepted
  163. From Grounding to Manipulation: Case Studies of Foundation Model Integration in Embodied Robotic Systemsaccepted
  164. From Hypothesis to Publication: A Comprehensive Survey of AI-Driven Research Support Systemsaccepted
  165. From Implicit Exploration to Structured Reasoning: Guideline and Refinement for LLMsaccepted
  166. From Input Perception to Predictive Insight: Modeling Model Blind Spots Before They Become Errorsaccepted
  167. From Insight to Exploit: Leveraging LLM Collaboration for Adaptive Adversarial Text Generationaccepted
  168. From KMMLU-Redux to Pro: A Professional Korean Benchmark Suite for LLM Evaluationaccepted
  169. From Knowledge to Treatment: Large Language Model Assisted Biomedical Concept Representation for Drug Repurposingaccepted
  170. From Language to Cognition: How LLMs Outgrow the Human Language Networkaccepted
  171. From Long to Lean: Performance-aware and Adaptive Chain-of-Thought Compression via Multi-round Refinementaccepted
  172. From Noise to Clarity: Filtering Real and LLM-Generated Samples for Enhanced Intent Detectionaccepted
  173. From Parameters to Performance: A Data-Driven Study on LLM Structure and Developmentaccepted
  174. From Personas to Talks: Revisiting the Impact of Personas on LLM-Synthesized Emotional Support Conversationsaccepted
  175. From Problem-Solving to Teaching Problem-Solving: Aligning LLMs with Pedagogy using Reinforcement Learningaccepted
  176. From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Modelsaccepted
  177. From Remembering to Metacognition: Do Existing Benchmarks Accurately Evaluate LLMs?accepted
  178. From Schema to State: Zero-Shot Scheme-Only Dialogue State Tracking via Diverse Synthetic Dialogue and Step-by-Step Distillationaccepted
  179. From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculationsaccepted
  180. From Shortcuts to Balance: Attribution Analysis of Speech-Text Feature Utilization in Distinguishing Original from Machine-Translated Textsaccepted
  181. From Surveys to Narratives: Rethinking Cultural Value Adaptation in LLMsaccepted
  182. From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognitionaccepted
  183. From Token to Action: State Machine Reasoning to Mitigate Overthinking in Information Retrievalaccepted
  184. From Tower to Spire: Adding the Speech Modality to a Translation-Specialist LLMaccepted
  185. From Unaligned to Aligned: Scaling Multilingual LLMs with Multi-Way Parallel Corporaaccepted
  186. From Understanding to Generation: An Efficient Shortcut for Evaluating Language Modelsaccepted
  187. From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Testaccepted
  188. From n-gram to Attention: How Model Architectures Learn and Propagate Bias in Language Modelingaccepted
  189. From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation modelaccepted
  190. FuseChat: Knowledge Fusion of Chat Modelsaccepted
  191. FusionDTI: Fine-grained Binding Discovery with Token-level Fusion for Drug-Target Interactionaccepted
  192. FuzzAug: Data Augmentation by Coverage-guided Fuzzing for Neural Test Generationaccepted
  193. Fuzzy Reasoning Chain (FRC): An Innovative Reasoning Framework from Fuzziness to Clarityaccepted
  194. F²Bench: An Open-ended Fairness Evaluation Benchmark for LLMs with Factuality Considerationsaccepted
  195. G2: Guided Generation for Enhanced Output Diversity in LLMsaccepted
  196. GAMIC: Graph-Aligned Molecular In-context Learning for Molecule Analysis via LLMsaccepted
  197. GAP: a Global Adaptive Pruning Method for Large Language Modelsaccepted
  198. GASE: Generatively Augmented Sentence Encodingaccepted
  199. GATEAU: Selecting Influential Samples for Long Context Alignmentaccepted
  200. GAttention: Gated Attention for the Detection of Abusive Languageaccepted
  201. GCML: Gradient Coherence Guided Meta-Learning for Cross-Domain Emerging Topic Rumor Detectionaccepted
  202. GDLLM: A Global Distance-aware Modeling Approach Based on Large Language Models for Event Temporal Relation Extractionaccepted
  203. GENUINE: Graph Enhanced Multi-level Uncertainty Estimation for Large Language Modelsaccepted
  204. GER-LLM: Efficient and Effective Geospatial Entity Resolution with Large Language Modelaccepted
  205. GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?accepted
  206. GLProtein: Global-and-Local Structure Aware Protein Representation Learningaccepted
  207. GRAD: Generative Retrieval-Aligned Demonstration Sampler for Efficient Few-Shot Reasoningaccepted
  208. GRADA: Graph-based Reranking against Adversarial Documents Attackaccepted
  209. GRADE: Generating multi-hop QA and fine-gRAined Difficulty matrix for RAG Evaluationaccepted
  210. GRAID: Synthetic Data Generation with Geometric Constraints and Multi-Agentic Reflection for Harmful Content Detectionaccepted
  211. GRASP: Replace Redundant Layers with Adaptive Singular Parameters for Efficient Model Compressionaccepted
  212. GRIL: Knowledge Graph Retrieval-Integrated Learning with Large Language Modelsaccepted
  213. GRIT: Guided Relational Integration for Efficient Multi-Table Understandingaccepted
  214. GRPO-Guided Modality Selection Enhanced LoRA-Tuned LLMs for Multimodal Emotion Recognitionaccepted
  215. GRPO-LEAD: A Difficulty-Aware Reinforcement Learning Approach for Concise Mathematical Reasoning in Language Modelsaccepted
  216. GRV-KBQA: A Three-Stage Framework for Knowledge Base Question Answering with Decoupled Logical Structure, Semantic Grounding and Structure-Aware Validationaccepted
  217. GTA: Supervised-Guided Reinforcement Learning for Text Classification with Large Language Modelsaccepted
  218. GUARD: Glocal Uncertainty-Aware Robust Decoding for Effective and Efficient Open-Ended Text Generationaccepted
  219. GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous Explorationaccepted
  220. Gamma-Guard: Lightweight Residual Adapters for Robust Guardrails in Large Language Modelsaccepted
  221. GeLoRA: Geometric Adaptive Ranks For Efficient LoRA Fine-tuningaccepted
  222. GenLink: Generation-Driven Schema-Linking via Multi-Model Learning for Text-to-SQLaccepted
  223. GenPTQ: Green Post-Training Quantization for Large-Scale ASR Models with Mixed-Precision Bit Allocationaccepted
  224. GenPilot: A Multi-Agent System for Test-Time Prompt Optimization in Image Generationaccepted
  225. GenPoE: Generative Passage-level Mixture of Experts for Knowledge Enhancement of LLMsaccepted
  226. Generation-Augmented Retrieval: Rethinking the Role of Large Language Models in Zero-Shot Relation Extractionaccepted
  227. Generative Annotation for ASR Named Entity Correctionaccepted
  228. Generative or Discriminative? Revisiting Text Classification in the Era of Transformersaccepted
  229. Generator-Assistant Stepwise Rollback Framework for Large Language Model Agentaccepted
  230. Genre Matters: How Text Types Interact with Decoding Strategies and Lexical Predictors in Shaping Reading Behavioraccepted
  231. GeoChain: Multimodal Chain-of-Thought for Geographic Reasoningaccepted
  232. GeoDANO: Geometric VLM with Domain Agnostic Vision Encoderaccepted
  233. GeoEdit: Geometric Knowledge Editing for Large Language Modelsaccepted
  234. GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric Reasoningaccepted
  235. Glider: Global and Local Instruction-Driven Expert Routeraccepted
  236. Glitter: A Multi-Sentence, Multi-Reference Benchmark for Gender-Fair German Machine Translationaccepted
  237. GmSLM : Generative Marmoset Spoken Language Modelingaccepted
  238. Golden Touchstone: A Comprehensive Bilingual Benchmark for Evaluating Financial Large Language Modelsaccepted
  239. Good Intentions Beyond ACL: Who Does NLP for Social Good, and Where?accepted
  240. Governance in Motion: Co-evolution of Constitutions and AI models for Scalable Safetyaccepted
  241. GraDaSE: Graph-Based Dataset Search with Examplesaccepted
  242. Graceful Forgetting in Generative Language Modelsaccepted
  243. Graders Should Cheat: Privileged Information Enables Expert-Level Automated Evaluationsaccepted
  244. Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrievalaccepted
  245. Grammar Pruning: Enabling Low-Latency Zero-Shot Task-Oriented Language Models for Edge AIaccepted
  246. Graph-Based Multi-Trait Essay Scoringaccepted
  247. Graph-Guided Textual Explanation Generation Frameworkaccepted
  248. Graph-R1: Incentivizing the Zero-Shot Graph Learning Capability in LLMs via Explicit Reasoningaccepted
  249. Graph-Reward-SQL: Execution-Free Reinforcement Learning for Text-to-SQL via Graph Matching and Stepwise Rewardaccepted
  250. GraphAgent: Agentic Graph Language Assistantaccepted
  251. GraphCheck: Multipath Fact-Checking with Entity-Relationship Graphsaccepted
  252. GraphKV: Breaking the Static Selection Paradigm with Graph-Based KV Cache Evictionaccepted
  253. GreekBarBench: A Challenging Benchmark for Free-Text Legal Reasoning and Citationsaccepted
  254. Grounded Semantic Role Labelling from Synthetic Multimodal Data for Situated Robot Commandsaccepted
  255. Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Modelsaccepted
  256. Grounding Multilingual Multimodal LLMs With Cultural Knowledgeaccepted
  257. Group-Aware Reinforcement Learning for Output Diversity in Large Language Modelsaccepted
  258. Group-SAE: Efficient Training of Sparse Autoencoders for Large Language Models via Layer Groupsaccepted
  259. Grouping Entities with Shared Properties using Multi-Facet Prompting and Property Embeddingsaccepted
  260. Guaranteed Guess: A Language Modeling Approach for CISC-to-RISC Transpilation with Testing Guaranteesaccepted
  261. Guess What I am Thinking: A Benchmark for Inner Thought Reasoning of Role-Playing Language Agentsaccepted
  262. GuessingGame: Measuring the Informativeness of Open-Ended Questions in Large Language Modelsaccepted
  263. GuiLoMo: Allocating Experts and Ranks for LoRA-MoE via Bilevel Optimization with GuidedSelection Vectorsaccepted
  264. Guiding Large Language Models for Biomedical Entity Linking via Restrictive and Contrastive Decodingaccepted
  265. HARE: an entity and relation centric evaluation framework for histopathology reportsaccepted
  266. HATECAT-TR: A Hate Speech Span Detection and Categorization Dataset for Turkishaccepted
  267. HAWK: Highlighting Entity-aware Knowledge for Alleviating Information Sparsity in Long Contextsaccepted
  268. HD-PiSSA: High-Rank Distributed Orthogonal Adaptationaccepted
  269. HDiff: Confidence-Guided Denoising Diffusion for Robust Hyper-relational Link Predictionaccepted
  270. HEAL: A Hypothesis-Based Preference-Aware Analysis Frameworkaccepted
  271. HEAL: An Empirical Study on Hallucinations in Embodied Agents Driven by Large Language Modelsaccepted
  272. HEAL: Hybrid Enhancement with LLM-based Agents for Text-attributed Hypergraph Self-supervised Representation Learningaccepted
  273. HELENE: Hessian Layer-wise Clipping and Gradient Annealing for Accelerating Fine-tuning LLM with Zeroth-order Optimizationaccepted
  274. HESEIA: A community-based dataset for evaluating social biases in large language models, co-designed in real school settings in Latin Americaaccepted
  275. HGAdapter: Hypergraph-based Adapters in Language Models for Code Summarization and Clone Detectionaccepted
  276. HICode: Hierarchical Inductive Coding with LLMsaccepted
  277. HIRAG: Hierarchical-Thought Instruction-Tuning Retrieval-Augmented Generationaccepted
  278. HMCL: Task-Optimal Text Representation Adaptation through Hierarchical Contrastive Learningaccepted
  279. HMoE: Heterogeneous Mixture of Experts for Language Modelingaccepted
  280. HS-STaR: Hierarchical Sampling for Self-Taught Reasoners via Difficulty Estimation and Budget Reallocationaccepted
  281. HVGuard: Utilizing Multimodal Large Language Models for Hateful Video Detectionaccepted
  282. HYDRA: A Multi-Head Encoder-only Architecture for Hierarchical Text Classificationaccepted
  283. Hallucination Detection in LLMs Using Spectral Features of Attention Mapsaccepted
  284. Hallucination Detection in Structured Query Generation via LLM Self-Debatingaccepted
  285. Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreationaccepted
  286. HapticCap: A Multimodal Dataset and Task for Understanding User Experience of Vibration Haptic Signalsaccepted
  287. Hard Negatives, Hard Lessons: Revisiting Training Data Quality for Robust Information Retrieval with LLMsaccepted
  288. Hardware-Aware Parallel Prompt Decoding for Memory-Efficient Acceleration of LLM Inferenceaccepted
  289. Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encodingaccepted
  290. Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Modelsaccepted
  291. HealthCards: Exploring Text-to-Image Generation as Visual Aids for Healthcare Knowledge Democratizing and Educationaccepted
  292. HebID: Detecting Social Identities in Hebrew-language Political Textaccepted
  293. HetGCoT: Heterogeneous Graph-Enhanced Chain-of-Thought LLM Reasoning for Academic Question Answeringaccepted
  294. HiMATE: A Hierarchical Multi-Agent Framework for Machine Translation Evaluationaccepted
  295. Hidden Ghost Hand: Unveiling Backdoor Vulnerabilities in MLLM-Powered Mobile GUI Agentsaccepted
  296. Hidden in Plain Sight: Reasoning in Underspecified and Misspecified Scenarios for Multimodal LLMsaccepted
  297. HierPrompt: Zero-Shot Hierarchical Text Classification with LLM-Enhanced Prototypesaccepted
  298. Hierarchical Bracketing Encodings Work for Dependency Graphsaccepted
  299. Hierarchical Reward Modeling for Fault Localization in Large Code Repositoriesaccepted
  300. HighMATH: Evaluating Math Reasoning of Large Language Models in Breadth and Depthaccepted
  301. HomoGraphAdapter: A Homogeneous Graph Neural Network as an Effective Adapter for Vision-Language Modelsaccepted
  302. HookMoE: A learnable performance compensation strategy of Mixture-of-Experts for LLM inference accelerationaccepted
  303. Hope vs. Hate: Understanding User Interactions with LGBTQ+ News Content in Mainstream US News Media through the Lens of Hope Speechaccepted
  304. Hopscotch: Discovering and Skipping Redundancies in Language Modelsaccepted
  305. How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on tau-benchaccepted
  306. How Diversely Can Language Models Solve Problems? Exploring the Algorithmic Diversity of Model-Generated Codeaccepted
  307. How Do Large Language Models Perform on PDE Discovery: A Coarse-to-fine Perspectiveaccepted
  308. How Do Large Vision-Language Models See Text in Image? Unveiling the Distinctive Role of OCR Headsaccepted
  309. How Do Social Bots Participate in Misinformation Spread? A Comprehensive Dataset and Analysisaccepted
  310. How Does Cognitive Bias Affect Large Language Models? A Case Study on the Anchoring Effect in Price Negotiation Simulationsaccepted
  311. How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysisaccepted
  312. How Does Knowledge Selection Help Retrieval Augmented Generation?accepted
  313. How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparisonaccepted
  314. How Good are LLM-based Rerankers? An Empirical Analysis of State-of-the-Art Reranking Modelsaccepted
  315. How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmarkaccepted
  316. How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigationaccepted
  317. How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM Hallucinationaccepted
  318. How Much Do Large Language Models Know about Human Motion? A Case Study in 3D Avatar Controlaccepted
  319. How Persuasive Is Your Context?accepted
  320. How Private are Language Models in Abstractive Summarization?accepted
  321. How Real Are Synthetic Therapy Conversations? Evaluating Fidelity in Prolonged Exposure Dialoguesaccepted
  322. How Reliable is Multilingual LLM-as-a-Judge?accepted
  323. How Sampling Affects the Detectability of Machine-written texts: A Comprehensive Studyaccepted
  324. How Sememic Components Can Benefit Link Prediction for Lexico-Semantic Knowledge Graphs?accepted
  325. How Well Can Reasoning Models Identify and Recover from Unhelpful Thoughts?accepted
  326. How a Bilingual LM Becomes Bilingual: Tracing Internal Representations with Sparse Autoencodersaccepted
  327. How do Language Models Generate Slang: A Systematic Comparison between Human and Machine-Generated Slang Usagesaccepted
  328. How do Language Models Reshape Entity Alignment? A Survey of LM-Driven EA Methods: Advances, Benchmarks, and Futureaccepted
  329. How do autoregressive transformers solve full addition?accepted
  330. How to Generalize the Detection of AI-Generated Text: Confounding Neuronsaccepted
  331. How to Make Large Language Models Generate 100% Valid Molecules?accepted
  332. How to Protect Yourself from 5G Radiation? Investigating LLM Responses to Implicit Misinformationaccepted
  333. How to inject knowledge efficiently? Knowledge Infusion Scaling Law for Pre-training Large Language Modelsaccepted
  334. Human-Inspired Obfuscation for Model Unlearning: Local and Global Strategies with Hyperbolic Representationsaccepted
  335. Humanity’s Last Code Exam: Can Advanced LLMs Conquer Human’s Hardest Code Competition?accepted
  336. Humanizing Machines: Rethinking LLM Anthropomorphism Through a Multi-Level Framework of Designaccepted
  337. Humans Hallucinate Too: Language Models Identify and Correct Subjective Annotation Errors With Label-in-a-Haystack Promptsaccepted
  338. Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comicsaccepted
  339. HydraOpt: Navigating the Efficiency-Performance Trade-off of Adapter Mergingaccepted
  340. HydraRAG: Structured Cross-Source Enhanced Large Language Model Reasoningaccepted
  341. HypER: Literature-grounded Hypothesis Generation and Distillation with Provenanceaccepted
  342. HyperKGR: Knowledge Graph Reasoning in Hyperbolic Space with Graph Neural Network Encoding Symbolic Pathaccepted
  343. I-GUARD: Interpretability-Guided Parameter Optimization for Adversarial Defenseaccepted
  344. ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignmentaccepted
  345. ICL CIPHERS: Quantifying ”Learning” in In-Context Learning via Substitution Ciphersaccepted
  346. ICL-Bandit: Relevance Labeling in Advertisement Recommendation Systems via LLMaccepted
  347. ICLER: Intent CLassification with Enhanced Reasoningaccepted
  348. ICR: Iterative Clarification and Rewriting for Conversational Searchaccepted
  349. IG-Pruning: Input-Guided Block Pruning for Large Language Modelsaccepted
  350. IIET: Efficient Numerical Transformer via Implicit Iterative Euler Methodaccepted
  351. IL-PCSR: Legal Corpus for Prior Case and Statute Retrievalaccepted
  352. INDOORWORLD : Integrating Physical Task Solving and Social Simulation in A Heterogeneous Multi-Agent Environmentaccepted
  353. INREACT: An Inspire-Then-Reinforce Training Framework For Multimodal GUI Agentaccepted
  354. IP-Dialog: Evaluating Implicit Personalization in Dialogue Systems with Synthetic Dataaccepted
  355. IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agentsaccepted
  356. ISACL: Internal State Analyzer for Copyrighted Training Data Leakageaccepted
  357. Icon2: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulationaccepted
  358. Identification of Multiple Logical Interpretations in Counter-Argumentsaccepted
  359. Identifying & Interactively Refining Ambiguous User Goals for Data Visualization Code Generationaccepted
  360. Identifying Aspects in Peer Reviewsaccepted
  361. Identifying Noise in Human-Created Datasets using Training Dynamics from Generative Modelsaccepted
  362. Identifying Pre-training Data in LLMs: A Neuron Activation-Based Detection Frameworkaccepted
  363. Identifying Rare Languages in Common Crawl Data is a Needles-in-a-Haystack Problemaccepted
  364. Identifying Unlearned Data in LLMs via Membership Inference Attacksaccepted
  365. Identifying and Answering Questions with False Assumptions: An Interpretable Approachaccepted
  366. Idiosyncratic Versus Normative Modeling of Atypical Speech Recognition: Dysarthric Case Studiesaccepted
  367. Idola Tribus of AI: Large Language Models tend to perceive order where none existsaccepted
  368. Igniting Creative Writing in Small Language Models: LLM-as-a-Judge versus Multi-Agent Refined Rewardsaccepted
  369. Image Difference Captioning via Adversarial Preference Optimizationaccepted
  370. Image Embedding Sampling Method for Diverse Captioningaccepted
  371. Imagination and Contemplation: A Balanced Framework for Semantic-Augmented Multimodal Machine Translationaccepted
  372. ImpRAG: Retrieval-Augmented Generation with Implicit Queriesaccepted
  373. ImpliRet: Benchmarking the Implicit Fact Retrieval Challengeaccepted
  374. Implicit Behavioral Alignment of Language Agents in High-Stakes Crowd Simulationsaccepted
  375. Implicit Values Embedded in How Humans and LLMs Complete Subjective Everyday Tasksaccepted
  376. Improbable Bigrams Expose Vulnerabilities of Incomplete Tokens in Byte-Level Tokenizersaccepted
  377. Improve LLM-as-a-Judge Ability as a General Abilityaccepted
  378. Improving Alignment in LVLMs with Debiased Self-Judgmentaccepted
  379. Improving Chemical Understanding of LLMs via SMILES Parsingaccepted
  380. Improving Clustering with Positive Pairs Generated from LLM-Driven Labelsaccepted
  381. Improving Context Fidelity via Native Retrieval-Augmented Reasoningaccepted
  382. Improving Cross Lingual Transfer by Pretraining with Active Forgettingaccepted
  383. Improving Handshape Representations for Sign Language Processing: A Graph Neural Network Approachaccepted
  384. Improving Influence-based Instruction Tuning Data Selection for Balanced Learning of Diverse Capabilitiesaccepted
  385. Improving Informally Romanized Language Identificationaccepted
  386. Improving Instruct Models for Free: A Study on Partial Adaptationaccepted
  387. Improving LLM Reasoning through Interpretable Role-Playing Steeringaccepted
  388. Improving LLM-as-a-Judge Inference with the Judgment Distributionaccepted
  389. Improving Language Model Personas via Rationalization with Psychological Scaffoldsaccepted
  390. Improving Large Language Model Safety with Contrastive Representation Learningaccepted
  391. Improving Large Language Models Function Calling and Interpretability via Guided-Structured Templatesaccepted
  392. Improving Low-Resource Sequence Labeling with Knowledge Fusion and Contextual Label Explanationsaccepted
  393. Improving Multilingual Retrieval-Augmented Language Models through Dialectic Reasoning Argumentationsaccepted
  394. Improving Neutral Point-of-View Generation with Data- and Parameter-Efficient RLaccepted
  395. Improving Online Job Advertisement Analysis via Compositional Entity Extractionaccepted
  396. Improving Preference Alignment of LLM with Inference-Free Self-Refinementaccepted
  397. Improving Prompt Generalization for Cross-prompt Essay Trait Scoring from the Scoring-invariance Perspectiveaccepted
  398. Improving Reasoning Capabilities in Small Models through Mixture-of-layers Distillation with Stepwise Attention on Key Informationaccepted
  399. Improving Rule-based Reasoning in LLMs using Neurosymbolic Representationsaccepted
  400. Improving Task Diversity in Label Efficient Supervised Finetuning of LLMsaccepted
  401. Improving Zero-shot Sentence Decontextualisation with Content Selection and Planningaccepted
  402. Improving the Language Understanding Capabilities of Large Language Models Using Reinforcement Learningaccepted
  403. Improving the Quality of Web-mined Parallel Corpora of Low-Resource Languages using Debiasing Heuristicsaccepted
  404. In Benchmarks We Trust ... Or Not?accepted
  405. In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varietiesaccepted
  406. InFact: Informativeness Alignment for Improved LLM Factualityaccepted
  407. InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning Stylesaccepted
  408. Inceptive Transformers: Enhancing Contextual Representations through Multi-Scale Feature Learning Across Domains and Languagesaccepted
  409. Inclusive Leadership in the Age of AI: A Dataset and Comparative Study of LLMs vs. Real-Life Leaders in Workplace Action Planningaccepted
  410. Incorporating Diverse Perspectives in Cultural Alignment: Survey of Evaluation Benchmarks Through A Three-Dimensional Frameworkaccepted
  411. IndiGEC: Multilingual Grammar Error Correction for Low-Resource Indian Languagesaccepted
  412. IndoSafety: Culturally Grounded Safety for LLMs in Indonesian Languagesaccepted
  413. Inducing Argument Facets for Faithful Opinion Summarizationaccepted
  414. Inductive Reasoning on Few-Shot Knowledge Graphs with Task-Aware Language Modelsaccepted
  415. Inefficiencies of Meta Agents for Agent Designaccepted
  416. InfAL: Inference Time Adversarial Learning for Improving Research Ideationaccepted
  417. InfiMM-WebMath-40B: Advancing Multimodal Pre-Training for Enhanced Mathematical Reasoningaccepted
  418. Infini-gram mini: Exact n-gram Search at the Internet Scale with FM-Indexaccepted
  419. InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Showsaccepted
  420. InfoGain-RAG: Boosting Retrieval-Augmented Generation through Document Information Gain-based Reranking and Filteringaccepted
  421. Information Integration in Large Language Models is Gated by Linguistic Structural Markersaccepted
  422. Injecting Domain-Specific Knowledge into Large Language Models: A Comprehensive Surveyaccepted
  423. InsBank: Evolving Instruction Subset for Ongoing Alignmentaccepted
  424. Insights into using temporal coordinated behaviour to explore connections between social media posts and influenceaccepted
  425. Instability in Downstream Task Performance During LLM Pretrainingaccepted
  426. Instance-level Randomization: Toward More Stable LLM Evaluationsaccepted
  427. Instructing Large Language Models for Low-Resource Languages: A Systematic Study for Basqueaccepted
  428. InteGround: On the Evaluation of Verification and Retrieval Planning in Integrative Groundingaccepted
  429. Integral Transformer: Denoising Attention, Not Too Much Not Too Littleaccepted
  430. Intelligent Document Parsing: Towards End-to-end Document Parsing via Decoupled Content Parsing and Layout Groundingaccepted
  431. Intent-aware Schema Generation and Refinement for Literature Review Tablesaccepted
  432. IntentionFrame: A Semi-Structured, Multi-Aspect Framework for Fine-Grained Conversational Intention Understandingaccepted
  433. Inter-sentence Context Modeling and Structure-aware Representation Enhancement for Conversational Sentiment Quadruple Extractionaccepted
  434. InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models with Human Feedbackaccepted
  435. InterIDEAS: Philosophical Intertextuality via LLMsaccepted
  436. InteractSpeech: A Speech Dialogue Interaction Corpus for Spoken Dialogue Modelaccepted
  437. Interdisciplinary Research in Conversation: A Case Study in Computational Morphology for Language Documentationaccepted
  438. Interesting Culture: Social Relation Recognition from Videos via Culture De-confoundingaccepted
  439. Internal Chain-of-Thought: Empirical Evidence for Layer‐wise Subtask Scheduling in LLMsaccepted
  440. Internal states before wait modulate reasoning patternsaccepted
  441. Interpretability Analysis of Arithmetic In-Context Learning in Large Language Modelsaccepted
  442. Interpretable Mnemonic Generation for Kanji Learning via Expectation-Maximizationaccepted
  443. Interpretable Text Embeddings and Text Similarity Explanation: A Surveyaccepted
  444. Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safetyaccepted
  445. IntrEx: A Dataset for Modeling Engagement in Educational Conversationsaccepted
  446. Intrinsic Test of Unlearning Using Parametric Knowledge Tracesaccepted
  447. Introducing Spotlight: A Novel Approach for Generating Captivating Key Information from Documentsaccepted
  448. Investigating Controversy Framing across Topics on Social Mediaaccepted
  449. Investigating Dictionary Expansion for Video-based Sign Language Dictionariesaccepted
  450. Investigating How Pre-training Data Leakage Affects Models’ Reproduction and Detection Capabilitiesaccepted
  451. Investigating Multi-layer Representations for Dense Passage Retrievalaccepted
  452. Investigating Neurons and Heads in Transformer-based LLMs for Typographical Errorsaccepted
  453. Investigating Pedagogical Teacher and Student LLM Agents: Genetic Adaptation Meets Retrieval-Augmented Generation Across Learning Stylesaccepted
  454. Investigating Value-Reasoning Reliability in Small Large Language Modelsaccepted
  455. Investigating the Impact of Conceptual Metaphors on LLM-based NLI through Shapley Interactionsaccepted
  456. Investigating the interaction of linguistic and mathematical reasoning in language models using multilingual number puzzlesaccepted
  457. Invisible Entropy: Towards Safe and Efficient Low-Entropy LLM Watermarkingaccepted
  458. Invisible Prompts, Visible Threats: Malicious Font Injection in External Resources for Large Language Modelsaccepted
  459. Invoke Interfaces Only When Needed: Adaptive Invocation for Large Language Models in Question Answeringaccepted
  460. IoTMigrator: LLM-driven Embedded IoT Code Migration across Different OSes for Cloud-device Integrationaccepted
  461. Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understandingaccepted
  462. Is Safety Standard Same for Everyone? User-Specific Safety Evaluation of Large Language Modelsaccepted
  463. Is the Top Still Spinning? Evaluating Subjectivity in Narrative Understandingaccepted
  464. Iterative Multilingual Spectral Attribute Erasureaccepted
  465. Iterative Prompt Refinement for Safer Text-to-Image Generationaccepted
  466. It’s All About In-Context Learning! Teaching Extremely Low-Resource Languages to LLMsaccepted
  467. JI2S: Joint Influence‐Aware Instruction Data Selection for Efficient Fine‐Tuningaccepted
  468. JOLT-SQL: Joint Loss Tuning of Text-to-SQL with Confusion-aware Noisy Schema Samplingaccepted
  469. JUDGEBERT: Assessing Legal Meaning Preservation Between Sentencesaccepted
  470. JUREX-4E: Juridical Expert-Annotated Four-Element Knowledge Base for Legal Reasoningaccepted
  471. Jailbreak Attack Initializations as Extractors of Compliance Directionsaccepted
  472. Jailbreak Distillation: Renewable Safety Benchmarkingaccepted
  473. Jailbreak LLMs through Internal Stance Manipulationaccepted
  474. Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibilityaccepted
  475. Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Modelsaccepted
  476. Job Unfair: An Investigation of Gender and Occupational Bias in Free-Form Text Completions by LLMsaccepted
  477. Joint Enhancement of Relational Reasoning for Long-Context LLMsaccepted
  478. Joint Modeling of Entities and Discourse Relations for Coherence Assessmentaccepted
  479. Journalism-Guided Agentic In-context Learning for News Stance Detectionaccepted
  480. Judge and Improve: Towards a Better Reasoning of Knowledge Graphs with Large Language Modelsaccepted
  481. Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Modelsaccepted
  482. Judging with Many Minds: Do More Perspectives Mean Less Prejudice? On Bias Amplification and Resistance in Multi-Agent Based LLM-as-Judgeaccepted
  483. KAHAN: Knowledge-Augmented Hierarchical Analysis and Narration for Financial Data Narrationaccepted
  484. KBAlign: Efficient Self Adaptation on Specific Textual Knowledge Basesaccepted
  485. KBM: Delineating Knowledge Boundary for Adaptive Retrieval in Large Language Modelsaccepted
  486. KCS: Diversify Multi-hop Question Generation with Knowledge Composition Samplingaccepted
  487. KELE: A Multi-Agent Framework for Structured Socratic Teaching with Large Language Modelsaccepted
  488. KELE: Residual Knowledge Erasure for Enhanced Multi-hop Reasoning in Knowledge Editingaccepted
  489. KERAG: Knowledge-Enhanced Retrieval-Augmented Generation for Advanced Question Answeringaccepted
  490. KG-CQR: Leveraging Structured Relation Representations in Knowledge Graphs for Contextual Query Retrievalaccepted
  491. KG-RAG: Enhancing GUI Agent Decision-Making via Knowledge Graph-Driven Retrieval-Augmented Generationaccepted
  492. KGE Calibrator: An Efficient Probability Calibration Method of Knowledge Graph Embedding Models for Trustworthy Link Predictionaccepted
  493. KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language Modelsaccepted
  494. KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contextsaccepted
  495. KaeDe: Progressive Generation of Logical Forms via Knowledge-Aware Question Decomposition for Improved KBQAaccepted
  496. Keep Security! Benchmarking Security Policy Preservation in Large Language Model Contexts Against Indirect Attacks in Question Answeringaccepted
  497. Knowing More, Acting Better: Hierarchical Representation for Embodied Decision-Makingaccepted
  498. Knowledge Editing through Chain-of-Thoughtaccepted
  499. Knowledge Graph-Driven Memory Editing with Directional Interventionsaccepted
  500. Knowledge-Aware Co-Reasoning for Multidisciplinary Collaborationaccepted
  501. Knowledge-enhanced Multimodal ECG Representation Learning with Arbitrary-Lead Inputsaccepted
  502. Ko-LongRAG: A Korean Long-Context RAG Benchmark Built with a Retrieval-Free Approachaccepted
  503. KoACD: The First Korean Adolescent Dataset for Cognitive Distortion Analysis via Role-Switching Multi-LLM Negotiationaccepted
  504. KoBLEX: Open Legal Question Answering with Multi-hop Reasoningaccepted
  505. KoLEG: On-the-Fly Korean Legal Knowledge Editing with Continuous Retrievalaccepted
  506. Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidanceaccepted
  507. Krikri: Advancing Open Large Language Models for Greekaccepted
  508. KurTail : Kurtosis-based LLM Quantizationaccepted
  509. LAGCL4Rec: When LLMs Activate Interactions Potential in Graph Contrastive Learning for Recommendationaccepted
  510. LASER: An LLM-based ASR Scoring and Evaluation Rubricaccepted
  511. LATTE: Learning to Think with Vision Specialistsaccepted
  512. LAVa: Layer-wise KV Cache Eviction with Dynamic Budget Allocationaccepted
  513. LAWCAT: Efficient Distillation from Quadratic to Linear Attention with Convolution across Tokens for Long Context Modelingaccepted
  514. LC-Eval: A Bilingual Multi-Task Evaluation Benchmark for Long-Context Understandingaccepted
  515. LCAN: A Label-Aware Contrastive Attention Network for Multi-Intent Recognition and Slot Filling in Task-Oriented Dialogue Systemsaccepted
  516. LCES: Zero-shot Automated Essay Scoring via Pairwise Comparisons Using Large Language Modelsaccepted
  517. LEAF: Large Language Diffusion Model for Time Series Forecastingaccepted
  518. LEO-MINI: An Efficient Multimodal Large Language Model using Conditional Token Reduction and Mixture of Multi-Modal Expertsaccepted
  519. LGA: LLM-GNN Aggregation for Temporal Evolution Attribute Graph Predictionaccepted
  520. LIDDIA: Language-based Intelligent Drug Discovery Agentaccepted
  521. LIFTED: Multimodal Clinical Trial Outcome Prediction via Large Language Models and Mixture-of-Expertsaccepted
  522. LILaC: Late Interacting in Layered Component Graph for Open-domain Multimodal Multihop Retrievalaccepted
  523. LLM Agents at the Roundtable: A Multi-Perspective and Dialectical Reasoning Framework for Essay Scoringaccepted
  524. LLM Agents for Education: Advances and Applicationsaccepted
  525. LLM Bias Detection and Mitigation through the Lens of Desired Distributionsaccepted
  526. LLM Distillation for Efficient Few-Shot Multiple Choice Question Answeringaccepted
  527. LLM Jailbreak Detection for (Almost) Free!accepted
  528. LLM-Based Web Data Collection for Research Dataset Creationaccepted
  529. LLM-Driven Completeness and Consistency Evaluation for Cultural Heritage Data Augmentation in Cross-Modal Retrievalaccepted
  530. LLM-Driven Implicit Target Augmentation and Fine-Grained Contextual Modeling for Zero-Shot and Few-Shot Stance Detectionaccepted
  531. LLM-Guided Co-Training for Text Classificationaccepted
  532. LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognitionaccepted
  533. LLM-Independent Adaptive RAG: Let the Question Speak for Itselfaccepted
  534. LLM-OREF: An Open Relation Extraction Framework Based on Large Language Modelsaccepted
  535. LLM-based Conversational Recommendation Agents with Collaborative Verbalized Experienceaccepted
  536. LLM-based Open Domain Planning by Leveraging Entity-Attribute-Level Domain Modelsaccepted
  537. LLM-empowered Dynamic Prompt Routing for Vision-Language Models Tuning under Long-Tailed Distributionsaccepted
  538. LLMAP: LLM-Assisted Multi-Objective Route Planning with User Preferencesaccepted
  539. LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validationaccepted
  540. LLMs Behind the Scenes: Enabling Narrative Scene Illustrationaccepted
  541. LLMs Can Compensate for Deficiencies in Visual Representationsaccepted
  542. LLMs Don’t Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanationsaccepted
  543. LLMs Reproduce Stereotypes of Sexual and Gender Minoritiesaccepted
  544. LLMs are Better Than You Think: Label-Guided In-Context Learning for Named Entity Recognitionaccepted
  545. LLMs are Privacy Erasableaccepted
  546. LLMs as World Models: Data-Driven and Human-Centered Pre-Event Simulation for Disaster Impact Assessmentaccepted
  547. LLMs as a synthesis between symbolic and distributed approaches to languageaccepted
  548. LLMs cannot spot math errors, even when allowed to peek into the solutionaccepted
  549. LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?accepted
  550. LLMsPark: A Benchmark for Evaluating Large Language Models in Strategic Gaming Contextsaccepted
  551. LLaMP: Large Language Model Made Powerful for High-fidelity Materials Knowledge Retrievalaccepted
  552. LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learningaccepted
  553. LM-Searcher: Cross-domain Neural Architecture Search with LLMs via Unified Numerical Encodingaccepted
  554. LM2Protein: A Structure-to-Token Protein Large Language Modelaccepted
  555. LMR-BENCH: Evaluating LLM Agent’s Ability on Reproducing Language Modeling Researchaccepted
  556. LMUNIT: Fine-grained Evaluation with Natural Language Unit Testsaccepted
  557. LNE-Blocking: An Efficient Framework for Contamination Mitigation Evaluation on Large Language Modelsaccepted
  558. LOHRec: Leveraging Order and Hierarchy in Generative Sequential Recommendationaccepted
  559. LORAXBENCH: A Multitask, Multilingual Benchmark Suite for 20 Indonesian Languagesaccepted
  560. LORE: Continual Logit Rewriting Fosters Faithful Generationaccepted
  561. LRPLAN: A Multi-Agent Collaboration of Large Language and Reasoning Models for Planning with Implicit & Explicit Constraintsaccepted
  562. LSRL: Process-Supervised GRPO on Latent Recurrent States Improves Mathematical Reasoningaccepted
  563. LUME: LLM Unlearning with Multitask Evaluationsaccepted
  564. LVLMs are Bad at Overhearing Human Referential Communicationaccepted
  565. LaMDAgent: An Autonomous Framework for Post-Training Pipeline Optimization via LLM Agentsaccepted
  566. LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profilesaccepted
  567. LaMP-QA: A Benchmark for Personalized Long-form Question Answeringaccepted
  568. LaMP-Val: Large Language Models Empower Personalized Valuation in Auctionaccepted
  569. Label Set Optimization via Activation Distribution Kurtosis for Zero-Shot Classification with Generative Modelsaccepted
  570. LangProBe: a Language Program Benchmarkaccepted
  571. Language Mixing in Reasoning Language Models: Patterns, Impact, and Internal Causesaccepted
  572. Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformersaccepted
  573. Language Models Can Easily Learn to Reason from Demonstrationsaccepted
  574. Language Models Can be Efficiently Steered via Minimal Embedding Layer Transformationsaccepted
  575. Language Models Identify Ambiguities and Exploit Loopholesaccepted
  576. Language Models as Causal Effect Generatorsaccepted
  577. Language Models as Continuous Self-Evolving Data Engineersaccepted
  578. Language models can learn implicit multi-hop reasoning, but only if they have lots of training dataaccepted
  579. Language-Guided Temporal Token Pruning for Efficient VideoLLM Processingaccepted
  580. Language-Informed Synthesis of Rational Agent Models for Grounded Theory-of-Mind Reasoning On-the-flyaccepted
  581. Language-Specific Layer Matters: Efficient Multilingual Enhancement for Large Vision-Language Modelsaccepted
  582. Language-to-Space Programming for Training-Free 3D Visual Groundingaccepted
  583. Languages Still Left Behind: Toward a Better Multilingual Machine Translation Benchmarkaccepted
  584. Large Language Model Agents in Finance: A Survey Bridging Research, Practice, and Real-World Deploymentaccepted
  585. Large Language Model Evaluation via Matrix Nuclear-Normaccepted
  586. Large Language Models Badly Generalize across Option Length, Problem Types, and Irrelevant Noun Replacementsaccepted
  587. Large Language Models Discriminate Against Speakers of German Dialectsaccepted
  588. Large Language Models Do Multi-Label Classification Differentlyaccepted
  589. Large Language Models Have Intrinsic Meta-Cognition, but Need a Good Lensaccepted
  590. Large Language Models Meet Knowledge Graphs for Question Answering: Synthesis and Opportunitiesaccepted
  591. Large Language Models Might Not Care What You Are Saying: Prompt Format Beats Descriptionsaccepted
  592. Large Language Models Threaten Language’s Epistemic and Communicative Foundationsaccepted
  593. Large Language Models as Reader for Bias Detectionaccepted
  594. Large Language Models as Realistic Microservice Trace Generatorsaccepted
  595. Large Language Models for Automated Literature Review: An Evaluation of Reference Generation, Abstract Writing, and Review Compositionaccepted
  596. Large Language Models for Controllable Multi-property Multi-objective Molecule Optimizationaccepted
  597. Large Language Models for Multilingual Previously Fact-Checked Claim Detectionaccepted
  598. Large Language Models with Temporal Reasoning for Longitudinal Clinical Summarization and Predictionaccepted
  599. Large Vision-Language Model Alignment and Misalignment: A Survey Through the Lens of Explainabilityaccepted
  600. LastingBench: Defend Benchmarks Against Knowledge Leakageaccepted
  601. Latent Inter-User Difference Modeling for LLM Personalizationaccepted
  602. Layer Duplication in LLMsaccepted
  603. Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignmentaccepted
  604. Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledgeaccepted
  605. Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech Representationsaccepted
  606. Layered Insights: Generalizable Analysis of Human Authorial Style by Leveraging All Transformer Layersaccepted
  607. LeTS: Learning to Think-and-Search via Process-and-Outcome Reward Hybridizationaccepted
  608. Leaky Thoughts: Large Reasoning Models Are Not Private Thinkersaccepted
  609. LeanK: Learnable K Cache Channel Pruning for Efficient Decodingaccepted
  610. Learn and Unlearn: Addressing Misinformation in Multilingual LLMsaccepted
  611. Learning API Functionality from In-Context Demonstrations for Tool-based Agentsaccepted
  612. Learning Contextual Retrieval for Robust Conversational Searchaccepted
  613. Learning Is Not A Race: Improving Retrieval in Language Models via Equal Learningaccepted
  614. Learning Like Humans: Advancing LLM Reasoning Capabilities via Adaptive Difficulty Curriculum Learning and Expert-Guided Self-Reformulationaccepted
  615. Learning SQL Like a Human: Structure-Aware Curriculum Learning for Text-to-SQL Generationaccepted
  616. Learning Subjective Label Distributions via Sociocultural Descriptorsaccepted
  617. Learning Trajectories of Figurative Language for Pre-Trained Language Modelsaccepted
  618. Learning What to Remember: Adaptive Probabilistic Memory Retention for Memory-Efficient Language Modelsaccepted
  619. Learning from Diverse Reasoning Paths with Routing and Collaborationaccepted
  620. Learning from Few Samples: A Novel Approach for High-Quality Malcode Generationaccepted
  621. Learning to Align: Addressing Character Frequency Distribution Shifts in Handwritten Text Recognitionaccepted
  622. Learning to Ask: When LLM Agents Meet Unclear Instructionaccepted
  623. Learning to Describe Implicit Changes: Noise-robust Pre-training for Image Difference Captioningaccepted
  624. Learning to Instruct: Fine-Tuning a Task-Aware Instruction Optimizer for Black-Box LLMsaccepted
  625. Learning to See through Sound: From VggCaps to Multi2Cap for Richer Automated Audio Captioningaccepted
  626. Legal Fact Prediction: The Missing Piece in Legal Judgment Predictionaccepted
  627. Legal Mathematical Reasoning with LLMs: Procedural Alignment through Two-Stage Reinforcement Learningaccepted
  628. LegalSearchLM: Rethinking Legal Case Retrieval as Legal Elements Generationaccepted
  629. LegoSLM: Connecting LLM with Speech Encoder using CTC Posteriorsaccepted
  630. Lemma Dilemma: On Lemma Generation Without Domain- or Language-Specific Training Dataaccepted
  631. Lemmatization as a Classification Task: Results from Arabic across Multiple Genresaccepted
  632. Lemmatization of Polish Multi-word Expressionsaccepted
  633. Length Representations in Large Language Modelsaccepted
  634. Less Is More? Examining Fairness in Pruned Large Language Models for Summarising Opinionsaccepted
  635. Less Is MuRE: Revisiting Shallow Knowledge Graph Embeddingsaccepted
  636. Less is More: The Effectiveness of Compact Typological Language Representationsaccepted
  637. Let Them Down Easy! Contextual Effects of LLM Guardrails on User Perceptions and Preferencesaccepted
  638. Let’s Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models’ Understanding of Sportsaccepted
  639. Let’s Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM’s Math Capabilityaccepted
  640. Leveraging 3D Gaussian for Temporal Knowledge Graph Embeddingaccepted
  641. Leveraging Cognitive Complexity of Texts for Contextualization in Dense Retrievalaccepted
  642. Leveraging High-Resource English Corpora for Cross-lingual Domain Adaptation in Low-Resource Japanese Medicine via Continued Pre-trainingaccepted
  643. Leveraging Knowledge Graph-Enhanced LLMs for Context-Aware Medical Consultationaccepted
  644. Leveraging Large Models to Evaluate Novel Content: A Case Study on Advertisement Creativityaccepted
  645. Leveraging Loanword Constraints for Improving Machine Translation in a Low-Resource Multilingual Contextaccepted
  646. Leveraging Multilingual Training for Authorship Representation: Enhancing Generalization across Languages and Domainsaccepted
  647. Leveraging Semantic Triples for Private Document Generation with Local Differential Privacy Guaranteesaccepted
  648. Leveraging Text-to-Text Transformers as Classifier Chain for Few-Shot Multi-Label Classificationaccepted
  649. Leveraging Unpaired Feedback for Long-Term LLM-based Recommendation Tuningaccepted
  650. Leveraging What’s Overfixed: Post-Correction via LLM Grammatical Error Overcorrectionaccepted
  651. Leveraging the Cross-Domain & Cross-Linguistic Corpus for Low Resource NMT: A Case Study On Bhili-Hindi-English Parallel Corpusaccepted
  652. LexTime: A Benchmark for Temporal Ordering of Legal Eventsaccepted
  653. LiTEx: A Linguistic Taxonomy of Explanations for Understanding Within-Label Variation in Natural Language Inferenceaccepted
  654. LiTransProQA: An LLM-based Literary Translation Evaluation Metric with Professional Question Answeringaccepted
  655. Liaozhai through the Looking-Glass: On Paratextual Explicitation of Culture-Bound Terms in Machine Translationaccepted
  656. Lifelong Knowledge Editing requires Better Regularizationaccepted
  657. LightRAG: Simple and Fast Retrieval-Augmented Generationaccepted
  658. LightThinker: Thinking Step-by-Step Compressionaccepted
  659. Likelihood Variance as Text Importance for Resampling Texts to Map Language Modelsaccepted
  660. LimRank: Less is More for Reasoning-Intensive Information Rerankingaccepted
  661. LimaCost: Data Valuation for Instruction Tuning of Large Language Modelsaccepted
  662. Linear Steerability in Language Models: When It Emerges and How It Evolvesaccepted
  663. Linear-Time Demonstration Selection for In-Context Learning via Gradient Estimationaccepted
  664. LingGym: How Far Are LLMs from Thinking Like Field Linguists?accepted
  665. LinguaLens: Towards Interpreting Linguistic Mechanisms of Large Language Models via Sparse Auto-Encoderaccepted
  666. Linguistic Alignment Predicts Learning in Small Group Tutoring Sessionsaccepted
  667. Linguistic Neuron Overlap Patterns to Facilitate Cross-lingual Transfer on Low-resource Languagesaccepted
  668. Linguistic and Embedding-Based Profiling of Texts Generated by Humans and Large Language Modelsaccepted
  669. Linguistically-Controlled Paraphrase Generationaccepted
  670. LinkAlign: Scalable Schema Linking for Real-World Large-Scale Multi-Database Text-to-SQLaccepted
  671. LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximationaccepted
  672. LiteraryQA: Towards Effective Evaluation of Long-document Narrative QAaccepted
  673. LlmFixer: Fix the Helpfulness of Defensive Large Language Modelsaccepted
  674. LoCt-Instruct: An Automatic Pipeline for Constructing Datasets of Logical Continuous Instructionsaccepted
  675. LoRA-MGPO: Mitigating Double Descent in Low-Rank Adaptation via Momentum-Guided Perturbation Optimizationaccepted
  676. LoRA-PAR: A Flexible Dual-System LoRA Partitioning Approach to Efficient LLM Fine-Tuningaccepted
  677. LoRACoE: Improving Large Language Model via Composition-based LoRA Expertaccepted
  678. LoRATK: LoRA Once, Backdoor Everywhere in the Share-and-Play Ecosystemaccepted
  679. LoRE-Merging: Exploring Low-Rank Estimation For Large Language Model Mergingaccepted
  680. LoRaDA: Low-Rank Direct Attention Adaptation for Efficient LLM Fine-tuningaccepted
  681. LoSiA: Efficient High-Rank Fine-Tuning via Subnet Localization and Optimizationaccepted
  682. Local Normalization Distortion and the Thermodynamic Formalism of Decoding Strategies for Large Language Modelsaccepted
  683. Localizing Malicious Outputs from CodeLLMaccepted
  684. Locate-then-Merge: Neuron-Level Parameter Fusion for Mitigating Catastrophic Forgetting in Multimodal LLMsaccepted
  685. Lock on Target! Precision Unlearning via Directional Controlaccepted
  686. LogiCoL: Logically-Informed Contrastive Learning for Set-based Dense Retrievalaccepted
  687. LogiDynamics: Unraveling the Dynamics of Inductive, Abductive and Deductive Logical Inferences in LLM Reasoningaccepted
  688. Logic-Thinker: Teaching Large Language Models to Think more Logically.accepted
  689. Logic: Long-form Outline Generation via Imitative and Critical Self-refinementaccepted
  690. LogicTree: Structured Proof Exploration for Coherent and Rigorous Logical Reasoning with Large Language Modelsaccepted
  691. Logical Reasoning with Outcome Reward Models for Test-Time Scalingaccepted
  692. Logit Space Constrained Fine-Tuning for Mitigating Hallucinations in LLM-Based Recommender Systemsaccepted
  693. Logits-Based Finetuningaccepted
  694. Logos as a Well-Tempered Pre-train for Sign Language Recognitionaccepted
  695. Long Chain-of-Thought Fine-tuning via Understanding-to-Reasoning Transitionaccepted
  696. Long-Form Information Alignment Evaluation Beyond Atomic Factsaccepted
  697. Long-context Language Models Fail in Basic Retrieval Tasks Without Sufficient Reasoning Stepsaccepted
  698. LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architectureaccepted
  699. LongTableBench: Benchmarking Long-Context Table Reasoning across Real-World Formats and Domainsaccepted
  700. LongTail-Swap: benchmarking language models’ abilities on rare wordsaccepted
  701. LongWeave: A Long-Form Generation Benchmark Bridging Real-World Relevance and Verifiabilityaccepted
  702. Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Modelsaccepted
  703. Look Beyond Feeling: Unveiling Latent Needs from Implicit Expressions for Proactive Emotional Supportaccepted
  704. Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Queryaccepted
  705. Looking Beyond Text: Reducing Language Bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidanceaccepted
  706. Looking Beyond the Pixels: Evaluating Visual Metaphor Understanding in VLMsaccepted
  707. Lost in Embeddings: Information Loss in Vision–Language Modelsaccepted
  708. Low-Confidence Gold: Refining Low-Confidence Samples for Efficient Instruction Tuningaccepted
  709. Low-Hallucination and Efficient Coreference Resolution with LLMsaccepted
  710. Low-Resource Languages LLM Disinformation is Within Reach: The Case of Walliserdeutschaccepted
  711. LyapLock: Bounded Knowledge Preservation in Sequential Large Language Model Editingaccepted
  712. M-ABSA: A Multilingual Dataset for Aspect-Based Sentiment Analysisaccepted
  713. M-BRe: Discovering Training Samples for Relation Extraction from Unlabeled Texts with Large Language Modelsaccepted
  714. M-Help: Using Social Media Data to Detect Mental Health Help-Seeking Signalsaccepted
  715. M-LongDoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Frameworkaccepted
  716. M-Ped: Multi-Prompt Ensemble Decoding for Large Language Modelsaccepted
  717. M-Wanda: Improving One-Shot Pruning for Multilingual LLMsaccepted
  718. M2Edit: Locate and Edit Multi-Granularity Knowledge in Multimodal Large Language Modelaccepted
  719. M3Retrieve: Benchmarking Multimodal Retrieval for Medicineaccepted
  720. MA-DPR: Manifold-aware Distance Metrics for Dense Passage Retrievalaccepted
  721. MA-GTS: A Multi-Agent Framework for Solving Complex Graph Problems in Real-World Applicationsaccepted
  722. MAC-Tuning: LLM Multi-Compositional Problem Reasoning with Enhanced Knowledge Boundary Awarenessaccepted
  723. MADAWSD: Multi-Agent Debate Framework for Adversarial Word Sense Disambiguationaccepted
  724. MADD: Multi-Agent Drug Discovery Orchestraaccepted
  725. MAFMO: Multi-modal Adaptive Fusion with Meta-template Optimization for Vision-Language Modelsaccepted
  726. MAGIC: A Multi-Hop and Graph-Based Benchmark for Inter-Context Conflicts in Retrieval-Augmented Generationaccepted
  727. MAIN: Mutual Alignment Is Necessary for instruction tuningaccepted
  728. MAKAR: a Multi-Agent framework based Knowledge-Augmented Reasoning for Grounded Multimodal Named Entity Recognitionaccepted
  729. MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMsaccepted
  730. MANTA: A Scalable Pipeline for Transmuting Massive Web Corpora into Instruction Datasetsaccepted
  731. MARIO-0.5B: A Multi-Agent Lightweight Model for Real-Time Open Information Extraction in Low-Resource Settingsaccepted
  732. MARS-Bench: A Multi-turn Athletic Real-world Scenario Benchmark for Dialogue Evaluationaccepted
  733. MASSIVE-Agents: A Benchmark for Multilingual Function-Calling in 52 Languagesaccepted
  734. MASSV: Multimodal Adaptation and Self-Data Distillation for Speculative Decoding of Vision-Language Modelsaccepted
  735. MASTER: Multi-Agent Security Through Exploration of Roles and Topological Structures - A Comprehensive Frameworkaccepted
  736. MATCH: Task-Driven Code Evaluation through Contrastive Learningaccepted
  737. MAVL: A Multilingual Audio-Video Lyrics Dataset for Animated Song Translationaccepted
  738. MAgICoRe: Multi-Agent, Iterative, Coarse-to-Fine Refinement for Reasoningaccepted
  739. MAviS: A Multimodal Conversational Assistant For Avian Speciesaccepted
  740. MC2: A Minimum-Coverage and Dataset-Agnostic Framework for Compositional Generalization of LLMs on Semantic Parsingaccepted
  741. MCIP: Protecting MCP Safety via Model Contextual Integrity Protocolaccepted
  742. MCTS-RAG: Enhancing Retrieval-Augmented Generation with Monte Carlo Tree Searchaccepted
  743. MCiteBench: A Multimodal Benchmark for Generating Text with Citationsaccepted
  744. MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarizationaccepted
  745. MEBench: Benchmarking Large Language Models for Cross-Document Multi-Entity Question Answeringaccepted
  746. MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapperaccepted
  747. MERMAID: Multi-perspective Self-reflective Agents with Generative Augmentation for Emotion Recognitionaccepted
  748. METok: Multi-Stage Event-based Token Compression for Efficient Long Video Understandingaccepted
  749. MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregationaccepted
  750. MFTCXplain: A Multilingual Benchmark Dataset for Evaluating the Moral Reasoning of LLMs through Multi-hop Hate Speech Explanationaccepted
  751. MIND: Towards Immersive Psychological Healing with Multi-Agent Inner Dialogueaccepted
  752. MIO: A Foundation Model on Multimodal Tokensaccepted
  753. MIRROR: Multimodal Cognitive Reframing Therapy for Rolling with Resistanceaccepted
  754. ML-Promise: A Multilingual Dataset for Corporate Promise Verificationaccepted
  755. MLAlgo-Bench: Can Machines Implement Machine Learning Algorithms?accepted
  756. MLWQ: Efficient Small Language Model Deployment via Multi-Level Weight Quantizationaccepted
  757. MM-CRITIC: A Holistic Evaluation of Large Multimodal Models as Multimodal Critiqueaccepted
  758. MMA: Cross-Domain Knowledge Integration via Mixture of Multi-Domain Agentsaccepted
  759. MMAG: Multimodal Learning for Mucus Anomaly Grading in Nasal Endoscopy via Semantic Attribute Promptingaccepted
  760. MMAPG: A Training-Free Framework for Multimodal Multi-hop Question Answering via Adaptive Planning Graphsaccepted
  761. MMDocIR: Benchmarking Multimodal Retrieval for Long Documentsaccepted
  762. MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluationaccepted
  763. MMPlanner: Zero-Shot Multimodal Procedural Planning with Chain-of-Thought Object State Reasoningaccepted
  764. MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts?accepted
  765. MOLE: Metadata Extraction and Validation in Scientific Papers Using LLMsaccepted
  766. MONAQ: Multi-Objective Neural Architecture Querying for Time-Series Analysis on Resource-Constrained Devicesaccepted
  767. MOSAIC: Modeling Social AI for Content Dissemination and Regulation in Multi-Agent Simulationsaccepted
  768. MPCG: Multi-Round Persona-Conditioned Generation for Modeling the Evolution of Misinformation with LLMsaccepted
  769. MPO: Boosting LLM Agents with Meta Plan Optimizationaccepted
  770. MPRF: Interpretable Stance Detection through Multi-Path Reasoning Frameworkaccepted
  771. MPTA: MultiTask Personalization Assessmentaccepted
  772. MR. Judge: Multimodal Reasoner as a Judgeaccepted
  773. MRAG: A Modular Retrieval Framework for Time-Sensitive Question Answeringaccepted
  774. MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMsaccepted
  775. MS-RAG: Simple and Effective Multi-Semantic Retrieval-Augmented Generationaccepted
  776. MT-Mol: Multi Agent System with Tool-based Reasoning for Molecular Optimizationaccepted
  777. MT-R1-Zero: Advancing LLM-based Machine Translation via R1-Zero-like Reinforcement Learningaccepted
  778. MT-RewardTree: A Comprehensive Framework for Advancing LLM-Based Machine Translation via Reward Modelingaccepted
  779. MTabVQA: Evaluating Multi-Tabular Reasoning of Language Models in Visual Spaceaccepted
  780. MUCAR: Benchmarking Multilingual Cross-Modal Ambiguity Resolution for Multimodal Large Language Modelsaccepted
  781. MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Languageaccepted
  782. MULTIGUARD: An Efficient Approach for AI Safety Moderation Across Languages and Modalitiesaccepted
  783. MULTITAT: Benchmarking Multilingual Table-and-Text Question Answeringaccepted
  784. MULTIVOX: A Benchmark for Evaluating Voice Assistants for Multimodal Interactionsaccepted
  785. MUSE: MCTS-Driven Red Teaming Framework for Enhanced Multi-Turn Dialogue Safety in Large Language Modelsaccepted
  786. MUZO: Leveraging Multiple Queries and Momentum for Zeroth-Order Fine-Tuning of Large Language Modelsaccepted
  787. MV-CLAM: Multi-View Molecular Interpretation with Cross-Modal Projection via Language Modelaccepted
  788. MaGiX: A Multi-Granular Adaptive Graph Intelligence Framework for Enhancing Cross-Lingual RAGaccepted
  789. MaZO: Masked Zeroth-Order Optimization for Multi-Task Fine-Tuning of Large Language Modelsaccepted
  790. Machine-generated text detection prevents language model collapseaccepted
  791. Mahānāma: A Unique Testbed for Literary Entity Discovery and Linkingaccepted
  792. Make Every Letter Count: Building Dialect Variation Dictionaries from Monolingual Corporaaccepted
  793. Making Every Step Effective: Jailbreaking Large Vision-Language Models Through Hierarchical KV Equalizationaccepted
  794. Making VLMs More Robot-Friendly: Self-Critical Distillation of Low-Level Procedural Reasoningaccepted
  795. Mamba Drafters for Speculative Decodingaccepted
  796. ManuSearch: Democratizing Deep Search in Large Language Models with a Transparent and Open Multi-Agent Frameworkaccepted
  797. Mapping Toxic Comments Across Demographics: A Dataset from German Public Broadcastingaccepted
  798. Mapping semantic networks to Dutch word embeddings as a diagnostic tool for cognitive declineaccepted
  799. Mapping the Minds of LLMs: A Graph-Based Analysis of Reasoning LLMsaccepted
  800. MarathiEmoExplain: A Dataset for Sentiment, Emotion, and Explanation in Low-Resource Marathiaccepted
  801. MaskCD: Mitigating LVLM Hallucinations by Image Head Masked Contrastive Decodingaccepted
  802. Masked Diffusion Captioning for Visual Feature Learningaccepted
  803. Massive Supervised Fine-tuning Experiments Reveal How Data, Layer, and Training Factors Shape LLM Alignment Qualityaccepted
  804. MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutorsaccepted
  805. Matter-of-Fact: A Benchmark for Verifying the Feasibility of Literature-Supported Claims in Materials Scienceaccepted
  806. Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biasesaccepted
  807. Measuring Chain of Thought Faithfulness by Unlearning Reasoning Stepsaccepted
  808. Measuring Lexical Diversity of Synthetic Data Generated through Fine-Grained Persona Promptingaccepted
  809. Measuring Risk of Bias in Biomedical Reports: The RoBBR Benchmarkaccepted
  810. Measuring Sycophancy of Language Models in Multi-turn Dialoguesaccepted
  811. Measuring and Mitigating Media Outlet Name Bias in Large Language Modelsaccepted
  812. Measuring scalar constructs in social science with LLMsaccepted
  813. Measuring the Effect of Disfluency in Multilingual Knowledge Probing Benchmarksaccepted
  814. Mechanisms vs. Outcomes: Probing for Syntax Fails to Explain Performance on Targeted Syntactic Evaluationsaccepted
  815. Mechanistic Understanding and Mitigation of Language Confusion in English-Centric Large Language Modelsaccepted
  816. Med-PRM: Medical Reasoning Models with Stepwise, Guideline-verified Process Rewardsaccepted
  817. Med-VRAgent: A Framework for Medical Visual Reasoning-Enhanced Agentsaccepted
  818. MedCOD: Enhancing English-to-Spanish Medical Translation of Large Language Models Using Enriched Chain-of-Dictionary Frameworkaccepted
  819. MedEBench: Diagnosing Reliability in Text-Guided Medical Image Editingaccepted
  820. MedFact: A Large-scale Chinese Dataset for Evidence-based Medical Fact-checking of LLM Responsesaccepted
  821. MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Modelsaccepted
  822. MedLinkDE – MedDRA Entity Linking for German with Guided Chain of Thought Reasoningaccepted
  823. MediVLM: A Vision Language Model for Radiology Report Generation from Medical Imagesaccepted
  824. Media Source Matters More Than Content: Unveiling Political Bias in LLM-Generated Citationsaccepted
  825. MemInsight: Autonomous Memory Augmentation for LLM Agentsaccepted
  826. Membership and Memorization in LLM Knowledge Distillationaccepted
  827. MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Modelsaccepted
  828. MemeIntel: Explainable Detection of Propagandistic and Hateful Memesaccepted
  829. MemeInterpret: Towards an All-in-One Dataset for Meme Understandingaccepted
  830. MemeReaCon: Probing Contextual Meme Understanding in Large Vision-Language Modelsaccepted
  831. Memorization or Reasoning? Exploring the Idiom Understanding of LLMsaccepted
  832. Memorization ≠ Understanding: Do Large Language Models Have the Ability of Scenario Cognition?accepted
  833. Memory OS of AI Agentaccepted
  834. Memory-QA: Answering Recall Questions Based on Multimodal Memoriesaccepted
  835. Memory-enhanced Large Language Model for Cross-lingual Dependency Parsing via Deep Hierarchical Syntax Understandingaccepted
  836. MentalGLM Series: Explainable Large Language Models for Mental Health Analysis on Chinese Social Mediaaccepted
  837. Merge then Realign: Simple and Effective Modality-Incremental Continual Learning for Multimodal LLMsaccepted
  838. Merger-as-a-Stealer: Stealing Targeted PII from Aligned LLMs with Model Mergingaccepted
  839. MessIRve: A Large-Scale Spanish Information Retrieval Datasetaccepted
  840. Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judgeaccepted
  841. Meta-Semantics Augmented Few-Shot Relational Learningaccepted
  842. MetaFaith: Faithful Natural Language Uncertainty Expression in LLMsaccepted
  843. MetaLadder: Ascending Mathematical Solution Quality via Analogical-Problem Reasoning Transferaccepted
  844. MetaMixSpeech: Meta Task Augmentation for Low-Resource Speech Recognitionaccepted
  845. Metric Calculating Benchmark: Code-Verifiable Complicate Instruction Following Benchmark for Large Language Modelsaccepted
  846. MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learningaccepted
  847. MiLQ: Benchmarking IR Models for Bilingual Web Search with Mixed Language Queriesaccepted
  848. MicroEdit: Neuron-level Knowledge Disentanglement and Localization in Lifelong Model Editingaccepted
  849. MidPO: Dual Preference Optimization for Safety and Helpfulness in Large Language Models via a Mixture of Experts Frameworkaccepted
  850. Middo: Model-Informed Dynamic Data Optimization for Enhanced LLM Fine-Tuning via Closed-Loop Learningaccepted
  851. Mind the Blind Spots: A Focus-Level Evaluation Framework for LLM Reviewsaccepted
  852. Mind the Dialect: NLP Advancements Uncover Fairness Disparities for Arabic Users in Recommendation Systemsaccepted
  853. Mind the Gap: A Closer Look at Tokenization for Multiple-Choice Question Answering with LLMsaccepted
  854. Mind the Gap: How BabyLMs Learn Filler-Gap Dependenciesaccepted
  855. Mind the Inclusivity Gap: Multilingual Gender-Neutral Translation Evaluation with mGeNTEaccepted
  856. Mind the Style Gap: Meta-Evaluation of Style and Attribute Transfer Metricsaccepted
  857. Mind the Value-Action Gap: Do LLMs Act in Alignment with Their Values?accepted
  858. Minimal Ranks, Maximum Confidence: Parameter-efficient Uncertainty Quantification for LoRAaccepted
  859. Minimal, Local, and Robust: Embedding-Only Edits for Implicit Bias in T2I Modelsaccepted
  860. Mining the Past with Dual Criteria: Integrating Three types of Historical Information for Context-aware Event Forecastingaccepted
  861. Misalignment Attack on Text-to-Image Models via Text Embedding Optimization and Inversionaccepted
  862. MisinfoBench: A Multi-Dimensional Benchmark for Evaluating LLMs’ Resilience to Misinformationaccepted
  863. Mitigate One, Skew Another? Tackling Intersectional Biases in Text-to-Image Modelsaccepted
  864. Mitigating Attention Localization in Small Scale: Self-Attention Refinement via One-step Belief Propagationaccepted
  865. Mitigating Biases in Language Models via Bias Unlearningaccepted
  866. Mitigating Catastrophic Forgetting in Large Language Models with Forgetting-aware Pruningaccepted
  867. Mitigating Gender Bias via Fostering Exploratory Thinking in LLMsaccepted
  868. Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality Aligningaccepted
  869. Mitigating Hallucination in Large Vision-Language Models through Aligning Attention Distribution to Information Flowaccepted
  870. Mitigating Hallucinations in LM-Based TTS Models via Distribution Alignment Using GFlowNetsaccepted
  871. Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinationsaccepted
  872. Mitigating Hallucinations in Large Vision-Language Models via Entity-Centric Multimodal Preference Optimizationaccepted
  873. Mitigating Hallucinations in Vision-Language Models through Image-Guided Head Suppressionaccepted
  874. Mitigating Interviewer Bias in Multimodal Depression Detection: An Approach with Adversarial Learning and Contextual Positional Encodingaccepted
  875. Mitigating Object Hallucinations in MLLMs via Multi-Frequency Perturbationsaccepted
  876. Mitigating Sequential Dependencies: A Survey of Algorithms and Systems for Generation-Refinement Frameworks in Autoregressive Modelsaccepted
  877. Mitigating Spurious Correlations via Counterfactual Contrastive Learningaccepted
  878. Mitigating Visual Knowledge Forgetting in MLLM Instruction-tuning via Modality-decoupled Gradient Descentaccepted
  879. Mitigating the Privacy Issues in Retrieval-Augmented Generation (RAG) via Pure Synthetic Dataaccepted
  880. MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corporaaccepted
  881. Mixed Signals: Decoding VLMs’ Reasoning and Underlying Bias in Vision-Language Conflictaccepted
  882. Mixing Inference-time Experts for Enhancing LLM Reasoningaccepted
  883. Mixture of Languages: Improved Multilingual Encoders Through Language Groupingaccepted
  884. Mixture of Length and Pruning Experts for Knowledge Graphs Reasoningaccepted
  885. Mixture of LoRA Experts for Continual Information Extraction with LLMsaccepted
  886. Mixture of Weight-shared Heterogeneous Group Attention Experts for Dynamic Token-wise KV Optimizationaccepted
  887. Mixture-of-Clustered-Experts: Advancing Expert Specialization and Generalization in Instruction Tuningaccepted
  888. MoLoRAG: Bootstrapping Document Understanding via Multi-modal Logic-aware Retrievalaccepted
  889. MoMentS: A Comprehensive Multimodal Benchmark for Theory of Mindaccepted
  890. MoMoE: Mixture of Moderation Experts Framework for AI-Assisted Online Governanceaccepted
  891. MoR: Better Handling Diverse Queries with a Mixture of Sparse, Dense, and Human Retrieversaccepted
  892. MoRoVoc: A Large Dataset for Geographical Variation Identification of the Spoken Romanian Languageaccepted
  893. MoSEs: Uncertainty-Aware AI-Generated Text Detection via Mixture of Stylistics Experts with Conditional Thresholdsaccepted
  894. MoVa: Towards Generalizable Classification of Human Morals and Valuesaccepted
  895. MoVoC: Morphology-Aware Subword Construction for Ge’ez Script Languagesaccepted
  896. MobiZO: Enabling Efficient LLM Fine-Tuning at the Edge via Inference Enginesaccepted
  897. ModRWKV: Transformer Multimodality in Linear Timeaccepted
  898. ModalPrompt: Towards Efficient Multimodal Continual Instruction Tuning with Dual-Modality Guided Promptaccepted
  899. Model Calibration for Emotion Detectionaccepted
  900. Model Consistency as a Cheap yet Predictive Proxy for LLM Elo Scoresaccepted
  901. Model Unlearning via Sparse Autoencoder Subspace Guided Projectionsaccepted
  902. Model-Based Ranking of Source Languages for Zero-Shot Cross-Lingual Transferaccepted
  903. Model-based Large Language Model Customization as Serviceaccepted
  904. ModelCitizens: Representing Community Voices in Online Safetyaccepted
  905. Modeling Bottom-up Information Quality during Language Processingaccepted
  906. Modeling Subjectivity in Cognitive Appraisal with Language Modelsaccepted
  907. Modeling, Evaluating, and Embodying Personality in LLMs: A Surveyaccepted
  908. ModelingAgent: Bridging LLMs and Mathematical Modeling for Real-World Challengesaccepted
  909. MolErr2Fix: Benchmarking LLM Trustworthiness in Chemistry via Modular Error Detection, Localization, Explanation, and Correctionaccepted
  910. Molecular String Representation Preferences in Pretrained LLMs: A Comparative Study in Zero- & Few-Shot Molecular Property Predictionaccepted
  911. Mondrian: A Framework for Logical Abstract (Re)Structuringaccepted
  912. Morables: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fablesaccepted
  913. Moral Framing in Politics (MFiP): A new resource and models for moral framingaccepted
  914. More Documents, Same Length: Isolating the Challenge of Multiple Documents in RAGaccepted
  915. More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compressionaccepted
  916. Morpheme Induction for Emergent Languageaccepted
  917. MotivGraph-SoIQ: Integrating Motivational Knowledge Graphs and Socratic Dialogue for Enhanced LLM Ideationaccepted
  918. MovieCORE: COgnitive REasoning in Moviesaccepted
  919. MrGuard: A Multilingual Reasoning Guardrail for Universal LLM Safetyaccepted
  920. MuCAL: Contrastive Alignment for Preference-Driven KG-to-Text Generationaccepted
  921. MuTIS: Enhancing Reasoning Efficiency through Multi Turn Intervention Sampling in Reinforcement Learningaccepted
  922. Multi-Agent Autonomous Driving Systems with Large Language Models: A Survey of Recent Advances, Resources, and Future Directionsaccepted
  923. Multi-Document Event Extraction Using Large and Small Language Modelsaccepted
  924. Multi-Domain Explainability of Preferencesaccepted
  925. Multi-Frequency Contrastive Decoding: Alleviating Hallucinations for Large Vision-Language Modelsaccepted
  926. Multi-LMentry: Can Multilingual LLMs Solve Elementary Tasks Across Languages?accepted
  927. Multi-Modal Framing Analysis of Newsaccepted
  928. Multi-Surrogate-Objective Optimization for Neural Topic Modelsaccepted
  929. Multi-level Diagnosis and Evaluation for Robust Tabular Feature Engineering with Large Language Modelsaccepted
  930. Multi-perspective Analysis of Large Language Model Domain Specialization: An Experiment in Accounting Audit Procedures Generationaccepted
  931. Multi-token Mask-filling and Implicit Discourse Relationsaccepted
  932. Multi-view-guided Passage Reranking with Large Language Modelsaccepted
  933. MultiAgentESC: A LLM-based Multi-Agent Collaboration Framework for Emotional Support Conversationaccepted
  934. MultiClaimNet: A Massively Multilingual Dataset of Fact-Checked Claim Clustersaccepted
  935. MultiConIR: Towards Multi-Condition Information Retrievalaccepted
  936. MultiDocFusion : Hierarchical and Multimodal Chunking Pipeline for Enhanced RAG on Long Industrial Documentsaccepted
  937. MultiLingPoT: Boosting Mathematical Reasoning in LLMs through Multilingual Program Integrationaccepted
  938. MultiLogicNMR(er): A Benchmark and Neural-Symbolic Framework for Non-monotonic Reasoning with Multiple Extensionsaccepted
  939. MultiMatch: Multihead Consistency Regularization Matching for Semi-Supervised Text Classificationaccepted
  940. MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translationaccepted
  941. MultiPL-MoE: Multi-Programming-Lingual Extension of Large Language Models through Hybrid Mixture-of-Expertsaccepted
  942. Multifaceted Evaluation of Audio-Visual Capability for MLLMs: Effectiveness, Efficiency, Generalizability and Robustnessaccepted
  943. Multilingual Collaborative Defense for Large Language Modelsaccepted
  944. Multilingual Data Filtering using Synthetic Data from Large Language Modelsaccepted
  945. Multilingual Datasets for Custom Input Extraction and Explanation Requests Parsing in Conversational XAI Systemsaccepted
  946. Multilingual Dialogue Generation and Localization with Dialogue Act Scriptingaccepted
  947. Multilingual Federated Low-Rank Adaptation for Collaborative Content Anomaly Detection across Multilingual Social Media Participantsaccepted
  948. Multilingual Generative Retrieval via Cross-lingual Semantic Compressionaccepted
  949. Multilingual Knowledge Graph Completion via Efficient Multilingual Knowledge Sharingaccepted
  950. Multilingual Language Model Pretraining using Machine-translated Dataaccepted
  951. Multilingual Pretraining for Pixel Language Modelsaccepted
  952. Multilingual Prompting for Improving LLM Generation Diversityaccepted
  953. Multilingual Verbalisation of Knowledge Graphsaccepted
  954. Multilingual vs Crosslingual Retrieval of Fact-Checked Claims: A Tale of Two Approachesaccepted
  955. Multilinguality Does not Make Sense: Investigating Factors Behind Zero-Shot Cross-Lingual Transfer in Sense-Aware Tasksaccepted
  956. Multimedia Event Extraction with LLM Knowledge Editingaccepted
  957. Multimodal Document-level Triple Extraction via Dynamic Graph Enhancement and Relation-Aware Reflectionaccepted
  958. Multimodal Emotion Recognition in Conversations: A Survey of Methods, Trends, Challenges and Prospectsaccepted
  959. Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesisaccepted
  960. Multimodal Language Models See Better When They Look Shalloweraccepted
  961. Multimodal Neural Machine Translation: A Survey of the State of the Artaccepted
  962. Multimodal UNcommonsense: From Odd to Ordinary and Ordinary to Oddaccepted
  963. MusKGC: A Flexible Multi-source Knowledge Enhancement Framework for Open-World Knowledge Graph Completionaccepted
  964. MuseScorer: Idea Originality Scoring At Scaleaccepted
  965. MythTriage: Scalable Detection of Opioid Use Disorder Myths on a Video-Sharing Platformaccepted
  966. N-CORE: N-View Consistency Regularization for Disentangled Representation Learning in Nonverbal Vocalizationsaccepted
  967. NAP2: A Benchmark for Naturalness and Privacy-Preserving Text Rewriting by Learning from Humanaccepted
  968. NER Retriever: Zero-Shot Named Entity Retrieval with Type-Aware Embeddingsaccepted
  969. NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Callsaccepted
  970. NEXUS: Network Exploration for eXploiting Unsafe Sequences in Multi-Turn LLM Jailbreaksaccepted
  971. NILE: Internal Consistency Alignment in Large Language Modelsaccepted
  972. NIM: Neuro-symbolic Ideographic Metalanguage for Inclusive Communicationaccepted
  973. NL-Debugging: Exploiting Natural Language as an Intermediate Representation for Code Debuggingaccepted
  974. NL2Lean: Translating Natural Language into Lean 4 through Multi-Aspect Reinforcement Learningaccepted
  975. NLKI: A Lightweight Natural Language Knowledge Integration Framework for Improving Small VLMs in Commonsense VQA Tasksaccepted
  976. NLP Needs Diversity outside of ‘Diversity’accepted
  977. NLP-ADBench: NLP Anomaly Detection Benchmarkaccepted
  978. NLoRA: Nyström-Initiated Low-Rank Adaptation for Large Language Modelsaccepted
  979. NOVA-63: Native Omni-lingual Versatile Assessments of 63 Disciplinesaccepted
  980. NOVER: Incentive Training for Language Models via Verifier-Free Reinforcement Learningaccepted
  981. NUMINA: A Natural Understanding Benchmark for Multi-dimensional Intelligence and Numerical Reasoning Abilitiesaccepted
  982. NUTMEG: Separating Signal From Noise in Annotator Disagreementaccepted
  983. NarratEX Dataset: Explaining the Dominant Narratives in News Textsaccepted
  984. Natural Context Drift Undermines the Natural Language Understanding of Large Language Modelsaccepted
  985. Navigating the Unknown: Intent Classification and Out-of-Distribution Detection Using Large Language Modelsaccepted
  986. NeLLCom-Lex: A Neural-agent Framework to Study the Interplay between Lexical Systems and Language Useaccepted
  987. NeighXLM: Enhancing Cross-Lingual Transfer in Low-Resource Languages via Neighbor-Augmented Contrastive Pretrainingaccepted
  988. Nested Named Entity Recognition as Single-Pass Sequence Labelingaccepted
  989. Neural Topic Modeling via Contextual and Graph Information Fusionaccepted
  990. NeuroAda: Activating Each Neuron’s Potential for Parameter-Efficient Fine-Tuningaccepted
  991. Neuron-Level Differentiation of Memorization and Generalization in Large Language Modelsaccepted
  992. Neutral Is Not Unbiased: Evaluating Implicit and Intersectional Identity Bias in LLMs Through Structured Narrative Scenariosaccepted
  993. Nexus: Adaptive Upcycling to Efficiently Pretrain Mixture of Expertsaccepted
  994. NileChat: Towards Linguistically Diverse and Culturally Aware LLMs for Local Communitiesaccepted
  995. Nine Ways to Break Copyright Law and Why Our LLM Won’t: A Fair Use Aligned Generation Frameworkaccepted
  996. NitiBench: Benchmarking LLM Frameworks on Thai Legal Question Answering Capabilitiesaccepted
  997. No Black Boxes: Interpretable and Interactable Predictive Healthcare with Knowledge-Enhanced Agentic Causal Discoveryaccepted
  998. No Free Lunch: Retrieval-Augmented Generation Undermines Fairness in LLMs, Even for Vigilant Usersaccepted
  999. No Need for Explanations: LLMs can implicitly learn from mistakes in-contextaccepted
  1000. Noise, Adaptation, and Strategy: Assessing LLM Fidelity in Decision-Makingaccepted

Looking for submission deadlines instead? See the conference deadline calendar.