← All conferences

EMNLP 2025 Accepted Papers

The full list of 3,211 papers accepted at EMNLP 2025 (Conference on Empirical Methods in Natural Language Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

accepted: 3,211
  1. (Almost) Free Modality Stitching of Foundation Modelsaccepted
  2. 1+1>2: A Synergistic Sparse and Low-Rank Compression Method for Large Language Modelsaccepted
  3. 2Columns1Row: A Russian Benchmark for Textual and Multimodal Table Understanding and Reasoningaccepted
  4. 3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillationaccepted
  5. 3DS: Medical Domain Adaptation of LLMs via Decomposed Difficulty-based Data Selectionaccepted
  6. 3MDBench: Medical Multimodal Multi-agent Dialogue Benchmarkaccepted
  7. 3R: Enhancing Sentence Representation Learning via Redundant Representation Reductionaccepted
  8. A Benchmark for Hindi Verb-Argument Structure Alternationsaccepted
  9. A Benchmark for Translations Across Styles and Language Variantsaccepted
  10. A Case Against Implicit Standards: Homophone Normalization in Machine Translation for Languages that use the Ge’ez Script.accepted
  11. A Category-Theoretic Approach to Neural-Symbolic Task Planning with Bidirectional Searchaccepted
  12. A Causal Lens for Evaluating Faithfulness Metricsaccepted
  13. A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Modelsaccepted
  14. A Comparison of Independent and Joint Fine-tuning Strategies for Retrieval-Augmented Generationaccepted
  15. A Comprehensive Framework to Operationalize Social Stereotypes for Responsible AI Evaluationsaccepted
  16. A Comprehensive Literary Chinese Reading Comprehension Dataset with an Evidence Curation Based Solutionaccepted
  17. A Comprehensive Survey on Learning from Rewards for Large Language Models: Reward Models and Learning Strategiesaccepted
  18. A Comprehensive Survey on the Trustworthiness of Large Language Models in Healthcareaccepted
  19. A Comprehensive Taxonomy of Negation for NLP and Neural Retrieversaccepted
  20. A Computational Simulation of Language Production in First Language Acquisitionaccepted
  21. A Culturally-diverse Multilingual Multimodal Video Benchmark & Modelaccepted
  22. A Decoupled Multi-Agent Framework for Complex Text Style Transferaccepted
  23. A Dynamic Fusion Model for Consistent Crisis Responseaccepted
  24. A Fully Probabilistic Perspective on Large Language Model Unlearning: Evaluation and Optimizationaccepted
  25. A Generalizable Rhetorical Strategy Annotation Model Using LLM-based Debate Simulation and Labellingaccepted
  26. A Generative Framework for Personalized Sticker Retrievalaccepted
  27. A Generative Pre-Trained Language Model for Channel Prediction in Wireless Communications Systemsaccepted
  28. A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Usersaccepted
  29. A Graph-Theoretical Framework for Analyzing the Behavior of Causal Language Modelsaccepted
  30. A Group Fairness Lens for Large Language Modelsaccepted
  31. A Head to Predict and a Head to Question: Pre-trained Uncertainty Quantification Heads for Hallucination Detection in LLM Outputsaccepted
  32. A Knapsack by Any Other Name: Presentation impacts LLM performance on NP-hard problemsaccepted
  33. A Knowledge-driven Adaptive Collaboration of LLMs for Enhancing Medical Decision-makingaccepted
  34. A Middle Path for On-Premises LLM Deployment: Preserving Privacy Without Sacrificing Model Confidentialityaccepted
  35. A Monte-Carlo Sampling Framework For Reliable Evaluation of Large Language Models Using Behavioral Analysisaccepted
  36. A Multi-Agent Framework with Automated Decision Rule Optimization for Cross-Domain Misinformation Detectionaccepted
  37. A Multi-Level Benchmark for Causal Language Understanding in Social Media Discourseaccepted
  38. A Multilingual, Culture-First Approach to Addressing Misgendering in LLM Applicationsaccepted
  39. A Necessary Step toward Faithfulness: Measuring and Improving Consistency in Free-Text Explanationsaccepted
  40. A Position Paper on the Automatic Generation of Machine Learning Leaderboardsaccepted
  41. A Probabilistic Inference Scaling Theory for LLM Self-Correctionaccepted
  42. A Rigorous Evaluation of LLM Data Generation Strategies for Low-Resource Languagesaccepted
  43. A Sequential Multi-Stage Approach for Code Vulnerability Detection via Confidence- and Collaboration-based Decision Makingaccepted
  44. A Similarity Measure for Comparing Conversational Dynamicsaccepted
  45. A Simple Yet Effective Method for Non-Refusing Context Relevant Fine-grained Safety Steering in LLMsaccepted
  46. A Structured Framework for Evaluating and Enhancing Interpretive Capabilities of Multimodal LLMs in Culturally Situated Tasksaccepted
  47. A Survey of Cognitive Distortion Detection and Classification in NLPaccepted
  48. A Survey of Link Prediction in N-ary Knowledge Graphsaccepted
  49. A Survey of Multilingual Reasoning in Language Modelsaccepted
  50. A Survey of Pun Generation: Datasets, Evaluations and Methodologiesaccepted
  51. A Survey of RAG-Reasoning Systems in Large Language Modelsaccepted
  52. A Survey on LLM-powered Agents for Recommender Systemsaccepted
  53. A Survey on LLMs for Story Generationaccepted
  54. A Survey on Multi-modal Intent Recognition: Recent Advances and New Frontiersaccepted
  55. A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Modelsaccepted
  56. A Survey on Training-free Alignment of Large Language Modelsaccepted
  57. A Symbolic Adversarial Learning Framework for Evolving Fake News Generation and Detectionaccepted
  58. A Systematic Analysis of Base Model Choice for Reward Modelingaccepted
  59. A Systematic Survey of Automatic Prompt Optimization Techniquesaccepted
  60. A Systematic Survey of Claim Verification: Corpora, Systems, and Case Studiesaccepted
  61. A Text-Based Recommender System that Leverages Explicit Affective State Preferencesaccepted
  62. A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolationaccepted
  63. A Unified Framework for N-ary Property Information Extraction in Materials Scienceaccepted
  64. A Zero-Shot Neuro-Symbolic Approach for Complex Knowledge Graph Question Answeringaccepted
  65. ACEBench: A Comprehensive Evaluation of LLM Tool Usageaccepted
  66. ACING: Actor-Critic for Instruction Learning in Black-Box LLMsaccepted
  67. AELC: Adaptive Entity Linking with LLM-Driven Contextualizationaccepted
  68. AFRIDOC-MT: Document-level MT Corpus for African Languagesaccepted
  69. AGENTVIGIL: Automatic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agentsaccepted
  70. AI Argues Differently: Distinct Argumentative and Linguistic Patterns of LLMs in Persuasive Contextsaccepted
  71. AI Chatbots as Professional Service Agents: Developing a Professional Identityaccepted
  72. AI Knows Where You Are: Exposure, Bias, and Inference in Multimodal Geolocation with KoreaGEOaccepted
  73. AI Sees Your Location—But With A Bias Toward The Wealthy Worldaccepted
  74. AIMMerging: Adaptive Iterative Model Merging Using Training Trajectories for Language Model Continual Learningaccepted
  75. AIP: Subverting Retrieval-Augmented Generation via Adversarial Instructional Promptaccepted
  76. AIR: Complex Instruction Generation via Automatic Iterative Refinementaccepted
  77. AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Scienceaccepted
  78. ALLabel: Three-stage Active Learning for LLM-based Entity Recognition using Demonstration Retrievalaccepted
  79. ALRPHFS: Adversarially Learned Risk Patterns with Hierarchical Fast & Slow Reasoning for Robust Agent Defenseaccepted
  80. AMACE: Automatic Multi-Agent Chart Evolution for Iteratively Tailored Chart Generationaccepted
  81. AMANDA: Agentic Medical Knowledge Augmentation for Data-Efficient Medical Visual Question Answeringaccepted
  82. AMIA: Automatic Masking and Joint Intention Analysis Makes LVLMs Robust Jailbreak Defendersaccepted
  83. AMQ: Enabling AutoML for Mixed-precision Weight-Only Quantization of Large Language Modelsaccepted
  84. APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transportaccepted
  85. AQuilt: Weaving Logic and Self-Inspection into Low-Cost, High-Relevance Data Synthesis for Specialist LLMsaccepted
  86. AROMA: Autonomous Rank-one Matrix Adaptationaccepted
  87. ARXSA: A General Negative Feedback Control Theory in Vision-Language Modelsaccepted
  88. ASD-iLLM:An Intervention Large Language Model for Autistic Children based on Real Clinical Dialogue Intervention Datasetaccepted
  89. ASTPrompter: Preference-Aligned Automated Language Model Red-Teaming to Generate Low-Perplexity Unsafe Promptsaccepted
  90. ASTRA: A Negotiation Agent with Adaptive and Strategic Reasoning via Tool-integrated Action for Dynamic Offer Optimizationaccepted
  91. AbsVis – Benchmarking How Humans and Vision-Language Models “See” Abstract Concepts in Imagesaccepted
  92. AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Modelsaccepted
  93. Accelerate Parallelizable Reasoning via Parallel Decoding within One Sequenceaccepted
  94. Accelerated Test-Time Scaling with Model-Free Speculative Samplingaccepted
  95. Accelerating LLM Reasoning via Early Rejection with Partial Reward Modelingaccepted
  96. Accept or Deny? Evaluating LLM Fairness and Performance in Loan Approval across Table-to-Text Serialization Approachesaccepted
  97. AccessEval: Benchmarking Disability Bias in Large Language Modelsaccepted
  98. Acquiescence Bias in Large Language Modelsaccepted
  99. ActionStudio: A Lightweight Framework for Data and Training of Large Action Modelsaccepted
  100. Active Domain Knowledge Acquisition with 100-Dollar Budget: Enhancing LLMs via Cost-Efficient, Expert-Involved Interaction in Sensitive Domainsaccepted
  101. Active Layer-Contrastive Decoding Reduces Hallucination in Large Language Model Generationaccepted
  102. Active Learning for Multidialectal Arabic POS Taggingaccepted
  103. AdDriftBench: A Benchmark for Detecting Data Drift and Label Drift in Short Video Advertisingaccepted
  104. AdaRewriter: Unleashing the Power of Prompting-based Conversational Query Reformulation via Test-Time Adaptationaccepted
  105. AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defenderaccepted
  106. AdaTP: Attention-Debiased Token Pruning for Video Large Language Modelsaccepted
  107. AdamS: Momentum Itself Can Be A Normalizer for LLM Pretraining and Post-trainingaccepted
  108. AdaptFlow: Adaptive Workflow Optimization via Meta-Learningaccepted
  109. AdaptMerge: Inference Time Adaptive Visual and Language-Guided Token Merging for Efficient Large Multimodal Modelsaccepted
  110. AdaptThink: Reasoning Models Can Learn When to Thinkaccepted
  111. Adapting Bias Evaluation to Domain Contexts using Generative Modelsaccepted
  112. Adapting Large Language Models for Character-based Augmentative and Alternative Communicationaccepted
  113. Adaptive LLM Routing under Budget Constraintsaccepted
  114. Adaptive Platt Scaling with Causal Interpretations for Self-Reflective Language Model Uncertainty Estimatesaccepted
  115. Adaptive Preference Optimization with Uncertainty-aware Utility Anchoraccepted
  116. Adaptive Schema-aware Event Extraction with Retrieval-Augmented Generationaccepted
  117. Adaptively profiling models with task elicitationaccepted
  118. Add-One-In: Incremental Sample Selection for Large Language Models via a Choice-Based Greedy Paradigmaccepted
  119. Addition in Four Movements: Mapping Layer-wise Information Trajectories in LLMsaccepted
  120. Addressing Tokenization Inconsistency in Steganography and Watermarking Based on Large Language Modelsaccepted
  121. Advancing Arabic Diacritization: Improved Datasets, Benchmarking, and State-of-the-Art Modelsaccepted
  122. Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Modelsaccepted
  123. Advancing Oversight Reasoning across Languages for Audit Sycophantic Behaviour via X-Agentaccepted
  124. Advancing Reasoning with Off-the-Shelf LLMs: A Semantic Structure Perspectiveaccepted
  125. Adversarial Attacks Against Automated Fact-Checking: A Surveyaccepted
  126. Adversary-Aware DPO: Enhancing Safety Alignment in Vision Language Models via Adversarial Trainingaccepted
  127. AesBiasBench: Evaluating Bias and Alignment in Multimodal Language Models for Personalized Image Aesthetic Assessmentaccepted
  128. Africa Health Check: Probing Cultural Bias in Medical LLMsaccepted
  129. AfroXLMR-Social: Adapting Pre-trained Language Models for African Languages Social Media Textaccepted
  130. Agent Laboratory: Using LLM Agents as Research Assistantsaccepted
  131. Agent Trading Arena: A Study on Numerical Understanding in LLM-Based Agentsaccepted
  132. Agent-as-Judge for Factual Summarization of Long Narrativesaccepted
  133. Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Modelsaccepted
  134. AgentDrug: Utilizing Large Language Models in an Agentic Workflow for Zero-Shot Molecular Editingaccepted
  135. AgentInit: Initializing LLM-based Multi-Agent Systems via Diversity and Expertise Orchestration for Effective and Efficient Collaborationaccepted
  136. AgentPro: Enhancing LLM Agents with Automated Process Supervisionaccepted
  137. AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Drivingaccepted
  138. Agentic Medical Knowledge Graphs Enhance Medical Question Answering: Bridging the Gap Between LLMs and Evolving Medical Knowledgeaccepted
  139. Agentic-R1: Distilled Dual-Strategy Reasoningaccepted
  140. Agentic-ToM: Cognition-Inspired Agentic Processing For Enhancing Theory of Mind Reasoningaccepted
  141. AirRAG: Autonomous Strategic Planning and Reasoning Steer Retrieval Augmented Generationaccepted
  142. Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQAaccepted
  143. AlignX: Advancing Multilingual Large Language Models with Multilingual Representation Alignmentaccepted
  144. Aligning Black-Box LLMs for Aspect Sentiment Quad Predictionaccepted
  145. Aligning Dialogue Agents with Global Feedback via Large Language Model Multimodal Reward Decompositionaccepted
  146. Aligning Text/Speech Representations from Multimodal Models with MEG Brain Activity During Listeningaccepted
  147. Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representationsaccepted
  148. Alignment for Efficient Tool Calling of Large Language Modelsaccepted
  149. Alignment with Fill-In-the-Middle for Enhancing Code Generationaccepted
  150. Alignment-Augmented Speculative Decoding with Alignment Sampling and Conditional Verificationaccepted
  151. All Roads Lead to Rome: Graph-Based Confidence Estimation for Large Language Model Reasoningaccepted
  152. All for One: LLMs Solve Mental Math at the Last Token With Information Transferred From Other Tokensaccepted
  153. All-in-one: Understanding and Generation in Multimodal Reasoning with the MAIA Benchmarkaccepted
  154. Alleviating Performance Degradation Caused by Out-of-Distribution Issues in Embedding-Based Retrievalaccepted
  155. AlphaOne: Reasoning Models Thinking Slow and Fast at Test Timeaccepted
  156. Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimizationaccepted
  157. AmpleHate: Amplifying the Attention for Versatile Implicit Hate Detectionaccepted
  158. Amulet: Putting Complex Multi-Turn Conversations on the Stand with LLM Juriesaccepted
  159. An Empirical Study of LLM Reasoning Ability Under Strict Output Length Constraintaccepted
  160. An Empirical Study of Position Bias in Modern Information Retrievalaccepted
  161. An Empirical Study on Strong-Weak Model Collaboration for Repo-level Code Generationaccepted
  162. An Evaluation Resource for Grounding Translation Errorsaccepted
  163. An Improved, Strong Baseline for Pre-Trained Large Language Models as Task-Oriented Dialogue Systemsaccepted
  164. An Interdisciplinary Approach to Human-Centered Machine Translationaccepted
  165. An LLM-based Temporal-spatial Data Generation and Fusion Approach for Early Detection of Late Onset Alzheimer’s Disease (LOAD) Stagings Especially in Chinese and English-speaking Populationsaccepted
  166. An Orthogonal High-Rank Adaptation for Large Language Modelsaccepted
  167. Analysing Chain of Thought Dynamics: Active Guidance or Unfaithful Post-hoc Rationalisation?accepted
  168. Analyzing Dialectical Biases in LLMs for Knowledge and Reasoning Benchmarksaccepted
  169. Analyzing Gambling Addictions: A Spanish Corpus for Understanding Pathological Behavioraccepted
  170. Analyzing Uncertainty of LLM-as-a-Judge: Interval Evaluations with Conformal Predictionaccepted
  171. Analyzing and Modeling LLM Response Lengths with Extreme Value Theory: Anchoring Effects and Hybrid Distributionsaccepted
  172. Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levelsaccepted
  173. Analyzing values about gendered language reform in LLMs’ revisionsaccepted
  174. Anatomy of a Feeling: Narrating Embodied Emotions via Large Vision-Language Modelsaccepted
  175. AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularityaccepted
  176. Anchoring-Guidance Fine-Tuning (AnGFT): Elevating Professional Response Quality in Role-Playing Conversational Agentsaccepted
  177. Anecdoctoring: Automated Red-Teaming Across Language and Placeaccepted
  178. Angular Dispersion Accelerates k-Nearest Neighbors Machine Translationaccepted
  179. Annotating Training Data for Conditional Semantic Textual Similarity Measurement using Large Language Modelsaccepted
  180. Annotation-Efficient Language Model Alignment via Diverse and Representative Response Textsaccepted
  181. Answer Convergence as a Signal for Early Stopping in Reasoningaccepted
  182. Answering Narrative-Driven Recommendation Queries via a Retrieve–Rank Paradigm and the OCG-Agentaccepted
  183. AnyMAC: Cascading Flexible Multi-Agent Collaboration via Next-Agent Predictionaccepted
  184. AraEval: An Arabic Multi-Task Evaluation Suite for Large Language Modelsaccepted
  185. AraReasoner: Evaluating Reasoning-Based LLMs for Arabic NLPaccepted
  186. AraSafe: Benchmarking Safety in Arabic LLMsaccepted
  187. Are Checklists Really Useful for Automatic Evaluation of Generative Tasks?accepted
  188. Are Economists Always More Introverted? Analyzing Consistency in Persona-Assigned LLMsaccepted
  189. Are Generative Models Underconfident? Better Quality Estimation with Boosted Model Probabilityaccepted
  190. Are Knowledge and Reference in Multilingual Language Models Cross-Lingually Consistent?accepted
  191. Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performanceaccepted
  192. Are LLMs Empathetic to All? Investigating the Influence of Multi-Demographic Personas on a Model’s Empathyaccepted
  193. Are Language Models Consequentialist or Deontological Moral Reasoners?accepted
  194. Are Large Language Models Chronically Online Surfers? A Dataset for Chinese Internet Meme Explanationaccepted
  195. Are Stereotypes Leading LLMs’ Zero-Shot Stance Detection ?accepted
  196. Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Studyaccepted
  197. Are the Reasoning Models Good at Automated Essay Scoring?accepted
  198. Are you sure? Measuring models bias in content moderation through uncertaintyaccepted
  199. Arena-lite: Efficient and Reliable Large Language Model Evaluation via Tournament-Based Direct Comparisonsaccepted
  200. ArgCMV: An Argument Summarization Benchmark for the LLM-eraaccepted
  201. Argument Summarization and its Evaluation in the Era of Large Language Modelsaccepted
  202. Artificial Impressions: Evaluating Large Language Model Behavior Through the Lens of Trait Impressionsaccepted
  203. Ask Patients with Patience: Enabling LLMs for Human-Centric Medical Dialogue with Grounded Reasoningaccepted
  204. AskToAct: Enhancing LLMs Tool Use via Self-Correcting Clarificationaccepted
  205. Aspect-Oriented Summarization for Psychiatric Short-Term Readmission Predictionaccepted
  206. Aspect-based Sentiment Analysis via Synthetic Image Generationaccepted
  207. Assay2Mol: Large Language Model-based Drug Design Using BioAssay Contextaccepted
  208. Assess and Prompt: A Generative RL Framework for Improving Engagement in Online Mental Health Communitiesaccepted
  209. Assessing French Readability for Adults with Low Literacy: A Global and Local Perspectiveaccepted
  210. Assessing LLM Reasoning Steps via Principal Knowledge Groundingaccepted
  211. Assessing Socio-Cultural Alignment and Technical Safety of Sovereign LLMsaccepted
  212. Assessing and Mitigating Medical Knowledge Drift and Conflicts in Large Language Modelsaccepted
  213. Assessing effective de-escalation of crisis conversations using transformer-based models and trend statisticsaccepted
  214. Assessing the Role of Data Quality in Training Bilingual Language Modelsaccepted
  215. Assessing the Sensitivity and Alignment of FOL Closeness Metricsaccepted
  216. Assistant-Guided Mitigation of Teacher Preference Bias in LLM-as-a-Judgeaccepted
  217. AssistedDS: Benchmarking How External Domain Knowledge Assists LLMs in Automated Data Scienceaccepted
  218. AssoCiAm: A Benchmark for Evaluating Association Thinking while Circumventing Ambiguityaccepted
  219. Assumed Identities: Quantifying Gender Bias in Machine Translation of Gender-Ambiguous Occupational Termsaccepted
  220. Astra: Efficient Transformer Architecture and Contrastive Dynamics Learning for Embodied Instruction Followingaccepted
  221. Attack as Defense: Safeguarding Large Vision-Language Models from Jailbreaking by Adversarial Attacksaccepted
  222. Attacking Misinformation Detection Using Adversarial Examples Generated by Language Modelsaccepted
  223. Attacks by Content: Automated Fact-checking is an AI Security Issueaccepted
  224. Attention Consistency for LLMs Explanationaccepted
  225. Attention Eclipse: Manipulating Attention to Bypass LLM Safety-Alignmentaccepted
  226. Attention-guided Self-reflection for Zero-shot Hallucination Detection in Large Language Modelsaccepted
  227. AttnComp: Attention-Guided Adaptive Context Compression for Retrieval-Augmented Generationaccepted
  228. Attributes as Textual Genes: Leveraging LLMs as Genetic Algorithm Simulators for Conditional Synthetic Data Generationaccepted
  229. Attribution and Application of Multiple Neurons in Multimodal Large Language Modelsaccepted
  230. Audio-Aware Large Language Models as Judges for Speaking Stylesaccepted
  231. Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Modelsaccepted
  232. Audio-centric Video Understanding Benchmark without Text Shortcutaccepted
  233. Augment before You Try: Knowledge-Enhanced Table Question Answering via Table Expansionaccepted
  234. Augmenting Multi-Agent Communication with State Delta Trajectoryaccepted
  235. AuraDial: A Large-Scale Human-Centric Dialogue Dataset for Chinese AI Psychological Counselingaccepted
  236. Auto-SLURP: A Benchmark Dataset for Evaluating Multi-Agent Frameworks in Smart Personal Assistantaccepted
  237. AutoCT: Automating Interpretable Clinical Trial Prediction with LLM Agentsaccepted
  238. AutoDCWorkflow: LLM-based Data Cleaning Workflow Auto-Generation and Benchmarkaccepted
  239. AutoEvolve: Automatically Evolving Queries for Applicable and Scalable Retrieval-Augmented Generation Benchmarkingaccepted
  240. AutoMIR: Effective Zero-Shot Medical Information Retrieval without Relevance Labelsaccepted
  241. AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-Scientistsaccepted
  242. AutoSpec: An Agentic Framework for Automatically Drafting Patent Specificationaccepted
  243. Autoformalization in the Wild: Assessing LLMs on Real-World Mathematical Definitionsaccepted
  244. Automate Strategy Finding with LLM in Quant Investmentaccepted
  245. Automated Creativity Evaluation for Large Language Models: A Reference-Based Approachaccepted
  246. Automated Knowledge Graph Construction using Large Language Models and Sentence Complexity Modellingaccepted
  247. Automating Alternative Generation in Decision-Makingaccepted
  248. Automating Steering for Safe Multimodal Large Language Modelsaccepted
  249. Automating eHMI Action Design with LLMs for Automated Vehicle Communicationaccepted
  250. Avoidance Decoding for Diverse Multi-Branch Story Generationaccepted
  251. Avoiding Knowledge Edit Skipping in Multi-hop Question Answering with Guided Decompositionaccepted
  252. B-REASO: A Multi-Level Multi-Faceted Bengali Evaluation Suite for Foundation Modelsaccepted
  253. BAGELS: Benchmarking the Automated Generation and Extraction of Limitations from Scholarly Textaccepted
  254. BANMIME : Misogyny Detection with Metaphor Explanation on Bangla Memesaccepted
  255. BBScoreV2: Learning Time-Evolution and Latent Alignment from Stochastic Representationaccepted
  256. BIRD: Bronze Inscription Restoration and Datingaccepted
  257. BOUQuET : dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translationaccepted
  258. BRIT: Bidirectional Retrieval over Unified Image-Text Graphaccepted
  259. BRSpeech-DF: A Deep Fake Synthetic Speech Dataset for Portuguese Zero-Shot TTSaccepted
  260. BSFA: Leveraging the Subspace Dichotomy to Accelerate Neural Network Trainingaccepted
  261. BTC-SAM: Leveraging LLMs for Generation of Bias Test Cases for Sentiment Analysis Modelsaccepted
  262. BTS: Harmonizing Specialized Experts into a Generalist LLMaccepted
  263. BTW: A Non-Parametric Variance Stabilization Framework for Multimodal Model Integrationaccepted
  264. BYOKG-RAG: Multi-Strategy Graph Retrieval for Knowledge Graph Question Answeringaccepted
  265. BabyLM’s First Constructions: Causal interventions provide a signal of learningaccepted
  266. Back Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language Modelsaccepted
  267. Backdoor-Powered Prompt Injection Attacks Nullify Defense Methodsaccepted
  268. BacktrackAgent: Enhancing GUI Agent with Error Detection and Backtracking Mechanismaccepted
  269. Bag of Tricks for Sparse Mixture-of-Experts: A Benchmark Across Reasoning, Efficiency, and Safetyaccepted
  270. Balanced Multi-Factor In-Context Learning for Multilingual Large Language Modelsaccepted
  271. Balcony: A Lightweight Approach to Dynamic Inference of Generative Language Modelsaccepted
  272. BanglaByT5: Byte-Level Modelling for Banglaaccepted
  273. BannerAgency: Advertising Banner Design with Multimodal LLM Agentsaccepted
  274. BannerBench: Benchmarking Vision Language Models for Multi-Ad Selection with Human Preferencesaccepted
  275. Batched Self-Consistency Improves LLM Relevance Assessment and Rankingaccepted
  276. BcQLM: Efficient Vision-Language Understanding with Distilled Q-Gated Cross-Modal Fusionaccepted
  277. BeSimulator: A Large Language Model Powered Text-based Behavior Simulatoraccepted
  278. BehaviorSFT: Behavioral Token Conditioning for Health Agents Across the Proactivity Spectrumaccepted
  279. BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Modelsaccepted
  280. Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarksaccepted
  281. Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Dataaccepted
  282. Benchmarking Critical Questions Generation: A Challenging Reasoning Task for Large Language Modelsaccepted
  283. Benchmarking Debiasing Methods for LLM-based Parameter Estimatesaccepted
  284. Benchmarking Foundation Models with Retrieval-Augmented Generation in Olympic-Level Physics Problem Solvingaccepted
  285. Benchmarking LLMs for Translating Classical Chinese Poetry: Evaluating Adequacy, Fluency, and Eleganceaccepted
  286. Benchmarking LLMs on Semantic Overlap Summarizationaccepted
  287. Benchmarking Large Language Models Under Data Contamination: A Survey from Static to Dynamic Evaluationaccepted
  288. Benchmarking Large Language Models for Cryptanalysis and Side-Channel Vulnerabilitiesaccepted
  289. Benchmarking Uncertainty Metrics for LLM Target-Aware Searchaccepted
  290. Benchmarking and Improving LLM Robustness for Personalized Generationaccepted
  291. Benchmarking and Mitigating MCQA Selection Bias of Large Vision-Language Modelsaccepted
  292. Benchmarking for Domain-Specific LLMs: A Case Study on Academia and Beyondaccepted
  293. Benchmarking the Detection of LLMs-Generated Modern Chinese Poetryaccepted
  294. Beneath the Facade: Probing Safety Vulnerabilities in LLMs via Auto-Generated Jailbreak Promptsaccepted
  295. Beyond A Single AI Cluster: A Survey of Decentralized LLM Trainingaccepted
  296. Beyond Averages: Learning with Annotator Disagreement in STSaccepted
  297. Beyond Binary Preferences: Semi-Online Label-Free GRACE-KTO with Group-Wise Adaptive Calibration for High-Quality Long-Text Generationaccepted
  298. Beyond Checkmate: Exploring the Creative Choke Points for AI Generated Textsaccepted
  299. Beyond Coarse Labels: Fine-Grained Problem Augmentation and Multi-Dimensional Feedback for Emotional Support Conversationaccepted
  300. Beyond Content: How Grammatical Gender Shapes Visual Representation in Text-to-Image Modelsaccepted
  301. Beyond Contrastive Learning: Synthetic Data Enables List-wise Training with Multiple Levels of Relevanceaccepted
  302. Beyond Correctness: Confidence-Aware Reward Modeling for Enhancing Large Language Model Reasoningaccepted
  303. Beyond Demographics: Enhancing Cultural Value Survey Simulation with Multi-Stage Personality-Driven Cognitive Reasoningaccepted
  304. Beyond Demonstrations: Dynamic Vector Construction from Latent Representationsaccepted
  305. Beyond Distribution: Investigating Language Models’ Understanding of Sino-Korean Morphemesaccepted
  306. Beyond Fixed-Length Calibration for Post-Training Compression of LLMsaccepted
  307. Beyond Function-Level Search: Repository-Aware Dual-Encoder Code Retrieval with Adversarial Verificationaccepted
  308. Beyond Guilt: Legal Judgment Prediction with Trichotomous Reasoningaccepted
  309. Beyond Hate Speech: NLP’s Challenges and Opportunities in Uncovering Dehumanizing Languageaccepted
  310. Beyond Human Labels: A Multi-Linguistic Auto-Generated Benchmark for Evaluating Large Language Models on Resume Parsingaccepted
  311. Beyond Inherent Cognition Biases in LLM-Based Event Forecasting: A Multi-Cognition Agentic Frameworkaccepted
  312. Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencodersaccepted
  313. Beyond Linear Steering: Unified Multi-Attribute Control for Language Modelsaccepted
  314. Beyond Online Sampling: Bridging Offline-to-Online Alignment via Dynamic Data Transformation for LLMsaccepted
  315. Beyond Outlining: Heterogeneous Recursive Planning for Adaptive Long-form Writing with Language Modelsaccepted
  316. Beyond Pairwise: Global Zero-shot Temporal Graph Generationaccepted
  317. Beyond Seen Data: Improving KBQA Generalization Through Schema-Guided Logical Form Generationaccepted
  318. Beyond Self-Reports: Multi-Observer Agents for Personality Assessment in Large Language Modelsaccepted
  319. Beyond Single Frames: Can LMMs Comprehend Implicit Narratives in Comic Strip?accepted
  320. Beyond Single-User Dialogue: Assessing Multi-User Dialogue State Tracking Capabilities of Large Language Modelsaccepted
  321. Beyond Spurious Signals: Debiasing Multimodal Large Language Models via Counterfactual Inference and Adaptive Expert Routingaccepted
  322. Beyond Static Testbeds: An Interaction-Centric Agent Simulation Platform for Dynamic Recommender Systemsaccepted
  323. Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Directionaccepted
  324. Beyond Task-Oriented and Chitchat Dialogues: Proactive and Transition-Aware Conversational Agentsaccepted
  325. Beyond Text: Unveiling Privacy Vulnerabilities in Multi-modal Retrieval-Augmented Generationaccepted
  326. Beyond WER: Probing Whisper’s Sub‐token Decoder Across Diverse Language Resource Levelsaccepted
  327. Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoningaccepted
  328. Beyond the Leaderboard: Understanding Performance Disparities in Large Language Models via Model Diffingaccepted
  329. Beyond the Scientific Document: A Citation-Aware Multi-Granular Summarization Approach with Heterogeneous Graphsaccepted
  330. Beyond the Score: Uncertainty-Calibrated LLMs for Automated Essay Assessmentaccepted
  331. Beyond the Surface: A Solution-Aware Retrieval Model for Competition-level Code Generationaccepted
  332. Beyond the Surface: Measuring Self-Preference in LLM Judgmentsaccepted
  333. Beyond the Textual: Generating Coherent Visual Options for MCQsaccepted
  334. Bhaasha, Bhāṣā, Zaban: A Survey for Low-Resourced Languages in South Asia – Current Stage and Challengesaccepted
  335. BiMax: Bidirectional MaxSim Score for Document-Level Alignmentaccepted
  336. BiMediX2 : Bio-Medical EXpert LMM for Diverse Medical Modalitiesaccepted
  337. Bias Analysis and Mitigation through Protected Attribute Detection and Regard Classificationaccepted
  338. Bias Beware: The Impact of Cognitive Biases on LLM-Driven Product Recommendationsaccepted
  339. Bias Mitigation or Cultural Commonsense? Evaluating LLMs with a Japanese Datasetaccepted
  340. Bias after Prompting: Persistent Discrimination in Large Language Modelsaccepted
  341. BiasFilter: An Inference-Time Debiasing Framework for Large Language Modelsaccepted
  342. Biased Tales: Cultural and Topic Bias in Generating Children’s Storiesaccepted
  343. Biology-Instructions: A Dataset and Benchmark for Multi-Omics Sequence Understanding Capability of Large Language Modelsaccepted
  344. Bit-Flip Error Resilience in LLMs: A Comprehensive Analysis and Defense Frameworkaccepted
  345. Bitune: Leveraging Bidirectional Attention to Improve Decoder-Only LLMsaccepted
  346. Blind Men and the Elephant: Diverse Perspectives on Gender Stereotypes in Benchmark Datasetsaccepted
  347. Bold Claims or Self-Doubt? Factuality Hallucination Type Detection via Belief Stateaccepted
  348. Boosting Data Utilization for Multilingual Dense Retrievalaccepted
  349. Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language Modelsaccepted
  350. Both Text and Images Leaked! A Systematic Analysis of Data Contamination in Multimodal LLMaccepted
  351. Boundary Matters: Leveraging Structured Text Plots for Long Text Outline Generationaccepted
  352. BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain Tasksaccepted
  353. BrainLoc: Brain Signal-Based Object Detection with Multi-modal Alignmentaccepted
  354. Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMsaccepted
  355. Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction Amplificationaccepted
  356. Breaking Bad Tokens: Detoxification of LLMs Using Sparse Autoencodersaccepted
  357. Breaking Token Into Concepts: Exploring Extreme Compression in Token Representation Via Compositional Shared Semanticsaccepted
  358. Breaking the Attention Trap in Code LLMs: A Rejection Sampling Approach to Enhance Code Execution Predictionaccepted
  359. Breaking the Noise Barrier: LLM-Guided Semantic Filtering and Enhancement for Multi-Modal Entity Alignmentaccepted
  360. Breaking the Reviewer: Assessing the Vulnerability of Large Language Models in Automated Peer Review Under Textual Adversarial Attacksaccepted
  361. Bridging External and Parametric Knowledge: Mitigating Hallucination of LLMs with Shared-Private Semantic Synergy in Dual-Stream Knowledgeaccepted
  362. Bridging Semantic and Modality Gaps in Zero-Shot Captioning via Retrieval from Synthetic Dataaccepted
  363. Bridging the Capability Gap: Joint Alignment Tuning for Harmonizing LLM-based Multi-Agent Systemsaccepted
  364. Bridging the Creativity Understanding Gap: Small-Scale Human Alignment Enables Expert-Level Humor Ranking in LLMsaccepted
  365. Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoningaccepted
  366. Bridging the Editing Gap in LLMs: FineEdit for Precise and Targeted Text Modificationsaccepted
  367. Bridging the Gap Between Molecule and Textual Descriptions via Substructure-aware Alignmentaccepted
  368. Bringing Pedagogy into Focus: Evaluating Virtual Teaching Assistants’ Question-Answering in Asynchronous Learning Environmentsaccepted
  369. Brittle Minds, Fixable Activations: Understanding Belief Representations in Language Modelsaccepted
  370. Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparencyaccepted
  371. Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systemsaccepted
  372. C3: A Bilingual Benchmark for Spoken Dialogue Models Exploring Challenges in Complex Conversationsaccepted
  373. CAARMA: Class Augmentation with Adversarial Mixup Regularizationaccepted
  374. CAC-CoT: Connector-Aware Compact Chain-of-Thought for Efficient Reasoning Data Synthesis Across Dual-System Cognitive Tasksaccepted
  375. CAFE: Retrieval Head-based Coarse-to-Fine Information Seeking to Enhance Multi-Document QA Capabilityaccepted
  376. CAIR: Counterfactual-based Agent Influence Ranker for Agentic AI Workflowsaccepted
  377. CANDY: Benchmarking LLMs’ Limitations and Assistive Potential in Chinese Misinformation Fact-Checkingaccepted
  378. CAPE: Context-Aware Personality Evaluation Framework for Large Language Modelsaccepted
  379. CARD: Cross-modal Agent Framework for Generative and Editable Residential Designaccepted
  380. CARE: A Disagreement Detection Framework with Concept Alignment and Reasoning Enhancementaccepted
  381. CARE: Multilingual Human Preference Learning for Cultural Awarenessaccepted
  382. CARFT: Boosting LLM Reasoning via Contrastive Learning with Annotated Chain-of-Thought-based Reinforced Fine-Tuningaccepted
  383. CARMA: Enhanced Compositionality in LLMs via Advanced Regularisation and Mutual Information Alignmentaccepted
  384. CARVQ: Corrective Adaptor with Group Residual Vector Quantization for LLM Embedding Compressionaccepted
  385. CAT: Causal Attention Tuning For Injecting Fine-grained Causal Knowledge into Large Language Modelsaccepted
  386. CATCH: A Novel Data Synthesis Framework for High Therapy Fidelity and Memory-Driven Planning Chain of Thought in AI Counselingaccepted
  387. CAVE : Detecting and Explaining Commonsense Anomalies in Visual Environmentsaccepted
  388. CBP-Tuning: Efficient Local Customization for Black-box Large Language Modelsaccepted
  389. CCG: Rare-Label Prediction via Neural SEM–Driven Causal Gameaccepted
  390. CCL-XCoT: An Efficient Cross-Lingual Knowledge Transfer Method for Mitigating Hallucination Generationaccepted
  391. CCQA: Generating Question from Solution Can Improve Inference-Time Reasoning in SLMsaccepted
  392. CDT: A Comprehensive Capability Framework for Large Language Models Across Cognition, Domain, and Taskaccepted
  393. CEMTM: Contextual Embedding-based Multimodal Topic Modelingaccepted
  394. CESRec: Constructing Pseudo Interactions for Sequential Recommendation via Conversational Feedbackaccepted
  395. CHENGYU-BENCH: Benchmarking Large Language Models for Chinese Idiom Understanding and Useaccepted
  396. CHURRO: Making History Readable with an Open-Weight Large Vision-Language Model for High-Accuracy, Low-Cost Historical Text Recognitionaccepted
  397. CIE: Controlling Language Model Text Generations Using Continuous Signalsaccepted
  398. CIFLEX: Contextual Instruction Flow for Sub-task Execution in Multi-Turn Interactions with a Single On-Device LLMaccepted
  399. CIKT: A Collaborative and Iterative Knowledge Tracing Framework with Large Language Modelsaccepted
  400. CIVET: Systematic Evaluation of Understanding in VLMsaccepted
  401. CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers?accepted
  402. CLEAR: A Clinically Grounded Tabular Framework for Radiology Report Evaluationaccepted
  403. CLEAR: A Comprehensive Linguistic Evaluation of Argument Rewriting by Large Language Modelsaccepted
  404. CLEAR: A Framework Enabling Large Language Models to Discern Confusing Legal Paragraphsaccepted
  405. CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcyclingaccepted
  406. CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcyclingaccepted
  407. CLLMate: A Multimodal Benchmark for Weather and Climate Events Forecastingaccepted
  408. CLMTracing: Black-box User-level Watermarking for Code Language Model Tracingaccepted
  409. CLaw: Benchmarking Chinese Legal Knowledge in Large Language Models - A Fine-grained Corpus and Reasoning Analysisaccepted
  410. CM-Align: Consistency-based Multilingual Alignment for Large Language Modelsaccepted
  411. CMHG: A Dataset and Benchmark for Headline Generation of Minority Languages in Chinaaccepted
  412. CMT-Eval: A Novel Chinese Multi-turn Dialogue Evaluation Dataset Addressing Real-world Conversational Challengesaccepted
  413. CMedCalc-Bench: A Fine-Grained Benchmark for Chinese Medical Calculations in LLMaccepted
  414. COAS2W: A Chinese Older-Adults Spoken-to-Written Transformation Corpus with Context Awarenessaccepted
  415. COCO-Tree: Compositional Hierarchical Concept Trees for Enhanced Reasoning in Vision-Language Modelsaccepted
  416. CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillationaccepted
  417. COLA: Collaborative Multi-Agent Framework with Dynamic Task Scheduling for GUI Automationaccepted
  418. COM-BOM: Bayesian Exemplar Search for Efficiently Exploring the Accuracy-Calibration Pareto Frontieraccepted
  419. COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixingaccepted
  420. COUNTDOWN: Contextually Sparse Activation Filtering Out Unnecessary Weights in Down Projectionaccepted
  421. CPO: Addressing Reward Ambiguity in Role-playing Dialogue via Comparative Policy Optimizationaccepted
  422. CR4-NarrEmote: An Open Vocabulary Dataset of Narrative Emotions Derived Using Citizen Scienceaccepted
  423. CREPE: Rapid Chest X-ray Report Evaluation by Predicting Multi-category Error Countsaccepted
  424. CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenariosaccepted
  425. CROP: Contextual Region-Oriented Visual Token Pruningaccepted
  426. CTCC: A Robust and Stealthy Fingerprinting Framework for Large Language Models via Cross-Turn Contextual Correlation Backdooraccepted
  427. CURE: Controlled Unlearning for Robust Embeddings — Mitigating Conceptual Shortcuts in Pre-Trained Language Modelsaccepted
  428. CYCLE-INSTRUCT: Fully Seed-Free Instruction Tuning via Dual Self-Training and Cycle Consistencyaccepted
  429. CaKE: Circuit-aware Editing Enables Generalizable Knowledge Learnersaccepted
  430. CaMMT: Benchmarking Culturally Aware Multimodal Machine Translationaccepted
  431. CaTER: A Framework for Context-aware Topology Entity Retrieval Contrastive Learning in End-to-End Task-Oriented Dialogue Systemsaccepted
  432. Cache Saver: A Modular Framework for Efficient, Affordable, and Reproducible LLM Inferenceaccepted
  433. Cache-Efficient Posterior Sampling for Reinforcement Learning with LLM-Derived Priors Across Discrete and Continuous Domainsaccepted
  434. Cache-of-Thought: Master-Apprentice Framework for Cost-Effective Vision Language Model Reasoningaccepted
  435. Cacheback: Speculative Decoding With Nothing But Cacheaccepted
  436. Calibrating LLM Confidence by Probing Perturbed Representation Stabilityaccepted
  437. Calibrating LLMs for Text-to-SQL Parsing by Leveraging Sub-clause Frequenciesaccepted
  438. Calibrating Pseudo-Labeling with Class Distribution for Semi-supervised Text Classificationaccepted
  439. Calibrating Verbal Uncertainty as a Linear Feature to Reduce Hallucinationsaccepted
  440. Calibration Across Layers: Understanding Calibration Evolution in LLMsaccepted
  441. CalligraphicOCR for Chinese Calligraphy Recognitionaccepted
  442. Can Code-Switched Texts Activate a Knowledge Switch in LLMs? A Case Study on English-Korean Code-Switchingaccepted
  443. Can Federated Learning Safeguard Private Data in LLM Training? Vulnerabilities, Attacks, and Defense Evaluationaccepted
  444. Can GRPO Boost Complex Multimodal Table Understanding?accepted
  445. Can LLM Agents Maintain a Persona in Discourse?accepted
  446. Can LLMs Be Efficient Predictors of Conversational Derailment?accepted
  447. Can LLMs Explain Themselves Counterfactually?accepted
  448. Can LLMs Express Personality Across Cultures? Introducing CulturalPersonas for Evaluating Trait Alignmentaccepted
  449. Can LLMs Extract Frame-Semantic Arguments?accepted
  450. Can LLMs Find a Needle in a Haystack? A Look at Anomaly Detection Language Modelingaccepted
  451. Can LLMs Generate and Solve Linguistic Olympiad Puzzles?accepted
  452. Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environmentsaccepted
  453. Can LLMs Judge Debates? Evaluating Non-Linear Reasoning via Argumentation Theory Semanticsaccepted
  454. Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computationaccepted
  455. Can LLMs Truly Plan? A Comprehensive Evaluation of Planning Capabilitiesaccepted
  456. Can LLMs be Good Graph Judge for Knowledge Graph Construction?accepted
  457. Can LLMs be Literary Companions?: Analysing LLMs on Bengali Figures of Speech Identificationaccepted
  458. Can LLMs simulate the same correct solutions to free-response math problems as real students?accepted
  459. Can Language Models Follow Multiple Turns of Entangled Instructions?accepted
  460. Can Large Language Models Act as Ensembler for Multi-GNNs?accepted
  461. Can Large Language Models Be Good Language Teachers?accepted
  462. Can Large Language Models Identify Implicit Suicidal Ideation? An Empirical Evaluationaccepted
  463. Can Large Language Models Outperform Non-Experts in Poetry Evaluation? A Comparative Study Using the Consensual Assessment Techniqueaccepted
  464. Can Large Language Models Personalize Dialogues to Generational Styles?accepted
  465. Can Large Language Models Tackle Graph Partitioning?accepted
  466. Can Large Language Models Translate Spoken-Only Languages through International Phonetic Transcription?accepted
  467. Can Large Language Models Translate Unseen Languages in Underrepresented Scripts?accepted
  468. Can Large Language Models Unlock Novel Scientific Research Ideas?accepted
  469. Can Large Language Models Win the International Mathematical Games?accepted
  470. Can Large Language Models be Effective Online Opinion Miners?accepted
  471. Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterizationaccepted
  472. Can Multiple Responses from an LLM Reveal the Sources of Its Uncertainty?accepted
  473. Can Out-of-Distribution Evaluations Uncover Reliance on Prediction Shortcuts? A Case Study in Question Answeringaccepted
  474. Can Prompts Rewind Time for LLMs? Evaluating the Effectiveness of Prompted Knowledge Cutoffsaccepted
  475. Can Role Vectors Affect LLM Behaviour?accepted
  476. Can VLMs Recall Factual Associations From Visual References?accepted
  477. Can Vision-Language Models Solve Visual Math Equations?accepted
  478. Can We Edit LLMs for Long-Tail Biomedical Knowledge?accepted
  479. Can We Steer Reasoning Direction by Thinking Intervention?accepted
  480. Can You Trick the Grader? Adversarial Persuasion of LLM Judgesaccepted
  481. Can an Individual Manipulate the Collective Decisions of Multi-Agents?accepted
  482. Can you SPLICE it together? A Human Curated Benchmark for Probing Visual Reasoning in VLMsaccepted
  483. Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimizationaccepted
  484. Capturing Latent Modal Association For Multimodal Entity Alignmentaccepted
  485. Cardiverse: Harnessing LLMs for Novel Card Game Prototypingaccepted
  486. Case-Based Decision-Theoretic Decoding with Quality Memoriesaccepted
  487. Castle: Causal Cascade Updates in Relational Databases with Large Language Modelsaccepted
  488. Catch Me If You Can? Not Yet: LLMs Still Struggle to Imitate the Implicit Writing Styles of Everyday Authorsaccepted
  489. Causal Interventions Reveal Shared Structure Across English Filler–Gap Constructionsaccepted
  490. Causal Representation Learning from Multimodal Clinical Records under Non-Random Modality Missingnessaccepted
  491. Causal Tree Extraction from Medical Case Reports: A Novel Task for Experts-like Text Comprehensionaccepted
  492. Causal-LLM: A Unified One-Shot Framework for Prompt- and Data-Driven Causal Graph Discoveryaccepted
  493. CausalMACE: Causality Empowered Multi-Agents in Minecraft Cooperative Tasksaccepted
  494. CausalVLBench: Benchmarking Visual Causal Reasoning in Large Vision-Language Modelsaccepted
  495. Certainty in Uncertainty: Reasoning over Uncertain Knowledge Graphs with Statistical Guaranteesaccepted
  496. Certified Mitigation of Worst-Case LLM Copyright Infringementaccepted
  497. Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agentsaccepted
  498. Chain of Strategy Optimization Makes Large Language Models Better Emotional Supporteraccepted
  499. Chain-of-Interactions: Multi-step Iterative ICL Framework for Abstractive Task-Oriented Dialogue Summarization of Conversational AI Interactionsaccepted
  500. Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captionsaccepted
  501. Chain-of-Thought Matters: Improving Long-Context Language Models with Reasoning Path Supervisionaccepted
  502. Chain-of-Thought Prompting Obscures Hallucination Cues in Large Language Models: An Empirical Evaluationaccepted
  503. Challenging the Evaluator: LLM Sycophancy Under User Rebuttalaccepted
  504. Chameleon LLMs: User Personas Influence Chatbot Personality Shiftsaccepted
  505. Character is Destiny: Can Persona-assigned Language Models Make Personal Choices?accepted
  506. CharacterCraft: Bridging the Literature-Reality Dialogue Gap for Practical Role-Playing Agentsaccepted
  507. Characterizing Positional Bias in Large Language Models: A Multi-Model Evaluation of Prompt Order Effectsaccepted
  508. Chart2Code53: A Large-Scale Diverse and Complex Dataset for Enhancing Chart-to-Code Generationaccepted
  509. ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinementaccepted
  510. ChartM3: A Multi-Stage Code-Driven Pipeline for Constructing Multi-Dimensional and Multi-Step Visual Reasoning Data in Chart Comprehensionaccepted
  511. ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question Answeringaccepted
  512. Charting the Landscape of African NLP: Mapping Progress and Shaping the Road Aheadaccepted
  513. Chat-Driven Text Generation and Interaction for Person Retrievalaccepted
  514. ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Modelaccepted
  515. Chatbot To Help Patients Understand Their Healthaccepted
  516. CheckEval: A reliable LLM-as-a-Judge framework for evaluating text generation using checklistsaccepted
  517. Child-Directed Language Does Not Consistently Boost Syntax Learning in Language Modelsaccepted
  518. Chinese Toxic Language Mitigation via Sentiment Polarity Consistent Rewritesaccepted
  519. Choosing a Model, Shaping a Future: Comparing LLM Perspectives on Sustainability and its Relationship with AIaccepted
  520. ChronoBias: A Benchmark for Evaluating Time-conditional Group Bias in the Time-sensitive Knowledge of Large Language Modelsaccepted
  521. Circuit Complexity Bounds for RoPE-based Transformer Architectureaccepted
  522. CiteBART: Learning to Generate Citations for Local Citation Recommendationaccepted
  523. CityEQA: A Hierarchical LLM Agent on Embodied Question Answering Benchmark in City Spaceaccepted
  524. ClaimGen-CN: A Large-scale Chinese Dataset for Legal Claim Generationaccepted
  525. ClimateViz: A Benchmark for Statistical Reasoning and Fact Verification on Scientific Chartsaccepted
  526. ClueAnchor: Clue-Anchored Knowledge Reasoning Exploration and Optimization for Retrieval-Augmented Generationaccepted
  527. ClusterUCB: Efficient Gradient-Based Data Selection for Targeted Fine-Tuning of LLMsaccepted
  528. Co-Eval: Augmenting LLM-based Evaluation with Machine Metricsaccepted
  529. Co-Evolving LLMs and Embedding Models via Density-Guided Preference Optimization for Text Clusteringaccepted
  530. CoAT: Chain-of-Associated-Thoughts Framework for Enhancing Large Language Models Reasoningaccepted
  531. CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triplesaccepted
  532. CoBia: Constructed Conversations Can Trigger Otherwise Concealed Societal Biases in LLMsaccepted
  533. CoCoA: Confidence- and Context-Aware Adaptive Decoding for Resolving Knowledge Conflicts in Large Language Modelsaccepted
  534. CoEvo: Coevolution of LLM and Retrieval Model for Domain-Specific Information Retrievalaccepted
  535. CoEx – Co-evolving World-model and Explorationaccepted
  536. CoLA: Compute-Efficient Pre-Training of LLMs via Low-Rank Activationaccepted
  537. CoMAT: Chain of Mathematically Annotated Thought Improves Mathematical Reasoningaccepted
  538. CoMMIT: Coordinated Multimodal Instruction Tuningaccepted
  539. CoMoE: Contrastive Representation for Mixture-of-Experts in Parameter-Efficient Fine-tuningaccepted
  540. CoPL: Collaborative Preference Learning for Personalizing LLMsaccepted
  541. CoRAG: Enhancing Hybrid Retrieval-Augmented Generation through a Cooperative Retriever Architectureaccepted
  542. CoRanking: Collaborative Ranking with Small and Large Ranking Agentsaccepted
  543. CoT-RAG: Integrating Chain of Thought and Retrieval-Augmented Generation to Enhance Reasoning in Large Language Modelsaccepted
  544. CoTD-PO: Chain-of-Thought Distillation with Preference Optimizationaccepted
  545. CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Modelsaccepted
  546. CoVoGER: A Multilingual Multitask Benchmark for Speech-to-text Generative Error Correction with Large Language Modelsaccepted
  547. Coarse-to-Fine Grounded Memory for LLM Agent Planningaccepted
  548. Code Execution as Grounded Supervision for LLM Reasoningaccepted
  549. Code Like Humans: A Multi-Agent Solution for Medical Codingaccepted
  550. Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMsaccepted
  551. CodeArena: Evaluating and Aligning CodeLLMs on Human Preferenceaccepted
  552. CodeComplex: Dataset for Worst-Case Time Complexity Predictionaccepted
  553. CodeContests+: High-Quality Test Case Generation for Competitive Programmingaccepted
  554. CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languagesaccepted
  555. CodeRAG: Finding Relevant and Necessary Knowledge for Retrieval-Augmented Repository-Level Code Completionaccepted
  556. CodeSSM: Towards State Space Models for Code Understandingaccepted
  557. CogAtom: From Cognitive Atoms to Olympiad-level Mathematical Reasoning in Large Language Modelsaccepted
  558. CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewardsaccepted
  559. Cognitive Linguistic Identity Fusion Score (CLIFS): A Scalable Cognition‐Informed Approach to Quantifying Identity Fusion from Textaccepted
  560. Cognitive-Level Adaptive Generation via Capability-Aware Retrieval and Style Adaptationaccepted
  561. Coherence of Argumentative Dialogue Snippets: A New Method for Large Scale Evaluation with an Application to Inference Anchoring Theoryaccepted
  562. Collab-Overcooked: Benchmarking and Evaluating Large Language Models as Collaborative Agentsaccepted
  563. Collaborative Beam Search: Enhancing LLM Reasoning via Collective Consensusaccepted
  564. Collaborative Rational Speech Act: Pragmatic Reasoning for Multi-Turn Dialogaccepted
  565. Columbo: Expanding Abbreviated Column Names for Tabular Data Using Large Language Modelsaccepted
  566. Combining Constrained and Unconstrained Decoding via Boosting: BoostCD and Its Application to Information Extractionaccepted
  567. ComicScene154: A Scene Dataset for Comic Analysisaccepted
  568. CompKBQA: Component-wise Task Decomposition for Knowledge Base Question Answeringaccepted
  569. Comparing Apples to Oranges: A Dataset & Analysis of LLM Humour Understanding from Traditional Puns to Topical Jokesaccepted
  570. Comparing Specialised Small and General Large Language Models on Text Classification: 100 Labelled Samples to Achieve Break-Even Performanceaccepted
  571. Comparing human and LLM politeness strategies in free productionaccepted
  572. CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Rewardaccepted
  573. Complex Numerical Reasoning with Numerical Semantic Pre-training Frameworkaccepted
  574. ComplexTempQA: A 100m Dataset for Complex Temporal Question Answeringaccepted
  575. Composable Cross-prompt Essay Scoring by Merging Modelsaccepted
  576. Compositional Generalisation for Explainable Hate Speech Detectionaccepted
  577. Compositional Translation: A Novel LLM-based Approach for Low-resource Machine Translationaccepted
  578. Compound AI Systems Optimization: A Survey of Methods, Challenges, and Future Directionsaccepted
  579. Comprehensive Evaluation on Lexical Normalization: Boundary-Aware Approaches for Unsegmented Languagesaccepted
  580. Comprehensive and Efficient Distillation for Lightweight Sentiment Analysis Modelsaccepted
  581. Computational Analysis of Character Development in Holocaust Testimoniesaccepted
  582. Computational Analysis of Conversation Dynamics through Participant Responsivityaccepted
  583. ConCISE: Confidence-guided Compression in Step-by-step Efficient Reasoningaccepted
  584. ConText-LE: Cross-Distribution Generalization for Longitudinal Experiential Data via Narrative-Based LLM Representationsaccepted
  585. Conan-Embedding-v2: Training an LLM from Scratch for Text Embeddingsaccepted
  586. Concept-pedia: a Wide-coverage Semantically-annotated Multimodal Datasetaccepted
  587. ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Modelsaccepted
  588. CondAmbigQA: A Benchmark and Dataset for Conditional Ambiguous Question Answeringaccepted
  589. CondenseLM: LLMs-driven Text Dataset Condensation via Reward Matchingaccepted
  590. Conditional [MASK] Discrete Diffusion Language Modelaccepted
  591. Confidence-guided Refinement Reasoning for Zero-shot Question Answeringaccepted
  592. Conflict-Aware Soft Prompting for Retrieval-Augmented Generationaccepted
  593. Conflicting Needles in a Haystack: How LLMs behave when faced with contradictory informationaccepted
  594. Conflicts in Texts: Data, Implications and Challengesaccepted
  595. Confounding Factors in Relating Model Performance to Morphologyaccepted
  596. Confusion is the Final Barrier: Rethinking Jailbreak Evaluation and Investigating the Real Misuse Threat of LLMsaccepted
  597. Connecting the Knowledge Dots: Retrieval-augmented Knowledge Connection for Commonsense Reasoningaccepted
  598. Consistent Discourse-level Temporal Relation Extraction Using Large Language Modelsaccepted
  599. ConsistentChat: Building Skeleton-Guided Consistent Multi-Turn Dialogues for Large Language Models from Scratchaccepted
  600. Constrained Non-negative Matrix Factorization for Guided Topic Modeling of Minority Topicsaccepted
  601. ConstraintLLM: A Neuro-Symbolic Framework for Industrial-Level Constraint Programmingaccepted
  602. Constructing Your Model’s Value Distinction: Towards LLM Alignment with Anchor Words Tuningaccepted
  603. Constructions are Revealed in Word Distributionsaccepted
  604. Context Copying Modulation: The Role of Entropy Neurons in Managing Parametric and Contextual Knowledge Conflictsaccepted
  605. Context Length Alone Hurts LLM Performance Despite Perfect Retrievalaccepted
  606. Context Minimization for Resource-Constrained Text Classification: Optimizing Performance-Efficiency Trade-offs through Linguistic Featuresaccepted
  607. Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learningaccepted
  608. Context and POS in Action: A Comparative Study of Chinese Homonym Disambiguation in Human and Language Modelsaccepted
  609. Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document Embeddingsaccepted
  610. Context-Aware Hierarchical Taxonomy Generation for Scientific Papers via LLM-Guided Multi-Aspect Clusteringaccepted
  611. Context-Aware Membership Inference Attacks against Pre-trained Large Language Modelsaccepted
  612. Context-Aware Reasoning On Parametric Knowledge for Inferring Causal Variablesaccepted
  613. Context-aware Biases for Length Extrapolationaccepted
  614. Continuous-Time Attention: PDE-Guided Mechanisms for Long-Sequence Transformersaccepted
  615. Continuously Steering LLMs Sensitivity to Contextual Knowledge with Proxy Modelsaccepted
  616. Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3Daccepted
  617. ControlText: Unlocking Controllable Fonts in Multilingual Text Rendering without Font Annotationsaccepted
  618. Controllable Memorization in LLMs via Weight Pruningaccepted
  619. Controlled Generation for Private Synthetic Textaccepted
  620. Controlled Retrieval-augmented Context Evaluation for Long-form RAGaccepted
  621. Conv-Basis: A New Paradigm for Efficient Attention Inference and Gradient Computation in Transformersaccepted
  622. ConvSearch-R1: Enhancing Query Reformulation for Conversational Search with Reasoning via Reinforcement Learningaccepted
  623. Convergence and Divergence of Language Models under Different Random Seedsaccepted
  624. Conversational Education at Scale: A Multi-LLM Agent Workflow for Procedural Learning and Pedagogic Quality Assessmentaccepted
  625. Convolutional LoRA Aggregation for Unseen Tasks Adaptationaccepted
  626. CopySpec: Accelerating LLMs with Speculative Copy-and-Pasteaccepted
  627. Correct-Detect: Balancing Performance and Ambiguity Through the Lens of Coreference Resolution in LLMsaccepted
  628. Correlation-Aware Example Selection for In-Context Learning with Nonsymmetric Determinantal Point Processesaccepted
  629. Corrupted but Not Broken: Understanding and Mitigating the Negative Impacts of Corrupted Data in Visual Instruction Tuningaccepted
  630. Cost-Optimal Grouped-Query Attention for Long-Context Modelingaccepted
  631. CourtReasoner: Can LLM Agents Reason Like Judges?accepted
  632. Crafting Customisable Characters with LLMs: A Persona-Driven Role-Playing Agent Frameworkaccepted
  633. Creative Preference Optimizationaccepted
  634. Creativity in LLM-based Multi-Agent Systems: A Surveyaccepted
  635. Crisp: Cognitive Restructuring of Negative Thoughts through Multi-turn Supportive Dialoguesaccepted
  636. Cross-Cultural Transfer of Commonsense Reasoning in LLMs: Evidence from the Arab Worldaccepted
  637. Cross-Document Cross-Lingual NLI via RST-Enhanced Graph Fusion and Interpretability Predictionaccepted
  638. Cross-MoE: An Efficient Temporal Prediction Framework Integrating Textual Modalityaccepted
  639. Cross-domain Rumor Detection via Test-Time Adaptation and Large Language Modelsaccepted
  640. CrossQG: Improving Difficulty-Controllable Question Generation through Consistency Enhancementaccepted
  641. CrystalICL: Enabling In-Context Learning for Crystal Generationaccepted
  642. CtrlNews: LLM-based Multi-Agent Controllable News Writing via Knowledge Gravitational Fieldaccepted
  643. CulturalFrames: Assessing Cultural Expectation Alignment in Text-to-Image Models and Evaluation Metricsaccepted
  644. Culture Cartography: Mapping the Landscape of Cultural Knowledgeaccepted
  645. Culture is Everywhere: A Call for Intentionally Cultural Evaluationaccepted
  646. CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesisaccepted
  647. Curr-ReFT: Overcoming Training Bottlenecks in Small-scale Vision-Language Models via Curriculum Reinforcement Finetuningaccepted
  648. Current Semantic-change Quantification Methods Struggle with Discovery in the Wildaccepted
  649. Curse of Knowledge: Your Guidance and Provided Knowledge are biasing LLM Judges in Complex Evaluationaccepted
  650. Cut the Deadwood Out: Backdoor Purification via Guided Module Substitutionaccepted
  651. D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question Decompositionaccepted
  652. D-RAG: Differentiable Retrieval-Augmented Generation for Knowledge Graph Question Answeringaccepted
  653. D2CS - Documents Graph Clustering using LLM supervisionaccepted
  654. DA-Pred: Performance Prediction for Text Summarization under Domain-Shift and Instruct-Tuningaccepted
  655. DAC: Decomposed Automation Correction for Text-to-SQLaccepted
  656. DAMON: A Dialogue-Aware MCTS Framework for Jailbreaking Large Language Modelsaccepted
  657. DAPE-BR: Distance-Aware Positional Encoding for Mitigating Object Hallucination in LVLMsaccepted
  658. DART: Distilling Autoregressive Reasoning to Silent Thoughtaccepted
  659. DASA-Trans-STM: Adaptive Efficient Transformer for Short Text Matching using Data Augmentation and Semantic Awarenessaccepted
  660. DAVIS: Planning Agent with Knowledge Graph-Powered Inner Monologueaccepted
  661. DB-Explore: Automated Database Exploration and Instruction Synthesis for Text-to-SQLaccepted
  662. DCIS: Efficient Length Extrapolation of LLMs via Divide-and-Conquer Scaling Factor Searchaccepted
  663. DCMKC: A Dual Consistency Matching Approach for Multi-hop Question Answering in LLMsaccepted
  664. DCP: Dual-Cue Pruning for Efficient Large Vision-Language Modelsaccepted
  665. DCR: Quantifying Data Contamination in LLMs Evaluationaccepted
  666. DCRM: A Heuristic to Measure Response Pair Quality in Preference Optimizationaccepted
  667. DDO: Dual-Decision Optimization for LLM-Based Medical Consultation via Multi-Agent Collaborationaccepted
  668. DEBATE, TRAIN, EVOLVE: Self‐Evolution of Language Model Reasoningaccepted
  669. DEL-ToM: Inference-Time Scaling for Theory-of-Mind Reasoning via Dynamic Epistemic Logicaccepted
  670. DELOC: Document Element Localizeraccepted
  671. DICE: Structured Reasoning in LLMs through SLM-Guided Chain-of-Thought Correctionaccepted
  672. DICP: Deep In-Context Prompt for Event Causality Identificationaccepted
  673. DIDS: Domain Impact-aware Data Sampling for Large Language Model Trainingaccepted
  674. DINT Transformeraccepted
  675. DIPLomA: Efficient Adaptation of Instructed LLMs to Low-Resource Languages via Post-Training Delta Mergingaccepted
  676. DISCO Balances the Scales: Adaptive Domain- and Difficulty-Aware Reinforcement Learning on Imbalanced Dataaccepted
  677. DIWALI - Diversity and Inclusivity aWare cuLture specific Items for India: Dataset and Assessment of LLMs for Cultural Text Adaptation in Indian Contextaccepted
  678. DLIR: Spherical Adaptation for Cross-Lingual Knowledge Transfer of Sociological Concepts Alignmentaccepted
  679. DLPO: Towards a Robust, Efficient, and Generalizable Prompt Optimization Framework from a Deep-Learning Perspectiveaccepted
  680. DLTKG: Denoising Logic-based Temporal Knowledge Graph Reasoningaccepted
  681. DM-Codec: Distilling Multimodal Representations for Speech Tokenizationaccepted
  682. DMDTEval: An Evaluation and Analysis of LLMs on Disambiguation in Multi-domain Translationaccepted
  683. DORM: Preference Data Weights Optimization for Reward Modeling in LLM Alignmentaccepted
  684. DP-GTR: Differentially Private Prompt Protection via Group Text Rewritingaccepted
  685. DPED: Multi-Layer Noise Distillation for Privacy-Preserving Text Embeddingsaccepted
  686. DPF-CM: A Data Processing Framework with Privacy-Preserving Vector Databases for Chinese Medical LLMs Training and Deploymentaccepted
  687. DRBO: Mitigating the Bottleneck Effect via Dynamic Reward Balancing in Multi-reward LLM Optimizationaccepted
  688. DRES: Fake news detection by dynamic representation and ensemble selectionaccepted
  689. DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models’ Understanding on Indian Cultureaccepted
  690. DS-MHP: Improving Chain-of-Thought through Dynamic Subgraph-Guided Multi-Hop Pathaccepted
  691. DSCD: Large Language Model Detoxification with Self-Constrained Decodingaccepted
  692. DSG-MCTS: A Dynamic Strategy-Guided Monte Carlo Tree Search for Diversified Reasoning in Large Language Modelsaccepted
  693. DSMoE: Matrix-Partitioned Experts with Dynamic Routing for Computation-Efficient Dense LLMsaccepted
  694. DSVD: Dynamic Self-Verify Decoding for Faithful Generation in Large Language Modelsaccepted
  695. DTDES-KGE: Dual-Teacher Knowledge Distillation with Distinct Embedding Spaces for Knowledge Graph Embeddingsaccepted
  696. DaMoC: Efficiently Selecting the Optimal Large Language Model for Fine-tuning Domain Tasks Based on Data and Model Compressionaccepted
  697. Dagger Behind Smile: Fool LLMs with a Happy Ending Storyaccepted
  698. Data Augmentation for Maltese NLP using Transliterated and Machine Translated Arabic Dataaccepted
  699. Data Descriptions from Large Language Models with Influence Estimationaccepted
  700. Data Doping or True Intelligence? Evaluating the Transferability of Injected Knowledge in LLMsaccepted
  701. Data Drives Unstable Hierarchical Generalization in LMsaccepted
  702. Data or Language Supervision: What Makes CLIP Better than DINO?accepted
  703. Data to Defense: The Role of Curation in Aligning Large Language Models Against Safety Compromiseaccepted
  704. Data-Efficient Hate Speech Detection via Cross-Lingual Nearest Neighbor Retrieval with Limited Labeled Dataaccepted
  705. Data-Efficient Selection via Grammatical Complexity in Continual Pre-training of Domain-Specific LLMsaccepted
  706. Data-scarce Behavior Editing of Language Modelsaccepted
  707. Database-Augmented Query Representation for Information Retrievalaccepted
  708. DatawiseAgent: A Notebook-Centric LLM Agent Framework for Adaptive and Robust Data Science Automationaccepted
  709. Date Fragments: A Hidden Bottleneck of Tokenization for Temporal Reasoningaccepted
  710. David vs. Goliath: Cost-Efficient Financial QA via Cascaded Multi-Agent Reasoningaccepted
  711. DeAR: Dual-Stage Document Reranking with Reasoning Agents via LLM Distillationaccepted
  712. DeCoRe: Decoding by Contrasting Retrieval Heads to Mitigate Hallucinationsaccepted
  713. DeFT-X: Denoised Sparse Fine-Tuning for Zero-Shot Cross-Lingual Transferaccepted
  714. DeKeyNLU: Enhancing Natural Language to SQL Generation through Task Decomposition and Keyword Extractionaccepted
  715. DeMAC: Enhancing Multi-Agent Coordination with Dynamic DAG and Manager-Player Feedbackaccepted
  716. DebUnc: Improving Large Language Model Agent Communication With Uncertainty Metricsaccepted
  717. Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluationaccepted
  718. Debate-to-Detect: Reformulating Misinformation Detection as a Real-World Debate with Large Language Modelsaccepted
  719. Debating for Better Reasoning in Vision-Language Modelsaccepted
  720. Debiasing Multilingual LLMs in Cross-lingual Latent Spaceaccepted
  721. DecisionFlow: Advancing Large Language Model as Principled Decision Makeraccepted
  722. Decoding Dense Embeddings: Sparse Autoencoders for Interpreting and Discretizing Dense Retrievalaccepted
  723. Decoding Uncertainty: The Impact of Decoding Strategies for Uncertainty Estimation in Large Language Modelsaccepted
  724. Decoding in Latent Spaces for Efficient Inference in LLM-based Recommendationaccepted
  725. Decoding the Rule Book: Extracting Hidden Moderation Criteria from Reddit Communitiesaccepted
  726. DecoupleSearch: Decouple Planning and Search via Hierarchical Reward Modelingaccepted
  727. Decoupled Proxy Alignment: Mitigating Language Prior Conflict for Multimodal Alignment in MLLMsaccepted
  728. DecoupledESC: Enhancing Emotional Support Generation via Strategy-Response Decoupled Preference Optimizationaccepted
  729. Deep Associations, High Creativity: A Simple yet Effective Metric for Evaluating Large Language Modelsaccepted
  730. DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environmentsaccepted
  731. DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuningaccepted
  732. DeepWell-Adol: A Scalable Expert-Based Dialogue Corpus for Adolescent Positive Mental Health and Wellbeing Promotionaccepted
  733. Defending against Indirect Prompt Injection by Instruction Detectionaccepted
  734. Definition Generation for Word Meaning Modeling: Monolingual, Multilingual, and Cross-Lingual Perspectivesaccepted
  735. Dementia Through Different Eyes: Explainable Modeling of Human and LLM Perceptions for Early Awarenessaccepted
  736. DemonAgent: Dynamically Encrypted Multi-Backdoor Implantation Attack on LLM-based Agentaccepted
  737. Demystifying Domain-adaptive Post-training for Financial LLMsaccepted
  738. Demystifying Multilingual Reasoning in Process Reward Modelingaccepted
  739. Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfallsaccepted
  740. Demystifying optimized prompts in language modelsaccepted
  741. Dense Retrievers Can Fail on Simple Queries: Revealing The Granularity Dilemma of Embeddingsaccepted
  742. Dependency Parsing-Based Syntactic Enhancement of Relation Extraction in Scientific Textsaccepted
  743. Deriving Strategic Market Insights with Large Language Models: A Benchmark for Forward Counterfactual Generationaccepted
  744. DesignCLIP: Multimodal Learning with CLIP for Design Patent Understandingaccepted
  745. Detecting Continuously Evolving Scam Calls under Limited Annotation: A LLM-Augmented Expert Rule Frameworkaccepted
  746. Detecting Corpus-Level Knowledge Inconsistencies in Wikipedia with Large Language Modelsaccepted
  747. Detecting Knowledge Boundary of Vision Large Language Models by Sampling-Based Inferenceaccepted
  748. Detecting LLM Hallucination Through Layer-wise Information Deficiency: Analysis of Ambiguous Prompts and Unanswerable Questionsaccepted
  749. Detecting Legal Citations in United Kingdom Court Judgmentsaccepted
  750. Detecting Stealthy Backdoor Samples based on Intra-class Distance for Large Language Modelsaccepted
  751. Detoxifying Large Language Models via the Diversity of Toxic Samplesaccepted
  752. Developing and Utilizing a Large-Scale Cantonese Dataset for Multi-Tasking in Large Language Modelsaccepted
  753. DiCoRe: Enhancing Zero-shot Event Detection via Divergent-Convergent LLM Reasoningaccepted
  754. DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoningaccepted
  755. DiNaM: Disinformation Narrative Mining with Large Language Modelsaccepted
  756. Diagnosing Memorization in Chain-of-Thought Reasoning, One Token at a Timeaccepted
  757. Diagnosing Moral Reasoning Acquisition in Language Models: Pragmatics and Generalizationaccepted
  758. Diagram-Driven Course Questions Generationaccepted
  759. Dial-In LLM: Human-Aligned LLM-in-the-loop Intent Clustering for Customer Service Dialoguesaccepted
  760. Dialect-SQL: An Adaptive Framework for Bridging the Dialect Gap in Text-to-SQLaccepted
  761. Dialectal Toxicity Detection: Evaluating LLM-as-a-Judge Consistency Across Language Varietiesaccepted
  762. Differentiated Vision: Unveiling Entity-Specific Visual Modality Requirements for Multimodal Knowledge Graphaccepted
  763. Diffusion vs. Autoregressive Language Models: A Text Embedding Perspectiveaccepted
  764. DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreakaccepted
  765. DiplomacyAgent: Do LLMs Balance Interests and Ethical Principles in International Events?accepted
  766. Dipper: Diversity in Prompts for Producing Large Language Model Ensembles in Reasoning Tasksaccepted
  767. Direct Judgement Preference Optimizationaccepted
  768. Direct Value Optimization: Improving Chain-of-Thought Reasoning in LLMs with Refined Valuesaccepted
  769. DisLoRA: Task-specific Low-Rank Adaptation via Orthogonal Basis from Singular Value Decompositionaccepted
  770. Disambiguation in Conversational Question Answering in the Era of LLMs and Agents: A Surveyaccepted
  771. DisastIR: A Comprehensive Information Retrieval Benchmark for Disaster Managementaccepted
  772. DischargeSim: A Simulation Benchmark for Educational Doctor–Patient Communication at Dischargeaccepted
  773. DiscoSG: Towards Discourse-Level Text Scene Graph Parsing through Iterative Graph Refinementaccepted
  774. Discourse Heuristics For Paradoxically Moral Self-Correctionaccepted
  775. Discourse-Driven Code-Switching: Analyzing the Role of Content and Communicative Function in Spanish-English Bilingual Speechaccepted
  776. Discovering Semantic Subdimensions through Disentangled Conceptual Representationsaccepted
  777. Discrepancy Detection at the Data Level: Toward Consistent Multilingual Question Answeringaccepted
  778. Discrete Minds in a Continuous World: Do Language Models Know Time Passes?accepted
  779. Discriminating Form and Meaning in Multilingual Models with Minimal-Pair ABX Tasksaccepted
  780. Discursive Circuits: How Do Language Models Understand Discourse Relations?accepted
  781. Disentangled Information Bottleneck for Adversarial Text Defenseaccepted
  782. Disentangling Language Understanding and Reasoning Structures in Cross-lingual Chain-of-Thought Promptingaccepted
  783. Disentangling Subjectivity and Uncertainty for Hate Speech Annotation and Modeling using Gazeaccepted
  784. Dissecting Logical Reasoning in LLMs: A Fine-Grained Evaluation and Supervision Studyaccepted
  785. Dissecting Persona-Driven Reasoning in Language Models via Activation Patchingaccepted
  786. Distill Visual Chart Reasoning Ability from LLMs to MLLMsaccepted
  787. Distilling Many-Shot In-Context Learning into a Cheat Sheetaccepted
  788. Distinguishing fair from unfair compositional generalization tasksaccepted
  789. Distributed LLM Serving on Consumer-Grade GPUs by Reconciling Computation and Communicationaccepted
  790. Distribution Prompting: Understanding the Expressivity of Language Models Through the Next-Token Distributions They Can Produceaccepted
  791. Distributional Surgery for Language Model Activationsaccepted
  792. DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Modelsaccepted
  793. DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenesaccepted
  794. DivScore: Zero-Shot Detection of LLM-Generated Text in Specialized Domainsaccepted
  795. Diverse Multi-tool Aggregation with Large Language Models for Enhanced Math Reasoningaccepted
  796. Diverse, not Short: A Length-Controlled Data Selection Strategy for Improving Response Diversity of Language Modelsaccepted
  797. Divide, Optimize, Merge: Scalable Fine-Grained Generative Optimization for LLM Agentsaccepted
  798. Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Modelsaccepted
  799. DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text Generationaccepted
  800. Do All Autoregressive Transformers Remember Facts the Same Way? A Cross-Architecture Analysis of Recall Mechanismsaccepted
  801. Do BERT-Like Bidirectional Models Still Perform Better on Text Classification in the Era of LLMs?accepted
  802. Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluationaccepted
  803. Do Code Semantics Help? A Comprehensive Study on Execution Trace-Based Information for Code Large Language Modelsaccepted
  804. Do Influence Functions Work on Large Language Models?accepted
  805. Do It Yourself (DIY): Modifying Images for Poems in a Zero-Shot Setting Using Weighted Prompt Manipulationaccepted
  806. Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitionsaccepted
  807. Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMsaccepted
  808. Do LLMs Behave as Claimed? Investigating How LLMs Follow Their Own Claims using Counterfactual Questionsaccepted
  809. Do LLMs Encode Frame Semantics? Evidence from Frame Identificationaccepted
  810. Do LLMs Know and Understand Domain Conceptual Knowledge?accepted
  811. Do LLMs Understand Wine Descriptors Across Cultures? A Benchmark for Cultural Adaptations of Wine Reviewsaccepted
  812. Do LVLMs Know What They Know? A Systematic Study of Knowledge Boundary Perception in LVLMsaccepted
  813. Do Large Language Models Truly Grasp Addition? A Rule-Focused Diagnostic Using Two-Integer Arithmeticaccepted
  814. Do Large Language Models Understand Word Senses?accepted
  815. Do Large Language Models excel in Complex Logical Reasoning with Formal Language?accepted
  816. Do RAG Systems Really Suffer From Positional Bias?accepted
  817. Do Slides Help? Multi-modal Context for Automatic Transcription of Conference Talksaccepted
  818. Do We Know What LLMs Don’t Know? A Study of Consistency in Knowledge Probingaccepted
  819. Do We Really Need All Those Dimensions? An Intrinsic Evaluation Framework for Compressed Embeddingsaccepted
  820. Do What? Teaching Vision-Language-Action Models to Reject the Impossibleaccepted
  821. Do You Know About My Nation? Investigating Multilingual Language Models’ Cultural Literacy Through Factual Knowledgeaccepted
  822. Doc2Chart: Intent-Driven Zero-Shot Chart Generation from Documentsaccepted
  823. DocAgent: An Agentic Framework for Multi-Modal Long-Context Document Understandingaccepted
  824. DocAssistant: Integrating Key-region Reading and Step-wise Reasoning for Robust Document Visual Question Answeringaccepted
  825. DocMMIR: A Framework for Document Multi-modal Information Retrievalaccepted
  826. DocReRank: Single-Page Hard Negative Query Generation for Training Multi-Modal RAG Rerankersaccepted
  827. Does Acceleration Cause Hidden Instability in Vision Language Models? Uncovering Instance-Level Divergence Through a Large-Scale Empirical Studyaccepted
  828. Does Context Matter? A Prosodic Comparison of English and Spanish in Monolingual and Multilingual Discourse Settingsaccepted
  829. Does It Run and Is That Enough? Revisiting Text-to-Chart Generation with a Multi-Agent Approachaccepted
  830. Does Localization Inform Unlearning? A Rigorous Examination of Local Parameter Attribution for Knowledge Unlearning in Language Modelsaccepted
  831. Does Reasoning Introduce Bias? A Study of Social Bias Evaluation and Mitigation in LLM Reasoningaccepted
  832. Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models?accepted
  833. Does quantization affect models’ performance on long-context tasks?accepted
  834. Domain Pre-training Impact on Representationsaccepted
  835. DongbaMIE: A Multimodal Information Extraction Dataset for Evaluating Semantic Understanding of Dongba Pictogramsaccepted
  836. Don’t Sweat the Small Stuff: Segment-Level Meta-Evaluation Based on Pairwise Difference Correlationaccepted
  837. Don’t Take the Premise for Granted: Evaluating the Premise Critique Ability of Large Language Modelsaccepted
  838. Doubling Your Data in Minutes: Ultra-fast Tabular Data Generation via LLM-Induced Dependency Graphsaccepted
  839. Dovetail: A CPU/GPU Heterogeneous Speculative Decoding for LLM inferenceaccepted
  840. DrAgent: Empowering Large Language Models as Medical Agents for Multi-hop Medical Reasoningaccepted
  841. DrDiff: Dynamic Routing Diffusion with Hierarchical Attention for Breaking the Efficiency-Quality Trade-offaccepted
  842. DrFrattn: Directly Learn Adaptive Policy from Attention for Simultaneous Machine Translationaccepted
  843. DrKGC: Dynamic Subgraph Retrieval-Augmented LLMs for Knowledge Graph Completion across General and Biomedical Domainsaccepted
  844. Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generationaccepted
  845. Dream to Chat: Model-based Reinforcement Learning on Dialogues with User Belief Modelingaccepted
  846. Drift-Adapter: A Practical Approach to Near Zero-Downtime Embedding Model Upgrades in Vector Databasesaccepted
  847. Drift: Decoding-time Personalized Alignments with Implicit User Preferencesaccepted
  848. Drivel-ology: Challenging LLMs with Interpreting Nonsense with Depthaccepted
  849. Droid: A Resource Suite for AI-Generated Code Detectionaccepted
  850. DroidCall: A Dataset for LLM-powered Android Intent Invocationaccepted
  851. Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMsaccepted
  852. Dual-Path Counterfactual Integration for Multimodal Aspect-Based Sentiment Classificationaccepted
  853. Dual-Path Dynamic Fusion with Learnable Query for Multimodal Sentiment Analysisaccepted
  854. Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbingaccepted
  855. DyePack: Provably Flagging Test Set Contamination in LLMs Using Backdoorsaccepted
  856. Dynamic Collaboration of Multi-Language Models based on Minimal Complete Semantic Unitsaccepted
  857. Dynamic Energy-Based Contrastive Learning with Multi-Stage Knowledge Verification for Event Causality Identificationaccepted
  858. Dynamic Evaluation for Oversensitivity in LLMsaccepted
  859. Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE Adaptationaccepted
  860. Dynamic Injection of Entity Knowledge into Dense Retrieversaccepted
  861. Dynamic Jointly Batch Selection for Data Efficient Machine Translation Fine-Tuningaccepted
  862. Dynamic Model-Bank Test-Time Adaptation for Automatic Speech Recognitionaccepted
  863. Dynamic Retriever for In-Context Knowledge Editing via Policy Optimizationaccepted
  864. Dynamic Simulation Framework for Disinformation Dissemination and Correction With Social Botsaccepted
  865. DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMsaccepted
  866. DynamicNER: A Dynamic, Multilingual, and Fine-Grained Dataset for LLM-based Named Entity Recognitionaccepted
  867. Dyve: Thinking Fast and Slow for Dynamic Process Verificationaccepted
  868. E-Verify: A Paradigm Shift to Scalable Embedding-based Factuality Verificationaccepted
  869. E2LLM: Encoder Elongated Large Language Models for Long-Context Understanding and Reasoningaccepted
  870. ECC: An Emotion-Cause Conversation Dataset for Empathy Responseaccepted
  871. ECO Decoding: Entropy-Based Control for Controllability and Fluency in Controllable Dialogue Generationaccepted
  872. EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understandingaccepted
  873. EIFBENCH: Extremely Complex Instruction Following Benchmark for Large Language Modelsaccepted
  874. EMNLP: Educator-role Moral and Normative Large Language Models Profilingaccepted
  875. EMO-RL: Emotion-Rule-Based Reinforcement Learning Enhanced Audio-Language Model for Generalized Speech Emotion Recognitionaccepted
  876. EMO: Embedding Model Distillation via Intra-Model Relation and Optimal Transport Alignmentsaccepted
  877. EQA-RM: A Generative Embodied Reward Model with Test-time Scalingaccepted
  878. ESC-Judge: A Framework for Comparing Emotional Support Conversational Agentsaccepted
  879. ESGenius: Benchmarking LLMs on Environmental, Social, and Governance (ESG) and Sustainability Knowledgeaccepted
  880. ET-MIER: Entity Type-guided Key Mention Identification and Evidence Retrieval for Document-level Relation Extractionaccepted
  881. EZ-VC: Easy Zero-shot Any-to-Any Voice Conversionaccepted
  882. Easy as PIE? Identifying Multi-Word Expressions with LLMsaccepted
  883. EasyRec: Simple yet Effective Language Models for Recommendationaccepted
  884. Echoes of Agreement: Argument Driven Sycophancy in Large Language modelsaccepted
  885. EcoLANG: Efficient and Effective Agent Communication Language Induction for Social Simulationaccepted
  886. EcoLoRA: Communication-Efficient Federated Fine-Tuning of Large Language Modelsaccepted
  887. EcoSafeRAG: Efficient Security through Context Analysis in Retrieval-Augmented Generationaccepted
  888. EcoTune: Token-Efficient Multi-Fidelity Hyperparameter Optimization for Large Language Model Inferenceaccepted
  889. EditID: Training-Free Editable ID Customization for Text-to-Image Generationaccepted
  890. Editing Across Languages: A Survey of Multilingual Knowledge Editingaccepted
  891. EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMsaccepted
  892. EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture Videosaccepted
  893. Effective Red-Teaming of Policy-Adherent Agentsaccepted
  894. Efficient Beam Search for Large Language Models Using Trie-Based Decodingaccepted
  895. Efficient Compositional Multi-tasking for On-device Large Language Modelsaccepted
  896. Efficient Context Selection for Long-Context QA: No Tuning, No Iteration, Just Adaptive‐kaccepted
  897. Efficient Dynamic Clustering-Based Document Compression for Retrieval-Augmented-Generationaccepted
  898. Efficient Integration of External Knowledge to LLM-based World Models via Retrieval-Augmented Generation and Reinforcement Learningaccepted
  899. Efficient Latent Semantic Clustering for Scaling Test-Time Computation of LLMsaccepted
  900. Efficient Layer-wise LLM Fine-tuning for Revision Intention Predictionaccepted
  901. Efficient Model Development through Fine-tuning Transferaccepted
  902. Efficient Real-time Refinement of Language Model Text Generationaccepted
  903. Efficient Unstructured Pruning of Mamba State-Space Models for Resource-Constrained Environmentsaccepted
  904. EfficientXLang: Towards Improving Token Efficiency Through Cross-Lingual Reasoningaccepted
  905. Efficiently Editing Mixture-of-Experts Models with Compressed Expertsaccepted
  906. Efficiently Selecting Response Generation Strategies for Synthetic Data Construction by Self-Aligned Perplexityaccepted
  907. Eliciting Implicit Acoustic Styles from Open-domain Instructions to Facilitate Fine-grained Controllable Generation of Speechaccepted
  908. Elucidating Mechanisms of Demographic Bias in LLMs for Healthcareaccepted
  909. EmByte: Decomposition and Compression Learning for Small yet Private NLPaccepted
  910. Embedding Domain Knowledge for Large Language Models via Reinforcement Learning from Augmented Generationaccepted
  911. Embedding-Free RAGaccepted
  912. Emergent morpho-phonological representations in self-supervised speech modelsaccepted
  913. EmoAgent: Assessing and Safeguarding Human-AI Interaction for Mental Health Safetyaccepted
  914. EmoBench-UA: A Benchmark Dataset for Emotion Detection in Ukrainianaccepted
  915. EmoGist: Efficient In-Context Learning for Visual Emotion Understandingaccepted
  916. Emotion Transfer with Enhanced Prototype for Unseen Emotion Recognition in Conversationaccepted
  917. Emphasising Structured Information: Integrating Abstract Meaning Representation into LLMs for Enhanced Open-Domain Dialogue Evaluationaccepted
  918. Empowering GraphRAG with Knowledge Filtering and Integrationaccepted
  919. Empowering Math Problem Generation and Reasoning for Large Language Model via Synthetic Data based Continual Learning Frameworkaccepted
  920. EnAnchored-X2X: English-Anchored Optimization for Many-to-Many Translationaccepted
  921. EnDive: A Cross-Dialect Benchmark for Fairness and Performance in Large Language Modelsaccepted
  922. End-to-End Learnable Psychiatric Scale Guided Risky Post Screening for Depression Detection on Social Mediaaccepted
  923. End-to-End Optimization for Multimodal Retrieval-Augmented Generation via Reward Backpropagationaccepted
  924. English as Defense Proxy: Mitigating Multilingual Jailbreak via Eliciting English Safety Knowledgeaccepted
  925. Enhanced Noun-Noun Compound Interpretation through Textual Enrichmentaccepted
  926. Enhancing Attributed Question Answering using Tailored Progressive Curriculum Learningaccepted
  927. Enhancing Chain-of-Thought Reasoning via Neuron Activation Differential Analysisaccepted
  928. Enhancing Chinese Offensive Language Detection with Homophonic Perturbationaccepted
  929. Enhancing Domain-Specific Encoder Models with LLM-Generated Data: How to Leverage Ontologies, and How to Do Without Themaccepted
  930. Enhancing Efficiency and Exploration in Reinforcement Learning for LLMsaccepted
  931. Enhancing Goal-oriented Proactive Dialogue Systems via Dynamic Multi-dimensional Consistency Optimizationaccepted
  932. Enhancing Hate Speech Classifiers through a Gradient-assisted Counterfactual Text Generation Strategyaccepted
  933. Enhancing LLM Knowledge Learning through Generalizationaccepted
  934. Enhancing LLM Language Adaption through Cross-lingual In-Context Pre-trainingaccepted
  935. Enhancing LLM Text Detection with Retrieved Contexts and Logits Distribution Consistencyaccepted
  936. Enhancing LLM-Based Persuasion Simulations with Cultural and Speaker-Specific Informationaccepted
  937. Enhancing LLM-Based Social Bot via an Adversarial Learning Frameworkaccepted
  938. Enhancing Large Language Model for Knowledge Graph Completion via Structure-Aware Alignment-Tuningaccepted
  939. Enhancing Large Vision-Language Models with Ultra-Detailed Image Caption Generationaccepted
  940. Enhancing Logical Reasoning in Language Models via Symbolically-Guided Monte Carlo Process Supervisionaccepted
  941. Enhancing Model Privacy in Federated Learning with Random Masking and Quantizationaccepted
  942. Enhancing Multi-Agent Debate System Performance via Confidence Expressionaccepted
  943. Enhancing Partially Relevant Video Retrieval with Robust Alignment Learningaccepted
  944. Enhancing RAG Efficiency with Adaptive Context Compressionaccepted
  945. Enhancing RLHF with Human Gaze Modelingaccepted
  946. Enhancing Reasoning Abilities of Small LLMs with Cognitive Alignmentaccepted
  947. Enhancing Recommendation Explanations through User-Centric Refinementaccepted
  948. Enhancing SQL Table Acquisition with Reverse Engineering for Text-to-SQLaccepted
  949. Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio Encodersaccepted
  950. Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generationaccepted
  951. Enhancing Study-Level Inference from Clinical Trial Papers via Reinforcement Learning-Based Numeric Reasoningaccepted
  952. Enhancing Time Awareness in Generative Recommendationaccepted
  953. Enrich-on-Graph: Query-Graph Alignment for Complex Reasoning with LLM Enrichingaccepted
  954. Enriching Patent Claim Generation with European Patent Datasetaccepted
  955. Ensembling Prompting Strategies for Zero-Shot Hierarchical Text Classification with Large Language Modelsaccepted
  956. Entity Profile Generation and Reasoning with LLMs for Entity Alignmentaccepted
  957. EoT: Evolution of Thoughts for Complex Reasoning Tasksaccepted
  958. Equal Truth: Rumor Detection with Invariant Group Fairnessaccepted
  959. EquiBench: Benchmarking Large Language Models’ Reasoning about Program Semantics via Equivalence Checkingaccepted
  960. Equipping Retrieval-Augmented Large Language Models with Document Structure Awarenessaccepted
  961. Error Classification of Large Language Models on Math Word Problems: A Dynamically Adaptive Frameworkaccepted
  962. Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervisionaccepted
  963. Estimating LLM Consistency: A User Baseline vs Surrogate Metricsaccepted
  964. Estimating Machine Translation Difficultyaccepted
  965. EuroGEST: Investigating gender stereotypes in multilingual language modelsaccepted
  966. Evaluating Automatic Speech Recognition Systems for Korean Meteorological Expertsaccepted
  967. Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humansaccepted
  968. Evaluating Cognitive-Behavioral Fixation via Multimodal User Viewing Patterns on Social Mediaaccepted
  969. Evaluating Compound AI Systems through Behaviors, Not Benchmarksaccepted
  970. Evaluating Cultural Knowledge and Reasoning in LLMs Through Persian Allusionsaccepted
  971. Evaluating Evaluation Metrics – The Mirage of Hallucination Detectionaccepted
  972. Evaluating Fairness in Large Vision-Language Models Across Diverse Demographic Attributes and Promptsaccepted
  973. Evaluating LLM-Generated Diagrams as Graphsaccepted
  974. Evaluating Language Translation Models by Playing Telephoneaccepted
  975. Evaluating Large Language Models for Belief Inference: Mapping Belief Networks at Scaleaccepted
  976. Evaluating Large Language Models for Cross-Lingual Retrievalaccepted
  977. Evaluating Large Language Models for Detecting Antisemitismaccepted
  978. Evaluating NL2SQL via SQL2NLaccepted
  979. Evaluating Robustness of Large Audio Language Models to Audio Injection: An Empirical Studyaccepted
  980. Evaluating Spatiotemporal Consistency in Automatically Generated Sewing Instructionsaccepted
  981. Evaluating Step-by-step Reasoning Traces: A Surveyaccepted
  982. Evaluating Taxonomy Free Character Role Labeling (TF-CRL) in News Stories using Large Language Modelsaccepted
  983. Evaluating Test-Time Scaling LLMs for Legal Reasoning: OpenAI o1, DeepSeek-R1, and Beyondaccepted
  984. Evaluating Text Generation Quality Using Spectral Distances of Surprisalaccepted
  985. Evaluating Uncertainty Quantification Methods in Argumentative Large Language Modelsaccepted
  986. Evaluating and Aligning Human Economic Risk Preferences in LLMsaccepted
  987. Evaluating distillation methods for data-efficient syntax learningaccepted
  988. Evaluating the Creativity of LLMs in Persian Literary Text Generationaccepted
  989. Evaluating the Effectiveness and Scalability of LLM-Based Data Augmentation for Retrievalaccepted
  990. Evaluating the Evaluators: Are readability metrics good measures of readability?accepted
  991. Evaluating the Robustness and Accuracy of Text Watermarking Under Real-World Cross-Lingual Manipulationsaccepted
  992. Evaluation and Facilitation of Online Discussions in the LLM Era: A Surveyaccepted
  993. Evaluation of Text-to-Image Generation from a Creativity Perspectiveaccepted
  994. EventRelBench: A Comprehensive Benchmark for Evaluating Event Relation Understanding in Large Language Modelsaccepted
  995. EverTracer: Hunting Stolen Large Language Models via Stealthy and Robust Probabilistic Fingerprintaccepted
  996. EvolKV: Evolutionary KV Cache Compression for LLM Inferenceaccepted
  997. Evolution in Simulation: AI-Agent School with Dual Memory for High-Fidelity Educational Dynamicsaccepted
  998. EvolveSearch: An Iterative Self-Evolving Search Agentaccepted
  999. Evolving Chinese Spelling Correction with Corrector-Verifier Collaborationaccepted
  1000. Evolving Stances on Reproducibility: A Longitudinal Study of NLP and ML Researchers’ Views and Experience of Reproducibilityaccepted

Looking for submission deadlines instead? See the conference deadline calendar.