EMNLP 2025 Accepted Papers
The full list of 3,211 papers accepted at EMNLP 2025 (Conference on Empirical Methods in Natural Language Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
accepted: 3,211
- (Almost) Free Modality Stitching of Foundation Modelsaccepted
- 1+1>2: A Synergistic Sparse and Low-Rank Compression Method for Large Language Modelsaccepted
- 2Columns1Row: A Russian Benchmark for Textual and Multimodal Table Understanding and Reasoningaccepted
- 3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillationaccepted
- 3DS: Medical Domain Adaptation of LLMs via Decomposed Difficulty-based Data Selectionaccepted
- 3MDBench: Medical Multimodal Multi-agent Dialogue Benchmarkaccepted
- 3R: Enhancing Sentence Representation Learning via Redundant Representation Reductionaccepted
- A Benchmark for Hindi Verb-Argument Structure Alternationsaccepted
- A Benchmark for Translations Across Styles and Language Variantsaccepted
- A Case Against Implicit Standards: Homophone Normalization in Machine Translation for Languages that use the Ge’ez Script.accepted
- A Category-Theoretic Approach to Neural-Symbolic Task Planning with Bidirectional Searchaccepted
- A Causal Lens for Evaluating Faithfulness Metricsaccepted
- A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Modelsaccepted
- A Comparison of Independent and Joint Fine-tuning Strategies for Retrieval-Augmented Generationaccepted
- A Comprehensive Framework to Operationalize Social Stereotypes for Responsible AI Evaluationsaccepted
- A Comprehensive Literary Chinese Reading Comprehension Dataset with an Evidence Curation Based Solutionaccepted
- A Comprehensive Survey on Learning from Rewards for Large Language Models: Reward Models and Learning Strategiesaccepted
- A Comprehensive Survey on the Trustworthiness of Large Language Models in Healthcareaccepted
- A Comprehensive Taxonomy of Negation for NLP and Neural Retrieversaccepted
- A Computational Simulation of Language Production in First Language Acquisitionaccepted
- A Culturally-diverse Multilingual Multimodal Video Benchmark & Modelaccepted
- A Decoupled Multi-Agent Framework for Complex Text Style Transferaccepted
- A Dynamic Fusion Model for Consistent Crisis Responseaccepted
- A Fully Probabilistic Perspective on Large Language Model Unlearning: Evaluation and Optimizationaccepted
- A Generalizable Rhetorical Strategy Annotation Model Using LLM-based Debate Simulation and Labellingaccepted
- A Generative Framework for Personalized Sticker Retrievalaccepted
- A Generative Pre-Trained Language Model for Channel Prediction in Wireless Communications Systemsaccepted
- A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Usersaccepted
- A Graph-Theoretical Framework for Analyzing the Behavior of Causal Language Modelsaccepted
- A Group Fairness Lens for Large Language Modelsaccepted
- A Head to Predict and a Head to Question: Pre-trained Uncertainty Quantification Heads for Hallucination Detection in LLM Outputsaccepted
- A Knapsack by Any Other Name: Presentation impacts LLM performance on NP-hard problemsaccepted
- A Knowledge-driven Adaptive Collaboration of LLMs for Enhancing Medical Decision-makingaccepted
- A Middle Path for On-Premises LLM Deployment: Preserving Privacy Without Sacrificing Model Confidentialityaccepted
- A Monte-Carlo Sampling Framework For Reliable Evaluation of Large Language Models Using Behavioral Analysisaccepted
- A Multi-Agent Framework with Automated Decision Rule Optimization for Cross-Domain Misinformation Detectionaccepted
- A Multi-Level Benchmark for Causal Language Understanding in Social Media Discourseaccepted
- A Multilingual, Culture-First Approach to Addressing Misgendering in LLM Applicationsaccepted
- A Necessary Step toward Faithfulness: Measuring and Improving Consistency in Free-Text Explanationsaccepted
- A Position Paper on the Automatic Generation of Machine Learning Leaderboardsaccepted
- A Probabilistic Inference Scaling Theory for LLM Self-Correctionaccepted
- A Rigorous Evaluation of LLM Data Generation Strategies for Low-Resource Languagesaccepted
- A Sequential Multi-Stage Approach for Code Vulnerability Detection via Confidence- and Collaboration-based Decision Makingaccepted
- A Similarity Measure for Comparing Conversational Dynamicsaccepted
- A Simple Yet Effective Method for Non-Refusing Context Relevant Fine-grained Safety Steering in LLMsaccepted
- A Structured Framework for Evaluating and Enhancing Interpretive Capabilities of Multimodal LLMs in Culturally Situated Tasksaccepted
- A Survey of Cognitive Distortion Detection and Classification in NLPaccepted
- A Survey of Link Prediction in N-ary Knowledge Graphsaccepted
- A Survey of Multilingual Reasoning in Language Modelsaccepted
- A Survey of Pun Generation: Datasets, Evaluations and Methodologiesaccepted
- A Survey of RAG-Reasoning Systems in Large Language Modelsaccepted
- A Survey on LLM-powered Agents for Recommender Systemsaccepted
- A Survey on LLMs for Story Generationaccepted
- A Survey on Multi-modal Intent Recognition: Recent Advances and New Frontiersaccepted
- A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Modelsaccepted
- A Survey on Training-free Alignment of Large Language Modelsaccepted
- A Symbolic Adversarial Learning Framework for Evolving Fake News Generation and Detectionaccepted
- A Systematic Analysis of Base Model Choice for Reward Modelingaccepted
- A Systematic Survey of Automatic Prompt Optimization Techniquesaccepted
- A Systematic Survey of Claim Verification: Corpora, Systems, and Case Studiesaccepted
- A Text-Based Recommender System that Leverages Explicit Affective State Preferencesaccepted
- A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolationaccepted
- A Unified Framework for N-ary Property Information Extraction in Materials Scienceaccepted
- A Zero-Shot Neuro-Symbolic Approach for Complex Knowledge Graph Question Answeringaccepted
- ACEBench: A Comprehensive Evaluation of LLM Tool Usageaccepted
- ACING: Actor-Critic for Instruction Learning in Black-Box LLMsaccepted
- AELC: Adaptive Entity Linking with LLM-Driven Contextualizationaccepted
- AFRIDOC-MT: Document-level MT Corpus for African Languagesaccepted
- AGENTVIGIL: Automatic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agentsaccepted
- AI Argues Differently: Distinct Argumentative and Linguistic Patterns of LLMs in Persuasive Contextsaccepted
- AI Chatbots as Professional Service Agents: Developing a Professional Identityaccepted
- AI Knows Where You Are: Exposure, Bias, and Inference in Multimodal Geolocation with KoreaGEOaccepted
- AI Sees Your Location—But With A Bias Toward The Wealthy Worldaccepted
- AIMMerging: Adaptive Iterative Model Merging Using Training Trajectories for Language Model Continual Learningaccepted
- AIP: Subverting Retrieval-Augmented Generation via Adversarial Instructional Promptaccepted
- AIR: Complex Instruction Generation via Automatic Iterative Refinementaccepted
- AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Scienceaccepted
- ALLabel: Three-stage Active Learning for LLM-based Entity Recognition using Demonstration Retrievalaccepted
- ALRPHFS: Adversarially Learned Risk Patterns with Hierarchical Fast & Slow Reasoning for Robust Agent Defenseaccepted
- AMACE: Automatic Multi-Agent Chart Evolution for Iteratively Tailored Chart Generationaccepted
- AMANDA: Agentic Medical Knowledge Augmentation for Data-Efficient Medical Visual Question Answeringaccepted
- AMIA: Automatic Masking and Joint Intention Analysis Makes LVLMs Robust Jailbreak Defendersaccepted
- AMQ: Enabling AutoML for Mixed-precision Weight-Only Quantization of Large Language Modelsaccepted
- APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transportaccepted
- AQuilt: Weaving Logic and Self-Inspection into Low-Cost, High-Relevance Data Synthesis for Specialist LLMsaccepted
- AROMA: Autonomous Rank-one Matrix Adaptationaccepted
- ARXSA: A General Negative Feedback Control Theory in Vision-Language Modelsaccepted
- ASD-iLLM:An Intervention Large Language Model for Autistic Children based on Real Clinical Dialogue Intervention Datasetaccepted
- ASTPrompter: Preference-Aligned Automated Language Model Red-Teaming to Generate Low-Perplexity Unsafe Promptsaccepted
- ASTRA: A Negotiation Agent with Adaptive and Strategic Reasoning via Tool-integrated Action for Dynamic Offer Optimizationaccepted
- AbsVis – Benchmarking How Humans and Vision-Language Models “See” Abstract Concepts in Imagesaccepted
- AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Modelsaccepted
- Accelerate Parallelizable Reasoning via Parallel Decoding within One Sequenceaccepted
- Accelerated Test-Time Scaling with Model-Free Speculative Samplingaccepted
- Accelerating LLM Reasoning via Early Rejection with Partial Reward Modelingaccepted
- Accept or Deny? Evaluating LLM Fairness and Performance in Loan Approval across Table-to-Text Serialization Approachesaccepted
- AccessEval: Benchmarking Disability Bias in Large Language Modelsaccepted
- Acquiescence Bias in Large Language Modelsaccepted
- ActionStudio: A Lightweight Framework for Data and Training of Large Action Modelsaccepted
- Active Domain Knowledge Acquisition with 100-Dollar Budget: Enhancing LLMs via Cost-Efficient, Expert-Involved Interaction in Sensitive Domainsaccepted
- Active Layer-Contrastive Decoding Reduces Hallucination in Large Language Model Generationaccepted
- Active Learning for Multidialectal Arabic POS Taggingaccepted
- AdDriftBench: A Benchmark for Detecting Data Drift and Label Drift in Short Video Advertisingaccepted
- AdaRewriter: Unleashing the Power of Prompting-based Conversational Query Reformulation via Test-Time Adaptationaccepted
- AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defenderaccepted
- AdaTP: Attention-Debiased Token Pruning for Video Large Language Modelsaccepted
- AdamS: Momentum Itself Can Be A Normalizer for LLM Pretraining and Post-trainingaccepted
- AdaptFlow: Adaptive Workflow Optimization via Meta-Learningaccepted
- AdaptMerge: Inference Time Adaptive Visual and Language-Guided Token Merging for Efficient Large Multimodal Modelsaccepted
- AdaptThink: Reasoning Models Can Learn When to Thinkaccepted
- Adapting Bias Evaluation to Domain Contexts using Generative Modelsaccepted
- Adapting Large Language Models for Character-based Augmentative and Alternative Communicationaccepted
- Adaptive LLM Routing under Budget Constraintsaccepted
- Adaptive Platt Scaling with Causal Interpretations for Self-Reflective Language Model Uncertainty Estimatesaccepted
- Adaptive Preference Optimization with Uncertainty-aware Utility Anchoraccepted
- Adaptive Schema-aware Event Extraction with Retrieval-Augmented Generationaccepted
- Adaptively profiling models with task elicitationaccepted
- Add-One-In: Incremental Sample Selection for Large Language Models via a Choice-Based Greedy Paradigmaccepted
- Addition in Four Movements: Mapping Layer-wise Information Trajectories in LLMsaccepted
- Addressing Tokenization Inconsistency in Steganography and Watermarking Based on Large Language Modelsaccepted
- Advancing Arabic Diacritization: Improved Datasets, Benchmarking, and State-of-the-Art Modelsaccepted
- Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Modelsaccepted
- Advancing Oversight Reasoning across Languages for Audit Sycophantic Behaviour via X-Agentaccepted
- Advancing Reasoning with Off-the-Shelf LLMs: A Semantic Structure Perspectiveaccepted
- Adversarial Attacks Against Automated Fact-Checking: A Surveyaccepted
- Adversary-Aware DPO: Enhancing Safety Alignment in Vision Language Models via Adversarial Trainingaccepted
- AesBiasBench: Evaluating Bias and Alignment in Multimodal Language Models for Personalized Image Aesthetic Assessmentaccepted
- Africa Health Check: Probing Cultural Bias in Medical LLMsaccepted
- AfroXLMR-Social: Adapting Pre-trained Language Models for African Languages Social Media Textaccepted
- Agent Laboratory: Using LLM Agents as Research Assistantsaccepted
- Agent Trading Arena: A Study on Numerical Understanding in LLM-Based Agentsaccepted
- Agent-as-Judge for Factual Summarization of Long Narrativesaccepted
- Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Modelsaccepted
- AgentDrug: Utilizing Large Language Models in an Agentic Workflow for Zero-Shot Molecular Editingaccepted
- AgentInit: Initializing LLM-based Multi-Agent Systems via Diversity and Expertise Orchestration for Effective and Efficient Collaborationaccepted
- AgentPro: Enhancing LLM Agents with Automated Process Supervisionaccepted
- AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Drivingaccepted
- Agentic Medical Knowledge Graphs Enhance Medical Question Answering: Bridging the Gap Between LLMs and Evolving Medical Knowledgeaccepted
- Agentic-R1: Distilled Dual-Strategy Reasoningaccepted
- Agentic-ToM: Cognition-Inspired Agentic Processing For Enhancing Theory of Mind Reasoningaccepted
- AirRAG: Autonomous Strategic Planning and Reasoning Steer Retrieval Augmented Generationaccepted
- Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQAaccepted
- AlignX: Advancing Multilingual Large Language Models with Multilingual Representation Alignmentaccepted
- Aligning Black-Box LLMs for Aspect Sentiment Quad Predictionaccepted
- Aligning Dialogue Agents with Global Feedback via Large Language Model Multimodal Reward Decompositionaccepted
- Aligning Text/Speech Representations from Multimodal Models with MEG Brain Activity During Listeningaccepted
- Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representationsaccepted
- Alignment for Efficient Tool Calling of Large Language Modelsaccepted
- Alignment with Fill-In-the-Middle for Enhancing Code Generationaccepted
- Alignment-Augmented Speculative Decoding with Alignment Sampling and Conditional Verificationaccepted
- All Roads Lead to Rome: Graph-Based Confidence Estimation for Large Language Model Reasoningaccepted
- All for One: LLMs Solve Mental Math at the Last Token With Information Transferred From Other Tokensaccepted
- All-in-one: Understanding and Generation in Multimodal Reasoning with the MAIA Benchmarkaccepted
- Alleviating Performance Degradation Caused by Out-of-Distribution Issues in Embedding-Based Retrievalaccepted
- AlphaOne: Reasoning Models Thinking Slow and Fast at Test Timeaccepted
- Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimizationaccepted
- AmpleHate: Amplifying the Attention for Versatile Implicit Hate Detectionaccepted
- Amulet: Putting Complex Multi-Turn Conversations on the Stand with LLM Juriesaccepted
- An Empirical Study of LLM Reasoning Ability Under Strict Output Length Constraintaccepted
- An Empirical Study of Position Bias in Modern Information Retrievalaccepted
- An Empirical Study on Strong-Weak Model Collaboration for Repo-level Code Generationaccepted
- An Evaluation Resource for Grounding Translation Errorsaccepted
- An Improved, Strong Baseline for Pre-Trained Large Language Models as Task-Oriented Dialogue Systemsaccepted
- An Interdisciplinary Approach to Human-Centered Machine Translationaccepted
- An LLM-based Temporal-spatial Data Generation and Fusion Approach for Early Detection of Late Onset Alzheimer’s Disease (LOAD) Stagings Especially in Chinese and English-speaking Populationsaccepted
- An Orthogonal High-Rank Adaptation for Large Language Modelsaccepted
- Analysing Chain of Thought Dynamics: Active Guidance or Unfaithful Post-hoc Rationalisation?accepted
- Analyzing Dialectical Biases in LLMs for Knowledge and Reasoning Benchmarksaccepted
- Analyzing Gambling Addictions: A Spanish Corpus for Understanding Pathological Behavioraccepted
- Analyzing Uncertainty of LLM-as-a-Judge: Interval Evaluations with Conformal Predictionaccepted
- Analyzing and Modeling LLM Response Lengths with Extreme Value Theory: Anchoring Effects and Hybrid Distributionsaccepted
- Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levelsaccepted
- Analyzing values about gendered language reform in LLMs’ revisionsaccepted
- Anatomy of a Feeling: Narrating Embodied Emotions via Large Vision-Language Modelsaccepted
- AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularityaccepted
- Anchoring-Guidance Fine-Tuning (AnGFT): Elevating Professional Response Quality in Role-Playing Conversational Agentsaccepted
- Anecdoctoring: Automated Red-Teaming Across Language and Placeaccepted
- Angular Dispersion Accelerates k-Nearest Neighbors Machine Translationaccepted
- Annotating Training Data for Conditional Semantic Textual Similarity Measurement using Large Language Modelsaccepted
- Annotation-Efficient Language Model Alignment via Diverse and Representative Response Textsaccepted
- Answer Convergence as a Signal for Early Stopping in Reasoningaccepted
- Answering Narrative-Driven Recommendation Queries via a Retrieve–Rank Paradigm and the OCG-Agentaccepted
- AnyMAC: Cascading Flexible Multi-Agent Collaboration via Next-Agent Predictionaccepted
- AraEval: An Arabic Multi-Task Evaluation Suite for Large Language Modelsaccepted
- AraReasoner: Evaluating Reasoning-Based LLMs for Arabic NLPaccepted
- AraSafe: Benchmarking Safety in Arabic LLMsaccepted
- Are Checklists Really Useful for Automatic Evaluation of Generative Tasks?accepted
- Are Economists Always More Introverted? Analyzing Consistency in Persona-Assigned LLMsaccepted
- Are Generative Models Underconfident? Better Quality Estimation with Boosted Model Probabilityaccepted
- Are Knowledge and Reference in Multilingual Language Models Cross-Lingually Consistent?accepted
- Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performanceaccepted
- Are LLMs Empathetic to All? Investigating the Influence of Multi-Demographic Personas on a Model’s Empathyaccepted
- Are Language Models Consequentialist or Deontological Moral Reasoners?accepted
- Are Large Language Models Chronically Online Surfers? A Dataset for Chinese Internet Meme Explanationaccepted
- Are Stereotypes Leading LLMs’ Zero-Shot Stance Detection ?accepted
- Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Studyaccepted
- Are the Reasoning Models Good at Automated Essay Scoring?accepted
- Are you sure? Measuring models bias in content moderation through uncertaintyaccepted
- Arena-lite: Efficient and Reliable Large Language Model Evaluation via Tournament-Based Direct Comparisonsaccepted
- ArgCMV: An Argument Summarization Benchmark for the LLM-eraaccepted
- Argument Summarization and its Evaluation in the Era of Large Language Modelsaccepted
- Artificial Impressions: Evaluating Large Language Model Behavior Through the Lens of Trait Impressionsaccepted
- Ask Patients with Patience: Enabling LLMs for Human-Centric Medical Dialogue with Grounded Reasoningaccepted
- AskToAct: Enhancing LLMs Tool Use via Self-Correcting Clarificationaccepted
- Aspect-Oriented Summarization for Psychiatric Short-Term Readmission Predictionaccepted
- Aspect-based Sentiment Analysis via Synthetic Image Generationaccepted
- Assay2Mol: Large Language Model-based Drug Design Using BioAssay Contextaccepted
- Assess and Prompt: A Generative RL Framework for Improving Engagement in Online Mental Health Communitiesaccepted
- Assessing French Readability for Adults with Low Literacy: A Global and Local Perspectiveaccepted
- Assessing LLM Reasoning Steps via Principal Knowledge Groundingaccepted
- Assessing Socio-Cultural Alignment and Technical Safety of Sovereign LLMsaccepted
- Assessing and Mitigating Medical Knowledge Drift and Conflicts in Large Language Modelsaccepted
- Assessing effective de-escalation of crisis conversations using transformer-based models and trend statisticsaccepted
- Assessing the Role of Data Quality in Training Bilingual Language Modelsaccepted
- Assessing the Sensitivity and Alignment of FOL Closeness Metricsaccepted
- Assistant-Guided Mitigation of Teacher Preference Bias in LLM-as-a-Judgeaccepted
- AssistedDS: Benchmarking How External Domain Knowledge Assists LLMs in Automated Data Scienceaccepted
- AssoCiAm: A Benchmark for Evaluating Association Thinking while Circumventing Ambiguityaccepted
- Assumed Identities: Quantifying Gender Bias in Machine Translation of Gender-Ambiguous Occupational Termsaccepted
- Astra: Efficient Transformer Architecture and Contrastive Dynamics Learning for Embodied Instruction Followingaccepted
- Attack as Defense: Safeguarding Large Vision-Language Models from Jailbreaking by Adversarial Attacksaccepted
- Attacking Misinformation Detection Using Adversarial Examples Generated by Language Modelsaccepted
- Attacks by Content: Automated Fact-checking is an AI Security Issueaccepted
- Attention Consistency for LLMs Explanationaccepted
- Attention Eclipse: Manipulating Attention to Bypass LLM Safety-Alignmentaccepted
- Attention-guided Self-reflection for Zero-shot Hallucination Detection in Large Language Modelsaccepted
- AttnComp: Attention-Guided Adaptive Context Compression for Retrieval-Augmented Generationaccepted
- Attributes as Textual Genes: Leveraging LLMs as Genetic Algorithm Simulators for Conditional Synthetic Data Generationaccepted
- Attribution and Application of Multiple Neurons in Multimodal Large Language Modelsaccepted
- Audio-Aware Large Language Models as Judges for Speaking Stylesaccepted
- Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Modelsaccepted
- Audio-centric Video Understanding Benchmark without Text Shortcutaccepted
- Augment before You Try: Knowledge-Enhanced Table Question Answering via Table Expansionaccepted
- Augmenting Multi-Agent Communication with State Delta Trajectoryaccepted
- AuraDial: A Large-Scale Human-Centric Dialogue Dataset for Chinese AI Psychological Counselingaccepted
- Auto-SLURP: A Benchmark Dataset for Evaluating Multi-Agent Frameworks in Smart Personal Assistantaccepted
- AutoCT: Automating Interpretable Clinical Trial Prediction with LLM Agentsaccepted
- AutoDCWorkflow: LLM-based Data Cleaning Workflow Auto-Generation and Benchmarkaccepted
- AutoEvolve: Automatically Evolving Queries for Applicable and Scalable Retrieval-Augmented Generation Benchmarkingaccepted
- AutoMIR: Effective Zero-Shot Medical Information Retrieval without Relevance Labelsaccepted
- AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-Scientistsaccepted
- AutoSpec: An Agentic Framework for Automatically Drafting Patent Specificationaccepted
- Autoformalization in the Wild: Assessing LLMs on Real-World Mathematical Definitionsaccepted
- Automate Strategy Finding with LLM in Quant Investmentaccepted
- Automated Creativity Evaluation for Large Language Models: A Reference-Based Approachaccepted
- Automated Knowledge Graph Construction using Large Language Models and Sentence Complexity Modellingaccepted
- Automating Alternative Generation in Decision-Makingaccepted
- Automating Steering for Safe Multimodal Large Language Modelsaccepted
- Automating eHMI Action Design with LLMs for Automated Vehicle Communicationaccepted
- Avoidance Decoding for Diverse Multi-Branch Story Generationaccepted
- Avoiding Knowledge Edit Skipping in Multi-hop Question Answering with Guided Decompositionaccepted
- B-REASO: A Multi-Level Multi-Faceted Bengali Evaluation Suite for Foundation Modelsaccepted
- BAGELS: Benchmarking the Automated Generation and Extraction of Limitations from Scholarly Textaccepted
- BANMIME : Misogyny Detection with Metaphor Explanation on Bangla Memesaccepted
- BBScoreV2: Learning Time-Evolution and Latent Alignment from Stochastic Representationaccepted
- BIRD: Bronze Inscription Restoration and Datingaccepted
- BOUQuET : dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translationaccepted
- BRIT: Bidirectional Retrieval over Unified Image-Text Graphaccepted
- BRSpeech-DF: A Deep Fake Synthetic Speech Dataset for Portuguese Zero-Shot TTSaccepted
- BSFA: Leveraging the Subspace Dichotomy to Accelerate Neural Network Trainingaccepted
- BTC-SAM: Leveraging LLMs for Generation of Bias Test Cases for Sentiment Analysis Modelsaccepted
- BTS: Harmonizing Specialized Experts into a Generalist LLMaccepted
- BTW: A Non-Parametric Variance Stabilization Framework for Multimodal Model Integrationaccepted
- BYOKG-RAG: Multi-Strategy Graph Retrieval for Knowledge Graph Question Answeringaccepted
- BabyLM’s First Constructions: Causal interventions provide a signal of learningaccepted
- Back Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language Modelsaccepted
- Backdoor-Powered Prompt Injection Attacks Nullify Defense Methodsaccepted
- BacktrackAgent: Enhancing GUI Agent with Error Detection and Backtracking Mechanismaccepted
- Bag of Tricks for Sparse Mixture-of-Experts: A Benchmark Across Reasoning, Efficiency, and Safetyaccepted
- Balanced Multi-Factor In-Context Learning for Multilingual Large Language Modelsaccepted
- Balcony: A Lightweight Approach to Dynamic Inference of Generative Language Modelsaccepted
- BanglaByT5: Byte-Level Modelling for Banglaaccepted
- BannerAgency: Advertising Banner Design with Multimodal LLM Agentsaccepted
- BannerBench: Benchmarking Vision Language Models for Multi-Ad Selection with Human Preferencesaccepted
- Batched Self-Consistency Improves LLM Relevance Assessment and Rankingaccepted
- BcQLM: Efficient Vision-Language Understanding with Distilled Q-Gated Cross-Modal Fusionaccepted
- BeSimulator: A Large Language Model Powered Text-based Behavior Simulatoraccepted
- BehaviorSFT: Behavioral Token Conditioning for Health Agents Across the Proactivity Spectrumaccepted
- BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Modelsaccepted
- Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarksaccepted
- Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Dataaccepted
- Benchmarking Critical Questions Generation: A Challenging Reasoning Task for Large Language Modelsaccepted
- Benchmarking Debiasing Methods for LLM-based Parameter Estimatesaccepted
- Benchmarking Foundation Models with Retrieval-Augmented Generation in Olympic-Level Physics Problem Solvingaccepted
- Benchmarking LLMs for Translating Classical Chinese Poetry: Evaluating Adequacy, Fluency, and Eleganceaccepted
- Benchmarking LLMs on Semantic Overlap Summarizationaccepted
- Benchmarking Large Language Models Under Data Contamination: A Survey from Static to Dynamic Evaluationaccepted
- Benchmarking Large Language Models for Cryptanalysis and Side-Channel Vulnerabilitiesaccepted
- Benchmarking Uncertainty Metrics for LLM Target-Aware Searchaccepted
- Benchmarking and Improving LLM Robustness for Personalized Generationaccepted
- Benchmarking and Mitigating MCQA Selection Bias of Large Vision-Language Modelsaccepted
- Benchmarking for Domain-Specific LLMs: A Case Study on Academia and Beyondaccepted
- Benchmarking the Detection of LLMs-Generated Modern Chinese Poetryaccepted
- Beneath the Facade: Probing Safety Vulnerabilities in LLMs via Auto-Generated Jailbreak Promptsaccepted
- Beyond A Single AI Cluster: A Survey of Decentralized LLM Trainingaccepted
- Beyond Averages: Learning with Annotator Disagreement in STSaccepted
- Beyond Binary Preferences: Semi-Online Label-Free GRACE-KTO with Group-Wise Adaptive Calibration for High-Quality Long-Text Generationaccepted
- Beyond Checkmate: Exploring the Creative Choke Points for AI Generated Textsaccepted
- Beyond Coarse Labels: Fine-Grained Problem Augmentation and Multi-Dimensional Feedback for Emotional Support Conversationaccepted
- Beyond Content: How Grammatical Gender Shapes Visual Representation in Text-to-Image Modelsaccepted
- Beyond Contrastive Learning: Synthetic Data Enables List-wise Training with Multiple Levels of Relevanceaccepted
- Beyond Correctness: Confidence-Aware Reward Modeling for Enhancing Large Language Model Reasoningaccepted
- Beyond Demographics: Enhancing Cultural Value Survey Simulation with Multi-Stage Personality-Driven Cognitive Reasoningaccepted
- Beyond Demonstrations: Dynamic Vector Construction from Latent Representationsaccepted
- Beyond Distribution: Investigating Language Models’ Understanding of Sino-Korean Morphemesaccepted
- Beyond Fixed-Length Calibration for Post-Training Compression of LLMsaccepted
- Beyond Function-Level Search: Repository-Aware Dual-Encoder Code Retrieval with Adversarial Verificationaccepted
- Beyond Guilt: Legal Judgment Prediction with Trichotomous Reasoningaccepted
- Beyond Hate Speech: NLP’s Challenges and Opportunities in Uncovering Dehumanizing Languageaccepted
- Beyond Human Labels: A Multi-Linguistic Auto-Generated Benchmark for Evaluating Large Language Models on Resume Parsingaccepted
- Beyond Inherent Cognition Biases in LLM-Based Event Forecasting: A Multi-Cognition Agentic Frameworkaccepted
- Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencodersaccepted
- Beyond Linear Steering: Unified Multi-Attribute Control for Language Modelsaccepted
- Beyond Online Sampling: Bridging Offline-to-Online Alignment via Dynamic Data Transformation for LLMsaccepted
- Beyond Outlining: Heterogeneous Recursive Planning for Adaptive Long-form Writing with Language Modelsaccepted
- Beyond Pairwise: Global Zero-shot Temporal Graph Generationaccepted
- Beyond Seen Data: Improving KBQA Generalization Through Schema-Guided Logical Form Generationaccepted
- Beyond Self-Reports: Multi-Observer Agents for Personality Assessment in Large Language Modelsaccepted
- Beyond Single Frames: Can LMMs Comprehend Implicit Narratives in Comic Strip?accepted
- Beyond Single-User Dialogue: Assessing Multi-User Dialogue State Tracking Capabilities of Large Language Modelsaccepted
- Beyond Spurious Signals: Debiasing Multimodal Large Language Models via Counterfactual Inference and Adaptive Expert Routingaccepted
- Beyond Static Testbeds: An Interaction-Centric Agent Simulation Platform for Dynamic Recommender Systemsaccepted
- Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Directionaccepted
- Beyond Task-Oriented and Chitchat Dialogues: Proactive and Transition-Aware Conversational Agentsaccepted
- Beyond Text: Unveiling Privacy Vulnerabilities in Multi-modal Retrieval-Augmented Generationaccepted
- Beyond WER: Probing Whisper’s Sub‐token Decoder Across Diverse Language Resource Levelsaccepted
- Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoningaccepted
- Beyond the Leaderboard: Understanding Performance Disparities in Large Language Models via Model Diffingaccepted
- Beyond the Scientific Document: A Citation-Aware Multi-Granular Summarization Approach with Heterogeneous Graphsaccepted
- Beyond the Score: Uncertainty-Calibrated LLMs for Automated Essay Assessmentaccepted
- Beyond the Surface: A Solution-Aware Retrieval Model for Competition-level Code Generationaccepted
- Beyond the Surface: Measuring Self-Preference in LLM Judgmentsaccepted
- Beyond the Textual: Generating Coherent Visual Options for MCQsaccepted
- Bhaasha, Bhāṣā, Zaban: A Survey for Low-Resourced Languages in South Asia – Current Stage and Challengesaccepted
- BiMax: Bidirectional MaxSim Score for Document-Level Alignmentaccepted
- BiMediX2 : Bio-Medical EXpert LMM for Diverse Medical Modalitiesaccepted
- Bias Analysis and Mitigation through Protected Attribute Detection and Regard Classificationaccepted
- Bias Beware: The Impact of Cognitive Biases on LLM-Driven Product Recommendationsaccepted
- Bias Mitigation or Cultural Commonsense? Evaluating LLMs with a Japanese Datasetaccepted
- Bias after Prompting: Persistent Discrimination in Large Language Modelsaccepted
- BiasFilter: An Inference-Time Debiasing Framework for Large Language Modelsaccepted
- Biased Tales: Cultural and Topic Bias in Generating Children’s Storiesaccepted
- Biology-Instructions: A Dataset and Benchmark for Multi-Omics Sequence Understanding Capability of Large Language Modelsaccepted
- Bit-Flip Error Resilience in LLMs: A Comprehensive Analysis and Defense Frameworkaccepted
- Bitune: Leveraging Bidirectional Attention to Improve Decoder-Only LLMsaccepted
- Blind Men and the Elephant: Diverse Perspectives on Gender Stereotypes in Benchmark Datasetsaccepted
- Bold Claims or Self-Doubt? Factuality Hallucination Type Detection via Belief Stateaccepted
- Boosting Data Utilization for Multilingual Dense Retrievalaccepted
- Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language Modelsaccepted
- Both Text and Images Leaked! A Systematic Analysis of Data Contamination in Multimodal LLMaccepted
- Boundary Matters: Leveraging Structured Text Plots for Long Text Outline Generationaccepted
- BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain Tasksaccepted
- BrainLoc: Brain Signal-Based Object Detection with Multi-modal Alignmentaccepted
- Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMsaccepted
- Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction Amplificationaccepted
- Breaking Bad Tokens: Detoxification of LLMs Using Sparse Autoencodersaccepted
- Breaking Token Into Concepts: Exploring Extreme Compression in Token Representation Via Compositional Shared Semanticsaccepted
- Breaking the Attention Trap in Code LLMs: A Rejection Sampling Approach to Enhance Code Execution Predictionaccepted
- Breaking the Noise Barrier: LLM-Guided Semantic Filtering and Enhancement for Multi-Modal Entity Alignmentaccepted
- Breaking the Reviewer: Assessing the Vulnerability of Large Language Models in Automated Peer Review Under Textual Adversarial Attacksaccepted
- Bridging External and Parametric Knowledge: Mitigating Hallucination of LLMs with Shared-Private Semantic Synergy in Dual-Stream Knowledgeaccepted
- Bridging Semantic and Modality Gaps in Zero-Shot Captioning via Retrieval from Synthetic Dataaccepted
- Bridging the Capability Gap: Joint Alignment Tuning for Harmonizing LLM-based Multi-Agent Systemsaccepted
- Bridging the Creativity Understanding Gap: Small-Scale Human Alignment Enables Expert-Level Humor Ranking in LLMsaccepted
- Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoningaccepted
- Bridging the Editing Gap in LLMs: FineEdit for Precise and Targeted Text Modificationsaccepted
- Bridging the Gap Between Molecule and Textual Descriptions via Substructure-aware Alignmentaccepted
- Bringing Pedagogy into Focus: Evaluating Virtual Teaching Assistants’ Question-Answering in Asynchronous Learning Environmentsaccepted
- Brittle Minds, Fixable Activations: Understanding Belief Representations in Language Modelsaccepted
- Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparencyaccepted
- Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systemsaccepted
- C3: A Bilingual Benchmark for Spoken Dialogue Models Exploring Challenges in Complex Conversationsaccepted
- CAARMA: Class Augmentation with Adversarial Mixup Regularizationaccepted
- CAC-CoT: Connector-Aware Compact Chain-of-Thought for Efficient Reasoning Data Synthesis Across Dual-System Cognitive Tasksaccepted
- CAFE: Retrieval Head-based Coarse-to-Fine Information Seeking to Enhance Multi-Document QA Capabilityaccepted
- CAIR: Counterfactual-based Agent Influence Ranker for Agentic AI Workflowsaccepted
- CANDY: Benchmarking LLMs’ Limitations and Assistive Potential in Chinese Misinformation Fact-Checkingaccepted
- CAPE: Context-Aware Personality Evaluation Framework for Large Language Modelsaccepted
- CARD: Cross-modal Agent Framework for Generative and Editable Residential Designaccepted
- CARE: A Disagreement Detection Framework with Concept Alignment and Reasoning Enhancementaccepted
- CARE: Multilingual Human Preference Learning for Cultural Awarenessaccepted
- CARFT: Boosting LLM Reasoning via Contrastive Learning with Annotated Chain-of-Thought-based Reinforced Fine-Tuningaccepted
- CARMA: Enhanced Compositionality in LLMs via Advanced Regularisation and Mutual Information Alignmentaccepted
- CARVQ: Corrective Adaptor with Group Residual Vector Quantization for LLM Embedding Compressionaccepted
- CAT: Causal Attention Tuning For Injecting Fine-grained Causal Knowledge into Large Language Modelsaccepted
- CATCH: A Novel Data Synthesis Framework for High Therapy Fidelity and Memory-Driven Planning Chain of Thought in AI Counselingaccepted
- CAVE : Detecting and Explaining Commonsense Anomalies in Visual Environmentsaccepted
- CBP-Tuning: Efficient Local Customization for Black-box Large Language Modelsaccepted
- CCG: Rare-Label Prediction via Neural SEM–Driven Causal Gameaccepted
- CCL-XCoT: An Efficient Cross-Lingual Knowledge Transfer Method for Mitigating Hallucination Generationaccepted
- CCQA: Generating Question from Solution Can Improve Inference-Time Reasoning in SLMsaccepted
- CDT: A Comprehensive Capability Framework for Large Language Models Across Cognition, Domain, and Taskaccepted
- CEMTM: Contextual Embedding-based Multimodal Topic Modelingaccepted
- CESRec: Constructing Pseudo Interactions for Sequential Recommendation via Conversational Feedbackaccepted
- CHENGYU-BENCH: Benchmarking Large Language Models for Chinese Idiom Understanding and Useaccepted
- CHURRO: Making History Readable with an Open-Weight Large Vision-Language Model for High-Accuracy, Low-Cost Historical Text Recognitionaccepted
- CIE: Controlling Language Model Text Generations Using Continuous Signalsaccepted
- CIFLEX: Contextual Instruction Flow for Sub-task Execution in Multi-Turn Interactions with a Single On-Device LLMaccepted
- CIKT: A Collaborative and Iterative Knowledge Tracing Framework with Large Language Modelsaccepted
- CIVET: Systematic Evaluation of Understanding in VLMsaccepted
- CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers?accepted
- CLEAR: A Clinically Grounded Tabular Framework for Radiology Report Evaluationaccepted
- CLEAR: A Comprehensive Linguistic Evaluation of Argument Rewriting by Large Language Modelsaccepted
- CLEAR: A Framework Enabling Large Language Models to Discern Confusing Legal Paragraphsaccepted
- CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcyclingaccepted
- CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcyclingaccepted
- CLLMate: A Multimodal Benchmark for Weather and Climate Events Forecastingaccepted
- CLMTracing: Black-box User-level Watermarking for Code Language Model Tracingaccepted
- CLaw: Benchmarking Chinese Legal Knowledge in Large Language Models - A Fine-grained Corpus and Reasoning Analysisaccepted
- CM-Align: Consistency-based Multilingual Alignment for Large Language Modelsaccepted
- CMHG: A Dataset and Benchmark for Headline Generation of Minority Languages in Chinaaccepted
- CMT-Eval: A Novel Chinese Multi-turn Dialogue Evaluation Dataset Addressing Real-world Conversational Challengesaccepted
- CMedCalc-Bench: A Fine-Grained Benchmark for Chinese Medical Calculations in LLMaccepted
- COAS2W: A Chinese Older-Adults Spoken-to-Written Transformation Corpus with Context Awarenessaccepted
- COCO-Tree: Compositional Hierarchical Concept Trees for Enhanced Reasoning in Vision-Language Modelsaccepted
- CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillationaccepted
- COLA: Collaborative Multi-Agent Framework with Dynamic Task Scheduling for GUI Automationaccepted
- COM-BOM: Bayesian Exemplar Search for Efficiently Exploring the Accuracy-Calibration Pareto Frontieraccepted
- COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixingaccepted
- COUNTDOWN: Contextually Sparse Activation Filtering Out Unnecessary Weights in Down Projectionaccepted
- CPO: Addressing Reward Ambiguity in Role-playing Dialogue via Comparative Policy Optimizationaccepted
- CR4-NarrEmote: An Open Vocabulary Dataset of Narrative Emotions Derived Using Citizen Scienceaccepted
- CREPE: Rapid Chest X-ray Report Evaluation by Predicting Multi-category Error Countsaccepted
- CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenariosaccepted
- CROP: Contextual Region-Oriented Visual Token Pruningaccepted
- CTCC: A Robust and Stealthy Fingerprinting Framework for Large Language Models via Cross-Turn Contextual Correlation Backdooraccepted
- CURE: Controlled Unlearning for Robust Embeddings — Mitigating Conceptual Shortcuts in Pre-Trained Language Modelsaccepted
- CYCLE-INSTRUCT: Fully Seed-Free Instruction Tuning via Dual Self-Training and Cycle Consistencyaccepted
- CaKE: Circuit-aware Editing Enables Generalizable Knowledge Learnersaccepted
- CaMMT: Benchmarking Culturally Aware Multimodal Machine Translationaccepted
- CaTER: A Framework for Context-aware Topology Entity Retrieval Contrastive Learning in End-to-End Task-Oriented Dialogue Systemsaccepted
- Cache Saver: A Modular Framework for Efficient, Affordable, and Reproducible LLM Inferenceaccepted
- Cache-Efficient Posterior Sampling for Reinforcement Learning with LLM-Derived Priors Across Discrete and Continuous Domainsaccepted
- Cache-of-Thought: Master-Apprentice Framework for Cost-Effective Vision Language Model Reasoningaccepted
- Cacheback: Speculative Decoding With Nothing But Cacheaccepted
- Calibrating LLM Confidence by Probing Perturbed Representation Stabilityaccepted
- Calibrating LLMs for Text-to-SQL Parsing by Leveraging Sub-clause Frequenciesaccepted
- Calibrating Pseudo-Labeling with Class Distribution for Semi-supervised Text Classificationaccepted
- Calibrating Verbal Uncertainty as a Linear Feature to Reduce Hallucinationsaccepted
- Calibration Across Layers: Understanding Calibration Evolution in LLMsaccepted
- CalligraphicOCR for Chinese Calligraphy Recognitionaccepted
- Can Code-Switched Texts Activate a Knowledge Switch in LLMs? A Case Study on English-Korean Code-Switchingaccepted
- Can Federated Learning Safeguard Private Data in LLM Training? Vulnerabilities, Attacks, and Defense Evaluationaccepted
- Can GRPO Boost Complex Multimodal Table Understanding?accepted
- Can LLM Agents Maintain a Persona in Discourse?accepted
- Can LLMs Be Efficient Predictors of Conversational Derailment?accepted
- Can LLMs Explain Themselves Counterfactually?accepted
- Can LLMs Express Personality Across Cultures? Introducing CulturalPersonas for Evaluating Trait Alignmentaccepted
- Can LLMs Extract Frame-Semantic Arguments?accepted
- Can LLMs Find a Needle in a Haystack? A Look at Anomaly Detection Language Modelingaccepted
- Can LLMs Generate and Solve Linguistic Olympiad Puzzles?accepted
- Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environmentsaccepted
- Can LLMs Judge Debates? Evaluating Non-Linear Reasoning via Argumentation Theory Semanticsaccepted
- Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computationaccepted
- Can LLMs Truly Plan? A Comprehensive Evaluation of Planning Capabilitiesaccepted
- Can LLMs be Good Graph Judge for Knowledge Graph Construction?accepted
- Can LLMs be Literary Companions?: Analysing LLMs on Bengali Figures of Speech Identificationaccepted
- Can LLMs simulate the same correct solutions to free-response math problems as real students?accepted
- Can Language Models Follow Multiple Turns of Entangled Instructions?accepted
- Can Large Language Models Act as Ensembler for Multi-GNNs?accepted
- Can Large Language Models Be Good Language Teachers?accepted
- Can Large Language Models Identify Implicit Suicidal Ideation? An Empirical Evaluationaccepted
- Can Large Language Models Outperform Non-Experts in Poetry Evaluation? A Comparative Study Using the Consensual Assessment Techniqueaccepted
- Can Large Language Models Personalize Dialogues to Generational Styles?accepted
- Can Large Language Models Tackle Graph Partitioning?accepted
- Can Large Language Models Translate Spoken-Only Languages through International Phonetic Transcription?accepted
- Can Large Language Models Translate Unseen Languages in Underrepresented Scripts?accepted
- Can Large Language Models Unlock Novel Scientific Research Ideas?accepted
- Can Large Language Models Win the International Mathematical Games?accepted
- Can Large Language Models be Effective Online Opinion Miners?accepted
- Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterizationaccepted
- Can Multiple Responses from an LLM Reveal the Sources of Its Uncertainty?accepted
- Can Out-of-Distribution Evaluations Uncover Reliance on Prediction Shortcuts? A Case Study in Question Answeringaccepted
- Can Prompts Rewind Time for LLMs? Evaluating the Effectiveness of Prompted Knowledge Cutoffsaccepted
- Can Role Vectors Affect LLM Behaviour?accepted
- Can VLMs Recall Factual Associations From Visual References?accepted
- Can Vision-Language Models Solve Visual Math Equations?accepted
- Can We Edit LLMs for Long-Tail Biomedical Knowledge?accepted
- Can We Steer Reasoning Direction by Thinking Intervention?accepted
- Can You Trick the Grader? Adversarial Persuasion of LLM Judgesaccepted
- Can an Individual Manipulate the Collective Decisions of Multi-Agents?accepted
- Can you SPLICE it together? A Human Curated Benchmark for Probing Visual Reasoning in VLMsaccepted
- Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimizationaccepted
- Capturing Latent Modal Association For Multimodal Entity Alignmentaccepted
- Cardiverse: Harnessing LLMs for Novel Card Game Prototypingaccepted
- Case-Based Decision-Theoretic Decoding with Quality Memoriesaccepted
- Castle: Causal Cascade Updates in Relational Databases with Large Language Modelsaccepted
- Catch Me If You Can? Not Yet: LLMs Still Struggle to Imitate the Implicit Writing Styles of Everyday Authorsaccepted
- Causal Interventions Reveal Shared Structure Across English Filler–Gap Constructionsaccepted
- Causal Representation Learning from Multimodal Clinical Records under Non-Random Modality Missingnessaccepted
- Causal Tree Extraction from Medical Case Reports: A Novel Task for Experts-like Text Comprehensionaccepted
- Causal-LLM: A Unified One-Shot Framework for Prompt- and Data-Driven Causal Graph Discoveryaccepted
- CausalMACE: Causality Empowered Multi-Agents in Minecraft Cooperative Tasksaccepted
- CausalVLBench: Benchmarking Visual Causal Reasoning in Large Vision-Language Modelsaccepted
- Certainty in Uncertainty: Reasoning over Uncertain Knowledge Graphs with Statistical Guaranteesaccepted
- Certified Mitigation of Worst-Case LLM Copyright Infringementaccepted
- Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agentsaccepted
- Chain of Strategy Optimization Makes Large Language Models Better Emotional Supporteraccepted
- Chain-of-Interactions: Multi-step Iterative ICL Framework for Abstractive Task-Oriented Dialogue Summarization of Conversational AI Interactionsaccepted
- Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captionsaccepted
- Chain-of-Thought Matters: Improving Long-Context Language Models with Reasoning Path Supervisionaccepted
- Chain-of-Thought Prompting Obscures Hallucination Cues in Large Language Models: An Empirical Evaluationaccepted
- Challenging the Evaluator: LLM Sycophancy Under User Rebuttalaccepted
- Chameleon LLMs: User Personas Influence Chatbot Personality Shiftsaccepted
- Character is Destiny: Can Persona-assigned Language Models Make Personal Choices?accepted
- CharacterCraft: Bridging the Literature-Reality Dialogue Gap for Practical Role-Playing Agentsaccepted
- Characterizing Positional Bias in Large Language Models: A Multi-Model Evaluation of Prompt Order Effectsaccepted
- Chart2Code53: A Large-Scale Diverse and Complex Dataset for Enhancing Chart-to-Code Generationaccepted
- ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinementaccepted
- ChartM3: A Multi-Stage Code-Driven Pipeline for Constructing Multi-Dimensional and Multi-Step Visual Reasoning Data in Chart Comprehensionaccepted
- ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question Answeringaccepted
- Charting the Landscape of African NLP: Mapping Progress and Shaping the Road Aheadaccepted
- Chat-Driven Text Generation and Interaction for Person Retrievalaccepted
- ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Modelaccepted
- Chatbot To Help Patients Understand Their Healthaccepted
- CheckEval: A reliable LLM-as-a-Judge framework for evaluating text generation using checklistsaccepted
- Child-Directed Language Does Not Consistently Boost Syntax Learning in Language Modelsaccepted
- Chinese Toxic Language Mitigation via Sentiment Polarity Consistent Rewritesaccepted
- Choosing a Model, Shaping a Future: Comparing LLM Perspectives on Sustainability and its Relationship with AIaccepted
- ChronoBias: A Benchmark for Evaluating Time-conditional Group Bias in the Time-sensitive Knowledge of Large Language Modelsaccepted
- Circuit Complexity Bounds for RoPE-based Transformer Architectureaccepted
- CiteBART: Learning to Generate Citations for Local Citation Recommendationaccepted
- CityEQA: A Hierarchical LLM Agent on Embodied Question Answering Benchmark in City Spaceaccepted
- ClaimGen-CN: A Large-scale Chinese Dataset for Legal Claim Generationaccepted
- ClimateViz: A Benchmark for Statistical Reasoning and Fact Verification on Scientific Chartsaccepted
- ClueAnchor: Clue-Anchored Knowledge Reasoning Exploration and Optimization for Retrieval-Augmented Generationaccepted
- ClusterUCB: Efficient Gradient-Based Data Selection for Targeted Fine-Tuning of LLMsaccepted
- Co-Eval: Augmenting LLM-based Evaluation with Machine Metricsaccepted
- Co-Evolving LLMs and Embedding Models via Density-Guided Preference Optimization for Text Clusteringaccepted
- CoAT: Chain-of-Associated-Thoughts Framework for Enhancing Large Language Models Reasoningaccepted
- CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triplesaccepted
- CoBia: Constructed Conversations Can Trigger Otherwise Concealed Societal Biases in LLMsaccepted
- CoCoA: Confidence- and Context-Aware Adaptive Decoding for Resolving Knowledge Conflicts in Large Language Modelsaccepted
- CoEvo: Coevolution of LLM and Retrieval Model for Domain-Specific Information Retrievalaccepted
- CoEx – Co-evolving World-model and Explorationaccepted
- CoLA: Compute-Efficient Pre-Training of LLMs via Low-Rank Activationaccepted
- CoMAT: Chain of Mathematically Annotated Thought Improves Mathematical Reasoningaccepted
- CoMMIT: Coordinated Multimodal Instruction Tuningaccepted
- CoMoE: Contrastive Representation for Mixture-of-Experts in Parameter-Efficient Fine-tuningaccepted
- CoPL: Collaborative Preference Learning for Personalizing LLMsaccepted
- CoRAG: Enhancing Hybrid Retrieval-Augmented Generation through a Cooperative Retriever Architectureaccepted
- CoRanking: Collaborative Ranking with Small and Large Ranking Agentsaccepted
- CoT-RAG: Integrating Chain of Thought and Retrieval-Augmented Generation to Enhance Reasoning in Large Language Modelsaccepted
- CoTD-PO: Chain-of-Thought Distillation with Preference Optimizationaccepted
- CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Modelsaccepted
- CoVoGER: A Multilingual Multitask Benchmark for Speech-to-text Generative Error Correction with Large Language Modelsaccepted
- Coarse-to-Fine Grounded Memory for LLM Agent Planningaccepted
- Code Execution as Grounded Supervision for LLM Reasoningaccepted
- Code Like Humans: A Multi-Agent Solution for Medical Codingaccepted
- Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMsaccepted
- CodeArena: Evaluating and Aligning CodeLLMs on Human Preferenceaccepted
- CodeComplex: Dataset for Worst-Case Time Complexity Predictionaccepted
- CodeContests+: High-Quality Test Case Generation for Competitive Programmingaccepted
- CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languagesaccepted
- CodeRAG: Finding Relevant and Necessary Knowledge for Retrieval-Augmented Repository-Level Code Completionaccepted
- CodeSSM: Towards State Space Models for Code Understandingaccepted
- CogAtom: From Cognitive Atoms to Olympiad-level Mathematical Reasoning in Large Language Modelsaccepted
- CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewardsaccepted
- Cognitive Linguistic Identity Fusion Score (CLIFS): A Scalable Cognition‐Informed Approach to Quantifying Identity Fusion from Textaccepted
- Cognitive-Level Adaptive Generation via Capability-Aware Retrieval and Style Adaptationaccepted
- Coherence of Argumentative Dialogue Snippets: A New Method for Large Scale Evaluation with an Application to Inference Anchoring Theoryaccepted
- Collab-Overcooked: Benchmarking and Evaluating Large Language Models as Collaborative Agentsaccepted
- Collaborative Beam Search: Enhancing LLM Reasoning via Collective Consensusaccepted
- Collaborative Rational Speech Act: Pragmatic Reasoning for Multi-Turn Dialogaccepted
- Columbo: Expanding Abbreviated Column Names for Tabular Data Using Large Language Modelsaccepted
- Combining Constrained and Unconstrained Decoding via Boosting: BoostCD and Its Application to Information Extractionaccepted
- ComicScene154: A Scene Dataset for Comic Analysisaccepted
- CompKBQA: Component-wise Task Decomposition for Knowledge Base Question Answeringaccepted
- Comparing Apples to Oranges: A Dataset & Analysis of LLM Humour Understanding from Traditional Puns to Topical Jokesaccepted
- Comparing Specialised Small and General Large Language Models on Text Classification: 100 Labelled Samples to Achieve Break-Even Performanceaccepted
- Comparing human and LLM politeness strategies in free productionaccepted
- CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Rewardaccepted
- Complex Numerical Reasoning with Numerical Semantic Pre-training Frameworkaccepted
- ComplexTempQA: A 100m Dataset for Complex Temporal Question Answeringaccepted
- Composable Cross-prompt Essay Scoring by Merging Modelsaccepted
- Compositional Generalisation for Explainable Hate Speech Detectionaccepted
- Compositional Translation: A Novel LLM-based Approach for Low-resource Machine Translationaccepted
- Compound AI Systems Optimization: A Survey of Methods, Challenges, and Future Directionsaccepted
- Comprehensive Evaluation on Lexical Normalization: Boundary-Aware Approaches for Unsegmented Languagesaccepted
- Comprehensive and Efficient Distillation for Lightweight Sentiment Analysis Modelsaccepted
- Computational Analysis of Character Development in Holocaust Testimoniesaccepted
- Computational Analysis of Conversation Dynamics through Participant Responsivityaccepted
- ConCISE: Confidence-guided Compression in Step-by-step Efficient Reasoningaccepted
- ConText-LE: Cross-Distribution Generalization for Longitudinal Experiential Data via Narrative-Based LLM Representationsaccepted
- Conan-Embedding-v2: Training an LLM from Scratch for Text Embeddingsaccepted
- Concept-pedia: a Wide-coverage Semantically-annotated Multimodal Datasetaccepted
- ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Modelsaccepted
- CondAmbigQA: A Benchmark and Dataset for Conditional Ambiguous Question Answeringaccepted
- CondenseLM: LLMs-driven Text Dataset Condensation via Reward Matchingaccepted
- Conditional [MASK] Discrete Diffusion Language Modelaccepted
- Confidence-guided Refinement Reasoning for Zero-shot Question Answeringaccepted
- Conflict-Aware Soft Prompting for Retrieval-Augmented Generationaccepted
- Conflicting Needles in a Haystack: How LLMs behave when faced with contradictory informationaccepted
- Conflicts in Texts: Data, Implications and Challengesaccepted
- Confounding Factors in Relating Model Performance to Morphologyaccepted
- Confusion is the Final Barrier: Rethinking Jailbreak Evaluation and Investigating the Real Misuse Threat of LLMsaccepted
- Connecting the Knowledge Dots: Retrieval-augmented Knowledge Connection for Commonsense Reasoningaccepted
- Consistent Discourse-level Temporal Relation Extraction Using Large Language Modelsaccepted
- ConsistentChat: Building Skeleton-Guided Consistent Multi-Turn Dialogues for Large Language Models from Scratchaccepted
- Constrained Non-negative Matrix Factorization for Guided Topic Modeling of Minority Topicsaccepted
- ConstraintLLM: A Neuro-Symbolic Framework for Industrial-Level Constraint Programmingaccepted
- Constructing Your Model’s Value Distinction: Towards LLM Alignment with Anchor Words Tuningaccepted
- Constructions are Revealed in Word Distributionsaccepted
- Context Copying Modulation: The Role of Entropy Neurons in Managing Parametric and Contextual Knowledge Conflictsaccepted
- Context Length Alone Hurts LLM Performance Despite Perfect Retrievalaccepted
- Context Minimization for Resource-Constrained Text Classification: Optimizing Performance-Efficiency Trade-offs through Linguistic Featuresaccepted
- Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learningaccepted
- Context and POS in Action: A Comparative Study of Chinese Homonym Disambiguation in Human and Language Modelsaccepted
- Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document Embeddingsaccepted
- Context-Aware Hierarchical Taxonomy Generation for Scientific Papers via LLM-Guided Multi-Aspect Clusteringaccepted
- Context-Aware Membership Inference Attacks against Pre-trained Large Language Modelsaccepted
- Context-Aware Reasoning On Parametric Knowledge for Inferring Causal Variablesaccepted
- Context-aware Biases for Length Extrapolationaccepted
- Continuous-Time Attention: PDE-Guided Mechanisms for Long-Sequence Transformersaccepted
- Continuously Steering LLMs Sensitivity to Contextual Knowledge with Proxy Modelsaccepted
- Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3Daccepted
- ControlText: Unlocking Controllable Fonts in Multilingual Text Rendering without Font Annotationsaccepted
- Controllable Memorization in LLMs via Weight Pruningaccepted
- Controlled Generation for Private Synthetic Textaccepted
- Controlled Retrieval-augmented Context Evaluation for Long-form RAGaccepted
- Conv-Basis: A New Paradigm for Efficient Attention Inference and Gradient Computation in Transformersaccepted
- ConvSearch-R1: Enhancing Query Reformulation for Conversational Search with Reasoning via Reinforcement Learningaccepted
- Convergence and Divergence of Language Models under Different Random Seedsaccepted
- Conversational Education at Scale: A Multi-LLM Agent Workflow for Procedural Learning and Pedagogic Quality Assessmentaccepted
- Convolutional LoRA Aggregation for Unseen Tasks Adaptationaccepted
- CopySpec: Accelerating LLMs with Speculative Copy-and-Pasteaccepted
- Correct-Detect: Balancing Performance and Ambiguity Through the Lens of Coreference Resolution in LLMsaccepted
- Correlation-Aware Example Selection for In-Context Learning with Nonsymmetric Determinantal Point Processesaccepted
- Corrupted but Not Broken: Understanding and Mitigating the Negative Impacts of Corrupted Data in Visual Instruction Tuningaccepted
- Cost-Optimal Grouped-Query Attention for Long-Context Modelingaccepted
- CourtReasoner: Can LLM Agents Reason Like Judges?accepted
- Crafting Customisable Characters with LLMs: A Persona-Driven Role-Playing Agent Frameworkaccepted
- Creative Preference Optimizationaccepted
- Creativity in LLM-based Multi-Agent Systems: A Surveyaccepted
- Crisp: Cognitive Restructuring of Negative Thoughts through Multi-turn Supportive Dialoguesaccepted
- Cross-Cultural Transfer of Commonsense Reasoning in LLMs: Evidence from the Arab Worldaccepted
- Cross-Document Cross-Lingual NLI via RST-Enhanced Graph Fusion and Interpretability Predictionaccepted
- Cross-MoE: An Efficient Temporal Prediction Framework Integrating Textual Modalityaccepted
- Cross-domain Rumor Detection via Test-Time Adaptation and Large Language Modelsaccepted
- CrossQG: Improving Difficulty-Controllable Question Generation through Consistency Enhancementaccepted
- CrystalICL: Enabling In-Context Learning for Crystal Generationaccepted
- CtrlNews: LLM-based Multi-Agent Controllable News Writing via Knowledge Gravitational Fieldaccepted
- CulturalFrames: Assessing Cultural Expectation Alignment in Text-to-Image Models and Evaluation Metricsaccepted
- Culture Cartography: Mapping the Landscape of Cultural Knowledgeaccepted
- Culture is Everywhere: A Call for Intentionally Cultural Evaluationaccepted
- CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesisaccepted
- Curr-ReFT: Overcoming Training Bottlenecks in Small-scale Vision-Language Models via Curriculum Reinforcement Finetuningaccepted
- Current Semantic-change Quantification Methods Struggle with Discovery in the Wildaccepted
- Curse of Knowledge: Your Guidance and Provided Knowledge are biasing LLM Judges in Complex Evaluationaccepted
- Cut the Deadwood Out: Backdoor Purification via Guided Module Substitutionaccepted
- D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question Decompositionaccepted
- D-RAG: Differentiable Retrieval-Augmented Generation for Knowledge Graph Question Answeringaccepted
- D2CS - Documents Graph Clustering using LLM supervisionaccepted
- DA-Pred: Performance Prediction for Text Summarization under Domain-Shift and Instruct-Tuningaccepted
- DAC: Decomposed Automation Correction for Text-to-SQLaccepted
- DAMON: A Dialogue-Aware MCTS Framework for Jailbreaking Large Language Modelsaccepted
- DAPE-BR: Distance-Aware Positional Encoding for Mitigating Object Hallucination in LVLMsaccepted
- DART: Distilling Autoregressive Reasoning to Silent Thoughtaccepted
- DASA-Trans-STM: Adaptive Efficient Transformer for Short Text Matching using Data Augmentation and Semantic Awarenessaccepted
- DAVIS: Planning Agent with Knowledge Graph-Powered Inner Monologueaccepted
- DB-Explore: Automated Database Exploration and Instruction Synthesis for Text-to-SQLaccepted
- DCIS: Efficient Length Extrapolation of LLMs via Divide-and-Conquer Scaling Factor Searchaccepted
- DCMKC: A Dual Consistency Matching Approach for Multi-hop Question Answering in LLMsaccepted
- DCP: Dual-Cue Pruning for Efficient Large Vision-Language Modelsaccepted
- DCR: Quantifying Data Contamination in LLMs Evaluationaccepted
- DCRM: A Heuristic to Measure Response Pair Quality in Preference Optimizationaccepted
- DDO: Dual-Decision Optimization for LLM-Based Medical Consultation via Multi-Agent Collaborationaccepted
- DEBATE, TRAIN, EVOLVE: Self‐Evolution of Language Model Reasoningaccepted
- DEL-ToM: Inference-Time Scaling for Theory-of-Mind Reasoning via Dynamic Epistemic Logicaccepted
- DELOC: Document Element Localizeraccepted
- DICE: Structured Reasoning in LLMs through SLM-Guided Chain-of-Thought Correctionaccepted
- DICP: Deep In-Context Prompt for Event Causality Identificationaccepted
- DIDS: Domain Impact-aware Data Sampling for Large Language Model Trainingaccepted
- DINT Transformeraccepted
- DIPLomA: Efficient Adaptation of Instructed LLMs to Low-Resource Languages via Post-Training Delta Mergingaccepted
- DISCO Balances the Scales: Adaptive Domain- and Difficulty-Aware Reinforcement Learning on Imbalanced Dataaccepted
- DIWALI - Diversity and Inclusivity aWare cuLture specific Items for India: Dataset and Assessment of LLMs for Cultural Text Adaptation in Indian Contextaccepted
- DLIR: Spherical Adaptation for Cross-Lingual Knowledge Transfer of Sociological Concepts Alignmentaccepted
- DLPO: Towards a Robust, Efficient, and Generalizable Prompt Optimization Framework from a Deep-Learning Perspectiveaccepted
- DLTKG: Denoising Logic-based Temporal Knowledge Graph Reasoningaccepted
- DM-Codec: Distilling Multimodal Representations for Speech Tokenizationaccepted
- DMDTEval: An Evaluation and Analysis of LLMs on Disambiguation in Multi-domain Translationaccepted
- DORM: Preference Data Weights Optimization for Reward Modeling in LLM Alignmentaccepted
- DP-GTR: Differentially Private Prompt Protection via Group Text Rewritingaccepted
- DPED: Multi-Layer Noise Distillation for Privacy-Preserving Text Embeddingsaccepted
- DPF-CM: A Data Processing Framework with Privacy-Preserving Vector Databases for Chinese Medical LLMs Training and Deploymentaccepted
- DRBO: Mitigating the Bottleneck Effect via Dynamic Reward Balancing in Multi-reward LLM Optimizationaccepted
- DRES: Fake news detection by dynamic representation and ensemble selectionaccepted
- DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models’ Understanding on Indian Cultureaccepted
- DS-MHP: Improving Chain-of-Thought through Dynamic Subgraph-Guided Multi-Hop Pathaccepted
- DSCD: Large Language Model Detoxification with Self-Constrained Decodingaccepted
- DSG-MCTS: A Dynamic Strategy-Guided Monte Carlo Tree Search for Diversified Reasoning in Large Language Modelsaccepted
- DSMoE: Matrix-Partitioned Experts with Dynamic Routing for Computation-Efficient Dense LLMsaccepted
- DSVD: Dynamic Self-Verify Decoding for Faithful Generation in Large Language Modelsaccepted
- DTDES-KGE: Dual-Teacher Knowledge Distillation with Distinct Embedding Spaces for Knowledge Graph Embeddingsaccepted
- DaMoC: Efficiently Selecting the Optimal Large Language Model for Fine-tuning Domain Tasks Based on Data and Model Compressionaccepted
- Dagger Behind Smile: Fool LLMs with a Happy Ending Storyaccepted
- Data Augmentation for Maltese NLP using Transliterated and Machine Translated Arabic Dataaccepted
- Data Descriptions from Large Language Models with Influence Estimationaccepted
- Data Doping or True Intelligence? Evaluating the Transferability of Injected Knowledge in LLMsaccepted
- Data Drives Unstable Hierarchical Generalization in LMsaccepted
- Data or Language Supervision: What Makes CLIP Better than DINO?accepted
- Data to Defense: The Role of Curation in Aligning Large Language Models Against Safety Compromiseaccepted
- Data-Efficient Hate Speech Detection via Cross-Lingual Nearest Neighbor Retrieval with Limited Labeled Dataaccepted
- Data-Efficient Selection via Grammatical Complexity in Continual Pre-training of Domain-Specific LLMsaccepted
- Data-scarce Behavior Editing of Language Modelsaccepted
- Database-Augmented Query Representation for Information Retrievalaccepted
- DatawiseAgent: A Notebook-Centric LLM Agent Framework for Adaptive and Robust Data Science Automationaccepted
- Date Fragments: A Hidden Bottleneck of Tokenization for Temporal Reasoningaccepted
- David vs. Goliath: Cost-Efficient Financial QA via Cascaded Multi-Agent Reasoningaccepted
- DeAR: Dual-Stage Document Reranking with Reasoning Agents via LLM Distillationaccepted
- DeCoRe: Decoding by Contrasting Retrieval Heads to Mitigate Hallucinationsaccepted
- DeFT-X: Denoised Sparse Fine-Tuning for Zero-Shot Cross-Lingual Transferaccepted
- DeKeyNLU: Enhancing Natural Language to SQL Generation through Task Decomposition and Keyword Extractionaccepted
- DeMAC: Enhancing Multi-Agent Coordination with Dynamic DAG and Manager-Player Feedbackaccepted
- DebUnc: Improving Large Language Model Agent Communication With Uncertainty Metricsaccepted
- Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluationaccepted
- Debate-to-Detect: Reformulating Misinformation Detection as a Real-World Debate with Large Language Modelsaccepted
- Debating for Better Reasoning in Vision-Language Modelsaccepted
- Debiasing Multilingual LLMs in Cross-lingual Latent Spaceaccepted
- DecisionFlow: Advancing Large Language Model as Principled Decision Makeraccepted
- Decoding Dense Embeddings: Sparse Autoencoders for Interpreting and Discretizing Dense Retrievalaccepted
- Decoding Uncertainty: The Impact of Decoding Strategies for Uncertainty Estimation in Large Language Modelsaccepted
- Decoding in Latent Spaces for Efficient Inference in LLM-based Recommendationaccepted
- Decoding the Rule Book: Extracting Hidden Moderation Criteria from Reddit Communitiesaccepted
- DecoupleSearch: Decouple Planning and Search via Hierarchical Reward Modelingaccepted
- Decoupled Proxy Alignment: Mitigating Language Prior Conflict for Multimodal Alignment in MLLMsaccepted
- DecoupledESC: Enhancing Emotional Support Generation via Strategy-Response Decoupled Preference Optimizationaccepted
- Deep Associations, High Creativity: A Simple yet Effective Metric for Evaluating Large Language Modelsaccepted
- DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environmentsaccepted
- DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuningaccepted
- DeepWell-Adol: A Scalable Expert-Based Dialogue Corpus for Adolescent Positive Mental Health and Wellbeing Promotionaccepted
- Defending against Indirect Prompt Injection by Instruction Detectionaccepted
- Definition Generation for Word Meaning Modeling: Monolingual, Multilingual, and Cross-Lingual Perspectivesaccepted
- Dementia Through Different Eyes: Explainable Modeling of Human and LLM Perceptions for Early Awarenessaccepted
- DemonAgent: Dynamically Encrypted Multi-Backdoor Implantation Attack on LLM-based Agentaccepted
- Demystifying Domain-adaptive Post-training for Financial LLMsaccepted
- Demystifying Multilingual Reasoning in Process Reward Modelingaccepted
- Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfallsaccepted
- Demystifying optimized prompts in language modelsaccepted
- Dense Retrievers Can Fail on Simple Queries: Revealing The Granularity Dilemma of Embeddingsaccepted
- Dependency Parsing-Based Syntactic Enhancement of Relation Extraction in Scientific Textsaccepted
- Deriving Strategic Market Insights with Large Language Models: A Benchmark for Forward Counterfactual Generationaccepted
- DesignCLIP: Multimodal Learning with CLIP for Design Patent Understandingaccepted
- Detecting Continuously Evolving Scam Calls under Limited Annotation: A LLM-Augmented Expert Rule Frameworkaccepted
- Detecting Corpus-Level Knowledge Inconsistencies in Wikipedia with Large Language Modelsaccepted
- Detecting Knowledge Boundary of Vision Large Language Models by Sampling-Based Inferenceaccepted
- Detecting LLM Hallucination Through Layer-wise Information Deficiency: Analysis of Ambiguous Prompts and Unanswerable Questionsaccepted
- Detecting Legal Citations in United Kingdom Court Judgmentsaccepted
- Detecting Stealthy Backdoor Samples based on Intra-class Distance for Large Language Modelsaccepted
- Detoxifying Large Language Models via the Diversity of Toxic Samplesaccepted
- Developing and Utilizing a Large-Scale Cantonese Dataset for Multi-Tasking in Large Language Modelsaccepted
- DiCoRe: Enhancing Zero-shot Event Detection via Divergent-Convergent LLM Reasoningaccepted
- DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoningaccepted
- DiNaM: Disinformation Narrative Mining with Large Language Modelsaccepted
- Diagnosing Memorization in Chain-of-Thought Reasoning, One Token at a Timeaccepted
- Diagnosing Moral Reasoning Acquisition in Language Models: Pragmatics and Generalizationaccepted
- Diagram-Driven Course Questions Generationaccepted
- Dial-In LLM: Human-Aligned LLM-in-the-loop Intent Clustering for Customer Service Dialoguesaccepted
- Dialect-SQL: An Adaptive Framework for Bridging the Dialect Gap in Text-to-SQLaccepted
- Dialectal Toxicity Detection: Evaluating LLM-as-a-Judge Consistency Across Language Varietiesaccepted
- Differentiated Vision: Unveiling Entity-Specific Visual Modality Requirements for Multimodal Knowledge Graphaccepted
- Diffusion vs. Autoregressive Language Models: A Text Embedding Perspectiveaccepted
- DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreakaccepted
- DiplomacyAgent: Do LLMs Balance Interests and Ethical Principles in International Events?accepted
- Dipper: Diversity in Prompts for Producing Large Language Model Ensembles in Reasoning Tasksaccepted
- Direct Judgement Preference Optimizationaccepted
- Direct Value Optimization: Improving Chain-of-Thought Reasoning in LLMs with Refined Valuesaccepted
- DisLoRA: Task-specific Low-Rank Adaptation via Orthogonal Basis from Singular Value Decompositionaccepted
- Disambiguation in Conversational Question Answering in the Era of LLMs and Agents: A Surveyaccepted
- DisastIR: A Comprehensive Information Retrieval Benchmark for Disaster Managementaccepted
- DischargeSim: A Simulation Benchmark for Educational Doctor–Patient Communication at Dischargeaccepted
- DiscoSG: Towards Discourse-Level Text Scene Graph Parsing through Iterative Graph Refinementaccepted
- Discourse Heuristics For Paradoxically Moral Self-Correctionaccepted
- Discourse-Driven Code-Switching: Analyzing the Role of Content and Communicative Function in Spanish-English Bilingual Speechaccepted
- Discovering Semantic Subdimensions through Disentangled Conceptual Representationsaccepted
- Discrepancy Detection at the Data Level: Toward Consistent Multilingual Question Answeringaccepted
- Discrete Minds in a Continuous World: Do Language Models Know Time Passes?accepted
- Discriminating Form and Meaning in Multilingual Models with Minimal-Pair ABX Tasksaccepted
- Discursive Circuits: How Do Language Models Understand Discourse Relations?accepted
- Disentangled Information Bottleneck for Adversarial Text Defenseaccepted
- Disentangling Language Understanding and Reasoning Structures in Cross-lingual Chain-of-Thought Promptingaccepted
- Disentangling Subjectivity and Uncertainty for Hate Speech Annotation and Modeling using Gazeaccepted
- Dissecting Logical Reasoning in LLMs: A Fine-Grained Evaluation and Supervision Studyaccepted
- Dissecting Persona-Driven Reasoning in Language Models via Activation Patchingaccepted
- Distill Visual Chart Reasoning Ability from LLMs to MLLMsaccepted
- Distilling Many-Shot In-Context Learning into a Cheat Sheetaccepted
- Distinguishing fair from unfair compositional generalization tasksaccepted
- Distributed LLM Serving on Consumer-Grade GPUs by Reconciling Computation and Communicationaccepted
- Distribution Prompting: Understanding the Expressivity of Language Models Through the Next-Token Distributions They Can Produceaccepted
- Distributional Surgery for Language Model Activationsaccepted
- DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Modelsaccepted
- DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenesaccepted
- DivScore: Zero-Shot Detection of LLM-Generated Text in Specialized Domainsaccepted
- Diverse Multi-tool Aggregation with Large Language Models for Enhanced Math Reasoningaccepted
- Diverse, not Short: A Length-Controlled Data Selection Strategy for Improving Response Diversity of Language Modelsaccepted
- Divide, Optimize, Merge: Scalable Fine-Grained Generative Optimization for LLM Agentsaccepted
- Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Modelsaccepted
- DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text Generationaccepted
- Do All Autoregressive Transformers Remember Facts the Same Way? A Cross-Architecture Analysis of Recall Mechanismsaccepted
- Do BERT-Like Bidirectional Models Still Perform Better on Text Classification in the Era of LLMs?accepted
- Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluationaccepted
- Do Code Semantics Help? A Comprehensive Study on Execution Trace-Based Information for Code Large Language Modelsaccepted
- Do Influence Functions Work on Large Language Models?accepted
- Do It Yourself (DIY): Modifying Images for Poems in a Zero-Shot Setting Using Weighted Prompt Manipulationaccepted
- Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitionsaccepted
- Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMsaccepted
- Do LLMs Behave as Claimed? Investigating How LLMs Follow Their Own Claims using Counterfactual Questionsaccepted
- Do LLMs Encode Frame Semantics? Evidence from Frame Identificationaccepted
- Do LLMs Know and Understand Domain Conceptual Knowledge?accepted
- Do LLMs Understand Wine Descriptors Across Cultures? A Benchmark for Cultural Adaptations of Wine Reviewsaccepted
- Do LVLMs Know What They Know? A Systematic Study of Knowledge Boundary Perception in LVLMsaccepted
- Do Large Language Models Truly Grasp Addition? A Rule-Focused Diagnostic Using Two-Integer Arithmeticaccepted
- Do Large Language Models Understand Word Senses?accepted
- Do Large Language Models excel in Complex Logical Reasoning with Formal Language?accepted
- Do RAG Systems Really Suffer From Positional Bias?accepted
- Do Slides Help? Multi-modal Context for Automatic Transcription of Conference Talksaccepted
- Do We Know What LLMs Don’t Know? A Study of Consistency in Knowledge Probingaccepted
- Do We Really Need All Those Dimensions? An Intrinsic Evaluation Framework for Compressed Embeddingsaccepted
- Do What? Teaching Vision-Language-Action Models to Reject the Impossibleaccepted
- Do You Know About My Nation? Investigating Multilingual Language Models’ Cultural Literacy Through Factual Knowledgeaccepted
- Doc2Chart: Intent-Driven Zero-Shot Chart Generation from Documentsaccepted
- DocAgent: An Agentic Framework for Multi-Modal Long-Context Document Understandingaccepted
- DocAssistant: Integrating Key-region Reading and Step-wise Reasoning for Robust Document Visual Question Answeringaccepted
- DocMMIR: A Framework for Document Multi-modal Information Retrievalaccepted
- DocReRank: Single-Page Hard Negative Query Generation for Training Multi-Modal RAG Rerankersaccepted
- Does Acceleration Cause Hidden Instability in Vision Language Models? Uncovering Instance-Level Divergence Through a Large-Scale Empirical Studyaccepted
- Does Context Matter? A Prosodic Comparison of English and Spanish in Monolingual and Multilingual Discourse Settingsaccepted
- Does It Run and Is That Enough? Revisiting Text-to-Chart Generation with a Multi-Agent Approachaccepted
- Does Localization Inform Unlearning? A Rigorous Examination of Local Parameter Attribution for Knowledge Unlearning in Language Modelsaccepted
- Does Reasoning Introduce Bias? A Study of Social Bias Evaluation and Mitigation in LLM Reasoningaccepted
- Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models?accepted
- Does quantization affect models’ performance on long-context tasks?accepted
- Domain Pre-training Impact on Representationsaccepted
- DongbaMIE: A Multimodal Information Extraction Dataset for Evaluating Semantic Understanding of Dongba Pictogramsaccepted
- Don’t Sweat the Small Stuff: Segment-Level Meta-Evaluation Based on Pairwise Difference Correlationaccepted
- Don’t Take the Premise for Granted: Evaluating the Premise Critique Ability of Large Language Modelsaccepted
- Doubling Your Data in Minutes: Ultra-fast Tabular Data Generation via LLM-Induced Dependency Graphsaccepted
- Dovetail: A CPU/GPU Heterogeneous Speculative Decoding for LLM inferenceaccepted
- DrAgent: Empowering Large Language Models as Medical Agents for Multi-hop Medical Reasoningaccepted
- DrDiff: Dynamic Routing Diffusion with Hierarchical Attention for Breaking the Efficiency-Quality Trade-offaccepted
- DrFrattn: Directly Learn Adaptive Policy from Attention for Simultaneous Machine Translationaccepted
- DrKGC: Dynamic Subgraph Retrieval-Augmented LLMs for Knowledge Graph Completion across General and Biomedical Domainsaccepted
- Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generationaccepted
- Dream to Chat: Model-based Reinforcement Learning on Dialogues with User Belief Modelingaccepted
- Drift-Adapter: A Practical Approach to Near Zero-Downtime Embedding Model Upgrades in Vector Databasesaccepted
- Drift: Decoding-time Personalized Alignments with Implicit User Preferencesaccepted
- Drivel-ology: Challenging LLMs with Interpreting Nonsense with Depthaccepted
- Droid: A Resource Suite for AI-Generated Code Detectionaccepted
- DroidCall: A Dataset for LLM-powered Android Intent Invocationaccepted
- Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMsaccepted
- Dual-Path Counterfactual Integration for Multimodal Aspect-Based Sentiment Classificationaccepted
- Dual-Path Dynamic Fusion with Learnable Query for Multimodal Sentiment Analysisaccepted
- Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbingaccepted
- DyePack: Provably Flagging Test Set Contamination in LLMs Using Backdoorsaccepted
- Dynamic Collaboration of Multi-Language Models based on Minimal Complete Semantic Unitsaccepted
- Dynamic Energy-Based Contrastive Learning with Multi-Stage Knowledge Verification for Event Causality Identificationaccepted
- Dynamic Evaluation for Oversensitivity in LLMsaccepted
- Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE Adaptationaccepted
- Dynamic Injection of Entity Knowledge into Dense Retrieversaccepted
- Dynamic Jointly Batch Selection for Data Efficient Machine Translation Fine-Tuningaccepted
- Dynamic Model-Bank Test-Time Adaptation for Automatic Speech Recognitionaccepted
- Dynamic Retriever for In-Context Knowledge Editing via Policy Optimizationaccepted
- Dynamic Simulation Framework for Disinformation Dissemination and Correction With Social Botsaccepted
- DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMsaccepted
- DynamicNER: A Dynamic, Multilingual, and Fine-Grained Dataset for LLM-based Named Entity Recognitionaccepted
- Dyve: Thinking Fast and Slow for Dynamic Process Verificationaccepted
- E-Verify: A Paradigm Shift to Scalable Embedding-based Factuality Verificationaccepted
- E2LLM: Encoder Elongated Large Language Models for Long-Context Understanding and Reasoningaccepted
- ECC: An Emotion-Cause Conversation Dataset for Empathy Responseaccepted
- ECO Decoding: Entropy-Based Control for Controllability and Fluency in Controllable Dialogue Generationaccepted
- EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understandingaccepted
- EIFBENCH: Extremely Complex Instruction Following Benchmark for Large Language Modelsaccepted
- EMNLP: Educator-role Moral and Normative Large Language Models Profilingaccepted
- EMO-RL: Emotion-Rule-Based Reinforcement Learning Enhanced Audio-Language Model for Generalized Speech Emotion Recognitionaccepted
- EMO: Embedding Model Distillation via Intra-Model Relation and Optimal Transport Alignmentsaccepted
- EQA-RM: A Generative Embodied Reward Model with Test-time Scalingaccepted
- ESC-Judge: A Framework for Comparing Emotional Support Conversational Agentsaccepted
- ESGenius: Benchmarking LLMs on Environmental, Social, and Governance (ESG) and Sustainability Knowledgeaccepted
- ET-MIER: Entity Type-guided Key Mention Identification and Evidence Retrieval for Document-level Relation Extractionaccepted
- EZ-VC: Easy Zero-shot Any-to-Any Voice Conversionaccepted
- Easy as PIE? Identifying Multi-Word Expressions with LLMsaccepted
- EasyRec: Simple yet Effective Language Models for Recommendationaccepted
- Echoes of Agreement: Argument Driven Sycophancy in Large Language modelsaccepted
- EcoLANG: Efficient and Effective Agent Communication Language Induction for Social Simulationaccepted
- EcoLoRA: Communication-Efficient Federated Fine-Tuning of Large Language Modelsaccepted
- EcoSafeRAG: Efficient Security through Context Analysis in Retrieval-Augmented Generationaccepted
- EcoTune: Token-Efficient Multi-Fidelity Hyperparameter Optimization for Large Language Model Inferenceaccepted
- EditID: Training-Free Editable ID Customization for Text-to-Image Generationaccepted
- Editing Across Languages: A Survey of Multilingual Knowledge Editingaccepted
- EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMsaccepted
- EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture Videosaccepted
- Effective Red-Teaming of Policy-Adherent Agentsaccepted
- Efficient Beam Search for Large Language Models Using Trie-Based Decodingaccepted
- Efficient Compositional Multi-tasking for On-device Large Language Modelsaccepted
- Efficient Context Selection for Long-Context QA: No Tuning, No Iteration, Just Adaptive‐kaccepted
- Efficient Dynamic Clustering-Based Document Compression for Retrieval-Augmented-Generationaccepted
- Efficient Integration of External Knowledge to LLM-based World Models via Retrieval-Augmented Generation and Reinforcement Learningaccepted
- Efficient Latent Semantic Clustering for Scaling Test-Time Computation of LLMsaccepted
- Efficient Layer-wise LLM Fine-tuning for Revision Intention Predictionaccepted
- Efficient Model Development through Fine-tuning Transferaccepted
- Efficient Real-time Refinement of Language Model Text Generationaccepted
- Efficient Unstructured Pruning of Mamba State-Space Models for Resource-Constrained Environmentsaccepted
- EfficientXLang: Towards Improving Token Efficiency Through Cross-Lingual Reasoningaccepted
- Efficiently Editing Mixture-of-Experts Models with Compressed Expertsaccepted
- Efficiently Selecting Response Generation Strategies for Synthetic Data Construction by Self-Aligned Perplexityaccepted
- Eliciting Implicit Acoustic Styles from Open-domain Instructions to Facilitate Fine-grained Controllable Generation of Speechaccepted
- Elucidating Mechanisms of Demographic Bias in LLMs for Healthcareaccepted
- EmByte: Decomposition and Compression Learning for Small yet Private NLPaccepted
- Embedding Domain Knowledge for Large Language Models via Reinforcement Learning from Augmented Generationaccepted
- Embedding-Free RAGaccepted
- Emergent morpho-phonological representations in self-supervised speech modelsaccepted
- EmoAgent: Assessing and Safeguarding Human-AI Interaction for Mental Health Safetyaccepted
- EmoBench-UA: A Benchmark Dataset for Emotion Detection in Ukrainianaccepted
- EmoGist: Efficient In-Context Learning for Visual Emotion Understandingaccepted
- Emotion Transfer with Enhanced Prototype for Unseen Emotion Recognition in Conversationaccepted
- Emphasising Structured Information: Integrating Abstract Meaning Representation into LLMs for Enhanced Open-Domain Dialogue Evaluationaccepted
- Empowering GraphRAG with Knowledge Filtering and Integrationaccepted
- Empowering Math Problem Generation and Reasoning for Large Language Model via Synthetic Data based Continual Learning Frameworkaccepted
- EnAnchored-X2X: English-Anchored Optimization for Many-to-Many Translationaccepted
- EnDive: A Cross-Dialect Benchmark for Fairness and Performance in Large Language Modelsaccepted
- End-to-End Learnable Psychiatric Scale Guided Risky Post Screening for Depression Detection on Social Mediaaccepted
- End-to-End Optimization for Multimodal Retrieval-Augmented Generation via Reward Backpropagationaccepted
- English as Defense Proxy: Mitigating Multilingual Jailbreak via Eliciting English Safety Knowledgeaccepted
- Enhanced Noun-Noun Compound Interpretation through Textual Enrichmentaccepted
- Enhancing Attributed Question Answering using Tailored Progressive Curriculum Learningaccepted
- Enhancing Chain-of-Thought Reasoning via Neuron Activation Differential Analysisaccepted
- Enhancing Chinese Offensive Language Detection with Homophonic Perturbationaccepted
- Enhancing Domain-Specific Encoder Models with LLM-Generated Data: How to Leverage Ontologies, and How to Do Without Themaccepted
- Enhancing Efficiency and Exploration in Reinforcement Learning for LLMsaccepted
- Enhancing Goal-oriented Proactive Dialogue Systems via Dynamic Multi-dimensional Consistency Optimizationaccepted
- Enhancing Hate Speech Classifiers through a Gradient-assisted Counterfactual Text Generation Strategyaccepted
- Enhancing LLM Knowledge Learning through Generalizationaccepted
- Enhancing LLM Language Adaption through Cross-lingual In-Context Pre-trainingaccepted
- Enhancing LLM Text Detection with Retrieved Contexts and Logits Distribution Consistencyaccepted
- Enhancing LLM-Based Persuasion Simulations with Cultural and Speaker-Specific Informationaccepted
- Enhancing LLM-Based Social Bot via an Adversarial Learning Frameworkaccepted
- Enhancing Large Language Model for Knowledge Graph Completion via Structure-Aware Alignment-Tuningaccepted
- Enhancing Large Vision-Language Models with Ultra-Detailed Image Caption Generationaccepted
- Enhancing Logical Reasoning in Language Models via Symbolically-Guided Monte Carlo Process Supervisionaccepted
- Enhancing Model Privacy in Federated Learning with Random Masking and Quantizationaccepted
- Enhancing Multi-Agent Debate System Performance via Confidence Expressionaccepted
- Enhancing Partially Relevant Video Retrieval with Robust Alignment Learningaccepted
- Enhancing RAG Efficiency with Adaptive Context Compressionaccepted
- Enhancing RLHF with Human Gaze Modelingaccepted
- Enhancing Reasoning Abilities of Small LLMs with Cognitive Alignmentaccepted
- Enhancing Recommendation Explanations through User-Centric Refinementaccepted
- Enhancing SQL Table Acquisition with Reverse Engineering for Text-to-SQLaccepted
- Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio Encodersaccepted
- Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generationaccepted
- Enhancing Study-Level Inference from Clinical Trial Papers via Reinforcement Learning-Based Numeric Reasoningaccepted
- Enhancing Time Awareness in Generative Recommendationaccepted
- Enrich-on-Graph: Query-Graph Alignment for Complex Reasoning with LLM Enrichingaccepted
- Enriching Patent Claim Generation with European Patent Datasetaccepted
- Ensembling Prompting Strategies for Zero-Shot Hierarchical Text Classification with Large Language Modelsaccepted
- Entity Profile Generation and Reasoning with LLMs for Entity Alignmentaccepted
- EoT: Evolution of Thoughts for Complex Reasoning Tasksaccepted
- Equal Truth: Rumor Detection with Invariant Group Fairnessaccepted
- EquiBench: Benchmarking Large Language Models’ Reasoning about Program Semantics via Equivalence Checkingaccepted
- Equipping Retrieval-Augmented Large Language Models with Document Structure Awarenessaccepted
- Error Classification of Large Language Models on Math Word Problems: A Dynamically Adaptive Frameworkaccepted
- Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervisionaccepted
- Estimating LLM Consistency: A User Baseline vs Surrogate Metricsaccepted
- Estimating Machine Translation Difficultyaccepted
- EuroGEST: Investigating gender stereotypes in multilingual language modelsaccepted
- Evaluating Automatic Speech Recognition Systems for Korean Meteorological Expertsaccepted
- Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humansaccepted
- Evaluating Cognitive-Behavioral Fixation via Multimodal User Viewing Patterns on Social Mediaaccepted
- Evaluating Compound AI Systems through Behaviors, Not Benchmarksaccepted
- Evaluating Cultural Knowledge and Reasoning in LLMs Through Persian Allusionsaccepted
- Evaluating Evaluation Metrics – The Mirage of Hallucination Detectionaccepted
- Evaluating Fairness in Large Vision-Language Models Across Diverse Demographic Attributes and Promptsaccepted
- Evaluating LLM-Generated Diagrams as Graphsaccepted
- Evaluating Language Translation Models by Playing Telephoneaccepted
- Evaluating Large Language Models for Belief Inference: Mapping Belief Networks at Scaleaccepted
- Evaluating Large Language Models for Cross-Lingual Retrievalaccepted
- Evaluating Large Language Models for Detecting Antisemitismaccepted
- Evaluating NL2SQL via SQL2NLaccepted
- Evaluating Robustness of Large Audio Language Models to Audio Injection: An Empirical Studyaccepted
- Evaluating Spatiotemporal Consistency in Automatically Generated Sewing Instructionsaccepted
- Evaluating Step-by-step Reasoning Traces: A Surveyaccepted
- Evaluating Taxonomy Free Character Role Labeling (TF-CRL) in News Stories using Large Language Modelsaccepted
- Evaluating Test-Time Scaling LLMs for Legal Reasoning: OpenAI o1, DeepSeek-R1, and Beyondaccepted
- Evaluating Text Generation Quality Using Spectral Distances of Surprisalaccepted
- Evaluating Uncertainty Quantification Methods in Argumentative Large Language Modelsaccepted
- Evaluating and Aligning Human Economic Risk Preferences in LLMsaccepted
- Evaluating distillation methods for data-efficient syntax learningaccepted
- Evaluating the Creativity of LLMs in Persian Literary Text Generationaccepted
- Evaluating the Effectiveness and Scalability of LLM-Based Data Augmentation for Retrievalaccepted
- Evaluating the Evaluators: Are readability metrics good measures of readability?accepted
- Evaluating the Robustness and Accuracy of Text Watermarking Under Real-World Cross-Lingual Manipulationsaccepted
- Evaluation and Facilitation of Online Discussions in the LLM Era: A Surveyaccepted
- Evaluation of Text-to-Image Generation from a Creativity Perspectiveaccepted
- EventRelBench: A Comprehensive Benchmark for Evaluating Event Relation Understanding in Large Language Modelsaccepted
- EverTracer: Hunting Stolen Large Language Models via Stealthy and Robust Probabilistic Fingerprintaccepted
- EvolKV: Evolutionary KV Cache Compression for LLM Inferenceaccepted
- Evolution in Simulation: AI-Agent School with Dual Memory for High-Fidelity Educational Dynamicsaccepted
- EvolveSearch: An Iterative Self-Evolving Search Agentaccepted
- Evolving Chinese Spelling Correction with Corrector-Verifier Collaborationaccepted
- Evolving Stances on Reproducibility: A Longitudinal Study of NLP and ML Researchers’ Views and Experience of Reproducibilityaccepted
EMNLP accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.