EMNLP 2025 Accepted Papers
The full list of 3,211 papers accepted at EMNLP 2025 (Conference on Empirical Methods in Natural Language Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
accepted: 3,211
- Examining False Positives under Inference Scaling for Mathematical Reasoningaccepted
- Examining Multilingual Embedding Models Cross-Lingually Through LLM-Generated Adversarial Examplesaccepted
- ExeCoder: Empowering Large Language Models with Executability Representation for Code Translationaccepted
- ExeSQL: Self-Taught Text-to-SQL Models with Execution-Driven Bootstrapping for SQL Dialectsaccepted
- ExpandR: Teaching Dense Retrievers Beyond Queries with LLM Guidanceaccepted
- Expanding before Inferring: Enhancing Factuality in Large Language Models through Premature Layers Interpolationaccepted
- Expectation Preference Optimization: Reliable Preference Estimation for Improving the Reasoning Capability of Large Language Modelsaccepted
- ExpertGenQA: Open-ended QA generation in Specialized Domainsaccepted
- Explainability and Interpretability of Multilingual Large Language Models: A Surveyaccepted
- Explainable Chain-of-Thought Reasoning: An Empirical Analysis on State-Aware Reasoning Dynamicsaccepted
- Explainable Text Classification with LLMs: Enhancing Performance through Dialectical Prompting and Explanation-Guided Trainingaccepted
- Explaining Differences Between Model Pairs in Natural Language through Sample Learningaccepted
- Explaining Length Bias in LLM-Based Preference Evaluationsaccepted
- Explaining novel senses using definition generation with open language modelsaccepted
- Explicit Learning and the LLM in Machine Translationaccepted
- Exploiting Prompt-induced Confidence for Black-Box Attacks on LLMsaccepted
- Exploration-Driven Reinforcement Learning for Expert Routing Improvement in Mixture-of-Experts Language Modelsaccepted
- Exploring Artificial Image Generation for Stance Detectionaccepted
- Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignmentaccepted
- Exploring Changes in Nation Perception with Nationality-Assigned Personas in LLMsaccepted
- Exploring Context Strategies in LLMs for Discourse-Aware Machine Translationaccepted
- Exploring Deductive and Inductive Reasoning Capabilities of Large Language Models in Procedural Planningaccepted
- Exploring Hyperbolic Hierarchical Structure for Multimodal Rumor Detectionaccepted
- Exploring Large Language Models for Detecting Mental Disordersaccepted
- Exploring Model Kinship for Merging Large Language Modelsaccepted
- Exploring Paraphrasing Strategies for CEFR A1-Level Constraints in LLMsaccepted
- Exploring Quality and Diversity in Synthetic Data Generation for Argument Miningaccepted
- Exploring Response Uncertainty in MLLMs: An Empirical Evaluation under Misleading Scenariosaccepted
- Exploring and Controlling Diversity in LLM-Agent Conversationaccepted
- Exploring and Detecting Self-disclosure in Multi-modal posts on Chinese Social Mediaaccepted
- Exploring and Evaluating Multimodal Knowledge Reasoning Consistency of Multimodal Large Language Modelsaccepted
- Exploring morphology-aware tokenization: A case study on Spanish language modelingaccepted
- Exploring the Generalizability of Factual Hallucination Mitigation via Enhancing Precise Knowledge Utilizationaccepted
- Exploring the Hidden Capacity of LLMs for One-Step Text Generationaccepted
- Exploring the Hidden Reasoning Process of Large Language Models by Misleading Themaccepted
- Exploring the Impact of Personality Traits on LLM Bias and Toxicityaccepted
- Exploring the Limitations of Mamba in COPY and CoT Reasoningaccepted
- Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulationaccepted
- Extending Automatic Machine Translation Evaluation to Book-Length Documentsaccepted
- Extracting Conceptual Spaces from LLMs Using Prototype Embeddingsaccepted
- Extracting Linguistic Information from Large Language Models: Syntactic Relations and Derivational Knowledgeaccepted
- Extracting and Combining Abilities For Building Multi-lingual Ability-enhanced Large Language Modelsaccepted
- Extractive Fact Decomposition for Interpretable Natural Language Inference in one Forward Passaccepted
- ExtremeAIGC: Benchmarking LMM Vulnerability to AI-Generated Extremist Contentaccepted
- F2TEval: Human-Aligned Multi-Dimensional Evaluation for Figure-to-Text Taskaccepted
- FACTCHECKMATE: Preemptively Detecting and Mitigating Hallucinations in LMsaccepted
- FAEDKV: Infinite-Window Fourier Transform for Unbiased KV Cache Compressionaccepted
- FANS: Formal Answer Selection for LLM Natural Language Math Reasoning Using Lean4accepted
- FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs’ Responsiveness to Human Feedbackaccepted
- FC-Attack: Jailbreaking Multimodal Large Language Models via Auto-Generated Flowchartsaccepted
- FESTA: Functionally Equivalent Sampling for Trust Assessment of Multimodal LLMsaccepted
- FG-PRM: Fine-grained Hallucination Detection and Mitigation in Language Model Mathematical Reasoningaccepted
- FIER: Fine-Grained and Efficient KV Cache Retrieval for Long-context LLM Inferenceaccepted
- FIRE: Flexible Integration of Data Quality Ratings for Effective Pretrainingaccepted
- FISTAPruner: Layer-wise Post-training Pruning for Large Language Modelsaccepted
- FLAIRR-TS - Forecasting LLM-Agents with Iterative Refinement and Retrieval for Time Seriesaccepted
- FLAMES: Improving LLM Math Reasoning via a Fine-Grained Analysis of the Data Synthesis Pipelineaccepted
- FLARE: Faithful Logic-Aided Reasoning and Explorationaccepted
- FLRC: Fine-grained Low-Rank Compressor for Efficient LLM Inferenceaccepted
- FLUID QA: A Multilingual Benchmark for Figurative Language Usage in Dialogue across English, Chinese, and Koreanaccepted
- FM2DS: Few-Shot Multimodal Multihop Data Synthesis with Knowledge Distillation for Question Answeringaccepted
- FNSCC: Fuzzy Neighborhood-Aware Self-Supervised Contrastive Clustering for Short Textaccepted
- FOSSIL: Harnessing Feedback on Suboptimal Samples for Data-Efficient Generalisation with Imitation Learning for Embodied Vision-and-Language Tasksaccepted
- FSTs vs ICL: Generalisation in LLMs for an under-resourced languageaccepted
- FaMTEB: Massive Text Embedding Benchmark in Persian Languageaccepted
- FaST: Feature-aware Sampling and Tuning for Personalized Preference Alignment with Limited Dataaccepted
- FaStFact: Faster, Stronger Long-Form Factuality Evaluations in LLMsaccepted
- FacLens: Transferable Probe for Foreseeing Non-Factuality in Fact-Seeking Question Answering of Large Language Modelsaccepted
- Facilitating Cognitive Accessibility with LLMs: A Multi-Task Approach to Easy-to-Read Text Generationaccepted
- Facilitating Cross-lingual Transfer of Empathy through Language-independent Latent Diffusion: A Case Study in Chineseaccepted
- Facilitating Long Context Understanding via Supervised Chain-of-Thought Reasoningaccepted
- Fact Verification on Knowledge Graph via Programmatic Graph Reasoningaccepted
- FactReasoner: A Probabilistic Approach to Long-Form Factuality Assessment for Large Language Modelsaccepted
- Facts Fade Fast: Evaluating Memorization of Outdated Medical Knowledge in Large Language Modelsaccepted
- Factuality Beyond Coherence: Evaluating LLM Watermarking Methods for Medical Textsaccepted
- Fair Text-Attributed Graph Representation Learningaccepted
- Fair or Framed? Political Bias in News Articles Generated by LLMsaccepted
- FairCoT: Enhancing Fairness in Text-to-Image Generation via Chain of Thought Reasoning with Multimodal Large Language Modelsaccepted
- FairGen: Controlling Sensitive Attributes for Fair Generations in Diffusion Models via Adaptive Latent Guidanceaccepted
- Fairness in Automatic Speech Recognition Isn’t a One-Size-Fits-Allaccepted
- FaithUn: Toward Faithful Forgetting in Language Models by Investigating the Interconnectedness of Knowledgeaccepted
- FakeSV-VLM: Taming VLM for Detecting Fake Short-Video News via Progressive Mixture-Of-Experts Adapteraccepted
- False Friends Are Not Foes: Investigating Vocabulary Overlap in Multilingual Language Modelsaccepted
- Familiarity-Aware Evidence Compression for Retrieval-Augmented Generationaccepted
- Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMsaccepted
- Fast Quiet-STaR: Thinking Without Thought Tokensaccepted
- Fast, Not Fancy: Rethinking G2P with Rich Data and Statistical Modelsaccepted
- FastCuRL: Curriculum Reinforcement Learning with Stage-wise Context Scaling for Efficient Training R1-like Reasoning Modelsaccepted
- Faster In-Context Learning for LLMs via N-Gram Trie Speculative Decodingaccepted
- Faster and Better LLMs via Latency-Aware Test-Time Scalingaccepted
- Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Modelsaccepted
- FedCoT: Federated Chain-of-Thought Distillation for Large Language Modelsaccepted
- FedMABench: Benchmarking Mobile GUI Agents on Decentralized Heterogeneous User Dataaccepted
- Federated Retrieval-Augmented Generation: A Systematic Mapping Studyaccepted
- Feel the Difference? A Comparative Analysis of Emotional Arcs in Real and LLM-Generated CBT Sessionsaccepted
- Ferret: Faster and Effective Automated Red Teaming with Reward-Based Scoring Techniqueaccepted
- Few-Shot Learning Translation from New Languagesaccepted
- Few-Shot Open-Set Classification via Reasoning-Aware Decompositionaccepted
- FiRST: Finetuning Router-Selective Transformers for Input-Adaptive Latency Reductionaccepted
- FicSim: A Dataset for Multi-Faceted Semantic Similarity in Long-Form Fictionaccepted
- FigEx: Aligned Extraction of Scientific Figures and Captionsaccepted
- FilBench: Can LLMs Understand and Generate Filipino?accepted
- FillerSpeech: Towards Human-Like Text-to-Speech Synthesis with Filler Insertion and Filler Style Controlaccepted
- FinGEAR: Financial Mapping-Guided Enhanced Answer Retrievalaccepted
- FinGrAct: A Framework for FINe-GRrained Evaluation of ACTionability in Explainable Automatic Fact-Checkingaccepted
- FinHEAR: Human Expertise and Adaptive Risk-Aware Temporal Reasoning for Financial Decision-Makingaccepted
- FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answeringaccepted
- FinMTEB: Finance Massive Text Embedding Benchmarkaccepted
- FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial Domainaccepted
- FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domainaccepted
- Financial Risk Relation Identification through Dual-view Adaptationaccepted
- Finding your MUSE: Mining Unexpected Solutions Engineaccepted
- Fine-Tuned Thoughts: Leveraging Chain-of-Thought Reasoning for Industrial Asset Health Monitoringaccepted
- Fine-Tuning Encoder-Decoder Models with Contrastive Learning for In-Context Distractor Generationaccepted
- Fine-tuning LLMs with Cross-Attention-based Weight Decay for Bias Mitigationaccepted
- Finetuning LLMs for Human Behavior Prediction in Social Science Experimentsaccepted
- Fingerprinting LLMs through Survey Item Factor Correlation: A Case Study on Humor Style Questionnaireaccepted
- Firewall Routing: Blocking Leads to Better Hybrid Inference for LLMsaccepted
- FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Gamesaccepted
- Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMsaccepted
- FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantizationaccepted
- Flexible Thinking for Multimodal Emotional Support Conversation via Reinforcement Learningaccepted
- Flexible-length Text Infilling for Discrete Diffusion Modelsaccepted
- Flexibly Utilize Memory for Long-Term Conversation via a Fragment-then-Compose Frameworkaccepted
- FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Modelsaccepted
- FlowMalTrans: Unsupervised Binary Code Translation for Malware Detection Using Flow-Adapter Architectureaccepted
- FoREST: Frame of Reference Evaluation in Spatial Reasoning Tasksaccepted
- Follow the Flow: Fine-grained Flowchart Attribution with Neurosymbolic Agentsaccepted
- Following Length Constraints in Instructionsaccepted
- Following Occam’s Razor: Dynamic Combination of Structured Knowledge for Multi-Hop Question Answering using LLMsaccepted
- Following the Autoregressive Nature of LLM Embeddings via Compression and Alignmentaccepted
- FoodSafeSum: Enabling Natural Language Processing Applications for Food Safety Document Summarization and Analysisaccepted
- Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluationaccepted
- Foot-In-The-Door: A Multi-turn Jailbreak for LLMsaccepted
- For a Fistful of Puns: Evaluating a Puns in Multiword Expressions Identification Algorithm Without Dedicated Datasetaccepted
- ForestCast: Open-Ended Event Forecasting with Semantic News Forestaccepted
- Forewarned is Forearmed: Pre-Synthesizing Jailbreak-like Instructions to Enhance LLM Safety Guardrail to Potential Attacksaccepted
- Forget What You Know about LLMs Evaluations - LLMs are Like a Chameleonaccepted
- Forget for Get: A Lightweight Two-phase Gradient Method for Knowledge Editing in Large Language Modelsaccepted
- Forget the Unneeded: Backdooring Large Language Models via Contrastive-enhanced Machine Unlearningaccepted
- Formalizing Style in Personal Narrativesaccepted
- FormosanBench: Benchmarking Low-Resource Austronesian Languages in the Era of Large Language Modelsaccepted
- FractalLLM: Lossless Self-Speculative Decoding with Layer Embedded Self-Compressionaccepted
- Frame First, Then Extract: A Frame-Semantic Reasoning Pipeline for Zero-Shot Relation Triplet Extractionaccepted
- FrameEOL: Semantic Frame Induction using Causal Language Modelsaccepted
- Frequency & Compositionality in Emergent Communicationaccepted
- Friend or Foe? A Computational Investigation of Semantic False Friends across Romance Languagesaccepted
- FroM: Frobenius Norm-Based Data-Free Adaptive Model Mergingaccepted
- From A and B to A+B: Can Large Language Models Solve Compositional Math Problems?accepted
- From Automation to Autonomy: A Survey on Large Language Models in Scientific Discoveryaccepted
- From Benchmark to Better Embeddings: Leveraging Synonym Substitution to Enhance Multimodal Models in Ukrainianaccepted
- From Capabilities to Performance: Evaluating Key Functional Properties of LLM Architectures in Penetration Testingaccepted
- From Characters to Tokens: Dynamic Grouping with Hierarchical BPEaccepted
- From Charts to Fair Narratives: Uncovering and Mitigating Geo-Economic Biases in Chart-to-Textaccepted
- From Chat Logs to Collective Insights: Aggregative Question Answeringaccepted
- From Confidence to Collapse in LLM Factual Robustnessaccepted
- From Cross-Task Examples to In-Task Prompts: A Graph-Based Pseudo-Labeling Framework for In-context Learningaccepted
- From General Reward to Targeted Reward: Improving Open-ended Long-context Generation Modelsaccepted
- From Generation to Detection: A Multimodal Multi-Task Dataset for Benchmarking Health Misinformationaccepted
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeaccepted
- From Generic Empathy to Personalized Emotional Support: A Self-Evolution Framework for User Preference Alignmentaccepted
- From Ground Trust to Truth: Disparities in Offensive Language Judgments on Contemporary Korean Political Discourseaccepted
- From Grounding to Manipulation: Case Studies of Foundation Model Integration in Embodied Robotic Systemsaccepted
- From Hypothesis to Publication: A Comprehensive Survey of AI-Driven Research Support Systemsaccepted
- From Implicit Exploration to Structured Reasoning: Guideline and Refinement for LLMsaccepted
- From Input Perception to Predictive Insight: Modeling Model Blind Spots Before They Become Errorsaccepted
- From Insight to Exploit: Leveraging LLM Collaboration for Adaptive Adversarial Text Generationaccepted
- From KMMLU-Redux to Pro: A Professional Korean Benchmark Suite for LLM Evaluationaccepted
- From Knowledge to Treatment: Large Language Model Assisted Biomedical Concept Representation for Drug Repurposingaccepted
- From Language to Cognition: How LLMs Outgrow the Human Language Networkaccepted
- From Long to Lean: Performance-aware and Adaptive Chain-of-Thought Compression via Multi-round Refinementaccepted
- From Noise to Clarity: Filtering Real and LLM-Generated Samples for Enhanced Intent Detectionaccepted
- From Parameters to Performance: A Data-Driven Study on LLM Structure and Developmentaccepted
- From Personas to Talks: Revisiting the Impact of Personas on LLM-Synthesized Emotional Support Conversationsaccepted
- From Problem-Solving to Teaching Problem-Solving: Aligning LLMs with Pedagogy using Reinforcement Learningaccepted
- From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Modelsaccepted
- From Remembering to Metacognition: Do Existing Benchmarks Accurately Evaluate LLMs?accepted
- From Schema to State: Zero-Shot Scheme-Only Dialogue State Tracking via Diverse Synthetic Dialogue and Step-by-Step Distillationaccepted
- From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculationsaccepted
- From Shortcuts to Balance: Attribution Analysis of Speech-Text Feature Utilization in Distinguishing Original from Machine-Translated Textsaccepted
- From Surveys to Narratives: Rethinking Cultural Value Adaptation in LLMsaccepted
- From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognitionaccepted
- From Token to Action: State Machine Reasoning to Mitigate Overthinking in Information Retrievalaccepted
- From Tower to Spire: Adding the Speech Modality to a Translation-Specialist LLMaccepted
- From Unaligned to Aligned: Scaling Multilingual LLMs with Multi-Way Parallel Corporaaccepted
- From Understanding to Generation: An Efficient Shortcut for Evaluating Language Modelsaccepted
- From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Testaccepted
- From n-gram to Attention: How Model Architectures Learn and Propagate Bias in Language Modelingaccepted
- From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation modelaccepted
- FuseChat: Knowledge Fusion of Chat Modelsaccepted
- FusionDTI: Fine-grained Binding Discovery with Token-level Fusion for Drug-Target Interactionaccepted
- FuzzAug: Data Augmentation by Coverage-guided Fuzzing for Neural Test Generationaccepted
- Fuzzy Reasoning Chain (FRC): An Innovative Reasoning Framework from Fuzziness to Clarityaccepted
- F²Bench: An Open-ended Fairness Evaluation Benchmark for LLMs with Factuality Considerationsaccepted
- G2: Guided Generation for Enhanced Output Diversity in LLMsaccepted
- GAMIC: Graph-Aligned Molecular In-context Learning for Molecule Analysis via LLMsaccepted
- GAP: a Global Adaptive Pruning Method for Large Language Modelsaccepted
- GASE: Generatively Augmented Sentence Encodingaccepted
- GATEAU: Selecting Influential Samples for Long Context Alignmentaccepted
- GAttention: Gated Attention for the Detection of Abusive Languageaccepted
- GCML: Gradient Coherence Guided Meta-Learning for Cross-Domain Emerging Topic Rumor Detectionaccepted
- GDLLM: A Global Distance-aware Modeling Approach Based on Large Language Models for Event Temporal Relation Extractionaccepted
- GENUINE: Graph Enhanced Multi-level Uncertainty Estimation for Large Language Modelsaccepted
- GER-LLM: Efficient and Effective Geospatial Entity Resolution with Large Language Modelaccepted
- GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?accepted
- GLProtein: Global-and-Local Structure Aware Protein Representation Learningaccepted
- GRAD: Generative Retrieval-Aligned Demonstration Sampler for Efficient Few-Shot Reasoningaccepted
- GRADA: Graph-based Reranking against Adversarial Documents Attackaccepted
- GRADE: Generating multi-hop QA and fine-gRAined Difficulty matrix for RAG Evaluationaccepted
- GRAID: Synthetic Data Generation with Geometric Constraints and Multi-Agentic Reflection for Harmful Content Detectionaccepted
- GRASP: Replace Redundant Layers with Adaptive Singular Parameters for Efficient Model Compressionaccepted
- GRIL: Knowledge Graph Retrieval-Integrated Learning with Large Language Modelsaccepted
- GRIT: Guided Relational Integration for Efficient Multi-Table Understandingaccepted
- GRPO-Guided Modality Selection Enhanced LoRA-Tuned LLMs for Multimodal Emotion Recognitionaccepted
- GRPO-LEAD: A Difficulty-Aware Reinforcement Learning Approach for Concise Mathematical Reasoning in Language Modelsaccepted
- GRV-KBQA: A Three-Stage Framework for Knowledge Base Question Answering with Decoupled Logical Structure, Semantic Grounding and Structure-Aware Validationaccepted
- GTA: Supervised-Guided Reinforcement Learning for Text Classification with Large Language Modelsaccepted
- GUARD: Glocal Uncertainty-Aware Robust Decoding for Effective and Efficient Open-Ended Text Generationaccepted
- GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous Explorationaccepted
- Gamma-Guard: Lightweight Residual Adapters for Robust Guardrails in Large Language Modelsaccepted
- GeLoRA: Geometric Adaptive Ranks For Efficient LoRA Fine-tuningaccepted
- GenLink: Generation-Driven Schema-Linking via Multi-Model Learning for Text-to-SQLaccepted
- GenPTQ: Green Post-Training Quantization for Large-Scale ASR Models with Mixed-Precision Bit Allocationaccepted
- GenPilot: A Multi-Agent System for Test-Time Prompt Optimization in Image Generationaccepted
- GenPoE: Generative Passage-level Mixture of Experts for Knowledge Enhancement of LLMsaccepted
- Generation-Augmented Retrieval: Rethinking the Role of Large Language Models in Zero-Shot Relation Extractionaccepted
- Generative Annotation for ASR Named Entity Correctionaccepted
- Generative or Discriminative? Revisiting Text Classification in the Era of Transformersaccepted
- Generator-Assistant Stepwise Rollback Framework for Large Language Model Agentaccepted
- Genre Matters: How Text Types Interact with Decoding Strategies and Lexical Predictors in Shaping Reading Behavioraccepted
- GeoChain: Multimodal Chain-of-Thought for Geographic Reasoningaccepted
- GeoDANO: Geometric VLM with Domain Agnostic Vision Encoderaccepted
- GeoEdit: Geometric Knowledge Editing for Large Language Modelsaccepted
- GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric Reasoningaccepted
- Glider: Global and Local Instruction-Driven Expert Routeraccepted
- Glitter: A Multi-Sentence, Multi-Reference Benchmark for Gender-Fair German Machine Translationaccepted
- GmSLM : Generative Marmoset Spoken Language Modelingaccepted
- Golden Touchstone: A Comprehensive Bilingual Benchmark for Evaluating Financial Large Language Modelsaccepted
- Good Intentions Beyond ACL: Who Does NLP for Social Good, and Where?accepted
- Governance in Motion: Co-evolution of Constitutions and AI models for Scalable Safetyaccepted
- GraDaSE: Graph-Based Dataset Search with Examplesaccepted
- Graceful Forgetting in Generative Language Modelsaccepted
- Graders Should Cheat: Privileged Information Enables Expert-Level Automated Evaluationsaccepted
- Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrievalaccepted
- Grammar Pruning: Enabling Low-Latency Zero-Shot Task-Oriented Language Models for Edge AIaccepted
- Graph-Based Multi-Trait Essay Scoringaccepted
- Graph-Guided Textual Explanation Generation Frameworkaccepted
- Graph-R1: Incentivizing the Zero-Shot Graph Learning Capability in LLMs via Explicit Reasoningaccepted
- Graph-Reward-SQL: Execution-Free Reinforcement Learning for Text-to-SQL via Graph Matching and Stepwise Rewardaccepted
- GraphAgent: Agentic Graph Language Assistantaccepted
- GraphCheck: Multipath Fact-Checking with Entity-Relationship Graphsaccepted
- GraphKV: Breaking the Static Selection Paradigm with Graph-Based KV Cache Evictionaccepted
- GreekBarBench: A Challenging Benchmark for Free-Text Legal Reasoning and Citationsaccepted
- Grounded Semantic Role Labelling from Synthetic Multimodal Data for Situated Robot Commandsaccepted
- Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Modelsaccepted
- Grounding Multilingual Multimodal LLMs With Cultural Knowledgeaccepted
- Group-Aware Reinforcement Learning for Output Diversity in Large Language Modelsaccepted
- Group-SAE: Efficient Training of Sparse Autoencoders for Large Language Models via Layer Groupsaccepted
- Grouping Entities with Shared Properties using Multi-Facet Prompting and Property Embeddingsaccepted
- Guaranteed Guess: A Language Modeling Approach for CISC-to-RISC Transpilation with Testing Guaranteesaccepted
- Guess What I am Thinking: A Benchmark for Inner Thought Reasoning of Role-Playing Language Agentsaccepted
- GuessingGame: Measuring the Informativeness of Open-Ended Questions in Large Language Modelsaccepted
- GuiLoMo: Allocating Experts and Ranks for LoRA-MoE via Bilevel Optimization with GuidedSelection Vectorsaccepted
- Guiding Large Language Models for Biomedical Entity Linking via Restrictive and Contrastive Decodingaccepted
- HARE: an entity and relation centric evaluation framework for histopathology reportsaccepted
- HATECAT-TR: A Hate Speech Span Detection and Categorization Dataset for Turkishaccepted
- HAWK: Highlighting Entity-aware Knowledge for Alleviating Information Sparsity in Long Contextsaccepted
- HD-PiSSA: High-Rank Distributed Orthogonal Adaptationaccepted
- HDiff: Confidence-Guided Denoising Diffusion for Robust Hyper-relational Link Predictionaccepted
- HEAL: A Hypothesis-Based Preference-Aware Analysis Frameworkaccepted
- HEAL: An Empirical Study on Hallucinations in Embodied Agents Driven by Large Language Modelsaccepted
- HEAL: Hybrid Enhancement with LLM-based Agents for Text-attributed Hypergraph Self-supervised Representation Learningaccepted
- HELENE: Hessian Layer-wise Clipping and Gradient Annealing for Accelerating Fine-tuning LLM with Zeroth-order Optimizationaccepted
- HESEIA: A community-based dataset for evaluating social biases in large language models, co-designed in real school settings in Latin Americaaccepted
- HGAdapter: Hypergraph-based Adapters in Language Models for Code Summarization and Clone Detectionaccepted
- HICode: Hierarchical Inductive Coding with LLMsaccepted
- HIRAG: Hierarchical-Thought Instruction-Tuning Retrieval-Augmented Generationaccepted
- HMCL: Task-Optimal Text Representation Adaptation through Hierarchical Contrastive Learningaccepted
- HMoE: Heterogeneous Mixture of Experts for Language Modelingaccepted
- HS-STaR: Hierarchical Sampling for Self-Taught Reasoners via Difficulty Estimation and Budget Reallocationaccepted
- HVGuard: Utilizing Multimodal Large Language Models for Hateful Video Detectionaccepted
- HYDRA: A Multi-Head Encoder-only Architecture for Hierarchical Text Classificationaccepted
- Hallucination Detection in LLMs Using Spectral Features of Attention Mapsaccepted
- Hallucination Detection in Structured Query Generation via LLM Self-Debatingaccepted
- Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreationaccepted
- HapticCap: A Multimodal Dataset and Task for Understanding User Experience of Vibration Haptic Signalsaccepted
- Hard Negatives, Hard Lessons: Revisiting Training Data Quality for Robust Information Retrieval with LLMsaccepted
- Hardware-Aware Parallel Prompt Decoding for Memory-Efficient Acceleration of LLM Inferenceaccepted
- Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encodingaccepted
- Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Modelsaccepted
- HealthCards: Exploring Text-to-Image Generation as Visual Aids for Healthcare Knowledge Democratizing and Educationaccepted
- HebID: Detecting Social Identities in Hebrew-language Political Textaccepted
- HetGCoT: Heterogeneous Graph-Enhanced Chain-of-Thought LLM Reasoning for Academic Question Answeringaccepted
- HiMATE: A Hierarchical Multi-Agent Framework for Machine Translation Evaluationaccepted
- Hidden Ghost Hand: Unveiling Backdoor Vulnerabilities in MLLM-Powered Mobile GUI Agentsaccepted
- Hidden in Plain Sight: Reasoning in Underspecified and Misspecified Scenarios for Multimodal LLMsaccepted
- HierPrompt: Zero-Shot Hierarchical Text Classification with LLM-Enhanced Prototypesaccepted
- Hierarchical Bracketing Encodings Work for Dependency Graphsaccepted
- Hierarchical Reward Modeling for Fault Localization in Large Code Repositoriesaccepted
- HighMATH: Evaluating Math Reasoning of Large Language Models in Breadth and Depthaccepted
- HomoGraphAdapter: A Homogeneous Graph Neural Network as an Effective Adapter for Vision-Language Modelsaccepted
- HookMoE: A learnable performance compensation strategy of Mixture-of-Experts for LLM inference accelerationaccepted
- Hope vs. Hate: Understanding User Interactions with LGBTQ+ News Content in Mainstream US News Media through the Lens of Hope Speechaccepted
- Hopscotch: Discovering and Skipping Redundancies in Language Modelsaccepted
- How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on tau-benchaccepted
- How Diversely Can Language Models Solve Problems? Exploring the Algorithmic Diversity of Model-Generated Codeaccepted
- How Do Large Language Models Perform on PDE Discovery: A Coarse-to-fine Perspectiveaccepted
- How Do Large Vision-Language Models See Text in Image? Unveiling the Distinctive Role of OCR Headsaccepted
- How Do Social Bots Participate in Misinformation Spread? A Comprehensive Dataset and Analysisaccepted
- How Does Cognitive Bias Affect Large Language Models? A Case Study on the Anchoring Effect in Price Negotiation Simulationsaccepted
- How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysisaccepted
- How Does Knowledge Selection Help Retrieval Augmented Generation?accepted
- How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparisonaccepted
- How Good are LLM-based Rerankers? An Empirical Analysis of State-of-the-Art Reranking Modelsaccepted
- How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmarkaccepted
- How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigationaccepted
- How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM Hallucinationaccepted
- How Much Do Large Language Models Know about Human Motion? A Case Study in 3D Avatar Controlaccepted
- How Persuasive Is Your Context?accepted
- How Private are Language Models in Abstractive Summarization?accepted
- How Real Are Synthetic Therapy Conversations? Evaluating Fidelity in Prolonged Exposure Dialoguesaccepted
- How Reliable is Multilingual LLM-as-a-Judge?accepted
- How Sampling Affects the Detectability of Machine-written texts: A Comprehensive Studyaccepted
- How Sememic Components Can Benefit Link Prediction for Lexico-Semantic Knowledge Graphs?accepted
- How Well Can Reasoning Models Identify and Recover from Unhelpful Thoughts?accepted
- How a Bilingual LM Becomes Bilingual: Tracing Internal Representations with Sparse Autoencodersaccepted
- How do Language Models Generate Slang: A Systematic Comparison between Human and Machine-Generated Slang Usagesaccepted
- How do Language Models Reshape Entity Alignment? A Survey of LM-Driven EA Methods: Advances, Benchmarks, and Futureaccepted
- How do autoregressive transformers solve full addition?accepted
- How to Generalize the Detection of AI-Generated Text: Confounding Neuronsaccepted
- How to Make Large Language Models Generate 100% Valid Molecules?accepted
- How to Protect Yourself from 5G Radiation? Investigating LLM Responses to Implicit Misinformationaccepted
- How to inject knowledge efficiently? Knowledge Infusion Scaling Law for Pre-training Large Language Modelsaccepted
- Human-Inspired Obfuscation for Model Unlearning: Local and Global Strategies with Hyperbolic Representationsaccepted
- Humanity’s Last Code Exam: Can Advanced LLMs Conquer Human’s Hardest Code Competition?accepted
- Humanizing Machines: Rethinking LLM Anthropomorphism Through a Multi-Level Framework of Designaccepted
- Humans Hallucinate Too: Language Models Identify and Correct Subjective Annotation Errors With Label-in-a-Haystack Promptsaccepted
- Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comicsaccepted
- HydraOpt: Navigating the Efficiency-Performance Trade-off of Adapter Mergingaccepted
- HydraRAG: Structured Cross-Source Enhanced Large Language Model Reasoningaccepted
- HypER: Literature-grounded Hypothesis Generation and Distillation with Provenanceaccepted
- HyperKGR: Knowledge Graph Reasoning in Hyperbolic Space with Graph Neural Network Encoding Symbolic Pathaccepted
- I-GUARD: Interpretability-Guided Parameter Optimization for Adversarial Defenseaccepted
- ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignmentaccepted
- ICL CIPHERS: Quantifying ”Learning” in In-Context Learning via Substitution Ciphersaccepted
- ICL-Bandit: Relevance Labeling in Advertisement Recommendation Systems via LLMaccepted
- ICLER: Intent CLassification with Enhanced Reasoningaccepted
- ICR: Iterative Clarification and Rewriting for Conversational Searchaccepted
- IG-Pruning: Input-Guided Block Pruning for Large Language Modelsaccepted
- IIET: Efficient Numerical Transformer via Implicit Iterative Euler Methodaccepted
- IL-PCSR: Legal Corpus for Prior Case and Statute Retrievalaccepted
- INDOORWORLD : Integrating Physical Task Solving and Social Simulation in A Heterogeneous Multi-Agent Environmentaccepted
- INREACT: An Inspire-Then-Reinforce Training Framework For Multimodal GUI Agentaccepted
- IP-Dialog: Evaluating Implicit Personalization in Dialogue Systems with Synthetic Dataaccepted
- IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agentsaccepted
- ISACL: Internal State Analyzer for Copyrighted Training Data Leakageaccepted
- Icon2: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulationaccepted
- Identification of Multiple Logical Interpretations in Counter-Argumentsaccepted
- Identifying & Interactively Refining Ambiguous User Goals for Data Visualization Code Generationaccepted
- Identifying Aspects in Peer Reviewsaccepted
- Identifying Noise in Human-Created Datasets using Training Dynamics from Generative Modelsaccepted
- Identifying Pre-training Data in LLMs: A Neuron Activation-Based Detection Frameworkaccepted
- Identifying Rare Languages in Common Crawl Data is a Needles-in-a-Haystack Problemaccepted
- Identifying Unlearned Data in LLMs via Membership Inference Attacksaccepted
- Identifying and Answering Questions with False Assumptions: An Interpretable Approachaccepted
- Idiosyncratic Versus Normative Modeling of Atypical Speech Recognition: Dysarthric Case Studiesaccepted
- Idola Tribus of AI: Large Language Models tend to perceive order where none existsaccepted
- Igniting Creative Writing in Small Language Models: LLM-as-a-Judge versus Multi-Agent Refined Rewardsaccepted
- Image Difference Captioning via Adversarial Preference Optimizationaccepted
- Image Embedding Sampling Method for Diverse Captioningaccepted
- Imagination and Contemplation: A Balanced Framework for Semantic-Augmented Multimodal Machine Translationaccepted
- ImpRAG: Retrieval-Augmented Generation with Implicit Queriesaccepted
- ImpliRet: Benchmarking the Implicit Fact Retrieval Challengeaccepted
- Implicit Behavioral Alignment of Language Agents in High-Stakes Crowd Simulationsaccepted
- Implicit Values Embedded in How Humans and LLMs Complete Subjective Everyday Tasksaccepted
- Improbable Bigrams Expose Vulnerabilities of Incomplete Tokens in Byte-Level Tokenizersaccepted
- Improve LLM-as-a-Judge Ability as a General Abilityaccepted
- Improving Alignment in LVLMs with Debiased Self-Judgmentaccepted
- Improving Chemical Understanding of LLMs via SMILES Parsingaccepted
- Improving Clustering with Positive Pairs Generated from LLM-Driven Labelsaccepted
- Improving Context Fidelity via Native Retrieval-Augmented Reasoningaccepted
- Improving Cross Lingual Transfer by Pretraining with Active Forgettingaccepted
- Improving Handshape Representations for Sign Language Processing: A Graph Neural Network Approachaccepted
- Improving Influence-based Instruction Tuning Data Selection for Balanced Learning of Diverse Capabilitiesaccepted
- Improving Informally Romanized Language Identificationaccepted
- Improving Instruct Models for Free: A Study on Partial Adaptationaccepted
- Improving LLM Reasoning through Interpretable Role-Playing Steeringaccepted
- Improving LLM-as-a-Judge Inference with the Judgment Distributionaccepted
- Improving Language Model Personas via Rationalization with Psychological Scaffoldsaccepted
- Improving Large Language Model Safety with Contrastive Representation Learningaccepted
- Improving Large Language Models Function Calling and Interpretability via Guided-Structured Templatesaccepted
- Improving Low-Resource Sequence Labeling with Knowledge Fusion and Contextual Label Explanationsaccepted
- Improving Multilingual Retrieval-Augmented Language Models through Dialectic Reasoning Argumentationsaccepted
- Improving Neutral Point-of-View Generation with Data- and Parameter-Efficient RLaccepted
- Improving Online Job Advertisement Analysis via Compositional Entity Extractionaccepted
- Improving Preference Alignment of LLM with Inference-Free Self-Refinementaccepted
- Improving Prompt Generalization for Cross-prompt Essay Trait Scoring from the Scoring-invariance Perspectiveaccepted
- Improving Reasoning Capabilities in Small Models through Mixture-of-layers Distillation with Stepwise Attention on Key Informationaccepted
- Improving Rule-based Reasoning in LLMs using Neurosymbolic Representationsaccepted
- Improving Task Diversity in Label Efficient Supervised Finetuning of LLMsaccepted
- Improving Zero-shot Sentence Decontextualisation with Content Selection and Planningaccepted
- Improving the Language Understanding Capabilities of Large Language Models Using Reinforcement Learningaccepted
- Improving the Quality of Web-mined Parallel Corpora of Low-Resource Languages using Debiasing Heuristicsaccepted
- In Benchmarks We Trust ... Or Not?accepted
- In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varietiesaccepted
- InFact: Informativeness Alignment for Improved LLM Factualityaccepted
- InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning Stylesaccepted
- Inceptive Transformers: Enhancing Contextual Representations through Multi-Scale Feature Learning Across Domains and Languagesaccepted
- Inclusive Leadership in the Age of AI: A Dataset and Comparative Study of LLMs vs. Real-Life Leaders in Workplace Action Planningaccepted
- Incorporating Diverse Perspectives in Cultural Alignment: Survey of Evaluation Benchmarks Through A Three-Dimensional Frameworkaccepted
- IndiGEC: Multilingual Grammar Error Correction for Low-Resource Indian Languagesaccepted
- IndoSafety: Culturally Grounded Safety for LLMs in Indonesian Languagesaccepted
- Inducing Argument Facets for Faithful Opinion Summarizationaccepted
- Inductive Reasoning on Few-Shot Knowledge Graphs with Task-Aware Language Modelsaccepted
- Inefficiencies of Meta Agents for Agent Designaccepted
- InfAL: Inference Time Adversarial Learning for Improving Research Ideationaccepted
- InfiMM-WebMath-40B: Advancing Multimodal Pre-Training for Enhanced Mathematical Reasoningaccepted
- Infini-gram mini: Exact n-gram Search at the Internet Scale with FM-Indexaccepted
- InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Showsaccepted
- InfoGain-RAG: Boosting Retrieval-Augmented Generation through Document Information Gain-based Reranking and Filteringaccepted
- Information Integration in Large Language Models is Gated by Linguistic Structural Markersaccepted
- Injecting Domain-Specific Knowledge into Large Language Models: A Comprehensive Surveyaccepted
- InsBank: Evolving Instruction Subset for Ongoing Alignmentaccepted
- Insights into using temporal coordinated behaviour to explore connections between social media posts and influenceaccepted
- Instability in Downstream Task Performance During LLM Pretrainingaccepted
- Instance-level Randomization: Toward More Stable LLM Evaluationsaccepted
- Instructing Large Language Models for Low-Resource Languages: A Systematic Study for Basqueaccepted
- InteGround: On the Evaluation of Verification and Retrieval Planning in Integrative Groundingaccepted
- Integral Transformer: Denoising Attention, Not Too Much Not Too Littleaccepted
- Intelligent Document Parsing: Towards End-to-end Document Parsing via Decoupled Content Parsing and Layout Groundingaccepted
- Intent-aware Schema Generation and Refinement for Literature Review Tablesaccepted
- IntentionFrame: A Semi-Structured, Multi-Aspect Framework for Fine-Grained Conversational Intention Understandingaccepted
- Inter-sentence Context Modeling and Structure-aware Representation Enhancement for Conversational Sentiment Quadruple Extractionaccepted
- InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models with Human Feedbackaccepted
- InterIDEAS: Philosophical Intertextuality via LLMsaccepted
- InteractSpeech: A Speech Dialogue Interaction Corpus for Spoken Dialogue Modelaccepted
- Interdisciplinary Research in Conversation: A Case Study in Computational Morphology for Language Documentationaccepted
- Interesting Culture: Social Relation Recognition from Videos via Culture De-confoundingaccepted
- Internal Chain-of-Thought: Empirical Evidence for Layer‐wise Subtask Scheduling in LLMsaccepted
- Internal states before wait modulate reasoning patternsaccepted
- Interpretability Analysis of Arithmetic In-Context Learning in Large Language Modelsaccepted
- Interpretable Mnemonic Generation for Kanji Learning via Expectation-Maximizationaccepted
- Interpretable Text Embeddings and Text Similarity Explanation: A Surveyaccepted
- Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safetyaccepted
- IntrEx: A Dataset for Modeling Engagement in Educational Conversationsaccepted
- Intrinsic Test of Unlearning Using Parametric Knowledge Tracesaccepted
- Introducing Spotlight: A Novel Approach for Generating Captivating Key Information from Documentsaccepted
- Investigating Controversy Framing across Topics on Social Mediaaccepted
- Investigating Dictionary Expansion for Video-based Sign Language Dictionariesaccepted
- Investigating How Pre-training Data Leakage Affects Models’ Reproduction and Detection Capabilitiesaccepted
- Investigating Multi-layer Representations for Dense Passage Retrievalaccepted
- Investigating Neurons and Heads in Transformer-based LLMs for Typographical Errorsaccepted
- Investigating Pedagogical Teacher and Student LLM Agents: Genetic Adaptation Meets Retrieval-Augmented Generation Across Learning Stylesaccepted
- Investigating Value-Reasoning Reliability in Small Large Language Modelsaccepted
- Investigating the Impact of Conceptual Metaphors on LLM-based NLI through Shapley Interactionsaccepted
- Investigating the interaction of linguistic and mathematical reasoning in language models using multilingual number puzzlesaccepted
- Invisible Entropy: Towards Safe and Efficient Low-Entropy LLM Watermarkingaccepted
- Invisible Prompts, Visible Threats: Malicious Font Injection in External Resources for Large Language Modelsaccepted
- Invoke Interfaces Only When Needed: Adaptive Invocation for Large Language Models in Question Answeringaccepted
- IoTMigrator: LLM-driven Embedded IoT Code Migration across Different OSes for Cloud-device Integrationaccepted
- Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understandingaccepted
- Is Safety Standard Same for Everyone? User-Specific Safety Evaluation of Large Language Modelsaccepted
- Is the Top Still Spinning? Evaluating Subjectivity in Narrative Understandingaccepted
- Iterative Multilingual Spectral Attribute Erasureaccepted
- Iterative Prompt Refinement for Safer Text-to-Image Generationaccepted
- It’s All About In-Context Learning! Teaching Extremely Low-Resource Languages to LLMsaccepted
- JI2S: Joint Influence‐Aware Instruction Data Selection for Efficient Fine‐Tuningaccepted
- JOLT-SQL: Joint Loss Tuning of Text-to-SQL with Confusion-aware Noisy Schema Samplingaccepted
- JUDGEBERT: Assessing Legal Meaning Preservation Between Sentencesaccepted
- JUREX-4E: Juridical Expert-Annotated Four-Element Knowledge Base for Legal Reasoningaccepted
- Jailbreak Attack Initializations as Extractors of Compliance Directionsaccepted
- Jailbreak Distillation: Renewable Safety Benchmarkingaccepted
- Jailbreak LLMs through Internal Stance Manipulationaccepted
- Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibilityaccepted
- Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Modelsaccepted
- Job Unfair: An Investigation of Gender and Occupational Bias in Free-Form Text Completions by LLMsaccepted
- Joint Enhancement of Relational Reasoning for Long-Context LLMsaccepted
- Joint Modeling of Entities and Discourse Relations for Coherence Assessmentaccepted
- Journalism-Guided Agentic In-context Learning for News Stance Detectionaccepted
- Judge and Improve: Towards a Better Reasoning of Knowledge Graphs with Large Language Modelsaccepted
- Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Modelsaccepted
- Judging with Many Minds: Do More Perspectives Mean Less Prejudice? On Bias Amplification and Resistance in Multi-Agent Based LLM-as-Judgeaccepted
- KAHAN: Knowledge-Augmented Hierarchical Analysis and Narration for Financial Data Narrationaccepted
- KBAlign: Efficient Self Adaptation on Specific Textual Knowledge Basesaccepted
- KBM: Delineating Knowledge Boundary for Adaptive Retrieval in Large Language Modelsaccepted
- KCS: Diversify Multi-hop Question Generation with Knowledge Composition Samplingaccepted
- KELE: A Multi-Agent Framework for Structured Socratic Teaching with Large Language Modelsaccepted
- KELE: Residual Knowledge Erasure for Enhanced Multi-hop Reasoning in Knowledge Editingaccepted
- KERAG: Knowledge-Enhanced Retrieval-Augmented Generation for Advanced Question Answeringaccepted
- KG-CQR: Leveraging Structured Relation Representations in Knowledge Graphs for Contextual Query Retrievalaccepted
- KG-RAG: Enhancing GUI Agent Decision-Making via Knowledge Graph-Driven Retrieval-Augmented Generationaccepted
- KGE Calibrator: An Efficient Probability Calibration Method of Knowledge Graph Embedding Models for Trustworthy Link Predictionaccepted
- KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language Modelsaccepted
- KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contextsaccepted
- KaeDe: Progressive Generation of Logical Forms via Knowledge-Aware Question Decomposition for Improved KBQAaccepted
- Keep Security! Benchmarking Security Policy Preservation in Large Language Model Contexts Against Indirect Attacks in Question Answeringaccepted
- Knowing More, Acting Better: Hierarchical Representation for Embodied Decision-Makingaccepted
- Knowledge Editing through Chain-of-Thoughtaccepted
- Knowledge Graph-Driven Memory Editing with Directional Interventionsaccepted
- Knowledge-Aware Co-Reasoning for Multidisciplinary Collaborationaccepted
- Knowledge-enhanced Multimodal ECG Representation Learning with Arbitrary-Lead Inputsaccepted
- Ko-LongRAG: A Korean Long-Context RAG Benchmark Built with a Retrieval-Free Approachaccepted
- KoACD: The First Korean Adolescent Dataset for Cognitive Distortion Analysis via Role-Switching Multi-LLM Negotiationaccepted
- KoBLEX: Open Legal Question Answering with Multi-hop Reasoningaccepted
- KoLEG: On-the-Fly Korean Legal Knowledge Editing with Continuous Retrievalaccepted
- Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidanceaccepted
- Krikri: Advancing Open Large Language Models for Greekaccepted
- KurTail : Kurtosis-based LLM Quantizationaccepted
- LAGCL4Rec: When LLMs Activate Interactions Potential in Graph Contrastive Learning for Recommendationaccepted
- LASER: An LLM-based ASR Scoring and Evaluation Rubricaccepted
- LATTE: Learning to Think with Vision Specialistsaccepted
- LAVa: Layer-wise KV Cache Eviction with Dynamic Budget Allocationaccepted
- LAWCAT: Efficient Distillation from Quadratic to Linear Attention with Convolution across Tokens for Long Context Modelingaccepted
- LC-Eval: A Bilingual Multi-Task Evaluation Benchmark for Long-Context Understandingaccepted
- LCAN: A Label-Aware Contrastive Attention Network for Multi-Intent Recognition and Slot Filling in Task-Oriented Dialogue Systemsaccepted
- LCES: Zero-shot Automated Essay Scoring via Pairwise Comparisons Using Large Language Modelsaccepted
- LEAF: Large Language Diffusion Model for Time Series Forecastingaccepted
- LEO-MINI: An Efficient Multimodal Large Language Model using Conditional Token Reduction and Mixture of Multi-Modal Expertsaccepted
- LGA: LLM-GNN Aggregation for Temporal Evolution Attribute Graph Predictionaccepted
- LIDDIA: Language-based Intelligent Drug Discovery Agentaccepted
- LIFTED: Multimodal Clinical Trial Outcome Prediction via Large Language Models and Mixture-of-Expertsaccepted
- LILaC: Late Interacting in Layered Component Graph for Open-domain Multimodal Multihop Retrievalaccepted
- LLM Agents at the Roundtable: A Multi-Perspective and Dialectical Reasoning Framework for Essay Scoringaccepted
- LLM Agents for Education: Advances and Applicationsaccepted
- LLM Bias Detection and Mitigation through the Lens of Desired Distributionsaccepted
- LLM Distillation for Efficient Few-Shot Multiple Choice Question Answeringaccepted
- LLM Jailbreak Detection for (Almost) Free!accepted
- LLM-Based Web Data Collection for Research Dataset Creationaccepted
- LLM-Driven Completeness and Consistency Evaluation for Cultural Heritage Data Augmentation in Cross-Modal Retrievalaccepted
- LLM-Driven Implicit Target Augmentation and Fine-Grained Contextual Modeling for Zero-Shot and Few-Shot Stance Detectionaccepted
- LLM-Guided Co-Training for Text Classificationaccepted
- LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognitionaccepted
- LLM-Independent Adaptive RAG: Let the Question Speak for Itselfaccepted
- LLM-OREF: An Open Relation Extraction Framework Based on Large Language Modelsaccepted
- LLM-based Conversational Recommendation Agents with Collaborative Verbalized Experienceaccepted
- LLM-based Open Domain Planning by Leveraging Entity-Attribute-Level Domain Modelsaccepted
- LLM-empowered Dynamic Prompt Routing for Vision-Language Models Tuning under Long-Tailed Distributionsaccepted
- LLMAP: LLM-Assisted Multi-Objective Route Planning with User Preferencesaccepted
- LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validationaccepted
- LLMs Behind the Scenes: Enabling Narrative Scene Illustrationaccepted
- LLMs Can Compensate for Deficiencies in Visual Representationsaccepted
- LLMs Don’t Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanationsaccepted
- LLMs Reproduce Stereotypes of Sexual and Gender Minoritiesaccepted
- LLMs are Better Than You Think: Label-Guided In-Context Learning for Named Entity Recognitionaccepted
- LLMs are Privacy Erasableaccepted
- LLMs as World Models: Data-Driven and Human-Centered Pre-Event Simulation for Disaster Impact Assessmentaccepted
- LLMs as a synthesis between symbolic and distributed approaches to languageaccepted
- LLMs cannot spot math errors, even when allowed to peek into the solutionaccepted
- LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?accepted
- LLMsPark: A Benchmark for Evaluating Large Language Models in Strategic Gaming Contextsaccepted
- LLaMP: Large Language Model Made Powerful for High-fidelity Materials Knowledge Retrievalaccepted
- LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learningaccepted
- LM-Searcher: Cross-domain Neural Architecture Search with LLMs via Unified Numerical Encodingaccepted
- LM2Protein: A Structure-to-Token Protein Large Language Modelaccepted
- LMR-BENCH: Evaluating LLM Agent’s Ability on Reproducing Language Modeling Researchaccepted
- LMUNIT: Fine-grained Evaluation with Natural Language Unit Testsaccepted
- LNE-Blocking: An Efficient Framework for Contamination Mitigation Evaluation on Large Language Modelsaccepted
- LOHRec: Leveraging Order and Hierarchy in Generative Sequential Recommendationaccepted
- LORAXBENCH: A Multitask, Multilingual Benchmark Suite for 20 Indonesian Languagesaccepted
- LORE: Continual Logit Rewriting Fosters Faithful Generationaccepted
- LRPLAN: A Multi-Agent Collaboration of Large Language and Reasoning Models for Planning with Implicit & Explicit Constraintsaccepted
- LSRL: Process-Supervised GRPO on Latent Recurrent States Improves Mathematical Reasoningaccepted
- LUME: LLM Unlearning with Multitask Evaluationsaccepted
- LVLMs are Bad at Overhearing Human Referential Communicationaccepted
- LaMDAgent: An Autonomous Framework for Post-Training Pipeline Optimization via LLM Agentsaccepted
- LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profilesaccepted
- LaMP-QA: A Benchmark for Personalized Long-form Question Answeringaccepted
- LaMP-Val: Large Language Models Empower Personalized Valuation in Auctionaccepted
- Label Set Optimization via Activation Distribution Kurtosis for Zero-Shot Classification with Generative Modelsaccepted
- LangProBe: a Language Program Benchmarkaccepted
- Language Mixing in Reasoning Language Models: Patterns, Impact, and Internal Causesaccepted
- Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformersaccepted
- Language Models Can Easily Learn to Reason from Demonstrationsaccepted
- Language Models Can be Efficiently Steered via Minimal Embedding Layer Transformationsaccepted
- Language Models Identify Ambiguities and Exploit Loopholesaccepted
- Language Models as Causal Effect Generatorsaccepted
- Language Models as Continuous Self-Evolving Data Engineersaccepted
- Language models can learn implicit multi-hop reasoning, but only if they have lots of training dataaccepted
- Language-Guided Temporal Token Pruning for Efficient VideoLLM Processingaccepted
- Language-Informed Synthesis of Rational Agent Models for Grounded Theory-of-Mind Reasoning On-the-flyaccepted
- Language-Specific Layer Matters: Efficient Multilingual Enhancement for Large Vision-Language Modelsaccepted
- Language-to-Space Programming for Training-Free 3D Visual Groundingaccepted
- Languages Still Left Behind: Toward a Better Multilingual Machine Translation Benchmarkaccepted
- Large Language Model Agents in Finance: A Survey Bridging Research, Practice, and Real-World Deploymentaccepted
- Large Language Model Evaluation via Matrix Nuclear-Normaccepted
- Large Language Models Badly Generalize across Option Length, Problem Types, and Irrelevant Noun Replacementsaccepted
- Large Language Models Discriminate Against Speakers of German Dialectsaccepted
- Large Language Models Do Multi-Label Classification Differentlyaccepted
- Large Language Models Have Intrinsic Meta-Cognition, but Need a Good Lensaccepted
- Large Language Models Meet Knowledge Graphs for Question Answering: Synthesis and Opportunitiesaccepted
- Large Language Models Might Not Care What You Are Saying: Prompt Format Beats Descriptionsaccepted
- Large Language Models Threaten Language’s Epistemic and Communicative Foundationsaccepted
- Large Language Models as Reader for Bias Detectionaccepted
- Large Language Models as Realistic Microservice Trace Generatorsaccepted
- Large Language Models for Automated Literature Review: An Evaluation of Reference Generation, Abstract Writing, and Review Compositionaccepted
- Large Language Models for Controllable Multi-property Multi-objective Molecule Optimizationaccepted
- Large Language Models for Multilingual Previously Fact-Checked Claim Detectionaccepted
- Large Language Models with Temporal Reasoning for Longitudinal Clinical Summarization and Predictionaccepted
- Large Vision-Language Model Alignment and Misalignment: A Survey Through the Lens of Explainabilityaccepted
- LastingBench: Defend Benchmarks Against Knowledge Leakageaccepted
- Latent Inter-User Difference Modeling for LLM Personalizationaccepted
- Layer Duplication in LLMsaccepted
- Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignmentaccepted
- Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledgeaccepted
- Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech Representationsaccepted
- Layered Insights: Generalizable Analysis of Human Authorial Style by Leveraging All Transformer Layersaccepted
- LeTS: Learning to Think-and-Search via Process-and-Outcome Reward Hybridizationaccepted
- Leaky Thoughts: Large Reasoning Models Are Not Private Thinkersaccepted
- LeanK: Learnable K Cache Channel Pruning for Efficient Decodingaccepted
- Learn and Unlearn: Addressing Misinformation in Multilingual LLMsaccepted
- Learning API Functionality from In-Context Demonstrations for Tool-based Agentsaccepted
- Learning Contextual Retrieval for Robust Conversational Searchaccepted
- Learning Is Not A Race: Improving Retrieval in Language Models via Equal Learningaccepted
- Learning Like Humans: Advancing LLM Reasoning Capabilities via Adaptive Difficulty Curriculum Learning and Expert-Guided Self-Reformulationaccepted
- Learning SQL Like a Human: Structure-Aware Curriculum Learning for Text-to-SQL Generationaccepted
- Learning Subjective Label Distributions via Sociocultural Descriptorsaccepted
- Learning Trajectories of Figurative Language for Pre-Trained Language Modelsaccepted
- Learning What to Remember: Adaptive Probabilistic Memory Retention for Memory-Efficient Language Modelsaccepted
- Learning from Diverse Reasoning Paths with Routing and Collaborationaccepted
- Learning from Few Samples: A Novel Approach for High-Quality Malcode Generationaccepted
- Learning to Align: Addressing Character Frequency Distribution Shifts in Handwritten Text Recognitionaccepted
- Learning to Ask: When LLM Agents Meet Unclear Instructionaccepted
- Learning to Describe Implicit Changes: Noise-robust Pre-training for Image Difference Captioningaccepted
- Learning to Instruct: Fine-Tuning a Task-Aware Instruction Optimizer for Black-Box LLMsaccepted
- Learning to See through Sound: From VggCaps to Multi2Cap for Richer Automated Audio Captioningaccepted
- Legal Fact Prediction: The Missing Piece in Legal Judgment Predictionaccepted
- Legal Mathematical Reasoning with LLMs: Procedural Alignment through Two-Stage Reinforcement Learningaccepted
- LegalSearchLM: Rethinking Legal Case Retrieval as Legal Elements Generationaccepted
- LegoSLM: Connecting LLM with Speech Encoder using CTC Posteriorsaccepted
- Lemma Dilemma: On Lemma Generation Without Domain- or Language-Specific Training Dataaccepted
- Lemmatization as a Classification Task: Results from Arabic across Multiple Genresaccepted
- Lemmatization of Polish Multi-word Expressionsaccepted
- Length Representations in Large Language Modelsaccepted
- Less Is More? Examining Fairness in Pruned Large Language Models for Summarising Opinionsaccepted
- Less Is MuRE: Revisiting Shallow Knowledge Graph Embeddingsaccepted
- Less is More: The Effectiveness of Compact Typological Language Representationsaccepted
- Let Them Down Easy! Contextual Effects of LLM Guardrails on User Perceptions and Preferencesaccepted
- Let’s Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models’ Understanding of Sportsaccepted
- Let’s Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM’s Math Capabilityaccepted
- Leveraging 3D Gaussian for Temporal Knowledge Graph Embeddingaccepted
- Leveraging Cognitive Complexity of Texts for Contextualization in Dense Retrievalaccepted
- Leveraging High-Resource English Corpora for Cross-lingual Domain Adaptation in Low-Resource Japanese Medicine via Continued Pre-trainingaccepted
- Leveraging Knowledge Graph-Enhanced LLMs for Context-Aware Medical Consultationaccepted
- Leveraging Large Models to Evaluate Novel Content: A Case Study on Advertisement Creativityaccepted
- Leveraging Loanword Constraints for Improving Machine Translation in a Low-Resource Multilingual Contextaccepted
- Leveraging Multilingual Training for Authorship Representation: Enhancing Generalization across Languages and Domainsaccepted
- Leveraging Semantic Triples for Private Document Generation with Local Differential Privacy Guaranteesaccepted
- Leveraging Text-to-Text Transformers as Classifier Chain for Few-Shot Multi-Label Classificationaccepted
- Leveraging Unpaired Feedback for Long-Term LLM-based Recommendation Tuningaccepted
- Leveraging What’s Overfixed: Post-Correction via LLM Grammatical Error Overcorrectionaccepted
- Leveraging the Cross-Domain & Cross-Linguistic Corpus for Low Resource NMT: A Case Study On Bhili-Hindi-English Parallel Corpusaccepted
- LexTime: A Benchmark for Temporal Ordering of Legal Eventsaccepted
- LiTEx: A Linguistic Taxonomy of Explanations for Understanding Within-Label Variation in Natural Language Inferenceaccepted
- LiTransProQA: An LLM-based Literary Translation Evaluation Metric with Professional Question Answeringaccepted
- Liaozhai through the Looking-Glass: On Paratextual Explicitation of Culture-Bound Terms in Machine Translationaccepted
- Lifelong Knowledge Editing requires Better Regularizationaccepted
- LightRAG: Simple and Fast Retrieval-Augmented Generationaccepted
- LightThinker: Thinking Step-by-Step Compressionaccepted
- Likelihood Variance as Text Importance for Resampling Texts to Map Language Modelsaccepted
- LimRank: Less is More for Reasoning-Intensive Information Rerankingaccepted
- LimaCost: Data Valuation for Instruction Tuning of Large Language Modelsaccepted
- Linear Steerability in Language Models: When It Emerges and How It Evolvesaccepted
- Linear-Time Demonstration Selection for In-Context Learning via Gradient Estimationaccepted
- LingGym: How Far Are LLMs from Thinking Like Field Linguists?accepted
- LinguaLens: Towards Interpreting Linguistic Mechanisms of Large Language Models via Sparse Auto-Encoderaccepted
- Linguistic Alignment Predicts Learning in Small Group Tutoring Sessionsaccepted
- Linguistic Neuron Overlap Patterns to Facilitate Cross-lingual Transfer on Low-resource Languagesaccepted
- Linguistic and Embedding-Based Profiling of Texts Generated by Humans and Large Language Modelsaccepted
- Linguistically-Controlled Paraphrase Generationaccepted
- LinkAlign: Scalable Schema Linking for Real-World Large-Scale Multi-Database Text-to-SQLaccepted
- LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximationaccepted
- LiteraryQA: Towards Effective Evaluation of Long-document Narrative QAaccepted
- LlmFixer: Fix the Helpfulness of Defensive Large Language Modelsaccepted
- LoCt-Instruct: An Automatic Pipeline for Constructing Datasets of Logical Continuous Instructionsaccepted
- LoRA-MGPO: Mitigating Double Descent in Low-Rank Adaptation via Momentum-Guided Perturbation Optimizationaccepted
- LoRA-PAR: A Flexible Dual-System LoRA Partitioning Approach to Efficient LLM Fine-Tuningaccepted
- LoRACoE: Improving Large Language Model via Composition-based LoRA Expertaccepted
- LoRATK: LoRA Once, Backdoor Everywhere in the Share-and-Play Ecosystemaccepted
- LoRE-Merging: Exploring Low-Rank Estimation For Large Language Model Mergingaccepted
- LoRaDA: Low-Rank Direct Attention Adaptation for Efficient LLM Fine-tuningaccepted
- LoSiA: Efficient High-Rank Fine-Tuning via Subnet Localization and Optimizationaccepted
- Local Normalization Distortion and the Thermodynamic Formalism of Decoding Strategies for Large Language Modelsaccepted
- Localizing Malicious Outputs from CodeLLMaccepted
- Locate-then-Merge: Neuron-Level Parameter Fusion for Mitigating Catastrophic Forgetting in Multimodal LLMsaccepted
- Lock on Target! Precision Unlearning via Directional Controlaccepted
- LogiCoL: Logically-Informed Contrastive Learning for Set-based Dense Retrievalaccepted
- LogiDynamics: Unraveling the Dynamics of Inductive, Abductive and Deductive Logical Inferences in LLM Reasoningaccepted
- Logic-Thinker: Teaching Large Language Models to Think more Logically.accepted
- Logic: Long-form Outline Generation via Imitative and Critical Self-refinementaccepted
- LogicTree: Structured Proof Exploration for Coherent and Rigorous Logical Reasoning with Large Language Modelsaccepted
- Logical Reasoning with Outcome Reward Models for Test-Time Scalingaccepted
- Logit Space Constrained Fine-Tuning for Mitigating Hallucinations in LLM-Based Recommender Systemsaccepted
- Logits-Based Finetuningaccepted
- Logos as a Well-Tempered Pre-train for Sign Language Recognitionaccepted
- Long Chain-of-Thought Fine-tuning via Understanding-to-Reasoning Transitionaccepted
- Long-Form Information Alignment Evaluation Beyond Atomic Factsaccepted
- Long-context Language Models Fail in Basic Retrieval Tasks Without Sufficient Reasoning Stepsaccepted
- LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architectureaccepted
- LongTableBench: Benchmarking Long-Context Table Reasoning across Real-World Formats and Domainsaccepted
- LongTail-Swap: benchmarking language models’ abilities on rare wordsaccepted
- LongWeave: A Long-Form Generation Benchmark Bridging Real-World Relevance and Verifiabilityaccepted
- Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Modelsaccepted
- Look Beyond Feeling: Unveiling Latent Needs from Implicit Expressions for Proactive Emotional Supportaccepted
- Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Queryaccepted
- Looking Beyond Text: Reducing Language Bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidanceaccepted
- Looking Beyond the Pixels: Evaluating Visual Metaphor Understanding in VLMsaccepted
- Lost in Embeddings: Information Loss in Vision–Language Modelsaccepted
- Low-Confidence Gold: Refining Low-Confidence Samples for Efficient Instruction Tuningaccepted
- Low-Hallucination and Efficient Coreference Resolution with LLMsaccepted
- Low-Resource Languages LLM Disinformation is Within Reach: The Case of Walliserdeutschaccepted
- LyapLock: Bounded Knowledge Preservation in Sequential Large Language Model Editingaccepted
- M-ABSA: A Multilingual Dataset for Aspect-Based Sentiment Analysisaccepted
- M-BRe: Discovering Training Samples for Relation Extraction from Unlabeled Texts with Large Language Modelsaccepted
- M-Help: Using Social Media Data to Detect Mental Health Help-Seeking Signalsaccepted
- M-LongDoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Frameworkaccepted
- M-Ped: Multi-Prompt Ensemble Decoding for Large Language Modelsaccepted
- M-Wanda: Improving One-Shot Pruning for Multilingual LLMsaccepted
- M2Edit: Locate and Edit Multi-Granularity Knowledge in Multimodal Large Language Modelaccepted
- M3Retrieve: Benchmarking Multimodal Retrieval for Medicineaccepted
- MA-DPR: Manifold-aware Distance Metrics for Dense Passage Retrievalaccepted
- MA-GTS: A Multi-Agent Framework for Solving Complex Graph Problems in Real-World Applicationsaccepted
- MAC-Tuning: LLM Multi-Compositional Problem Reasoning with Enhanced Knowledge Boundary Awarenessaccepted
- MADAWSD: Multi-Agent Debate Framework for Adversarial Word Sense Disambiguationaccepted
- MADD: Multi-Agent Drug Discovery Orchestraaccepted
- MAFMO: Multi-modal Adaptive Fusion with Meta-template Optimization for Vision-Language Modelsaccepted
- MAGIC: A Multi-Hop and Graph-Based Benchmark for Inter-Context Conflicts in Retrieval-Augmented Generationaccepted
- MAIN: Mutual Alignment Is Necessary for instruction tuningaccepted
- MAKAR: a Multi-Agent framework based Knowledge-Augmented Reasoning for Grounded Multimodal Named Entity Recognitionaccepted
- MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMsaccepted
- MANTA: A Scalable Pipeline for Transmuting Massive Web Corpora into Instruction Datasetsaccepted
- MARIO-0.5B: A Multi-Agent Lightweight Model for Real-Time Open Information Extraction in Low-Resource Settingsaccepted
- MARS-Bench: A Multi-turn Athletic Real-world Scenario Benchmark for Dialogue Evaluationaccepted
- MASSIVE-Agents: A Benchmark for Multilingual Function-Calling in 52 Languagesaccepted
- MASSV: Multimodal Adaptation and Self-Data Distillation for Speculative Decoding of Vision-Language Modelsaccepted
- MASTER: Multi-Agent Security Through Exploration of Roles and Topological Structures - A Comprehensive Frameworkaccepted
- MATCH: Task-Driven Code Evaluation through Contrastive Learningaccepted
- MAVL: A Multilingual Audio-Video Lyrics Dataset for Animated Song Translationaccepted
- MAgICoRe: Multi-Agent, Iterative, Coarse-to-Fine Refinement for Reasoningaccepted
- MAviS: A Multimodal Conversational Assistant For Avian Speciesaccepted
- MC2: A Minimum-Coverage and Dataset-Agnostic Framework for Compositional Generalization of LLMs on Semantic Parsingaccepted
- MCIP: Protecting MCP Safety via Model Contextual Integrity Protocolaccepted
- MCTS-RAG: Enhancing Retrieval-Augmented Generation with Monte Carlo Tree Searchaccepted
- MCiteBench: A Multimodal Benchmark for Generating Text with Citationsaccepted
- MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarizationaccepted
- MEBench: Benchmarking Large Language Models for Cross-Document Multi-Entity Question Answeringaccepted
- MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapperaccepted
- MERMAID: Multi-perspective Self-reflective Agents with Generative Augmentation for Emotion Recognitionaccepted
- METok: Multi-Stage Event-based Token Compression for Efficient Long Video Understandingaccepted
- MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregationaccepted
- MFTCXplain: A Multilingual Benchmark Dataset for Evaluating the Moral Reasoning of LLMs through Multi-hop Hate Speech Explanationaccepted
- MIND: Towards Immersive Psychological Healing with Multi-Agent Inner Dialogueaccepted
- MIO: A Foundation Model on Multimodal Tokensaccepted
- MIRROR: Multimodal Cognitive Reframing Therapy for Rolling with Resistanceaccepted
- ML-Promise: A Multilingual Dataset for Corporate Promise Verificationaccepted
- MLAlgo-Bench: Can Machines Implement Machine Learning Algorithms?accepted
- MLWQ: Efficient Small Language Model Deployment via Multi-Level Weight Quantizationaccepted
- MM-CRITIC: A Holistic Evaluation of Large Multimodal Models as Multimodal Critiqueaccepted
- MMA: Cross-Domain Knowledge Integration via Mixture of Multi-Domain Agentsaccepted
- MMAG: Multimodal Learning for Mucus Anomaly Grading in Nasal Endoscopy via Semantic Attribute Promptingaccepted
- MMAPG: A Training-Free Framework for Multimodal Multi-hop Question Answering via Adaptive Planning Graphsaccepted
- MMDocIR: Benchmarking Multimodal Retrieval for Long Documentsaccepted
- MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluationaccepted
- MMPlanner: Zero-Shot Multimodal Procedural Planning with Chain-of-Thought Object State Reasoningaccepted
- MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts?accepted
- MOLE: Metadata Extraction and Validation in Scientific Papers Using LLMsaccepted
- MONAQ: Multi-Objective Neural Architecture Querying for Time-Series Analysis on Resource-Constrained Devicesaccepted
- MOSAIC: Modeling Social AI for Content Dissemination and Regulation in Multi-Agent Simulationsaccepted
- MPCG: Multi-Round Persona-Conditioned Generation for Modeling the Evolution of Misinformation with LLMsaccepted
- MPO: Boosting LLM Agents with Meta Plan Optimizationaccepted
- MPRF: Interpretable Stance Detection through Multi-Path Reasoning Frameworkaccepted
- MPTA: MultiTask Personalization Assessmentaccepted
- MR. Judge: Multimodal Reasoner as a Judgeaccepted
- MRAG: A Modular Retrieval Framework for Time-Sensitive Question Answeringaccepted
- MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMsaccepted
- MS-RAG: Simple and Effective Multi-Semantic Retrieval-Augmented Generationaccepted
- MT-Mol: Multi Agent System with Tool-based Reasoning for Molecular Optimizationaccepted
- MT-R1-Zero: Advancing LLM-based Machine Translation via R1-Zero-like Reinforcement Learningaccepted
- MT-RewardTree: A Comprehensive Framework for Advancing LLM-Based Machine Translation via Reward Modelingaccepted
- MTabVQA: Evaluating Multi-Tabular Reasoning of Language Models in Visual Spaceaccepted
- MUCAR: Benchmarking Multilingual Cross-Modal Ambiguity Resolution for Multimodal Large Language Modelsaccepted
- MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Languageaccepted
- MULTIGUARD: An Efficient Approach for AI Safety Moderation Across Languages and Modalitiesaccepted
- MULTITAT: Benchmarking Multilingual Table-and-Text Question Answeringaccepted
- MULTIVOX: A Benchmark for Evaluating Voice Assistants for Multimodal Interactionsaccepted
- MUSE: MCTS-Driven Red Teaming Framework for Enhanced Multi-Turn Dialogue Safety in Large Language Modelsaccepted
- MUZO: Leveraging Multiple Queries and Momentum for Zeroth-Order Fine-Tuning of Large Language Modelsaccepted
- MV-CLAM: Multi-View Molecular Interpretation with Cross-Modal Projection via Language Modelaccepted
- MaGiX: A Multi-Granular Adaptive Graph Intelligence Framework for Enhancing Cross-Lingual RAGaccepted
- MaZO: Masked Zeroth-Order Optimization for Multi-Task Fine-Tuning of Large Language Modelsaccepted
- Machine-generated text detection prevents language model collapseaccepted
- Mahānāma: A Unique Testbed for Literary Entity Discovery and Linkingaccepted
- Make Every Letter Count: Building Dialect Variation Dictionaries from Monolingual Corporaaccepted
- Making Every Step Effective: Jailbreaking Large Vision-Language Models Through Hierarchical KV Equalizationaccepted
- Making VLMs More Robot-Friendly: Self-Critical Distillation of Low-Level Procedural Reasoningaccepted
- Mamba Drafters for Speculative Decodingaccepted
- ManuSearch: Democratizing Deep Search in Large Language Models with a Transparent and Open Multi-Agent Frameworkaccepted
- Mapping Toxic Comments Across Demographics: A Dataset from German Public Broadcastingaccepted
- Mapping semantic networks to Dutch word embeddings as a diagnostic tool for cognitive declineaccepted
- Mapping the Minds of LLMs: A Graph-Based Analysis of Reasoning LLMsaccepted
- MarathiEmoExplain: A Dataset for Sentiment, Emotion, and Explanation in Low-Resource Marathiaccepted
- MaskCD: Mitigating LVLM Hallucinations by Image Head Masked Contrastive Decodingaccepted
- Masked Diffusion Captioning for Visual Feature Learningaccepted
- Massive Supervised Fine-tuning Experiments Reveal How Data, Layer, and Training Factors Shape LLM Alignment Qualityaccepted
- MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutorsaccepted
- Matter-of-Fact: A Benchmark for Verifying the Feasibility of Literature-Supported Claims in Materials Scienceaccepted
- Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biasesaccepted
- Measuring Chain of Thought Faithfulness by Unlearning Reasoning Stepsaccepted
- Measuring Lexical Diversity of Synthetic Data Generated through Fine-Grained Persona Promptingaccepted
- Measuring Risk of Bias in Biomedical Reports: The RoBBR Benchmarkaccepted
- Measuring Sycophancy of Language Models in Multi-turn Dialoguesaccepted
- Measuring and Mitigating Media Outlet Name Bias in Large Language Modelsaccepted
- Measuring scalar constructs in social science with LLMsaccepted
- Measuring the Effect of Disfluency in Multilingual Knowledge Probing Benchmarksaccepted
- Mechanisms vs. Outcomes: Probing for Syntax Fails to Explain Performance on Targeted Syntactic Evaluationsaccepted
- Mechanistic Understanding and Mitigation of Language Confusion in English-Centric Large Language Modelsaccepted
- Med-PRM: Medical Reasoning Models with Stepwise, Guideline-verified Process Rewardsaccepted
- Med-VRAgent: A Framework for Medical Visual Reasoning-Enhanced Agentsaccepted
- MedCOD: Enhancing English-to-Spanish Medical Translation of Large Language Models Using Enriched Chain-of-Dictionary Frameworkaccepted
- MedEBench: Diagnosing Reliability in Text-Guided Medical Image Editingaccepted
- MedFact: A Large-scale Chinese Dataset for Evidence-based Medical Fact-checking of LLM Responsesaccepted
- MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Modelsaccepted
- MedLinkDE – MedDRA Entity Linking for German with Guided Chain of Thought Reasoningaccepted
- MediVLM: A Vision Language Model for Radiology Report Generation from Medical Imagesaccepted
- Media Source Matters More Than Content: Unveiling Political Bias in LLM-Generated Citationsaccepted
- MemInsight: Autonomous Memory Augmentation for LLM Agentsaccepted
- Membership and Memorization in LLM Knowledge Distillationaccepted
- MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Modelsaccepted
- MemeIntel: Explainable Detection of Propagandistic and Hateful Memesaccepted
- MemeInterpret: Towards an All-in-One Dataset for Meme Understandingaccepted
- MemeReaCon: Probing Contextual Meme Understanding in Large Vision-Language Modelsaccepted
- Memorization or Reasoning? Exploring the Idiom Understanding of LLMsaccepted
- Memorization ≠ Understanding: Do Large Language Models Have the Ability of Scenario Cognition?accepted
- Memory OS of AI Agentaccepted
- Memory-QA: Answering Recall Questions Based on Multimodal Memoriesaccepted
- Memory-enhanced Large Language Model for Cross-lingual Dependency Parsing via Deep Hierarchical Syntax Understandingaccepted
- MentalGLM Series: Explainable Large Language Models for Mental Health Analysis on Chinese Social Mediaaccepted
- Merge then Realign: Simple and Effective Modality-Incremental Continual Learning for Multimodal LLMsaccepted
- Merger-as-a-Stealer: Stealing Targeted PII from Aligned LLMs with Model Mergingaccepted
- MessIRve: A Large-Scale Spanish Information Retrieval Datasetaccepted
- Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judgeaccepted
- Meta-Semantics Augmented Few-Shot Relational Learningaccepted
- MetaFaith: Faithful Natural Language Uncertainty Expression in LLMsaccepted
- MetaLadder: Ascending Mathematical Solution Quality via Analogical-Problem Reasoning Transferaccepted
- MetaMixSpeech: Meta Task Augmentation for Low-Resource Speech Recognitionaccepted
- Metric Calculating Benchmark: Code-Verifiable Complicate Instruction Following Benchmark for Large Language Modelsaccepted
- MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learningaccepted
- MiLQ: Benchmarking IR Models for Bilingual Web Search with Mixed Language Queriesaccepted
- MicroEdit: Neuron-level Knowledge Disentanglement and Localization in Lifelong Model Editingaccepted
- MidPO: Dual Preference Optimization for Safety and Helpfulness in Large Language Models via a Mixture of Experts Frameworkaccepted
- Middo: Model-Informed Dynamic Data Optimization for Enhanced LLM Fine-Tuning via Closed-Loop Learningaccepted
- Mind the Blind Spots: A Focus-Level Evaluation Framework for LLM Reviewsaccepted
- Mind the Dialect: NLP Advancements Uncover Fairness Disparities for Arabic Users in Recommendation Systemsaccepted
- Mind the Gap: A Closer Look at Tokenization for Multiple-Choice Question Answering with LLMsaccepted
- Mind the Gap: How BabyLMs Learn Filler-Gap Dependenciesaccepted
- Mind the Inclusivity Gap: Multilingual Gender-Neutral Translation Evaluation with mGeNTEaccepted
- Mind the Style Gap: Meta-Evaluation of Style and Attribute Transfer Metricsaccepted
- Mind the Value-Action Gap: Do LLMs Act in Alignment with Their Values?accepted
- Minimal Ranks, Maximum Confidence: Parameter-efficient Uncertainty Quantification for LoRAaccepted
- Minimal, Local, and Robust: Embedding-Only Edits for Implicit Bias in T2I Modelsaccepted
- Mining the Past with Dual Criteria: Integrating Three types of Historical Information for Context-aware Event Forecastingaccepted
- Misalignment Attack on Text-to-Image Models via Text Embedding Optimization and Inversionaccepted
- MisinfoBench: A Multi-Dimensional Benchmark for Evaluating LLMs’ Resilience to Misinformationaccepted
- Mitigate One, Skew Another? Tackling Intersectional Biases in Text-to-Image Modelsaccepted
- Mitigating Attention Localization in Small Scale: Self-Attention Refinement via One-step Belief Propagationaccepted
- Mitigating Biases in Language Models via Bias Unlearningaccepted
- Mitigating Catastrophic Forgetting in Large Language Models with Forgetting-aware Pruningaccepted
- Mitigating Gender Bias via Fostering Exploratory Thinking in LLMsaccepted
- Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality Aligningaccepted
- Mitigating Hallucination in Large Vision-Language Models through Aligning Attention Distribution to Information Flowaccepted
- Mitigating Hallucinations in LM-Based TTS Models via Distribution Alignment Using GFlowNetsaccepted
- Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinationsaccepted
- Mitigating Hallucinations in Large Vision-Language Models via Entity-Centric Multimodal Preference Optimizationaccepted
- Mitigating Hallucinations in Vision-Language Models through Image-Guided Head Suppressionaccepted
- Mitigating Interviewer Bias in Multimodal Depression Detection: An Approach with Adversarial Learning and Contextual Positional Encodingaccepted
- Mitigating Object Hallucinations in MLLMs via Multi-Frequency Perturbationsaccepted
- Mitigating Sequential Dependencies: A Survey of Algorithms and Systems for Generation-Refinement Frameworks in Autoregressive Modelsaccepted
- Mitigating Spurious Correlations via Counterfactual Contrastive Learningaccepted
- Mitigating Visual Knowledge Forgetting in MLLM Instruction-tuning via Modality-decoupled Gradient Descentaccepted
- Mitigating the Privacy Issues in Retrieval-Augmented Generation (RAG) via Pure Synthetic Dataaccepted
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corporaaccepted
- Mixed Signals: Decoding VLMs’ Reasoning and Underlying Bias in Vision-Language Conflictaccepted
- Mixing Inference-time Experts for Enhancing LLM Reasoningaccepted
- Mixture of Languages: Improved Multilingual Encoders Through Language Groupingaccepted
- Mixture of Length and Pruning Experts for Knowledge Graphs Reasoningaccepted
- Mixture of LoRA Experts for Continual Information Extraction with LLMsaccepted
- Mixture of Weight-shared Heterogeneous Group Attention Experts for Dynamic Token-wise KV Optimizationaccepted
- Mixture-of-Clustered-Experts: Advancing Expert Specialization and Generalization in Instruction Tuningaccepted
- MoLoRAG: Bootstrapping Document Understanding via Multi-modal Logic-aware Retrievalaccepted
- MoMentS: A Comprehensive Multimodal Benchmark for Theory of Mindaccepted
- MoMoE: Mixture of Moderation Experts Framework for AI-Assisted Online Governanceaccepted
- MoR: Better Handling Diverse Queries with a Mixture of Sparse, Dense, and Human Retrieversaccepted
- MoRoVoc: A Large Dataset for Geographical Variation Identification of the Spoken Romanian Languageaccepted
- MoSEs: Uncertainty-Aware AI-Generated Text Detection via Mixture of Stylistics Experts with Conditional Thresholdsaccepted
- MoVa: Towards Generalizable Classification of Human Morals and Valuesaccepted
- MoVoC: Morphology-Aware Subword Construction for Ge’ez Script Languagesaccepted
- MobiZO: Enabling Efficient LLM Fine-Tuning at the Edge via Inference Enginesaccepted
- ModRWKV: Transformer Multimodality in Linear Timeaccepted
- ModalPrompt: Towards Efficient Multimodal Continual Instruction Tuning with Dual-Modality Guided Promptaccepted
- Model Calibration for Emotion Detectionaccepted
- Model Consistency as a Cheap yet Predictive Proxy for LLM Elo Scoresaccepted
- Model Unlearning via Sparse Autoencoder Subspace Guided Projectionsaccepted
- Model-Based Ranking of Source Languages for Zero-Shot Cross-Lingual Transferaccepted
- Model-based Large Language Model Customization as Serviceaccepted
- ModelCitizens: Representing Community Voices in Online Safetyaccepted
- Modeling Bottom-up Information Quality during Language Processingaccepted
- Modeling Subjectivity in Cognitive Appraisal with Language Modelsaccepted
- Modeling, Evaluating, and Embodying Personality in LLMs: A Surveyaccepted
- ModelingAgent: Bridging LLMs and Mathematical Modeling for Real-World Challengesaccepted
- MolErr2Fix: Benchmarking LLM Trustworthiness in Chemistry via Modular Error Detection, Localization, Explanation, and Correctionaccepted
- Molecular String Representation Preferences in Pretrained LLMs: A Comparative Study in Zero- & Few-Shot Molecular Property Predictionaccepted
- Mondrian: A Framework for Logical Abstract (Re)Structuringaccepted
- Morables: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fablesaccepted
- Moral Framing in Politics (MFiP): A new resource and models for moral framingaccepted
- More Documents, Same Length: Isolating the Challenge of Multiple Documents in RAGaccepted
- More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compressionaccepted
- Morpheme Induction for Emergent Languageaccepted
- MotivGraph-SoIQ: Integrating Motivational Knowledge Graphs and Socratic Dialogue for Enhanced LLM Ideationaccepted
- MovieCORE: COgnitive REasoning in Moviesaccepted
- MrGuard: A Multilingual Reasoning Guardrail for Universal LLM Safetyaccepted
- MuCAL: Contrastive Alignment for Preference-Driven KG-to-Text Generationaccepted
- MuTIS: Enhancing Reasoning Efficiency through Multi Turn Intervention Sampling in Reinforcement Learningaccepted
- Multi-Agent Autonomous Driving Systems with Large Language Models: A Survey of Recent Advances, Resources, and Future Directionsaccepted
- Multi-Document Event Extraction Using Large and Small Language Modelsaccepted
- Multi-Domain Explainability of Preferencesaccepted
- Multi-Frequency Contrastive Decoding: Alleviating Hallucinations for Large Vision-Language Modelsaccepted
- Multi-LMentry: Can Multilingual LLMs Solve Elementary Tasks Across Languages?accepted
- Multi-Modal Framing Analysis of Newsaccepted
- Multi-Surrogate-Objective Optimization for Neural Topic Modelsaccepted
- Multi-level Diagnosis and Evaluation for Robust Tabular Feature Engineering with Large Language Modelsaccepted
- Multi-perspective Analysis of Large Language Model Domain Specialization: An Experiment in Accounting Audit Procedures Generationaccepted
- Multi-token Mask-filling and Implicit Discourse Relationsaccepted
- Multi-view-guided Passage Reranking with Large Language Modelsaccepted
- MultiAgentESC: A LLM-based Multi-Agent Collaboration Framework for Emotional Support Conversationaccepted
- MultiClaimNet: A Massively Multilingual Dataset of Fact-Checked Claim Clustersaccepted
- MultiConIR: Towards Multi-Condition Information Retrievalaccepted
- MultiDocFusion : Hierarchical and Multimodal Chunking Pipeline for Enhanced RAG on Long Industrial Documentsaccepted
- MultiLingPoT: Boosting Mathematical Reasoning in LLMs through Multilingual Program Integrationaccepted
- MultiLogicNMR(er): A Benchmark and Neural-Symbolic Framework for Non-monotonic Reasoning with Multiple Extensionsaccepted
- MultiMatch: Multihead Consistency Regularization Matching for Semi-Supervised Text Classificationaccepted
- MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translationaccepted
- MultiPL-MoE: Multi-Programming-Lingual Extension of Large Language Models through Hybrid Mixture-of-Expertsaccepted
- Multifaceted Evaluation of Audio-Visual Capability for MLLMs: Effectiveness, Efficiency, Generalizability and Robustnessaccepted
- Multilingual Collaborative Defense for Large Language Modelsaccepted
- Multilingual Data Filtering using Synthetic Data from Large Language Modelsaccepted
- Multilingual Datasets for Custom Input Extraction and Explanation Requests Parsing in Conversational XAI Systemsaccepted
- Multilingual Dialogue Generation and Localization with Dialogue Act Scriptingaccepted
- Multilingual Federated Low-Rank Adaptation for Collaborative Content Anomaly Detection across Multilingual Social Media Participantsaccepted
- Multilingual Generative Retrieval via Cross-lingual Semantic Compressionaccepted
- Multilingual Knowledge Graph Completion via Efficient Multilingual Knowledge Sharingaccepted
- Multilingual Language Model Pretraining using Machine-translated Dataaccepted
- Multilingual Pretraining for Pixel Language Modelsaccepted
- Multilingual Prompting for Improving LLM Generation Diversityaccepted
- Multilingual Verbalisation of Knowledge Graphsaccepted
- Multilingual vs Crosslingual Retrieval of Fact-Checked Claims: A Tale of Two Approachesaccepted
- Multilinguality Does not Make Sense: Investigating Factors Behind Zero-Shot Cross-Lingual Transfer in Sense-Aware Tasksaccepted
- Multimedia Event Extraction with LLM Knowledge Editingaccepted
- Multimodal Document-level Triple Extraction via Dynamic Graph Enhancement and Relation-Aware Reflectionaccepted
- Multimodal Emotion Recognition in Conversations: A Survey of Methods, Trends, Challenges and Prospectsaccepted
- Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesisaccepted
- Multimodal Language Models See Better When They Look Shalloweraccepted
- Multimodal Neural Machine Translation: A Survey of the State of the Artaccepted
- Multimodal UNcommonsense: From Odd to Ordinary and Ordinary to Oddaccepted
- MusKGC: A Flexible Multi-source Knowledge Enhancement Framework for Open-World Knowledge Graph Completionaccepted
- MuseScorer: Idea Originality Scoring At Scaleaccepted
- MythTriage: Scalable Detection of Opioid Use Disorder Myths on a Video-Sharing Platformaccepted
- N-CORE: N-View Consistency Regularization for Disentangled Representation Learning in Nonverbal Vocalizationsaccepted
- NAP2: A Benchmark for Naturalness and Privacy-Preserving Text Rewriting by Learning from Humanaccepted
- NER Retriever: Zero-Shot Named Entity Retrieval with Type-Aware Embeddingsaccepted
- NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Callsaccepted
- NEXUS: Network Exploration for eXploiting Unsafe Sequences in Multi-Turn LLM Jailbreaksaccepted
- NILE: Internal Consistency Alignment in Large Language Modelsaccepted
- NIM: Neuro-symbolic Ideographic Metalanguage for Inclusive Communicationaccepted
- NL-Debugging: Exploiting Natural Language as an Intermediate Representation for Code Debuggingaccepted
- NL2Lean: Translating Natural Language into Lean 4 through Multi-Aspect Reinforcement Learningaccepted
- NLKI: A Lightweight Natural Language Knowledge Integration Framework for Improving Small VLMs in Commonsense VQA Tasksaccepted
- NLP Needs Diversity outside of ‘Diversity’accepted
- NLP-ADBench: NLP Anomaly Detection Benchmarkaccepted
- NLoRA: Nyström-Initiated Low-Rank Adaptation for Large Language Modelsaccepted
- NOVA-63: Native Omni-lingual Versatile Assessments of 63 Disciplinesaccepted
- NOVER: Incentive Training for Language Models via Verifier-Free Reinforcement Learningaccepted
- NUMINA: A Natural Understanding Benchmark for Multi-dimensional Intelligence and Numerical Reasoning Abilitiesaccepted
- NUTMEG: Separating Signal From Noise in Annotator Disagreementaccepted
- NarratEX Dataset: Explaining the Dominant Narratives in News Textsaccepted
- Natural Context Drift Undermines the Natural Language Understanding of Large Language Modelsaccepted
- Navigating the Unknown: Intent Classification and Out-of-Distribution Detection Using Large Language Modelsaccepted
- NeLLCom-Lex: A Neural-agent Framework to Study the Interplay between Lexical Systems and Language Useaccepted
- NeighXLM: Enhancing Cross-Lingual Transfer in Low-Resource Languages via Neighbor-Augmented Contrastive Pretrainingaccepted
- Nested Named Entity Recognition as Single-Pass Sequence Labelingaccepted
- Neural Topic Modeling via Contextual and Graph Information Fusionaccepted
- NeuroAda: Activating Each Neuron’s Potential for Parameter-Efficient Fine-Tuningaccepted
- Neuron-Level Differentiation of Memorization and Generalization in Large Language Modelsaccepted
- Neutral Is Not Unbiased: Evaluating Implicit and Intersectional Identity Bias in LLMs Through Structured Narrative Scenariosaccepted
- Nexus: Adaptive Upcycling to Efficiently Pretrain Mixture of Expertsaccepted
- NileChat: Towards Linguistically Diverse and Culturally Aware LLMs for Local Communitiesaccepted
- Nine Ways to Break Copyright Law and Why Our LLM Won’t: A Fair Use Aligned Generation Frameworkaccepted
- NitiBench: Benchmarking LLM Frameworks on Thai Legal Question Answering Capabilitiesaccepted
- No Black Boxes: Interpretable and Interactable Predictive Healthcare with Knowledge-Enhanced Agentic Causal Discoveryaccepted
- No Free Lunch: Retrieval-Augmented Generation Undermines Fairness in LLMs, Even for Vigilant Usersaccepted
- No Need for Explanations: LLMs can implicitly learn from mistakes in-contextaccepted
- Noise, Adaptation, and Strategy: Assessing LLM Fidelity in Decision-Makingaccepted
EMNLP accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.