EMNLP 2023 Accepted Papers
The full list of 2,009 papers accepted at EMNLP 2023 (Conference on Empirical Methods in Natural Language Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
Long Main: 862Long Findings: 818Short Findings: 197Short Main: 132
- Intra-Event and Inter-Event Dependency-Aware Graph Network for Event Argument ExtractionLong Findings
- Introducing Rhetorical Parallelism Detection: A New Task with Datasets, Metrics, and BaselinesLong Main
- Intuitive Multilingual Audio-Visual Speech Recognition with a Single-Trained ModelShort Findings
- InvGC: Robust Cross-Modal Retrieval by Inverse Graph ConvolutionLong Findings
- Inverse Reinforcement Learning for Text SummarizationLong Findings
- Inverse Scaling Can Become U-ShapedShort Main
- Investigating Bias in Multilingual Language Models: Cross-Lingual Transfer of Debiasing TechniquesShort Main
- Investigating Efficiently Extending Transformers for Long Input SummarizationLong Main
- Investigating Multilingual Coreference Resolution by Universal AnnotationsLong Findings
- Investigating Online Community Engagement through StancetakingLong Findings
- Investigating the Effect of Pre-finetuning BERT Models on NLI Involving PresuppositionsLong Findings
- Investigating the Effectiveness of Multiple Expert Models CollaborationShort Findings
- Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking AgentsLong Main
- Is ChatGPT a Financial Expert? Evaluating Language Models on Financial Natural Language ProcessingShort Findings
- Is ChatGPT a General-Purpose Natural Language Processing Task Solver?Long Main
- Is ChatGPT a Good Causal Reasoner? A Comprehensive EvaluationLong Findings
- Is ChatGPT a Good Multi-Party Conversation Solver?Long Findings
- Is ChatGPT the ultimate Data Augmentation Algorithm?Short Findings
- Is Explanation the Cure? Misinformation Mitigation in the Short Term and Long TermShort Findings
- Is GPT-4 a Good Data Analyst?Long Findings
- Is Probing All You Need? Indicator Tasks as an Alternative to Probing Embedding SpacesLong Findings
- Is Robustness Transferable across Languages in Multilingual Neural Machine Translation?Long Findings
- Is a Prestigious Job the same as a Prestigious Country? A Case Study on Multilingual Sentence Embeddings and European CountriesShort Findings
- Is the Answer in the Text? Challenging ChatGPT with Evidence Retrieval from Instructive TextShort Findings
- Isotropic Representation Can Improve Zero-Shot Cross-Lingual Transfer on Multilingual Language ModelsLong Findings
- Isotropy-Enhanced Conditional Masked Language ModelsLong Findings
- It Ain't Over: A Multi-aspect Diverse Math Word Problem DatasetLong Main
- JASMINE: Arabic GPT Models for Few-Shot LearningLong Main
- JWSign: A Highly Multilingual Corpus of Bible Translations for more Diversity in Sign Language ProcessingLong Findings
- Joint Entity and Relation Extraction with Span Pruning and Hypergraph Neural NetworksLong Main
- Joint Geometrical and Statistical Domain Adaptation for Cross-domain Code Vulnerability DetectionLong Main
- Joint Semantic and Strategy Matching for Persuasive DialogueLong Findings
- JointMatch: A Unified Approach for Diverse and Collaborative Pseudo-Labeling to Semi-Supervised Text ClassificationLong Main
- Just Adjust One Prompt: Enhancing In-Context Dialogue Scoring via Constructing the Optimal Subgraph of Demonstrations and PromptsLong Main
- Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human FeedbackShort Main
- K-HATERS: A Hate Speech Detection Corpus in Korean with Target-Specific RatingsLong Findings
- KAPALM: Knowledge grAPh enhAnced Language Models for Fake News DetectionLong Findings
- KBioXLM: A Knowledge-anchored Biomedical Multilingual Pretrained Language ModelLong Findings
- KCTS: Knowledge-Constrained Tree Search Decoding with Token-Level Hallucination DetectionLong Main
- KEBAP: Korean Error Explainable Benchmark Dataset for ASR and Post-processingLong Main
- KEPL: Knowledge Enhanced Prompt Learning for Chinese Hypernym-Hyponym ExtractionLong Main
- KEPLET: Knowledge-Enhanced Pretrained Language Model with Topic Entity AwarenessLong Findings
- KG-GPT: A General Framework for Reasoning on Knowledge Graphs Using Large Language ModelsShort Findings
- KICGPT: Large Language Model with Knowledge in Context for Knowledge Graph CompletionLong Findings
- KRLS: Improving End-to-End Response Generation in Task Oriented Dialog with Reinforced Keywords LearningLong Main
- KeFVP: Knowledge-enhanced Financial Volatility PredictionLong Findings
- Knowledge Corpus Error in Question AnsweringShort Findings
- Knowledge Distillation ≈ Label Smoothing: Fact or Fallacy?Short Main
- Knowledge Graph Compression Enhances Diverse Commonsense GenerationLong Main
- Knowledge Rumination for Pre-trained Language ModelsLong Main
- Knowledge is a Region in Weight Space for Fine-tuned Language ModelsLong Findings
- Knowledge-Augmented Language Model VerificationLong Main
- Knowledge-Selective Pretraining for Attribute Value ExtractionLong Findings
- LACMA: Language-Aligning Contrastive Learning with Meta-Actions for Embodied Instruction FollowingLong Main
- LATENTLOGIC: Learning Logic Rules in Latent Space over Knowledge GraphsShort Findings
- LDM$^2$: A Large Decision Model Imitating Human Cognition with Dynamic Memory EnhancementLong Findings
- LEGO: A Multi-agent Collaborative Framework with Role-playing and Iterative Feedback for Causality Explanation GenerationLong Findings
- LEXTREME: A Multi-Lingual and Multi-Task Benchmark for the Legal DomainLong Findings
- LIMIT: Language Identification, Misidentification, and Translation using Hierarchical Models in 350+ LanguagesLong Main
- LINC: A Neurosymbolic Approach for Logical Reasoning by Combining Language Models with First-Order Logic ProversLong Main
- LLM aided semi-supervision for efficient Extractive Dialog SummarizationShort Findings
- LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language ModelsLong Main
- LLM-FP4: 4-Bit Floating-Point Quantized TransformersLong Main
- LLM-enhanced Self-training for Cross-domain Constituency ParsingLong Main
- LLM-in-the-loop: Leveraging Large Language Model for Thematic AnalysisShort Findings
- LLM-powered Data Augmentation for Enhanced Cross-lingual PerformanceLong Main
- LLMDet: A Third Party Large Language Models Generated Text Detection ToolLong Findings
- LLMLingua: Compressing Prompts for Accelerated Inference of Large Language ModelsLong Main
- LLMaAA: Making Large Language Models as Active AnnotatorsLong Findings
- LLMs -- the Good, the Bad or the Indispensable?: A Use Case on Legal Statute Prediction and Legal Judgment Prediction on Indian Court CasesShort Findings
- LM vs LM: Detecting Factual Errors via Cross ExaminationLong Main
- LMGQS: A Large-scale Dataset for Query-focused SummarizationLong Findings
- Label Words are Anchors: An Information Flow Perspective for Understanding In-Context LearningLong Main
- Language Guided Visual Question Answering: Elevate Your Multimodal Language Model Using Knowledge-Enriched PromptsShort Findings
- Language Model Quality Correlates with Psychometric Predictive Power in Multiple LanguagesShort Main
- Language Model is Suitable for Correction of Handwritten Mathematical Expressions RecognitionLong Main
- Language Models with RationalityLong Main
- Language Representation Projection: Can We Transfer Factual Knowledge across Languages in Multilingual Language Models?Short Main
- Language and Mental Health: Measures of Emotion Dynamics from Text as Linguistic Biosocial MarkersLong Main
- Language-Agnostic Bias Detection in Language Models with Bias ProbingLong Findings
- Large Language Model Is Not a Good Few-shot Information Extractor, but a Good Reranker for Hard Samples!Long Findings
- Large Language Models Are Better Adversaries: Exploring Generative Clean-Label Backdoor Attacks Against Text ClassifiersLong Findings
- Large Language Models Can Self-ImproveLong Main
- Large Language Models Know Your Contextual Search Intent: A Prompting Framework for Conversational SearchLong Findings
- Large Language Models Meet Harry Potter: A Dataset for Aligning Dialogue Agents with CharactersLong Findings
- Large Language Models Meet Open-World Intent Discovery and Recognition: An Evaluation of ChatGPTLong Main
- Large Language Models Only Pass Primary School Exams in Indonesia: A Comprehensive Test on IndoMMLULong Main
- Large Language Models and Multimodal Retrieval for Visual Word Sense DisambiguationLong Main
- Large Language Models are Better Reasoners with Self-VerificationLong Findings
- Large Language Models are Complex Table ParsersLong Main
- Large Language Models are Not Yet Human-Level Evaluators for Abstractive SummarizationLong Findings
- Large Language Models are Temporal and Causal Reasoners for Video Question AnsweringLong Main
- Large Language Models are biased to overestimate profoundnessShort Main
- Large Language Models as Source Planner for Personalized Knowledge-grounded DialoguesLong Findings
- Large Language Models: The Need for Nuance in Current Debates and a Pragmatic Perspective on UnderstandingLong Main
- Large-Scale and Multi-Perspective Opinion Summarization with Diverse Review SubsetsLong Findings
- Large-scale similarity search with Optimal TransportShort Main
- Larger Probes Tell a Different Story: Extending Psycholinguistic Datasets Via In-Context LearningShort Main
- Late Fusion of Transformers for Sentiment Analysis of Code-Switched DataShort Findings
- LayoutDIT: Layout-Aware End-to-End Document Image Translation with Multi-Step Conductive DecoderLong Findings
- Lazy-k Decoding: Constrained Decoding for Information ExtractionLong Main
- Leap-of-Thought: Accelerating Transformers via Dynamic Token RoutingLong Main
- Learn From One Specialized Sub-Teacher: One-to-One Mapping for Feature-Based Knowledge DistillationLong Findings
- Learn Your Tokens: Word-Pooled Tokenization for Language ModelingLong Findings
- Learn and Consolidate: Continual Adaptation for Zero-Shot and Multilingual Neural Machine TranslationLong Main
- Learning Co-Speech Gesture for Multimodal Aphasia Type DetectionLong Main
- Learning Dynamic Representations for Discourse Dependency ParsingLong Findings
- Learning Easily Updated General Purpose Text Representations with Adaptable Task-Specific PrefixShort Findings
- Learning Interpretable Style Embeddings via Prompting LLMsLong Findings
- Learning Knowledge-Enhanced Contextual Language Representations for Domain Natural Language UnderstandingLong Main
- Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment AnalysisLong Main
- Learning Preference Model for LLMs via Automatic Preference Data GenerationLong Main
- Learning Retrieval Augmentation for Personalized Dialogue GenerationLong Main
- Learning Semantic Role Labeling from Compatible Label SequencesLong Findings
- Learning from Mistakes via Cooperative Study Assistant for Large Language ModelsLong Main
- Learning the Visualness of Text Using Large Vision-Language ModelsLong Main
- Learning to Abstract with Nonparametric Variational Information BottleneckShort Findings
- Learning to Compose Representations of Different Encoder Layers towards Improving Compositional GeneralizationLong Findings
- Learning to Correct Noisy Labels for Fine-Grained Entity Typing via Co-Prediction Prompt TuningLong Findings
- Learning to Describe for Predicting Zero-shot Drug-Drug InteractionsLong Main
- Learning to Follow Object-Centric Image Editing Instructions FaithfullyLong Findings
- Learning to Predict Task Transferability via Soft PromptLong Main
- Learning to Rank Context for Named Entity Recognition Using a Synthetic DatasetLong Main
- Learning to Rank Generation with Pairwise Partial RewardsLong Main
- Learning to love diligent trolls: Accounting for rater effects in the dialogue safety taskShort Findings
- Learning under Label Proportions for Text ClassificationLong Findings
- Legally Enforceable Hate Speech Detection for Public ForumsLong Findings
- Length is a Curse and a Blessing for Document-level SemanticsLong Main
- Length-Adaptive Distillation: Customizing Small Language Model for Dynamic Token PruningLong Findings
- Less than One-shot: Named Entity Recognition via Extremely Weak SupervisionLong Findings
- Let GPT be a Math Tutor: Teaching Math Word Problem Solvers with Customized Exercise GenerationLong Main
- Let's Sample Step by Step: Adaptive-Consistency for Efficient Reasoning and Coding with LLMsLong Main
- Let's Synthesize Step by Step: Iterative Dataset Synthesis with Large Language Models by Extrapolating Errors from Small ModelsLong Findings
- Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-ThoughtLong Main
- Leveraging Contrastive Learning and Knowledge Distillation for Incomplete Modality Rumor DetectionLong Findings
- Leveraging GPT-4 for Automatic Translation Post-EditingLong Findings
- Leveraging Multiple Teachers for Test-Time Adaptation of Language-Guided ClassifiersLong Findings
- Leveraging Structured Information for Explainable Multi-hop Question Answering and ReasoningLong Findings
- Lexical Repetitions Lead to Rote Learning: Unveiling the Impact of Lexical Overlap in Train and Test Reference SummariesLong Findings
- Licon: A Diverse, Controllable and Challenging Linguistic Concept Learning BenchmarkLong Findings
- Lifelong Sequence Generation with Dynamic Module Expansion and AdaptationLong Main
- Linear-Time Modeling of Linguistic Structure: An Order-Theoretic PerspectiveLong Main
- Ling-CL: Understanding NLP Models through Linguistic CurriculaLong Main
- Linguistic Compression in Single-Sentence Human-Written SummariesLong Findings
- Linguistically Motivated Sign Language SegmentationLong Findings
- Linking Surface Facts to Large-Scale Knowledge GraphsLong Main
- Lion: Adversarial Distillation of Proprietary Large Language ModelsLong Main
- Localizing Active Objects from Egocentric Vision with Symbolic World KnowledgeLong Main
- Locally Differentially Private Document Generation Using Zero Shot PromptingLong Findings
- Location-Aware Visual Question Generation with Lightweight ModelsLong Main
- Log-FGAER: Logic-Guided Fine-Grained Address Entity Recognition from Multi-Turn Spoken DialogueLong Main
- LogiCoT: Logical Chain-of-Thought Instruction TuningLong Findings
- Logic Unveils Truth, While Disguise Obscures It: Transition Logic Augmented Response Selection for Multi-Turn DialogueLong Findings
- Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical ReasoningLong Findings
- LogicAttack: Adversarial Attacks for Evaluating Logical Consistency of Natural Language InferenceShort Findings
- Long-Horizon Dialogue Understanding for Role Identification in the Game of Avalon with Large Language ModelsLong Findings
- Long-Range Language Modeling with Selective CacheLong Findings
- Longtriever: a Pre-trained Long Text Encoder for Dense Document RetrievalLong Main
- Look-back Decoding for Open-Ended Text GenerationLong Main
- Loose lips sink ships: Mitigating Length Bias in Reinforcement Learning from Human FeedbackLong Findings
- Low-Resource Comparative Opinion Quintuple Extraction by Data Augmentation with PromptingShort Findings
- M$^3$Seg: A Maximum-Minimum Mutual Information Paradigm for Unsupervised Topic Segmentation in ASR TranscriptsShort Main
- M2C: Towards Automatic Multimodal Manga ComplementShort Findings
- M2DF: Multi-grained Multi-curriculum Denoising Framework for Multimodal Aspect-based Sentiment AnalysisLong Main
- MADNet: Maximizing Addressee Deduction Expectation for Multi-Party Conversation GenerationLong Main
- MAF: Multi-Aspect Feedback for Improving Reasoning in Large Language ModelsLong Main
- MAGNIFICo: Evaluating the In-Context Learning Ability of Large Language Models to Generalize to Novel InterpretationsLong Main
- MAPO: Boosting Large Language Model Performance with Model-Adaptive Prompt OptimizationLong Findings
- MCC-KD: Multi-CoT Consistent Knowledge DistillationLong Findings
- MCLF: A Multi-grained Contrastive Learning Framework for ASR-robust Spoken Language UnderstandingLong Findings
- MEEP: Is this Engaging? Prompting Large Language Models for Dialogue Evaluation in Multilingual SettingsLong Findings
- MEGA: Multilingual Evaluation of Generative AILong Main
- MEGClass: Extremely Weakly Supervised Text Classification via Mutually-Enhancing Text GranularitiesLong Findings
- MILDSum: A Novel Benchmark Dataset for Multilingual Summarization of Indian Legal Case JudgmentsShort Main
- MISCA: A Joint Model for Multiple Intent Detection and Slot Filling with Intent-Slot Co-AttentionLong Findings
- MM-Reasoner: A Multi-Modal Knowledge-Aware Framework for Knowledge-Based Visual Question AnsweringLong Findings
- MMNMT: Modularizing Multilingual Neural Machine Translation with Flexibly Assembled MoE and Dense BlocksLong Main
- MPrompt: Exploring Multi-level Prompt Tuning for Machine Reading ComprehensionLong Findings
- MProto: Multi-Prototype Network with Denoised Optimal Transport for Distantly Supervised Named Entity RecognitionLong Main
- MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop QuestionsLong Main
- MRRL: Modifying the Reference via Reinforcement Learning for Non-Autoregressive Joint Multiple Intent Detection and Slot FillingLong Findings
- MSCFFN: A New FFN with Multi-Space Cross to Accelerate TransformerShort Findings
- MT2: Towards a Multi-Task Machine Translation Model with Translation-Specific In-Context LearningLong Main
- MTGER: Multi-view Temporal Graph Enhanced Temporal Reasoning over Time-Involved DocumentLong Findings
- MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection BenchmarkLong Main
- MUX-PLMs: Data Multiplexing for High-throughput Language ModelsLong Findings
- MaNtLE: Model-agnostic Natural Language ExplainerLong Main
- MaXM: Towards Multilingual Visual Question AnsweringLong Findings
- MacLaSa: Multi-Aspect Controllable Text Generation via Efficient Sampling from Compact Latent SpaceLong Findings
- Macedon: Minimizing Representation Coding Rate Reduction for Cross-Lingual Natural Language UnderstandingLong Findings
- Machine Reading Comprehension using Case-based ReasoningLong Findings
- MailEx: Email Event and Argument ExtractionLong Main
- Make Every Example Count: On the Stability and Utility of Self-Influence for Learning from Noisy NLP DatasetsLong Main
- Make Your Decision Convincing! A Unified Two-Stage Framework: Self-Attribution and Decision-MakingLong Findings
- Making Body Movement in Sign Language Corpus Accessible for Linguists and Machines with Three-Dimensional Normalization of MediaPipeLong Findings
- Making Large Language Models Better Data CreatorsLong Main
- Mandarin classifier systems optimize to accommodate communicative pressuresLong Findings
- Manifold-Preserving Transformers are Effective for Short-Long Range EncodingLong Findings
- Manipulating the Perceived Personality Traits of Language ModelsLong Findings
- MarkQA: A large scale KBQA dataset with numerical reasoningLong Main
- Masked Path Modeling for Vision-and-Language NavigationLong Findings
- MathDial: A Dialogue Tutoring Dataset with Rich Pedagogical Properties Grounded in Math Reasoning ProblemsLong Findings
- MeaeQ: Mount Model Extraction Attacks with Efficient QueriesLong Main
- Measure Children's Mindreading Ability with Machine ReadingLong Findings
- Measuring Faithful and Plausible Visual Grounding in VQALong Findings
- Measuring Pointwise $\mathcal{V}$-Usable Information In-Context-lyLong Findings
- Measuring and Mitigating Constraint Violations of In-Context Learning for Utterance-to-API Semantic ParsingLong Findings
- Measuring and Narrowing the Compositionality Gap in Language ModelsLong Findings
- Measuring bias in Instruction-Following models with P-ATLong Findings
- Measuring the Knowledge Acquisition-Utilization Gap in Pretrained Language ModelsLong Findings
- MedEval: A Multi-Level, Multi-Task, and Multi-Domain Medical Benchmark for Language Model EvaluationLong Main
- MediaHG: Rethinking Eye-catchy Features in Social Media Headline GenerationLong Main
- Medical Text Simplification: Optimizing for Readability with Unlikelihood Training and Reranked Beam Search DecodingShort Findings
- MemeCap: A Dataset for Captioning and Interpreting MemesLong Main
- Memorisation Cartography: Mapping out the Memorisation-Generalisation Continuum in Neural Machine TranslationLong Main
- Memory-Based Invariance Learning for Out-of-Domain Text ClassificationLong Main
- MenatQA: A New Dataset for Testing the Temporal Comprehension and Reasoning Abilities of Large Language ModelsLong Findings
- Merging Experts into One: Improving Computational Efficiency of Mixture of ExpertsShort Main
- Merging Generated and Retrieved Knowledge for Open-Domain QALong Main
- Meta-Learning Online Adaptation of Language ModelsLong Main
- Meta-Learning of Prompt Generation for Lightweight Prompt Engineering on Language-Model-as-a-ServiceLong Findings
- MetaReVision: Meta-Learning with Retrieval for Visually Grounded Compositional Concept AcquisitionLong Findings
- Methodological Insights in Detecting Subtle Semantic Shifts with Contextualized and Static Language ModelsLong Findings
- Mind the Gap Between Conversations for Improved Long-Term Dialogue GenerationLong Findings
- Mind the Gap: Automated Corpus Creation for Enthymeme Detection and Reconstruction in Learner ArgumentsLong Findings
- MindGames: Targeting Theory of Mind in Large Language Models with Dynamic Epistemic Modal LogicShort Findings
- MingOfficial: A Ming Official Career Dataset and a Historical Context-Aware Representation Learning FrameworkLong Main
- Miracle: Towards Personalized Dialogue Generation with Latent-Space Multiple Personal Attribute ControlLong Findings
- Mirages. On Anthropomorphism in Dialogue SystemsLong Main
- Mirror: A Universal Framework for Various Information Extraction TasksLong Main
- Misery Loves Complexity: Exploring Linguistic Complexity in the Context of Emotion DetectionLong Findings
- Mitigating Backdoor Poisoning Attacks through the Lens of Spurious CorrelationLong Main
- Mitigating Biases in Hate Speech Detection from A Causal PerspectiveLong Findings
- Mitigating Data Imbalance and Representation Degeneration in Multilingual Machine TranslationLong Findings
- Mitigating Framing Bias with Polarity Minimization LossShort Findings
- Mitigating Intrinsic Named Entity-Related Hallucinations of Abstractive Text SummarizationLong Findings
- Mitigating Temporal Misalignment by Discarding Outdated FactsLong Main
- MixEdit: Revisiting Data Augmentation and Beyond for Grammatical Error CorrectionLong Findings
- MixTEA: Semi-supervised Entity Alignment with Mixture TeachingLong Findings
- Mixture of Soft Prompts for Controllable Data GenerationLong Findings
- Mixture-of-Linguistic-Experts Adapters for Improving and Interpreting Pre-trained Language ModelsLong Findings
- MoPe: Model Perturbation based Privacy Attacks on Language ModelsLong Main
- MoT: Memory-of-Thought Enables ChatGPT to Self-ImproveLong Main
- Model-tuning Via Prompts Makes NLP Models Adversarially RobustLong Main
- Modeling Conceptual Attribute Likeness and Domain Inconsistency for Metaphor DetectionLong Main
- Modeling Empathic Similarity in Personal NarrativesLong Main
- Modeling Highlighting of Metaphors in Multitask Contrastive Learning ParadigmsLong Findings
- Modeling Legal Reasoning: LM Annotation at the Edge of Human AgreementLong Main
- Models See Hallucinations: Evaluating the Factuality in Video CaptioningLong Main
- MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal AdapterLong Main
- Monte Carlo Thought Search: Large Language Model Querying for Complex Scientific Reasoning in Catalyst DesignShort Findings
- MoqaGPT : Zero-Shot Multi-modal Open-domain Question Answering with Large Language ModelLong Findings
- More than Votes? Voting and Language based Partisanship in the US Supreme CourtShort Findings
- MuG: A Multimodal Classification Benchmark on Game Data with Tabular, Textual, and Visual FieldsLong Findings
- Mulan: A Multi-Level Alignment Model for Video Question AnsweringLong Findings
- Multi-Defendant Legal Judgment Prediction via Hierarchical ReasoningLong Findings
- Multi-Granularity Information Interaction Framework for Incomplete Utterance RewritingShort Findings
- Multi-Modal Knowledge Graph Transformer Framework for Multi-Modal Entity AlignmentLong Findings
- Multi-Source Multi-Type Knowledge Exploration and Exploitation for Dialogue GenerationLong Main
- Multi-Source Probing for Open-Domain Conversational UnderstandingLong Main
- Multi-Stage Pre-training Enhanced by ChatGPT for Multi-Scenario Multi-Domain Dialogue SummarizationLong Findings
- Multi-Task Knowledge Distillation with Embedding Constraints for Scholarly Keyphrase Boundary ClassificationLong Main
- Multi-Task Learning of Query Generation and Classification for Generative Conversational Question RewritingLong Findings
- Multi-User MultiWOZ: Task-Oriented Dialogues among Multiple UsersLong Findings
- Multi-label and Multi-target Sampling of Machine Annotation for Computational Stance DetectionShort Findings
- Multi-level Adaptive Contrastive Learning for Knowledge Internalization in Dialogue GenerationLong Main
- Multi-level Contrastive Learning for Script-based Character UnderstandingLong Main
- Multi-step Jailbreaking Privacy Attacks on ChatGPTLong Findings
- Multi-view Contrastive Learning for Entity Typing over Knowledge GraphsLong Main
- MultiCMET: A Novel Chinese Benchmark for Understanding Multimodal MetaphorLong Findings
- MultiCoNER v2: a Large Multilingual dataset for Fine-grained and Noisy Named Entity RecognitionShort Findings
- MultiTurnCleanup: A Benchmark for Multi-Turn Spoken Conversational Transcript CleanupShort Main
- Multilingual Coarse Political Stance Classification of Media. The Editorial Line of a ChatGPT and Bard NewspaperShort Findings
- Multilingual Generation and Answering of Questions from Texts and Knowledge GraphsLong Findings
- Multilingual Holistic Bias: Extending Descriptors and Patterns to Unveil Demographic Biases in Languages at ScaleLong Main
- Multilingual Large Language Models Are Not (Yet) Code-SwitchersLong Main
- Multilingual Lottery Tickets to Pretrain Language ModelsLong Findings
- Multilingual Pixel Representations for Translation and Effective Cross-lingual TransferLong Main
- Multilingual Simplification of Medical TextsLong Main
- Multilingual estimation of political-party positioning: From label aggregation to long-input TransformersLong Main
- Multimodal Automated Fact-Checking: A SurveyLong Findings
- Multimodal Embodied Plan Prediction Augmented with Synthetic Embodied DialogueLong Main
- Multitask Multimodal Prompted Training for Interactive Embodied Task CompletionLong Main
- Multiview Clickbait Detection via Jointly Modeling Subjective and Objective PreferenceLong Findings
- NAIL: Lexical Retrieval Indices with Efficient Non-Autoregressive DecodersLong Main
- NASH: A Simple Unified Framework of Structured Pruning for Accelerating Encoder-Decoder Language ModelsLong Findings
- NERetrieve: Dataset for Next Generation Named Entity Recognition and RetrievalLong Findings
- NERvous About My Health: Constructing a Bengali Medical Named Entity Recognition DatasetShort Findings
- NEWTON: Are Large Language Models Capable of Physical Reasoning?Long Findings
- NLI4CT: Multi-Evidence Natural Language Inference for Clinical Trial ReportsLong Main
- NLMs: Augmenting Negation in Language ModelsLong Findings
- NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each BenchmarkShort Findings
- NORMSAGE: Multi-Lingual Multi-Cultural Norm Discovery from Conversations On-the-FlyLong Main
- NameGuess: Column Name Expansion for Tabular DataLong Main
- Narrative Order Aware Story Generation via Bidirectional Pretraining Model with Optimal Transport RewardLong Findings
- Narrative Style and the Spread of Health Misinformation on TwitterLong Findings
- NarrativeXL: a Large-scale Dataset for Long-Term Memory ModelsLong Findings
- Natural Disaster Tweets Classification Using Multimodal DataLong Main
- Natural Language Annotations for Reasoning about Program SemanticsShort Findings
- Natural Language Decompositions of Implicit Content Enable Better Text RepresentationsLong Main
- Natural Response Generation for Chinese Reading ComprehensionLong Findings
- Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language ModelsLong Main
- Nearest Neighbor Machine Translation is Meta-Optimizer on Output Projection LayerLong Main
- NeuSTIP: A Neuro-Symbolic Model for Link and Time Prediction in Temporal Knowledge GraphsLong Main
- Neuro-Symbolic Sentiment Analysis with Dynamic Word Sense DisambiguationLong Findings
- New Datasets and Controllable Iterative Data Augmentation Method for Code-switching ASR Error CorrectionLong Findings
- No offence, Bert - I insult only humans! Multilingual sentence-level attack on toxicity detection networksShort Findings
- Noise-Robust Fine-Tuning of Pretrained Language Models via External GuidanceLong Findings
- Noise-Robust Semi-Supervised Learning for Distantly Supervised Relation ExtractionLong Findings
- Noisy Exemplars Make Large Language Models More Robust: A Domain-Agnostic Behavioral AnalysisShort Main
- Noisy Pair Corrector for Dense RetrievalLong Findings
- Noisy Self-Training with Synthetic Queries for Dense RetrievalLong Findings
- Non-Autoregressive Document-Level Machine TranslationLong Findings
- Non-Autoregressive Math Word Problem Solver with Unified Tree StructureLong Main
- Non-Autoregressive Sentence OrderingLong Findings
- Non-Compositionality in Sentiment: New Data and AnalysesShort Findings
- Non-Programmers Can Label Programs Indirectly via Active Examples: A Case Study with Text-to-SQLLong Main
- Non-autoregressive Streaming Transformer for Simultaneous TranslationLong Main
- Non-autoregressive Text Editing with Copy-aware Latent AlignmentsLong Main
- Non-compositional Expression Generation Based on Curriculum Learning and Continual LearningLong Findings
- Non-parallel Accent Transfer based on Fine-grained Controllable Accent ModellingLong Findings
- Norm of Word Embedding Encodes Information GainLong Main
- NormDial: A Comparable Bilingual Synthetic Dialog Dataset for Modeling Social Norm Adherence and ViolationShort Main
- Normal-Abnormal Decoupling Memory for Medical Report GenerationLong Findings
- Not All Demonstration Examples are Equally Beneficial: Reweighting Demonstration Examples for In-Context LearningLong Findings
- Not All Languages Are Created Equal in LLMs: Improving Multilingual Capability by Cross-Lingual-Thought PromptingLong Findings
- Not all quantifiers are equal: Probing Transformer-based language models' understanding of generalised quantifiersLong Main
- NovaCOMET: Open Commonsense Foundation Models with Symbolic Knowledge DistillationLong Findings
- Novel Relation Detection: Discovering Unknown Relation Types via Multi-Strategy Self-Supervised LearningLong Findings
- Novel Slot Detection With an Incremental SettingLong Findings
- ORCHID: A Chinese Debate Corpus for Target-Independent Stance Detection and Argumentative Dialogue SummarizationLong Main
- On Bilingual Lexicon Induction with Large Language ModelsLong Main
- On Evaluation of Bangla Word AnalogiesShort Main
- On Event Individuation for Document-Level Information ExtractionShort Findings
- On General Language UnderstandingShort Findings
- On Robustness of Finetuned Transformer-based NLP ModelsLong Findings
- On Surgical Fine-tuning for Language EncodersShort Findings
- On Task-personalized Multimodal Few-shot Learning for Visually-rich Document Entity RetrievalLong Findings
- On Uncertainty Calibration and Selective Generation in Probabilistic Neural Summarization: A Benchmark StudyShort Findings
- On the Automatic Generation and Simplification of Children's StoriesLong Main
- On the Benefits of Learning to Route in Mixture-of-Experts ModelsLong Main
- On the Calibration of Large Language Models and AlignmentLong Findings
- On the Challenges of Using Black-Box APIs for Toxicity Evaluation in ResearchLong Main
- On the Dimensionality of Sentence EmbeddingsLong Findings
- On the Impact of Cross-Domain Data on German Language ModelsLong Findings
- On the Representational Capacity of Recurrent Neural Language ModelsLong Main
- On the Risk of Misinformation Pollution with Large Language ModelsLong Findings
- On the Transferability of Visually Grounded PCFGsLong Findings
- On the Zero-Shot Generalization of Machine-Generated Text DetectorsShort Findings
- Once Upon a ${\it Time}$ in ${\it Graph}$: Relative-Time Pretraining for Complex Temporal ReasoningLong Main
- Once is Enough: A Light-Weight Cross-Attention for Fast Sentence Pair ModelingShort Main
- One For All $\&$ All For One: Bypassing Hyperparameter Tuning with Model Averaging for Cross-Lingual TransferShort Findings
- One-Model-Connects-All: A Unified Graph Pre-Training Model for Online Community ModelingLong Findings
- Oolong: Investigating What Makes Transfer Learning Hard with Controlled StudiesShort Main
- Open Domain Multi-document Summarization: A Comprehensive Study of Model Brittleness under RetrievalLong Findings
- Open Information Extraction via ChunksLong Main
- Open-Ended Instructable Embodied Agents with Memory-Augmented Large Language ModelsLong Findings
- Open-ended Commonsense Reasoning with Unrestricted Answer CandidatesLong Findings
- Open-source Large Language Models are Strong Zero-shot Query Likelihood Models for Document RankingShort Findings
- Open-world Semi-supervised Generalized Relation Discovery Aligned in a Real-world SettingLong Main
- OpenAsp: A Benchmark for Multi-document Open Aspect-based SummarizationLong Main
- Optimized Tokenization for Transcribed Error CorrectionLong Main
- Optimizing Retrieval-augmented Reader Models via Token EliminationLong Main
- Orca: A Few-shot Benchmark for Chinese Conversational Machine Reading ComprehensionLong Findings
- Orthogonal Subspace Learning for Language Model Continual LearningLong Findings
- OssCSE: Overcoming Surface Structure Bias in Contrastive Learning for Unsupervised Sentence EmbeddingLong Main
- Out-of-Distribution Generalization in Natural Language Processing: Past, Present, and FutureLong Main
- Outlier Dimensions Encode Task Specific KnowledgeShort Main
- Outlier Suppression+: Accurate quantization of large language models by equivalent and effective shifting and scalingLong Main
- PAC-tuning: Fine-tuning Pre-trained Language Models with PAC-driven Perturbed Gradient DescentLong Main
- PALS: Personalized Active Learning for Subjective Tasks in NLPLong Main
- PARROT: Zero-Shot Narrative Reading Comprehension via Parallel ReadingLong Findings
- PCMID: Multi-Intent Detection through Supervised Prototypical Contrastive LearningLong Findings
- PEFTDebias : Capturing debiasing information using PEFTsShort Main
- PHD: Pixel-Based Language Modeling of Historical DocumentsLong Main
- PIEClass: Weakly-Supervised Text Classification with Prompting and Noise-Robust Iterative Ensemble TrainingLong Main
- PIVOINE: Instruction Tuning for Open-world Entity ProfilingLong Findings
- PMIndiaSum: Multilingual and Cross-lingual Headline Summarization for Languages in IndiaLong Findings
- POE: Process of Elimination for Multiple Choice ReasoningShort Main
- POSQA: Probe the World Models of LLMs with Size ComparisonsLong Findings
- PR-MCS: Perturbation Robust Metric for MultiLingual Image CaptioningLong Findings
- PRCA: Fitting Black-Box Large Language Models for Retrieval Question Answering via Pluggable Reward-Driven Contextual AdapterLong Main
- PRESTO: A Multilingual Dataset for Parsing Realistic Task-Oriented DialogsLong Main
- PROSE: A Pronoun Omission Solution for Chinese-English Spoken Language TranslationLong Main
- PROTEGE: Prompt-based Diverse Question Generation from Web ArticlesLong Findings
- PTP: Boosting Stability and Performance of Prompt Tuning with Perturbation-Based RegularizerLong Main
- PUNR: Pre-training with User Behavior Modeling for News RecommendationLong Findings
- PaRaDe: Passage Ranking using Demonstrations with LLMsShort Findings
- Parameter Efficient Multi-task Fine-tuning by Learning to Transfer Token-wise PromptsLong Findings
- Parameter-Efficient Cross-lingual Transfer of Vision and Language Models via Translation-based AlignmentLong Findings
- Parameter-Efficient Language Model Tuning with Active Learning in Low-Resource SettingsLong Main
- Parameter-Efficient Prompt Tuning Makes Generalized and Calibrated Neural Text RetrieversLong Findings
- Parameter-efficient Tuning for Large Language Model without Calculating Its GradientsLong Main
- Paraphrase Types for Generation and DetectionLong Main
- ParroT: Translating during Chat using Large Language Models tuned with Human Translation and FeedbackLong Findings
- Penalty Decoding: Well Suppress the Self-Reinforcement Effect in Open-Ended Text GenerationShort Main
- People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language DetectionLong Main
- Perceptual Structure in the absence of grounding: the impact of abstractedness and subjectivity in color language for LLMsShort Findings
- PersonaLM: Language Model Personalization via Domain-distributed Span Aggregated K-Nearest N-gram Retrieval AugmentationLong Findings
- Personalized Distillation: Empowering Open-Sourced LLMs with Adaptive Learning for Code GenerationLong Main
- PerturbScore: Connecting Discrete and Continuous Perturbations in NLPLong Findings
- Physician Detection of Clinical Harm in Machine Translation: Quality Estimation Aids in Reliance and Backtranslation Identifies Critical ErrorsLong Main
- Pit One Against Many: Leveraging Attention-head Embeddings for Parameter-efficient Multi-head AttentionLong Findings
- PivotFEC: Enhancing Few-shot Factual Error Correction with a Pivot Task Approach using Large Language ModelsLong Findings
- Plan, Verify and Switch: Integrated Reasoning with Diverse X-of-ThoughtsLong Main
- PlugMed: Improving Specificity in Patient-Centered Medical Dialogue Generation using In-Context LearningLong Findings
- Pointwise Mutual Information Based Metric and Decoding Strategy for Faithful Generation in Document Grounded DialogsLong Main
- Poisoning Retrieval Corpora by Injecting Adversarial PassagesShort Main
- Polar Ducks and Where to Find Them: Enhancing Entity Linking with Duck Typing and Polar Box EmbeddingsLong Main
- Polyglot or Not? Measuring Multilingual Encyclopedic Knowledge in Foundation ModelsShort Main
- Post-hoc Utterance Refining Method by Entity Mining for Faithful Knowledge Grounded ConversationsLong Main
- Practical Computational Power of Linear Transformers and Their Recurrent and Self-Referential ExtensionsShort Main
- Pragmatic Reasoning Unlocks Quantifier Semantics for Foundation ModelsLong Main
- Pragmatics in Language Grounding: Phenomena, Tasks, and Modeling ApproachesLong Findings
- Pre-Trained Language Models Augmented with Synthetic Scanpaths for Natural Language UnderstandingShort Main
- Pre-trained Speech Processing Models Contain Human-Like Biases that Propagate to Speech Emotion RecognitionLong Findings
- Pre-training Intent-Aware Encoders for Zero- and Few-Shot Intent ClassificationLong Main
- Pre-training Language Models for Comparative ReasoningLong Main
- Pre-training Multi-task Contrastive Learning Models for Scientific Literature UnderstandingLong Findings
- PreWoMe: Exploiting Presuppositions as Working Memory for Long Form Question AnsweringShort Main
- Precedent-Enhanced Legal Judgment Prediction with LLM and Domain-Model CollaborationLong Main
- Predict the Future from the Past? On the Temporal Data Distribution Shift in Financial Sentiment ClassificationsLong Main
- Predictive Chemistry Augmented with Text RetrievalLong Main
- Prefix-Tuning Based Unsupervised Text Style TransferLong Findings
- Preserving Knowledge Invariance: Rethinking Robustness Evaluation of Open Information ExtractionLong Main
- Preserving Privacy Through Dememorization: An Unlearning Technique For Mitigating Memorization Risks In Language ModelsLong Main
- Pretraining Language Models with Text-Attributed Heterogeneous GraphsLong Findings
- Primacy Effect of ChatGPTShort Main
- Privacy Implications of Retrieval-Based Language ModelsLong Main
- Probabilistic Tree-of-thought Reasoning for Answering Knowledge-intensive Complex QuestionsLong Findings
- Probing LLMs for Joint Encoding of Linguistic CategoriesLong Findings
- Probing LLMs for hate speech detection: strengths and vulnerabilitiesLong Findings
- Probing Representations for Document-level Event ExtractionShort Findings
- Probing the “Creativity” of Large Language Models: Can models produce divergent semantic association?Short Findings
- Program Translation via Code DistillationLong Main
- Promoting Topic Coherence and Inter-Document Consorts in Multi-Document Summarization via Simplicial Complex and Sheaf GraphLong Main
- Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language ModelsLong Main
- Prompt-Based Editing for Text Style TransferLong Findings
- Prompt-Based Monte-Carlo Tree Search for Goal-oriented Dialogue Policy PlanningShort Main
- Prompt-based Logical Semantics Enhancement for Implicit Discourse Relation RecognitionLong Main
- PromptARA: Improving Deep Representation in Hybrid Automatic Readability Assessment with Prompt and Orthogonal ProjectionLong Findings
- PromptMix: A Class Boundary Augmentation Method for Large Language Model DistillationLong Main
- PromptST: Abstract Prompt Learning for End-to-End Speech TranslationLong Main
- Prompting ChatGPT in MNER: Enhanced Multimodal Named Entity Recognition with Auxiliary Refined KnowledgeLong Findings
- Prompting Large Language Models with Chain-of-Thought for Few-Shot Knowledge Base Question GenerationLong Main
- Prompting Scientific Names for Zero-Shot Species RecognitionShort Main
- Prompting and Evaluating Large Language Models for Proactive Dialogues: Clarification, Target-guided, and Non-collaborationLong Findings
- Prompting is not a substitute for probability measurements in large language modelsLong Main
- Prompting with Pseudo-Code InstructionsLong Main
- Proto-lm: A Prototypical Network-Based Framework for Built-in Interpretability in Large Language ModelsLong Findings
- Prototype-based HyperAdapter for Sample-Efficient Multi-task TuningLong Main
- Pseudointelligence: A Unifying Lens on Language Model EvaluationShort Findings
- PsyAttention: Psychological Attention Model for Personality DetectionLong Findings
- PsyCoT: Psychological Questionnaire as Powerful Chain-of-Thought for Personality DetectionLong Findings
- Pushdown Layers: Encoding Recursive Structure in Transformer Language ModelsLong Main
- QA-NatVer: Question Answering for Natural Logic-based Fact VerificationLong Main
- QADYNAMICS: Training Dynamics-Driven Synthetic QA Diagnostic for Zero-Shot Commonsense Question AnsweringShort Findings
- QTSumm: Query-Focused Summarization over Tabular DataLong Main
- QUADRo: Dataset and Models for QUestion-Answer Database RetrievalLong Findings
- QUDeval: The Evaluation of Questions Under Discussion Discourse ParsingLong Main
- Qualitative Code Suggestion: A Human-Centric Approach to Qualitative CodingLong Findings
- Quality Estimation-Assisted Automatic Post-EditingLong Findings
- Quantifying Character Similarity with Vision TransformersLong Main
- Quantifying the Dialect Gap and its Correlates Across LanguagesLong Findings
- Quantifying the redundancy between prosody and textLong Main
- Query Rewriting in Retrieval-Augmented Large Language ModelsLong Main
- Query-as-context Pre-training for Dense Passage RetrievalLong Main
- Query-based Image Captioning from Multi-context 360° ImagesLong Findings
- Query2Triple: Unified Query Encoding for Answering Diverse Complex Queries over Knowledge GraphsLong Findings
- Query2doc: Query Expansion with Large Language ModelsShort Main
- Question Answering as Programming for Solving Time-Sensitive QuestionsLong Main
- Quick Back-Translation for Unsupervised Machine TranslationLong Findings
- R$^3$ Prompting: Review, Rephrase and Resolve for Chain-of-Thought Reasoning in Large Language Models under Noisy ContextLong Findings
- R2H: Building Multimodal Navigation Helpers that Respond to Help RequestsLong Main
- RAPL: A Relation-Aware Prototype Learning Approach for Few-Shot Document-Level Relation ExtractionLong Main
- RECAL: Sample-Relation Guided Confidence Calibration over Tabular DataLong Findings
- RECAP: Towards Precise Radiology Report Generation via Dynamic Disease Progression ReasoningLong Findings
- ROBBIE: Robust Bias Evaluation of Large Generative Language ModelsLong Main
- ROME: Evaluating Pre-trained Vision-Language Models on Reasoning beyond Visual Common SenseLong Findings
- RSVP: Customer Intent Detection via Agent Response Contrastive and Generative Pre-TrainingLong Findings
- RWKV: Reinventing RNNs for the Transformer EraLong Findings
- RainProof: An Umbrella to Shield Text Generator from Out-Of-Distribution DataLong Main
- Random Entity Quantization for Parameter-Efficient Compositional Knowledge Graph RepresentationLong Main
- Ranking LLM-Generated Loop Invariants for Program VerificationShort Findings
- Rather a Nurse than a Physician - Contrastive Explanations under InvestigationLong Main
- Rationale-Enhanced Language Models are Better Continual Relation LearnersShort Main
- Re$^3$Dial: Retrieve, Reorganize and Rescale Conversations for Long-Turn Open-Domain Dialogue Pre-trainingLong Main
- Re-Examining Summarization Evaluation across Multiple Quality CriteriaShort Findings
- Re-ViLM: Retrieval-Augmented Visual Language Model for Zero and Few-Shot Image CaptioningLong Findings
- Re-weighting Tokens: A Simple and Effective Active Learning Strategy for Named Entity RecognitionShort Findings
- ReCEval: Evaluating Reasoning Chains via Correctness and InformativenessLong Main
- ReFSQL: A Retrieval-Augmentation Framework for Text-to-SQL GenerationLong Findings
- ReLM: Leveraging Language Models for Enhanced Chemical Reaction PredictionShort Findings
- ReSee: Responding through Seeing Fine-grained Visual Knowledge in Open-domain DialogueLong Main
- ReTAG: Reasoning Aware Table to Analytic Text GenerationLong Main
- ReadPrompt: A Readable Prompting Method for Reliable Knowledge ProbingLong Findings
- Reading Books is Great, But Not if You Are Driving! Visually Grounded Reasoning about Defeasible Commonsense NormsLong Main
- Reading Order Matters: Information Extraction from Visually-rich Documents by Token Path PredictionLong Main
- RealBehavior: A Framework for Faithfully Characterizing Foundation Models’ Human-like Behavior MechanismsLong Findings
- Reasoning Makes Good Annotators : An Automatic Task-specific Rules Distilling Framework for Low-resource Relation ExtractionLong Findings
- Reasoning about Ambiguous Definite DescriptionsShort Findings
- Reasoning with Language Model is Planning with World ModelLong Main
- ReasoningLM: Enabling Structural Subgraph Reasoning in Pre-trained Language Models for Question Answering over Knowledge GraphLong Main
- Reconstruct Before Summarize: An Efficient Two-Step Framework for Condensing and Summarizing Meeting TranscriptsLong Main
- Recurrent Neural Language Models as Probabilistic Finite-state AutomataLong Main
- Reduce Human Labor On Evaluating Conversational Information Retrieval System: A Human-Machine Collaboration ApproachLong Main
- Reducing Sequence Length by Predicting Edit Spans with Large Language ModelsLong Main
- Reducing Spurious Correlations in Aspect-based Sentiment Analysis with Explanation from Large Language ModelsLong Findings
- RefGPT: Dialogue Generation of GPT, by GPT, and for GPTLong Findings
- Referring Image Segmentation via Joint Mask Contextual Embedding Learning and Progressive Alignment NetworkLong Main
- RegaVAE: A Retrieval-Augmented Gaussian Mixture Variational Auto-Encoder for Language ModelingLong Findings
- Regulation and NLP (RegNLP): Taming Large Language ModelsLong Main
- Reinforced Target-driven Conversational PromotionLong Main
- Reinforcement Replaces Supervision: Query focused Summarization using Deep Reinforcement LearningLong Main
- Relation-Aware Question Answering for Heterogeneous Knowledge GraphsLong Findings
- Remember what you did so you know what to do nextShort Findings
- RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and GenerationLong Main
- Representation Projection Invariance Mitigates Representation CollapseLong Findings
- Representative Demonstration Selection for In-Context Learning with Two-Stage Determinantal Point ProcessLong Main
- Representativeness as a Forgotten Lesson for Multilingual and Code-switched Data Collection and PreparationLong Findings
- Responsible AI Considerations in Text Summarization Research: A Review of Current PracticesLong Findings
- Rethinking Model Selection and Decoding for Keyphrase Generation with Pre-trained Sequence-to-Sequence ModelsLong Main
- Rethinking Negative Pairs in Code SearchLong Main
- Rethinking Word-Level Auto-Completion in Computer-Aided TranslationLong Main
- Rethinking and Improving Multi-task Learning for End-to-end Speech TranslationLong Main
- Rethinking the Construction of Effective Metrics for Understanding the Mechanisms of Pretrained Language ModelsLong Findings
- Rethinking the Evaluation for Conversational Recommendation in the Era of Large Language ModelsLong Main
- Retrieval-Augmented Few-shot Text ClassificationLong Findings
- Retrieval-Augmented Parsing for Complex Graphs by Exploiting Structure and UncertaintyLong Findings
- Retrieval-Generation Alignment for End-to-End Task-Oriented Dialogue SystemLong Main
- Retrieval-based Knowledge Transfer: An Effective Approach for Extreme Large Language Model CompressionLong Findings
- Retrieving Multimodal Information for Augmented Generation: A SurveyLong Findings
- Retrofitting Light-weight Language Models for Emotions using Supervised Contrastive LearningLong Main
- Revisiting Automated Topic Model Evaluation with Large Language ModelsShort Main
- Revisiting Block-based Quantisation: What is Important for Sub-8-bit LLM Inference?Long Main
- Revisiting De-Identification of Electronic Medical Records: Evaluation of Within- and Cross-Hospital GeneralizationShort Main
- Revisiting Entropy Rate Constancy in TextLong Findings
- Revisiting Instruction Fine-tuned Model Evaluation to Guide Industrial ApplicationsShort Main
- Revisiting Large Language Models as Zero-shot Relation ExtractorsLong Findings
- Revisiting Machine Translation for Cross-lingual ClassificationLong Main
- Revisiting Source Context in Nearest Neighbor Machine TranslationLong Main
- Revisiting Sparse Retrieval for Few-shot Entity LinkingShort Main
- Revisiting the Knowledge Injection FrameworksLong Main
- Revisiting the Optimality of Word LengthsLong Main
- Reward-Augmented Decoding: Efficient Controlled Text Generation With a Unidirectional Reward ModelShort Main
- RexUIE: A Recursive Method with Explicit Schema Instructor for Universal Information ExtractionLong Findings
- RoAST: Robustifying Language Models via Adversarial Perturbation with Selective TrainingLong Findings
- RoBoCoP: A Comprehensive ROmance BOrrowing COgnate Package and Benchmark for Multilingual Cognate IdentificationLong Main
- RoMQA: A Benchmark for Robust, Multi-evidence, Multi-answer Question AnsweringLong Findings
- Robust Prompt Optimization for Large Language Models Against Distribution ShiftsLong Main
- RobustEmbed: Robust Sentence Embeddings Using Self-Supervised Contrastive Pre-TrainingLong Findings
- RobustGEC: Robust Grammatical Error Correction Against Subtle Context PerturbationLong Main
- Robustness Tests for Automatic Machine Translation Metrics with Adversarial AttacksShort Findings
- Robustness of Named-Entity Replacements for In-Context LearningShort Findings
- Role of Context in Unsupervised Sentence Representation Learning: the Case of Dialog Act ModelingShort Findings
- Roles of Scaling and Instruction Tuning in Language Perception: Model vs. Human AttentionLong Findings
- Romanization-based Large-scale Adaptation of Multilingual Language ModelsShort Findings
- Rumor Detection on Social Media with Crowd Intelligence and ChatGPT-Assisted NetworksLong Main
- S2abEL: A Dataset for Entity Linking from Scientific TablesLong Main
- SAC$^3$: Reliable Hallucination Detection in Black-Box Language Models via Semantic-aware Cross-check ConsistencyLong Findings
- SAMRank: Unsupervised Keyphrase Extraction using Self-Attention Map in BERT and GPT-2Long Main
- SCENE: Self-Labeled Counterfactuals for Extrapolating to Negative ExamplesLong Main
- SCITAB: A Challenging Benchmark for Compositional Reasoning and Claim Verification on Scientific TablesLong Main
- SDOH-NLI: a Dataset for Inferring Social Determinants of Health from Clinical NotesShort Findings
- SEAHORSE: A Multilingual, Multifaceted Dataset for Summarization EvaluationLong Main
- SEER : A Knapsack approach to Exemplar Selection for In-Context HybridQALong Main
- SELFOOD: Self-Supervised Out-Of-Distribution Detection via Learning to RankLong Findings
- SGP-TOD: Building Task Bots Effortlessly via Schema-Guided LLM PromptingLong Findings
- SHARCS: Efficient Transformers Through Routing with Dynamic Width Sub-networksShort Findings
- SIR-ABSC: Incorporating Syntax into RoBERTa-based Sentiment Analysis Models with a Special Aggregator TokenLong Findings
- SKD-NER: Continual Named Entity Recognition via Span-based Knowledge Distillation with Reinforcement LearningLong Main
- SLOG: A Structural Generalization Benchmark for Semantic ParsingLong Main
- SMoP: Towards Efficient and Effective Prompt Tuning with Sparse Mixture-of-PromptsShort Main
- SODA: Million-scale Dialogue Distillation with Social Commonsense ContextualizationLong Main
- SOUL: Towards Sentiment and Opinion Understanding of LanguageShort Main
- SPT: Learning to Selectively Insert Prompts for Better Prompt TuningLong Main
- STAIR: Learning Sparse Text and Image Representation in Grounded TokensLong Main
- STEER: Unified Style Transfer with Expert ReinforcementLong Findings
- STINMatch: Semi-Supervised Semantic-Topological Iteration Network for Financial Risk Detection via News Label DiffusionLong Main
- SUT: Active Defects Probing for Transcompiler ModelsShort Main
- SWEET - Weakly Supervised Person Name Extraction for Fighting Human TraffickingLong Findings
- SYMPTOMIFY: Transforming Symptom Annotations with Language Model Knowledge HarvestingLong Findings
- Salespeople vs SalesBot: Exploring the Role of Educational Value in Conversational Recommender SystemsLong Findings
- Scalable-DSC: A Structural Template Prompt Approach to Scalable Dialogue State CorrectionLong Main
- Scaling Law for Document Neural Machine TranslationLong Findings
- Scaling Laws vs Model Architectures: How does Inductive Bias Influence Scaling?Long Findings
- Scaling Vision-Language Models with Sparse Mixture of ExpertsLong Findings
- ScanDL: A Diffusion Model for Generating Synthetic Scanpaths on TextsLong Main
- ScdNER: Span-Based Consistency-Aware Document-Level Named Entity RecognitionShort Main
- Scene Graph Enhanced Pseudo-Labeling for Referring Expression ComprehensionLong Findings
- Schema-adaptable Knowledge Graph ConstructionLong Findings
- SciRepEval: A Multi-Format Benchmark for Scientific Document RepresentationsLong Main
- Search Augmented Instruction LearningLong Findings
- Seeing through the mess: evolutionary dynamics of lexical polysemyLong Main
- SegAugment: Maximizing the Utility of Speech Translation Data with Segmentation-based AugmentationsLong Findings
- Segmented Recurrent Transformer: An Efficient Sequence-to-Sequence ModelLong Findings
- Select, Prompt, Filter: Distilling Large Language Models for Summarizing ConversationsShort Main
- Selecting Key Views for Zero-Shot Entity LinkingLong Findings
- Selective Demonstrations for Cross-domain Text-to-SQLLong Findings
- Selective Labeling: How to Radically Lower Data-Labeling Costs for Document Extraction ModelsLong Main
- Selectively Answering Ambiguous QuestionsLong Main
- Self-Detoxifying Language Models via Toxification ReversalLong Main
- Self-Ensemble of $N$-best Generation Hypotheses by Lexically Constrained DecodingShort Main
- Self-Evolution Learning for Mixup: Enhance Data Augmentation on Few-Shot Text Classification TasksLong Main
- Self-ICL: Zero-Shot In-Context Learning with Self-Generated DemonstrationsLong Main
- Self-Improvement of Non-autoregressive Model via Sequence-Level DistillationLong Main
- Self-Influence Guided Data Reweighting for Language Model Pre-trainingLong Main
- Self-Knowledge Guided Retrieval Augmentation for Large Language ModelsLong Findings
- Self-Polish: Enhance Reasoning in Large Language Models via Problem RefinementLong Findings
- Self-Supervised Behavior Cloned Transformers are Path Crawlers for Text GamesShort Findings
- Self-Supervised Rule Learning to Link Text Segments to Relational Elements of Structured KnowledgeLong Findings
- Self-prompted Chain-of-Thought on Large Language Models for Open-domain Multi-hop ReasoningLong Findings
- Self-supervised Meta-Prompt Learning with Meta-Gradient Regularization for Few-shot GeneralizationLong Findings
- Self-supervised Post-processing Method to Enrich Pretrained Word VectorsShort Findings
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsLong Main
- Semantic Decomposition of Question and SQL for Text-to-SQL ParsingLong Findings
- Semantic Parsing by Large Language Models for Intricate Updating Strategies of Zero-Shot Dialogue State TrackingShort Findings
- Semantic Similarity Covariance Matrix ShrinkageLong Findings
- Semantic Space Grounded Weighted Decoding for Multi-Attribute Controllable Dialogue GenerationLong Main
- Semantic matching for text classification with complex class descriptionsLong Main
- Semi-Structured Object Sequence EncodersLong Findings
- Semi-automatic Data Enhancement for Document-Level Relation Extraction with Distant Supervision from Large Language ModelsShort Main
- Semi-supervised multimodal coreference resolution in image narrationsLong Main
- SentiStream: A Co-Training Framework for Adaptive Online Sentiment Analysis in Evolving Data StreamsLong Main
- Sentiment Analysis on Streaming User Reviews via Dual-Channel Dynamic Graph Neural NetworkLong Main
- Seq2seq is All You Need for Coreference ResolutionLong Main
- SeqXGPT: Sentence-Level AI-Generated Text DetectionLong Main
- Set Learning for Generative Information ExtractionShort Main
- Shall We Pretrain Autoregressive Language Models with Retrieval? A Comprehensive StudyLong Main
- Show, Write, and Retrieve: Entity-aware Article Generation and RetrievalLong Findings
- SiMFy: A Simple Yet Effective Approach for Temporal Knowledge Graph ReasoningLong Findings
- SimCKP: Simple Contrastive Learning of Keyphrase RepresentationsLong Findings
- SimCSE++: Improving Contrastive Learning for Sentence Embeddings from Two PerspectivesLong Main
- Simple Hardware-Efficient PCFGs with Independent Left and Right ProductionsShort Findings
- Simple Temporal Adaptation to Changing Label Sets: Hashtag Prediction via Dense KNNShort Main
- Simple and Effective Input Reformulations for TranslationLong Main
- Simpler neural networks prefer subregular languagesLong Findings
- Simplicity Level Estimate (SLE): A Learned Reference-Less Metric for Sentence SimplificationShort Main
- Simultaneous Machine Translation with Tailored ReferenceLong Findings
- Skill-Based Few-Shot Selection for In-Context LearningLong Main
- Small Language Models Fine-tuned to Coordinate Larger Language Models improve Complex ReasoningLong Main
- Smart “Chef”: Verifying the Effect of Role-based Paraphrasing for Aspect Term ExtractionShort Findings
- SmartSpanNER: Making SpanNER Robust in Low Resource ScenariosLong Findings
- Social Commonsense-Guided Search Query Generation for Open-Domain Knowledge-Powered ConversationsShort Findings
- Sociocultural Norm Similarities and Differences via Situational Alignment and Explainable Textual EntailmentLong Main
- Solving Hard Analogy Questions with Relation Embedding ChainsLong Main
- Solving the Right Problem is Key for Translational NLP: A Case Study in UMLS Vocabulary InsertionLong Findings
- Somali Information Retrieval Corpus: Bridging the Gap between Query Translation and Dedicated Language ResourcesShort Main
- SoulChat: Improving LLMs' Empathy, Listening, and Comfort Abilities through Fine-tuning with Multi-turn Empathy ConversationsShort Findings
- Sound of Story: Multi-modal Storytelling with AudioLong Findings
- Sources of Hallucination by Large Language Models on Inference TasksLong Findings
- SpEL: Structured Prediction for Entity LinkingLong Main
- Sparse Black-Box Multimodal Attack for Vision-Language Adversary GenerationLong Findings
- Sparse Frame Grouping Network with Action Centered for Untrimmed Video Paragraph CaptioningLong Findings
- Sparse Low-rank Adaptation of Pre-trained Language ModelsLong Main
- Sparse Universal TransformerLong Main
- Speak, Memory: An Archaeology of Books Known to ChatGPT/GPT-4Long Main
- Speaking Style Conversion in the Waveform Domain Using Discrete Self-Supervised UnitsLong Findings
- Specialist or Generalist? Instruction Tuning for Specific NLP TasksLong Main
- Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq GenerationLong Findings
- Speech Recognition and Meaning Interpretation: Towards Disambiguation of Structurally Ambiguous Spoken Utterances in IndonesianLong Main
- Speech-enriched Memory for Inference-time Adaptation of ASR Models to Word DictionariesLong Main
- SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational AbilitiesLong Findings
- Spoiler Detection as Semantic Text MatchingShort Main
- Stance Detection on Social Media with Background KnowledgeLong Main
- Standardizing Distress Analysis: Emotion-Driven Distress Identification and Cause Extraction (DICE) in Multimodal Online PostsLong Main
- Statistical Depth for Ranking and Characterizing Transformer-Based Text EmbeddingsLong Main
- Statistically Profiling Biases in Natural Language Reasoning Datasets and ModelsLong Findings
- SteerLM: Attribute Conditioned SFT as an (User-Steerable) Alternative to RLHFLong Findings
- Steering Large Language Models for Machine Translation with Finetuning and In-Context LearningShort Findings
- StereoMap: Quantifying the Awareness of Human-like Stereotypes in Large Language ModelsLong Main
- Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation BenchmarksShort Main
- StoryAnalogy: Deriving Story-level Analogies from Large Language Models to Unlock Analogical UnderstandingLong Main
- StrAE: Autoencoding for Pre-Trained Embeddings using Explicit StructureLong Main
- Strong and Efficient Baselines for Open Domain Conversational Question AnsweringShort Findings
- Struct-XLM: A Structure Discovery Multilingual Language Model for Enhancing Cross-lingual Transfer through Reinforcement LearningLong Main
- StructGPT: A General Framework for Large Language Model to Reason over Structured DataLong Main
- Structural Priming Demonstrates Abstract Grammatical Representations in Multilingual Language ModelsLong Main
- Structural generalization in COGS: Supertagging is (almost) all you needLong Main
- Structure-aware Knowledge Graph-to-text Generation with Planning Selection and Similarity DistinctionLong Main
- Style-Aware Radiology Report Generation with RadGraph and Few-Shot PromptingLong Findings
- StyleBART: Decorate Pretrained Model with Style Adapters for Unsupervised Stylistic Headline GenerationLong Findings
- Stylized Dialogue Generation with Feature-Guided Knowledge AugmentationLong Findings
- Sub-network Discovery and Soft-masking for Continual Learning of Mixed TasksLong Findings
- Subspace Chronicles: How Linguistic Information Emerges, Shifts and Interacts during Language Model TrainingLong Findings
- SummEdits: Measuring LLM Ability at Factual Reasoning Through The Lens of SummarizationLong Main
- SummIt: Iterative Text Summarization via ChatGPTLong Findings
- Summarizing Multiple Documents with Conversational Structure for Meta-Review GenerationLong Findings
- SuperDialseg: A Large-scale Dataset for Supervised Dialogue SegmentationLong Main
- SuperTweetEval: A Challenging, Unified and Heterogeneous Benchmark for Social Media NLP ResearchLong Findings
- Superlim: A Swedish Language Understanding Evaluation BenchmarkLong Main
- Support or Refute: Analyzing the Stance of Evidence to Detect Out-of-Context Mis- and DisinformationLong Main
- Survival of the Most Influential Prompts: Efficient Black-Box Prompt Search via Clustering and PruningLong Findings
- Syllogistic Reasoning for Legal Judgment AnalysisLong Main
- Symbol tuning improves in-context learning in language modelsLong Main
- Symbolic Planning and Code Generation for Grounded DialogueLong Main
- Symbolization, Prompt, and Classification: A Framework for Implicit Speaker Identification in NovelsLong Findings
- Syntactic Substitutability as Unsupervised Dependency SyntaxLong Main
- Syntax Matters: Towards Spoken Language Understanding via Syntax-Aware AttentionShort Findings
- Syntax-Aware Retrieval Augmented Code GenerationLong Findings
- Synthesize, if you do not have: Effective Synthetic Dataset Creation Strategies for Self-Supervised Opinion Summarization in E-commerceShort Findings
- Synthetic Data Generation with Large Language Models for Text Classification: Potential and LimitationsLong Main
- System Combination via Quality Estimation for Grammatical Error CorrectionLong Main
- Systematic Assessment of Factual Knowledge in Large Language ModelsShort Findings
- Systematic word meta-sense extensionLong Main
- T-Projection: High Quality Annotation Projection for Sequence Labeling TasksLong Findings
- T5Score: Discriminative Fine-tuning of Generative Evaluation MetricsLong Findings
- TADI: Topic-aware Attention and Powerful Dual-encoder Interaction for Recall in News RecommendationLong Findings
- TATA: Stance Detection via Topic-Agnostic and Topic-Aware EmbeddingsLong Main
- TCFLE-8: a Corpus of Learner Written Productions for French as a Foreign Language and its Application to Automated Essay ScoringLong Main
- TCRA-LLM: Token Compression Retrieval Augmented Large Language Model for Inference Cost ReductionLong Findings
- TELeR: A General Taxonomy of LLM Prompts for Benchmarking Complex TasksShort Findings
- TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language UnderstandingLong Findings
- TK-KNN: A Balanced Distance-Based Pseudo Labeling Approach for Semi-Supervised Intent ClassificationLong Findings
- TLM: Token-Level Masking for TransformersLong Main
- TOD-Flow: Modeling the Structure of Task-Oriented DialoguesLong Main
- TR-Rules: Rule-based Model for Link Forecasting on Temporal Knowledge Graph Considering Temporal RedundancyLong Findings
- TRAMS: Training-free Memory Selection for Long-range Language ModelingShort Findings
- TRAVEL: Tag-Aware Conversational FAQ Retrieval via Reinforcement LearningLong Main
- TRIGO: Benchmarking Formal Mathematical Proof Reduction for Generative Language ModelsLong Main
- TRIP: Accelerating Document-level Multilingual Pre-training via Triangular Document-level Pre-training on Parallel Data TripletsLong Findings
- TSTR: Target Similarity Tuning Meets the Real WorldShort Findings
- TaTA: A Multilingual Table-to-Text Dataset for African LanguagesLong Findings
- TabPrompt: Graph-based Pre-training and Prompting for Few-shot Table UnderstandingLong Findings
- TacoPrompt: A Collaborative Multi-Task Prompt Learning Method for Self-Supervised Taxonomy CompletionLong Main
- Tagging-Assisted Generation Model with Encoder and Decoder Supervision for Aspect Sentiment Triplet ExtractionLong Main
- Take a Closer Look at Multilinguality! Improve Multilingual Pre-Training Using Monolingual Corpora OnlyLong Findings
- TalkUp: Paving the Way for Understanding Empowering LanguageLong Findings
- Target-Agnostic Gender-Aware Contrastive Learning for Mitigating Bias in Multilingual Machine TranslationLong Main
- Target-Aware Spatio-Temporal Reasoning via Answering Questions in Dynamic Audio-Visual ScenariosLong Findings
- Target-oriented Proactive Dialogue Systems with Personalization: Problem Formulation and Dataset CurationShort Main
- Target-to-Source Augmentation for Aspect Sentiment Triplet ExtractionLong Main
- Task-Adaptive Tokenization: Enhancing Long-Form Text Generation Efficacy in Mental Health and BeyondLong Main
- Task-Agnostic Low-Rank Adapters for Unseen English DialectsLong Main
- Task-Attentive Transformer Architecture for Continual Learning of Vision-and-Language Tasks Using Knowledge DistillationLong Findings
- Task-Aware Self-Supervised Framework for Dialogue Discourse ParsingLong Findings
- Task-Level Thinking Steps Help Large Language Models for Challenging Classification TaskLong Main
- TaskWeb: Selecting Better Source Tasks for Multi-task NLPLong Main
- Taxonomy Expansion for Named Entity RecognitionLong Main
- Teacher Perception of Automatically Extracted Grammar Concepts for L2 Language LearningLong Findings
- TempTabQA: Temporal Question Answering for Semi-Structured TablesLong Main
- Temporal Extrapolation and Knowledge Transfer for Lifelong Temporal Knowledge Graph ReasoningLong Findings
- Temporal Knowledge Graph Forecasting Without Knowledge Using In-Context LearningLong Main
- Temporal Knowledge Graph Reasoning Based on N-tuple ModelingLong Findings
- Test-Time Self-Adaptive Small Language Models for Question AnsweringShort Findings
- Test-time Augmentation for Factual ProbingShort Findings
- Text Augmented Spatial Aware Zero-shot Referring Image SegmentationLong Findings
- Text Classification via Large Language ModelsLong Findings
- Text Embeddings Reveal (Almost) As Much As TextLong Main
- Text Fact TransferLong Main
- Text Rendering Strategies for Pixel Language ModelsLong Main
- Text Representation Distillation via Information Bottleneck PrincipleLong Main
- Text encoders bottleneck compositionality in contrastive vision-language modelsLong Main
- Text-Transport: Toward Learning Causal Effects of Natural LanguageLong Main
- Text-guided 3D Human Generation from 2D CollectionsLong Findings
- Text2Tree: Aligning Text Representation to the Label Tree Hierarchy for Imbalanced Medical ClassificationLong Findings
- TextMixer: Mixing Multiple Inputs for Privacy-Preserving InferenceLong Findings
- That was the last straw, we need more: Are Translation Systems Sensitive to Disambiguating Context?Long Findings
- The ACL OCL Corpus: Advancing Open Science in Computational LinguisticsLong Main
- The Art of SOCRATIC QUESTIONING: Recursive Thinking with Large Language ModelsLong Main
- The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal ModelsLong Main
- The Benefits of Label-Description Training for Zero-Shot Text ClassificationLong Main
- The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-TuningLong Main
- The Cost of Compression: Investigating the Impact of Compression on Parametric Knowledge in Language ModelsLong Findings
- The Curious Case of Hallucinatory (Un)answerability: Finding Truths in the Hidden States of Over-Confident Large Language ModelsLong Main
- The Distributional Hypothesis Does Not Fully Explain the Benefits of Masked Language Model PretrainingLong Main
- The Effect of Scaling, Retrieval Augmentation and Form on the Factual Consistency of Language ModelsLong Main
- The Framework Tax: Disparities Between Inference Efficiency in NLP Research and DeploymentLong Main
- The Intended Uses of Automated Fact-Checking Artefacts: Why, How and WhoLong Findings
- The Internal State of an LLM Knows When It's LyingLong Findings
- The Interpreter Understands Your Meaning: End-to-end Spoken Language Understanding Aided by Speech TranslationLong Findings
- The Iron(ic) Melting Pot: Reviewing Human Evaluation in Humour, Irony and Sarcasm GenerationLong Findings
- The Law and NLP: Bridging Disciplinary DisconnectsShort Findings
- The Less the Merrier? Investigating Language Representation in Multilingual ModelsLong Findings
- The Linearity of the Effect of Surprisal on Reading Times across LanguagesShort Findings
- The Locality and Symmetry of Positional EncodingsLong Findings
- The PEACE-Reviews dataset: Modeling Cognitive Appraisals in Emotion Text AnalysisLong Findings
- The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and ValuesLong Main
- The Past, Present, and Future of Typological Databases in NLPShort Findings
- The Sentiment Problem: A Critical Survey towards Deconstructing Sentiment AnalysisLong Main
- The Shifted and The Overlooked: A Task-oriented Investigation of User-GPT InteractionsLong Main
- The Skipped Beat: A Study of Sociopragmatic Understanding in LLMs for 64 LanguagesLong Main
- The Troubling Emergence of Hallucination in Large Language Models - An Extensive Definition, Quantification, and Prescriptive RemediationsLong Main
- The Truth, The Whole Truth, and Nothing but the Truth: A New Benchmark Dataset for Hebrew Text Credibility AssessmentLong Findings
- The Vault: A Comprehensive Multilingual Dataset for Advancing Code Understanding and GenerationLong Findings
- The language of prompting: What linguistic properties make a prompt successful?Long Findings
- The neural dynamics of word recognition and integrationLong Main
- The student becomes the master: Outperforming GPT3 on Scientific Factual Error CorrectionLong Findings
- TheoremQA: A Theorem-driven Question Answering DatasetLong Main
- Theory of Mind for Multi-Agent Collaboration via Large Language ModelsLong Main
- This Reads Like That: Deep Learning for Interpretable Natural Language ProcessingShort Main
- This is not a Dataset: A Large Negation Benchmark to Challenge Large Language ModelsLong Main
- Thorny Roses: Investigating the Dual Use Dilemma in Natural Language ProcessingLong Findings
- Three Questions Concerning the Use of Large Language Models to Facilitate Mathematics LearningShort Findings
- Three Stream Based Multi-level Event Contrastive Learning for Text-Video Event ExtractionLong Main
- Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie CalibrationLong Main
- Time-Aware Language Modeling for Historical Text DatingLong Findings
- Time-Considerable Dialogue Models via Reranking by Time DependencyLong Findings
- To Build Our Future, We Must Know Our Past: Contextualizing Paradigm Shifts in Natural Language ProcessingLong Main
- ToViLaG: Your Visual-Language Generative Model is Also An EvildoerLong Main
- Token Prediction as Implicit Classification to Identify LLM-Generated TextShort Main
- TokenDrop + BucketSampler: Towards Efficient Padding-free Fine-tuning of Language ModelsLong Findings
- Tokenization Consistency Matters for Generative Models on Extractive NLP TasksShort Findings
- TopWORDS-Poetry: Simultaneous Text Segmentation and Word Discovery for Classical Chinese Poetry via Bayesian InferenceLong Main
- Topic-DPR: Topic-based Prompts for Dense Passage RetrievalLong Findings
- Topic-Informed Dialogue Summarization using Topic Distribution and Prompt-based ModelingShort Findings
- Toward Human Readable Prompt Tuning: Kubrick’s The Shining is a good movie, and a good prompt too?Long Findings
- Toward Joint Language Modeling for Speech Units and TextLong Findings
- Toward a Critical Toponymy Framework for Named Entity Recognition: A Case Study of Airbnb in New York CityLong Main
- Towards A Holistic Landscape of Situated Theory of Mind in Large Language ModelsLong Findings
- Towards A Unified View of Sparse Feed-Forward Network in Pretraining Large Language ModelLong Main
- Towards Anytime Fine-tuning: Continually Pre-trained Language Models with Hypernetwork PromptsLong Findings
- Towards Being Parameter-Efficient: A Stratified Sparsely Activated Transformer with Dynamic CapacityLong Findings
- Towards Better Representations for Multi-Label Text Classification with Multi-granularity InformationLong Findings
- Towards Building More Robust NER datasets: An Empirical Study on NER Dataset Bias from a Dataset Difficulty ViewLong Main
- Towards Concept-Aware Large Language ModelsLong Findings
- Towards Conceptualization of ``Fair Explanation'': Disparate Impacts of anti-Asian Hate Speech Explanations on Content ModeratorsLong Main
- Towards Detecting Contextual Real-Time Toxicity for In-Game ChatLong Findings
- Towards Enhancing Relational Rules for Knowledge Graph Link PredictionLong Findings
- Towards Example-Based NMT with Multi-Levenshtein TransformersLong Main
- Towards Formality-Aware Neural Machine Translation by Leveraging Context InformationShort Findings
- Towards General Error Diagnosis via Behavioral Testing in Machine TranslationLong Findings
- Towards Informative Few-Shot Prompt with Maximum Information Gain for In-Context LearningLong Findings
- Towards Informative Open-ended Text Generation with Dynamic Knowledge TriplesLong Findings
- Towards Interpretable Mental Health Analysis with Large Language ModelsLong Main
- Towards LLM-driven Dialogue State TrackingLong Main
- Towards Low-Resource Automatic Program Repair with Meta-Learning and Pretrained Language ModelsLong Main
- Towards Making the Most of ChatGPT for Machine TranslationLong Findings
- Towards Mitigating LLM Hallucination via Self ReflectionLong Findings
- Towards Multilingual Interlinear Morphological GlossingLong Findings
- Towards Noise-Tolerant Speech-Referring Video Object Segmentation: Bridging Speech and TextLong Main
- Towards Reliable Misinformation Mitigation: Generalization, Uncertainty, and GPT-4Long Main
- Towards Robust Pruning: An Adaptive Knowledge-Retention Pruning Strategy for Language ModelsLong Main
- Towards Unsupervised Recognition of Token-level Semantic Differences in Related DocumentsShort Main
- Towards Zero-shot Learning for End-to-end Cross-modal Translation ModelsShort Findings
- Towards Zero-shot Relation Extraction in Web Mining: A Multimodal Approach with Relative XML PathLong Findings
- Towards a Better Understanding of Variations in Zero-Shot Neural Machine Translation PerformanceLong Main
- Towards a Deep Understanding of Multilingual End-to-End Speech TranslationLong Findings
- Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language ModelsLong Main
- Towards a Unified Conversational Recommendation System: Multi-task Learning via Contextualized Knowledge DistillationLong Main
- Towards a Unified Framework for Reference Retrieval and Related Work GenerationLong Findings
- Towards large language model-based personal agents in the enterprise: Current trends and open problemsLong Findings
- ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI ConversationShort Findings
- Toxicity in Multilingual Machine Translation at ScaleLong Findings
- Toxicity in chatgpt: Analyzing persona-assigned language modelsLong Findings
- Toxicity, Morality, and Speech Act Guided Stance DetectionLong Findings
- Training Simultaneous Speech Translation with Robust and Random Wait-k-Tokens StrategyLong Main
- Transcending Scaling Laws with 0.1% Extra ComputeLong Main
- Transductive Learning for Textual Few-Shot Classification in API-based Embedding ModelsLong Main
- Transfer-Free Data-Efficient Multilingual Slot LabelingLong Main
- Transformer Working Memory Enables Regular Language Reasoning And Natural Language Length ExtrapolationLong Findings
- Transformer-Based Language Model Surprisal Predicts Human Reading Times Best with About Two Billion Training TokensShort Findings
- Transformer-based Live Update Generation for Soccer Matches from Microblog PostsShort Main
- Transitioning Representations between Languages for Cross-lingual Event Detection via Langevin DynamicsShort Findings
- Translating away Translationese without Parallel DataLong Main
- Transparency at the Source: Evaluating and Interpreting Language Models With Access to the True DistributionLong Findings
- Tree Prompting: Efficient Task Adaptation without Fine-TuningLong Main
- Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language ModelsShort Main
- TrojanSQL: SQL Injection against Natural Language Interface to DatabaseLong Main
- TrueTeacher: Learning Factual Consistency Evaluation with Large Language ModelsLong Main
- Tuna: Instruction Tuning using Feedback from Large Language ModelsLong Findings
- Tunable Soft Prompts are Messengers in Federated LearningLong Findings
- Turn-Level Active Learning for Dialogue State TrackingLong Main
- Two Directions for Clinical Data Generation with Large Language Models: Data-to-Label and Label-to-DataLong Findings
- Type-Aware Decomposed Framework for Few-Shot Named Entity RecognitionLong Findings
- UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of RerankersLong Main
- ULF: Unsupervised Labeling Function Correction using Cross-Validation for Weak SupervisionShort Main
- UPRISE: Universal Prompt Retrieval for Improving Zero-Shot EvaluationLong Main
- UPTON: Preventing Authorship Leakage from Public Text Release via Data PoisoningLong Findings
- UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language ModelLong Findings
- USB: A Unified Summarization Benchmark Across Tasks and DomainsLong Findings
- Ultra-Fine Entity Typing with Prior Knowledge about Labels: A Simple Clustering Based StrategyLong Findings
- Uncertainty Guided Global Memory Improves Multi-Hop Question AnsweringShort Main
- Uncertainty-aware Parameter-Efficient Self-training for Semi-supervised Language UnderstandingLong Findings
- Uncovering Limitations in Text-to-Image Generation: A Contrastive Approach with Structured Semantic AlignmentLong Findings
- Uncovering the Root of Hate Speech: A Dataset for Identifying Hate Instigating SpeechLong Findings
- Understanding Compositional Data Augmentation in Typologically Diverse Morphological InflectionLong Main
- Understanding Computational Models of Semantic Change: New Insights from the Speech CommunityShort Main
- Understanding HTML with Large Language ModelsLong Findings
- Understanding Translationese in Cross-Lingual SummarizationLong Findings
- Understanding the Effect of Model Compression on Social Bias in Large Language ModelsShort Main
- Understanding the Inner-workings of Language Models Through Representation DissimilarityShort Main
- Understanding the Role of Input Token Characters in Language Models: How Does Information Loss Affect Performance?Long Main
- UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and ReasoningLong Main
- UniMath: A Foundational and Multimodal Mathematical ReasonerShort Main
- Unified Low-Resource Sequence Labeling by Sample-Aware Dynamic Sparse FinetuningLong Main
- Unified Representation for Non-compositional and Compositional ExpressionsLong Findings
- Uniform Complexity for Text GenerationLong Findings
- Unifying Cross-Lingual Transfer across Scenarios of Resource ScarcityLong Main
- Unifying Discrete and Continuous Representations for Unsupervised Paraphrase GenerationLong Main
- Unifying Text, Tables, and Images for Multimodal Question AnsweringLong Findings
- Universal Domain Adaptation for Robust Handling of Distributional Shifts in NLPLong Findings
- Universal Self-Adaptive PromptingLong Main
- Unlearn What You Want to Forget: Efficient Unlearning for LLMsLong Main
- Unleashing the Multilingual Encoder Potential: Boosting Zero-Shot Performance via Probability CalibrationShort Findings
- Unleashing the Power of Language Models in Text-Attributed GraphLong Findings
- Unmasking the Hidden Meaning: Bridging Implicit and Explicit Hate Speech Embedding RepresentationsShort Findings
- Unnatural Error Correction: GPT-4 Can Almost Perfectly Handle Unnatural Scrambled TextShort Main
- Unnatural language processing: How do language models handle machine-generated prompts?Long Findings
- Unraveling Downstream Gender Bias from Large Language Models: A Study on AI Educational Writing AssistanceLong Findings
- Unraveling Feature Extraction Mechanisms in Neural NetworksLong Main
- Unsupervised Binary Code Translation with Application to Code Clone Detection and Vulnerability DiscoveryLong Findings
- Unsupervised Candidate Answer Extraction through Differentiable Masker-Reconstructor ModelLong Findings
- Unsupervised Grammatical Error Correction Rivaling Supervised MethodsLong Main
- Unsupervised Lexical Simplification with Context AugmentationShort Findings
- Unsupervised Sounding Pixel LearningLong Main
- Unveiling the Essence of Poetry: Introducing a Comprehensive Dataset and Benchmark for Poem SummarizationShort Main
- Unveiling the Implicit Toxicity in Large Language ModelsLong Main
- Unveiling the Multi-Annotation Process: Examining the Influence of Annotation Quantity and Instance Difficulty on Model PerformanceLong Findings
- Unveiling the Power of Argument Arrangement in Online Persuasive DiscussionsLong Findings
- Using Artificial French Data to Understand the Emergence of Gender Bias in Transformer Language ModelsShort Main
- Using In-Context Learning to Improve Dialogue SafetyLong Findings
- Using Interpretation Methods for Model EnhancementLong Main
- Using LLM for Improving Key Event Discovery: Temporal-Guided News Stream Clustering with Event SummariesShort Findings
- VECHR: A Dataset for Explainable and Robust Classification of Vulnerability Type in the European Court of Human RightsShort Main
- VER: Unifying Verbalizing Entities and RelationsLong Findings
- VERVE: Template-based ReflectiVE Rewriting for MotiVational IntErviewingLong Findings
- VIBE: Topic-Driven Temporal Adaptation for Twitter ClassificationLong Main
- VIP5: Towards Multimodal Foundation Models for RecommendationLong Findings
- VIPHY: Probing “Visible” Physical Commonsense KnowledgeLong Findings
- VISIT: Visualizing and Interpreting the Semantic Information Flow of TransformersLong Findings
- VISTA: Visual-Textual Knowledge Graph Representation LearningLong Findings
- VLIS: Unimodal Language Models Guide Multimodal Language GenerationLong Main
- Values, Ethics, Morals? On the Use of Moral Concepts in NLP ResearchLong Findings
- Variance Matters: Detecting Semantic Differences without Corpus/Word AlignmentLong Main
- Variator: Accelerating Pre-trained Models with Plug-and-Play Compression ModulesLong Findings
- Vector-Quantized Prompt Learning for Paraphrase GenerationLong Findings
- Vera: A General-Purpose Plausibility Estimation Model for Commonsense StatementsLong Main
- Verb Conjugation in Transformers Is Determined by Linear Encodings of Subject NumberShort Findings
- ViPE: Visualise Pretty-much EverythingLong Main
- ViSoBERT: A Pre-Trained Language Model for Vietnamese Social Media Text ProcessingLong Main
- ViStruct: Visual Structural Knowledge Extraction via Curriculum Guided Code-Vision RepresentationLong Main
- Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is OffensiveLong Main
- Vicinal Risk Minimization for Few-Shot Cross-lingual Transfer in Abusive Language DetectionLong Main
- Video-Helpful Multimodal Machine TranslationLong Main
- Video-Text Retrieval by Supervised Sparse Multi-Grained LearningLong Findings
- Viewing Knowledge Transfer in Multilingual Machine Translation Through a Representational LensLong Findings
- Vision-Enhanced Semantic Entity Recognition in Document Images via Visually-Asymmetric Consistency LearningLong Main
- Visual Elements Mining as Prompts for Instruction Learning for Target-Oriented Multimodal Sentiment ClassificationLong Findings
- Visual Storytelling with Question-Answer PlansLong Findings
- Visually Grounded Continual Language Learning with Selective SpecializationLong Findings
- Visually-Situated Natural Language Understanding with Contrastive Reading Model and Frozen Large Language ModelsLong Main
- VivesDebate-Speech: A Corpus of Spoken Argumentation to Leverage Audio Features for Argument MiningShort Main
- WSDMS: Debunk Fake News via Weakly Supervised Detection of Misinforming Sentences with Contextualized Social WisdomLong Main
- Watermarking LLMs with Weight QuantizationLong Findings
- Watermarking PLMs on Classification Tasks by Combining Contrastive Learning with Weight PerturbationLong Findings
- We Are What We Repeatedly Do: Inducing and Deploying Habitual Schemas in Persona-Based ResponsesLong Main
- We Need to Talk About Reproducibility in NLP Model ComparisonLong Main
- We are Who We Cite: Bridges of Influence Between Natural Language Processing and Other Academic FieldsLong Main
- We're Afraid Language Models Aren't Modeling AmbiguityLong Main
- Weakly Supervised Semantic Parsing with Execution-based Spurious Program FilteringLong Main
- Weakly-Supervised Learning of Visual Relations in Multimodal PretrainingLong Main
- Well Begun is Half Done: Generator-agnostic Knowledge Pre-Selection for Knowledge-Grounded DialogueLong Main
- What Comes Next? Evaluating Uncertainty in Neural Text Generators Against Human Production VariabilityLong Main
- What Else Do I Need to Know? The Effect of Background Information on Users’ Reliance on QA SystemsLong Main
- What Makes Chain-of-Thought Prompting Effective? A Counterfactual StudyLong Findings
- What Makes it Ok to Set a Fire? Iterative Self-distillation of Contexts and Rationales for Disambiguating Defeasible Social and Moral SituationsLong Findings
- What do Deck Chairs and Sun Hats Have in Common? Uncovering Shared Properties in Large Concept VocabulariesShort Main
- What to Read in a Contract? Party-Specific Summarization of Legal Obligations, Entitlements, and ProhibitionsLong Main
- What's "up" with vision-language models? Investigating their struggle with spatial reasoningLong Main
- When Do Decompositions Help for Machine Reading?Short Main
- When Language Models Fall in Love: Animacy Processing in Transformer Language ModelsLong Main
- When Reviewers Lock Horns: Finding Disagreements in Scientific Peer ReviewsShort Main
- When and Why Does Bias Mitigation Work?Long Findings
- When are Lemons Purple? The Concept Association Bias of Vision-Language ModelsLong Main
- When it Rains, it Pours: Modeling Media Storms and the News EcosystemLong Findings
- When the Majority is Wrong: Modeling Annotator Disagreement for Subjective TasksLong Main
- Where to start? Analyzing the potential value of intermediate modelsLong Main
- Whispering LLaMA: A Cross-Modal Generative Error Correction Framework for Speech RecognitionShort Main
- Who Wrote it and Why? Prompting Large-Language Models for Authorship VerificationShort Findings
- Who is Speaking? Speaker-Aware Multiparty Dialogue Act ClassificationLong Findings
- Why LLMs Hallucinate, and How to Get (Evidential) Closure: Perceptual, Intensional, and Extensional Learning for Faithful Natural Language GenerationLong Main
- Why Should This Article Be Deleted? Transparent Stance Detection in Multilingual Wikipedia Editor DiscussionsLong Main
- WiCE: Real-World Entailment for Claims in WikipediaLong Main
- WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on WikipediaLong Findings
- Women Wearing Lipstick: Measuring the Bias Between an Object and Its Related GenderShort Findings
- WordNet Is All You Need: A Surprisingly Effective Unsupervised Method for Graded Lexical EntailmentShort Findings
- Words, Subwords, and Morphemes: What Really Matters in the Surprisal-Reading Time Relationship?Short Findings
- X-SNS: Cross-Lingual Transfer Prediction through Sub-Network SimilarityLong Findings
- XLM-V: Overcoming the Vocabulary Bottleneck in Multilingual Masked Language ModelsLong Main
- XLS-R fine-tuning on noisy word boundaries for unsupervised speech segmentation into wordsShort Findings
- XTREME-UP: A User-Centric Scarce-Data Benchmark for Under-Represented LanguagesLong Findings
- You Are What You Annotate: Towards Better Models through Annotator RepresentationsLong Findings
- You Told Me That Joke Twice: A Systematic Investigation of Transferability and Robustness of Humor Detection ModelsLong Main
- ZARA: Improving Few-Shot Self-Rationalization for Small Language ModelsLong Findings
- ZEROTOP: Zero-Shot Task-Oriented Semantic Parsing using Large Language ModelsShort Main
- ZGUL: Zero-shot Generalization to Unseen Languages using Multi-source Ensembling of Language AdaptersLong Main
- Zero-Shot Data Maps. Efficient Dataset Cartography Without Model TrainingLong Findings
- Zero-Shot-BERT-Adapters: a Zero-Shot Pipeline for Unknown Intent DetectionLong Findings
- Zero-shot Faithfulness Evaluation for Text Summarization with Foundation Language ModelLong Main
- Zero-shot Sharpness-Aware Quantization for Pre-trained Language ModelsLong Main
- Zero-shot Topical Text Classification with LLMs - an Experimental StudyLong Findings
- ZeroSCROLLS: A Zero-Shot Benchmark for Long Text UnderstandingLong Findings
- `Don't Get Too Technical with Me': A Discourse Structure-Based Framework for Automatic Science JournalismLong Main
- clembench: Using Game Play to Evaluate Chat-Optimized Language Models as Conversational AgentsLong Main
- e-THERAPIST: I suggest you to cultivate a mindset of positivity and nurture uplifting thoughtsLong Main
- impact of sample selection on in-context learning for entity extraction from scientific writingLong Findings
- kNN-CM: A Non-parametric Inference-Phase Adaptation of Parametric Text ClassifiersLong Findings
- mAggretriever: A Simple yet Effective Approach to Zero-Shot Multilingual Dense RetrievalShort Main
- mLongT5: A Multilingual and Efficient Text-To-Text Transformer for Longer SequencesShort Findings
EMNLP accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.