EMNLP 2023 Accepted Papers
The full list of 2,009 papers accepted at EMNLP 2023 (Conference on Empirical Methods in Natural Language Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
Long Main: 862Long Findings: 818Short Findings: 197Short Main: 132
- "A Tale of Two Movements": Identifying and Comparing Perspectives in \#BlackLivesMatter and \#BlueLivesMatter Movements-related Tweets using Weakly Supervised Graph-based Structured PredictionLong Findings
- "Are Your Explanations Reliable?" Investigating the Stability of LIME in Explaining Text Classifiers by Marrying XAI and Adversarial AttackLong Main
- "Fifty Shades of Bias": Normative Ratings of Gender Bias in GPT Generated English TextLong Main
- "You Are An Expert Linguistic Annotator": Limits of LLMs as Analyzers of Abstract Meaning RepresentationShort Findings
- $\textbf{\emph{CLMSM}}$: A Multi-Task Learning Framework for Pre-training on Procedural TextLong Findings
- $\textit{From Chaos to Clarity}$: Claim Normalization to Empower Fact-CheckingLong Findings
- $\textit{Lost in Translation, Found in Spans}$: Identifying Claims in Multilingual Social MediaLong Main
- $\textit{SelectNoise:}$ Unsupervised Noise Injection to Enable Zero-Shot Machine Translation for Extremely Low-resource LanguagesLong Findings
- $\textit{Swap and Predict}$ -- Predicting the Semantic Changes in Words across Corpora by Context SwappingLong Findings
- $\textit{``Don't Take This Out of Context!''}$ On the Need for Contextual Models and Evaluations for Stylistic RewritingLong Main
- $k$NN-LM Does Not Improve Open-ended Text GenerationLong Main
- 'Person' == Light-skinned, Western Man, and Sexualization of Women of Color: Stereotypes in Stable DiffusionLong Findings
- 1-PAGER: One Pass Answer Generation and Evidence RetrievalLong Findings
- 2INER: Instructive and In-Context Learning on Few-Shot Named Entity RecognitionLong Findings
- 3DRP-Net: 3D Relative Position-aware Network for 3D Visual GroundingLong Main
- 4 and 7-bit Labeling for Projective and Non-Projective Dependency TreesShort Main
- A Benchmark for Semi-Inductive Link Prediction in Knowledge GraphsShort Findings
- A Black-Box Attack on Code Models via Representation Nearest Neighbor SearchLong Findings
- A Boundary Offset Prediction Network for Named Entity RecognitionLong Findings
- A Causal View of Entity Bias in (Large) Language ModelsLong Findings
- A Challenging Multimodal Video Summary: Simultaneously Extracting and Generating Keyframe-Caption Pairs from VideoLong Main
- A Cheaper and Better Diffusion Language Model with Soft-Masked NoiseLong Main
- A Closer Look into Using Large Language Models for Automatic EvaluationShort Findings
- A Comprehensive Evaluation of Biomedical Entity Linking ModelsLong Main
- A Comprehensive Evaluation of Large Language Models on Legal Judgment PredictionLong Findings
- A Comprehensive Evaluation of Tool-Assisted Generation StrategiesLong Findings
- A Computational Interface to Translate Strategic Intent from Unstructured Language in a Low-Data SettingLong Findings
- A Confederacy of Models: a Comprehensive Evaluation of LLMs on Creative WritingLong Findings
- A Critical Analysis of Document Out-of-Distribution DetectionLong Findings
- A Dataset for Investigating the Impact of Context for Offensive Language Detection in TweetsShort Findings
- A Deeper (Autoregressive) Approach to Non-Convergent Discourse ParsingLong Main
- A Diachronic Analysis of Paradigm Shifts in NLP Research: When, How, and Why?Long Main
- A Diachronic Perspective on User Trust in AI under UncertaintyLong Main
- A Diffusion Weighted Graph Framework for New Intent DiscoveryLong Main
- A Fair and In-Depth Evaluation of Existing End-to-End Entity Linking SystemsLong Main
- A Fine-Grained Taxonomy of Replies to Hate SpeechLong Main
- A Framework for Bidirectional Decoding: Case Study in Morphological InflectionLong Findings
- A Framework for Exploring Player Perceptions of LLM-Generated Dialogue in Commercial Video GamesLong Findings
- A Framework for Vision-Language Warm-up Tasks in Multimodal Dialogue ModelsLong Main
- A Frustratingly Easy Plug-and-Play Detection-and-Reasoning Module for Chinese Spelling CheckLong Findings
- A Frustratingly Easy Post-Training Quantization Scheme for LLMsLong Main
- A Generation-based Deductive Method for Math Word ProblemsLong Main
- A Hierarchical Encoding-Decoding Scheme for Abstractive Multi-document SummarizationLong Findings
- A Joint Matrix Factorization Analysis of Multilingual RepresentationsLong Findings
- A Language Model with Limited Memory Capacity Captures Interference in Human Sentence ProcessingLong Findings
- A Lightweight Method to Generate Unanswerable Questions in EnglishShort Findings
- A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation AnalysisLong Main
- A Multi-Modal Multilingual Benchmark for Document Image ClassificationLong Findings
- A Multi-Task Dataset for Assessing Discourse Coherence in Chinese Essays: Structure, Theme, and Logic AnalysisLong Main
- A New Benchmark and Reverse Validation Method for Passage-level Hallucination DetectionLong Findings
- A Novel Contrastive Learning Method for Clickbait Detection on RoCliCo: A Romanian Clickbait Corpus of News ArticlesShort Findings
- A Parallel Corpus for Vietnamese Central-Northern Dialect Text TransferLong Findings
- A Predictive Factor Analysis of Social Biases and Task-Performance in Pretrained Masked Language ModelsLong Main
- A Quality-based Syntactic Template Retriever for Syntactically-Controlled Paraphrase GenerationLong Main
- A Query-Parallel Machine Reading Comprehension Framework for Low-resource NERLong Findings
- A Question Answering Framework for Decontextualizing User-facing Snippets from Scientific DocumentsLong Main
- A Read-and-Select Framework for Zero-shot Entity LinkingLong Findings
- A Reference-free Segmentation Quality Index (SegReFree)Long Findings
- A Rewriting Approach for Gender Inclusivity in PortugueseLong Findings
- A Rose by Any Other Name would not Smell as Sweet: Social Bias in Names MistranslationLong Main
- A Scalable Framework for Table of Contents Extraction from Complex ESG Annual ReportsLong Main
- A Self-enhancement Multitask Framework for Unsupervised Aspect Category DetectionLong Main
- A Sequence-to-Structure Approach to Document-level Targeted Sentiment AnalysisLong Findings
- A Simple Baseline for Knowledge-Based Visual Question AnsweringShort Main
- A Spectral Viewpoint on Continual Relation ExtractionShort Findings
- A State-Vector Framework for Dataset EffectsLong Main
- A Structure-Aware Generative Adversarial Network for Bilingual Lexicon InductionLong Findings
- A Study on Accessing Linguistic Information in Pre-Trained Language Models by Using PromptsShort Main
- A Suite of Generative Tasks for Multi-Level Multimodal Webpage UnderstandingLong Main
- A Systematic Study of Performance Disparities in Multilingual Task-Oriented Dialogue SystemsLong Main
- A Tale of Pronouns: Interpretability Informs Gender Bias Mitigation for Fairer Instruction-Tuned Machine TranslationLong Main
- A Thorough Examination on Zero-shot Dense RetrievalLong Findings
- A Training-Free Debiasing Framework with Counterfactual Reasoning for Conversational Emotion DetectionLong Main
- A Unified Framework for Synaesthesia AnalysisLong Findings
- A Unified View of Evaluation Metrics for Structured PredictionLong Main
- A Video Is Worth 4096 Tokens: Verbalize Story Videos To Understand Them In Zero ShotLong Main
- A Word Sense Distribution-based approach for Semantic Change PredictionLong Findings
- A Zero-Shot Language Agent for Computer Control with Structured ReflectionLong Findings
- A linear time approximation of Wasserstein distance with word embedding selectionLong Main
- ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life VideosLong Main
- ACT-SQL: In-Context Learning for Text-to-SQL with Automatically-Generated Chain-of-ThoughtLong Findings
- ACTOR: Active Learning with Annotator-specific Classification Heads to Embrace Human Label VariationShort Main
- AD-NLP: A Benchmark for Anomaly Detection in Natural Language ProcessingLong Main
- ALCUNA: Large Language Models Meet New KnowledgeLong Main
- ALDi: Quantifying the Arabic Level of Dialectness of TextLong Main
- AMR Parsing is Far from Solved: GrAPES, the Granular AMR Parsing Evaluation SuiteLong Main
- AMR Parsing with Causal Hierarchical Attention and PointersLong Main
- API-Assisted Code Generation for Question Answering on Varied Table StructuresLong Main
- API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMsLong Main
- APP: Adaptive Prototypical Pseudo-Labeling for Few-shot OOD DetectionLong Findings
- APoLLo : Unified Adapter and Prompt Learning for Vision Language ModelsLong Main
- APrompt: Attention Prompt Tuning for Efficient Adaptation of Pre-trained Language ModelsLong Main
- ARKitSceneRefer: Text-based Localization of Small Objects in Diverse Real-World 3D Indoor ScenesLong Findings
- ART: rule bAsed futuRe-inference deducTionLong Main
- ASPIRO: Any-shot Structured Parsing-error-Induced ReprOmpting for Consistent Data-to-Text GenerationShort Findings
- ASSERT: Automated Safety Scenario Red Teaming for Evaluating the Robustness of Large Language ModelsLong Findings
- ATFormer: A Learned Performance Model with Transfer Learning Across Devices for Deep Learning Tensor ProgramsLong Main
- Absolute Position Embedding Learns Sinusoid-like Waves for Attention Based on Relative PositionLong Main
- Abstractive Open Information ExtractionLong Main
- Accelerating Multiple Intent Detection and Slot Filling via Targeted Knowledge DistillationLong Findings
- Accelerating Toeplitz Neural Network with Constant-time Inference ComplexityLong Main
- Accented Speech Recognition With Accent-specific CodebooksLong Main
- Accuracy is not enough: Evaluating Personalization in SummarizersLong Findings
- Active Instruction Tuning: Improving Cross-Task Generalization by Training on Prompt Sensitive TasksLong Main
- Active Learning Principles for In-Context Learning with Large Language ModelsLong Findings
- Active Learning for Natural Language GenerationLong Main
- Active Retrieval Augmented GenerationLong Main
- AdaSent: Efficient Domain-Adapted Sentence Embeddings for Few-Shot ClassificationLong Main
- AdaTranS: Adapting with Boundary-based Shrinking for End-to-End Speech TranslationShort Findings
- Adapt in Contexts: Retrieval-Augmented Domain Adaptation via In-Context LearningLong Main
- Adaptation with Self-Evaluation to Improve Selective Prediction in LLMsLong Findings
- Adapter Pruning using Tropical CharacterizationShort Findings
- Adapter-TST: A Parameter Efficient Method for Multiple-Attribute Text Style TransferLong Findings
- Adapting Language Models to Compress ContextsLong Main
- Adapting Pretrained Text-to-Text Models for Long Text SequencesLong Findings
- Adaptive End-to-End Metric Learning for Zero-Shot Cross-Domain Slot FillingLong Main
- Adaptive Gating in Mixture-of-Experts based Language ModelsLong Main
- Adaptive Hinge Balance Loss for Document-Level Relation ExtractionShort Findings
- Adaptive Policy with Wait-k Model for Simultaneous TranslationLong Main
- Adaptive Structure Induction for Aspect-based Sentiment Analysis with Spectral PerspectiveLong Findings
- Adaptive Textual Label Noise Learning based on Pre-trained ModelsLong Findings
- Addressing Linguistic Bias through a Contrastive Analysis of Academic Writing in the NLP DomainLong Main
- Addressing NER Annotation Noises with Uncertainty-Guided Tree-Structured CRFsLong Main
- Addressing the Length Bias Challenge in Document-Level Neural Machine TranslationLong Findings
- Advancements in Arabic Grammatical Error Detection and Correction: An Empirical InvestigationLong Main
- Adversarial Robustness for Large Language NER models using Disentanglement and Word AttributionsLong Findings
- Adversarial Text Generation by Search and LearningLong Findings
- Affective and Dynamic Beam Search for Story GenerationLong Findings
- AfriSenti: A Twitter Sentiment Analysis Benchmark for African LanguagesLong Main
- Air-Decoding: Attribute Distribution Reconstruction for Decoding-Time Controllable Text GenerationLong Main
- Aligning Language Models to User OpinionsLong Findings
- Aligning Large Language Models through Synthetic FeedbackLong Main
- Aligning Predictive Uncertainty with Clarification Questions in Grounded DialogLong Findings
- Alignment Precedes Fusion: Open-Vocabulary Named Entity Recognition as Context-Type Semantic MatchingLong Findings
- All Things Considered: Detecting Partisan Events from News Media with Cross-Article ComparisonLong Main
- Allies: Prompting Large Language Model with Beam SearchLong Findings
- Always the Best Fit: Adaptive Domain Gap Filling from Causal Perspective for Few-Shot Relation ExtractionShort Findings
- An Adaptive Prompt Generation Framework for Task-oriented Dialogue SystemLong Findings
- An Attribution Method for Siamese EncodersShort Main
- An Empirical Investigation of Implicit and Explicit Knowledge-Enhanced Methods for Ad Hoc Dataset RetrievalLong Findings
- An Empirical Study of Frame Selection for Text-to-Video RetrievalLong Findings
- An Empirical Study of Instruction-tuning Large Language Models in ChineseLong Findings
- An Empirical Study of Multimodal Model MergingLong Findings
- An Empirical Study of Translation Hypothesis Ensembling with Large Language ModelsLong Main
- An Empirical Study on Multiple Knowledge from ChatGPT for Emotion Recognition in ConversationsLong Findings
- An Exploration of Left-Corner TransformationsLong Main
- An Expression Tree Decoding Strategy for Mathematical Equation GenerationLong Main
- An Integrative Survey on Mental Health Conversational Agents to Bridge Computer Science and Medical PerspectivesLong Main
- An Intent-based and Annotation-free Method for Duplicate Question Detection in CQA ForumsLong Findings
- An Investigation of LLMs’ Inefficacy in Understanding Converse RelationsLong Main
- An Iteratively Parallel Generation Method with the Pre-Filling Strategy for Document-level Event ExtractionLong Main
- Analysing State-Backed Propaganda Websites: a New Dataset and Linguistic StudyShort Main
- Analysis of Style-Shifting on Social Media: Using Neural Language Model Conditioned by Social MeaningsLong Findings
- Analyzing Cognitive Plausibility of Subword TokenizationShort Main
- Analyzing Film Adaptation through Narrative AlignmentLong Main
- Analyzing Modular Approaches for Visual Question DecompositionLong Main
- Analyzing Norm Violations in Live-Stream ChatLong Main
- Anaphor Assisted Document-Level Relation ExtractionLong Main
- Anchoring Fine-tuning of Sentence Transformer with Semantic Label Information for Efficient Truly Few-shot ClassificationShort Main
- AniEE: A Dataset of Animal Experimental Literature for Event ExtractionLong Findings
- Annotation Sensitivity: Training Data Collection Methods Affect Model PerformanceLong Findings
- Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence GroundingLong Findings
- Answer-state Recurrent Relational Network (AsRRN) for Constructed Response Assessment and Feedback GroupingLong Findings
- Answering Questions by Meta-Reasoning over Multiple Chains of ThoughtLong Main
- Appraising the Potential Uses and Harms of LLMs for Medical Systematic ReviewsLong Main
- Approximating CKY with TransformersLong Findings
- Approximating Two-Layer Feedforward Networks for Efficient TransformersLong Findings
- Arabic Mini-ClimateGPT : A Climate Change and Sustainability Tailored Arabic LLMShort Findings
- Architectural Sweet Spots for Modeling Human Label Variation by the Example of Argument Quality: It’s Best to Relate Perspectives!Long Main
- Are All Steps Equally Important? Benchmarking Essentiality Detection in Event ProcessesShort Main
- Are Compressed Language Models Less Subgroup Robust?Short Main
- Are Embedded Potatoes Still Vegetables? On the Limitations of WordNet Embeddings for Lexical SemanticsLong Main
- Are Language Models Worse than Humans at Following Prompts? It's ComplicatedShort Findings
- Are NLP Models Good at Tracing Thoughts: An Overview of Narrative UnderstandingLong Findings
- Are Personalized Stochastic Parrots More Dangerous? Evaluating Persona Biases in Dialogue SystemsLong Findings
- Are Structural Concepts Universal in Transformer Language Models? Towards Interpretable Cross-Lingual GeneralizationLong Findings
- Argue with Me Tersely: Towards Sentence-Level Counter-Argument GenerationLong Main
- Argument mining as a multi-hop generative machine reading comprehension taskLong Findings
- Argument-based Detection and Classification of Fallacies in Political DebatesLong Main
- Ask Language Model to Clean Your Noisy Translation DataLong Findings
- Ask To The Point: Open-Domain Entity-Centric Question GenerationLong Findings
- Asking Clarification Questions to Handle Ambiguity in Open-Domain QALong Findings
- Aspect-Category Enhanced Learning with a Neural Coherence Model for Implicit Sentiment AnalysisLong Findings
- Aspect-to-Scope Oriented Multi-view Contrastive Learning for Aspect-based Sentiment AnalysisLong Findings
- Assessing Privacy Risks in Language Models: A Case Study on Summarization TasksLong Findings
- Assessing Step-by-Step Reasoning against Lexical Negation: A Case Study on SyllogismShort Main
- Attack Prompt Generation for Red Teaming and Defending Large Language ModelsLong Findings
- Attention-Enhancing Backdoor Attacks Against BERT-based ModelsLong Findings
- Augmenting Zero-Shot Dense Retrievers with Plug-in Mixture-of-MemoriesLong Main
- Auto Search Indexer for End-to-End Document RetrievalLong Findings
- Auto-Instruct: Automatic Instruction Generation and Ranking for Black-Box Language ModelsLong Findings
- AutoPlan: Automatic Planning of Interactive Decision-Making Tasks With Large Language ModelsLong Findings
- AutoTrial: Prompting Language Models for Clinical Trial DesignLong Main
- Automated Few-Shot Classification with Instruction-Finetuned Language ModelsLong Findings
- Automatic Analysis of Substantiation in Scientific Peer ReviewsLong Findings
- Automatic Debate Evaluation with Argumentation Semantics and Natural Language Argument Graph NetworksLong Main
- Automatic Evaluate Dialogue Appropriateness by Using Dialogue ActLong Findings
- Automatic Evaluation of Attribution by Large Language ModelsLong Findings
- Automatic Model Selection with Large Language Models for ReasoningLong Findings
- Automatic Prompt Augmentation and Selection with Chain-of-Thought from Labeled DataLong Findings
- Automatic Prompt Optimization with "Gradient Descent" and Beam SearchLong Main
- Automatic Pronunciation Assessment - A ReviewLong Findings
- Automatic Transcription of Handwritten Old Occitan LanguageLong Main
- Axiomatic Preference Modeling for Longform Question AnsweringLong Main
- BERT Has More to Offer: BERT Layers Combination Yields Better Sentence EmbeddingsShort Findings
- BERTie Bott's Every Flavor Labels: A Tasty Introduction to Semantic Role Labeling for GalicianLong Main
- BERTwich: Extending BERT’s Capabilities to Model Dialectal and Noisy TextLong Findings
- BLESS: Benchmarking Large Language Models on Sentence SimplificationLong Main
- BLM-s/lE: A structured dataset of English spray-load verb alternations for testing generalization in LLMsLong Findings
- BRAINTEASER: Lateral Thinking Puzzles for Large Language ModelsLong Main
- BYOC: Personalized Few-Shot Classification with Co-Authored Class DescriptionsLong Findings
- Back Transcription as a Method for Evaluating Robustness of Natural Language Understanding Models to Speech Recognition ErrorsLong Main
- Background Summarization of Event TimelinesLong Main
- Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat DataLong Main
- Balance Act: Mitigating Hubness in Cross-Modal Retrieval with Query and Gallery BanksLong Main
- Balaur: Language Model Pretraining with Lexical Semantic RelationsLong Findings
- BanLemma: A Word Formation Dependent Rule and Dictionary Based Bangla LemmatizerLong Findings
- BanglaAbuseMeme: A Dataset for Bengali Abusive Meme ClassificationLong Main
- BasahaCorpus: An Expanded Linguistic Resource for Readability Assessment in Central Philippine LanguagesShort Main
- Battle of the Large Language Models: Dolly vs LLaMA vs Vicuna vs Guanaco vs Bard vs ChatGPT - A Text-to-SQL Parsing ComparisonLong Findings
- Bayesian Multi-Task Transfer Learning for Soft Prompt TuningLong Findings
- Be Selfish, But Wisely: Investigating the Impact of Agent Personality in Mixed-Motive Human-Agent InteractionsLong Main
- Beat LLMs at Their Own Game: Zero-Shot LLM-Generated Text Detection via Querying ChatGPTShort Main
- Benchmarking and Improving Text-to-SQL Generation under AmbiguityLong Main
- Beneath Surface Similarity: Large Language Models Make Reasonable Scientific Analogies after Structure AbductionLong Findings
- Beneath the Surface: Unveiling Harmful Memes with Multimodal Reasoning Distilled from Large Language ModelsLong Findings
- Best of Both Worlds: Towards Improving Temporal Knowledge Base Question Answering via Targeted Fact ExtractionShort Main
- Better Quality Pre-training Data and T5 Models for African LanguagesShort Main
- Better Together: Enhancing Generative Knowledge Graph Completion with Language Models and Neighborhood InformationShort Findings
- Beware of Model Collapse! Fast and Stable Test-time Adaptation for Robust Question AnsweringLong Main
- Beyond Candidates : Adaptive Dialogue Agent Utilizing Persona and KnowledgeLong Findings
- Beyond Denouncing Hate: Strategies for Countering Implied Biases and Stereotypes in LanguageLong Findings
- Beyond Detection: A Defend-and-Summarize Strategy for Robust and Interpretable Rumor Analysis on Social MediaLong Main
- Beyond Factuality: A Comprehensive Evaluation of Large Language Models as Knowledge GeneratorsLong Main
- Beyond Labels: Empowering Human Annotators with Natural Language Explanations through a Novel Active-Learning ArchitectureLong Findings
- Beyond Layout Embedding: Layout Attention with Gaussian Biases for Structured Document UnderstandingLong Findings
- Beyond Shared Vocabulary: Increasing Representational Word Similarities across Languages for Multilingual Machine TranslationLong Main
- Beyond Testers’ Biases: Guiding Model Testing with Knowledge Bases using LLMsLong Findings
- Bi-Drop: Enhancing Fine-tuning Generalization via Synchronous sub-net Estimation and OptimizationLong Findings
- BiSPN: Generating Entity Set and Relation Set Coherently in One PassLong Findings
- Bias Neutralization in Non-Parallel Texts: A Cyclic Approach with Auxiliary GuidanceLong Main
- BiasX: “Thinking Slow” in Toxic Content Moderation with Explanations of Implied Social BiasesShort Main
- BioDEX: Large-Scale Biomedical Adverse Drug Event Extraction for Real-World PharmacovigilanceLong Findings
- BioFEG: Generate Latent Features for Biomedical Entity LinkingLong Main
- BioPlanner: Automatic Evaluation of LLMs on Protocol Planning in BiologyLong Main
- BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language AssociationsLong Main
- Biomedical Named Entity Recognition via Dictionary-based Synonym GeneralizationLong Main
- Bipartite Graph Pre-training for Unsupervised Extractive Summarization with Graph Convolutional Auto-EncodersLong Findings
- Black-Box Tuning of Vision-Language Models with Effective Gradient ApproximationLong Findings
- Blackbird language matrices (BLM), a new task for rule-like generalization in neural networks: Can Large Language Models pass the test?Long Findings
- Boosting Inference Efficiency: Unleashing the Power of Parameter-Shared Pre-trained Language ModelsLong Findings
- Boosting Prompt-Based Self-Training With Mapping-Free Automatic Verbalizer for Multi-Class ClassificationLong Findings
- Boosting Summarization with Normalizing Flows and Aggressive TrainingLong Main
- Boot and Switch: Alternating Distillation for Zero-Shot Dense RetrievalLong Findings
- Bootstrapping Small \& High Performance Language Models with Unmasking-Removal Training PolicyShort Main
- BotPercent: Estimating Bot Populations in Twitter CommunitiesLong Findings
- Breaking Boundaries in Retrieval Systems: Unsupervised Domain Adaptation with Denoise-FinetuningLong Findings
- Breaking the Language Barrier: Improving Cross-Lingual Reasoning with Structured Self-AttentionLong Findings
- Breaking through Deterministic Barriers: Randomized Pruning Mask Generation and SelectionLong Findings
- Bridging Background Knowledge Gaps in Translation with Automatic ExplicitationLong Main
- Bridging Continuous and Discrete Spaces: Interpretable Sentence Representation Learning via Compositional OperationsLong Main
- Bridging Information-Theoretic and Geometric Compression in Language ModelsLong Main
- Bridging the Digital Divide: Performance Variation across Socio-Economic Factors in Vision-Language ModelsLong Main
- Bridging the Gap between Synthetic and Authentic Images for Multimodal Machine TranslationLong Main
- Building Multi-domain Dialog State Trackers from Single-domain DialogsLong Main
- Building Persona Consistent Dialogue Agents with Offline Reinforcement LearningLong Main
- Byte Pair Encoding for Symbolic MusicLong Main
- ByteSized32: A Corpus and Challenge Task for Generating Task-Specific World Models Expressed as Text GamesLong Main
- C-STS: Conditional Semantic Textual SimilarityLong Main
- C2D2 Dataset: A Resource for the Cognitive Distortion Analysis and Its Impact on Mental HealthLong Findings
- CAPSTONE: Curriculum Sampling for Dense Retrieval with Document ExpansionLong Main
- CAR: Conceptualization-Augmented Reasoner for Zero-Shot Commonsense Question AnsweringLong Findings
- CASE: Commonsense-Augmented Score with an Expanded Answer SpaceLong Findings
- CASSI: Contextual and Semantic Structure-based Interpolation Augmentation for Low-Resource NERLong Findings
- CCEval: A Representative Evaluation Benchmark for the Chinese-centric Multilingual Machine TranslationShort Findings
- CCIM: Cross-modal Cross-lingual Interactive Image TranslationShort Findings
- CCSRD: Content-Centric Speech Representation Disentanglement Learning for End-to-End Speech TranslationLong Findings
- CESAR: Automatic Induction of Compositional Instructions for Multi-turn DialogsLong Main
- CHEF in the Language Kitchen: A Generative Data Augmentation Leveraging Korean Morpheme IngredientsLong Main
- CHiLL: Zero-shot Custom Interpretable Feature Extraction from Clinical Notes with Large Language ModelsLong Findings
- CITB: A Benchmark for Continual Instruction TuningLong Findings
- CLAD-ST: Contrastive Learning with Adversarial Data for Robust Speech TranslationShort Main
- CLAIR: Evaluating Image Captions with Large Language ModelsShort Main
- CLASS: A Design Framework for Building Intelligent Tutoring Systems Based on Learning Science principlesLong Findings
- CLEME: Debiasing Multi-reference Evaluation for Grammatical Error CorrectionLong Main
- CLEVR-Implicit: A Diagnostic Dataset for Implicit Reasoning in Referring Expression ComprehensionLong Main
- COFFEE: Counterfactual Fairness for Personalized Text Generation in Explainable RecommendationLong Main
- COHESENTIA: A Novel Benchmark of Incremental versus Holistic Assessment of Coherence in Generated TextsLong Main
- COMET-M: Reasoning about Multiple Events in Complex SentencesLong Findings
- CONTRASTE: Supervised Contrastive Pre-training With Aspect-based Prompts For Aspect Sentiment Triplet ExtractionLong Findings
- CORE: A Few-Shot Company Relation Classification Dataset for Robust Domain Adaptation.Long Main
- COUNT: COntrastive UNlikelihood Text Style Transfer for Text DetoxificationShort Findings
- COVID-19 Vaccine Misinformation in Middle Income CountriesLong Main
- CP-BCS: Binary Code Summarization Guided by Control Flow Graph and Pseudo CodeLong Main
- CQE: A Comprehensive Quantity ExtractorLong Main
- CRAB: Assessing the Strength of Causal Relationships Between Real-world EventsLong Main
- CREATOR: Tool Creation for Disentangling Abstract and Concrete Reasoning of Large Language ModelsLong Findings
- CRT-QA: A Dataset of Complex Reasoning Question Answering over Tabular DataLong Main
- CRUSH4SQL: Collective Retrieval Using Schema Hallucination For Text2SQLLong Main
- CRaSh: Clustering, Removing, and Sharing Enhance Fine-tuning without Full Large Language ModelLong Main
- CReTIHC: Designing Causal Reasoning Tasks about Temporal Interventions and Hallucinated ConfoundingsShort Findings
- CRoW: Benchmarking Commonsense Reasoning in Real-World TasksLong Main
- CS2W: A Chinese Spoken-to-Written Style Conversion Dataset with Multiple Conversion TypesLong Main
- CT-GAT: Cross-Task Generative Adversarial Attack based on TransferabilityLong Main
- CTQScorer: Combining Multiple Features for In-context Example Selection for Machine TranslationLong Findings
- Cabbage Sweeter than Cake? Analysing the Potential of Large Language Models for Learning Conceptual SpacesShort Main
- Cache me if you Can: an Online Cost-aware Teacher-Student framework to Reduce the Calls to Large Language ModelsShort Findings
- Calc-X and Calcformers: Empowering Arithmetical Chain-of-Thought through Interaction with Symbolic SystemsShort Main
- Calibrated Seq2seq Models for Efficient and Generalizable Ultra-fine Entity TypingLong Findings
- Can Brain Signals Reveal Inner Alignment with Human Languages?Short Findings
- Can ChatGPT Assess Human Personalities? A General Evaluation FrameworkLong Findings
- Can ChatGPT Defend its Belief in Truth? Evaluating LLM Reasoning via DebateLong Findings
- Can ChatGPT Perform Reasoning Using the IRAC Method in Analyzing Legal Scenarios Like a Lawyer?Long Findings
- Can Foundation Models Watch, Talk and Guide You Step by Step to Make a Cake?Long Findings
- Can LLMs Facilitate Interpretation of Pre-trained Language Models?Long Main
- Can Language Models Laugh at YouTube Short-form Videos?Long Main
- Can Language Models Understand Physical Concepts?Long Main
- Can Large Language Models Capture Dissenting Human Voices?Long Main
- Can Large Language Models Fix Data Annotation Errors? An Empirical Study Using Debatepedia for Query-Focused Text SummarizationShort Findings
- Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?Long Main
- Can Retriever-Augmented Language Models Reason? The Blame Game Between the Retriever and the Language ModelLong Findings
- Can We Edit Factual Knowledge by In-Context Learning?Long Main
- Can We Edit Multimodal Large Language Models?Long Main
- Can You Follow Me? Testing Situational Understanding for ChatGPTLong Main
- Can you Summarize my learnings? Towards Perspective-based Educational Dialogue SummarizationLong Findings
- CaseEncoder: A Knowledge-enhanced Pre-trained Model for Legal Case EncodingLong Main
- Causal Document-Grounded Dialogue Pre-trainingLong Main
- Causal Inference from Text: Unveiling Interactions between VariablesLong Findings
- Causal Intervention for Abstractive Related Work GenerationLong Findings
- Causal Intervention-based Few-Shot Named Entity RecognitionLong Findings
- Causal Reasoning through Two Cognition Layers for Improving Generalization in Visual Question AnsweringLong Main
- Centering the Margins: Outlier-Based Identification of Harmed Populations in Toxicity DetectionLong Main
- Chain of Thought with Explicit Evidence Reasoning for Few-shot Relation ExtractionLong Findings
- Chain-of-Questions Training with Latent Answers for Robust Multistep Question AnsweringLong Main
- Chain-of-Thought Embeddings for Stance Detection on Social MediaShort Findings
- Chain-of-Thought Reasoning in Tabular Language ModelsLong Findings
- Chain-of-Thought Tuning: Masked Language Models can also Think Step By Step in Natural Language UnderstandingLong Main
- Challenges in Context-Aware Neural Machine TranslationLong Main
- Character-LLM: A Trainable Agent for Role-PlayingLong Main
- Characterizing Mechanisms for Factual Recall in Language ModelsLong Main
- Characterizing and Verifying Scientific Claims: Qualitative Causal Structure is All You NeedLong Main
- ChatCoT: Tool-Augmented Chain-of-Thought Reasoning on Chat-based Large Language ModelsLong Findings
- ChatEdit: Towards Multi-turn Interactive Facial Image Editing via DialogueLong Main
- ChatGPT Beyond English: Towards a Comprehensive Evaluation of Large Language Models in Multilingual LearningLong Findings
- ChatGPT to Replace Crowdsourcing of Paraphrases for Intent Classification: Higher Diversity and Comparable Model RobustnessLong Main
- Chinese Lexical Substitution: Dataset and MethodLong Main
- Chinese Metaphorical Relation ExtractionLong Findings
- Citance-Contextualized Summarization of Scientific PapersLong Findings
- CiteBench: A Benchmark for Scientific Citation Text GenerationLong Main
- CleanCoNLL: A Nearly Noise-Free Named Entity Recognition DatasetLong Main
- ClimateBERT-NetZero: Detecting and Assessing Net Zero and Reduction TargetsShort Main
- Clinical Contradiction DetectionLong Main
- Closed Boundary Learning for Classification Tasks with the Universum ClassLong Findings
- ClozEx: A Task toward Generation of English Cloze ExplanationLong Findings
- ClusterLLM: Large Language Models as a Guide for Text ClusteringLong Main
- ClusterPrompt: Cluster Semantic Enhanced Prompt Learning for New Intent DiscoveryLong Findings
- Clustering Pseudo Language Family in Multilingual Translation Models with Fisher Information MatrixShort Main
- Co$^2$PT: Mitigating Bias in Pre-trained Language Models through Counterfactual Contrastive Prompt TuningLong Findings
- Co-training and Co-distillation for Quality Improvement and Compression of Language ModelsLong Findings
- CoAnnotating: Uncertainty-Guided Work Allocation between Human and Large Language Models for Data AnnotationLong Main
- CoEdIT: Text Editing by Task-Specific Instruction TuningLong Findings
- CoF-CoT: Enhancing Large Language Models with Coarse-to-Fine Chain-of-Thought Prompting for Multi-domain NLU TasksShort Main
- CoLT5: Faster Long-Range Transformers with Conditional ComputationLong Main
- CoMPosT: Characterizing and Evaluating Caricature in LLM SimulationsLong Main
- CoRec: An Easy Approach for Coordination RecognitionShort Main
- CoSyn: Detecting Implicit Hate Speech in Online Conversations Using a Context Synergized Hyperbolic NetworkLong Main
- CoVariance-based Causal Debiasing for Entity and Relation ExtractionLong Findings
- Coarse-to-Fine Contrastive Learning in Image-Text-Graph Space for Improved Vision-Language CompositionalityLong Main
- Coarse-to-Fine Dual Encoders are Better Frame Identification LearnersLong Findings
- Code-Switching Metrics Using Intonation UnitsLong Main
- Code-Switching with Word Senses for Pretraining in Neural Machine TranslationLong Findings
- CodeBERTScore: Evaluating Code Generation with Pretrained Models of CodeLong Main
- CodeFusion: A Pre-trained Diffusion Model for Code GenerationShort Main
- CodeT5+: Open Code Large Language Models for Code Understanding and GenerationLong Main
- CodeTransOcean: A Comprehensive Multilingual Benchmark for Code TranslationLong Findings
- Cognate Transformer for Automated Phonological Reconstruction and Cognate Reflex PredictionLong Main
- Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?Short Main
- Coherent Entity Disambiguation via Modeling Topic and Categorical DependencyLong Findings
- Collaborative Generative AI: Integrating GPT-k for Efficient Editing in Text-to-Image GenerationShort Main
- CombLM: Adapting Black-Box Language Models through Small Fine-Tuned ModelsLong Main
- Combining Counting Processes and Classification Improves a Stopping Rule for Technology Assisted ReviewShort Findings
- Combining Denoising Autoencoders with Contrastive Learning to fine-tune Transformer ModelsLong Main
- Comparing Biases and the Impact of Multilingual Training across Multiple LanguagesLong Main
- Comparing Prompt-Based and Standard Fine-Tuning for Urdu Text ClassificationShort Findings
- Comparing Styles across LanguagesLong Main
- Comparing the Evaluation and Production of Loophole Behavior in Humans and Large Language ModelsLong Findings
- CompleQA: Benchmarking the Impacts of Knowledge Graph Completion Methods on Question AnsweringShort Findings
- Complex Event Schema Induction with Knowledge-Enriched Diffusion ModelLong Findings
- Complexity-Guided Curriculum Learning for Text GraphsLong Findings
- Compositional Generalization for Data-to-Text GenerationLong Findings
- CompoundPiece: Evaluating and Improving Decompounding Performance of Language ModelsLong Main
- Compressing Context to Enhance Inference Efficiency of Large Language ModelsLong Main
- Compressing and Debiasing Vision-Language Pre-Trained Models for Visual Question AnsweringLong Main
- ConPrompt: Pre-training a Language Model with Machine-Generated Data for Implicit Hate Speech DetectionLong Findings
- Conceptor-Aided Debiasing of Large Language ModelsLong Main
- Conceptual structure coheres in human cognition but not in large language modelsLong Main
- Condensing Multilingual Knowledge with Lightweight Language-Specific ModulesLong Main
- Conditional Natural Language InferenceLong Findings
- Conditioning on Dialog Acts improves Empathy Style TransferLong Findings
- Confidence-based Ensembling of Perspective-aware ModelsLong Main
- Conic10K: A Challenging Math Problem Understanding and Reasoning DatasetLong Findings
- Connecting degree and polarity: An artificial language learning studyLong Main
- Connecting the Dots: What Graph-Based Text Representations Work Best for Text Classification using Graph Neural Networks?Long Findings
- Consistency is Key: On Data-Efficient Modality Transfer in Speech TranslationShort Findings
- Consonant is all you need: a compact representation of English text for efficient NLPLong Findings
- Construction Artifacts in Metaphor Identification DatasetsShort Main
- Content- and Topology-Aware Representation Learning for Scientific Multi-LiteratureLong Main
- Context Compression for Auto-regressive Transformers with Sentinel TokensShort Main
- Context Quality Matters in Training Fusion-in-Decoder for Extractive Open-Domain Question AnsweringLong Findings
- Context-faithful Prompting for Large Language ModelsLong Findings
- Contextual Interaction for Argument Post Quality AssessmentLong Main
- Continual Dialogue State Tracking via Example-Guided Question AnsweringLong Main
- Continual Event Extraction with Semantic Confusion RectificationLong Main
- Continual Generalized Intent Discovery: Marching Towards Dynamic and Open-world Intent RecognitionLong Findings
- Continual Learning for Multilingual Neural Machine Translation via Dual Importance-based Model DivisionLong Main
- Continual Named Entity Recognition without Catastrophic ForgettingLong Main
- Continually Improving Extractive QA via Human FeedbackLong Main
- Contrastive Deterministic Autoencoders For Language ModelingLong Findings
- Contrastive Distant Supervision for Debiased and Denoised Machine Reading ComprehensionLong Findings
- Contrastive Learning for Inference in DialogueLong Main
- Contrastive Learning of Sentence Embeddings from ScratchLong Main
- Contrastive Learning-based Sentence Encoders Implicitly Weight Informative WordsShort Findings
- Contrastive Pre-training for Personalized Expert FindingLong Findings
- Controllable Chest X-Ray Report Generation from Longitudinal RepresentationsLong Findings
- Controllable Contrastive Generation for Multilingual Biomedical Entity LinkingLong Main
- Controlling Pre-trained Language Models for Grade-Specific Text SimplificationLong Main
- Conversation Chronicles: Towards Diverse Temporal and Relational Dynamics in Multi-Session ConversationsLong Main
- Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality InteractionLong Main
- Conversational Recommender System and Large Language Model Are Made for Each Other in E-commerce Pre-sales DialogueLong Findings
- Conversational Semantic Parsing using Dynamic Context GraphsLong Main
- Copyright Violations and Large Language ModelsShort Main
- CorefPrompt: Prompt-based Event Coreference Resolution by Measuring Event Type and Argument CompatibilitiesLong Main
- Counter Turing Test (CT2): AI-Generated Text Detection is Not as Easy as You May Think - Introducing AI Detectability Index (ADI)Long Main
- Countering Misinformation via Emotional Response GenerationLong Main
- Counting the Bugs in ChatGPT's Wugs: A Multilingual Investigation into the Morphological Capabilities of a Large Language ModelLong Main
- Coverage-based Example Selection for In-Context LearningLong Findings
- Critic-Driven Decoding for Mitigating Hallucinations in Data-to-text GenerationShort Main
- Cross-Cultural Analysis of Human Values, Morals, and Biases in Folk TalesLong Main
- Cross-Document Event Coreference Resolution on Discourse StructureLong Main
- Cross-Lingual Consistency of Factual Knowledge in Multilingual Language ModelsLong Main
- Cross-Lingual Cross-Target Stance Detection with Dual Knowledge Distillation FrameworkLong Main
- Cross-Modal Conceptualization in Bottleneck ModelsLong Main
- Cross-lingual Open-Retrieval Question Answering for African LanguagesLong Findings
- Cross-lingual Prompting: Improving Zero-shot Chain-of-Thought Reasoning across LanguagesLong Main
- Cross-lingual Transfer Can Worsen Bias in Sentiment AnalysisLong Main
- Cross-modality Data Augmentation for End-to-End Sign Language TranslationLong Findings
- Crossing the Aisle: Unveiling Partisan and Counter-Partisan Events in News ReportingShort Findings
- Crossing the Threshold: Idiomatic Machine Translation through Retrieval Augmentation and Loss WeightingLong Main
- Crosslingual Transfer Learning for Low-Resource Languages Based on Multilingual Colexification GraphsLong Findings
- Crystal: Introspective Reasoners Reinforced with Self-FeedbackLong Main
- Cue-CoT: Chain-of-thought Prompting for Responding to In-depth Dialogue Questions with LLMsLong Findings
- Cultural Compass: Predicting Transfer Learning Success in Offensive Language Detection with Cultural FeaturesLong Findings
- Cultural Concept Adaptation on Multimodal ReasoningLong Main
- Culturally Aware Natural Language InferenceLong Findings
- D$^2$TV: Dual Knowledge Distillation and Target-oriented Vision Modeling for Many-to-Many Multimodal SummarizationLong Findings
- DADA: Dialect Adaptation via Dynamic Aggregation of Linguistic RulesLong Main
- DALE: Generative Data Augmentation for Low-Resource Legal NLPLong Main
- DEPN: Detecting and Editing Privacy Neurons in Pretrained Language ModelsLong Main
- DISCO: A Large Scale Human Annotated Corpus for Disfluency Correction in Indo-European LanguagesLong Findings
- DIVE: Towards Descriptive and Diverse Visual Commonsense GenerationLong Main
- DNA: Denoised Neighborhood Aggregation for Fine-grained Category DiscoveryLong Main
- DPP-TTS: Diversifying prosodic features of speech via determinantal point processesLong Main
- DRAFT: Dense Retrieval Augmented Few-shot Topic classifier FrameworkLong Findings
- DREAM: Deployment of Recombination and Ensembles in Argument MiningLong Main
- DSI++: Updating Transformer Memory with New DocumentsLong Main
- DUMB: A Dutch Model BenchmarkLong Main
- DUnE: Dataset for Unified EditingLong Main
- Dancing Between Success and Failure: Edit-level Simplification Evaluation using SALSALong Main
- Data Augmentation for Code Translation with Comparable Corpora and Multiple ReferencesLong Findings
- Data Factors for Better Compositional GeneralizationLong Main
- Data Selection Curriculum for Abstractive Text SummarizationShort Findings
- Data Similarity is Not Enough to Explain Language Model PerformanceShort Main
- Data-efficient Active Learning for Structured Prediction with Partial Annotation and Self-TrainingLong Findings
- Dataset Bias Mitigation in Multiple-Choice Visual Question Answering and BeyondLong Findings
- DeCrisisMB: Debiased Semi-Supervised Learning for Crisis Tweet Classification via Memory BankLong Findings
- DeSIQ: Towards an Unbiased, Challenging Benchmark for Social Intelligence UnderstandingLong Main
- DeTiME: Diffusion-Enhanced Topic Modeling using Encoder-decoder based LLMLong Findings
- Debias NLU Datasets via Training-free PerturbationsLong Findings
- Debiasing Made State-of-the-art: Revisiting the Simple Seed-based Weak Supervision for Text ClassificationLong Main
- Debiasing Multimodal Models via Causal Information MinimizationLong Findings
- DecipherPref: Analyzing Influential Factors in Human Preference Judgments via GPT-4Long Main
- Deciphering Stereotypes in Pre-Trained Language ModelsLong Main
- DecoMT: Decomposed Prompting for Machine Translation Between Related Languages using Large Language ModelsLong Main
- Decoding Stumpers: Large Language Models vs. Human Problem-SolversShort Findings
- Decoding the Silent Majority: Inducing Belief Augmented Social Graph with Large Language Model for Response ForecastingLong Main
- Decomposed Prompt Tuning via Low-Rank ReparameterizationLong Findings
- Decomposing Complex Queries for Tip-of-the-tongue RetrievalLong Findings
- Deep Natural Language Feature Learning for Interpretable PredictionLong Main
- Defining a New NLP PlaygroundLong Findings
- Definitions Matter: Guiding GPT for Multi-label ClassificationShort Findings
- DeltaScore: Fine-Grained Story Evaluation with PerturbationsLong Findings
- DelucionQA: Detecting Hallucinations in Domain-specific Question AnsweringLong Findings
- DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language GroundingLong Findings
- DemoNSF: A Multi-task Demonstration-based Generative Framework for Noisy Slot Filling TaskShort Findings
- DemoSG: Demonstration-enhanced Schema-guided Generation for Low-resource Event ExtractionLong Findings
- Democratizing LLMs: An Exploration of Cost-Performance Trade-offs in Self-Refined Open-Source ModelsLong Findings
- Democratizing Reasoning Ability: Tailored Learning from Large Language ModelLong Main
- Demystifying Prompts in Language Models via Perplexity EstimationLong Findings
- Dense Retrieval as Indirect Supervision for Large-space Decision MakingLong Findings
- Density-Aware Prototypical Network for Few-Shot Relation ClassificationLong Findings
- DepNeCTI: Dependency-based Nested Compound Type Identification for SanskritLong Findings
- DepWiGNN: A Depth-wise Graph Neural Network for Multi-hop Spatial Reasoning in TextLong Findings
- Describe Me an Auklet: Generating Grounded Perceptual Category DescriptionsLong Main
- Descriptive Prompt Paraphrasing for Target-Oriented Multimodal Sentiment ClassificationLong Findings
- DetGPT: Detect What You Need via ReasoningLong Main
- DetectLLM: Leveraging Log Rank Information for Zero-Shot Detection of Machine-Generated TextLong Findings
- Detecting Erroneously Recognized Handwritten Byzantine TextLong Findings
- Detecting Propaganda Techniques in Code-Switched Social Media TextLong Main
- Detecting Syntactic Change with Pre-trained Transformer ModelsLong Findings
- Detecting and Mitigating Hallucinations in Multilingual SummarisationLong Main
- Detection of Multiple Mental Disorders from Social Media with Two-Stream Psychiatric ExpertsLong Main
- Detrimental Contexts in Open-Domain Question AnsweringLong Findings
- DiFair: A Benchmark for Disentangled Assessment of Gender Knowledge and BiasLong Findings
- DiNeR: A Large Realistic Dataset for Evaluating Compositional GeneralizationLong Main
- DiQAD: A Benchmark Dataset for Open-domain Dialogue Quality AssessmentLong Findings
- DiSTRICT: Dialogue State Tracking with Retriever Driven In-Context TuningLong Main
- DialCoT Meets PPO: Decomposing and Exploring Reasoning Paths in Smaller Language ModelsLong Main
- DialGuide: Aligning Dialogue Model Behavior with Developer GuidelinesLong Findings
- Dialect Transfer for Swiss German Speech TranslationLong Findings
- Dialect-to-Standard Normalization: A Large-Scale Multilingual EvaluationLong Findings
- DialogQAE: N-to-N Question Answer Pair Extraction from Customer Service ChatlogLong Findings
- Dialogizer: Context-aware Conversational-QA Dataset Generation from Textual SourcesLong Main
- Dialogue Act-Aided Backchannel Prediction Using Multi-Task LearningShort Findings
- Dialogue Chain-of-Thought Distillation for Commonsense-aware Conversational AgentsLong Main
- Dialogue Medical Information Extraction with Medical-Item Graph and Dialogue-Status Enriched RepresentationLong Findings
- Did You Mean...? Confidence-based Trade-offs in Semantic ParsingShort Main
- DiffS2UT: A Semantic Preserving Diffusion Model for Textless Direct Speech-to-Speech TranslationLong Main
- Difference-Masking: Choosing What to Mask in Continued PretrainingLong Findings
- DiffuSeq-v2: Bridging Discrete and Continuous Text Spaces for Accelerated Seq2Seq Diffusion ModelsShort Findings
- DiffuVST: Narrating Fictional Scenes with Global-History-Guided Denoising ModelsLong Findings
- Diffusion Language Model with Query-Document Relevance for Query-Focused SummarizationLong Findings
- DiffusionRet: Diffusion-Enhanced Generative Retriever using Constrained DecodingLong Findings
- DiffusionSL: Sequence Labeling via Tag Diffusion ProcessLong Findings
- Dimensions of Online Conflict: Towards Modeling AgonismLong Findings
- Dior-CVAE: Pre-trained Language Models and Diffusion Priors for Variational Dialog GenerationLong Findings
- DisCo: Distilled Student Models Co-training for Semi-supervised Text MiningLong Main
- Discourse Sense Flows: Modelling the Rhetorical Style of Documents across Various DomainsLong Findings
- Discourse Structures Guided Fine-grained Propaganda IdentificationLong Main
- Discovering Highly Influential Shortcut Reasoning: An Automated Template-Free ApproachShort Findings
- Discovering Universal Geometry in Embeddings with ICALong Main
- Disentangling Extraction and Reasoning in Multi-hop Spatial ReasoningLong Findings
- Disentangling Structure and Style: Political Bias Detection in News by Inducing Document HierarchyLong Findings
- Disentangling Transformer Language Models as Superposed Topic ModelsLong Main
- Disfluent Cues for Enhanced Speech Understanding in Large Language ModelsLong Findings
- Dissecting In-Context Learning of Translations in GPT-3Short Findings
- Dissecting Recall of Factual Associations in Auto-Regressive Language ModelsLong Main
- Distance-Based Propagation for Efficient Knowledge Graph ReasoningLong Main
- DistillCSE: Distilled Contrastive Learning for Sentence EmbeddingsLong Findings
- Distilling ChatGPT for Explainable Automated Student Answer AssessmentLong Findings
- Ditto: A Simple and Efficient Approach to Improve Sentence EmbeddingsShort Main
- Diversify Question Generation with Retrieval-Augmented Style TransferLong Main
- Diversifying language models for lesser-studied languages and language-usage contexts: A case of second language KoreanLong Findings
- Diversity Enhanced Narrative Question Generation for StorybooksLong Main
- Do All Languages Cost the Same? Tokenization in the Era of Commercial Language ModelsLong Main
- Do LLMs Understand Social Knowledge? Evaluating the Sociability of Large Language Models with SocKET BenchmarkLong Main
- Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language ModelsLong Main
- Do Stochastic Parrots have Feelings Too? Improving Neural Detection of Synthetic Text via Emotion RecognitionLong Findings
- Do “English” Named Entity Recognizers Work Well on Global Englishes?Long Findings
- DocAsRef: An Empirical Study on Repurposing Reference-based Summary Quality Metrics as Reference-free MetricsShort Findings
- DocSplit: Simple Contrastive Pretraining for Large Document EmbeddingsShort Findings
- DocTrack: A Visually-Rich Document Dataset Really Aligned with Human Eye Movement for Machine ReadingLong Findings
- Document-Level Machine Translation with Large Language ModelsLong Main
- Document-level Relationship Extraction by Bidirectional Constraints of Beta RulesLong Main
- Does Listener Gaze in Face-to-Face Interaction Follow the Entropy Rate Constancy Principle: An Empirical StudyShort Findings
- Does the Correctness of Factual Knowledge Matter for Factual Knowledge-Enhanced Pre-trained Language Models?Long Main
- Dolphin: A Challenging and Diverse Benchmark for Arabic NLGLong Findings
- Domain Adaptation for Conversational Query Production with the RAG Model FeedbackLong Findings
- Domain Adaptation for Sentiment Analysis Using Robust Internal RepresentationsLong Findings
- Domain Private Transformers for Multi-Domain Dialog SystemsShort Findings
- Don't waste a single annotation: improving single-label classifiers through soft labelsShort Findings
- Don’t Add, don’t Miss: Effective Content Preserving Generation from Pre-Selected Text SpansLong Findings
- Don’t Trust ChatGPT when your Question is not in English: A Study of Multilingual Abilities and Types of LLMsLong Main
- Doolittle: Benchmarks and Corpora for Academic Writing FormalizationLong Main
- Dr ChatGPT tell me what I want to hear: How different prompts impact health answer correctnessLong Main
- Drilling Down into the Discourse Structure with LLMs for Long Document Question AnsweringLong Findings
- Dual-Channel Span for Aspect Sentiment Triplet ExtractionLong Main
- Dual-Feedback Knowledge Retrieval for Task-Oriented Dialogue SystemsLong Main
- DueT: Image-Text Contrastive Transfer Learning with Dual-adapter TuningLong Main
- Dynamic Low-rank Estimation for Transformer-based Language ModelsLong Findings
- Dynamic Open-book Prompt for Conversational Recommender SystemLong Findings
- Dynamic Stance: Modeling Discussions by Labeling the InteractionsLong Findings
- Dynamic Stashing Quantization for Efficient Transformer TrainingShort Findings
- Dynamic Top-k Estimation Consolidates Disagreement between Feature Attribution MethodsShort Main
- Dynamic Voting for Efficient Reasoning in Large Language ModelsLong Findings
- Dynosaur: A Dynamic Growth Paradigm for Instruction-Tuning Data CurationLong Main
- E-CORE: Emotion Correlation Enhanced Empathetic Dialogue GenerationLong Main
- EARA: Improving Biomedical Semantic Textual Similarity with Entity-Aligned Attention and Retrieval AugmentationLong Findings
- ECHo: A Visio-Linguistic Dataset for Event Causality Inference via Human-Centric ReasoningLong Findings
- EDIS: Entity-Driven Image Search over Multimodal Web ContentLong Main
- EDeR: Towards Understanding Dependency Relations Between EventsLong Main
- EMO-KNOW: A Large Scale Dataset on Emotion-CauseShort Findings
- ESPVR: Entity Spans Position Visual Regions for Multimodal Named Entity RecognitionLong Findings
- EXPLAIN, EDIT, GENERATE: Rationale-Sensitive Counterfactual Data Augmentation for Multi-hop Fact VerificationLong Main
- EZ-STANCE: A Large Dataset for Zero-Shot Stance DetectionLong Findings
- EasyQuant: An Efficient Data-free Quantization Algorithm for LLMsLong Main
- Ecologically Valid Explanations for Label Variation in NLIShort Findings
- EconBERTa: Towards Robust Extraction of Named Entities in EconomicsLong Findings
- Editing Common Sense in TransformersLong Main
- Editing Large Language Models: Problems, Methods, and OpportunitiesLong Main
- Effects of Human Adversarial and Affable Samples on BERT GeneralizationLong Findings
- Effects of sub-word segmentation on performance of transformer language modelsLong Main
- Efficient Algorithms for Recognizing Weighted Tree-Adjoining LanguagesLong Main
- Efficient Classification of Long Documents via State-Space ModelsShort Main
- Efficient Continue Training of Temporal Language Model with Structural InformationLong Findings
- Efficient Cross-Task Prompt Tuning for Few-Shot Conversational Emotion RecognitionLong Findings
- Efficient Data Learning for Open Information Extraction with Pre-trained Language ModelsShort Findings
- Efficient Grammatical Error Correction Via Multi-Task Training and Optimized Training ScheduleLong Main
- Efficient Latent Variable Modeling for Knowledge-Grounded Dialogue GenerationLong Findings
- Efficient Long-Range Transformers: You Need to Attend More, but Not Necessarily at Every LayerLong Findings
- Efficient Multilingual Language Model Compression through Vocabulary TrimmingLong Findings
- Efficient k-NN Search with Cross-Encoders using Adaptive Multi-Round CUR DecompositionShort Findings
- Efficiently Enhancing Zero-Shot Performance of Instruction Following Model via Retrieval of Soft PromptLong Findings
- Elaborative Simplification as Implicit Questions Under DiscussionLong Main
- Emergence of Abstract State Representations in Embodied Sequence ModelingLong Main
- Emergent Inabilities? Inverse Scaling Over the Course of PretrainingShort Findings
- Empathy Intent Drives Empathy DetectionLong Main
- Empirical Study of Zero-Shot NER with ChatGPTLong Main
- Empower Nested Boolean Logic via Self-Supervised Curriculum LearningLong Main
- Empowering Psychotherapy with Large Language Models: Cognitive Distortion Detection through Diagnosis of Thought PromptingShort Findings
- Emptying the Ocean with a Spoon: Should We Edit Models?Short Findings
- Enabling Large Language Models to Generate Text with CitationsLong Main
- Enabling Unsupervised Neural Machine Translation with Word-level Visual RepresentationsLong Findings
- End-to-End Autoregressive Retrieval via Bootstrapping for Smart Reply SystemsLong Findings
- End-to-End Single-Channel Speaker-Turn Aware Conversational Speech TranslationLong Main
- End-to-end Adversarial Sample Generation for Data AugmentationLong Findings
- End-to-end Task-oriented Dialogue: A Survey of Tasks, Methods, and Future DirectionsLong Main
- Energy and Carbon Considerations of Fine-Tuning BERTShort Findings
- Enhanced Simultaneous Machine Translation with Word-level PoliciesLong Findings
- Enhancing Abstractiveness of Summarization Models through Calibrated DistillationLong Findings
- Enhancing Accessible Communication: from European Portuguese to Portuguese Sign LanguageShort Findings
- Enhancing Argument Structure Extraction with Efficient Leverage of Contextual InformationShort Findings
- Enhancing Biomedical Lay Summarisation with External Knowledge GraphsLong Main
- Enhancing Chat Language Models by Scaling High-quality Instructional ConversationsLong Main
- Enhancing Code-Switching for Cross-lingual SLU: A Unified View of Semantic and Grammatical CoherenceShort Main
- Enhancing Computation Efficiency in Large Language Models through Weight and Activation QuantizationLong Main
- Enhancing Conversational Search: Large Language Model-Aided Informative Query RewritingLong Findings
- Enhancing Emotion Recognition in Conversation via Multi-view Feature Alignment and MemorizationLong Findings
- Enhancing Generative Retrieval with Reinforcement Learning from Relevance FeedbackLong Main
- Enhancing Low-resource Fine-grained Named Entity Recognition by Leveraging Coarse-grained DatasetsLong Main
- Enhancing Neural Machine Translation with Semantic UnitsLong Findings
- Enhancing Reasoning Capabilities by Instruction Learning and Chain-of-Thoughts for Implicit Discourse Relation RecognitionShort Findings
- Enhancing Retrieval-Augmented Large Language Models with Iterative Retrieval-Generation SynergyLong Findings
- Enhancing Scalability of Pre-trained Language Models via Efficient Parameter SharingLong Findings
- Enhancing Structured Evidence Extraction for Fact VerificationLong Main
- Enhancing Task-oriented Dialogue Systems with Generative Post-processing NetworksLong Main
- Enhancing Text-to-SQL Capabilities of Large Language Models: A Study on Prompt Design StrategiesLong Findings
- Enhancing Textbooks with Visuals from the Web for Improved LearningLong Main
- Enhancing Uncertainty-Based Hallucination Detection with Stronger FocusLong Main
- Enhancing the Ranking Context of Dense Retrieval through Reciprocal Nearest NeighborsLong Main
- Ensemble-Instruct: Instruction Tuning Data Generation with a Heterogeneous Mixture of LMsLong Findings
- EntSUMv2: Dataset, Models and Evaluation for More Abstractive Entity-Centric SummarizationShort Main
- Entity Disambiguation on a Tight Labeling BudgetShort Findings
- Entity-Based Evaluation of Political Bias in Automatic SummarizationShort Findings
- EpiK-Eval: Evaluation for Language Models as Epistemic ModelsLong Main
- Epsilon Sampling Rocks: Investigating Sampling Strategies for Minimum Bayes Risk Decoding for Machine TranslationLong Findings
- Error Detection for Text-to-SQL Semantic ParsingLong Findings
- Establishing Trustworthiness: Rethinking Tasks and Model EvaluationShort Main
- Estimating Large Language Model Capabilities without Labeled Test DataLong Findings
- Ethical Reasoning over Moral Alignment: A Case and Framework for In-Context Ethical Policies in LLMsLong Findings
- EtiCor: Corpus for Analyzing LLMs for EtiquettesShort Main
- Evaluating Bias and Fairness in Gender-Neutral Pretrained Vision-and-Language ModelsLong Main
- Evaluating Cross-Domain Text-to-SQL Models and BenchmarksLong Main
- Evaluating Dependencies in Fact Editing for Language Models: Specificity and Implication AwarenessLong Findings
- Evaluating Emotion Arcs Across Languages: Bridging the Global Divide in Sentiment AnalysisLong Findings
- Evaluating Evaluation Metrics: A Framework for Analyzing NLG Evaluation Metrics using Measurement TheoryLong Main
- Evaluating Large Language Models on Controlled Generation TasksLong Main
- Evaluating Object Hallucination in Large Vision-Language ModelsLong Main
- Evaluating Parameter-Efficient Finetuning Approaches for Pre-trained Models on the Financial DomainShort Findings
- Evaluating Subjective Cognitive Appraisals of Emotions from Large Language ModelsLong Findings
- Evaluating Verifiability in Generative Search EnginesLong Findings
- Evaluating and Enhancing the Robustness of Code Pre-trained Models through Structure-Aware Adversarial Samples GenerationLong Findings
- Evaluating and Modeling Attribution for Cross-Lingual Question AnsweringLong Main
- Evaluating the Knowledge Base Completion Potential of GPTShort Findings
- Evaluating the Rationale Understanding of Critical Reasoning in Logical Reading ComprehensionLong Main
- Evaluation Metrics in the Era of GPT-4: Reliably Evaluating Large Language Models on Sequence to Sequence TasksShort Main
- Evaluation of African American Language Bias in Natural Language GenerationLong Main
- Event Causality Extraction via Implicit Cause-Effect InteractionsLong Main
- Event Ontology Completion with Hierarchical Structure Evolution NetworksLong Main
- Event-Location Tracking in Narratives: A Case Study on Holocaust TestimoniesLong Main
- Examining Inter-Consistency of Large Language Models Collaboration: An In-depth Analysis via DebateLong Findings
- Example-based Hypernetworks for Multi-source Adaptation to Unseen DomainsLong Findings
- Exchange-of-Thought: Enhancing Large Language Model Capabilities through Cross-Model CommunicationLong Main
- Execution-Based Evaluation for Open-Domain Code GenerationLong Findings
- ExpNote: Black-box Large Language Models are better Task Solvers with Experience NotebookShort Findings
- Expand, Highlight, Generate: RL-driven Document Generation for Passage RerankingLong Main
- Explain-then-translate: an analysis on improving program translation with self-generated explanationsLong Findings
- ExplainCPE: A Free-text Explanation Benchmark of Chinese Pharmacist ExaminationLong Findings
- Explainable Claim Verification via Knowledge-Grounded Reasoning with Large Language ModelsLong Findings
- Explaining Interactions Between Text SpansLong Main
- Explaining with Contrastive Phrasal Highlighting: A Case Study in Assisting Humans to Detect Translation DifferencesLong Main
- Explanation Selection Using Unlabeled Data for Chain-of-Thought PromptingLong Main
- Explicit Alignment and Many-to-many Entailment Based Reasoning for Conversational Machine ReadingLong Findings
- Explicit Planning Helps Language Models in Logical ReasoningLong Main
- Exploiting Asymmetry for Synthetic Training Data Generation: SynthIE and the Case of Information ExtractionLong Main
- Exploiting Contrastive Learning and Numerical Evidence for Confusing Legal Judgment PredictionLong Findings
- Exploiting Emotion-Semantic Correlations for Empathetic Response GenerationLong Findings
- Explore the Way: Exploring Reasoning Path by Bridging Entities for Effective Cross-Document Relation ExtractionShort Findings
- Explore-Instruct: Enhancing Domain-Specific Instruction Coverage through Active ExplorationLong Main
- Exploring All-In-One Knowledge Distillation Framework for Neural Machine TranslationLong Main
- Exploring Chain of Thought Style Prompting for Text-to-SQLLong Main
- Exploring Context-Aware Evaluation Metrics for Machine TranslationShort Findings
- Exploring Discourse Structure in Document-level Machine TranslationLong Main
- Exploring Graph Pre-training for Aspect-based Sentiment AnalysisLong Findings
- Exploring In-Context Learning for Knowledge Grounded Dialog GenerationLong Findings
- Exploring Jiu-Jitsu Argumentation for Writing Peer Review RebuttalsLong Main
- Exploring Large Language Models for Multi-Modal Out-of-Distribution DetectionLong Findings
- Exploring Linguistic Probes for Morphological InflectionShort Main
- Exploring Linguistic Properties of Monolingual BERTs with Typological Classification among LanguagesLong Findings
- Exploring the Boundaries of GPT-4 in RadiologyLong Main
- Exploring the Cognitive Knowledge Structure of Large Language Models: An Educational Diagnostic Assessment ApproachShort Findings
- Exploring the Effectiveness of Multi-Lingual Commonsense Knowledge-Aware Open-Domain Dialogue Response GenerationLong Findings
- Exploring the Impact of Corpus Diversity on Financial Pretrained Language ModelsShort Findings
- Exploring the Impact of Model Scaling on Parameter-Efficient TuningLong Main
- Exploring the Numerical Reasoning Capabilities of Language Models: A Comprehensive Analysis on Tabular DataLong Findings
- Exploring the Potential of Large Language Models in Generating Code-Tracing Questions for Introductory Programming CoursesShort Findings
- Exploring the Sensitivity of LLMs' Decision-Making Capabilities: Insights from Prompt Variations and HyperparametersShort Findings
- Expository Text Generation: Imitate, Retrieve, ParaphraseLong Main
- Extractive Summarization via ChatGPT for Faithful Summary GenerationShort Findings
- Extrapolating Multilingual Understanding Models as Multilingual GeneratorsLong Findings
- Eyes Show the Way: Modelling Gaze Behaviour for Hallucination DetectionLong Findings
- FACTIFY3M: A benchmark for multimodal fact verification with explainability through 5W Question-AnsweringLong Main
- FANToM: A Benchmark for Stress-testing Machine Theory of Mind in InteractionsLong Main
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationLong Main
- FFAEval: Evaluating Dialogue System via Free-For-All RankingLong Findings
- FLatS: Principled Out-of-Distribution Detection with Feature-Based Likelihood Ratio ScoreShort Main
- FOCUS: Effective Embedding Initialization for Monolingual Specialization of Multilingual ModelsLong Main
- FREDSum: A Dialogue Summarization Corpus for French Political DebatesLong Findings
- FaLA: Fast Linear Adaptation for Replacing Backbone Models on Edge DevicesLong Findings
- FaMeSumm: Investigating and Improving Faithfulness of Medical SummarizationLong Main
- FactKB: Generalizable Factuality Evaluation using Language Models Enhanced with Factual KnowledgeLong Main
- FactSpotter: Evaluating the Factual Faithfulness of Graph-to-Text GenerationLong Findings
- Factual Relation Discrimination for Factuality-oriented Abstractive SummarizationLong Findings
- Failures Pave the Way: Enhancing Large Language Models through Tuning-free Rule AccumulationLong Main
- Fair Text Classification with Wasserstein IndependenceLong Main
- Fair Without Leveling Down: A New Intersectional Fairness DefinitionLong Main
- Faithful Model Evaluation for Model-Based MetricsShort Main
- Fast and Accurate Factual Inconsistency Detection Over Long DocumentsLong Main
- Fast and Robust Early-Exiting Framework for Autoregressive Language Models with Synchronized Parallel DecodingLong Main
- Faster Minimum Bayes Risk Decoding with Confidence-based PruningShort Main
- FedID: Federated Interactive Distillation for Large-Scale Pretraining Language ModelsLong Main
- FedTherapist: Mental Health Monitoring with User-Generated Linguistic Expressions on Smartphones via Federated LearningShort Main
- Federated Learning of Large Language Models with Parameter-Efficient Prompt Tuning and Adaptive OptimizationLong Main
- Few-shot Unified Question Answering: Tuning Models or Prompts?Long Findings
- Fidelity-Enriched Contrastive Search: Reconciling the Faithfulness-Diversity Trade-Off in Text GenerationShort Main
- Fighting Fire with Fire: The Dual Role of LLMs in Crafting and Detecting Elusive DisinformationLong Main
- Filling the Image Information Gap for VQA: Prompting Large Language Models to Proactively Ask QuestionsLong Findings
- FinEntity: Entity-level Sentiment Classification for Financial TextsShort Main
- FinGPT: Large Generative Models for a Small LanguageLong Main
- Find-2-Find: Multitask Learning for Anaphora Resolution and Object LocalizationLong Main
- Finding Authentic Counterhate Arguments: A Case Study with Public FiguresLong Main
- Finding Common Ground: Annotating and Predicting Common Ground in Spoken ConversationsLong Findings
- Finding Support Examples for In-Context LearningLong Findings
- Fine-grained Conversational Decoding via Isotropic and Proximal SearchShort Main
- Fine-tuned LLMs Know More, Hallucinate Less with Few-Shot Sequence-to-Sequence Semantic Parsing over WikidataLong Main
- FinePrompt: Unveiling the Role of Finetuned Inductive Bias on Compositional Reasoning in GPT-4Short Findings
- Flatness-Aware Prompt Selection Improves Accuracy and Sample EfficiencyLong Findings
- Focus Your Attention (with Adaptive IIR Filters)Long Main
- Focus on the Core: Efficient Attention via Pruned Token Compression for Document ClassificationLong Findings
- For Generated Text, Is NLI-Neutral Text the Best Text?Short Findings
- FreeAL: Towards Human-Free Active Learning in the Era of Large Language ModelsLong Main
- Frequency Balanced Datasets Lead to Better Language ModelsLong Findings
- From Complex to Simple: Unraveling the Cognitive Tree for Reasoning with Small Language ModelsLong Findings
- From Dissonance to Insights: Dissecting Disagreements in Rationale Construction for Case Outcome ClassificationLong Main
- From Heuristic to Analytic: Cognitively Motivated Strategies for Coherent Physical Commonsense ReasoningLong Main
- From Multilingual Complexity to Emotional Clarity: Leveraging Commonsense to Unveil Emotions in Code-Mixed DialoguesLong Main
- From Parse-Execute to Parse-Execute-Refine: Improving Semantic Parser for Complex Question Answering over Knowledge BaseLong Main
- From Relevance to Utility: Evidence Retrieval with Feedback for Fact VerificationShort Findings
- From Simple to Complex: A Progressive Framework for Document-level Informative Argument ExtractionLong Findings
- From Speculation Detection to Trustworthy Relational Tuples in Information ExtractionLong Findings
- From Values to Opinions: Predicting Human Behaviors and Stances Using Value-Injected Large Language ModelsLong Main
- From Words to Wires: Generating Functioning Electronic Devices from Natural Language DescriptionsLong Findings
- From Wrong To Right: A Recursive Approach Towards Vision-Language ExplanationLong Main
- Frugal Prompting for Dialog ModelsLong Findings
- Fusing Temporal Graphs into Transformers for Time-Sensitive Question AnsweringLong Findings
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentLong Main
- G-SPEED: General SParse Efficient Editing MoDelLong Findings
- GATITOS: Using a New Multilingual Lexicon for Low-resource Machine TranslationLong Main
- GBT: Generative Boosting Training Approach for Paraphrase IdentificationLong Findings
- GD-COMET: A Geo-Diverse Commonsense Inference ModelShort Main
- GDA: Grammar-based Data Augmentation for Text Classification using Slot InformationLong Findings
- GEM: Gestalt Enhanced Markup Language Model for Web Understanding via Render TreeLong Main
- GEMINI: Controlling The Sentence-Level Summary Style in Abstractive Text SummarizationLong Main
- GLEN: General-Purpose Event Detection for Thousands of TypesLong Main
- GLEN: Generative Retrieval via Lexical Index LearningLong Main
- GLGR: Question-aware Global-to-Local Graph Reasoning for Multi-party Dialogue Reading ComprehensionLong Findings
- GNAT: A General Narrative Alignment ToolLong Main
- GPT Deciphering Fedspeak: Quantifying Dissent Among Hawks and DovesShort Findings
- GPT-4 as an Effective Zero-Shot Evaluator for Scientific Figure CaptionsShort Findings
- GPT-RE: In-context Learning for Relation Extraction using Large Language ModelsLong Main
- GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLPLong Main
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head CheckpointsShort Main
- GRACE: Discriminator-Guided Chain-of-Thought ReasoningLong Findings
- GRENADE: Graph-Centric Language Model for Self-Supervised Representation Learning on Text-Attributed GraphsLong Findings
- GRI: Graph-based Relative Isomorphism of Word Embedding SpacesLong Findings
- GROOViST: A Metric for Grounding Objects in Visual StorytellingShort Main
- GROVE: A Retrieval-augmented Complex Story Generation Framework with A Forest of EvidenceLong Findings
- GSAP-NER: A Novel Task, Corpus, and Baseline for Scholarly Entity Extraction Focused on Machine Learning Models and DatasetsLong Findings
- GTA: Gated Toxicity Avoidance for LM Performance PreservationLong Findings
- GazeVQA: A Video Question Answering Dataset for Multiview Eye-Gaze Task-Oriented CollaborationsLong Main
- GenKIE: Robust Generative Multimodal Document Key Information ExtractionLong Findings
- Gender Biases in Automatic Evaluation Metrics for Image CaptioningLong Main
- Generalizing Few-Shot Named Entity Recognizers to Unseen Domains with Type-Related FeaturesLong Findings
- Generating Commonsense Counterfactuals for Stable Relation ExtractionLong Main
- Generating Data for Symbolic Language with Large Language ModelsLong Main
- Generating Extractive Answers: Gated Recurrent Memory Reader for Conversational Question AnsweringShort Findings
- Generating Summaries with Controllable Readability LevelsLong Main
- Generating and Evaluating Tests for K-12 Students with Language Model Simulations: A Case Study on Sentence Reading EfficiencyLong Main
- Generative Adversarial Training with Perturbed Token Detection for Model RobustnessLong Main
- Generative Calibration for In-context LearningLong Findings
- Generative Emotion Cause Triplet Extraction in Conversations with Commonsense KnowledgeLong Findings
- Generative Spoken Language Model based on continuous word-sized audio tokensLong Main
- Generative Table Pre-training Empowers Models for Tabular PredictionLong Main
- GeoLM: Empowering Language Models for Geospatially Grounded Language UnderstandingLong Main
- Geographical Erasure in Language GenerationLong Findings
- Getting MoRE out of Mixture of Language Model Reasoning ExpertsLong Findings
- Give Me the Facts! A Survey on Factual Knowledge Probing in Pre-trained Language ModelsLong Findings
- Global Structure Knowledge-Guided Relation Extraction Method for Visually-Rich DocumentLong Findings
- Global Voices, Local Biases: Socio-Cultural Prejudices across LanguagesLong Main
- GlobalBench: A Benchmark for Global Progress in Natural Language ProcessingLong Main
- GlotLID: Language Identification for Low-Resource LanguagesLong Findings
- Goal-Driven Explainable Clustering via Language DescriptionsLong Main
- Gold: A Global and Local-aware Denoising Framework for Commonsense Knowledge Graph Noise DetectionLong Findings
- Good Meta-tasks Make A Better Cross-lingual Meta-transfer Learning for Low-resource LanguagesLong Findings
- Goodtriever: Adaptive Toxicity Mitigation with Retrieval-augmented ModelsLong Findings
- GradSim: Gradient-Based Language Grouping for Effective Multilingual TrainingLong Main
- Gradient-based Gradual Pruning for Language-Specific Multilingual Neural Machine TranslationLong Main
- Gradually Excavating External Knowledge for Implicit Complex Question AnsweringLong Findings
- Grammar-Constrained Decoding for Structured NLP Tasks without FinetuningLong Main
- Grammatical Error Correction via Mixed-Grained Weighted TrainingLong Findings
- Granularity Matters: Pathological Graph-driven Cross-modal Alignment for Brain CT Report GenerationLong Main
- Graph vs. Sequence: An Empirical Study on Knowledge Forms for Knowledge-Grounded DialogueLong Main
- GreedyCAS: Unsupervised Scientific Abstract Segmentation with Normalized Mutual InformationLong Main
- Grounded and well-rounded: a methodological approach to the study of cross-modal and cross-lingual groundingLong Findings
- Grounding Visual Illusions in Language: Do Vision-Language Models Perceive Illusions Like Humans?Long Main
- Guideline Learning for In-Context Information ExtractionLong Main
- Guiding LLM to Fool Itself: Automatically Manipulating Machine Reading Comprehension Shortcut TriggersShort Findings
- HANSEN: Human and AI Spoken Text Benchmark for Authorship AnalysisLong Findings
- HARE: Explainable Hate Speech Detection with Step-by-Step ReasoningShort Findings
- HEAR: Hearing Enhanced Audio Response for Video-grounded DialogueLong Findings
- HFMRE: Constructing Huffman Tree in Bags to Find Excellent Instances for Distantly Supervised Relation ExtractionLong Findings
- HPE: Answering Complex Questions over Text by Hybrid Question Parsing and ExecutionLong Findings
- HadSkip: Homotopic and Adaptive Layer Skipping of Pre-trained Language Models for Efficient InferenceLong Findings
- HalOmi: A Manually Annotated Benchmark for Multilingual Hallucination and Omission Detection in Machine TranslationLong Main
- Hallucination Detection for Generative Large Language Models by Bayesian Sequential EstimationLong Main
- Hallucination Detection for Grounded Instruction GenerationShort Findings
- Hallucination Mitigation in Natural Language Generation from Large-Scale Open-Domain Knowledge GraphsLong Main
- HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language ModelsLong Main
- Handshape-Aware Sign Language Recognition: Extended Datasets and Exploration of Handshape-Inclusive MethodsLong Findings
- Harnessing Black-Box Control to Boost Commonsense in LM's GenerationLong Main
- Harnessing Dataset Cartography for Improved Compositional Generalization in TransformersLong Findings
- Harnessing the Power of Large Language Models for Empathetic Response Generation: Empirical Investigations and ImprovementsLong Findings
- Harnessing the power of LLMs: Evaluating human-AI text co-creation through the lens of news headline generationLong Findings
- Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language ModelsLong Main
- HeQ: a Large and Diverse Hebrew Reading Comprehension BenchmarkLong Findings
- Hi Guys or Hi Folks? Benchmarking Gender-Neutral Machine Translation with the GeNTE CorpusLong Main
- Hi-ArG: Exploring the Integration of Hierarchical Argumentation Graphs in Language PretrainingLong Main
- Hi-ToM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language ModelsLong Findings
- HiCL: Hierarchical Contrastive Learning of Unsupervised Sentence EmbeddingsLong Findings
- HiddenTables and PyQTax: A Cooperative Game and Dataset For TableQA to Ensure Scale and Data Privacy Across a Myriad of TaxonomiesLong Main
- Hidding the Ghostwriters: An Adversarial Evaluation of AI-Generated Student Essay DetectionLong Main
- Hiding in Plain Sight: Tweets with Hate Speech Masked by HomoglyphsShort Findings
- Hierarchical Catalogue Generation for Literature Review: A BenchmarkLong Findings
- Hierarchical Enhancement Framework for Aspect-based Argument MiningLong Findings
- Hierarchical Fusion for Online Multimodal Dialog Act ClassificationLong Findings
- Hierarchical Pretraining on Multimodal Electronic Health RecordsLong Main
- Hierarchical Prompting Assists Large Language Model on Web NavigationShort Findings
- HierarchicalContrast: A Coarse-to-Fine Contrastive Learning Framework for Cross-Domain Zero-Shot Slot FillingLong Findings
- High-quality argumentative information in low resources approaches improve counter-narrative generationLong Findings
- HistAlign: Improving Context Dependency in Language Generation by Aligning with HistoryLong Main
- Holistic Inter-Annotator Agreement and Corpus Coherence Estimation in a Large-scale Multilingual Annotation CampaignLong Main
- Homophone Disambiguation Reveals Patterns of Context Mixing in Speech TransformersLong Main
- HoneyBee: Progressive Instruction Finetuning of Large Language Models for Materials ScienceLong Findings
- How Do Large Language Models Capture the Ever-changing World Knowledge? A Review of Recent AdvancesLong Main
- How Does Generative Retrieval Scale to Millions of Passages?Long Main
- How Many Demonstrations Do You Need for In-context Learning?Long Findings
- How Predictable Are Large Language Model Capabilities? A Case Study on BIG-benchLong Findings
- How Reliable Are AI-Generated-Text Detectors? An Assessment Framework Using Evasive Soft PromptsLong Findings
- How Well Do Text Embedding Models Understand Syntax?Long Findings
- How do languages influence each other? Studying cross-lingual data sharing during LM fine-tuningLong Main
- How to Determine the Most Powerful Pre-trained Language Model without Brute Force Fine-tuning? An Empirical SurveyLong Findings
- How to Enhance Causal Discrimination of Utterances: A Case on Affective ReasoningLong Main
- How to Train Your Dragon: Diverse Augmentation Towards Generalizable Dense RetrievalLong Findings
- HuatuoGPT, Towards Taming Language Model to Be a DoctorLong Findings
- Human Learning by Model Feedback: The Dynamics of Iterative Prompting with MidjourneyLong Main
- Human Raters Cannot Distinguish English Translations from Original English TextsShort Main
- HutCRS: Hierarchical User-Interest Tracking for Conversational Recommender SystemLong Main
- Hybrid Inverted Index Is a Robust Accelerator for Dense RetrievalLong Main
- HyperNetwork-based Decoupling to Improve Model Generalization for Few-Shot Relation ExtractionLong Main
- HyperRouter: Towards Efficient Training and Inference of Sparse Mixture of ExpertsShort Main
- Hyperpolyglot LLMs: Cross-Lingual Interpretability in Token EmbeddingsShort Main
- IAEval: A Comprehensive Evaluation of Instance Attribution on Natural Language UnderstandingLong Findings
- IAG: Induction-Augmented Generation Framework for Answering Reasoning QuestionsLong Main
- IBADR: an Iterative Bias-Aware Dataset Refinement Framework for Debiasing NLU modelsLong Main
- IC3: Image Captioning by Committee ConsensusLong Main
- ICU: Conquering Language Barriers in Vision-and-Language Modeling by Dividing the Tasks into Image Captioning and Language UnderstandingShort Findings
- IDTraffickers: An Authorship Attribution Dataset to link and connect Potential Human-Trafficking Operations on Text Escort AdvertisementsLong Main
- IEKG: A Commonsense Knowledge Graph for Idiomatic ExpressionsLong Main
- IMTLab: An Open-Source Platform for Building, Evaluating, and Diagnosing Interactive Machine Translation SystemsLong Main
- IMU2CLIP: Language-grounded Motion Sensor Translation with Multimodal Contrastive LearningShort Findings
- INA: An Integrative Approach for Enhancing Negotiation Strategies with Reward-Based Dialogue AgentLong Findings
- INFORM : Information eNtropy based multi-step reasoning FOR large language ModelsLong Main
- INGENIOUS: Using Informative Data Subsets for Efficient Pre-Training of Language ModelsLong Findings
- INSTRUCTSCORE: Towards Explainable Text Generation Evaluation with Automatic FeedbackLong Main
- INVITE: a Testbed of Automatically Generated Invalid Questions to Evaluate Large Language Models for HallucinationsShort Findings
- INarIG: Iterative Non-autoregressive Instruct Generation Model For Word-Level Auto CompletionLong Findings
- IRFL: Image Recognition of Figurative LanguageLong Findings
- IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language ModelsLong Findings
- Identification of Multimodal Stance Towards Frames of CommunicationLong Main
- Identifying Conspiracy Theories News based on Event Relation GraphLong Findings
- Identifying Informational Sources in News ArticlesLong Main
- Identifying Statements Crucial for Awareness of Interpretive Nonsense to Prevent Communication BreakdownsLong Main
- Identifying {Early Maladaptive Schemas} from Mental Health Question TextsShort Findings
- Ideology Takes Multiple Looks: A High-Quality Dataset for Multifaceted Ideology DetectionLong Main
- IfQA: A Dataset for Open-domain Question Answering under Counterfactual PresuppositionsLong Main
- Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs Through a Global Prompt Hacking CompetitionLong Main
- Image Manipulation via Multi-Hop Instructions - A New Dataset and Weakly-Supervised Neuro-Symbolic ApproachLong Main
- Image and Text: Fighting the same Battle? Super Resolution Learning for Imbalanced Text ClassificationLong Findings
- ImageNetVC: Zero- and Few-Shot Visual Commonsense Evaluation on 1000 ImageNet CategoriesLong Findings
- Impact of Co-occurrence on Factual Knowledge of Large Language ModelsLong Findings
- Implicit Sense-labeled Connective Recognition as Text GenerationShort Findings
- Impressions: Visual Semiotics and Aesthetic Impact UnderstandingLong Main
- Improved Pseudo Data for Machine Translation Quality Estimation with Constrained Beam SearchLong Main
- Improved Training of Deep Text ClusteringShort Findings
- Improved Unsupervised Chinese Word Segmentation Using Pre-trained Knowledge and Pseudo-labeling TransferShort Main
- Improving Bias Mitigation through Bias Experts in Natural Language UnderstandingLong Main
- Improving Biomedical Abstractive Summarisation with Knowledge Aggregation from Citation PapersLong Main
- Improving Chinese Pop Song and Hokkien Gezi Opera Singing Voice Synthesis by Enhancing Local ModelingLong Main
- Improving Consistency for Text Summarization with Energy FunctionsShort Findings
- Improving Contrastive Learning of Sentence Embeddings with Focal InfoNCEShort Findings
- Improving Conversational Recommendation Systems via Bias Analysis and Language-Model-Enhanced Data AugmentationLong Findings
- Improving Cross-lingual Transfer through Subtree-aware Word ReorderingLong Findings
- Improving Dialogue Discourse Parsing via Reply-to Structures of Addressee RecognitionLong Main
- Improving Diversity of Demographic Representation in Large Language Models via Collective-Critiques and Self-VotingLong Main
- Improving End-to-End Speech Processing by Efficient Text Data Utilization with Latent SynthesisLong Findings
- Improving Factual Consistency for Knowledge-Grounded Dialogue Systems via Knowledge Enhancement and AlignmentLong Findings
- Improving Image Captioning via Predicting Structured ConceptsLong Main
- Improving Input-label Mapping with Demonstration Replay for In-context LearningLong Findings
- Improving Language Models’ Meaning Understanding and Consistency by Learning Conceptual Roles from DictionaryLong Main
- Improving Long Document Topic Segmentation Models With Enhanced Coherence ModelingLong Main
- Improving Low-resource Question Answering by Augmenting Question InformationShort Findings
- Improving Multi-Criteria Chinese Word Segmentation through Learning Sentence RepresentationShort Findings
- Improving Multimodal Sentiment Analysis: Supervised Angular margin-based Contrastive Learning for Enhanced Fusion RepresentationLong Findings
- Improving Neural Machine Translation by Multi-Knowledge Integration with PromptingLong Findings
- Improving Pacing in Long-Form Story PlanningShort Findings
- Improving Question Generation with Multi-level Content PlanningLong Findings
- Improving Seq2Seq Grammatical Error Correction via Decoding InterventionsLong Findings
- Improving Sequential Model Editing with Fact RetrievalLong Findings
- Improving Span Representation by Efficient Span-Level AttentionShort Findings
- Improving Speech Translation by Fusing Speech and TextLong Findings
- Improving Summarization with Human EditsLong Main
- Improving Transformer-based Program Repair Model through False Behavior DiagnosisLong Main
- Improving Unsupervised Relation Extraction by Augmenting Diverse Sentence PairsLong Main
- Improving Zero-shot Reader by Reducing Distractions from Irrelevant Documents in Open-Domain Question AnsweringShort Findings
- Improving generalization in large langue model by learning prefix subspacesLong Findings
- Improving the Robustness of Summarization Models by Detecting and Removing Input NoiseLong Findings
- Improving word mover's distance by leveraging self-attention matrixLong Findings
- In What Languages are Generative Language Models the Most Formal? Analyzing Formality Distribution across LanguagesLong Findings
- In-Context Demonstration Selection with Cross Entropy DifferenceLong Findings
- In-Context Learning Creates Task VectorsShort Findings
- In-Image Neural Machine Translation with Segmented Pixel Sequence-to-Sequence ModelLong Findings
- In-context Learning for Few-shot Multimodal Named Entity RecognitionLong Findings
- Incorporating Object-Level Visual Context for Multimodal Fine-Grained Entity TypingLong Findings
- Incorporating Probing Signals into Multimodal Machine Translation via Visual Question-Answering PairsLong Findings
- Incorporating Structured Representations into Pretrained Vision \& Language Models Using Scene GraphsLong Main
- Incorporating Syntactic Knowledge into Pre-trained Language Model using Optimization for Overcoming Catastrophic ForgettingLong Findings
- Incorporating Worker Perspectives into MTurk Annotation Practices for NLPLong Main
- Increasing Coverage and Precision of Textual Information in Multilingual Knowledge GraphsLong Main
- Increasing Probability Mass on Answer Choices Does Not Always Improve AccuracyLong Main
- IndiSocialFT: Multilingual Word Representation for Indian languages in code-mixed environmentShort Findings
- Indicative Summarization of Long DiscussionsLong Main
- Inductive Relation Inference of Knowledge Graph Enhanced by Ontology InformationLong Findings
- Inference-Time Policy Adapters (IPA): Tailoring Extreme-Scale LMs without Fine-tuningLong Main
- Influence Scores at Scale for Efficient Language Data SamplingLong Main
- InfoCL: Alleviating Catastrophic Forgetting in Continual Text Classification from An Information Theoretic PerspectiveLong Findings
- InfoDiffusion: Information Entropy Aware Diffusion Process for Non-Autoregressive Text GenerationLong Findings
- Information Extraction from Legal Wills: How Well Does GPT-4 Do?Short Findings
- Information Value: Measuring Utterance Predictability as Distance from Plausible AlternativesLong Main
- InheritSumm: A General, Versatile and Compact Summarizer by Distilling from GPTLong Findings
- Injecting structural hints: Using language models to study inductive biases in language learningLong Findings
- InstOptima: Evolutionary Multi-objective Instruction Optimization via Large Language Model-based Instruction OperatorsShort Findings
- Instances and Labels: Hierarchy-aware Joint Supervised Contrastive Learning for Hierarchical Multi-Label Text ClassificationLong Findings
- Instruct and Extract: Instruction Tuning for On-Demand Information ExtractionLong Main
- InstructExcel: A Benchmark for Natural Language Instruction in ExcelLong Findings
- InstructSafety: A Unified Framework for Building Multidimensional and Explainable Safety Detector through Instruction TuningLong Findings
- Instructed Language Models with Retrievers Are Powerful Entity LinkersLong Main
- Instructive Dialogue Summarization with Query AggregationsLong Main
- InstructoR: Instructing Unsupervised Conversational Dense Retrieval with Large Language ModelsLong Findings
- InteMATs: Integrating Granularity-Specific Multilingual Adapters for Cross-Lingual TransferLong Findings
- Integrating Language Models into Direct Speech Translation: An Inference-Time Solution to Control Gender InflectionShort Main
- IntenDD: A Unified Contrastive Learning Approach for Intent Detection and DiscoveryLong Findings
- InterFair: Debiasing with Natural Language Feedback for Fair Interpretable PredictionsShort Main
- Interactive Text GenerationLong Main
- Interpreting Answers to Yes-No Questions in User-Generated ContentLong Findings
- Interpreting Embedding Spaces by ConceptualizationLong Main
- Interpreting Indirect Answers to Yes-No Questions in Multiple LanguagesLong Findings
- InterroLang: Exploring NLP Models and Datasets through Dialogue-based ExplanationsLong Findings
- Intersectional Stereotypes in Large Language Models: Dataset and AnalysisShort Findings
- Intervention-Based Alignment of Code Search with Execution FeedbackLong Findings
- Interventional RationalizationLong Main
- Interview Evaluation: A Novel Approach for Automatic Evaluation of Conversational Question Answering ModelsLong Main
EMNLP accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.