← All conferences

EMNLP 2023 Accepted Papers

The full list of 2,009 papers accepted at EMNLP 2023 (Conference on Empirical Methods in Natural Language Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

Long Main: 862Long Findings: 818Short Findings: 197Short Main: 132
  1. "A Tale of Two Movements": Identifying and Comparing Perspectives in \#BlackLivesMatter and \#BlueLivesMatter Movements-related Tweets using Weakly Supervised Graph-based Structured PredictionLong Findings
  2. "Are Your Explanations Reliable?" Investigating the Stability of LIME in Explaining Text Classifiers by Marrying XAI and Adversarial AttackLong Main
  3. "Fifty Shades of Bias": Normative Ratings of Gender Bias in GPT Generated English TextLong Main
  4. "You Are An Expert Linguistic Annotator": Limits of LLMs as Analyzers of Abstract Meaning RepresentationShort Findings
  5. $\textbf{\emph{CLMSM}}$: A Multi-Task Learning Framework for Pre-training on Procedural TextLong Findings
  6. $\textit{From Chaos to Clarity}$: Claim Normalization to Empower Fact-CheckingLong Findings
  7. $\textit{Lost in Translation, Found in Spans}$: Identifying Claims in Multilingual Social MediaLong Main
  8. $\textit{SelectNoise:}$ Unsupervised Noise Injection to Enable Zero-Shot Machine Translation for Extremely Low-resource LanguagesLong Findings
  9. $\textit{Swap and Predict}$ -- Predicting the Semantic Changes in Words across Corpora by Context SwappingLong Findings
  10. $\textit{``Don't Take This Out of Context!''}$ On the Need for Contextual Models and Evaluations for Stylistic RewritingLong Main
  11. $k$NN-LM Does Not Improve Open-ended Text GenerationLong Main
  12. 'Person' == Light-skinned, Western Man, and Sexualization of Women of Color: Stereotypes in Stable DiffusionLong Findings
  13. 1-PAGER: One Pass Answer Generation and Evidence RetrievalLong Findings
  14. 2INER: Instructive and In-Context Learning on Few-Shot Named Entity RecognitionLong Findings
  15. 3DRP-Net: 3D Relative Position-aware Network for 3D Visual GroundingLong Main
  16. 4 and 7-bit Labeling for Projective and Non-Projective Dependency TreesShort Main
  17. A Benchmark for Semi-Inductive Link Prediction in Knowledge GraphsShort Findings
  18. A Black-Box Attack on Code Models via Representation Nearest Neighbor SearchLong Findings
  19. A Boundary Offset Prediction Network for Named Entity RecognitionLong Findings
  20. A Causal View of Entity Bias in (Large) Language ModelsLong Findings
  21. A Challenging Multimodal Video Summary: Simultaneously Extracting and Generating Keyframe-Caption Pairs from VideoLong Main
  22. A Cheaper and Better Diffusion Language Model with Soft-Masked NoiseLong Main
  23. A Closer Look into Using Large Language Models for Automatic EvaluationShort Findings
  24. A Comprehensive Evaluation of Biomedical Entity Linking ModelsLong Main
  25. A Comprehensive Evaluation of Large Language Models on Legal Judgment PredictionLong Findings
  26. A Comprehensive Evaluation of Tool-Assisted Generation StrategiesLong Findings
  27. A Computational Interface to Translate Strategic Intent from Unstructured Language in a Low-Data SettingLong Findings
  28. A Confederacy of Models: a Comprehensive Evaluation of LLMs on Creative WritingLong Findings
  29. A Critical Analysis of Document Out-of-Distribution DetectionLong Findings
  30. A Dataset for Investigating the Impact of Context for Offensive Language Detection in TweetsShort Findings
  31. A Deeper (Autoregressive) Approach to Non-Convergent Discourse ParsingLong Main
  32. A Diachronic Analysis of Paradigm Shifts in NLP Research: When, How, and Why?Long Main
  33. A Diachronic Perspective on User Trust in AI under UncertaintyLong Main
  34. A Diffusion Weighted Graph Framework for New Intent DiscoveryLong Main
  35. A Fair and In-Depth Evaluation of Existing End-to-End Entity Linking SystemsLong Main
  36. A Fine-Grained Taxonomy of Replies to Hate SpeechLong Main
  37. A Framework for Bidirectional Decoding: Case Study in Morphological InflectionLong Findings
  38. A Framework for Exploring Player Perceptions of LLM-Generated Dialogue in Commercial Video GamesLong Findings
  39. A Framework for Vision-Language Warm-up Tasks in Multimodal Dialogue ModelsLong Main
  40. A Frustratingly Easy Plug-and-Play Detection-and-Reasoning Module for Chinese Spelling CheckLong Findings
  41. A Frustratingly Easy Post-Training Quantization Scheme for LLMsLong Main
  42. A Generation-based Deductive Method for Math Word ProblemsLong Main
  43. A Hierarchical Encoding-Decoding Scheme for Abstractive Multi-document SummarizationLong Findings
  44. A Joint Matrix Factorization Analysis of Multilingual RepresentationsLong Findings
  45. A Language Model with Limited Memory Capacity Captures Interference in Human Sentence ProcessingLong Findings
  46. A Lightweight Method to Generate Unanswerable Questions in EnglishShort Findings
  47. A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation AnalysisLong Main
  48. A Multi-Modal Multilingual Benchmark for Document Image ClassificationLong Findings
  49. A Multi-Task Dataset for Assessing Discourse Coherence in Chinese Essays: Structure, Theme, and Logic AnalysisLong Main
  50. A New Benchmark and Reverse Validation Method for Passage-level Hallucination DetectionLong Findings
  51. A Novel Contrastive Learning Method for Clickbait Detection on RoCliCo: A Romanian Clickbait Corpus of News ArticlesShort Findings
  52. A Parallel Corpus for Vietnamese Central-Northern Dialect Text TransferLong Findings
  53. A Predictive Factor Analysis of Social Biases and Task-Performance in Pretrained Masked Language ModelsLong Main
  54. A Quality-based Syntactic Template Retriever for Syntactically-Controlled Paraphrase GenerationLong Main
  55. A Query-Parallel Machine Reading Comprehension Framework for Low-resource NERLong Findings
  56. A Question Answering Framework for Decontextualizing User-facing Snippets from Scientific DocumentsLong Main
  57. A Read-and-Select Framework for Zero-shot Entity LinkingLong Findings
  58. A Reference-free Segmentation Quality Index (SegReFree)Long Findings
  59. A Rewriting Approach for Gender Inclusivity in PortugueseLong Findings
  60. A Rose by Any Other Name would not Smell as Sweet: Social Bias in Names MistranslationLong Main
  61. A Scalable Framework for Table of Contents Extraction from Complex ESG Annual ReportsLong Main
  62. A Self-enhancement Multitask Framework for Unsupervised Aspect Category DetectionLong Main
  63. A Sequence-to-Structure Approach to Document-level Targeted Sentiment AnalysisLong Findings
  64. A Simple Baseline for Knowledge-Based Visual Question AnsweringShort Main
  65. A Spectral Viewpoint on Continual Relation ExtractionShort Findings
  66. A State-Vector Framework for Dataset EffectsLong Main
  67. A Structure-Aware Generative Adversarial Network for Bilingual Lexicon InductionLong Findings
  68. A Study on Accessing Linguistic Information in Pre-Trained Language Models by Using PromptsShort Main
  69. A Suite of Generative Tasks for Multi-Level Multimodal Webpage UnderstandingLong Main
  70. A Systematic Study of Performance Disparities in Multilingual Task-Oriented Dialogue SystemsLong Main
  71. A Tale of Pronouns: Interpretability Informs Gender Bias Mitigation for Fairer Instruction-Tuned Machine TranslationLong Main
  72. A Thorough Examination on Zero-shot Dense RetrievalLong Findings
  73. A Training-Free Debiasing Framework with Counterfactual Reasoning for Conversational Emotion DetectionLong Main
  74. A Unified Framework for Synaesthesia AnalysisLong Findings
  75. A Unified View of Evaluation Metrics for Structured PredictionLong Main
  76. A Video Is Worth 4096 Tokens: Verbalize Story Videos To Understand Them In Zero ShotLong Main
  77. A Word Sense Distribution-based approach for Semantic Change PredictionLong Findings
  78. A Zero-Shot Language Agent for Computer Control with Structured ReflectionLong Findings
  79. A linear time approximation of Wasserstein distance with word embedding selectionLong Main
  80. ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life VideosLong Main
  81. ACT-SQL: In-Context Learning for Text-to-SQL with Automatically-Generated Chain-of-ThoughtLong Findings
  82. ACTOR: Active Learning with Annotator-specific Classification Heads to Embrace Human Label VariationShort Main
  83. AD-NLP: A Benchmark for Anomaly Detection in Natural Language ProcessingLong Main
  84. ALCUNA: Large Language Models Meet New KnowledgeLong Main
  85. ALDi: Quantifying the Arabic Level of Dialectness of TextLong Main
  86. AMR Parsing is Far from Solved: GrAPES, the Granular AMR Parsing Evaluation SuiteLong Main
  87. AMR Parsing with Causal Hierarchical Attention and PointersLong Main
  88. API-Assisted Code Generation for Question Answering on Varied Table StructuresLong Main
  89. API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMsLong Main
  90. APP: Adaptive Prototypical Pseudo-Labeling for Few-shot OOD DetectionLong Findings
  91. APoLLo : Unified Adapter and Prompt Learning for Vision Language ModelsLong Main
  92. APrompt: Attention Prompt Tuning for Efficient Adaptation of Pre-trained Language ModelsLong Main
  93. ARKitSceneRefer: Text-based Localization of Small Objects in Diverse Real-World 3D Indoor ScenesLong Findings
  94. ART: rule bAsed futuRe-inference deducTionLong Main
  95. ASPIRO: Any-shot Structured Parsing-error-Induced ReprOmpting for Consistent Data-to-Text GenerationShort Findings
  96. ASSERT: Automated Safety Scenario Red Teaming for Evaluating the Robustness of Large Language ModelsLong Findings
  97. ATFormer: A Learned Performance Model with Transfer Learning Across Devices for Deep Learning Tensor ProgramsLong Main
  98. Absolute Position Embedding Learns Sinusoid-like Waves for Attention Based on Relative PositionLong Main
  99. Abstractive Open Information ExtractionLong Main
  100. Accelerating Multiple Intent Detection and Slot Filling via Targeted Knowledge DistillationLong Findings
  101. Accelerating Toeplitz Neural Network with Constant-time Inference ComplexityLong Main
  102. Accented Speech Recognition With Accent-specific CodebooksLong Main
  103. Accuracy is not enough: Evaluating Personalization in SummarizersLong Findings
  104. Active Instruction Tuning: Improving Cross-Task Generalization by Training on Prompt Sensitive TasksLong Main
  105. Active Learning Principles for In-Context Learning with Large Language ModelsLong Findings
  106. Active Learning for Natural Language GenerationLong Main
  107. Active Retrieval Augmented GenerationLong Main
  108. AdaSent: Efficient Domain-Adapted Sentence Embeddings for Few-Shot ClassificationLong Main
  109. AdaTranS: Adapting with Boundary-based Shrinking for End-to-End Speech TranslationShort Findings
  110. Adapt in Contexts: Retrieval-Augmented Domain Adaptation via In-Context LearningLong Main
  111. Adaptation with Self-Evaluation to Improve Selective Prediction in LLMsLong Findings
  112. Adapter Pruning using Tropical CharacterizationShort Findings
  113. Adapter-TST: A Parameter Efficient Method for Multiple-Attribute Text Style TransferLong Findings
  114. Adapting Language Models to Compress ContextsLong Main
  115. Adapting Pretrained Text-to-Text Models for Long Text SequencesLong Findings
  116. Adaptive End-to-End Metric Learning for Zero-Shot Cross-Domain Slot FillingLong Main
  117. Adaptive Gating in Mixture-of-Experts based Language ModelsLong Main
  118. Adaptive Hinge Balance Loss for Document-Level Relation ExtractionShort Findings
  119. Adaptive Policy with Wait-k Model for Simultaneous TranslationLong Main
  120. Adaptive Structure Induction for Aspect-based Sentiment Analysis with Spectral PerspectiveLong Findings
  121. Adaptive Textual Label Noise Learning based on Pre-trained ModelsLong Findings
  122. Addressing Linguistic Bias through a Contrastive Analysis of Academic Writing in the NLP DomainLong Main
  123. Addressing NER Annotation Noises with Uncertainty-Guided Tree-Structured CRFsLong Main
  124. Addressing the Length Bias Challenge in Document-Level Neural Machine TranslationLong Findings
  125. Advancements in Arabic Grammatical Error Detection and Correction: An Empirical InvestigationLong Main
  126. Adversarial Robustness for Large Language NER models using Disentanglement and Word AttributionsLong Findings
  127. Adversarial Text Generation by Search and LearningLong Findings
  128. Affective and Dynamic Beam Search for Story GenerationLong Findings
  129. AfriSenti: A Twitter Sentiment Analysis Benchmark for African LanguagesLong Main
  130. Air-Decoding: Attribute Distribution Reconstruction for Decoding-Time Controllable Text GenerationLong Main
  131. Aligning Language Models to User OpinionsLong Findings
  132. Aligning Large Language Models through Synthetic FeedbackLong Main
  133. Aligning Predictive Uncertainty with Clarification Questions in Grounded DialogLong Findings
  134. Alignment Precedes Fusion: Open-Vocabulary Named Entity Recognition as Context-Type Semantic MatchingLong Findings
  135. All Things Considered: Detecting Partisan Events from News Media with Cross-Article ComparisonLong Main
  136. Allies: Prompting Large Language Model with Beam SearchLong Findings
  137. Always the Best Fit: Adaptive Domain Gap Filling from Causal Perspective for Few-Shot Relation ExtractionShort Findings
  138. An Adaptive Prompt Generation Framework for Task-oriented Dialogue SystemLong Findings
  139. An Attribution Method for Siamese EncodersShort Main
  140. An Empirical Investigation of Implicit and Explicit Knowledge-Enhanced Methods for Ad Hoc Dataset RetrievalLong Findings
  141. An Empirical Study of Frame Selection for Text-to-Video RetrievalLong Findings
  142. An Empirical Study of Instruction-tuning Large Language Models in ChineseLong Findings
  143. An Empirical Study of Multimodal Model MergingLong Findings
  144. An Empirical Study of Translation Hypothesis Ensembling with Large Language ModelsLong Main
  145. An Empirical Study on Multiple Knowledge from ChatGPT for Emotion Recognition in ConversationsLong Findings
  146. An Exploration of Left-Corner TransformationsLong Main
  147. An Expression Tree Decoding Strategy for Mathematical Equation GenerationLong Main
  148. An Integrative Survey on Mental Health Conversational Agents to Bridge Computer Science and Medical PerspectivesLong Main
  149. An Intent-based and Annotation-free Method for Duplicate Question Detection in CQA ForumsLong Findings
  150. An Investigation of LLMs’ Inefficacy in Understanding Converse RelationsLong Main
  151. An Iteratively Parallel Generation Method with the Pre-Filling Strategy for Document-level Event ExtractionLong Main
  152. Analysing State-Backed Propaganda Websites: a New Dataset and Linguistic StudyShort Main
  153. Analysis of Style-Shifting on Social Media: Using Neural Language Model Conditioned by Social MeaningsLong Findings
  154. Analyzing Cognitive Plausibility of Subword TokenizationShort Main
  155. Analyzing Film Adaptation through Narrative AlignmentLong Main
  156. Analyzing Modular Approaches for Visual Question DecompositionLong Main
  157. Analyzing Norm Violations in Live-Stream ChatLong Main
  158. Anaphor Assisted Document-Level Relation ExtractionLong Main
  159. Anchoring Fine-tuning of Sentence Transformer with Semantic Label Information for Efficient Truly Few-shot ClassificationShort Main
  160. AniEE: A Dataset of Animal Experimental Literature for Event ExtractionLong Findings
  161. Annotation Sensitivity: Training Data Collection Methods Affect Model PerformanceLong Findings
  162. Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence GroundingLong Findings
  163. Answer-state Recurrent Relational Network (AsRRN) for Constructed Response Assessment and Feedback GroupingLong Findings
  164. Answering Questions by Meta-Reasoning over Multiple Chains of ThoughtLong Main
  165. Appraising the Potential Uses and Harms of LLMs for Medical Systematic ReviewsLong Main
  166. Approximating CKY with TransformersLong Findings
  167. Approximating Two-Layer Feedforward Networks for Efficient TransformersLong Findings
  168. Arabic Mini-ClimateGPT : A Climate Change and Sustainability Tailored Arabic LLMShort Findings
  169. Architectural Sweet Spots for Modeling Human Label Variation by the Example of Argument Quality: It’s Best to Relate Perspectives!Long Main
  170. Are All Steps Equally Important? Benchmarking Essentiality Detection in Event ProcessesShort Main
  171. Are Compressed Language Models Less Subgroup Robust?Short Main
  172. Are Embedded Potatoes Still Vegetables? On the Limitations of WordNet Embeddings for Lexical SemanticsLong Main
  173. Are Language Models Worse than Humans at Following Prompts? It's ComplicatedShort Findings
  174. Are NLP Models Good at Tracing Thoughts: An Overview of Narrative UnderstandingLong Findings
  175. Are Personalized Stochastic Parrots More Dangerous? Evaluating Persona Biases in Dialogue SystemsLong Findings
  176. Are Structural Concepts Universal in Transformer Language Models? Towards Interpretable Cross-Lingual GeneralizationLong Findings
  177. Argue with Me Tersely: Towards Sentence-Level Counter-Argument GenerationLong Main
  178. Argument mining as a multi-hop generative machine reading comprehension taskLong Findings
  179. Argument-based Detection and Classification of Fallacies in Political DebatesLong Main
  180. Ask Language Model to Clean Your Noisy Translation DataLong Findings
  181. Ask To The Point: Open-Domain Entity-Centric Question GenerationLong Findings
  182. Asking Clarification Questions to Handle Ambiguity in Open-Domain QALong Findings
  183. Aspect-Category Enhanced Learning with a Neural Coherence Model for Implicit Sentiment AnalysisLong Findings
  184. Aspect-to-Scope Oriented Multi-view Contrastive Learning for Aspect-based Sentiment AnalysisLong Findings
  185. Assessing Privacy Risks in Language Models: A Case Study on Summarization TasksLong Findings
  186. Assessing Step-by-Step Reasoning against Lexical Negation: A Case Study on SyllogismShort Main
  187. Attack Prompt Generation for Red Teaming and Defending Large Language ModelsLong Findings
  188. Attention-Enhancing Backdoor Attacks Against BERT-based ModelsLong Findings
  189. Augmenting Zero-Shot Dense Retrievers with Plug-in Mixture-of-MemoriesLong Main
  190. Auto Search Indexer for End-to-End Document RetrievalLong Findings
  191. Auto-Instruct: Automatic Instruction Generation and Ranking for Black-Box Language ModelsLong Findings
  192. AutoPlan: Automatic Planning of Interactive Decision-Making Tasks With Large Language ModelsLong Findings
  193. AutoTrial: Prompting Language Models for Clinical Trial DesignLong Main
  194. Automated Few-Shot Classification with Instruction-Finetuned Language ModelsLong Findings
  195. Automatic Analysis of Substantiation in Scientific Peer ReviewsLong Findings
  196. Automatic Debate Evaluation with Argumentation Semantics and Natural Language Argument Graph NetworksLong Main
  197. Automatic Evaluate Dialogue Appropriateness by Using Dialogue ActLong Findings
  198. Automatic Evaluation of Attribution by Large Language ModelsLong Findings
  199. Automatic Model Selection with Large Language Models for ReasoningLong Findings
  200. Automatic Prompt Augmentation and Selection with Chain-of-Thought from Labeled DataLong Findings
  201. Automatic Prompt Optimization with "Gradient Descent" and Beam SearchLong Main
  202. Automatic Pronunciation Assessment - A ReviewLong Findings
  203. Automatic Transcription of Handwritten Old Occitan LanguageLong Main
  204. Axiomatic Preference Modeling for Longform Question AnsweringLong Main
  205. BERT Has More to Offer: BERT Layers Combination Yields Better Sentence EmbeddingsShort Findings
  206. BERTie Bott's Every Flavor Labels: A Tasty Introduction to Semantic Role Labeling for GalicianLong Main
  207. BERTwich: Extending BERT’s Capabilities to Model Dialectal and Noisy TextLong Findings
  208. BLESS: Benchmarking Large Language Models on Sentence SimplificationLong Main
  209. BLM-s/lE: A structured dataset of English spray-load verb alternations for testing generalization in LLMsLong Findings
  210. BRAINTEASER: Lateral Thinking Puzzles for Large Language ModelsLong Main
  211. BYOC: Personalized Few-Shot Classification with Co-Authored Class DescriptionsLong Findings
  212. Back Transcription as a Method for Evaluating Robustness of Natural Language Understanding Models to Speech Recognition ErrorsLong Main
  213. Background Summarization of Event TimelinesLong Main
  214. Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat DataLong Main
  215. Balance Act: Mitigating Hubness in Cross-Modal Retrieval with Query and Gallery BanksLong Main
  216. Balaur: Language Model Pretraining with Lexical Semantic RelationsLong Findings
  217. BanLemma: A Word Formation Dependent Rule and Dictionary Based Bangla LemmatizerLong Findings
  218. BanglaAbuseMeme: A Dataset for Bengali Abusive Meme ClassificationLong Main
  219. BasahaCorpus: An Expanded Linguistic Resource for Readability Assessment in Central Philippine LanguagesShort Main
  220. Battle of the Large Language Models: Dolly vs LLaMA vs Vicuna vs Guanaco vs Bard vs ChatGPT - A Text-to-SQL Parsing ComparisonLong Findings
  221. Bayesian Multi-Task Transfer Learning for Soft Prompt TuningLong Findings
  222. Be Selfish, But Wisely: Investigating the Impact of Agent Personality in Mixed-Motive Human-Agent InteractionsLong Main
  223. Beat LLMs at Their Own Game: Zero-Shot LLM-Generated Text Detection via Querying ChatGPTShort Main
  224. Benchmarking and Improving Text-to-SQL Generation under AmbiguityLong Main
  225. Beneath Surface Similarity: Large Language Models Make Reasonable Scientific Analogies after Structure AbductionLong Findings
  226. Beneath the Surface: Unveiling Harmful Memes with Multimodal Reasoning Distilled from Large Language ModelsLong Findings
  227. Best of Both Worlds: Towards Improving Temporal Knowledge Base Question Answering via Targeted Fact ExtractionShort Main
  228. Better Quality Pre-training Data and T5 Models for African LanguagesShort Main
  229. Better Together: Enhancing Generative Knowledge Graph Completion with Language Models and Neighborhood InformationShort Findings
  230. Beware of Model Collapse! Fast and Stable Test-time Adaptation for Robust Question AnsweringLong Main
  231. Beyond Candidates : Adaptive Dialogue Agent Utilizing Persona and KnowledgeLong Findings
  232. Beyond Denouncing Hate: Strategies for Countering Implied Biases and Stereotypes in LanguageLong Findings
  233. Beyond Detection: A Defend-and-Summarize Strategy for Robust and Interpretable Rumor Analysis on Social MediaLong Main
  234. Beyond Factuality: A Comprehensive Evaluation of Large Language Models as Knowledge GeneratorsLong Main
  235. Beyond Labels: Empowering Human Annotators with Natural Language Explanations through a Novel Active-Learning ArchitectureLong Findings
  236. Beyond Layout Embedding: Layout Attention with Gaussian Biases for Structured Document UnderstandingLong Findings
  237. Beyond Shared Vocabulary: Increasing Representational Word Similarities across Languages for Multilingual Machine TranslationLong Main
  238. Beyond Testers’ Biases: Guiding Model Testing with Knowledge Bases using LLMsLong Findings
  239. Bi-Drop: Enhancing Fine-tuning Generalization via Synchronous sub-net Estimation and OptimizationLong Findings
  240. BiSPN: Generating Entity Set and Relation Set Coherently in One PassLong Findings
  241. Bias Neutralization in Non-Parallel Texts: A Cyclic Approach with Auxiliary GuidanceLong Main
  242. BiasX: “Thinking Slow” in Toxic Content Moderation with Explanations of Implied Social BiasesShort Main
  243. BioDEX: Large-Scale Biomedical Adverse Drug Event Extraction for Real-World PharmacovigilanceLong Findings
  244. BioFEG: Generate Latent Features for Biomedical Entity LinkingLong Main
  245. BioPlanner: Automatic Evaluation of LLMs on Protocol Planning in BiologyLong Main
  246. BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language AssociationsLong Main
  247. Biomedical Named Entity Recognition via Dictionary-based Synonym GeneralizationLong Main
  248. Bipartite Graph Pre-training for Unsupervised Extractive Summarization with Graph Convolutional Auto-EncodersLong Findings
  249. Black-Box Tuning of Vision-Language Models with Effective Gradient ApproximationLong Findings
  250. Blackbird language matrices (BLM), a new task for rule-like generalization in neural networks: Can Large Language Models pass the test?Long Findings
  251. Boosting Inference Efficiency: Unleashing the Power of Parameter-Shared Pre-trained Language ModelsLong Findings
  252. Boosting Prompt-Based Self-Training With Mapping-Free Automatic Verbalizer for Multi-Class ClassificationLong Findings
  253. Boosting Summarization with Normalizing Flows and Aggressive TrainingLong Main
  254. Boot and Switch: Alternating Distillation for Zero-Shot Dense RetrievalLong Findings
  255. Bootstrapping Small \& High Performance Language Models with Unmasking-Removal Training PolicyShort Main
  256. BotPercent: Estimating Bot Populations in Twitter CommunitiesLong Findings
  257. Breaking Boundaries in Retrieval Systems: Unsupervised Domain Adaptation with Denoise-FinetuningLong Findings
  258. Breaking the Language Barrier: Improving Cross-Lingual Reasoning with Structured Self-AttentionLong Findings
  259. Breaking through Deterministic Barriers: Randomized Pruning Mask Generation and SelectionLong Findings
  260. Bridging Background Knowledge Gaps in Translation with Automatic ExplicitationLong Main
  261. Bridging Continuous and Discrete Spaces: Interpretable Sentence Representation Learning via Compositional OperationsLong Main
  262. Bridging Information-Theoretic and Geometric Compression in Language ModelsLong Main
  263. Bridging the Digital Divide: Performance Variation across Socio-Economic Factors in Vision-Language ModelsLong Main
  264. Bridging the Gap between Synthetic and Authentic Images for Multimodal Machine TranslationLong Main
  265. Building Multi-domain Dialog State Trackers from Single-domain DialogsLong Main
  266. Building Persona Consistent Dialogue Agents with Offline Reinforcement LearningLong Main
  267. Byte Pair Encoding for Symbolic MusicLong Main
  268. ByteSized32: A Corpus and Challenge Task for Generating Task-Specific World Models Expressed as Text GamesLong Main
  269. C-STS: Conditional Semantic Textual SimilarityLong Main
  270. C2D2 Dataset: A Resource for the Cognitive Distortion Analysis and Its Impact on Mental HealthLong Findings
  271. CAPSTONE: Curriculum Sampling for Dense Retrieval with Document ExpansionLong Main
  272. CAR: Conceptualization-Augmented Reasoner for Zero-Shot Commonsense Question AnsweringLong Findings
  273. CASE: Commonsense-Augmented Score with an Expanded Answer SpaceLong Findings
  274. CASSI: Contextual and Semantic Structure-based Interpolation Augmentation for Low-Resource NERLong Findings
  275. CCEval: A Representative Evaluation Benchmark for the Chinese-centric Multilingual Machine TranslationShort Findings
  276. CCIM: Cross-modal Cross-lingual Interactive Image TranslationShort Findings
  277. CCSRD: Content-Centric Speech Representation Disentanglement Learning for End-to-End Speech TranslationLong Findings
  278. CESAR: Automatic Induction of Compositional Instructions for Multi-turn DialogsLong Main
  279. CHEF in the Language Kitchen: A Generative Data Augmentation Leveraging Korean Morpheme IngredientsLong Main
  280. CHiLL: Zero-shot Custom Interpretable Feature Extraction from Clinical Notes with Large Language ModelsLong Findings
  281. CITB: A Benchmark for Continual Instruction TuningLong Findings
  282. CLAD-ST: Contrastive Learning with Adversarial Data for Robust Speech TranslationShort Main
  283. CLAIR: Evaluating Image Captions with Large Language ModelsShort Main
  284. CLASS: A Design Framework for Building Intelligent Tutoring Systems Based on Learning Science principlesLong Findings
  285. CLEME: Debiasing Multi-reference Evaluation for Grammatical Error CorrectionLong Main
  286. CLEVR-Implicit: A Diagnostic Dataset for Implicit Reasoning in Referring Expression ComprehensionLong Main
  287. COFFEE: Counterfactual Fairness for Personalized Text Generation in Explainable RecommendationLong Main
  288. COHESENTIA: A Novel Benchmark of Incremental versus Holistic Assessment of Coherence in Generated TextsLong Main
  289. COMET-M: Reasoning about Multiple Events in Complex SentencesLong Findings
  290. CONTRASTE: Supervised Contrastive Pre-training With Aspect-based Prompts For Aspect Sentiment Triplet ExtractionLong Findings
  291. CORE: A Few-Shot Company Relation Classification Dataset for Robust Domain Adaptation.Long Main
  292. COUNT: COntrastive UNlikelihood Text Style Transfer for Text DetoxificationShort Findings
  293. COVID-19 Vaccine Misinformation in Middle Income CountriesLong Main
  294. CP-BCS: Binary Code Summarization Guided by Control Flow Graph and Pseudo CodeLong Main
  295. CQE: A Comprehensive Quantity ExtractorLong Main
  296. CRAB: Assessing the Strength of Causal Relationships Between Real-world EventsLong Main
  297. CREATOR: Tool Creation for Disentangling Abstract and Concrete Reasoning of Large Language ModelsLong Findings
  298. CRT-QA: A Dataset of Complex Reasoning Question Answering over Tabular DataLong Main
  299. CRUSH4SQL: Collective Retrieval Using Schema Hallucination For Text2SQLLong Main
  300. CRaSh: Clustering, Removing, and Sharing Enhance Fine-tuning without Full Large Language ModelLong Main
  301. CReTIHC: Designing Causal Reasoning Tasks about Temporal Interventions and Hallucinated ConfoundingsShort Findings
  302. CRoW: Benchmarking Commonsense Reasoning in Real-World TasksLong Main
  303. CS2W: A Chinese Spoken-to-Written Style Conversion Dataset with Multiple Conversion TypesLong Main
  304. CT-GAT: Cross-Task Generative Adversarial Attack based on TransferabilityLong Main
  305. CTQScorer: Combining Multiple Features for In-context Example Selection for Machine TranslationLong Findings
  306. Cabbage Sweeter than Cake? Analysing the Potential of Large Language Models for Learning Conceptual SpacesShort Main
  307. Cache me if you Can: an Online Cost-aware Teacher-Student framework to Reduce the Calls to Large Language ModelsShort Findings
  308. Calc-X and Calcformers: Empowering Arithmetical Chain-of-Thought through Interaction with Symbolic SystemsShort Main
  309. Calibrated Seq2seq Models for Efficient and Generalizable Ultra-fine Entity TypingLong Findings
  310. Can Brain Signals Reveal Inner Alignment with Human Languages?Short Findings
  311. Can ChatGPT Assess Human Personalities? A General Evaluation FrameworkLong Findings
  312. Can ChatGPT Defend its Belief in Truth? Evaluating LLM Reasoning via DebateLong Findings
  313. Can ChatGPT Perform Reasoning Using the IRAC Method in Analyzing Legal Scenarios Like a Lawyer?Long Findings
  314. Can Foundation Models Watch, Talk and Guide You Step by Step to Make a Cake?Long Findings
  315. Can LLMs Facilitate Interpretation of Pre-trained Language Models?Long Main
  316. Can Language Models Laugh at YouTube Short-form Videos?Long Main
  317. Can Language Models Understand Physical Concepts?Long Main
  318. Can Large Language Models Capture Dissenting Human Voices?Long Main
  319. Can Large Language Models Fix Data Annotation Errors? An Empirical Study Using Debatepedia for Query-Focused Text SummarizationShort Findings
  320. Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?Long Main
  321. Can Retriever-Augmented Language Models Reason? The Blame Game Between the Retriever and the Language ModelLong Findings
  322. Can We Edit Factual Knowledge by In-Context Learning?Long Main
  323. Can We Edit Multimodal Large Language Models?Long Main
  324. Can You Follow Me? Testing Situational Understanding for ChatGPTLong Main
  325. Can you Summarize my learnings? Towards Perspective-based Educational Dialogue SummarizationLong Findings
  326. CaseEncoder: A Knowledge-enhanced Pre-trained Model for Legal Case EncodingLong Main
  327. Causal Document-Grounded Dialogue Pre-trainingLong Main
  328. Causal Inference from Text: Unveiling Interactions between VariablesLong Findings
  329. Causal Intervention for Abstractive Related Work GenerationLong Findings
  330. Causal Intervention-based Few-Shot Named Entity RecognitionLong Findings
  331. Causal Reasoning through Two Cognition Layers for Improving Generalization in Visual Question AnsweringLong Main
  332. Centering the Margins: Outlier-Based Identification of Harmed Populations in Toxicity DetectionLong Main
  333. Chain of Thought with Explicit Evidence Reasoning for Few-shot Relation ExtractionLong Findings
  334. Chain-of-Questions Training with Latent Answers for Robust Multistep Question AnsweringLong Main
  335. Chain-of-Thought Embeddings for Stance Detection on Social MediaShort Findings
  336. Chain-of-Thought Reasoning in Tabular Language ModelsLong Findings
  337. Chain-of-Thought Tuning: Masked Language Models can also Think Step By Step in Natural Language UnderstandingLong Main
  338. Challenges in Context-Aware Neural Machine TranslationLong Main
  339. Character-LLM: A Trainable Agent for Role-PlayingLong Main
  340. Characterizing Mechanisms for Factual Recall in Language ModelsLong Main
  341. Characterizing and Verifying Scientific Claims: Qualitative Causal Structure is All You NeedLong Main
  342. ChatCoT: Tool-Augmented Chain-of-Thought Reasoning on Chat-based Large Language ModelsLong Findings
  343. ChatEdit: Towards Multi-turn Interactive Facial Image Editing via DialogueLong Main
  344. ChatGPT Beyond English: Towards a Comprehensive Evaluation of Large Language Models in Multilingual LearningLong Findings
  345. ChatGPT to Replace Crowdsourcing of Paraphrases for Intent Classification: Higher Diversity and Comparable Model RobustnessLong Main
  346. Chinese Lexical Substitution: Dataset and MethodLong Main
  347. Chinese Metaphorical Relation ExtractionLong Findings
  348. Citance-Contextualized Summarization of Scientific PapersLong Findings
  349. CiteBench: A Benchmark for Scientific Citation Text GenerationLong Main
  350. CleanCoNLL: A Nearly Noise-Free Named Entity Recognition DatasetLong Main
  351. ClimateBERT-NetZero: Detecting and Assessing Net Zero and Reduction TargetsShort Main
  352. Clinical Contradiction DetectionLong Main
  353. Closed Boundary Learning for Classification Tasks with the Universum ClassLong Findings
  354. ClozEx: A Task toward Generation of English Cloze ExplanationLong Findings
  355. ClusterLLM: Large Language Models as a Guide for Text ClusteringLong Main
  356. ClusterPrompt: Cluster Semantic Enhanced Prompt Learning for New Intent DiscoveryLong Findings
  357. Clustering Pseudo Language Family in Multilingual Translation Models with Fisher Information MatrixShort Main
  358. Co$^2$PT: Mitigating Bias in Pre-trained Language Models through Counterfactual Contrastive Prompt TuningLong Findings
  359. Co-training and Co-distillation for Quality Improvement and Compression of Language ModelsLong Findings
  360. CoAnnotating: Uncertainty-Guided Work Allocation between Human and Large Language Models for Data AnnotationLong Main
  361. CoEdIT: Text Editing by Task-Specific Instruction TuningLong Findings
  362. CoF-CoT: Enhancing Large Language Models with Coarse-to-Fine Chain-of-Thought Prompting for Multi-domain NLU TasksShort Main
  363. CoLT5: Faster Long-Range Transformers with Conditional ComputationLong Main
  364. CoMPosT: Characterizing and Evaluating Caricature in LLM SimulationsLong Main
  365. CoRec: An Easy Approach for Coordination RecognitionShort Main
  366. CoSyn: Detecting Implicit Hate Speech in Online Conversations Using a Context Synergized Hyperbolic NetworkLong Main
  367. CoVariance-based Causal Debiasing for Entity and Relation ExtractionLong Findings
  368. Coarse-to-Fine Contrastive Learning in Image-Text-Graph Space for Improved Vision-Language CompositionalityLong Main
  369. Coarse-to-Fine Dual Encoders are Better Frame Identification LearnersLong Findings
  370. Code-Switching Metrics Using Intonation UnitsLong Main
  371. Code-Switching with Word Senses for Pretraining in Neural Machine TranslationLong Findings
  372. CodeBERTScore: Evaluating Code Generation with Pretrained Models of CodeLong Main
  373. CodeFusion: A Pre-trained Diffusion Model for Code GenerationShort Main
  374. CodeT5+: Open Code Large Language Models for Code Understanding and GenerationLong Main
  375. CodeTransOcean: A Comprehensive Multilingual Benchmark for Code TranslationLong Findings
  376. Cognate Transformer for Automated Phonological Reconstruction and Cognate Reflex PredictionLong Main
  377. Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?Short Main
  378. Coherent Entity Disambiguation via Modeling Topic and Categorical DependencyLong Findings
  379. Collaborative Generative AI: Integrating GPT-k for Efficient Editing in Text-to-Image GenerationShort Main
  380. CombLM: Adapting Black-Box Language Models through Small Fine-Tuned ModelsLong Main
  381. Combining Counting Processes and Classification Improves a Stopping Rule for Technology Assisted ReviewShort Findings
  382. Combining Denoising Autoencoders with Contrastive Learning to fine-tune Transformer ModelsLong Main
  383. Comparing Biases and the Impact of Multilingual Training across Multiple LanguagesLong Main
  384. Comparing Prompt-Based and Standard Fine-Tuning for Urdu Text ClassificationShort Findings
  385. Comparing Styles across LanguagesLong Main
  386. Comparing the Evaluation and Production of Loophole Behavior in Humans and Large Language ModelsLong Findings
  387. CompleQA: Benchmarking the Impacts of Knowledge Graph Completion Methods on Question AnsweringShort Findings
  388. Complex Event Schema Induction with Knowledge-Enriched Diffusion ModelLong Findings
  389. Complexity-Guided Curriculum Learning for Text GraphsLong Findings
  390. Compositional Generalization for Data-to-Text GenerationLong Findings
  391. CompoundPiece: Evaluating and Improving Decompounding Performance of Language ModelsLong Main
  392. Compressing Context to Enhance Inference Efficiency of Large Language ModelsLong Main
  393. Compressing and Debiasing Vision-Language Pre-Trained Models for Visual Question AnsweringLong Main
  394. ConPrompt: Pre-training a Language Model with Machine-Generated Data for Implicit Hate Speech DetectionLong Findings
  395. Conceptor-Aided Debiasing of Large Language ModelsLong Main
  396. Conceptual structure coheres in human cognition but not in large language modelsLong Main
  397. Condensing Multilingual Knowledge with Lightweight Language-Specific ModulesLong Main
  398. Conditional Natural Language InferenceLong Findings
  399. Conditioning on Dialog Acts improves Empathy Style TransferLong Findings
  400. Confidence-based Ensembling of Perspective-aware ModelsLong Main
  401. Conic10K: A Challenging Math Problem Understanding and Reasoning DatasetLong Findings
  402. Connecting degree and polarity: An artificial language learning studyLong Main
  403. Connecting the Dots: What Graph-Based Text Representations Work Best for Text Classification using Graph Neural Networks?Long Findings
  404. Consistency is Key: On Data-Efficient Modality Transfer in Speech TranslationShort Findings
  405. Consonant is all you need: a compact representation of English text for efficient NLPLong Findings
  406. Construction Artifacts in Metaphor Identification DatasetsShort Main
  407. Content- and Topology-Aware Representation Learning for Scientific Multi-LiteratureLong Main
  408. Context Compression for Auto-regressive Transformers with Sentinel TokensShort Main
  409. Context Quality Matters in Training Fusion-in-Decoder for Extractive Open-Domain Question AnsweringLong Findings
  410. Context-faithful Prompting for Large Language ModelsLong Findings
  411. Contextual Interaction for Argument Post Quality AssessmentLong Main
  412. Continual Dialogue State Tracking via Example-Guided Question AnsweringLong Main
  413. Continual Event Extraction with Semantic Confusion RectificationLong Main
  414. Continual Generalized Intent Discovery: Marching Towards Dynamic and Open-world Intent RecognitionLong Findings
  415. Continual Learning for Multilingual Neural Machine Translation via Dual Importance-based Model DivisionLong Main
  416. Continual Named Entity Recognition without Catastrophic ForgettingLong Main
  417. Continually Improving Extractive QA via Human FeedbackLong Main
  418. Contrastive Deterministic Autoencoders For Language ModelingLong Findings
  419. Contrastive Distant Supervision for Debiased and Denoised Machine Reading ComprehensionLong Findings
  420. Contrastive Learning for Inference in DialogueLong Main
  421. Contrastive Learning of Sentence Embeddings from ScratchLong Main
  422. Contrastive Learning-based Sentence Encoders Implicitly Weight Informative WordsShort Findings
  423. Contrastive Pre-training for Personalized Expert FindingLong Findings
  424. Controllable Chest X-Ray Report Generation from Longitudinal RepresentationsLong Findings
  425. Controllable Contrastive Generation for Multilingual Biomedical Entity LinkingLong Main
  426. Controlling Pre-trained Language Models for Grade-Specific Text SimplificationLong Main
  427. Conversation Chronicles: Towards Diverse Temporal and Relational Dynamics in Multi-Session ConversationsLong Main
  428. Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality InteractionLong Main
  429. Conversational Recommender System and Large Language Model Are Made for Each Other in E-commerce Pre-sales DialogueLong Findings
  430. Conversational Semantic Parsing using Dynamic Context GraphsLong Main
  431. Copyright Violations and Large Language ModelsShort Main
  432. CorefPrompt: Prompt-based Event Coreference Resolution by Measuring Event Type and Argument CompatibilitiesLong Main
  433. Counter Turing Test (CT2): AI-Generated Text Detection is Not as Easy as You May Think - Introducing AI Detectability Index (ADI)Long Main
  434. Countering Misinformation via Emotional Response GenerationLong Main
  435. Counting the Bugs in ChatGPT's Wugs: A Multilingual Investigation into the Morphological Capabilities of a Large Language ModelLong Main
  436. Coverage-based Example Selection for In-Context LearningLong Findings
  437. Critic-Driven Decoding for Mitigating Hallucinations in Data-to-text GenerationShort Main
  438. Cross-Cultural Analysis of Human Values, Morals, and Biases in Folk TalesLong Main
  439. Cross-Document Event Coreference Resolution on Discourse StructureLong Main
  440. Cross-Lingual Consistency of Factual Knowledge in Multilingual Language ModelsLong Main
  441. Cross-Lingual Cross-Target Stance Detection with Dual Knowledge Distillation FrameworkLong Main
  442. Cross-Modal Conceptualization in Bottleneck ModelsLong Main
  443. Cross-lingual Open-Retrieval Question Answering for African LanguagesLong Findings
  444. Cross-lingual Prompting: Improving Zero-shot Chain-of-Thought Reasoning across LanguagesLong Main
  445. Cross-lingual Transfer Can Worsen Bias in Sentiment AnalysisLong Main
  446. Cross-modality Data Augmentation for End-to-End Sign Language TranslationLong Findings
  447. Crossing the Aisle: Unveiling Partisan and Counter-Partisan Events in News ReportingShort Findings
  448. Crossing the Threshold: Idiomatic Machine Translation through Retrieval Augmentation and Loss WeightingLong Main
  449. Crosslingual Transfer Learning for Low-Resource Languages Based on Multilingual Colexification GraphsLong Findings
  450. Crystal: Introspective Reasoners Reinforced with Self-FeedbackLong Main
  451. Cue-CoT: Chain-of-thought Prompting for Responding to In-depth Dialogue Questions with LLMsLong Findings
  452. Cultural Compass: Predicting Transfer Learning Success in Offensive Language Detection with Cultural FeaturesLong Findings
  453. Cultural Concept Adaptation on Multimodal ReasoningLong Main
  454. Culturally Aware Natural Language InferenceLong Findings
  455. D$^2$TV: Dual Knowledge Distillation and Target-oriented Vision Modeling for Many-to-Many Multimodal SummarizationLong Findings
  456. DADA: Dialect Adaptation via Dynamic Aggregation of Linguistic RulesLong Main
  457. DALE: Generative Data Augmentation for Low-Resource Legal NLPLong Main
  458. DEPN: Detecting and Editing Privacy Neurons in Pretrained Language ModelsLong Main
  459. DISCO: A Large Scale Human Annotated Corpus for Disfluency Correction in Indo-European LanguagesLong Findings
  460. DIVE: Towards Descriptive and Diverse Visual Commonsense GenerationLong Main
  461. DNA: Denoised Neighborhood Aggregation for Fine-grained Category DiscoveryLong Main
  462. DPP-TTS: Diversifying prosodic features of speech via determinantal point processesLong Main
  463. DRAFT: Dense Retrieval Augmented Few-shot Topic classifier FrameworkLong Findings
  464. DREAM: Deployment of Recombination and Ensembles in Argument MiningLong Main
  465. DSI++: Updating Transformer Memory with New DocumentsLong Main
  466. DUMB: A Dutch Model BenchmarkLong Main
  467. DUnE: Dataset for Unified EditingLong Main
  468. Dancing Between Success and Failure: Edit-level Simplification Evaluation using SALSALong Main
  469. Data Augmentation for Code Translation with Comparable Corpora and Multiple ReferencesLong Findings
  470. Data Factors for Better Compositional GeneralizationLong Main
  471. Data Selection Curriculum for Abstractive Text SummarizationShort Findings
  472. Data Similarity is Not Enough to Explain Language Model PerformanceShort Main
  473. Data-efficient Active Learning for Structured Prediction with Partial Annotation and Self-TrainingLong Findings
  474. Dataset Bias Mitigation in Multiple-Choice Visual Question Answering and BeyondLong Findings
  475. DeCrisisMB: Debiased Semi-Supervised Learning for Crisis Tweet Classification via Memory BankLong Findings
  476. DeSIQ: Towards an Unbiased, Challenging Benchmark for Social Intelligence UnderstandingLong Main
  477. DeTiME: Diffusion-Enhanced Topic Modeling using Encoder-decoder based LLMLong Findings
  478. Debias NLU Datasets via Training-free PerturbationsLong Findings
  479. Debiasing Made State-of-the-art: Revisiting the Simple Seed-based Weak Supervision for Text ClassificationLong Main
  480. Debiasing Multimodal Models via Causal Information MinimizationLong Findings
  481. DecipherPref: Analyzing Influential Factors in Human Preference Judgments via GPT-4Long Main
  482. Deciphering Stereotypes in Pre-Trained Language ModelsLong Main
  483. DecoMT: Decomposed Prompting for Machine Translation Between Related Languages using Large Language ModelsLong Main
  484. Decoding Stumpers: Large Language Models vs. Human Problem-SolversShort Findings
  485. Decoding the Silent Majority: Inducing Belief Augmented Social Graph with Large Language Model for Response ForecastingLong Main
  486. Decomposed Prompt Tuning via Low-Rank ReparameterizationLong Findings
  487. Decomposing Complex Queries for Tip-of-the-tongue RetrievalLong Findings
  488. Deep Natural Language Feature Learning for Interpretable PredictionLong Main
  489. Defining a New NLP PlaygroundLong Findings
  490. Definitions Matter: Guiding GPT for Multi-label ClassificationShort Findings
  491. DeltaScore: Fine-Grained Story Evaluation with PerturbationsLong Findings
  492. DelucionQA: Detecting Hallucinations in Domain-specific Question AnsweringLong Findings
  493. DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language GroundingLong Findings
  494. DemoNSF: A Multi-task Demonstration-based Generative Framework for Noisy Slot Filling TaskShort Findings
  495. DemoSG: Demonstration-enhanced Schema-guided Generation for Low-resource Event ExtractionLong Findings
  496. Democratizing LLMs: An Exploration of Cost-Performance Trade-offs in Self-Refined Open-Source ModelsLong Findings
  497. Democratizing Reasoning Ability: Tailored Learning from Large Language ModelLong Main
  498. Demystifying Prompts in Language Models via Perplexity EstimationLong Findings
  499. Dense Retrieval as Indirect Supervision for Large-space Decision MakingLong Findings
  500. Density-Aware Prototypical Network for Few-Shot Relation ClassificationLong Findings
  501. DepNeCTI: Dependency-based Nested Compound Type Identification for SanskritLong Findings
  502. DepWiGNN: A Depth-wise Graph Neural Network for Multi-hop Spatial Reasoning in TextLong Findings
  503. Describe Me an Auklet: Generating Grounded Perceptual Category DescriptionsLong Main
  504. Descriptive Prompt Paraphrasing for Target-Oriented Multimodal Sentiment ClassificationLong Findings
  505. DetGPT: Detect What You Need via ReasoningLong Main
  506. DetectLLM: Leveraging Log Rank Information for Zero-Shot Detection of Machine-Generated TextLong Findings
  507. Detecting Erroneously Recognized Handwritten Byzantine TextLong Findings
  508. Detecting Propaganda Techniques in Code-Switched Social Media TextLong Main
  509. Detecting Syntactic Change with Pre-trained Transformer ModelsLong Findings
  510. Detecting and Mitigating Hallucinations in Multilingual SummarisationLong Main
  511. Detection of Multiple Mental Disorders from Social Media with Two-Stream Psychiatric ExpertsLong Main
  512. Detrimental Contexts in Open-Domain Question AnsweringLong Findings
  513. DiFair: A Benchmark for Disentangled Assessment of Gender Knowledge and BiasLong Findings
  514. DiNeR: A Large Realistic Dataset for Evaluating Compositional GeneralizationLong Main
  515. DiQAD: A Benchmark Dataset for Open-domain Dialogue Quality AssessmentLong Findings
  516. DiSTRICT: Dialogue State Tracking with Retriever Driven In-Context TuningLong Main
  517. DialCoT Meets PPO: Decomposing and Exploring Reasoning Paths in Smaller Language ModelsLong Main
  518. DialGuide: Aligning Dialogue Model Behavior with Developer GuidelinesLong Findings
  519. Dialect Transfer for Swiss German Speech TranslationLong Findings
  520. Dialect-to-Standard Normalization: A Large-Scale Multilingual EvaluationLong Findings
  521. DialogQAE: N-to-N Question Answer Pair Extraction from Customer Service ChatlogLong Findings
  522. Dialogizer: Context-aware Conversational-QA Dataset Generation from Textual SourcesLong Main
  523. Dialogue Act-Aided Backchannel Prediction Using Multi-Task LearningShort Findings
  524. Dialogue Chain-of-Thought Distillation for Commonsense-aware Conversational AgentsLong Main
  525. Dialogue Medical Information Extraction with Medical-Item Graph and Dialogue-Status Enriched RepresentationLong Findings
  526. Did You Mean...? Confidence-based Trade-offs in Semantic ParsingShort Main
  527. DiffS2UT: A Semantic Preserving Diffusion Model for Textless Direct Speech-to-Speech TranslationLong Main
  528. Difference-Masking: Choosing What to Mask in Continued PretrainingLong Findings
  529. DiffuSeq-v2: Bridging Discrete and Continuous Text Spaces for Accelerated Seq2Seq Diffusion ModelsShort Findings
  530. DiffuVST: Narrating Fictional Scenes with Global-History-Guided Denoising ModelsLong Findings
  531. Diffusion Language Model with Query-Document Relevance for Query-Focused SummarizationLong Findings
  532. DiffusionRet: Diffusion-Enhanced Generative Retriever using Constrained DecodingLong Findings
  533. DiffusionSL: Sequence Labeling via Tag Diffusion ProcessLong Findings
  534. Dimensions of Online Conflict: Towards Modeling AgonismLong Findings
  535. Dior-CVAE: Pre-trained Language Models and Diffusion Priors for Variational Dialog GenerationLong Findings
  536. DisCo: Distilled Student Models Co-training for Semi-supervised Text MiningLong Main
  537. Discourse Sense Flows: Modelling the Rhetorical Style of Documents across Various DomainsLong Findings
  538. Discourse Structures Guided Fine-grained Propaganda IdentificationLong Main
  539. Discovering Highly Influential Shortcut Reasoning: An Automated Template-Free ApproachShort Findings
  540. Discovering Universal Geometry in Embeddings with ICALong Main
  541. Disentangling Extraction and Reasoning in Multi-hop Spatial ReasoningLong Findings
  542. Disentangling Structure and Style: Political Bias Detection in News by Inducing Document HierarchyLong Findings
  543. Disentangling Transformer Language Models as Superposed Topic ModelsLong Main
  544. Disfluent Cues for Enhanced Speech Understanding in Large Language ModelsLong Findings
  545. Dissecting In-Context Learning of Translations in GPT-3Short Findings
  546. Dissecting Recall of Factual Associations in Auto-Regressive Language ModelsLong Main
  547. Distance-Based Propagation for Efficient Knowledge Graph ReasoningLong Main
  548. DistillCSE: Distilled Contrastive Learning for Sentence EmbeddingsLong Findings
  549. Distilling ChatGPT for Explainable Automated Student Answer AssessmentLong Findings
  550. Ditto: A Simple and Efficient Approach to Improve Sentence EmbeddingsShort Main
  551. Diversify Question Generation with Retrieval-Augmented Style TransferLong Main
  552. Diversifying language models for lesser-studied languages and language-usage contexts: A case of second language KoreanLong Findings
  553. Diversity Enhanced Narrative Question Generation for StorybooksLong Main
  554. Do All Languages Cost the Same? Tokenization in the Era of Commercial Language ModelsLong Main
  555. Do LLMs Understand Social Knowledge? Evaluating the Sociability of Large Language Models with SocKET BenchmarkLong Main
  556. Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language ModelsLong Main
  557. Do Stochastic Parrots have Feelings Too? Improving Neural Detection of Synthetic Text via Emotion RecognitionLong Findings
  558. Do “English” Named Entity Recognizers Work Well on Global Englishes?Long Findings
  559. DocAsRef: An Empirical Study on Repurposing Reference-based Summary Quality Metrics as Reference-free MetricsShort Findings
  560. DocSplit: Simple Contrastive Pretraining for Large Document EmbeddingsShort Findings
  561. DocTrack: A Visually-Rich Document Dataset Really Aligned with Human Eye Movement for Machine ReadingLong Findings
  562. Document-Level Machine Translation with Large Language ModelsLong Main
  563. Document-level Relationship Extraction by Bidirectional Constraints of Beta RulesLong Main
  564. Does Listener Gaze in Face-to-Face Interaction Follow the Entropy Rate Constancy Principle: An Empirical StudyShort Findings
  565. Does the Correctness of Factual Knowledge Matter for Factual Knowledge-Enhanced Pre-trained Language Models?Long Main
  566. Dolphin: A Challenging and Diverse Benchmark for Arabic NLGLong Findings
  567. Domain Adaptation for Conversational Query Production with the RAG Model FeedbackLong Findings
  568. Domain Adaptation for Sentiment Analysis Using Robust Internal RepresentationsLong Findings
  569. Domain Private Transformers for Multi-Domain Dialog SystemsShort Findings
  570. Don't waste a single annotation: improving single-label classifiers through soft labelsShort Findings
  571. Don’t Add, don’t Miss: Effective Content Preserving Generation from Pre-Selected Text SpansLong Findings
  572. Don’t Trust ChatGPT when your Question is not in English: A Study of Multilingual Abilities and Types of LLMsLong Main
  573. Doolittle: Benchmarks and Corpora for Academic Writing FormalizationLong Main
  574. Dr ChatGPT tell me what I want to hear: How different prompts impact health answer correctnessLong Main
  575. Drilling Down into the Discourse Structure with LLMs for Long Document Question AnsweringLong Findings
  576. Dual-Channel Span for Aspect Sentiment Triplet ExtractionLong Main
  577. Dual-Feedback Knowledge Retrieval for Task-Oriented Dialogue SystemsLong Main
  578. DueT: Image-Text Contrastive Transfer Learning with Dual-adapter TuningLong Main
  579. Dynamic Low-rank Estimation for Transformer-based Language ModelsLong Findings
  580. Dynamic Open-book Prompt for Conversational Recommender SystemLong Findings
  581. Dynamic Stance: Modeling Discussions by Labeling the InteractionsLong Findings
  582. Dynamic Stashing Quantization for Efficient Transformer TrainingShort Findings
  583. Dynamic Top-k Estimation Consolidates Disagreement between Feature Attribution MethodsShort Main
  584. Dynamic Voting for Efficient Reasoning in Large Language ModelsLong Findings
  585. Dynosaur: A Dynamic Growth Paradigm for Instruction-Tuning Data CurationLong Main
  586. E-CORE: Emotion Correlation Enhanced Empathetic Dialogue GenerationLong Main
  587. EARA: Improving Biomedical Semantic Textual Similarity with Entity-Aligned Attention and Retrieval AugmentationLong Findings
  588. ECHo: A Visio-Linguistic Dataset for Event Causality Inference via Human-Centric ReasoningLong Findings
  589. EDIS: Entity-Driven Image Search over Multimodal Web ContentLong Main
  590. EDeR: Towards Understanding Dependency Relations Between EventsLong Main
  591. EMO-KNOW: A Large Scale Dataset on Emotion-CauseShort Findings
  592. ESPVR: Entity Spans Position Visual Regions for Multimodal Named Entity RecognitionLong Findings
  593. EXPLAIN, EDIT, GENERATE: Rationale-Sensitive Counterfactual Data Augmentation for Multi-hop Fact VerificationLong Main
  594. EZ-STANCE: A Large Dataset for Zero-Shot Stance DetectionLong Findings
  595. EasyQuant: An Efficient Data-free Quantization Algorithm for LLMsLong Main
  596. Ecologically Valid Explanations for Label Variation in NLIShort Findings
  597. EconBERTa: Towards Robust Extraction of Named Entities in EconomicsLong Findings
  598. Editing Common Sense in TransformersLong Main
  599. Editing Large Language Models: Problems, Methods, and OpportunitiesLong Main
  600. Effects of Human Adversarial and Affable Samples on BERT GeneralizationLong Findings
  601. Effects of sub-word segmentation on performance of transformer language modelsLong Main
  602. Efficient Algorithms for Recognizing Weighted Tree-Adjoining LanguagesLong Main
  603. Efficient Classification of Long Documents via State-Space ModelsShort Main
  604. Efficient Continue Training of Temporal Language Model with Structural InformationLong Findings
  605. Efficient Cross-Task Prompt Tuning for Few-Shot Conversational Emotion RecognitionLong Findings
  606. Efficient Data Learning for Open Information Extraction with Pre-trained Language ModelsShort Findings
  607. Efficient Grammatical Error Correction Via Multi-Task Training and Optimized Training ScheduleLong Main
  608. Efficient Latent Variable Modeling for Knowledge-Grounded Dialogue GenerationLong Findings
  609. Efficient Long-Range Transformers: You Need to Attend More, but Not Necessarily at Every LayerLong Findings
  610. Efficient Multilingual Language Model Compression through Vocabulary TrimmingLong Findings
  611. Efficient k-NN Search with Cross-Encoders using Adaptive Multi-Round CUR DecompositionShort Findings
  612. Efficiently Enhancing Zero-Shot Performance of Instruction Following Model via Retrieval of Soft PromptLong Findings
  613. Elaborative Simplification as Implicit Questions Under DiscussionLong Main
  614. Emergence of Abstract State Representations in Embodied Sequence ModelingLong Main
  615. Emergent Inabilities? Inverse Scaling Over the Course of PretrainingShort Findings
  616. Empathy Intent Drives Empathy DetectionLong Main
  617. Empirical Study of Zero-Shot NER with ChatGPTLong Main
  618. Empower Nested Boolean Logic via Self-Supervised Curriculum LearningLong Main
  619. Empowering Psychotherapy with Large Language Models: Cognitive Distortion Detection through Diagnosis of Thought PromptingShort Findings
  620. Emptying the Ocean with a Spoon: Should We Edit Models?Short Findings
  621. Enabling Large Language Models to Generate Text with CitationsLong Main
  622. Enabling Unsupervised Neural Machine Translation with Word-level Visual RepresentationsLong Findings
  623. End-to-End Autoregressive Retrieval via Bootstrapping for Smart Reply SystemsLong Findings
  624. End-to-End Single-Channel Speaker-Turn Aware Conversational Speech TranslationLong Main
  625. End-to-end Adversarial Sample Generation for Data AugmentationLong Findings
  626. End-to-end Task-oriented Dialogue: A Survey of Tasks, Methods, and Future DirectionsLong Main
  627. Energy and Carbon Considerations of Fine-Tuning BERTShort Findings
  628. Enhanced Simultaneous Machine Translation with Word-level PoliciesLong Findings
  629. Enhancing Abstractiveness of Summarization Models through Calibrated DistillationLong Findings
  630. Enhancing Accessible Communication: from European Portuguese to Portuguese Sign LanguageShort Findings
  631. Enhancing Argument Structure Extraction with Efficient Leverage of Contextual InformationShort Findings
  632. Enhancing Biomedical Lay Summarisation with External Knowledge GraphsLong Main
  633. Enhancing Chat Language Models by Scaling High-quality Instructional ConversationsLong Main
  634. Enhancing Code-Switching for Cross-lingual SLU: A Unified View of Semantic and Grammatical CoherenceShort Main
  635. Enhancing Computation Efficiency in Large Language Models through Weight and Activation QuantizationLong Main
  636. Enhancing Conversational Search: Large Language Model-Aided Informative Query RewritingLong Findings
  637. Enhancing Emotion Recognition in Conversation via Multi-view Feature Alignment and MemorizationLong Findings
  638. Enhancing Generative Retrieval with Reinforcement Learning from Relevance FeedbackLong Main
  639. Enhancing Low-resource Fine-grained Named Entity Recognition by Leveraging Coarse-grained DatasetsLong Main
  640. Enhancing Neural Machine Translation with Semantic UnitsLong Findings
  641. Enhancing Reasoning Capabilities by Instruction Learning and Chain-of-Thoughts for Implicit Discourse Relation RecognitionShort Findings
  642. Enhancing Retrieval-Augmented Large Language Models with Iterative Retrieval-Generation SynergyLong Findings
  643. Enhancing Scalability of Pre-trained Language Models via Efficient Parameter SharingLong Findings
  644. Enhancing Structured Evidence Extraction for Fact VerificationLong Main
  645. Enhancing Task-oriented Dialogue Systems with Generative Post-processing NetworksLong Main
  646. Enhancing Text-to-SQL Capabilities of Large Language Models: A Study on Prompt Design StrategiesLong Findings
  647. Enhancing Textbooks with Visuals from the Web for Improved LearningLong Main
  648. Enhancing Uncertainty-Based Hallucination Detection with Stronger FocusLong Main
  649. Enhancing the Ranking Context of Dense Retrieval through Reciprocal Nearest NeighborsLong Main
  650. Ensemble-Instruct: Instruction Tuning Data Generation with a Heterogeneous Mixture of LMsLong Findings
  651. EntSUMv2: Dataset, Models and Evaluation for More Abstractive Entity-Centric SummarizationShort Main
  652. Entity Disambiguation on a Tight Labeling BudgetShort Findings
  653. Entity-Based Evaluation of Political Bias in Automatic SummarizationShort Findings
  654. EpiK-Eval: Evaluation for Language Models as Epistemic ModelsLong Main
  655. Epsilon Sampling Rocks: Investigating Sampling Strategies for Minimum Bayes Risk Decoding for Machine TranslationLong Findings
  656. Error Detection for Text-to-SQL Semantic ParsingLong Findings
  657. Establishing Trustworthiness: Rethinking Tasks and Model EvaluationShort Main
  658. Estimating Large Language Model Capabilities without Labeled Test DataLong Findings
  659. Ethical Reasoning over Moral Alignment: A Case and Framework for In-Context Ethical Policies in LLMsLong Findings
  660. EtiCor: Corpus for Analyzing LLMs for EtiquettesShort Main
  661. Evaluating Bias and Fairness in Gender-Neutral Pretrained Vision-and-Language ModelsLong Main
  662. Evaluating Cross-Domain Text-to-SQL Models and BenchmarksLong Main
  663. Evaluating Dependencies in Fact Editing for Language Models: Specificity and Implication AwarenessLong Findings
  664. Evaluating Emotion Arcs Across Languages: Bridging the Global Divide in Sentiment AnalysisLong Findings
  665. Evaluating Evaluation Metrics: A Framework for Analyzing NLG Evaluation Metrics using Measurement TheoryLong Main
  666. Evaluating Large Language Models on Controlled Generation TasksLong Main
  667. Evaluating Object Hallucination in Large Vision-Language ModelsLong Main
  668. Evaluating Parameter-Efficient Finetuning Approaches for Pre-trained Models on the Financial DomainShort Findings
  669. Evaluating Subjective Cognitive Appraisals of Emotions from Large Language ModelsLong Findings
  670. Evaluating Verifiability in Generative Search EnginesLong Findings
  671. Evaluating and Enhancing the Robustness of Code Pre-trained Models through Structure-Aware Adversarial Samples GenerationLong Findings
  672. Evaluating and Modeling Attribution for Cross-Lingual Question AnsweringLong Main
  673. Evaluating the Knowledge Base Completion Potential of GPTShort Findings
  674. Evaluating the Rationale Understanding of Critical Reasoning in Logical Reading ComprehensionLong Main
  675. Evaluation Metrics in the Era of GPT-4: Reliably Evaluating Large Language Models on Sequence to Sequence TasksShort Main
  676. Evaluation of African American Language Bias in Natural Language GenerationLong Main
  677. Event Causality Extraction via Implicit Cause-Effect InteractionsLong Main
  678. Event Ontology Completion with Hierarchical Structure Evolution NetworksLong Main
  679. Event-Location Tracking in Narratives: A Case Study on Holocaust TestimoniesLong Main
  680. Examining Inter-Consistency of Large Language Models Collaboration: An In-depth Analysis via DebateLong Findings
  681. Example-based Hypernetworks for Multi-source Adaptation to Unseen DomainsLong Findings
  682. Exchange-of-Thought: Enhancing Large Language Model Capabilities through Cross-Model CommunicationLong Main
  683. Execution-Based Evaluation for Open-Domain Code GenerationLong Findings
  684. ExpNote: Black-box Large Language Models are better Task Solvers with Experience NotebookShort Findings
  685. Expand, Highlight, Generate: RL-driven Document Generation for Passage RerankingLong Main
  686. Explain-then-translate: an analysis on improving program translation with self-generated explanationsLong Findings
  687. ExplainCPE: A Free-text Explanation Benchmark of Chinese Pharmacist ExaminationLong Findings
  688. Explainable Claim Verification via Knowledge-Grounded Reasoning with Large Language ModelsLong Findings
  689. Explaining Interactions Between Text SpansLong Main
  690. Explaining with Contrastive Phrasal Highlighting: A Case Study in Assisting Humans to Detect Translation DifferencesLong Main
  691. Explanation Selection Using Unlabeled Data for Chain-of-Thought PromptingLong Main
  692. Explicit Alignment and Many-to-many Entailment Based Reasoning for Conversational Machine ReadingLong Findings
  693. Explicit Planning Helps Language Models in Logical ReasoningLong Main
  694. Exploiting Asymmetry for Synthetic Training Data Generation: SynthIE and the Case of Information ExtractionLong Main
  695. Exploiting Contrastive Learning and Numerical Evidence for Confusing Legal Judgment PredictionLong Findings
  696. Exploiting Emotion-Semantic Correlations for Empathetic Response GenerationLong Findings
  697. Explore the Way: Exploring Reasoning Path by Bridging Entities for Effective Cross-Document Relation ExtractionShort Findings
  698. Explore-Instruct: Enhancing Domain-Specific Instruction Coverage through Active ExplorationLong Main
  699. Exploring All-In-One Knowledge Distillation Framework for Neural Machine TranslationLong Main
  700. Exploring Chain of Thought Style Prompting for Text-to-SQLLong Main
  701. Exploring Context-Aware Evaluation Metrics for Machine TranslationShort Findings
  702. Exploring Discourse Structure in Document-level Machine TranslationLong Main
  703. Exploring Graph Pre-training for Aspect-based Sentiment AnalysisLong Findings
  704. Exploring In-Context Learning for Knowledge Grounded Dialog GenerationLong Findings
  705. Exploring Jiu-Jitsu Argumentation for Writing Peer Review RebuttalsLong Main
  706. Exploring Large Language Models for Multi-Modal Out-of-Distribution DetectionLong Findings
  707. Exploring Linguistic Probes for Morphological InflectionShort Main
  708. Exploring Linguistic Properties of Monolingual BERTs with Typological Classification among LanguagesLong Findings
  709. Exploring the Boundaries of GPT-4 in RadiologyLong Main
  710. Exploring the Cognitive Knowledge Structure of Large Language Models: An Educational Diagnostic Assessment ApproachShort Findings
  711. Exploring the Effectiveness of Multi-Lingual Commonsense Knowledge-Aware Open-Domain Dialogue Response GenerationLong Findings
  712. Exploring the Impact of Corpus Diversity on Financial Pretrained Language ModelsShort Findings
  713. Exploring the Impact of Model Scaling on Parameter-Efficient TuningLong Main
  714. Exploring the Numerical Reasoning Capabilities of Language Models: A Comprehensive Analysis on Tabular DataLong Findings
  715. Exploring the Potential of Large Language Models in Generating Code-Tracing Questions for Introductory Programming CoursesShort Findings
  716. Exploring the Sensitivity of LLMs' Decision-Making Capabilities: Insights from Prompt Variations and HyperparametersShort Findings
  717. Expository Text Generation: Imitate, Retrieve, ParaphraseLong Main
  718. Extractive Summarization via ChatGPT for Faithful Summary GenerationShort Findings
  719. Extrapolating Multilingual Understanding Models as Multilingual GeneratorsLong Findings
  720. Eyes Show the Way: Modelling Gaze Behaviour for Hallucination DetectionLong Findings
  721. FACTIFY3M: A benchmark for multimodal fact verification with explainability through 5W Question-AnsweringLong Main
  722. FANToM: A Benchmark for Stress-testing Machine Theory of Mind in InteractionsLong Main
  723. FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationLong Main
  724. FFAEval: Evaluating Dialogue System via Free-For-All RankingLong Findings
  725. FLatS: Principled Out-of-Distribution Detection with Feature-Based Likelihood Ratio ScoreShort Main
  726. FOCUS: Effective Embedding Initialization for Monolingual Specialization of Multilingual ModelsLong Main
  727. FREDSum: A Dialogue Summarization Corpus for French Political DebatesLong Findings
  728. FaLA: Fast Linear Adaptation for Replacing Backbone Models on Edge DevicesLong Findings
  729. FaMeSumm: Investigating and Improving Faithfulness of Medical SummarizationLong Main
  730. FactKB: Generalizable Factuality Evaluation using Language Models Enhanced with Factual KnowledgeLong Main
  731. FactSpotter: Evaluating the Factual Faithfulness of Graph-to-Text GenerationLong Findings
  732. Factual Relation Discrimination for Factuality-oriented Abstractive SummarizationLong Findings
  733. Failures Pave the Way: Enhancing Large Language Models through Tuning-free Rule AccumulationLong Main
  734. Fair Text Classification with Wasserstein IndependenceLong Main
  735. Fair Without Leveling Down: A New Intersectional Fairness DefinitionLong Main
  736. Faithful Model Evaluation for Model-Based MetricsShort Main
  737. Fast and Accurate Factual Inconsistency Detection Over Long DocumentsLong Main
  738. Fast and Robust Early-Exiting Framework for Autoregressive Language Models with Synchronized Parallel DecodingLong Main
  739. Faster Minimum Bayes Risk Decoding with Confidence-based PruningShort Main
  740. FedID: Federated Interactive Distillation for Large-Scale Pretraining Language ModelsLong Main
  741. FedTherapist: Mental Health Monitoring with User-Generated Linguistic Expressions on Smartphones via Federated LearningShort Main
  742. Federated Learning of Large Language Models with Parameter-Efficient Prompt Tuning and Adaptive OptimizationLong Main
  743. Few-shot Unified Question Answering: Tuning Models or Prompts?Long Findings
  744. Fidelity-Enriched Contrastive Search: Reconciling the Faithfulness-Diversity Trade-Off in Text GenerationShort Main
  745. Fighting Fire with Fire: The Dual Role of LLMs in Crafting and Detecting Elusive DisinformationLong Main
  746. Filling the Image Information Gap for VQA: Prompting Large Language Models to Proactively Ask QuestionsLong Findings
  747. FinEntity: Entity-level Sentiment Classification for Financial TextsShort Main
  748. FinGPT: Large Generative Models for a Small LanguageLong Main
  749. Find-2-Find: Multitask Learning for Anaphora Resolution and Object LocalizationLong Main
  750. Finding Authentic Counterhate Arguments: A Case Study with Public FiguresLong Main
  751. Finding Common Ground: Annotating and Predicting Common Ground in Spoken ConversationsLong Findings
  752. Finding Support Examples for In-Context LearningLong Findings
  753. Fine-grained Conversational Decoding via Isotropic and Proximal SearchShort Main
  754. Fine-tuned LLMs Know More, Hallucinate Less with Few-Shot Sequence-to-Sequence Semantic Parsing over WikidataLong Main
  755. FinePrompt: Unveiling the Role of Finetuned Inductive Bias on Compositional Reasoning in GPT-4Short Findings
  756. Flatness-Aware Prompt Selection Improves Accuracy and Sample EfficiencyLong Findings
  757. Focus Your Attention (with Adaptive IIR Filters)Long Main
  758. Focus on the Core: Efficient Attention via Pruned Token Compression for Document ClassificationLong Findings
  759. For Generated Text, Is NLI-Neutral Text the Best Text?Short Findings
  760. FreeAL: Towards Human-Free Active Learning in the Era of Large Language ModelsLong Main
  761. Frequency Balanced Datasets Lead to Better Language ModelsLong Findings
  762. From Complex to Simple: Unraveling the Cognitive Tree for Reasoning with Small Language ModelsLong Findings
  763. From Dissonance to Insights: Dissecting Disagreements in Rationale Construction for Case Outcome ClassificationLong Main
  764. From Heuristic to Analytic: Cognitively Motivated Strategies for Coherent Physical Commonsense ReasoningLong Main
  765. From Multilingual Complexity to Emotional Clarity: Leveraging Commonsense to Unveil Emotions in Code-Mixed DialoguesLong Main
  766. From Parse-Execute to Parse-Execute-Refine: Improving Semantic Parser for Complex Question Answering over Knowledge BaseLong Main
  767. From Relevance to Utility: Evidence Retrieval with Feedback for Fact VerificationShort Findings
  768. From Simple to Complex: A Progressive Framework for Document-level Informative Argument ExtractionLong Findings
  769. From Speculation Detection to Trustworthy Relational Tuples in Information ExtractionLong Findings
  770. From Values to Opinions: Predicting Human Behaviors and Stances Using Value-Injected Large Language ModelsLong Main
  771. From Words to Wires: Generating Functioning Electronic Devices from Natural Language DescriptionsLong Findings
  772. From Wrong To Right: A Recursive Approach Towards Vision-Language ExplanationLong Main
  773. Frugal Prompting for Dialog ModelsLong Findings
  774. Fusing Temporal Graphs into Transformers for Time-Sensitive Question AnsweringLong Findings
  775. G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentLong Main
  776. G-SPEED: General SParse Efficient Editing MoDelLong Findings
  777. GATITOS: Using a New Multilingual Lexicon for Low-resource Machine TranslationLong Main
  778. GBT: Generative Boosting Training Approach for Paraphrase IdentificationLong Findings
  779. GD-COMET: A Geo-Diverse Commonsense Inference ModelShort Main
  780. GDA: Grammar-based Data Augmentation for Text Classification using Slot InformationLong Findings
  781. GEM: Gestalt Enhanced Markup Language Model for Web Understanding via Render TreeLong Main
  782. GEMINI: Controlling The Sentence-Level Summary Style in Abstractive Text SummarizationLong Main
  783. GLEN: General-Purpose Event Detection for Thousands of TypesLong Main
  784. GLEN: Generative Retrieval via Lexical Index LearningLong Main
  785. GLGR: Question-aware Global-to-Local Graph Reasoning for Multi-party Dialogue Reading ComprehensionLong Findings
  786. GNAT: A General Narrative Alignment ToolLong Main
  787. GPT Deciphering Fedspeak: Quantifying Dissent Among Hawks and DovesShort Findings
  788. GPT-4 as an Effective Zero-Shot Evaluator for Scientific Figure CaptionsShort Findings
  789. GPT-RE: In-context Learning for Relation Extraction using Large Language ModelsLong Main
  790. GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLPLong Main
  791. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head CheckpointsShort Main
  792. GRACE: Discriminator-Guided Chain-of-Thought ReasoningLong Findings
  793. GRENADE: Graph-Centric Language Model for Self-Supervised Representation Learning on Text-Attributed GraphsLong Findings
  794. GRI: Graph-based Relative Isomorphism of Word Embedding SpacesLong Findings
  795. GROOViST: A Metric for Grounding Objects in Visual StorytellingShort Main
  796. GROVE: A Retrieval-augmented Complex Story Generation Framework with A Forest of EvidenceLong Findings
  797. GSAP-NER: A Novel Task, Corpus, and Baseline for Scholarly Entity Extraction Focused on Machine Learning Models and DatasetsLong Findings
  798. GTA: Gated Toxicity Avoidance for LM Performance PreservationLong Findings
  799. GazeVQA: A Video Question Answering Dataset for Multiview Eye-Gaze Task-Oriented CollaborationsLong Main
  800. GenKIE: Robust Generative Multimodal Document Key Information ExtractionLong Findings
  801. Gender Biases in Automatic Evaluation Metrics for Image CaptioningLong Main
  802. Generalizing Few-Shot Named Entity Recognizers to Unseen Domains with Type-Related FeaturesLong Findings
  803. Generating Commonsense Counterfactuals for Stable Relation ExtractionLong Main
  804. Generating Data for Symbolic Language with Large Language ModelsLong Main
  805. Generating Extractive Answers: Gated Recurrent Memory Reader for Conversational Question AnsweringShort Findings
  806. Generating Summaries with Controllable Readability LevelsLong Main
  807. Generating and Evaluating Tests for K-12 Students with Language Model Simulations: A Case Study on Sentence Reading EfficiencyLong Main
  808. Generative Adversarial Training with Perturbed Token Detection for Model RobustnessLong Main
  809. Generative Calibration for In-context LearningLong Findings
  810. Generative Emotion Cause Triplet Extraction in Conversations with Commonsense KnowledgeLong Findings
  811. Generative Spoken Language Model based on continuous word-sized audio tokensLong Main
  812. Generative Table Pre-training Empowers Models for Tabular PredictionLong Main
  813. GeoLM: Empowering Language Models for Geospatially Grounded Language UnderstandingLong Main
  814. Geographical Erasure in Language GenerationLong Findings
  815. Getting MoRE out of Mixture of Language Model Reasoning ExpertsLong Findings
  816. Give Me the Facts! A Survey on Factual Knowledge Probing in Pre-trained Language ModelsLong Findings
  817. Global Structure Knowledge-Guided Relation Extraction Method for Visually-Rich DocumentLong Findings
  818. Global Voices, Local Biases: Socio-Cultural Prejudices across LanguagesLong Main
  819. GlobalBench: A Benchmark for Global Progress in Natural Language ProcessingLong Main
  820. GlotLID: Language Identification for Low-Resource LanguagesLong Findings
  821. Goal-Driven Explainable Clustering via Language DescriptionsLong Main
  822. Gold: A Global and Local-aware Denoising Framework for Commonsense Knowledge Graph Noise DetectionLong Findings
  823. Good Meta-tasks Make A Better Cross-lingual Meta-transfer Learning for Low-resource LanguagesLong Findings
  824. Goodtriever: Adaptive Toxicity Mitigation with Retrieval-augmented ModelsLong Findings
  825. GradSim: Gradient-Based Language Grouping for Effective Multilingual TrainingLong Main
  826. Gradient-based Gradual Pruning for Language-Specific Multilingual Neural Machine TranslationLong Main
  827. Gradually Excavating External Knowledge for Implicit Complex Question AnsweringLong Findings
  828. Grammar-Constrained Decoding for Structured NLP Tasks without FinetuningLong Main
  829. Grammatical Error Correction via Mixed-Grained Weighted TrainingLong Findings
  830. Granularity Matters: Pathological Graph-driven Cross-modal Alignment for Brain CT Report GenerationLong Main
  831. Graph vs. Sequence: An Empirical Study on Knowledge Forms for Knowledge-Grounded DialogueLong Main
  832. GreedyCAS: Unsupervised Scientific Abstract Segmentation with Normalized Mutual InformationLong Main
  833. Grounded and well-rounded: a methodological approach to the study of cross-modal and cross-lingual groundingLong Findings
  834. Grounding Visual Illusions in Language: Do Vision-Language Models Perceive Illusions Like Humans?Long Main
  835. Guideline Learning for In-Context Information ExtractionLong Main
  836. Guiding LLM to Fool Itself: Automatically Manipulating Machine Reading Comprehension Shortcut TriggersShort Findings
  837. HANSEN: Human and AI Spoken Text Benchmark for Authorship AnalysisLong Findings
  838. HARE: Explainable Hate Speech Detection with Step-by-Step ReasoningShort Findings
  839. HEAR: Hearing Enhanced Audio Response for Video-grounded DialogueLong Findings
  840. HFMRE: Constructing Huffman Tree in Bags to Find Excellent Instances for Distantly Supervised Relation ExtractionLong Findings
  841. HPE: Answering Complex Questions over Text by Hybrid Question Parsing and ExecutionLong Findings
  842. HadSkip: Homotopic and Adaptive Layer Skipping of Pre-trained Language Models for Efficient InferenceLong Findings
  843. HalOmi: A Manually Annotated Benchmark for Multilingual Hallucination and Omission Detection in Machine TranslationLong Main
  844. Hallucination Detection for Generative Large Language Models by Bayesian Sequential EstimationLong Main
  845. Hallucination Detection for Grounded Instruction GenerationShort Findings
  846. Hallucination Mitigation in Natural Language Generation from Large-Scale Open-Domain Knowledge GraphsLong Main
  847. HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language ModelsLong Main
  848. Handshape-Aware Sign Language Recognition: Extended Datasets and Exploration of Handshape-Inclusive MethodsLong Findings
  849. Harnessing Black-Box Control to Boost Commonsense in LM's GenerationLong Main
  850. Harnessing Dataset Cartography for Improved Compositional Generalization in TransformersLong Findings
  851. Harnessing the Power of Large Language Models for Empathetic Response Generation: Empirical Investigations and ImprovementsLong Findings
  852. Harnessing the power of LLMs: Evaluating human-AI text co-creation through the lens of news headline generationLong Findings
  853. Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language ModelsLong Main
  854. HeQ: a Large and Diverse Hebrew Reading Comprehension BenchmarkLong Findings
  855. Hi Guys or Hi Folks? Benchmarking Gender-Neutral Machine Translation with the GeNTE CorpusLong Main
  856. Hi-ArG: Exploring the Integration of Hierarchical Argumentation Graphs in Language PretrainingLong Main
  857. Hi-ToM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language ModelsLong Findings
  858. HiCL: Hierarchical Contrastive Learning of Unsupervised Sentence EmbeddingsLong Findings
  859. HiddenTables and PyQTax: A Cooperative Game and Dataset For TableQA to Ensure Scale and Data Privacy Across a Myriad of TaxonomiesLong Main
  860. Hidding the Ghostwriters: An Adversarial Evaluation of AI-Generated Student Essay DetectionLong Main
  861. Hiding in Plain Sight: Tweets with Hate Speech Masked by HomoglyphsShort Findings
  862. Hierarchical Catalogue Generation for Literature Review: A BenchmarkLong Findings
  863. Hierarchical Enhancement Framework for Aspect-based Argument MiningLong Findings
  864. Hierarchical Fusion for Online Multimodal Dialog Act ClassificationLong Findings
  865. Hierarchical Pretraining on Multimodal Electronic Health RecordsLong Main
  866. Hierarchical Prompting Assists Large Language Model on Web NavigationShort Findings
  867. HierarchicalContrast: A Coarse-to-Fine Contrastive Learning Framework for Cross-Domain Zero-Shot Slot FillingLong Findings
  868. High-quality argumentative information in low resources approaches improve counter-narrative generationLong Findings
  869. HistAlign: Improving Context Dependency in Language Generation by Aligning with HistoryLong Main
  870. Holistic Inter-Annotator Agreement and Corpus Coherence Estimation in a Large-scale Multilingual Annotation CampaignLong Main
  871. Homophone Disambiguation Reveals Patterns of Context Mixing in Speech TransformersLong Main
  872. HoneyBee: Progressive Instruction Finetuning of Large Language Models for Materials ScienceLong Findings
  873. How Do Large Language Models Capture the Ever-changing World Knowledge? A Review of Recent AdvancesLong Main
  874. How Does Generative Retrieval Scale to Millions of Passages?Long Main
  875. How Many Demonstrations Do You Need for In-context Learning?Long Findings
  876. How Predictable Are Large Language Model Capabilities? A Case Study on BIG-benchLong Findings
  877. How Reliable Are AI-Generated-Text Detectors? An Assessment Framework Using Evasive Soft PromptsLong Findings
  878. How Well Do Text Embedding Models Understand Syntax?Long Findings
  879. How do languages influence each other? Studying cross-lingual data sharing during LM fine-tuningLong Main
  880. How to Determine the Most Powerful Pre-trained Language Model without Brute Force Fine-tuning? An Empirical SurveyLong Findings
  881. How to Enhance Causal Discrimination of Utterances: A Case on Affective ReasoningLong Main
  882. How to Train Your Dragon: Diverse Augmentation Towards Generalizable Dense RetrievalLong Findings
  883. HuatuoGPT, Towards Taming Language Model to Be a DoctorLong Findings
  884. Human Learning by Model Feedback: The Dynamics of Iterative Prompting with MidjourneyLong Main
  885. Human Raters Cannot Distinguish English Translations from Original English TextsShort Main
  886. HutCRS: Hierarchical User-Interest Tracking for Conversational Recommender SystemLong Main
  887. Hybrid Inverted Index Is a Robust Accelerator for Dense RetrievalLong Main
  888. HyperNetwork-based Decoupling to Improve Model Generalization for Few-Shot Relation ExtractionLong Main
  889. HyperRouter: Towards Efficient Training and Inference of Sparse Mixture of ExpertsShort Main
  890. Hyperpolyglot LLMs: Cross-Lingual Interpretability in Token EmbeddingsShort Main
  891. IAEval: A Comprehensive Evaluation of Instance Attribution on Natural Language UnderstandingLong Findings
  892. IAG: Induction-Augmented Generation Framework for Answering Reasoning QuestionsLong Main
  893. IBADR: an Iterative Bias-Aware Dataset Refinement Framework for Debiasing NLU modelsLong Main
  894. IC3: Image Captioning by Committee ConsensusLong Main
  895. ICU: Conquering Language Barriers in Vision-and-Language Modeling by Dividing the Tasks into Image Captioning and Language UnderstandingShort Findings
  896. IDTraffickers: An Authorship Attribution Dataset to link and connect Potential Human-Trafficking Operations on Text Escort AdvertisementsLong Main
  897. IEKG: A Commonsense Knowledge Graph for Idiomatic ExpressionsLong Main
  898. IMTLab: An Open-Source Platform for Building, Evaluating, and Diagnosing Interactive Machine Translation SystemsLong Main
  899. IMU2CLIP: Language-grounded Motion Sensor Translation with Multimodal Contrastive LearningShort Findings
  900. INA: An Integrative Approach for Enhancing Negotiation Strategies with Reward-Based Dialogue AgentLong Findings
  901. INFORM : Information eNtropy based multi-step reasoning FOR large language ModelsLong Main
  902. INGENIOUS: Using Informative Data Subsets for Efficient Pre-Training of Language ModelsLong Findings
  903. INSTRUCTSCORE: Towards Explainable Text Generation Evaluation with Automatic FeedbackLong Main
  904. INVITE: a Testbed of Automatically Generated Invalid Questions to Evaluate Large Language Models for HallucinationsShort Findings
  905. INarIG: Iterative Non-autoregressive Instruct Generation Model For Word-Level Auto CompletionLong Findings
  906. IRFL: Image Recognition of Figurative LanguageLong Findings
  907. IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language ModelsLong Findings
  908. Identification of Multimodal Stance Towards Frames of CommunicationLong Main
  909. Identifying Conspiracy Theories News based on Event Relation GraphLong Findings
  910. Identifying Informational Sources in News ArticlesLong Main
  911. Identifying Statements Crucial for Awareness of Interpretive Nonsense to Prevent Communication BreakdownsLong Main
  912. Identifying {Early Maladaptive Schemas} from Mental Health Question TextsShort Findings
  913. Ideology Takes Multiple Looks: A High-Quality Dataset for Multifaceted Ideology DetectionLong Main
  914. IfQA: A Dataset for Open-domain Question Answering under Counterfactual PresuppositionsLong Main
  915. Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs Through a Global Prompt Hacking CompetitionLong Main
  916. Image Manipulation via Multi-Hop Instructions - A New Dataset and Weakly-Supervised Neuro-Symbolic ApproachLong Main
  917. Image and Text: Fighting the same Battle? Super Resolution Learning for Imbalanced Text ClassificationLong Findings
  918. ImageNetVC: Zero- and Few-Shot Visual Commonsense Evaluation on 1000 ImageNet CategoriesLong Findings
  919. Impact of Co-occurrence on Factual Knowledge of Large Language ModelsLong Findings
  920. Implicit Sense-labeled Connective Recognition as Text GenerationShort Findings
  921. Impressions: Visual Semiotics and Aesthetic Impact UnderstandingLong Main
  922. Improved Pseudo Data for Machine Translation Quality Estimation with Constrained Beam SearchLong Main
  923. Improved Training of Deep Text ClusteringShort Findings
  924. Improved Unsupervised Chinese Word Segmentation Using Pre-trained Knowledge and Pseudo-labeling TransferShort Main
  925. Improving Bias Mitigation through Bias Experts in Natural Language UnderstandingLong Main
  926. Improving Biomedical Abstractive Summarisation with Knowledge Aggregation from Citation PapersLong Main
  927. Improving Chinese Pop Song and Hokkien Gezi Opera Singing Voice Synthesis by Enhancing Local ModelingLong Main
  928. Improving Consistency for Text Summarization with Energy FunctionsShort Findings
  929. Improving Contrastive Learning of Sentence Embeddings with Focal InfoNCEShort Findings
  930. Improving Conversational Recommendation Systems via Bias Analysis and Language-Model-Enhanced Data AugmentationLong Findings
  931. Improving Cross-lingual Transfer through Subtree-aware Word ReorderingLong Findings
  932. Improving Dialogue Discourse Parsing via Reply-to Structures of Addressee RecognitionLong Main
  933. Improving Diversity of Demographic Representation in Large Language Models via Collective-Critiques and Self-VotingLong Main
  934. Improving End-to-End Speech Processing by Efficient Text Data Utilization with Latent SynthesisLong Findings
  935. Improving Factual Consistency for Knowledge-Grounded Dialogue Systems via Knowledge Enhancement and AlignmentLong Findings
  936. Improving Image Captioning via Predicting Structured ConceptsLong Main
  937. Improving Input-label Mapping with Demonstration Replay for In-context LearningLong Findings
  938. Improving Language Models’ Meaning Understanding and Consistency by Learning Conceptual Roles from DictionaryLong Main
  939. Improving Long Document Topic Segmentation Models With Enhanced Coherence ModelingLong Main
  940. Improving Low-resource Question Answering by Augmenting Question InformationShort Findings
  941. Improving Multi-Criteria Chinese Word Segmentation through Learning Sentence RepresentationShort Findings
  942. Improving Multimodal Sentiment Analysis: Supervised Angular margin-based Contrastive Learning for Enhanced Fusion RepresentationLong Findings
  943. Improving Neural Machine Translation by Multi-Knowledge Integration with PromptingLong Findings
  944. Improving Pacing in Long-Form Story PlanningShort Findings
  945. Improving Question Generation with Multi-level Content PlanningLong Findings
  946. Improving Seq2Seq Grammatical Error Correction via Decoding InterventionsLong Findings
  947. Improving Sequential Model Editing with Fact RetrievalLong Findings
  948. Improving Span Representation by Efficient Span-Level AttentionShort Findings
  949. Improving Speech Translation by Fusing Speech and TextLong Findings
  950. Improving Summarization with Human EditsLong Main
  951. Improving Transformer-based Program Repair Model through False Behavior DiagnosisLong Main
  952. Improving Unsupervised Relation Extraction by Augmenting Diverse Sentence PairsLong Main
  953. Improving Zero-shot Reader by Reducing Distractions from Irrelevant Documents in Open-Domain Question AnsweringShort Findings
  954. Improving generalization in large langue model by learning prefix subspacesLong Findings
  955. Improving the Robustness of Summarization Models by Detecting and Removing Input NoiseLong Findings
  956. Improving word mover's distance by leveraging self-attention matrixLong Findings
  957. In What Languages are Generative Language Models the Most Formal? Analyzing Formality Distribution across LanguagesLong Findings
  958. In-Context Demonstration Selection with Cross Entropy DifferenceLong Findings
  959. In-Context Learning Creates Task VectorsShort Findings
  960. In-Image Neural Machine Translation with Segmented Pixel Sequence-to-Sequence ModelLong Findings
  961. In-context Learning for Few-shot Multimodal Named Entity RecognitionLong Findings
  962. Incorporating Object-Level Visual Context for Multimodal Fine-Grained Entity TypingLong Findings
  963. Incorporating Probing Signals into Multimodal Machine Translation via Visual Question-Answering PairsLong Findings
  964. Incorporating Structured Representations into Pretrained Vision \& Language Models Using Scene GraphsLong Main
  965. Incorporating Syntactic Knowledge into Pre-trained Language Model using Optimization for Overcoming Catastrophic ForgettingLong Findings
  966. Incorporating Worker Perspectives into MTurk Annotation Practices for NLPLong Main
  967. Increasing Coverage and Precision of Textual Information in Multilingual Knowledge GraphsLong Main
  968. Increasing Probability Mass on Answer Choices Does Not Always Improve AccuracyLong Main
  969. IndiSocialFT: Multilingual Word Representation for Indian languages in code-mixed environmentShort Findings
  970. Indicative Summarization of Long DiscussionsLong Main
  971. Inductive Relation Inference of Knowledge Graph Enhanced by Ontology InformationLong Findings
  972. Inference-Time Policy Adapters (IPA): Tailoring Extreme-Scale LMs without Fine-tuningLong Main
  973. Influence Scores at Scale for Efficient Language Data SamplingLong Main
  974. InfoCL: Alleviating Catastrophic Forgetting in Continual Text Classification from An Information Theoretic PerspectiveLong Findings
  975. InfoDiffusion: Information Entropy Aware Diffusion Process for Non-Autoregressive Text GenerationLong Findings
  976. Information Extraction from Legal Wills: How Well Does GPT-4 Do?Short Findings
  977. Information Value: Measuring Utterance Predictability as Distance from Plausible AlternativesLong Main
  978. InheritSumm: A General, Versatile and Compact Summarizer by Distilling from GPTLong Findings
  979. Injecting structural hints: Using language models to study inductive biases in language learningLong Findings
  980. InstOptima: Evolutionary Multi-objective Instruction Optimization via Large Language Model-based Instruction OperatorsShort Findings
  981. Instances and Labels: Hierarchy-aware Joint Supervised Contrastive Learning for Hierarchical Multi-Label Text ClassificationLong Findings
  982. Instruct and Extract: Instruction Tuning for On-Demand Information ExtractionLong Main
  983. InstructExcel: A Benchmark for Natural Language Instruction in ExcelLong Findings
  984. InstructSafety: A Unified Framework for Building Multidimensional and Explainable Safety Detector through Instruction TuningLong Findings
  985. Instructed Language Models with Retrievers Are Powerful Entity LinkersLong Main
  986. Instructive Dialogue Summarization with Query AggregationsLong Main
  987. InstructoR: Instructing Unsupervised Conversational Dense Retrieval with Large Language ModelsLong Findings
  988. InteMATs: Integrating Granularity-Specific Multilingual Adapters for Cross-Lingual TransferLong Findings
  989. Integrating Language Models into Direct Speech Translation: An Inference-Time Solution to Control Gender InflectionShort Main
  990. IntenDD: A Unified Contrastive Learning Approach for Intent Detection and DiscoveryLong Findings
  991. InterFair: Debiasing with Natural Language Feedback for Fair Interpretable PredictionsShort Main
  992. Interactive Text GenerationLong Main
  993. Interpreting Answers to Yes-No Questions in User-Generated ContentLong Findings
  994. Interpreting Embedding Spaces by ConceptualizationLong Main
  995. Interpreting Indirect Answers to Yes-No Questions in Multiple LanguagesLong Findings
  996. InterroLang: Exploring NLP Models and Datasets through Dialogue-based ExplanationsLong Findings
  997. Intersectional Stereotypes in Large Language Models: Dataset and AnalysisShort Findings
  998. Intervention-Based Alignment of Code Search with Execution FeedbackLong Findings
  999. Interventional RationalizationLong Main
  1000. Interview Evaluation: A Novel Approach for Automatic Evaluation of Conversational Question Answering ModelsLong Main

Looking for submission deadlines instead? See the conference deadline calendar.