← All conferences

EMNLP 2023 Accepted Papers

The full list of 2,009 papers accepted at EMNLP 2023 (Conference on Empirical Methods in Natural Language Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

Long Main: 862Long Findings: 818Short Findings: 197Short Main: 132
  1. Intra-Event and Inter-Event Dependency-Aware Graph Network for Event Argument ExtractionLong Findings
  2. Introducing Rhetorical Parallelism Detection: A New Task with Datasets, Metrics, and BaselinesLong Main
  3. Intuitive Multilingual Audio-Visual Speech Recognition with a Single-Trained ModelShort Findings
  4. InvGC: Robust Cross-Modal Retrieval by Inverse Graph ConvolutionLong Findings
  5. Inverse Reinforcement Learning for Text SummarizationLong Findings
  6. Inverse Scaling Can Become U-ShapedShort Main
  7. Investigating Bias in Multilingual Language Models: Cross-Lingual Transfer of Debiasing TechniquesShort Main
  8. Investigating Efficiently Extending Transformers for Long Input SummarizationLong Main
  9. Investigating Multilingual Coreference Resolution by Universal AnnotationsLong Findings
  10. Investigating Online Community Engagement through StancetakingLong Findings
  11. Investigating the Effect of Pre-finetuning BERT Models on NLI Involving PresuppositionsLong Findings
  12. Investigating the Effectiveness of Multiple Expert Models CollaborationShort Findings
  13. Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking AgentsLong Main
  14. Is ChatGPT a Financial Expert? Evaluating Language Models on Financial Natural Language ProcessingShort Findings
  15. Is ChatGPT a General-Purpose Natural Language Processing Task Solver?Long Main
  16. Is ChatGPT a Good Causal Reasoner? A Comprehensive EvaluationLong Findings
  17. Is ChatGPT a Good Multi-Party Conversation Solver?Long Findings
  18. Is ChatGPT the ultimate Data Augmentation Algorithm?Short Findings
  19. Is Explanation the Cure? Misinformation Mitigation in the Short Term and Long TermShort Findings
  20. Is GPT-4 a Good Data Analyst?Long Findings
  21. Is Probing All You Need? Indicator Tasks as an Alternative to Probing Embedding SpacesLong Findings
  22. Is Robustness Transferable across Languages in Multilingual Neural Machine Translation?Long Findings
  23. Is a Prestigious Job the same as a Prestigious Country? A Case Study on Multilingual Sentence Embeddings and European CountriesShort Findings
  24. Is the Answer in the Text? Challenging ChatGPT with Evidence Retrieval from Instructive TextShort Findings
  25. Isotropic Representation Can Improve Zero-Shot Cross-Lingual Transfer on Multilingual Language ModelsLong Findings
  26. Isotropy-Enhanced Conditional Masked Language ModelsLong Findings
  27. It Ain't Over: A Multi-aspect Diverse Math Word Problem DatasetLong Main
  28. JASMINE: Arabic GPT Models for Few-Shot LearningLong Main
  29. JWSign: A Highly Multilingual Corpus of Bible Translations for more Diversity in Sign Language ProcessingLong Findings
  30. Joint Entity and Relation Extraction with Span Pruning and Hypergraph Neural NetworksLong Main
  31. Joint Geometrical and Statistical Domain Adaptation for Cross-domain Code Vulnerability DetectionLong Main
  32. Joint Semantic and Strategy Matching for Persuasive DialogueLong Findings
  33. JointMatch: A Unified Approach for Diverse and Collaborative Pseudo-Labeling to Semi-Supervised Text ClassificationLong Main
  34. Just Adjust One Prompt: Enhancing In-Context Dialogue Scoring via Constructing the Optimal Subgraph of Demonstrations and PromptsLong Main
  35. Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human FeedbackShort Main
  36. K-HATERS: A Hate Speech Detection Corpus in Korean with Target-Specific RatingsLong Findings
  37. KAPALM: Knowledge grAPh enhAnced Language Models for Fake News DetectionLong Findings
  38. KBioXLM: A Knowledge-anchored Biomedical Multilingual Pretrained Language ModelLong Findings
  39. KCTS: Knowledge-Constrained Tree Search Decoding with Token-Level Hallucination DetectionLong Main
  40. KEBAP: Korean Error Explainable Benchmark Dataset for ASR and Post-processingLong Main
  41. KEPL: Knowledge Enhanced Prompt Learning for Chinese Hypernym-Hyponym ExtractionLong Main
  42. KEPLET: Knowledge-Enhanced Pretrained Language Model with Topic Entity AwarenessLong Findings
  43. KG-GPT: A General Framework for Reasoning on Knowledge Graphs Using Large Language ModelsShort Findings
  44. KICGPT: Large Language Model with Knowledge in Context for Knowledge Graph CompletionLong Findings
  45. KRLS: Improving End-to-End Response Generation in Task Oriented Dialog with Reinforced Keywords LearningLong Main
  46. KeFVP: Knowledge-enhanced Financial Volatility PredictionLong Findings
  47. Knowledge Corpus Error in Question AnsweringShort Findings
  48. Knowledge Distillation ≈ Label Smoothing: Fact or Fallacy?Short Main
  49. Knowledge Graph Compression Enhances Diverse Commonsense GenerationLong Main
  50. Knowledge Rumination for Pre-trained Language ModelsLong Main
  51. Knowledge is a Region in Weight Space for Fine-tuned Language ModelsLong Findings
  52. Knowledge-Augmented Language Model VerificationLong Main
  53. Knowledge-Selective Pretraining for Attribute Value ExtractionLong Findings
  54. LACMA: Language-Aligning Contrastive Learning with Meta-Actions for Embodied Instruction FollowingLong Main
  55. LATENTLOGIC: Learning Logic Rules in Latent Space over Knowledge GraphsShort Findings
  56. LDM$^2$: A Large Decision Model Imitating Human Cognition with Dynamic Memory EnhancementLong Findings
  57. LEGO: A Multi-agent Collaborative Framework with Role-playing and Iterative Feedback for Causality Explanation GenerationLong Findings
  58. LEXTREME: A Multi-Lingual and Multi-Task Benchmark for the Legal DomainLong Findings
  59. LIMIT: Language Identification, Misidentification, and Translation using Hierarchical Models in 350+ LanguagesLong Main
  60. LINC: A Neurosymbolic Approach for Logical Reasoning by Combining Language Models with First-Order Logic ProversLong Main
  61. LLM aided semi-supervision for efficient Extractive Dialog SummarizationShort Findings
  62. LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language ModelsLong Main
  63. LLM-FP4: 4-Bit Floating-Point Quantized TransformersLong Main
  64. LLM-enhanced Self-training for Cross-domain Constituency ParsingLong Main
  65. LLM-in-the-loop: Leveraging Large Language Model for Thematic AnalysisShort Findings
  66. LLM-powered Data Augmentation for Enhanced Cross-lingual PerformanceLong Main
  67. LLMDet: A Third Party Large Language Models Generated Text Detection ToolLong Findings
  68. LLMLingua: Compressing Prompts for Accelerated Inference of Large Language ModelsLong Main
  69. LLMaAA: Making Large Language Models as Active AnnotatorsLong Findings
  70. LLMs -- the Good, the Bad or the Indispensable?: A Use Case on Legal Statute Prediction and Legal Judgment Prediction on Indian Court CasesShort Findings
  71. LM vs LM: Detecting Factual Errors via Cross ExaminationLong Main
  72. LMGQS: A Large-scale Dataset for Query-focused SummarizationLong Findings
  73. Label Words are Anchors: An Information Flow Perspective for Understanding In-Context LearningLong Main
  74. Language Guided Visual Question Answering: Elevate Your Multimodal Language Model Using Knowledge-Enriched PromptsShort Findings
  75. Language Model Quality Correlates with Psychometric Predictive Power in Multiple LanguagesShort Main
  76. Language Model is Suitable for Correction of Handwritten Mathematical Expressions RecognitionLong Main
  77. Language Models with RationalityLong Main
  78. Language Representation Projection: Can We Transfer Factual Knowledge across Languages in Multilingual Language Models?Short Main
  79. Language and Mental Health: Measures of Emotion Dynamics from Text as Linguistic Biosocial MarkersLong Main
  80. Language-Agnostic Bias Detection in Language Models with Bias ProbingLong Findings
  81. Large Language Model Is Not a Good Few-shot Information Extractor, but a Good Reranker for Hard Samples!Long Findings
  82. Large Language Models Are Better Adversaries: Exploring Generative Clean-Label Backdoor Attacks Against Text ClassifiersLong Findings
  83. Large Language Models Can Self-ImproveLong Main
  84. Large Language Models Know Your Contextual Search Intent: A Prompting Framework for Conversational SearchLong Findings
  85. Large Language Models Meet Harry Potter: A Dataset for Aligning Dialogue Agents with CharactersLong Findings
  86. Large Language Models Meet Open-World Intent Discovery and Recognition: An Evaluation of ChatGPTLong Main
  87. Large Language Models Only Pass Primary School Exams in Indonesia: A Comprehensive Test on IndoMMLULong Main
  88. Large Language Models and Multimodal Retrieval for Visual Word Sense DisambiguationLong Main
  89. Large Language Models are Better Reasoners with Self-VerificationLong Findings
  90. Large Language Models are Complex Table ParsersLong Main
  91. Large Language Models are Not Yet Human-Level Evaluators for Abstractive SummarizationLong Findings
  92. Large Language Models are Temporal and Causal Reasoners for Video Question AnsweringLong Main
  93. Large Language Models are biased to overestimate profoundnessShort Main
  94. Large Language Models as Source Planner for Personalized Knowledge-grounded DialoguesLong Findings
  95. Large Language Models: The Need for Nuance in Current Debates and a Pragmatic Perspective on UnderstandingLong Main
  96. Large-Scale and Multi-Perspective Opinion Summarization with Diverse Review SubsetsLong Findings
  97. Large-scale similarity search with Optimal TransportShort Main
  98. Larger Probes Tell a Different Story: Extending Psycholinguistic Datasets Via In-Context LearningShort Main
  99. Late Fusion of Transformers for Sentiment Analysis of Code-Switched DataShort Findings
  100. LayoutDIT: Layout-Aware End-to-End Document Image Translation with Multi-Step Conductive DecoderLong Findings
  101. Lazy-k Decoding: Constrained Decoding for Information ExtractionLong Main
  102. Leap-of-Thought: Accelerating Transformers via Dynamic Token RoutingLong Main
  103. Learn From One Specialized Sub-Teacher: One-to-One Mapping for Feature-Based Knowledge DistillationLong Findings
  104. Learn Your Tokens: Word-Pooled Tokenization for Language ModelingLong Findings
  105. Learn and Consolidate: Continual Adaptation for Zero-Shot and Multilingual Neural Machine TranslationLong Main
  106. Learning Co-Speech Gesture for Multimodal Aphasia Type DetectionLong Main
  107. Learning Dynamic Representations for Discourse Dependency ParsingLong Findings
  108. Learning Easily Updated General Purpose Text Representations with Adaptable Task-Specific PrefixShort Findings
  109. Learning Interpretable Style Embeddings via Prompting LLMsLong Findings
  110. Learning Knowledge-Enhanced Contextual Language Representations for Domain Natural Language UnderstandingLong Main
  111. Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment AnalysisLong Main
  112. Learning Preference Model for LLMs via Automatic Preference Data GenerationLong Main
  113. Learning Retrieval Augmentation for Personalized Dialogue GenerationLong Main
  114. Learning Semantic Role Labeling from Compatible Label SequencesLong Findings
  115. Learning from Mistakes via Cooperative Study Assistant for Large Language ModelsLong Main
  116. Learning the Visualness of Text Using Large Vision-Language ModelsLong Main
  117. Learning to Abstract with Nonparametric Variational Information BottleneckShort Findings
  118. Learning to Compose Representations of Different Encoder Layers towards Improving Compositional GeneralizationLong Findings
  119. Learning to Correct Noisy Labels for Fine-Grained Entity Typing via Co-Prediction Prompt TuningLong Findings
  120. Learning to Describe for Predicting Zero-shot Drug-Drug InteractionsLong Main
  121. Learning to Follow Object-Centric Image Editing Instructions FaithfullyLong Findings
  122. Learning to Predict Task Transferability via Soft PromptLong Main
  123. Learning to Rank Context for Named Entity Recognition Using a Synthetic DatasetLong Main
  124. Learning to Rank Generation with Pairwise Partial RewardsLong Main
  125. Learning to love diligent trolls: Accounting for rater effects in the dialogue safety taskShort Findings
  126. Learning under Label Proportions for Text ClassificationLong Findings
  127. Legally Enforceable Hate Speech Detection for Public ForumsLong Findings
  128. Length is a Curse and a Blessing for Document-level SemanticsLong Main
  129. Length-Adaptive Distillation: Customizing Small Language Model for Dynamic Token PruningLong Findings
  130. Less than One-shot: Named Entity Recognition via Extremely Weak SupervisionLong Findings
  131. Let GPT be a Math Tutor: Teaching Math Word Problem Solvers with Customized Exercise GenerationLong Main
  132. Let's Sample Step by Step: Adaptive-Consistency for Efficient Reasoning and Coding with LLMsLong Main
  133. Let's Synthesize Step by Step: Iterative Dataset Synthesis with Large Language Models by Extrapolating Errors from Small ModelsLong Findings
  134. Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-ThoughtLong Main
  135. Leveraging Contrastive Learning and Knowledge Distillation for Incomplete Modality Rumor DetectionLong Findings
  136. Leveraging GPT-4 for Automatic Translation Post-EditingLong Findings
  137. Leveraging Multiple Teachers for Test-Time Adaptation of Language-Guided ClassifiersLong Findings
  138. Leveraging Structured Information for Explainable Multi-hop Question Answering and ReasoningLong Findings
  139. Lexical Repetitions Lead to Rote Learning: Unveiling the Impact of Lexical Overlap in Train and Test Reference SummariesLong Findings
  140. Licon: A Diverse, Controllable and Challenging Linguistic Concept Learning BenchmarkLong Findings
  141. Lifelong Sequence Generation with Dynamic Module Expansion and AdaptationLong Main
  142. Linear-Time Modeling of Linguistic Structure: An Order-Theoretic PerspectiveLong Main
  143. Ling-CL: Understanding NLP Models through Linguistic CurriculaLong Main
  144. Linguistic Compression in Single-Sentence Human-Written SummariesLong Findings
  145. Linguistically Motivated Sign Language SegmentationLong Findings
  146. Linking Surface Facts to Large-Scale Knowledge GraphsLong Main
  147. Lion: Adversarial Distillation of Proprietary Large Language ModelsLong Main
  148. Localizing Active Objects from Egocentric Vision with Symbolic World KnowledgeLong Main
  149. Locally Differentially Private Document Generation Using Zero Shot PromptingLong Findings
  150. Location-Aware Visual Question Generation with Lightweight ModelsLong Main
  151. Log-FGAER: Logic-Guided Fine-Grained Address Entity Recognition from Multi-Turn Spoken DialogueLong Main
  152. LogiCoT: Logical Chain-of-Thought Instruction TuningLong Findings
  153. Logic Unveils Truth, While Disguise Obscures It: Transition Logic Augmented Response Selection for Multi-Turn DialogueLong Findings
  154. Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical ReasoningLong Findings
  155. LogicAttack: Adversarial Attacks for Evaluating Logical Consistency of Natural Language InferenceShort Findings
  156. Long-Horizon Dialogue Understanding for Role Identification in the Game of Avalon with Large Language ModelsLong Findings
  157. Long-Range Language Modeling with Selective CacheLong Findings
  158. Longtriever: a Pre-trained Long Text Encoder for Dense Document RetrievalLong Main
  159. Look-back Decoding for Open-Ended Text GenerationLong Main
  160. Loose lips sink ships: Mitigating Length Bias in Reinforcement Learning from Human FeedbackLong Findings
  161. Low-Resource Comparative Opinion Quintuple Extraction by Data Augmentation with PromptingShort Findings
  162. M$^3$Seg: A Maximum-Minimum Mutual Information Paradigm for Unsupervised Topic Segmentation in ASR TranscriptsShort Main
  163. M2C: Towards Automatic Multimodal Manga ComplementShort Findings
  164. M2DF: Multi-grained Multi-curriculum Denoising Framework for Multimodal Aspect-based Sentiment AnalysisLong Main
  165. MADNet: Maximizing Addressee Deduction Expectation for Multi-Party Conversation GenerationLong Main
  166. MAF: Multi-Aspect Feedback for Improving Reasoning in Large Language ModelsLong Main
  167. MAGNIFICo: Evaluating the In-Context Learning Ability of Large Language Models to Generalize to Novel InterpretationsLong Main
  168. MAPO: Boosting Large Language Model Performance with Model-Adaptive Prompt OptimizationLong Findings
  169. MCC-KD: Multi-CoT Consistent Knowledge DistillationLong Findings
  170. MCLF: A Multi-grained Contrastive Learning Framework for ASR-robust Spoken Language UnderstandingLong Findings
  171. MEEP: Is this Engaging? Prompting Large Language Models for Dialogue Evaluation in Multilingual SettingsLong Findings
  172. MEGA: Multilingual Evaluation of Generative AILong Main
  173. MEGClass: Extremely Weakly Supervised Text Classification via Mutually-Enhancing Text GranularitiesLong Findings
  174. MILDSum: A Novel Benchmark Dataset for Multilingual Summarization of Indian Legal Case JudgmentsShort Main
  175. MISCA: A Joint Model for Multiple Intent Detection and Slot Filling with Intent-Slot Co-AttentionLong Findings
  176. MM-Reasoner: A Multi-Modal Knowledge-Aware Framework for Knowledge-Based Visual Question AnsweringLong Findings
  177. MMNMT: Modularizing Multilingual Neural Machine Translation with Flexibly Assembled MoE and Dense BlocksLong Main
  178. MPrompt: Exploring Multi-level Prompt Tuning for Machine Reading ComprehensionLong Findings
  179. MProto: Multi-Prototype Network with Denoised Optimal Transport for Distantly Supervised Named Entity RecognitionLong Main
  180. MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop QuestionsLong Main
  181. MRRL: Modifying the Reference via Reinforcement Learning for Non-Autoregressive Joint Multiple Intent Detection and Slot FillingLong Findings
  182. MSCFFN: A New FFN with Multi-Space Cross to Accelerate TransformerShort Findings
  183. MT2: Towards a Multi-Task Machine Translation Model with Translation-Specific In-Context LearningLong Main
  184. MTGER: Multi-view Temporal Graph Enhanced Temporal Reasoning over Time-Involved DocumentLong Findings
  185. MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection BenchmarkLong Main
  186. MUX-PLMs: Data Multiplexing for High-throughput Language ModelsLong Findings
  187. MaNtLE: Model-agnostic Natural Language ExplainerLong Main
  188. MaXM: Towards Multilingual Visual Question AnsweringLong Findings
  189. MacLaSa: Multi-Aspect Controllable Text Generation via Efficient Sampling from Compact Latent SpaceLong Findings
  190. Macedon: Minimizing Representation Coding Rate Reduction for Cross-Lingual Natural Language UnderstandingLong Findings
  191. Machine Reading Comprehension using Case-based ReasoningLong Findings
  192. MailEx: Email Event and Argument ExtractionLong Main
  193. Make Every Example Count: On the Stability and Utility of Self-Influence for Learning from Noisy NLP DatasetsLong Main
  194. Make Your Decision Convincing! A Unified Two-Stage Framework: Self-Attribution and Decision-MakingLong Findings
  195. Making Body Movement in Sign Language Corpus Accessible for Linguists and Machines with Three-Dimensional Normalization of MediaPipeLong Findings
  196. Making Large Language Models Better Data CreatorsLong Main
  197. Mandarin classifier systems optimize to accommodate communicative pressuresLong Findings
  198. Manifold-Preserving Transformers are Effective for Short-Long Range EncodingLong Findings
  199. Manipulating the Perceived Personality Traits of Language ModelsLong Findings
  200. MarkQA: A large scale KBQA dataset with numerical reasoningLong Main
  201. Masked Path Modeling for Vision-and-Language NavigationLong Findings
  202. MathDial: A Dialogue Tutoring Dataset with Rich Pedagogical Properties Grounded in Math Reasoning ProblemsLong Findings
  203. MeaeQ: Mount Model Extraction Attacks with Efficient QueriesLong Main
  204. Measure Children's Mindreading Ability with Machine ReadingLong Findings
  205. Measuring Faithful and Plausible Visual Grounding in VQALong Findings
  206. Measuring Pointwise $\mathcal{V}$-Usable Information In-Context-lyLong Findings
  207. Measuring and Mitigating Constraint Violations of In-Context Learning for Utterance-to-API Semantic ParsingLong Findings
  208. Measuring and Narrowing the Compositionality Gap in Language ModelsLong Findings
  209. Measuring bias in Instruction-Following models with P-ATLong Findings
  210. Measuring the Knowledge Acquisition-Utilization Gap in Pretrained Language ModelsLong Findings
  211. MedEval: A Multi-Level, Multi-Task, and Multi-Domain Medical Benchmark for Language Model EvaluationLong Main
  212. MediaHG: Rethinking Eye-catchy Features in Social Media Headline GenerationLong Main
  213. Medical Text Simplification: Optimizing for Readability with Unlikelihood Training and Reranked Beam Search DecodingShort Findings
  214. MemeCap: A Dataset for Captioning and Interpreting MemesLong Main
  215. Memorisation Cartography: Mapping out the Memorisation-Generalisation Continuum in Neural Machine TranslationLong Main
  216. Memory-Based Invariance Learning for Out-of-Domain Text ClassificationLong Main
  217. MenatQA: A New Dataset for Testing the Temporal Comprehension and Reasoning Abilities of Large Language ModelsLong Findings
  218. Merging Experts into One: Improving Computational Efficiency of Mixture of ExpertsShort Main
  219. Merging Generated and Retrieved Knowledge for Open-Domain QALong Main
  220. Meta-Learning Online Adaptation of Language ModelsLong Main
  221. Meta-Learning of Prompt Generation for Lightweight Prompt Engineering on Language-Model-as-a-ServiceLong Findings
  222. MetaReVision: Meta-Learning with Retrieval for Visually Grounded Compositional Concept AcquisitionLong Findings
  223. Methodological Insights in Detecting Subtle Semantic Shifts with Contextualized and Static Language ModelsLong Findings
  224. Mind the Gap Between Conversations for Improved Long-Term Dialogue GenerationLong Findings
  225. Mind the Gap: Automated Corpus Creation for Enthymeme Detection and Reconstruction in Learner ArgumentsLong Findings
  226. MindGames: Targeting Theory of Mind in Large Language Models with Dynamic Epistemic Modal LogicShort Findings
  227. MingOfficial: A Ming Official Career Dataset and a Historical Context-Aware Representation Learning FrameworkLong Main
  228. Miracle: Towards Personalized Dialogue Generation with Latent-Space Multiple Personal Attribute ControlLong Findings
  229. Mirages. On Anthropomorphism in Dialogue SystemsLong Main
  230. Mirror: A Universal Framework for Various Information Extraction TasksLong Main
  231. Misery Loves Complexity: Exploring Linguistic Complexity in the Context of Emotion DetectionLong Findings
  232. Mitigating Backdoor Poisoning Attacks through the Lens of Spurious CorrelationLong Main
  233. Mitigating Biases in Hate Speech Detection from A Causal PerspectiveLong Findings
  234. Mitigating Data Imbalance and Representation Degeneration in Multilingual Machine TranslationLong Findings
  235. Mitigating Framing Bias with Polarity Minimization LossShort Findings
  236. Mitigating Intrinsic Named Entity-Related Hallucinations of Abstractive Text SummarizationLong Findings
  237. Mitigating Temporal Misalignment by Discarding Outdated FactsLong Main
  238. MixEdit: Revisiting Data Augmentation and Beyond for Grammatical Error CorrectionLong Findings
  239. MixTEA: Semi-supervised Entity Alignment with Mixture TeachingLong Findings
  240. Mixture of Soft Prompts for Controllable Data GenerationLong Findings
  241. Mixture-of-Linguistic-Experts Adapters for Improving and Interpreting Pre-trained Language ModelsLong Findings
  242. MoPe: Model Perturbation based Privacy Attacks on Language ModelsLong Main
  243. MoT: Memory-of-Thought Enables ChatGPT to Self-ImproveLong Main
  244. Model-tuning Via Prompts Makes NLP Models Adversarially RobustLong Main
  245. Modeling Conceptual Attribute Likeness and Domain Inconsistency for Metaphor DetectionLong Main
  246. Modeling Empathic Similarity in Personal NarrativesLong Main
  247. Modeling Highlighting of Metaphors in Multitask Contrastive Learning ParadigmsLong Findings
  248. Modeling Legal Reasoning: LM Annotation at the Edge of Human AgreementLong Main
  249. Models See Hallucinations: Evaluating the Factuality in Video CaptioningLong Main
  250. MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal AdapterLong Main
  251. Monte Carlo Thought Search: Large Language Model Querying for Complex Scientific Reasoning in Catalyst DesignShort Findings
  252. MoqaGPT : Zero-Shot Multi-modal Open-domain Question Answering with Large Language ModelLong Findings
  253. More than Votes? Voting and Language based Partisanship in the US Supreme CourtShort Findings
  254. MuG: A Multimodal Classification Benchmark on Game Data with Tabular, Textual, and Visual FieldsLong Findings
  255. Mulan: A Multi-Level Alignment Model for Video Question AnsweringLong Findings
  256. Multi-Defendant Legal Judgment Prediction via Hierarchical ReasoningLong Findings
  257. Multi-Granularity Information Interaction Framework for Incomplete Utterance RewritingShort Findings
  258. Multi-Modal Knowledge Graph Transformer Framework for Multi-Modal Entity AlignmentLong Findings
  259. Multi-Source Multi-Type Knowledge Exploration and Exploitation for Dialogue GenerationLong Main
  260. Multi-Source Probing for Open-Domain Conversational UnderstandingLong Main
  261. Multi-Stage Pre-training Enhanced by ChatGPT for Multi-Scenario Multi-Domain Dialogue SummarizationLong Findings
  262. Multi-Task Knowledge Distillation with Embedding Constraints for Scholarly Keyphrase Boundary ClassificationLong Main
  263. Multi-Task Learning of Query Generation and Classification for Generative Conversational Question RewritingLong Findings
  264. Multi-User MultiWOZ: Task-Oriented Dialogues among Multiple UsersLong Findings
  265. Multi-label and Multi-target Sampling of Machine Annotation for Computational Stance DetectionShort Findings
  266. Multi-level Adaptive Contrastive Learning for Knowledge Internalization in Dialogue GenerationLong Main
  267. Multi-level Contrastive Learning for Script-based Character UnderstandingLong Main
  268. Multi-step Jailbreaking Privacy Attacks on ChatGPTLong Findings
  269. Multi-view Contrastive Learning for Entity Typing over Knowledge GraphsLong Main
  270. MultiCMET: A Novel Chinese Benchmark for Understanding Multimodal MetaphorLong Findings
  271. MultiCoNER v2: a Large Multilingual dataset for Fine-grained and Noisy Named Entity RecognitionShort Findings
  272. MultiTurnCleanup: A Benchmark for Multi-Turn Spoken Conversational Transcript CleanupShort Main
  273. Multilingual Coarse Political Stance Classification of Media. The Editorial Line of a ChatGPT and Bard NewspaperShort Findings
  274. Multilingual Generation and Answering of Questions from Texts and Knowledge GraphsLong Findings
  275. Multilingual Holistic Bias: Extending Descriptors and Patterns to Unveil Demographic Biases in Languages at ScaleLong Main
  276. Multilingual Large Language Models Are Not (Yet) Code-SwitchersLong Main
  277. Multilingual Lottery Tickets to Pretrain Language ModelsLong Findings
  278. Multilingual Pixel Representations for Translation and Effective Cross-lingual TransferLong Main
  279. Multilingual Simplification of Medical TextsLong Main
  280. Multilingual estimation of political-party positioning: From label aggregation to long-input TransformersLong Main
  281. Multimodal Automated Fact-Checking: A SurveyLong Findings
  282. Multimodal Embodied Plan Prediction Augmented with Synthetic Embodied DialogueLong Main
  283. Multitask Multimodal Prompted Training for Interactive Embodied Task CompletionLong Main
  284. Multiview Clickbait Detection via Jointly Modeling Subjective and Objective PreferenceLong Findings
  285. NAIL: Lexical Retrieval Indices with Efficient Non-Autoregressive DecodersLong Main
  286. NASH: A Simple Unified Framework of Structured Pruning for Accelerating Encoder-Decoder Language ModelsLong Findings
  287. NERetrieve: Dataset for Next Generation Named Entity Recognition and RetrievalLong Findings
  288. NERvous About My Health: Constructing a Bengali Medical Named Entity Recognition DatasetShort Findings
  289. NEWTON: Are Large Language Models Capable of Physical Reasoning?Long Findings
  290. NLI4CT: Multi-Evidence Natural Language Inference for Clinical Trial ReportsLong Main
  291. NLMs: Augmenting Negation in Language ModelsLong Findings
  292. NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each BenchmarkShort Findings
  293. NORMSAGE: Multi-Lingual Multi-Cultural Norm Discovery from Conversations On-the-FlyLong Main
  294. NameGuess: Column Name Expansion for Tabular DataLong Main
  295. Narrative Order Aware Story Generation via Bidirectional Pretraining Model with Optimal Transport RewardLong Findings
  296. Narrative Style and the Spread of Health Misinformation on TwitterLong Findings
  297. NarrativeXL: a Large-scale Dataset for Long-Term Memory ModelsLong Findings
  298. Natural Disaster Tweets Classification Using Multimodal DataLong Main
  299. Natural Language Annotations for Reasoning about Program SemanticsShort Findings
  300. Natural Language Decompositions of Implicit Content Enable Better Text RepresentationsLong Main
  301. Natural Response Generation for Chinese Reading ComprehensionLong Findings
  302. Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language ModelsLong Main
  303. Nearest Neighbor Machine Translation is Meta-Optimizer on Output Projection LayerLong Main
  304. NeuSTIP: A Neuro-Symbolic Model for Link and Time Prediction in Temporal Knowledge GraphsLong Main
  305. Neuro-Symbolic Sentiment Analysis with Dynamic Word Sense DisambiguationLong Findings
  306. New Datasets and Controllable Iterative Data Augmentation Method for Code-switching ASR Error CorrectionLong Findings
  307. No offence, Bert - I insult only humans! Multilingual sentence-level attack on toxicity detection networksShort Findings
  308. Noise-Robust Fine-Tuning of Pretrained Language Models via External GuidanceLong Findings
  309. Noise-Robust Semi-Supervised Learning for Distantly Supervised Relation ExtractionLong Findings
  310. Noisy Exemplars Make Large Language Models More Robust: A Domain-Agnostic Behavioral AnalysisShort Main
  311. Noisy Pair Corrector for Dense RetrievalLong Findings
  312. Noisy Self-Training with Synthetic Queries for Dense RetrievalLong Findings
  313. Non-Autoregressive Document-Level Machine TranslationLong Findings
  314. Non-Autoregressive Math Word Problem Solver with Unified Tree StructureLong Main
  315. Non-Autoregressive Sentence OrderingLong Findings
  316. Non-Compositionality in Sentiment: New Data and AnalysesShort Findings
  317. Non-Programmers Can Label Programs Indirectly via Active Examples: A Case Study with Text-to-SQLLong Main
  318. Non-autoregressive Streaming Transformer for Simultaneous TranslationLong Main
  319. Non-autoregressive Text Editing with Copy-aware Latent AlignmentsLong Main
  320. Non-compositional Expression Generation Based on Curriculum Learning and Continual LearningLong Findings
  321. Non-parallel Accent Transfer based on Fine-grained Controllable Accent ModellingLong Findings
  322. Norm of Word Embedding Encodes Information GainLong Main
  323. NormDial: A Comparable Bilingual Synthetic Dialog Dataset for Modeling Social Norm Adherence and ViolationShort Main
  324. Normal-Abnormal Decoupling Memory for Medical Report GenerationLong Findings
  325. Not All Demonstration Examples are Equally Beneficial: Reweighting Demonstration Examples for In-Context LearningLong Findings
  326. Not All Languages Are Created Equal in LLMs: Improving Multilingual Capability by Cross-Lingual-Thought PromptingLong Findings
  327. Not all quantifiers are equal: Probing Transformer-based language models' understanding of generalised quantifiersLong Main
  328. NovaCOMET: Open Commonsense Foundation Models with Symbolic Knowledge DistillationLong Findings
  329. Novel Relation Detection: Discovering Unknown Relation Types via Multi-Strategy Self-Supervised LearningLong Findings
  330. Novel Slot Detection With an Incremental SettingLong Findings
  331. ORCHID: A Chinese Debate Corpus for Target-Independent Stance Detection and Argumentative Dialogue SummarizationLong Main
  332. On Bilingual Lexicon Induction with Large Language ModelsLong Main
  333. On Evaluation of Bangla Word AnalogiesShort Main
  334. On Event Individuation for Document-Level Information ExtractionShort Findings
  335. On General Language UnderstandingShort Findings
  336. On Robustness of Finetuned Transformer-based NLP ModelsLong Findings
  337. On Surgical Fine-tuning for Language EncodersShort Findings
  338. On Task-personalized Multimodal Few-shot Learning for Visually-rich Document Entity RetrievalLong Findings
  339. On Uncertainty Calibration and Selective Generation in Probabilistic Neural Summarization: A Benchmark StudyShort Findings
  340. On the Automatic Generation and Simplification of Children's StoriesLong Main
  341. On the Benefits of Learning to Route in Mixture-of-Experts ModelsLong Main
  342. On the Calibration of Large Language Models and AlignmentLong Findings
  343. On the Challenges of Using Black-Box APIs for Toxicity Evaluation in ResearchLong Main
  344. On the Dimensionality of Sentence EmbeddingsLong Findings
  345. On the Impact of Cross-Domain Data on German Language ModelsLong Findings
  346. On the Representational Capacity of Recurrent Neural Language ModelsLong Main
  347. On the Risk of Misinformation Pollution with Large Language ModelsLong Findings
  348. On the Transferability of Visually Grounded PCFGsLong Findings
  349. On the Zero-Shot Generalization of Machine-Generated Text DetectorsShort Findings
  350. Once Upon a ${\it Time}$ in ${\it Graph}$: Relative-Time Pretraining for Complex Temporal ReasoningLong Main
  351. Once is Enough: A Light-Weight Cross-Attention for Fast Sentence Pair ModelingShort Main
  352. One For All $\&$ All For One: Bypassing Hyperparameter Tuning with Model Averaging for Cross-Lingual TransferShort Findings
  353. One-Model-Connects-All: A Unified Graph Pre-Training Model for Online Community ModelingLong Findings
  354. Oolong: Investigating What Makes Transfer Learning Hard with Controlled StudiesShort Main
  355. Open Domain Multi-document Summarization: A Comprehensive Study of Model Brittleness under RetrievalLong Findings
  356. Open Information Extraction via ChunksLong Main
  357. Open-Ended Instructable Embodied Agents with Memory-Augmented Large Language ModelsLong Findings
  358. Open-ended Commonsense Reasoning with Unrestricted Answer CandidatesLong Findings
  359. Open-source Large Language Models are Strong Zero-shot Query Likelihood Models for Document RankingShort Findings
  360. Open-world Semi-supervised Generalized Relation Discovery Aligned in a Real-world SettingLong Main
  361. OpenAsp: A Benchmark for Multi-document Open Aspect-based SummarizationLong Main
  362. Optimized Tokenization for Transcribed Error CorrectionLong Main
  363. Optimizing Retrieval-augmented Reader Models via Token EliminationLong Main
  364. Orca: A Few-shot Benchmark for Chinese Conversational Machine Reading ComprehensionLong Findings
  365. Orthogonal Subspace Learning for Language Model Continual LearningLong Findings
  366. OssCSE: Overcoming Surface Structure Bias in Contrastive Learning for Unsupervised Sentence EmbeddingLong Main
  367. Out-of-Distribution Generalization in Natural Language Processing: Past, Present, and FutureLong Main
  368. Outlier Dimensions Encode Task Specific KnowledgeShort Main
  369. Outlier Suppression+: Accurate quantization of large language models by equivalent and effective shifting and scalingLong Main
  370. PAC-tuning: Fine-tuning Pre-trained Language Models with PAC-driven Perturbed Gradient DescentLong Main
  371. PALS: Personalized Active Learning for Subjective Tasks in NLPLong Main
  372. PARROT: Zero-Shot Narrative Reading Comprehension via Parallel ReadingLong Findings
  373. PCMID: Multi-Intent Detection through Supervised Prototypical Contrastive LearningLong Findings
  374. PEFTDebias : Capturing debiasing information using PEFTsShort Main
  375. PHD: Pixel-Based Language Modeling of Historical DocumentsLong Main
  376. PIEClass: Weakly-Supervised Text Classification with Prompting and Noise-Robust Iterative Ensemble TrainingLong Main
  377. PIVOINE: Instruction Tuning for Open-world Entity ProfilingLong Findings
  378. PMIndiaSum: Multilingual and Cross-lingual Headline Summarization for Languages in IndiaLong Findings
  379. POE: Process of Elimination for Multiple Choice ReasoningShort Main
  380. POSQA: Probe the World Models of LLMs with Size ComparisonsLong Findings
  381. PR-MCS: Perturbation Robust Metric for MultiLingual Image CaptioningLong Findings
  382. PRCA: Fitting Black-Box Large Language Models for Retrieval Question Answering via Pluggable Reward-Driven Contextual AdapterLong Main
  383. PRESTO: A Multilingual Dataset for Parsing Realistic Task-Oriented DialogsLong Main
  384. PROSE: A Pronoun Omission Solution for Chinese-English Spoken Language TranslationLong Main
  385. PROTEGE: Prompt-based Diverse Question Generation from Web ArticlesLong Findings
  386. PTP: Boosting Stability and Performance of Prompt Tuning with Perturbation-Based RegularizerLong Main
  387. PUNR: Pre-training with User Behavior Modeling for News RecommendationLong Findings
  388. PaRaDe: Passage Ranking using Demonstrations with LLMsShort Findings
  389. Parameter Efficient Multi-task Fine-tuning by Learning to Transfer Token-wise PromptsLong Findings
  390. Parameter-Efficient Cross-lingual Transfer of Vision and Language Models via Translation-based AlignmentLong Findings
  391. Parameter-Efficient Language Model Tuning with Active Learning in Low-Resource SettingsLong Main
  392. Parameter-Efficient Prompt Tuning Makes Generalized and Calibrated Neural Text RetrieversLong Findings
  393. Parameter-efficient Tuning for Large Language Model without Calculating Its GradientsLong Main
  394. Paraphrase Types for Generation and DetectionLong Main
  395. ParroT: Translating during Chat using Large Language Models tuned with Human Translation and FeedbackLong Findings
  396. Penalty Decoding: Well Suppress the Self-Reinforcement Effect in Open-Ended Text GenerationShort Main
  397. People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language DetectionLong Main
  398. Perceptual Structure in the absence of grounding: the impact of abstractedness and subjectivity in color language for LLMsShort Findings
  399. PersonaLM: Language Model Personalization via Domain-distributed Span Aggregated K-Nearest N-gram Retrieval AugmentationLong Findings
  400. Personalized Distillation: Empowering Open-Sourced LLMs with Adaptive Learning for Code GenerationLong Main
  401. PerturbScore: Connecting Discrete and Continuous Perturbations in NLPLong Findings
  402. Physician Detection of Clinical Harm in Machine Translation: Quality Estimation Aids in Reliance and Backtranslation Identifies Critical ErrorsLong Main
  403. Pit One Against Many: Leveraging Attention-head Embeddings for Parameter-efficient Multi-head AttentionLong Findings
  404. PivotFEC: Enhancing Few-shot Factual Error Correction with a Pivot Task Approach using Large Language ModelsLong Findings
  405. Plan, Verify and Switch: Integrated Reasoning with Diverse X-of-ThoughtsLong Main
  406. PlugMed: Improving Specificity in Patient-Centered Medical Dialogue Generation using In-Context LearningLong Findings
  407. Pointwise Mutual Information Based Metric and Decoding Strategy for Faithful Generation in Document Grounded DialogsLong Main
  408. Poisoning Retrieval Corpora by Injecting Adversarial PassagesShort Main
  409. Polar Ducks and Where to Find Them: Enhancing Entity Linking with Duck Typing and Polar Box EmbeddingsLong Main
  410. Polyglot or Not? Measuring Multilingual Encyclopedic Knowledge in Foundation ModelsShort Main
  411. Post-hoc Utterance Refining Method by Entity Mining for Faithful Knowledge Grounded ConversationsLong Main
  412. Practical Computational Power of Linear Transformers and Their Recurrent and Self-Referential ExtensionsShort Main
  413. Pragmatic Reasoning Unlocks Quantifier Semantics for Foundation ModelsLong Main
  414. Pragmatics in Language Grounding: Phenomena, Tasks, and Modeling ApproachesLong Findings
  415. Pre-Trained Language Models Augmented with Synthetic Scanpaths for Natural Language UnderstandingShort Main
  416. Pre-trained Speech Processing Models Contain Human-Like Biases that Propagate to Speech Emotion RecognitionLong Findings
  417. Pre-training Intent-Aware Encoders for Zero- and Few-Shot Intent ClassificationLong Main
  418. Pre-training Language Models for Comparative ReasoningLong Main
  419. Pre-training Multi-task Contrastive Learning Models for Scientific Literature UnderstandingLong Findings
  420. PreWoMe: Exploiting Presuppositions as Working Memory for Long Form Question AnsweringShort Main
  421. Precedent-Enhanced Legal Judgment Prediction with LLM and Domain-Model CollaborationLong Main
  422. Predict the Future from the Past? On the Temporal Data Distribution Shift in Financial Sentiment ClassificationsLong Main
  423. Predictive Chemistry Augmented with Text RetrievalLong Main
  424. Prefix-Tuning Based Unsupervised Text Style TransferLong Findings
  425. Preserving Knowledge Invariance: Rethinking Robustness Evaluation of Open Information ExtractionLong Main
  426. Preserving Privacy Through Dememorization: An Unlearning Technique For Mitigating Memorization Risks In Language ModelsLong Main
  427. Pretraining Language Models with Text-Attributed Heterogeneous GraphsLong Findings
  428. Primacy Effect of ChatGPTShort Main
  429. Privacy Implications of Retrieval-Based Language ModelsLong Main
  430. Probabilistic Tree-of-thought Reasoning for Answering Knowledge-intensive Complex QuestionsLong Findings
  431. Probing LLMs for Joint Encoding of Linguistic CategoriesLong Findings
  432. Probing LLMs for hate speech detection: strengths and vulnerabilitiesLong Findings
  433. Probing Representations for Document-level Event ExtractionShort Findings
  434. Probing the “Creativity” of Large Language Models: Can models produce divergent semantic association?Short Findings
  435. Program Translation via Code DistillationLong Main
  436. Promoting Topic Coherence and Inter-Document Consorts in Multi-Document Summarization via Simplicial Complex and Sheaf GraphLong Main
  437. Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language ModelsLong Main
  438. Prompt-Based Editing for Text Style TransferLong Findings
  439. Prompt-Based Monte-Carlo Tree Search for Goal-oriented Dialogue Policy PlanningShort Main
  440. Prompt-based Logical Semantics Enhancement for Implicit Discourse Relation RecognitionLong Main
  441. PromptARA: Improving Deep Representation in Hybrid Automatic Readability Assessment with Prompt and Orthogonal ProjectionLong Findings
  442. PromptMix: A Class Boundary Augmentation Method for Large Language Model DistillationLong Main
  443. PromptST: Abstract Prompt Learning for End-to-End Speech TranslationLong Main
  444. Prompting ChatGPT in MNER: Enhanced Multimodal Named Entity Recognition with Auxiliary Refined KnowledgeLong Findings
  445. Prompting Large Language Models with Chain-of-Thought for Few-Shot Knowledge Base Question GenerationLong Main
  446. Prompting Scientific Names for Zero-Shot Species RecognitionShort Main
  447. Prompting and Evaluating Large Language Models for Proactive Dialogues: Clarification, Target-guided, and Non-collaborationLong Findings
  448. Prompting is not a substitute for probability measurements in large language modelsLong Main
  449. Prompting with Pseudo-Code InstructionsLong Main
  450. Proto-lm: A Prototypical Network-Based Framework for Built-in Interpretability in Large Language ModelsLong Findings
  451. Prototype-based HyperAdapter for Sample-Efficient Multi-task TuningLong Main
  452. Pseudointelligence: A Unifying Lens on Language Model EvaluationShort Findings
  453. PsyAttention: Psychological Attention Model for Personality DetectionLong Findings
  454. PsyCoT: Psychological Questionnaire as Powerful Chain-of-Thought for Personality DetectionLong Findings
  455. Pushdown Layers: Encoding Recursive Structure in Transformer Language ModelsLong Main
  456. QA-NatVer: Question Answering for Natural Logic-based Fact VerificationLong Main
  457. QADYNAMICS: Training Dynamics-Driven Synthetic QA Diagnostic for Zero-Shot Commonsense Question AnsweringShort Findings
  458. QTSumm: Query-Focused Summarization over Tabular DataLong Main
  459. QUADRo: Dataset and Models for QUestion-Answer Database RetrievalLong Findings
  460. QUDeval: The Evaluation of Questions Under Discussion Discourse ParsingLong Main
  461. Qualitative Code Suggestion: A Human-Centric Approach to Qualitative CodingLong Findings
  462. Quality Estimation-Assisted Automatic Post-EditingLong Findings
  463. Quantifying Character Similarity with Vision TransformersLong Main
  464. Quantifying the Dialect Gap and its Correlates Across LanguagesLong Findings
  465. Quantifying the redundancy between prosody and textLong Main
  466. Query Rewriting in Retrieval-Augmented Large Language ModelsLong Main
  467. Query-as-context Pre-training for Dense Passage RetrievalLong Main
  468. Query-based Image Captioning from Multi-context 360° ImagesLong Findings
  469. Query2Triple: Unified Query Encoding for Answering Diverse Complex Queries over Knowledge GraphsLong Findings
  470. Query2doc: Query Expansion with Large Language ModelsShort Main
  471. Question Answering as Programming for Solving Time-Sensitive QuestionsLong Main
  472. Quick Back-Translation for Unsupervised Machine TranslationLong Findings
  473. R$^3$ Prompting: Review, Rephrase and Resolve for Chain-of-Thought Reasoning in Large Language Models under Noisy ContextLong Findings
  474. R2H: Building Multimodal Navigation Helpers that Respond to Help RequestsLong Main
  475. RAPL: A Relation-Aware Prototype Learning Approach for Few-Shot Document-Level Relation ExtractionLong Main
  476. RECAL: Sample-Relation Guided Confidence Calibration over Tabular DataLong Findings
  477. RECAP: Towards Precise Radiology Report Generation via Dynamic Disease Progression ReasoningLong Findings
  478. ROBBIE: Robust Bias Evaluation of Large Generative Language ModelsLong Main
  479. ROME: Evaluating Pre-trained Vision-Language Models on Reasoning beyond Visual Common SenseLong Findings
  480. RSVP: Customer Intent Detection via Agent Response Contrastive and Generative Pre-TrainingLong Findings
  481. RWKV: Reinventing RNNs for the Transformer EraLong Findings
  482. RainProof: An Umbrella to Shield Text Generator from Out-Of-Distribution DataLong Main
  483. Random Entity Quantization for Parameter-Efficient Compositional Knowledge Graph RepresentationLong Main
  484. Ranking LLM-Generated Loop Invariants for Program VerificationShort Findings
  485. Rather a Nurse than a Physician - Contrastive Explanations under InvestigationLong Main
  486. Rationale-Enhanced Language Models are Better Continual Relation LearnersShort Main
  487. Re$^3$Dial: Retrieve, Reorganize and Rescale Conversations for Long-Turn Open-Domain Dialogue Pre-trainingLong Main
  488. Re-Examining Summarization Evaluation across Multiple Quality CriteriaShort Findings
  489. Re-ViLM: Retrieval-Augmented Visual Language Model for Zero and Few-Shot Image CaptioningLong Findings
  490. Re-weighting Tokens: A Simple and Effective Active Learning Strategy for Named Entity RecognitionShort Findings
  491. ReCEval: Evaluating Reasoning Chains via Correctness and InformativenessLong Main
  492. ReFSQL: A Retrieval-Augmentation Framework for Text-to-SQL GenerationLong Findings
  493. ReLM: Leveraging Language Models for Enhanced Chemical Reaction PredictionShort Findings
  494. ReSee: Responding through Seeing Fine-grained Visual Knowledge in Open-domain DialogueLong Main
  495. ReTAG: Reasoning Aware Table to Analytic Text GenerationLong Main
  496. ReadPrompt: A Readable Prompting Method for Reliable Knowledge ProbingLong Findings
  497. Reading Books is Great, But Not if You Are Driving! Visually Grounded Reasoning about Defeasible Commonsense NormsLong Main
  498. Reading Order Matters: Information Extraction from Visually-rich Documents by Token Path PredictionLong Main
  499. RealBehavior: A Framework for Faithfully Characterizing Foundation Models’ Human-like Behavior MechanismsLong Findings
  500. Reasoning Makes Good Annotators : An Automatic Task-specific Rules Distilling Framework for Low-resource Relation ExtractionLong Findings
  501. Reasoning about Ambiguous Definite DescriptionsShort Findings
  502. Reasoning with Language Model is Planning with World ModelLong Main
  503. ReasoningLM: Enabling Structural Subgraph Reasoning in Pre-trained Language Models for Question Answering over Knowledge GraphLong Main
  504. Reconstruct Before Summarize: An Efficient Two-Step Framework for Condensing and Summarizing Meeting TranscriptsLong Main
  505. Recurrent Neural Language Models as Probabilistic Finite-state AutomataLong Main
  506. Reduce Human Labor On Evaluating Conversational Information Retrieval System: A Human-Machine Collaboration ApproachLong Main
  507. Reducing Sequence Length by Predicting Edit Spans with Large Language ModelsLong Main
  508. Reducing Spurious Correlations in Aspect-based Sentiment Analysis with Explanation from Large Language ModelsLong Findings
  509. RefGPT: Dialogue Generation of GPT, by GPT, and for GPTLong Findings
  510. Referring Image Segmentation via Joint Mask Contextual Embedding Learning and Progressive Alignment NetworkLong Main
  511. RegaVAE: A Retrieval-Augmented Gaussian Mixture Variational Auto-Encoder for Language ModelingLong Findings
  512. Regulation and NLP (RegNLP): Taming Large Language ModelsLong Main
  513. Reinforced Target-driven Conversational PromotionLong Main
  514. Reinforcement Replaces Supervision: Query focused Summarization using Deep Reinforcement LearningLong Main
  515. Relation-Aware Question Answering for Heterogeneous Knowledge GraphsLong Findings
  516. Remember what you did so you know what to do nextShort Findings
  517. RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and GenerationLong Main
  518. Representation Projection Invariance Mitigates Representation CollapseLong Findings
  519. Representative Demonstration Selection for In-Context Learning with Two-Stage Determinantal Point ProcessLong Main
  520. Representativeness as a Forgotten Lesson for Multilingual and Code-switched Data Collection and PreparationLong Findings
  521. Responsible AI Considerations in Text Summarization Research: A Review of Current PracticesLong Findings
  522. Rethinking Model Selection and Decoding for Keyphrase Generation with Pre-trained Sequence-to-Sequence ModelsLong Main
  523. Rethinking Negative Pairs in Code SearchLong Main
  524. Rethinking Word-Level Auto-Completion in Computer-Aided TranslationLong Main
  525. Rethinking and Improving Multi-task Learning for End-to-end Speech TranslationLong Main
  526. Rethinking the Construction of Effective Metrics for Understanding the Mechanisms of Pretrained Language ModelsLong Findings
  527. Rethinking the Evaluation for Conversational Recommendation in the Era of Large Language ModelsLong Main
  528. Retrieval-Augmented Few-shot Text ClassificationLong Findings
  529. Retrieval-Augmented Parsing for Complex Graphs by Exploiting Structure and UncertaintyLong Findings
  530. Retrieval-Generation Alignment for End-to-End Task-Oriented Dialogue SystemLong Main
  531. Retrieval-based Knowledge Transfer: An Effective Approach for Extreme Large Language Model CompressionLong Findings
  532. Retrieving Multimodal Information for Augmented Generation: A SurveyLong Findings
  533. Retrofitting Light-weight Language Models for Emotions using Supervised Contrastive LearningLong Main
  534. Revisiting Automated Topic Model Evaluation with Large Language ModelsShort Main
  535. Revisiting Block-based Quantisation: What is Important for Sub-8-bit LLM Inference?Long Main
  536. Revisiting De-Identification of Electronic Medical Records: Evaluation of Within- and Cross-Hospital GeneralizationShort Main
  537. Revisiting Entropy Rate Constancy in TextLong Findings
  538. Revisiting Instruction Fine-tuned Model Evaluation to Guide Industrial ApplicationsShort Main
  539. Revisiting Large Language Models as Zero-shot Relation ExtractorsLong Findings
  540. Revisiting Machine Translation for Cross-lingual ClassificationLong Main
  541. Revisiting Source Context in Nearest Neighbor Machine TranslationLong Main
  542. Revisiting Sparse Retrieval for Few-shot Entity LinkingShort Main
  543. Revisiting the Knowledge Injection FrameworksLong Main
  544. Revisiting the Optimality of Word LengthsLong Main
  545. Reward-Augmented Decoding: Efficient Controlled Text Generation With a Unidirectional Reward ModelShort Main
  546. RexUIE: A Recursive Method with Explicit Schema Instructor for Universal Information ExtractionLong Findings
  547. RoAST: Robustifying Language Models via Adversarial Perturbation with Selective TrainingLong Findings
  548. RoBoCoP: A Comprehensive ROmance BOrrowing COgnate Package and Benchmark for Multilingual Cognate IdentificationLong Main
  549. RoMQA: A Benchmark for Robust, Multi-evidence, Multi-answer Question AnsweringLong Findings
  550. Robust Prompt Optimization for Large Language Models Against Distribution ShiftsLong Main
  551. RobustEmbed: Robust Sentence Embeddings Using Self-Supervised Contrastive Pre-TrainingLong Findings
  552. RobustGEC: Robust Grammatical Error Correction Against Subtle Context PerturbationLong Main
  553. Robustness Tests for Automatic Machine Translation Metrics with Adversarial AttacksShort Findings
  554. Robustness of Named-Entity Replacements for In-Context LearningShort Findings
  555. Role of Context in Unsupervised Sentence Representation Learning: the Case of Dialog Act ModelingShort Findings
  556. Roles of Scaling and Instruction Tuning in Language Perception: Model vs. Human AttentionLong Findings
  557. Romanization-based Large-scale Adaptation of Multilingual Language ModelsShort Findings
  558. Rumor Detection on Social Media with Crowd Intelligence and ChatGPT-Assisted NetworksLong Main
  559. S2abEL: A Dataset for Entity Linking from Scientific TablesLong Main
  560. SAC$^3$: Reliable Hallucination Detection in Black-Box Language Models via Semantic-aware Cross-check ConsistencyLong Findings
  561. SAMRank: Unsupervised Keyphrase Extraction using Self-Attention Map in BERT and GPT-2Long Main
  562. SCENE: Self-Labeled Counterfactuals for Extrapolating to Negative ExamplesLong Main
  563. SCITAB: A Challenging Benchmark for Compositional Reasoning and Claim Verification on Scientific TablesLong Main
  564. SDOH-NLI: a Dataset for Inferring Social Determinants of Health from Clinical NotesShort Findings
  565. SEAHORSE: A Multilingual, Multifaceted Dataset for Summarization EvaluationLong Main
  566. SEER : A Knapsack approach to Exemplar Selection for In-Context HybridQALong Main
  567. SELFOOD: Self-Supervised Out-Of-Distribution Detection via Learning to RankLong Findings
  568. SGP-TOD: Building Task Bots Effortlessly via Schema-Guided LLM PromptingLong Findings
  569. SHARCS: Efficient Transformers Through Routing with Dynamic Width Sub-networksShort Findings
  570. SIR-ABSC: Incorporating Syntax into RoBERTa-based Sentiment Analysis Models with a Special Aggregator TokenLong Findings
  571. SKD-NER: Continual Named Entity Recognition via Span-based Knowledge Distillation with Reinforcement LearningLong Main
  572. SLOG: A Structural Generalization Benchmark for Semantic ParsingLong Main
  573. SMoP: Towards Efficient and Effective Prompt Tuning with Sparse Mixture-of-PromptsShort Main
  574. SODA: Million-scale Dialogue Distillation with Social Commonsense ContextualizationLong Main
  575. SOUL: Towards Sentiment and Opinion Understanding of LanguageShort Main
  576. SPT: Learning to Selectively Insert Prompts for Better Prompt TuningLong Main
  577. STAIR: Learning Sparse Text and Image Representation in Grounded TokensLong Main
  578. STEER: Unified Style Transfer with Expert ReinforcementLong Findings
  579. STINMatch: Semi-Supervised Semantic-Topological Iteration Network for Financial Risk Detection via News Label DiffusionLong Main
  580. SUT: Active Defects Probing for Transcompiler ModelsShort Main
  581. SWEET - Weakly Supervised Person Name Extraction for Fighting Human TraffickingLong Findings
  582. SYMPTOMIFY: Transforming Symptom Annotations with Language Model Knowledge HarvestingLong Findings
  583. Salespeople vs SalesBot: Exploring the Role of Educational Value in Conversational Recommender SystemsLong Findings
  584. Scalable-DSC: A Structural Template Prompt Approach to Scalable Dialogue State CorrectionLong Main
  585. Scaling Law for Document Neural Machine TranslationLong Findings
  586. Scaling Laws vs Model Architectures: How does Inductive Bias Influence Scaling?Long Findings
  587. Scaling Vision-Language Models with Sparse Mixture of ExpertsLong Findings
  588. ScanDL: A Diffusion Model for Generating Synthetic Scanpaths on TextsLong Main
  589. ScdNER: Span-Based Consistency-Aware Document-Level Named Entity RecognitionShort Main
  590. Scene Graph Enhanced Pseudo-Labeling for Referring Expression ComprehensionLong Findings
  591. Schema-adaptable Knowledge Graph ConstructionLong Findings
  592. SciRepEval: A Multi-Format Benchmark for Scientific Document RepresentationsLong Main
  593. Search Augmented Instruction LearningLong Findings
  594. Seeing through the mess: evolutionary dynamics of lexical polysemyLong Main
  595. SegAugment: Maximizing the Utility of Speech Translation Data with Segmentation-based AugmentationsLong Findings
  596. Segmented Recurrent Transformer: An Efficient Sequence-to-Sequence ModelLong Findings
  597. Select, Prompt, Filter: Distilling Large Language Models for Summarizing ConversationsShort Main
  598. Selecting Key Views for Zero-Shot Entity LinkingLong Findings
  599. Selective Demonstrations for Cross-domain Text-to-SQLLong Findings
  600. Selective Labeling: How to Radically Lower Data-Labeling Costs for Document Extraction ModelsLong Main
  601. Selectively Answering Ambiguous QuestionsLong Main
  602. Self-Detoxifying Language Models via Toxification ReversalLong Main
  603. Self-Ensemble of $N$-best Generation Hypotheses by Lexically Constrained DecodingShort Main
  604. Self-Evolution Learning for Mixup: Enhance Data Augmentation on Few-Shot Text Classification TasksLong Main
  605. Self-ICL: Zero-Shot In-Context Learning with Self-Generated DemonstrationsLong Main
  606. Self-Improvement of Non-autoregressive Model via Sequence-Level DistillationLong Main
  607. Self-Influence Guided Data Reweighting for Language Model Pre-trainingLong Main
  608. Self-Knowledge Guided Retrieval Augmentation for Large Language ModelsLong Findings
  609. Self-Polish: Enhance Reasoning in Large Language Models via Problem RefinementLong Findings
  610. Self-Supervised Behavior Cloned Transformers are Path Crawlers for Text GamesShort Findings
  611. Self-Supervised Rule Learning to Link Text Segments to Relational Elements of Structured KnowledgeLong Findings
  612. Self-prompted Chain-of-Thought on Large Language Models for Open-domain Multi-hop ReasoningLong Findings
  613. Self-supervised Meta-Prompt Learning with Meta-Gradient Regularization for Few-shot GeneralizationLong Findings
  614. Self-supervised Post-processing Method to Enrich Pretrained Word VectorsShort Findings
  615. SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsLong Main
  616. Semantic Decomposition of Question and SQL for Text-to-SQL ParsingLong Findings
  617. Semantic Parsing by Large Language Models for Intricate Updating Strategies of Zero-Shot Dialogue State TrackingShort Findings
  618. Semantic Similarity Covariance Matrix ShrinkageLong Findings
  619. Semantic Space Grounded Weighted Decoding for Multi-Attribute Controllable Dialogue GenerationLong Main
  620. Semantic matching for text classification with complex class descriptionsLong Main
  621. Semi-Structured Object Sequence EncodersLong Findings
  622. Semi-automatic Data Enhancement for Document-Level Relation Extraction with Distant Supervision from Large Language ModelsShort Main
  623. Semi-supervised multimodal coreference resolution in image narrationsLong Main
  624. SentiStream: A Co-Training Framework for Adaptive Online Sentiment Analysis in Evolving Data StreamsLong Main
  625. Sentiment Analysis on Streaming User Reviews via Dual-Channel Dynamic Graph Neural NetworkLong Main
  626. Seq2seq is All You Need for Coreference ResolutionLong Main
  627. SeqXGPT: Sentence-Level AI-Generated Text DetectionLong Main
  628. Set Learning for Generative Information ExtractionShort Main
  629. Shall We Pretrain Autoregressive Language Models with Retrieval? A Comprehensive StudyLong Main
  630. Show, Write, and Retrieve: Entity-aware Article Generation and RetrievalLong Findings
  631. SiMFy: A Simple Yet Effective Approach for Temporal Knowledge Graph ReasoningLong Findings
  632. SimCKP: Simple Contrastive Learning of Keyphrase RepresentationsLong Findings
  633. SimCSE++: Improving Contrastive Learning for Sentence Embeddings from Two PerspectivesLong Main
  634. Simple Hardware-Efficient PCFGs with Independent Left and Right ProductionsShort Findings
  635. Simple Temporal Adaptation to Changing Label Sets: Hashtag Prediction via Dense KNNShort Main
  636. Simple and Effective Input Reformulations for TranslationLong Main
  637. Simpler neural networks prefer subregular languagesLong Findings
  638. Simplicity Level Estimate (SLE): A Learned Reference-Less Metric for Sentence SimplificationShort Main
  639. Simultaneous Machine Translation with Tailored ReferenceLong Findings
  640. Skill-Based Few-Shot Selection for In-Context LearningLong Main
  641. Small Language Models Fine-tuned to Coordinate Larger Language Models improve Complex ReasoningLong Main
  642. Smart “Chef”: Verifying the Effect of Role-based Paraphrasing for Aspect Term ExtractionShort Findings
  643. SmartSpanNER: Making SpanNER Robust in Low Resource ScenariosLong Findings
  644. Social Commonsense-Guided Search Query Generation for Open-Domain Knowledge-Powered ConversationsShort Findings
  645. Sociocultural Norm Similarities and Differences via Situational Alignment and Explainable Textual EntailmentLong Main
  646. Solving Hard Analogy Questions with Relation Embedding ChainsLong Main
  647. Solving the Right Problem is Key for Translational NLP: A Case Study in UMLS Vocabulary InsertionLong Findings
  648. Somali Information Retrieval Corpus: Bridging the Gap between Query Translation and Dedicated Language ResourcesShort Main
  649. SoulChat: Improving LLMs' Empathy, Listening, and Comfort Abilities through Fine-tuning with Multi-turn Empathy ConversationsShort Findings
  650. Sound of Story: Multi-modal Storytelling with AudioLong Findings
  651. Sources of Hallucination by Large Language Models on Inference TasksLong Findings
  652. SpEL: Structured Prediction for Entity LinkingLong Main
  653. Sparse Black-Box Multimodal Attack for Vision-Language Adversary GenerationLong Findings
  654. Sparse Frame Grouping Network with Action Centered for Untrimmed Video Paragraph CaptioningLong Findings
  655. Sparse Low-rank Adaptation of Pre-trained Language ModelsLong Main
  656. Sparse Universal TransformerLong Main
  657. Speak, Memory: An Archaeology of Books Known to ChatGPT/GPT-4Long Main
  658. Speaking Style Conversion in the Waveform Domain Using Discrete Self-Supervised UnitsLong Findings
  659. Specialist or Generalist? Instruction Tuning for Specific NLP TasksLong Main
  660. Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq GenerationLong Findings
  661. Speech Recognition and Meaning Interpretation: Towards Disambiguation of Structurally Ambiguous Spoken Utterances in IndonesianLong Main
  662. Speech-enriched Memory for Inference-time Adaptation of ASR Models to Word DictionariesLong Main
  663. SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational AbilitiesLong Findings
  664. Spoiler Detection as Semantic Text MatchingShort Main
  665. Stance Detection on Social Media with Background KnowledgeLong Main
  666. Standardizing Distress Analysis: Emotion-Driven Distress Identification and Cause Extraction (DICE) in Multimodal Online PostsLong Main
  667. Statistical Depth for Ranking and Characterizing Transformer-Based Text EmbeddingsLong Main
  668. Statistically Profiling Biases in Natural Language Reasoning Datasets and ModelsLong Findings
  669. SteerLM: Attribute Conditioned SFT as an (User-Steerable) Alternative to RLHFLong Findings
  670. Steering Large Language Models for Machine Translation with Finetuning and In-Context LearningShort Findings
  671. StereoMap: Quantifying the Awareness of Human-like Stereotypes in Large Language ModelsLong Main
  672. Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation BenchmarksShort Main
  673. StoryAnalogy: Deriving Story-level Analogies from Large Language Models to Unlock Analogical UnderstandingLong Main
  674. StrAE: Autoencoding for Pre-Trained Embeddings using Explicit StructureLong Main
  675. Strong and Efficient Baselines for Open Domain Conversational Question AnsweringShort Findings
  676. Struct-XLM: A Structure Discovery Multilingual Language Model for Enhancing Cross-lingual Transfer through Reinforcement LearningLong Main
  677. StructGPT: A General Framework for Large Language Model to Reason over Structured DataLong Main
  678. Structural Priming Demonstrates Abstract Grammatical Representations in Multilingual Language ModelsLong Main
  679. Structural generalization in COGS: Supertagging is (almost) all you needLong Main
  680. Structure-aware Knowledge Graph-to-text Generation with Planning Selection and Similarity DistinctionLong Main
  681. Style-Aware Radiology Report Generation with RadGraph and Few-Shot PromptingLong Findings
  682. StyleBART: Decorate Pretrained Model with Style Adapters for Unsupervised Stylistic Headline GenerationLong Findings
  683. Stylized Dialogue Generation with Feature-Guided Knowledge AugmentationLong Findings
  684. Sub-network Discovery and Soft-masking for Continual Learning of Mixed TasksLong Findings
  685. Subspace Chronicles: How Linguistic Information Emerges, Shifts and Interacts during Language Model TrainingLong Findings
  686. SummEdits: Measuring LLM Ability at Factual Reasoning Through The Lens of SummarizationLong Main
  687. SummIt: Iterative Text Summarization via ChatGPTLong Findings
  688. Summarizing Multiple Documents with Conversational Structure for Meta-Review GenerationLong Findings
  689. SuperDialseg: A Large-scale Dataset for Supervised Dialogue SegmentationLong Main
  690. SuperTweetEval: A Challenging, Unified and Heterogeneous Benchmark for Social Media NLP ResearchLong Findings
  691. Superlim: A Swedish Language Understanding Evaluation BenchmarkLong Main
  692. Support or Refute: Analyzing the Stance of Evidence to Detect Out-of-Context Mis- and DisinformationLong Main
  693. Survival of the Most Influential Prompts: Efficient Black-Box Prompt Search via Clustering and PruningLong Findings
  694. Syllogistic Reasoning for Legal Judgment AnalysisLong Main
  695. Symbol tuning improves in-context learning in language modelsLong Main
  696. Symbolic Planning and Code Generation for Grounded DialogueLong Main
  697. Symbolization, Prompt, and Classification: A Framework for Implicit Speaker Identification in NovelsLong Findings
  698. Syntactic Substitutability as Unsupervised Dependency SyntaxLong Main
  699. Syntax Matters: Towards Spoken Language Understanding via Syntax-Aware AttentionShort Findings
  700. Syntax-Aware Retrieval Augmented Code GenerationLong Findings
  701. Synthesize, if you do not have: Effective Synthetic Dataset Creation Strategies for Self-Supervised Opinion Summarization in E-commerceShort Findings
  702. Synthetic Data Generation with Large Language Models for Text Classification: Potential and LimitationsLong Main
  703. System Combination via Quality Estimation for Grammatical Error CorrectionLong Main
  704. Systematic Assessment of Factual Knowledge in Large Language ModelsShort Findings
  705. Systematic word meta-sense extensionLong Main
  706. T-Projection: High Quality Annotation Projection for Sequence Labeling TasksLong Findings
  707. T5Score: Discriminative Fine-tuning of Generative Evaluation MetricsLong Findings
  708. TADI: Topic-aware Attention and Powerful Dual-encoder Interaction for Recall in News RecommendationLong Findings
  709. TATA: Stance Detection via Topic-Agnostic and Topic-Aware EmbeddingsLong Main
  710. TCFLE-8: a Corpus of Learner Written Productions for French as a Foreign Language and its Application to Automated Essay ScoringLong Main
  711. TCRA-LLM: Token Compression Retrieval Augmented Large Language Model for Inference Cost ReductionLong Findings
  712. TELeR: A General Taxonomy of LLM Prompts for Benchmarking Complex TasksShort Findings
  713. TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language UnderstandingLong Findings
  714. TK-KNN: A Balanced Distance-Based Pseudo Labeling Approach for Semi-Supervised Intent ClassificationLong Findings
  715. TLM: Token-Level Masking for TransformersLong Main
  716. TOD-Flow: Modeling the Structure of Task-Oriented DialoguesLong Main
  717. TR-Rules: Rule-based Model for Link Forecasting on Temporal Knowledge Graph Considering Temporal RedundancyLong Findings
  718. TRAMS: Training-free Memory Selection for Long-range Language ModelingShort Findings
  719. TRAVEL: Tag-Aware Conversational FAQ Retrieval via Reinforcement LearningLong Main
  720. TRIGO: Benchmarking Formal Mathematical Proof Reduction for Generative Language ModelsLong Main
  721. TRIP: Accelerating Document-level Multilingual Pre-training via Triangular Document-level Pre-training on Parallel Data TripletsLong Findings
  722. TSTR: Target Similarity Tuning Meets the Real WorldShort Findings
  723. TaTA: A Multilingual Table-to-Text Dataset for African LanguagesLong Findings
  724. TabPrompt: Graph-based Pre-training and Prompting for Few-shot Table UnderstandingLong Findings
  725. TacoPrompt: A Collaborative Multi-Task Prompt Learning Method for Self-Supervised Taxonomy CompletionLong Main
  726. Tagging-Assisted Generation Model with Encoder and Decoder Supervision for Aspect Sentiment Triplet ExtractionLong Main
  727. Take a Closer Look at Multilinguality! Improve Multilingual Pre-Training Using Monolingual Corpora OnlyLong Findings
  728. TalkUp: Paving the Way for Understanding Empowering LanguageLong Findings
  729. Target-Agnostic Gender-Aware Contrastive Learning for Mitigating Bias in Multilingual Machine TranslationLong Main
  730. Target-Aware Spatio-Temporal Reasoning via Answering Questions in Dynamic Audio-Visual ScenariosLong Findings
  731. Target-oriented Proactive Dialogue Systems with Personalization: Problem Formulation and Dataset CurationShort Main
  732. Target-to-Source Augmentation for Aspect Sentiment Triplet ExtractionLong Main
  733. Task-Adaptive Tokenization: Enhancing Long-Form Text Generation Efficacy in Mental Health and BeyondLong Main
  734. Task-Agnostic Low-Rank Adapters for Unseen English DialectsLong Main
  735. Task-Attentive Transformer Architecture for Continual Learning of Vision-and-Language Tasks Using Knowledge DistillationLong Findings
  736. Task-Aware Self-Supervised Framework for Dialogue Discourse ParsingLong Findings
  737. Task-Level Thinking Steps Help Large Language Models for Challenging Classification TaskLong Main
  738. TaskWeb: Selecting Better Source Tasks for Multi-task NLPLong Main
  739. Taxonomy Expansion for Named Entity RecognitionLong Main
  740. Teacher Perception of Automatically Extracted Grammar Concepts for L2 Language LearningLong Findings
  741. TempTabQA: Temporal Question Answering for Semi-Structured TablesLong Main
  742. Temporal Extrapolation and Knowledge Transfer for Lifelong Temporal Knowledge Graph ReasoningLong Findings
  743. Temporal Knowledge Graph Forecasting Without Knowledge Using In-Context LearningLong Main
  744. Temporal Knowledge Graph Reasoning Based on N-tuple ModelingLong Findings
  745. Test-Time Self-Adaptive Small Language Models for Question AnsweringShort Findings
  746. Test-time Augmentation for Factual ProbingShort Findings
  747. Text Augmented Spatial Aware Zero-shot Referring Image SegmentationLong Findings
  748. Text Classification via Large Language ModelsLong Findings
  749. Text Embeddings Reveal (Almost) As Much As TextLong Main
  750. Text Fact TransferLong Main
  751. Text Rendering Strategies for Pixel Language ModelsLong Main
  752. Text Representation Distillation via Information Bottleneck PrincipleLong Main
  753. Text encoders bottleneck compositionality in contrastive vision-language modelsLong Main
  754. Text-Transport: Toward Learning Causal Effects of Natural LanguageLong Main
  755. Text-guided 3D Human Generation from 2D CollectionsLong Findings
  756. Text2Tree: Aligning Text Representation to the Label Tree Hierarchy for Imbalanced Medical ClassificationLong Findings
  757. TextMixer: Mixing Multiple Inputs for Privacy-Preserving InferenceLong Findings
  758. That was the last straw, we need more: Are Translation Systems Sensitive to Disambiguating Context?Long Findings
  759. The ACL OCL Corpus: Advancing Open Science in Computational LinguisticsLong Main
  760. The Art of SOCRATIC QUESTIONING: Recursive Thinking with Large Language ModelsLong Main
  761. The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal ModelsLong Main
  762. The Benefits of Label-Description Training for Zero-Shot Text ClassificationLong Main
  763. The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-TuningLong Main
  764. The Cost of Compression: Investigating the Impact of Compression on Parametric Knowledge in Language ModelsLong Findings
  765. The Curious Case of Hallucinatory (Un)answerability: Finding Truths in the Hidden States of Over-Confident Large Language ModelsLong Main
  766. The Distributional Hypothesis Does Not Fully Explain the Benefits of Masked Language Model PretrainingLong Main
  767. The Effect of Scaling, Retrieval Augmentation and Form on the Factual Consistency of Language ModelsLong Main
  768. The Framework Tax: Disparities Between Inference Efficiency in NLP Research and DeploymentLong Main
  769. The Intended Uses of Automated Fact-Checking Artefacts: Why, How and WhoLong Findings
  770. The Internal State of an LLM Knows When It's LyingLong Findings
  771. The Interpreter Understands Your Meaning: End-to-end Spoken Language Understanding Aided by Speech TranslationLong Findings
  772. The Iron(ic) Melting Pot: Reviewing Human Evaluation in Humour, Irony and Sarcasm GenerationLong Findings
  773. The Law and NLP: Bridging Disciplinary DisconnectsShort Findings
  774. The Less the Merrier? Investigating Language Representation in Multilingual ModelsLong Findings
  775. The Linearity of the Effect of Surprisal on Reading Times across LanguagesShort Findings
  776. The Locality and Symmetry of Positional EncodingsLong Findings
  777. The PEACE-Reviews dataset: Modeling Cognitive Appraisals in Emotion Text AnalysisLong Findings
  778. The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and ValuesLong Main
  779. The Past, Present, and Future of Typological Databases in NLPShort Findings
  780. The Sentiment Problem: A Critical Survey towards Deconstructing Sentiment AnalysisLong Main
  781. The Shifted and The Overlooked: A Task-oriented Investigation of User-GPT InteractionsLong Main
  782. The Skipped Beat: A Study of Sociopragmatic Understanding in LLMs for 64 LanguagesLong Main
  783. The Troubling Emergence of Hallucination in Large Language Models - An Extensive Definition, Quantification, and Prescriptive RemediationsLong Main
  784. The Truth, The Whole Truth, and Nothing but the Truth: A New Benchmark Dataset for Hebrew Text Credibility AssessmentLong Findings
  785. The Vault: A Comprehensive Multilingual Dataset for Advancing Code Understanding and GenerationLong Findings
  786. The language of prompting: What linguistic properties make a prompt successful?Long Findings
  787. The neural dynamics of word recognition and integrationLong Main
  788. The student becomes the master: Outperforming GPT3 on Scientific Factual Error CorrectionLong Findings
  789. TheoremQA: A Theorem-driven Question Answering DatasetLong Main
  790. Theory of Mind for Multi-Agent Collaboration via Large Language ModelsLong Main
  791. This Reads Like That: Deep Learning for Interpretable Natural Language ProcessingShort Main
  792. This is not a Dataset: A Large Negation Benchmark to Challenge Large Language ModelsLong Main
  793. Thorny Roses: Investigating the Dual Use Dilemma in Natural Language ProcessingLong Findings
  794. Three Questions Concerning the Use of Large Language Models to Facilitate Mathematics LearningShort Findings
  795. Three Stream Based Multi-level Event Contrastive Learning for Text-Video Event ExtractionLong Main
  796. Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie CalibrationLong Main
  797. Time-Aware Language Modeling for Historical Text DatingLong Findings
  798. Time-Considerable Dialogue Models via Reranking by Time DependencyLong Findings
  799. To Build Our Future, We Must Know Our Past: Contextualizing Paradigm Shifts in Natural Language ProcessingLong Main
  800. ToViLaG: Your Visual-Language Generative Model is Also An EvildoerLong Main
  801. Token Prediction as Implicit Classification to Identify LLM-Generated TextShort Main
  802. TokenDrop + BucketSampler: Towards Efficient Padding-free Fine-tuning of Language ModelsLong Findings
  803. Tokenization Consistency Matters for Generative Models on Extractive NLP TasksShort Findings
  804. TopWORDS-Poetry: Simultaneous Text Segmentation and Word Discovery for Classical Chinese Poetry via Bayesian InferenceLong Main
  805. Topic-DPR: Topic-based Prompts for Dense Passage RetrievalLong Findings
  806. Topic-Informed Dialogue Summarization using Topic Distribution and Prompt-based ModelingShort Findings
  807. Toward Human Readable Prompt Tuning: Kubrick’s The Shining is a good movie, and a good prompt too?Long Findings
  808. Toward Joint Language Modeling for Speech Units and TextLong Findings
  809. Toward a Critical Toponymy Framework for Named Entity Recognition: A Case Study of Airbnb in New York CityLong Main
  810. Towards A Holistic Landscape of Situated Theory of Mind in Large Language ModelsLong Findings
  811. Towards A Unified View of Sparse Feed-Forward Network in Pretraining Large Language ModelLong Main
  812. Towards Anytime Fine-tuning: Continually Pre-trained Language Models with Hypernetwork PromptsLong Findings
  813. Towards Being Parameter-Efficient: A Stratified Sparsely Activated Transformer with Dynamic CapacityLong Findings
  814. Towards Better Representations for Multi-Label Text Classification with Multi-granularity InformationLong Findings
  815. Towards Building More Robust NER datasets: An Empirical Study on NER Dataset Bias from a Dataset Difficulty ViewLong Main
  816. Towards Concept-Aware Large Language ModelsLong Findings
  817. Towards Conceptualization of ``Fair Explanation'': Disparate Impacts of anti-Asian Hate Speech Explanations on Content ModeratorsLong Main
  818. Towards Detecting Contextual Real-Time Toxicity for In-Game ChatLong Findings
  819. Towards Enhancing Relational Rules for Knowledge Graph Link PredictionLong Findings
  820. Towards Example-Based NMT with Multi-Levenshtein TransformersLong Main
  821. Towards Formality-Aware Neural Machine Translation by Leveraging Context InformationShort Findings
  822. Towards General Error Diagnosis via Behavioral Testing in Machine TranslationLong Findings
  823. Towards Informative Few-Shot Prompt with Maximum Information Gain for In-Context LearningLong Findings
  824. Towards Informative Open-ended Text Generation with Dynamic Knowledge TriplesLong Findings
  825. Towards Interpretable Mental Health Analysis with Large Language ModelsLong Main
  826. Towards LLM-driven Dialogue State TrackingLong Main
  827. Towards Low-Resource Automatic Program Repair with Meta-Learning and Pretrained Language ModelsLong Main
  828. Towards Making the Most of ChatGPT for Machine TranslationLong Findings
  829. Towards Mitigating LLM Hallucination via Self ReflectionLong Findings
  830. Towards Multilingual Interlinear Morphological GlossingLong Findings
  831. Towards Noise-Tolerant Speech-Referring Video Object Segmentation: Bridging Speech and TextLong Main
  832. Towards Reliable Misinformation Mitigation: Generalization, Uncertainty, and GPT-4Long Main
  833. Towards Robust Pruning: An Adaptive Knowledge-Retention Pruning Strategy for Language ModelsLong Main
  834. Towards Unsupervised Recognition of Token-level Semantic Differences in Related DocumentsShort Main
  835. Towards Zero-shot Learning for End-to-end Cross-modal Translation ModelsShort Findings
  836. Towards Zero-shot Relation Extraction in Web Mining: A Multimodal Approach with Relative XML PathLong Findings
  837. Towards a Better Understanding of Variations in Zero-Shot Neural Machine Translation PerformanceLong Main
  838. Towards a Deep Understanding of Multilingual End-to-End Speech TranslationLong Findings
  839. Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language ModelsLong Main
  840. Towards a Unified Conversational Recommendation System: Multi-task Learning via Contextualized Knowledge DistillationLong Main
  841. Towards a Unified Framework for Reference Retrieval and Related Work GenerationLong Findings
  842. Towards large language model-based personal agents in the enterprise: Current trends and open problemsLong Findings
  843. ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI ConversationShort Findings
  844. Toxicity in Multilingual Machine Translation at ScaleLong Findings
  845. Toxicity in chatgpt: Analyzing persona-assigned language modelsLong Findings
  846. Toxicity, Morality, and Speech Act Guided Stance DetectionLong Findings
  847. Training Simultaneous Speech Translation with Robust and Random Wait-k-Tokens StrategyLong Main
  848. Transcending Scaling Laws with 0.1% Extra ComputeLong Main
  849. Transductive Learning for Textual Few-Shot Classification in API-based Embedding ModelsLong Main
  850. Transfer-Free Data-Efficient Multilingual Slot LabelingLong Main
  851. Transformer Working Memory Enables Regular Language Reasoning And Natural Language Length ExtrapolationLong Findings
  852. Transformer-Based Language Model Surprisal Predicts Human Reading Times Best with About Two Billion Training TokensShort Findings
  853. Transformer-based Live Update Generation for Soccer Matches from Microblog PostsShort Main
  854. Transitioning Representations between Languages for Cross-lingual Event Detection via Langevin DynamicsShort Findings
  855. Translating away Translationese without Parallel DataLong Main
  856. Transparency at the Source: Evaluating and Interpreting Language Models With Access to the True DistributionLong Findings
  857. Tree Prompting: Efficient Task Adaptation without Fine-TuningLong Main
  858. Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language ModelsShort Main
  859. TrojanSQL: SQL Injection against Natural Language Interface to DatabaseLong Main
  860. TrueTeacher: Learning Factual Consistency Evaluation with Large Language ModelsLong Main
  861. Tuna: Instruction Tuning using Feedback from Large Language ModelsLong Findings
  862. Tunable Soft Prompts are Messengers in Federated LearningLong Findings
  863. Turn-Level Active Learning for Dialogue State TrackingLong Main
  864. Two Directions for Clinical Data Generation with Large Language Models: Data-to-Label and Label-to-DataLong Findings
  865. Type-Aware Decomposed Framework for Few-Shot Named Entity RecognitionLong Findings
  866. UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of RerankersLong Main
  867. ULF: Unsupervised Labeling Function Correction using Cross-Validation for Weak SupervisionShort Main
  868. UPRISE: Universal Prompt Retrieval for Improving Zero-Shot EvaluationLong Main
  869. UPTON: Preventing Authorship Leakage from Public Text Release via Data PoisoningLong Findings
  870. UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language ModelLong Findings
  871. USB: A Unified Summarization Benchmark Across Tasks and DomainsLong Findings
  872. Ultra-Fine Entity Typing with Prior Knowledge about Labels: A Simple Clustering Based StrategyLong Findings
  873. Uncertainty Guided Global Memory Improves Multi-Hop Question AnsweringShort Main
  874. Uncertainty-aware Parameter-Efficient Self-training for Semi-supervised Language UnderstandingLong Findings
  875. Uncovering Limitations in Text-to-Image Generation: A Contrastive Approach with Structured Semantic AlignmentLong Findings
  876. Uncovering the Root of Hate Speech: A Dataset for Identifying Hate Instigating SpeechLong Findings
  877. Understanding Compositional Data Augmentation in Typologically Diverse Morphological InflectionLong Main
  878. Understanding Computational Models of Semantic Change: New Insights from the Speech CommunityShort Main
  879. Understanding HTML with Large Language ModelsLong Findings
  880. Understanding Translationese in Cross-Lingual SummarizationLong Findings
  881. Understanding the Effect of Model Compression on Social Bias in Large Language ModelsShort Main
  882. Understanding the Inner-workings of Language Models Through Representation DissimilarityShort Main
  883. Understanding the Role of Input Token Characters in Language Models: How Does Information Loss Affect Performance?Long Main
  884. UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and ReasoningLong Main
  885. UniMath: A Foundational and Multimodal Mathematical ReasonerShort Main
  886. Unified Low-Resource Sequence Labeling by Sample-Aware Dynamic Sparse FinetuningLong Main
  887. Unified Representation for Non-compositional and Compositional ExpressionsLong Findings
  888. Uniform Complexity for Text GenerationLong Findings
  889. Unifying Cross-Lingual Transfer across Scenarios of Resource ScarcityLong Main
  890. Unifying Discrete and Continuous Representations for Unsupervised Paraphrase GenerationLong Main
  891. Unifying Text, Tables, and Images for Multimodal Question AnsweringLong Findings
  892. Universal Domain Adaptation for Robust Handling of Distributional Shifts in NLPLong Findings
  893. Universal Self-Adaptive PromptingLong Main
  894. Unlearn What You Want to Forget: Efficient Unlearning for LLMsLong Main
  895. Unleashing the Multilingual Encoder Potential: Boosting Zero-Shot Performance via Probability CalibrationShort Findings
  896. Unleashing the Power of Language Models in Text-Attributed GraphLong Findings
  897. Unmasking the Hidden Meaning: Bridging Implicit and Explicit Hate Speech Embedding RepresentationsShort Findings
  898. Unnatural Error Correction: GPT-4 Can Almost Perfectly Handle Unnatural Scrambled TextShort Main
  899. Unnatural language processing: How do language models handle machine-generated prompts?Long Findings
  900. Unraveling Downstream Gender Bias from Large Language Models: A Study on AI Educational Writing AssistanceLong Findings
  901. Unraveling Feature Extraction Mechanisms in Neural NetworksLong Main
  902. Unsupervised Binary Code Translation with Application to Code Clone Detection and Vulnerability DiscoveryLong Findings
  903. Unsupervised Candidate Answer Extraction through Differentiable Masker-Reconstructor ModelLong Findings
  904. Unsupervised Grammatical Error Correction Rivaling Supervised MethodsLong Main
  905. Unsupervised Lexical Simplification with Context AugmentationShort Findings
  906. Unsupervised Sounding Pixel LearningLong Main
  907. Unveiling the Essence of Poetry: Introducing a Comprehensive Dataset and Benchmark for Poem SummarizationShort Main
  908. Unveiling the Implicit Toxicity in Large Language ModelsLong Main
  909. Unveiling the Multi-Annotation Process: Examining the Influence of Annotation Quantity and Instance Difficulty on Model PerformanceLong Findings
  910. Unveiling the Power of Argument Arrangement in Online Persuasive DiscussionsLong Findings
  911. Using Artificial French Data to Understand the Emergence of Gender Bias in Transformer Language ModelsShort Main
  912. Using In-Context Learning to Improve Dialogue SafetyLong Findings
  913. Using Interpretation Methods for Model EnhancementLong Main
  914. Using LLM for Improving Key Event Discovery: Temporal-Guided News Stream Clustering with Event SummariesShort Findings
  915. VECHR: A Dataset for Explainable and Robust Classification of Vulnerability Type in the European Court of Human RightsShort Main
  916. VER: Unifying Verbalizing Entities and RelationsLong Findings
  917. VERVE: Template-based ReflectiVE Rewriting for MotiVational IntErviewingLong Findings
  918. VIBE: Topic-Driven Temporal Adaptation for Twitter ClassificationLong Main
  919. VIP5: Towards Multimodal Foundation Models for RecommendationLong Findings
  920. VIPHY: Probing “Visible” Physical Commonsense KnowledgeLong Findings
  921. VISIT: Visualizing and Interpreting the Semantic Information Flow of TransformersLong Findings
  922. VISTA: Visual-Textual Knowledge Graph Representation LearningLong Findings
  923. VLIS: Unimodal Language Models Guide Multimodal Language GenerationLong Main
  924. Values, Ethics, Morals? On the Use of Moral Concepts in NLP ResearchLong Findings
  925. Variance Matters: Detecting Semantic Differences without Corpus/Word AlignmentLong Main
  926. Variator: Accelerating Pre-trained Models with Plug-and-Play Compression ModulesLong Findings
  927. Vector-Quantized Prompt Learning for Paraphrase GenerationLong Findings
  928. Vera: A General-Purpose Plausibility Estimation Model for Commonsense StatementsLong Main
  929. Verb Conjugation in Transformers Is Determined by Linear Encodings of Subject NumberShort Findings
  930. ViPE: Visualise Pretty-much EverythingLong Main
  931. ViSoBERT: A Pre-Trained Language Model for Vietnamese Social Media Text ProcessingLong Main
  932. ViStruct: Visual Structural Knowledge Extraction via Curriculum Guided Code-Vision RepresentationLong Main
  933. Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is OffensiveLong Main
  934. Vicinal Risk Minimization for Few-Shot Cross-lingual Transfer in Abusive Language DetectionLong Main
  935. Video-Helpful Multimodal Machine TranslationLong Main
  936. Video-Text Retrieval by Supervised Sparse Multi-Grained LearningLong Findings
  937. Viewing Knowledge Transfer in Multilingual Machine Translation Through a Representational LensLong Findings
  938. Vision-Enhanced Semantic Entity Recognition in Document Images via Visually-Asymmetric Consistency LearningLong Main
  939. Visual Elements Mining as Prompts for Instruction Learning for Target-Oriented Multimodal Sentiment ClassificationLong Findings
  940. Visual Storytelling with Question-Answer PlansLong Findings
  941. Visually Grounded Continual Language Learning with Selective SpecializationLong Findings
  942. Visually-Situated Natural Language Understanding with Contrastive Reading Model and Frozen Large Language ModelsLong Main
  943. VivesDebate-Speech: A Corpus of Spoken Argumentation to Leverage Audio Features for Argument MiningShort Main
  944. WSDMS: Debunk Fake News via Weakly Supervised Detection of Misinforming Sentences with Contextualized Social WisdomLong Main
  945. Watermarking LLMs with Weight QuantizationLong Findings
  946. Watermarking PLMs on Classification Tasks by Combining Contrastive Learning with Weight PerturbationLong Findings
  947. We Are What We Repeatedly Do: Inducing and Deploying Habitual Schemas in Persona-Based ResponsesLong Main
  948. We Need to Talk About Reproducibility in NLP Model ComparisonLong Main
  949. We are Who We Cite: Bridges of Influence Between Natural Language Processing and Other Academic FieldsLong Main
  950. We're Afraid Language Models Aren't Modeling AmbiguityLong Main
  951. Weakly Supervised Semantic Parsing with Execution-based Spurious Program FilteringLong Main
  952. Weakly-Supervised Learning of Visual Relations in Multimodal PretrainingLong Main
  953. Well Begun is Half Done: Generator-agnostic Knowledge Pre-Selection for Knowledge-Grounded DialogueLong Main
  954. What Comes Next? Evaluating Uncertainty in Neural Text Generators Against Human Production VariabilityLong Main
  955. What Else Do I Need to Know? The Effect of Background Information on Users’ Reliance on QA SystemsLong Main
  956. What Makes Chain-of-Thought Prompting Effective? A Counterfactual StudyLong Findings
  957. What Makes it Ok to Set a Fire? Iterative Self-distillation of Contexts and Rationales for Disambiguating Defeasible Social and Moral SituationsLong Findings
  958. What do Deck Chairs and Sun Hats Have in Common? Uncovering Shared Properties in Large Concept VocabulariesShort Main
  959. What to Read in a Contract? Party-Specific Summarization of Legal Obligations, Entitlements, and ProhibitionsLong Main
  960. What's "up" with vision-language models? Investigating their struggle with spatial reasoningLong Main
  961. When Do Decompositions Help for Machine Reading?Short Main
  962. When Language Models Fall in Love: Animacy Processing in Transformer Language ModelsLong Main
  963. When Reviewers Lock Horns: Finding Disagreements in Scientific Peer ReviewsShort Main
  964. When and Why Does Bias Mitigation Work?Long Findings
  965. When are Lemons Purple? The Concept Association Bias of Vision-Language ModelsLong Main
  966. When it Rains, it Pours: Modeling Media Storms and the News EcosystemLong Findings
  967. When the Majority is Wrong: Modeling Annotator Disagreement for Subjective TasksLong Main
  968. Where to start? Analyzing the potential value of intermediate modelsLong Main
  969. Whispering LLaMA: A Cross-Modal Generative Error Correction Framework for Speech RecognitionShort Main
  970. Who Wrote it and Why? Prompting Large-Language Models for Authorship VerificationShort Findings
  971. Who is Speaking? Speaker-Aware Multiparty Dialogue Act ClassificationLong Findings
  972. Why LLMs Hallucinate, and How to Get (Evidential) Closure: Perceptual, Intensional, and Extensional Learning for Faithful Natural Language GenerationLong Main
  973. Why Should This Article Be Deleted? Transparent Stance Detection in Multilingual Wikipedia Editor DiscussionsLong Main
  974. WiCE: Real-World Entailment for Claims in WikipediaLong Main
  975. WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on WikipediaLong Findings
  976. Women Wearing Lipstick: Measuring the Bias Between an Object and Its Related GenderShort Findings
  977. WordNet Is All You Need: A Surprisingly Effective Unsupervised Method for Graded Lexical EntailmentShort Findings
  978. Words, Subwords, and Morphemes: What Really Matters in the Surprisal-Reading Time Relationship?Short Findings
  979. X-SNS: Cross-Lingual Transfer Prediction through Sub-Network SimilarityLong Findings
  980. XLM-V: Overcoming the Vocabulary Bottleneck in Multilingual Masked Language ModelsLong Main
  981. XLS-R fine-tuning on noisy word boundaries for unsupervised speech segmentation into wordsShort Findings
  982. XTREME-UP: A User-Centric Scarce-Data Benchmark for Under-Represented LanguagesLong Findings
  983. You Are What You Annotate: Towards Better Models through Annotator RepresentationsLong Findings
  984. You Told Me That Joke Twice: A Systematic Investigation of Transferability and Robustness of Humor Detection ModelsLong Main
  985. ZARA: Improving Few-Shot Self-Rationalization for Small Language ModelsLong Findings
  986. ZEROTOP: Zero-Shot Task-Oriented Semantic Parsing using Large Language ModelsShort Main
  987. ZGUL: Zero-shot Generalization to Unseen Languages using Multi-source Ensembling of Language AdaptersLong Main
  988. Zero-Shot Data Maps. Efficient Dataset Cartography Without Model TrainingLong Findings
  989. Zero-Shot-BERT-Adapters: a Zero-Shot Pipeline for Unknown Intent DetectionLong Findings
  990. Zero-shot Faithfulness Evaluation for Text Summarization with Foundation Language ModelLong Main
  991. Zero-shot Sharpness-Aware Quantization for Pre-trained Language ModelsLong Main
  992. Zero-shot Topical Text Classification with LLMs - an Experimental StudyLong Findings
  993. ZeroSCROLLS: A Zero-Shot Benchmark for Long Text UnderstandingLong Findings
  994. `Don't Get Too Technical with Me': A Discourse Structure-Based Framework for Automatic Science JournalismLong Main
  995. clembench: Using Game Play to Evaluate Chat-Optimized Language Models as Conversational AgentsLong Main
  996. e-THERAPIST: I suggest you to cultivate a mindset of positivity and nurture uplifting thoughtsLong Main
  997. impact of sample selection on in-context learning for entity extraction from scientific writingLong Findings
  998. kNN-CM: A Non-parametric Inference-Phase Adaptation of Parametric Text ClassifiersLong Findings
  999. mAggretriever: A Simple yet Effective Approach to Zero-Shot Multilingual Dense RetrievalShort Main
  1000. mLongT5: A Multilingual and Efficient Text-To-Text Transformer for Longer SequencesShort Findings

Looking for submission deadlines instead? See the conference deadline calendar.