← Search

Junyi Jessy Li

37 accepted papers

2026

PLSemanticsBench: A Formal Semantics Reasoning Benchmark for Code

ICML 2026poster

Recent work asks whether large language models (LLMs) condition their reasoning on explicit rules rather than statistical regularities from pretraining. Program execution provides a canonical instance: formal semantics define behavior through sym- bolic transition rules that can be systematically al…

Cited by 0SourceScholar
2025

AstroVisBench: A Code Benchmark for Scientific Computing and Visualization in Astronomy

NeurIPS 2025poster

Large Language Models (LLMs) are being explored for applications in scientific research, including their capabilities to synthesize literature, answer research questions, generate research ideas, and even conduct computational experiments. Ultimately, our goal is for these to help scientists derive…

Cited by 0SourceScholar
2025

Behavioral Analysis of Information Salience in Large Language Models

ACL 2025finding

Large Language Models (LLMs) excel at text summarization, a task that requires models to select content based on its importance. However, the exact notion of salience that LLMs have internalized remains unclear. To bridge this gap, we introduce an explainable framework to systematically derive and i…

2025

Is It JUST Semantics? A Case Study of Discourse Particle Understanding in LLMs

ACL 2025finding

Discourse particles are crucial elements that subtly shape the meaning of text. These words, often polyfunctional, give rise to nuanced and often quite disparate semantic/discourse effects,as exemplified by the diverse uses of the particle *just* (e.g., exclusive, temporal, emphatic). This work inve…

Cited by 0SourcePDFScholar
2025

SPRI: Aligning Large Language Models with Context-Situated Principles

ICML 2025poster

Aligning Large Language Models to integrate and reflect human values, especially for tasks that demand intricate human oversight, is arduous since it is resource-intensive and time-consuming to depend on human expertise for context-specific guidance. Prior work has utilized predefined sets of rules…

Cited by 0SourcePDFScholar
2024

Detection and Measurement of Syntactic Templates in Generated Text

EMNLP 2024main

The diversity of text can be measured beyond word-level features, however existing diversity evaluation focuses primarily on word-level features. Here we propose a method for evaluating diversity over syntactic features to characterize general repetition in models, beyond frequent n-grams. Specifica…

2024

Do *they* mean ‘us’? Interpreting Referring Expression variation under Intergroup Bias

EMNLP 2024finding

The variations between in-group and out-group speech (intergroup bias) are subtle and could underlie many social phenomena like stereotype perpetuation and implicit bias. In this paper, we model intergroup bias as a tagging task on English sports comments from forums dedicated to fandom for NFL team…

2024

FactPICO: Factuality Evaluation for Plain Language Summarization of Medical Evidence

ACL 2024long

Plain language summarization with LLMs can be useful for improving textual accessibility of technical content. But how factual are these summaries in a high-stakes domain like medicine? This paper presents FactPICO, a factuality benchmark for plain language summarization of medical texts describing…

2024

InfoLossQA: Characterizing and Recovering Information Loss in Text Simplification

ACL 2024long

Text simplification aims to make technical texts more accessible to laypeople but often results in deletion of information and vagueness. This work proposes InfoLossQA, a framework to characterize and recover simplification-induced information loss in form of question-and-answer (QA) pairs. Building…

2024

Language Models (Mostly) Do Not Consider Emotion Triggers When Predicting Emotion

NAACL 2024short

Situations and events evoke emotions in humans, but to what extent do they inform the prediction of emotion detection models? This work investigates how well human-annotated emotion triggers correlate with features that models deemed salient in their prediction of emotions. First, we introduce a nov…

2024

Learning to Refine with Fine-Grained Natural Language Feedback

EMNLP 2024finding

Recent work has explored the capability of large language models (LLMs) to identify and correct errors in LLM-generated responses. These refinement approaches frequently evaluate what sizes of models are able to do refinement for what problems, but less attention is paid to what effective feedback f…

2024

Which questions should I answer? Salience Prediction of Inquisitive Questions

EMNLP 2024main

Inquisitive questions — open-ended, curiosity-driven questions people ask as they read — are an integral part of discourse processing and comprehension. Recent work in NLP has taken advantage of question generation capabilities of LLMs to enhance a wide range of applications. But the space of inquis…

2023

Counterfactual Probing for the Influence of Affect and Specificity on Intergroup Bias

ACL 2023findings

While existing work on studying bias in NLP focues on negative or pejorative language use, Govindarajan et al. (2023) offer a revised framing of bias in terms of intergroup social context, and its effects on language behavior. In this paper, we investigate if two pragmatic features (specificity and…

2023

Discourse Analysis via Questions and Answers: Parsing Dependency Structures of Questions Under Discussion

ACL 2023findings

Automatic discourse processing is bottlenecked by data: current discourse formalisms pose highly demanding annotation tasks involving large taxonomies of discourse relations, making them inaccessible to lay annotators. This work instead adopts the linguistic framework of Questions Under Discussion (…

2023

Elaborative Simplification as Implicit Questions Under Discussion

EMNLP 2023long main

Automated text simplification, a technique useful for making text more accessible to people such as children and emergent bilinguals, is often thought of as a monolingual translation task from complex sentences to simplified sentences using encoder-decoder models. This view fails to account for elab…

Cited by 0SourceScholar
2023

Evaluating Subjective Cognitive Appraisals of Emotions from Large Language Models

EMNLP 2023long findings

The emotions we experience involve complex processes; besides physiological aspects, research in psychology has studied cognitive appraisals where people assess their situations subjectively, according to their own values (Scherer, 2005). Thus, the same situation can often result in different emotio…

Cited by 0SourcecodeScholar
2023

Multilingual Simplification of Medical Texts

EMNLP 2023long main

Automated text simplification aims to produce simple versions of complex texts. This task is especially useful in the medical domain, where the latest medical findings are typically communicated via complex and technical articles. This creates barriers for laypeople seeking access to up-to-date med…

Cited by 0SourcecodeScholar
2023

QUDeval: The Evaluation of Questions Under Discussion Discourse Parsing

EMNLP 2023long main

Questions Under Discussion (QUD) is a versatile linguistic framework in which discourse progresses as continuously asking questions and answering them. Automatic parsing of a discourse to produce a QUD structure thus entails a complex question generation task: given a document and an answer sentence…

Cited by 0SourcecodeScholar
2023

Summarizing, Simplifying, and Synthesizing Medical Evidence using GPT-3 (with Varying Success)

ACL 2023short

Large language models, particularly GPT-3, are able to produce high quality summaries ofgeneral domain news articles in few- and zero-shot settings. However, it is unclear if such models are similarly capable in more specialized domains such as biomedicine. In this paper we enlist domain experts (in…

2023

Unsupervised Extractive Summarization of Emotion Triggers

ACL 2023long

Understanding what leads to emotions during large-scale crises is important as it can provide groundings for expressed emotions and subsequently improve the understanding of ongoing disasters. Recent approaches trained supervised models to both detect emotions and explain emotion triggers (events an…

2022

Discourse Comprehension: A Question Answering Framework to Represent Sentence Connections

EMNLP 2022main

While there has been substantial progress in text comprehension through simple factoid question answering, more holistic comprehension of a discourse still presents a major challenge (Dunietz et al., 2020). Someone critically reflecting on a text as they read it will pose curiosity-driven, often ope…

2022

Evaluating Factuality in Text Simplification

ACL 2022long

Automated simplification models aim to make input texts more readable. Such methods have the potential to make complex information accessible to a wider audience, e.g., providing access to recent medical literature which might otherwise be impenetrable for a lay reader. However, such models risk int…

2022

How Do We Answer Complex Questions: Discourse Structure of Long-form Answers

ACL 2022long

Long-form answers, consisting of multiple sentences, can provide nuanced and comprehensive answers to a broader set of questions. To better understand this complex and understudied task, we study the functional structure of long-form answers collected from three datasets, ELI5, WebGPT and Natural Qu…

2022

Impact of Evaluation Methodologies on Code Summarization

ACL 2022long

There has been a growing interest in developing machine learning (ML) models for code summarization tasks, e.g., comment generation and method naming. Despite substantial increase in the effectiveness of ML models, the evaluation methodologies, i.e., the way people split datasets into training, vali…

2022

Learning to Describe Solutions for Bug Reports Based on Developer Discussions

ACL 2022findings

When a software bug is reported, developers engage in a discussion to collaboratively resolve it. While the solution is likely formulated within the discussion, it is often buried in a large amount of text, making it difficult to comprehend and delaying its implementation. To expedite bug resolution…

2022

Political Ideology and Polarization: A Multi-dimensional Approach

NAACL 2022long

Analyzing ideology and polarization is of critical importance in advancing our grasp of modern politics. Recent research has made great strides towards understanding the ideological bias (i.e., stance) of news media along the left-right spectrum. In this work, we instead take a novel and more nuance…

2022

ProtoTEx: Explaining Model Decisions with Prototype Tensors

ACL 2022long

We present ProtoTEx, a novel white-box NLP classification architecture based on prototype networks (Li et al., 2018). ProtoTEx faithfully explains model decisions based on prototype tensors that encode latent clusters of training examples. At inference time, classification decisions are based on the…

2022

Text Simplification of College Admissions Instructions: A Professionally Simplified and Verified Corpus

COLING 2022main

Access to higher education is critical for minority populations and emergent bilingual students. However, the language used by higher education institutions to communicate with prospective students is often too complex; concretely, many institutions in the US publish admissions application instructi…

Cited by 4SourcePDFScholar
2022

The Role of Context and Uncertainty in Shallow Discourse Parsing

COLING 2022main

Discourse parsing has proven to be useful for a number of NLP tasks that require complex reasoning. However, over a decade since the advent of the Penn Discourse Treebank, predicting implicit discourse relations in text remains challenging. There are several possible reasons for this, and we hypothe…

Cited by 0SourcePDFScholar
2022

Using Developer Discussions to Guide Fixing Bugs in Software

EMNLP 2022finding

Automatically fixing software bugs is a challenging task. While recent work showed that natural language context is useful in guiding bug-fixing models, the approach required prompting developers to provide this context, which was simulated through commit messages written after the bug-fixing code c…

2022

Why Do You Feel This Way? Summarizing Triggers of Emotions in Social Media Posts

EMNLP 2022main

Crises such as the COVID-19 pandemic continuously threaten our world and emotionally affect billions of people worldwide in distinct ways. Understanding the triggers leading to people’s emotions is of crucial importance. Social media posts can be a good source of such analysis, yet these texts tend…

2021

Deep Just-In-Time Inconsistency Detection Between Comments and Source Code

AAAI 2021technical

Natural language comments convey key aspects of source code such as implementation, usage, and pre- and post-conditions. Failure to update comments accordingly when the corresponding code is modified introduces inconsistencies, which is known to lead to confusion and software bugs. In this paper, we…

Cited by 55SourcePDFScholar
2021

Did they answer? Subjective acts and intents in conversational discourse

NAACL 2021long

Discourse signals are often implicit, leaving it up to the interpreter to draw the required inferences. At the same time, discourse is embedded in a social context, meaning that interpreters apply their own assumptions and beliefs when resolving these inferences, leading to multiple, valid interpret…

2021

Paragraph-level Simplification of Medical Texts

NAACL 2021long

We consider the problem of learning to simplify medical texts. This is important because most reliable, up-to-date information in biomedicine is dense with jargon and thus practically inaccessible to the lay audience. Furthermore, manual simplification does not scale to the rapidly growing body of b…