← Search

Siyan Li

7 accepted papers

2025

Bottom-Up Synthesis of Knowledge-Grounded Task-Oriented Dialogues with Iteratively Self-Refined Prompts

NAACL 2025short

Training conversational question-answering (QA) systems demands a substantial amount of in-domain data, which is often scarce in practice. A common solution to this challenge is to generate synthetic data. Traditional methods typically follow a top-down approach, where a large language model (LLM) g…

Cited by 0SourcePDFScholar
2025

Investigating Personalized Driving Behaviors in Dilemma Zones: Analysis and Prediction of Stop-or-Go Decisions

RA-L 2025

Dilemma zones at signalized intersections present a commonly occurring yet unsolved challenge in traffic safety. The onsets of yellow-light prompts varied responses from drivers: some may brake abruptly, compromising ride comfort, while others may accelerate, increasing the likelihood of red-light v

Cited by 4SourceScholar
2025

PAPILLON: Privacy Preservation from Internet-based and Local Language Model Ensembles

NAACL 2025long

Users can divulge sensitive information to proprietary LLM providers, raising significant privacy concerns. While open-source models, hosted locally on the user’s machine, alleviate some concerns, models that users can host locally are often less capable than proprietary frontier models. Toward pres…

Cited by 5SourcePDFScholar
2024

Benchmarking and Improving Generator-Validator Consistency of Language Models

ICLR 2024poster

As of September 2023, ChatGPT correctly answers “what is 7+8” with 15, but when asked “7+8=15, True or False” it responds with “False”. This inconsistency between generating and validating an answer is prevalent in language models (LMs) and erodes trust. In this paper, we propose a framework for mea…

Cited by 29SourcePDFScholar
2022

Enhancing Self-Consistency and Performance of Pre-Trained Language Models through Natural Language Inference

EMNLP 2022main

While large pre-trained language models are powerful, their predictions often lack logical consistency across test inputs. For example, a state-of-the-art Macaw question-answering (QA) model answers <i>Yes</i> to <i>Is a sparrow a bird?</i> and <i>Does a bird have feet?</i> but answers <i>No</i> to…

Cited by 47SourcePDFScholar
2022

Inferring the Reader: Guiding Automated Story Generation with Commonsense Reasoning

EMNLP 2022finding

Transformer-based language model approaches to automated story generation currently provide state-of-the-art results. However, they still suffer from plot incoherence when generatingnarratives over time, and critically lack basiccommonsense reasoning. Furthermore, existing methods generally focus on…

2022

Systematicity in GPT-3’s Interpretation of Novel English Noun Compounds

EMNLP 2022finding

Levin et al. (2019) show experimentally that the interpretations of novel English noun compounds (e.g., stew skillet), while not fully compositional, are highly predictable based on whether the modifier and head refer to artifacts or natural kinds. Is the large language model GPT-3 governed by the s…

Cited by 14SourcePDFScholar