← Search

Cheng-Kuang Wu

8 accepted papers

2025

None of the Above, Less of the Right Parallel Patterns in Human and LLM Performance on Multi-Choice Questions Answering

ACL 2025finding

Multiple-choice exam questions with “None of the above” (NA) options have been extensively studied in educational testing, in which existing research suggests that they better assess true knowledge. However, their impact on Large Language Models (LLMs) evaluation remains underexplored. Through syste…

Cited by 0SourcePDFScholar
2024

I Need Help! Evaluating LLM’s Ability to Ask for Users’ Support: A Case Study on Text-to-SQL Generation

EMNLP 2024main

This study explores the proactive ability of LLMs to seek user support. We propose metrics to evaluate the trade-off between performance improvements and user burden, and investigate whether LLMs can determine when to request help under varying information availability. Our experiments show that wit…

2024

Let Me Speak Freely? A Study On The Impact Of Format Restrictions On Large Language Model Performance.

EMNLP 2024industry

Structured generation, the process of producing content in standardized formats like JSON and XML, is widely utilized in real-world applications to extract key output information from large language models (LLMs).This study investigates whether such constraints on generation space impact LLMs’ abili…

Cited by 6SourcePDFScholar
2024

StreamBench: Towards Benchmarking Continuous Improvement of Language Agents

NeurIPS 2024poster

Recent works have shown that large language model (LLM) agents are able to improve themselves from experience, which is an important ability for continuous enhancement post-deployment. However, existing benchmarks primarily evaluate their innate capabilities and do not assess their ability to improv…

2024

Unveiling Selection Biases: Exploring Order and Token Sensitivity in Large Language Models

ACL 2024findings

In this paper, we investigate the phenomena of “selection biases” in Large Language Models (LLMs), focusing on problems where models are tasked with choosing the optimal option from an ordered sequence. We delve into biases related to option order and token usage, which significantly impact LLMs’ de…

Cited by 19SourcePDFScholar
2023

Fidelity-Enriched Contrastive Search: Reconciling the Faithfulness-Diversity Trade-Off in Text Generation

EMNLP 2023short main

In this paper, we address the hallucination problem commonly found in natural language generation tasks. Language models often generate fluent and convincing content but can lack consistency with the provided source, resulting in potential inaccuracies. We propose a new decoding method called Fideli…

Cited by 0SourcecodeScholar
2023

Self-ICL: Zero-Shot In-Context Learning with Self-Generated Demonstrations

EMNLP 2023long main

Large language models (LLMs) have exhibited striking in-context learning (ICL) ability to adapt to target tasks with a few input-output demonstrations. For better ICL, different methods are proposed to select representative demonstrations from existing training corpora. However, such settings are no…

Cited by 0SourcecodeScholar
2023

ZARA: Improving Few-Shot Self-Rationalization for Small Language Models

EMNLP 2023long findings

Language models (LMs) that jointly generate end-task answers as well as free-text rationales are known as self-rationalization models. Recent works demonstrate great performance gain for self-rationalization by few-shot prompting LMs with rationale-augmented exemplars. However, the ability to benefi…

Cited by 0SourcecodeScholar