← Search

Yun-Nung Chen

34 accepted papers

2026

Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors

ICML 2026poster

Existing data poisoning attacks on retrieval-augmented generation (RAG) systems scale poorly because they require costly optimization of poisoned documents for each target phrase. We introduce Eyes-on-Me, a modular attack that decomposes an adversarial document into reusable **Attention Attractors**…

Cited by 0SourceScholar
2025

Compound AI Systems Optimization: A Survey of Methods, Challenges, and Future Directions

EMNLP 2025

Recent advancements in large language models (LLMs) and AI systems have led to a paradigm shift in the design and optimization of complex AI workflows. By integrating multiple components, compound AI systems have become increasingly adept at performing sophisticated tasks. However, as these systems

2025

Creativity in LLM-based Multi-Agent Systems: A Survey

EMNLP 2025

Large language model (LLM)-driven multi-agent systems (MAS) are transforming how humans and AIs collaboratively generate ideas and artifacts. While existing surveys provide comprehensive overviews of MAS infrastructures, they largely overlook the dimension of creativity , including how novel outputs

2025

LLMs are Biased Evaluators But Not Biased for Fact-Centric Retrieval Augmented Generation

ACL 2025finding

Recent studies have demonstrated that large language models (LLMs) exhibit significant biases in evaluation tasks, particularly in preferentially rating and favoring self-generated content. However, the extent to which this bias manifests in fact-oriented tasks, especially within retrieval-augmented…

2025

Mitigating Forgetting in LLM Fine-Tuning via Low-Perplexity Token Learning

NeurIPS 2025poster

Maintaining consistent model performance across domains is a fundamental challenge in machine learning. While recent work has explored using LLM-generated data for fine-tuning, its impact on cross-domain generalization remains poorly understood. This paper presents a systematic analysis revealing th…

Cited by 0SourceScholar
2025

None of the Above, Less of the Right Parallel Patterns in Human and LLM Performance on Multi-Choice Questions Answering

ACL 2025finding

Multiple-choice exam questions with “None of the above” (NA) options have been extensively studied in educational testing, in which existing research suggests that they better assess true knowledge. However, their impact on Large Language Models (LLMs) evaluation remains underexplored. Through syste…

Cited by 0SourcePDFScholar
2025

Transferring Textual Preferences to Vision-Language Understanding through Model Merging

ACL 2025short

Large vision-language models (LVLMs) perform outstandingly across various multimodal tasks. However, their ability to evaluate generated content remains limited, and training vision-language reward models (VLRMs) with preference data is computationally expensive. This paper explores a training-free…

Cited by 0SourcePDFScholar
2024

DogeRM: Equipping Reward Models with Domain Knowledge through Model Merging

EMNLP 2024main

Reinforcement learning from human feedback (RLHF) is a popular strategy for aligning large language models (LLMs) with desired behaviors. Reward modeling is a crucial step in RLHF. However, collecting paired preference data for training reward models is often costly and time-consuming, especially fo…

2024

Editing the Mind of Giants: An In-Depth Exploration of Pitfalls of Knowledge Editing in Large Language Models

EMNLP 2024finding

Knowledge editing is a rising technique for efficiently updating factual knowledge in large language models (LLMs) with minimal alteration of parameters. However, recent studies have identified side effects, such as knowledge distortion and the deterioration of general abilities, that have emerged a…

2024

Efficient Unseen Language Adaptation for Multilingual Pre-Trained Language Models

EMNLP 2024main

Multilingual pre-trained language models (mPLMs) have demonstrated notable effectiveness in zero-shot cross-lingual transfer tasks. Specifically, they can be fine-tuned solely on tasks in the source language and subsequently applied to tasks in the target language. However, for low-resource language…

2024

I Need Help! Evaluating LLM’s Ability to Ask for Users’ Support: A Case Study on Text-to-SQL Generation

EMNLP 2024main

This study explores the proactive ability of LLMs to seek user support. We propose metrics to evaluate the trade-off between performance improvements and user burden, and investigate whether LLMs can determine when to request help under varying information availability. Our experiments show that wit…

2024

Injecting Salesperson’s Dialogue Strategies in Large Language Models with Chain-of-Thought Reasoning

ACL 2024findings

Recent research in dialogue systems focuses on two main categories: task-oriented (TOD) and open-domain (chit-chat) dialogues. TOD systems help users complete specific tasks, while open-domain systems aim to create engaging conversations. However, user intents often emerge during interactions. A rec…

2024

Let Me Speak Freely? A Study On The Impact Of Format Restrictions On Large Language Model Performance.

EMNLP 2024industry

Structured generation, the process of producing content in standardized formats like JSON and XML, is widely utilized in real-world applications to extract key output information from large language models (LLMs).This study investigates whether such constraints on generation space impact LLMs’ abili…

Cited by 6SourcePDFScholar
2024

StreamBench: Towards Benchmarking Continuous Improvement of Language Agents

NeurIPS 2024poster

Recent works have shown that large language model (LLM) agents are able to improve themselves from experience, which is an important ability for continuous enhancement post-deployment. However, existing benchmarks primarily evaluate their innate capabilities and do not assess their ability to improv…

2024

Two Tales of Persona in LLMs: A Survey of Role-Playing and Personalization

EMNLP 2024finding

The concept of *persona*, originally adopted in dialogue literature, has re-surged as a promising framework for tailoring large language models (LLMs) to specific context (*e.g.*, personalized search, LLM-as-a-judge). However, the growing research on leveraging persona in LLMs is relatively disorgan…

2024

Visualizing Dialogues: Enhancing Image Selection through Dialogue Understanding with Large Language Models

ACL 2024findings

For dialogue systems, the utilization of multimodal dialogue responses, as opposed to relying solely on text-only responses, offers the capability to describe different concepts through various modalities. This enhances the effectiveness of communication and elevates the overall conversational exper…

2023

Self-ICL: Zero-Shot In-Context Learning with Self-Generated Demonstrations

EMNLP 2023long main

Large language models (LLMs) have exhibited striking in-context learning (ICL) ability to adapt to target tasks with a few input-output demonstrations. For better ICL, different methods are proposed to select representative demonstrations from existing training corpora. However, such settings are no…

Cited by 0SourcecodeScholar
2023

Zero-Shot Prompting for Implicit Intent Prediction and Recommendation with Commonsense Reasoning

ACL 2023findings

The current generation of intelligent assistants require explicit user requests to perform tasks or services, often leading to lengthy and complex conversations. In contrast, human assistants can infer multiple implicit intents from utterances via their commonsense knowledge, thereby simplifying int…

2022

SalesBot: Transitioning from Chit-Chat to Task-Oriented Dialogues

ACL 2022long

Dialogue systems are usually categorized into two types, open-domain and task-oriented. The first one focuses on chatting with users and making them engage in the conversations, where selecting a proper topic to fit the dialogue context is essential for a successful dialogue. The other one focuses o…

2021

Efficient Multi-Task Auxiliary Learning: Selecting Auxiliary Data by Feature Similarity

EMNLP 2021main

Multi-task auxiliary learning utilizes a set of relevant auxiliary tasks to improve the performance of a primary task. A common usage is to manually select multiple auxiliary tasks for multi-task learning on all data, which raises two issues: (1) selecting beneficial auxiliary tasks for a primary ta…

2021

Modeling Diagnostic Label Correlation for Automatic ICD Coding

NAACL 2021long

Given the clinical notes written in electronic health records (EHRs), it is challenging to predict the diagnostic codes which is formulated as a multi-label classification task. The large set of labels, the hierarchical dependency, and the imbalanced data make this prediction task extremely hard. Mo…

2020

TaylorGAN: Neighbor-Augmented Policy Update Towards Sample-Efficient Natural Language Generation

NeurIPS 2020poster

Score function-based natural language generation (NLG) approaches such as REINFORCE, in general, suffer from low sample efficiency and training instability problems. This is mainly due to the non-differentiable nature of the discrete space sampling and thus these methods have to treat the discrimina…

2019

Mitigating the Impact of Speech Recognition Errors on Spoken Question Answering by Adversarial Domain Adaptation

ICASSP 2019accepted

Spoken question answering (SQA) is challenging due to complex reasoning on top of the spoken documents. The recent studies have also shown the catastrophic impact of automatic speech recognition (ASR) errors on SQA. Therefore, this work proposes to mitigate the ASR errors by aligning the mismatch be…

Cited by 0SourceScholar
2019

Modeling Melodic Feature Dependency with Modularized Variational Auto-encoder

ICASSP 2019accepted

Automatic melody generation has been a long-time aspiration for both AI researchers and musicians. However, learning to generate euphonious melodies has turned out to be highly challenging. This paper introduces 1) a new variant of variational autoencoder (VAE), where the model structure is designed…

Cited by 0SourceScholar
2018

Adversarial Advantage Actor-Critic Model for Task-Completion Dialogue Policy Learning

ICASSP 2018accepted

This paper presents a new method - adversarial advantage actor-critic (Adversarial A2C), which significantly improves the efficiency of dialogue policy learning in task-completion dialogue systems. Inspired by generative adversarial networks (GAN), we train a discriminator to differentiate responses…

Cited by 0SourceScholar
2017

End-to-end joint learning of natural language understanding and dialogue manager

ICASSP 2017accepted

Natural language understanding and dialogue policy learning are both essential in conversational systems that predict the next system actions in response to a current user utterance. Conventional approaches aggregate separate models of natural language understanding (NLU) and system action predictio…

Cited by 0SourceScholar
2016

Unsupervised user intent modeling by feature-enriched matrix factorization

ICASSP 2016accepted

Spoken language interfaces are being incorporated into various devices such as smart phones and TVs. However, dialogue systems may fail to respond correctly when users' request functionality is not supported by currently installed apps. This paper proposes a feature-enriched matrix factorization (MF…

Cited by 0SourceScholar
2016

Zero-shot learning of intent embeddings for expansion by convolutional deep structured semantic models

ICASSP 2016accepted

The recent surge of intelligent personal assistants motivates spoken language understanding of dialogue systems. However, the domain constraint along with the inflexible intent schema remains a big issue. This paper focuses on the task of intent expansion, which helps remove the domain limit and mak…

Cited by 0SourceScholar