← Search

Yohan Jo

30 accepted papers

2026

Dual Mechanisms of Value Expression: Intrinsic vs. Prompted Values in Large Language Models

ICML 2026poster

Large language models can express values in two main ways: (1) $\textit{intrinsic}$ expression, reflecting the model's inherent values learned during training, and (2) $\textit{prompted}$ expression, elicited by explicit prompts. Given their widespread use in value alignment, it is paramount to clea…

Cited by 0SourceScholar
2026

Fine-Grained Multi Image Object Hallucination Benchmark

CVPR 2026

Multimodal Large Language Models (MLLMs) are increasingly deployed in multi-image scenarios requiring complex reasoning across visual contexts. However, current MLLMs remain fundamentally limited by object hallucination--generating plausible yet factually inconsistent descriptions about objects. Exi

Cited by 0SourceScholar
2026

Non-Collaborative User Simulators for Tool Agents

ICLR 2026poster

Tool agents interact with users through multi-turn dialogues to accomplish various tasks. Recent studies have adopted user simulation methods to develop these agents in multi-turn settings. However, existing user simulators tend to be agent-friendly, exhibiting only cooperative behaviors, which fail…

Cited by 0SourcecodeScholar
2026

ProPerSim: Developing Proactive and Personalized AI Assistants through User-Assistant Simulation

ICLR 2026poster

As large language models (LLMs) become increasingly integrated into daily life, there is growing demand for AI assistants that are not only reactive but also proactive and personalized. While recent advances have pushed forward proactivity and personalization individually, their combination remains…

Cited by 0SourceScholar
2026

SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents

ICLR 2026oral

Large Language Model (LLM) agents excel at multi-step, tool-augmented tasks. However, smart homes introduce distinct challenges, requiring agents to handle latent user intents, temporal dependencies, device constraints, scheduling, and more. The main bottlenecks for developing smart home agents with…

Cited by 0SourcecodeScholar
2025

Ask Optimal Questions: Aligning Large Language Models with Retriever’s Preference in Conversation

NAACL 2025findings

Conversational search, unlike single-turn retrieval tasks, requires understanding the current question within a dialogue context. The common approach of rewrite-then-retrieve aims to decontextualize questions to be self-sufficient for off-the-shelf retrievers, but most existing methods produce sub-o…

2025

Dialogue Systems for Emotional Support via Value Reinforcement

ACL 2025long

Emotional support dialogue systems aim to reduce help-seekers’ distress and help them overcome challenges. While human values—core beliefs that shape an individual’s priorities—are increasingly emphasized in contemporary psychological therapy for their role in fostering internal transformation and l…

Cited by 0SourcePDFScholar
2025

Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction

ACL 2025long

In designing multiple-choice questions (MCQs) in education, creating plausible distractors is crucial for identifying students’ misconceptions and gaps in knowledge and accurately assessing their understanding. However, prior studies on distractor generation have not paid sufficient attention to enh…

2025

Improving Dialogue State Tracking through Combinatorial Search for In-Context Examples

ACL 2025long

In dialogue state tracking (DST), in-context learning comprises a retriever that selects labeled dialogues as in-context examples and a DST model that uses these examples to infer the dialogue state of the query dialogue. Existing methods for constructing training data for retrievers suffer from thr…

2025

KMI: A Dataset of Korean Motivational Interviewing Dialogues for Psychotherapy

NAACL 2025long

The increasing demand for mental health services has led to the rise of AI-driven mental health chatbots, though challenges related to privacy, data collection, and expertise persist. Motivational Interviewing (MI) is gaining attention as a theoretical basis for boosting expertise in the development…

2025

Knowledge Tracing in Programming Education Integrating Students’ Questions

ACL 2025long

Knowledge tracing (KT) in programming education presents unique challenges due to the complexity of coding tasks and the diverse methods students use to solve problems. Although students’ questions often contain valuable signals about their understanding and misconceptions, traditional KT models oft…

2025

PVP: An Image Dataset for Personalized Visual Persuasion with Persuasion Strategies, Viewer Characteristics, and Persuasiveness Ratings

ACL 2025long

Visual persuasion, which uses visual elements to influence cognition and behaviors, is crucial in fields such as advertising and politicalcommunication. With recent advancements in artificial intelligence, there is growing potential to develop persuasive systems that automatically generate persuasiv…

2025

ToolDial: Multi-turn Dialogue Generation Method for Tool-Augmented Language Models

ICLR 2025poster

Tool-Augmented Language Models (TALMs) leverage external APIs to answer user queries across various domains. However, existing benchmark datasets for TALM research often feature simplistic dialogues that do not reflect real-world scenarios, such as the need for models to ask clarifying questions or…

2025

Towards Lifelong Dialogue Agents via Timeline-based Memory Management

NAACL 2025long

To achieve lifelong human-agent interaction, dialogue agents need to constantly memorize perceived information and properly retrieve it for response generation (RG). While prior studies focus on getting rid of outdated memories to improve retrieval quality, we argue that such memories provide rich,…

Cited by 0SourcePDFScholar
2025

Value Portrait: Assessing Language Models’ Values through Psychometrically and Ecologically Valid Items

ACL 2025long

The importance of benchmarks for assessing the values of language models has been pronounced due to the growing need of more authentic, human-aligned responses. However, existing benchmarks rely on human or machine annotations that are vulnerable to value-related biases. Furthermore, the tested scen…

Cited by 0SourcePDFScholar
2024

Mitigating Hallucination in Abstractive Summarization with Domain-Conditional Mutual Information

NAACL 2024findings

A primary challenge in abstractive summarization is hallucination—the phenomenon where a model generates plausible text that is absent in the source text. We hypothesize that the domain (or topic) of the source text triggers the model to generate text that is highly probable in the domain, neglectin…

2024

Model-based Preference Optimization in Abstractive Summarization without Human Feedback

EMNLP 2024main

In abstractive summarization, the challenge of producing concise and accurate summaries arises from the vast amount of information contained in the source document. Consequently, although Large Language Models (LLMs) can generate fluent text, they often introduce inaccuracies by hallucinating conten…

2024

Publicly Shareable Clinical Large Language Model Built on Synthetic Clinical Notes

ACL 2024findings

The development of large language models tailored for handling patients’ clinical notes is often hindered by the limited accessibility and usability of these notes due to strict privacy regulations.To address these challenges, we first create synthetic large-scale clinical notes using publicly avail…

2023

A Closer Look at the Intervention Procedure of Concept Bottleneck Models

ICML 2023poster

Concept bottleneck models (CBMs) are a class of interpretable neural network models that predict the target response of a given input based on its high-level concepts. Unlike the standard end-to-end models, CBMs enable domain experts to intervene on the predicted concepts and rectify any mistakes at…

2023

FactKG: Fact Verification via Reasoning on Knowledge Graphs

ACL 2023long

In real world applications, knowledge graphs (KG) are widely used in various domains (e.g. medical applications and dialogue agents). However, for fact verification, KGs have not been adequately utilized as a knowledge source. KGs can be a valuable knowledge source in fact verification due to their…

2023

From Values to Opinions: Predicting Human Behaviors and Stances Using Value-Injected Large Language Models

EMNLP 2023long main

Being able to predict people's opinions on issues and behaviors in realistic scenarios can be helpful in various domains, such as politics and marketing. However, conducting large-scale surveys like the European Social Survey to solicit people's opinions on individual issues can incur prohibitive co…

Cited by 0SourcecodeScholar
2023

KG-GPT: A General Framework for Reasoning on Knowledge Graphs Using Large Language Models

EMNLP 2023short findings

While large language models (LLMs) have made considerable advancements in understanding and generating unstructured text, their application in structured data remains underexplored. Particularly, using LLMs for complex reasoning tasks on knowledge graphs (KGs) remains largely untouched. To address t…

Cited by 0SourcecodeScholar
2023

Multi-User MultiWOZ: Task-Oriented Dialogues among Multiple Users

EMNLP 2023long findings

While most task-oriented dialogues assume conversations between the agent and one user at a time, dialogue systems are increasingly expected to communicate with multiple users simultaneously who make decisions collaboratively. To facilitate development of such systems, we release the Multi-User Mult…

Cited by 0SourcecodeScholar
2023

Open-WikiTable : Dataset for Open Domain Question Answering with Complex Reasoning over Table

ACL 2023findings

Despite recent interest in open domain question answering (ODQA) over tables, many studies still rely on datasets that are not truly optimal for the task with respect to utilizing structural nature of table. These datasets assume answers reside as a single cell value and do not necessitate exploring…

2022

Argument Mining for Review Helpfulness Prediction

EMNLP 2022main

The importance of reliably determining the helpfulness of product reviews is rising as both helpful and unhelpful reviews continue to accumulate on e-commerce websites. And argumentational features—such as the structure of arguments and the types of underlying elementary units—have shown to be promi…

Cited by 10SourcePDFScholar
2022

Status Biases in Deliberation Online: Evidence from a Randomized Experiment on ChangeMyView

EMNLP 2022finding

Status is widely used to incentivize user engagement online. However, visible status indicators could inadvertently bias online deliberation to favor high-status users. In this work, we design and deploy a randomized experiment on the ChangeMyView platform to quantify status biases in deliberation o…

2021

Knowledge-Enhanced Evidence Retrieval for Counterargument Generation

EMNLP 2021finding

Finding counterevidence to statements is key to many tasks, including counterargument generation. We build a system that, given a statement, retrieves counterevidence from diverse sources on the Web. At the core of this system is a natural language inference (NLI) model that determines whether a can…