← Search

Dong Won Lee

7 accepted papers

2026

Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models’ Social Reasoning

RSS 2026poster

Our work focuses on the social reasoning capabilities of foundational models for real-world human–robot interactions. We introduce the Social Human Robot Embodied Conversation (SHREC) Dataset, a large-scale benchmark of 400 real-world human-robot interaction videos and over 10K annotations, capturin…

Cited by 0SourceScholar
2025

Aligning Dialogue Agents with Global Feedback via Large Language Model Multimodal Reward Decomposition

EMNLP 2025

We propose a large language model based reward decomposition framework for aligning dialogue agents using only a single session-level feedback signal. We leverage the reasoning capabilities of a frozen, pretrained large language model (LLM) to infer fine-grained local implicit rewards by decomposing

Cited by 0SourcePDFScholar
2024

Global Reward to Local Rewards: Multimodal-Guided Decomposition for Improving Dialogue Agents

EMNLP 2024main

We describe an approach for aligning an LLM based dialogue agent for long-term social dialogue, where there is only a single global score given by the user at the end of the session. In this paper, we propose the usage of denser naturally-occurring multimodal communicative signals as local implicit…

2023

Lecture Presentations Multimodal Dataset: Towards Understanding Multimodality in Educational Videos

ICCV 2023poster

Many educational videos use slide presentations, a sequence of visual pages that contain text and figures accompanied by spoken language, which are constructed and presented carefully in order to optimally transfer knowledge to students. Previous studies in multimedia and psychology attribute the ef…

Cited by 11PDFcodeScholar
2023

MultiPar-T: Multiparty-Transformer for Capturing Contingent Behaviors in Group Conversations

IJCAI 2023poster

As we move closer to real-world social AI systems, AI agents must be able to deal with multiparty (group) conversations. Recognizing and interpreting multiparty behaviors is challenging, as the system must recognize individual behavioral cues, deal with the complexity of multiple streams of data fro…

2022

Low-Resource Adaptation for Personalized Co-Speech Gesture Generation

CVPR 2022poster

Personalizing an avatar for co-speech gesture generation from spoken language requires learning the idiosyncrasies of a person's gesture style from a small amount of data. Previous methods in gesture generation require large amounts of data for each speaker, which is often infeasible. We propose an…

Cited by 30PDFScholar
2020

Style Transfer for Co-Speech Gesture Animation: A Multi-Speaker Conditional-Mixture Approach

ECCV 2020poster

How can we teach robots or virtual assistants to gesture naturally? Can we go further and adapt the gesturing style to follow a specific speaker? Gestures that are naturally timed with corresponding speech during human communication are called co-speech gestures. A key challenge, called gesture styl…