← Search

Heeseung Yun

11 accepted papers

2025

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates

ACL 2025long

While pre-trained multimodal representations (e.g., CLIP) have shown impressive capabilities, they exhibit significant compositional vulnerabilities leading to counterintuitive judgments. We introduce Multimodal Adversarial Compositionality (MAC), a benchmark that leverages large language models (LL…

2025

FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games

EMNLP 2025

GUI agents powered by LLMs show promise in interacting with diverse digital environments. Among these, video games offer a valuable testbed due to their varied interfaces, with adventure games posing additional challenges through complex, narrative-driven interactions. Existing game benchmarks, howe

Cited by 0SourcePDFScholar
2025

Gaze Beyond the Frame: Forecasting Egocentric 3D Visual Span

NeurIPS 2025spotlight

People continuously perceive and interact with their surroundings based on underlying intentions that drive their exploration and behaviors. While research in egocentric user and scene understanding has focused primarily on motion and contact-based interaction, forecasting human visual perception it…

Cited by 0SourceScholar
2025

ReSpec: Relevance and Specificity Grounded Online Filtering for Learning on Video-Text Data Streams

CVPR 2025poster

The rapid growth of video-text data presents challenges in storage and computation during training. Online learning, which processes streaming data in real-time, offers a promising solution to these issues while also allowing swift adaptations in scenarios demanding real-time responsiveness. One str…

2023

Fusing Pre-Trained Language Models With Multimodal Prompts Through Reinforcement Learning

CVPR 2023poster

Language models are capable of commonsense reasoning: while domain-specific models can learn from explicit knowledge (e.g. commonsense graphs [6], ethical norms [25]), and larger models like GPT-3 manifest broad commonsense reasoning capacity. Can their knowledge be extended to multimodal inputs suc…

2022

Panoramic Vision Transformer for Saliency Detection in 360° Videos

ECCV 2022poster

"360° video saliency detection is one of the challenging benchmarks for 360° video understanding since non-negligible distortion and discontinuity occur in the projection of any format of 360° videos, and capture-worthy viewpoint in the omnidirectional sphere is ambiguous by nature. We present a…

2021

Pano-AVQA: Grounded Audio-Visual Question Answering on 360deg Videos

ICCV 2021poster

360deg videos convey holistic views for the surroundings of a scene. It provides audio-visual cues beyond predetermined normal field of views and displays distinctive spatial relations on a sphere. However, previous benchmark tasks for panoramic videos are still limited to evaluate the semantic unde…

Cited by 99PDFcodeScholar
2021

Transitional Adaptation of Pretrained Models for Visual Storytelling

CVPR 2021poster

Previous models for vision-to-language generation tasks usually pretrain a visual encoder and a language generator in the respective domains and jointly finetune them with the target task. However, this direct transfer practice may suffer from the discord between visual specificity and language flue…

Cited by 40PDFcodeScholar
2020

Character Grounding and Re-Identification in Story of Videos and Text Descriptions

ECCV 2020poster

We address character grounding and re-identification in multiple story-based videos like movies and associated text descriptions. In order to solve these related tasks in a mutually rewarding way, we propose a model named Character in Story Identification Network (CiSIN). Our method builds two seman…