← Search

Yunseok Jang

7 accepted papers

2025

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents

CVPR 2025poster

Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have sparked significant interest in developing GUI visual agents. We introduce MONDAY (Mobile OS Navigation Task Dataset for Agents from YouTube), a large-scale dataset of 313K annotated frames from 20K instructio…

2024

YTCommentQA: Video Question Answerability in Instructional Videos

AAAI 2024technical

Instructional videos provide detailed how-to guides for various tasks, with viewers often posing questions regarding the content. Addressing these questions is vital for comprehending the content, yet receiving immediate answers is difficult. While numerous computational models have been developed f…

2023

Unsupervised Task Graph Generation from Instructional Video Transcripts

ACL 2023findings

This work explores the problem of generating task graphs of real-world activities. Different from prior formulations, we consider a setting where text transcripts of instructional videos performing a real-world activity (e.g., making coffee) are provided and the goal is to identify the key steps rel…

Cited by 12SourcePDFScholar
2019

Adversarial Defense via Learning to Generate Diverse Attacks

ICCV 2019poster

With the remarkable success of deep learning, Deep Neural Networks (DNNs) have been applied as dominant tools to various machine learning domains. Despite this success, however, it has been found that DNNs are surprisingly vulnerable to malicious attacks; adding a small, perceptually indistinguishab…

Cited by 99PDFcodeScholar
2019

Diversity-Sensitive Conditional Generative Adversarial Networks

ICLR 2019poster

We propose a simple yet highly effective method that addresses the mode-collapse problem in the Conditional Generative Adversarial Network (cGAN). Although conditional distributions are multi-modal (i.e., having many modes) in practice, most cGAN approaches tend to learn an overly simplified distr…

Cited by 253SourcePDFScholar
2017

TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering

CVPR 2017spotlight

Vision and language understanding has emerged as a subject undergoing intense study in Artificial Intelligence. Among many tasks in this line of research, visual question answering (VQA) has been one of the most successful ones, where the goal is to learn a model that understands visual content at r…

Cited by 676PDFScholar