← Search

Jae Won Cho

8 accepted papers

2026

GUIDE: A Benchmark for Understanding and Assisting Users in Open-Ended GUI Tasks

CVPR 2026

Graphical User Interface (GUI) agents have the potential to assist users in interacting with complex software (e.g., PowerPoint, Photoshop). While prior research has primarily focused on automating user actions through clicks and keystrokes, this paradigm overlooks human intention, where users value

Cited by 0SourcecodeScholar
2026

Stay in your Lane: Role Specific Queries with Overlap Suppression Loss for Dense Video Captioning

CVPR 2026

Dense Video Captioning (DVC) is a challenging multimodal task that involves temporally localizing multiple events within a video and describing them with natural language. While query-based frameworks enable the simultaneous, end-to-end processing of localization and captioning, their reliance on sh

Cited by 0SourceScholar
2024

Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality

EMNLP 2024main

In this paper, we propose a new method to enhance compositional understanding in pre-trained vision and language models (VLMs) without sacrificing performance in zero-shot multi-modal tasks. Traditional fine-tuning approaches often improve compositional reasoning at the cost of degrading multi-modal…

2023

Generative Bias for Robust Visual Question Answering

CVPR 2023poster

The task of Visual Question Answering (VQA) is known to be plagued by the issue of VQA models exploiting biases within the dataset to make its final prediction. Various previous ensemble based debiasing methods have been proposed where an additional model is purposefully trained to be biased in orde…

2022

Investigating Top-k White-Box and Transferable Black-Box Attack

CVPR 2022poster

Existing works have identified the limitation of top-1 attack success rate (ASR) as a metric to evaluate the attack strength but exclusively investigated it in the white-box setting, while our work extends it to a more practical black-box setting: transferable attack. It is widely reported that stro…

Cited by 51PDFcodeScholar
2021

Correlate-and-Excite: Real-Time Stereo Matching via Guided Cost Volume Excitation

IROS 2021poster

Volumetric deep learning approach towards stereo matching aggregates a cost volume computed from input left and right images using 3D convolutions. Recent works showed that utilization of extracted image features and a spatially varying cost volume aggregation complements 3D convolutions. However, e…

Cited by 85SourcecodeScholar
2021

LabOR: Labeling Only if Required for Domain Adaptive Semantic Segmentation

ICCV 2021poster

Unsupervised Domain Adaptation (UDA) for semantic segmentation has been actively studied to mitigate the domain gap between label-rich source data and unlabeled target data. Despite these efforts, UDA still has a long way to go to reach the fully supervised performance. To this end, we propose a Lab…

Cited by 55PDFScholar
2021

Optical Flow Estimation from a Single Motion-blurred Image

AAAI 2021technical

In most of computer vision applications, motion blur is regarded as an undesirable artifact. However, it has been shown that motion blur in an image may have practical interests in fundamental computer vision problems. In this work, we propose a novel framework to estimate optical flow from a single…

Cited by 19SourcePDFScholar