← Search

Dohwan Ko

8 accepted papers

2026

DocPrune: Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruning

CVPR 2026

Recent advances in vision-language models have shown strong performance across diverse multimodal tasks, including document question answering that leverages structured visual cues from text, tables, and figures. However, unlike natural images, document images contain large backgrounds and only spar

Cited by 0SourceScholar
2026

MoE-GRPO: Optimizing Mixture-of-Experts via Reinforcement Learning in Vision-Language Models

CVPR 2026

Mixture-of-Experts (MoE) has emerged as an effective approach to reduce the computational overhead of Transformer architectures by sparsely activating a subset of parameters for each token while preserving high model capacity. This paradigm has recently been extended to Vision-Language Models (VLMs)

Cited by 0SourceScholar
2025

Bidirectional Likelihood Estimation with Multi-Modal Large Language Models for Text-Video Retrieval

ICCV 2025poster

Text-Video Retrieval aims to find the most relevant text (or video) candidate given a video (or text) query from large-scale online databases. Recent work leverages multi-modal large language models (MLLMs) to improve retrieval, especially for long or complex query-candidate pairs. However, we obser…

2024

LLaMo: Large Language Model-based Molecular Graph Assistant

NeurIPS 2024poster

Large Language Models (LLMs) have demonstrated remarkable generalization and instruction-following capabilities with instruction tuning. The advancements in LLMs and instruction tuning have led to the development of Large Vision-Language Models (LVLMs). However, the competency of the LLMs and instru…

2023

Large Language Models are Temporal and Causal Reasoners for Video Question Answering

EMNLP 2023long main

Large Language Models (LLMs) have shown remarkable performances on a wide range of natural language understanding and generation tasks. We observe that the LLMs provide effective priors in exploiting $\textit{linguistic shortcuts}$ for temporal and causal reasoning in Video Question Answering (Video…

Cited by 0SourcecodeScholar
2023

MELTR: Meta Loss Transformer for Learning To Fine-Tune Video Foundation Models

CVPR 2023poster

Foundation models have shown outstanding performance and generalization capabilities across domains. Since most studies on foundation models mainly focus on the pretraining phase, a naive strategy to minimize a single task-specific loss is adopted for fine-tuning. However, such fine-tuning methods d…

2023

Open-vocabulary Video Question Answering: A New Benchmark for Evaluating the Generalizability of Video Question Answering Models

ICCV 2023poster

Video Question Answering (VideoQA) is a challenging task that entails complex multi-modal reasoning. In contrast to multiple-choice VideoQA which aims to predict the answer given several options, the goal of open-ended VideoQA is to answer questions without restricting candidate answers. However, th…

Cited by 7PDFcodeScholar
2022

Video-Text Representation Learning via Differentiable Weak Temporal Alignment

CVPR 2022poster

Learning generic joint representations for video and text by a supervised method requires a prohibitively substantial amount of manually annotated video datasets. As a practical alternative, a large-scale but uncurated and narrated video dataset, HowTo100M, has recently been introduced. But it is st…

Cited by 25PDFcodeScholar