← Search

Hyunwoo J Kim

68 accepted papers

2026

DocPrune: Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruning

CVPR 2026

Recent advances in vision-language models have shown strong performance across diverse multimodal tasks, including document question answering that leverages structured visual cues from text, tables, and figures. However, unlike natural images, document images contain large backgrounds and only spar

Cited by 0SourceScholar
2026

Error as Signal: Stiffness-Aware Diffusion Sampling via Embedded Runge-Kutta Guidance

ICLR 2026poster

Classifier-Free Guidance (CFG) has established the foundation for guidance mechanisms in diffusion models, showing that well-designed guidance proxies significantly improve conditional generation and sample quality. Autoguidance (AG) has extended this idea, but it relies on an auxiliary network and…

Cited by 0SourcecodeScholar
2026

Improving Large Molecular Language Model via Relation-aware Multimodal Collaboration

AAAI 2026technical

Large language models (LLMs) have demonstrated their instruction-following capabilities and achieved powerful performance on various tasks. Inspired by their success, recent works in the molecular domain have led to the development of large molecular language models (LMLMs) that integrate 1D molecul

Cited by 0SourcePDFScholar
2026

MoE-GRPO: Optimizing Mixture-of-Experts via Reinforcement Learning in Vision-Language Models

CVPR 2026

Mixture-of-Experts (MoE) has emerged as an effective approach to reduce the computational overhead of Transformer architectures by sparsely activating a subset of parameters for each token while preserving high model capacity. This paradigm has recently been extended to Vision-Language Models (VLMs)

Cited by 0SourceScholar
2026

RegFormer: Transferable Relational Grounding for Efficient Weakly-Supervised Human-Object Interaction Detection

CVPR 2026

Weakly-supervised Human-Object Interaction (HOI) detection is essential for scalable scene understanding, as it learns interactions from only image-level annotations. Due to the lack of localization signals, prior works typically rely on an external object detector to generate candidate pairs and th

Cited by 0SourcecodeScholar
2026

SPRINT: Sparse-Dense Residual Fusion for Efficient Diffusion Transformers

ICLR 2026poster

Diffusion Transformers (DiTs) deliver state-of-the-art generative performance but their quadratic training cost with sequence length makes large-scale pretraining prohibitively expensive. Token dropping can reduce training cost, yet naïve strategies degrade representations, and existing methods are…

Cited by 0SourcecodeScholar
2026

TabFlash: Efficient Table Understanding with Progressive Question Conditioning and Token Focusing

AAAI 2026technical

Table images present unique challenges for effective and efficient understanding due to the need for question-specific focus and the presence of redundant background regions. Existing Multimodal Large Language Model (MLLM) approaches often overlook these characteristics, resulting in uninformative a

Cited by 0SourcePDFScholar
2026

Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization

AAAI 2026technical

Vision-Language Models (VLMs) have been widely used in various visual recognition tasks due to their remarkable generalization capabilities. As these models grow in size and complexity, fine-tuning becomes costly, emphasizing the need to reuse adaptation knowledge from

Cited by 0SourcePDFScholar
2025

Bidirectional Likelihood Estimation with Multi-Modal Large Language Models for Text-Video Retrieval

ICCV 2025poster

Text-Video Retrieval aims to find the most relevant text (or video) candidate given a video (or text) query from large-scale online databases. Recent work leverages multi-modal large language models (MLLMs) to improve retrieval, especially for long or complex query-candidate pairs. However, we obser…

2025

Blockwise Flow Matching: Improving Flow Matching Models For Efficient High-Quality Generation

NeurIPS 2025poster

Recently, Flow Matching models have pushed the boundaries of high-fidelity data generation across a wide range of domains. It typically employs a single large network to learn the entire generative trajectory from noise to data. Despite their effectiveness, this design struggles to capture distinct…

Cited by 0SourceScholar
2025

Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization

EMNLP 2025

In text-video retrieval, auxiliary captions are often used to enhance video understanding, bridging the gap between the modalities. While recent advances in multi-modal large language models (MLLMs) have enabled strong zero-shot caption generation, we observe that such captions tend to be generic an

2025

DeepVideo-R1: Video Reinforcement Fine-Tuning via Difficulty-aware Regressive GRPO

NeurIPS 2025poster

Recent works have demonstrated the effectiveness of reinforcement learning (RL)-based post-training for enhancing the reasoning capabilities of large language models (LLMs). In particular, Group Relative Policy Optimization (GRPO) has shown impressive success using a PPO-style reinforcement algorith…

Cited by 0SourceScholar
2025

EfficientViM: Efficient Vision Mamba with Hidden State Mixer based State Space Duality

CVPR 2025poster

For the deployment of neural networks in resource-constrained environments, prior works have built lightweight architectures with convolution and attention for capturing local and global dependencies, respectively. Recently, the state space model (SSM) has emerged as an effective operation for globa…

2025

Latent Bayesian Optimization via Autoregressive Normalizing Flows

ICLR 2025oral

Bayesian Optimization (BO) has been recognized for its effectiveness in optimizing expensive and complex objective functions. Recent advancements in Latent Bayesian Optimization (LBO) have shown promise by integrating generative models such as variational autoencoders (VAEs) to manage the complexity…

Cited by 1SourcePDFScholar
2025

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

ICML 2025poster

Multimodal Large Language Models (MLLMs) have shown promising progress in understanding and analyzing video content. However, processing long videos remains a significant challenge constrained by LLM's context size. To address this limitation, we propose \textbf{LongVU}, a spatiotemporal adaptive co…

2025

Multidimensional Adaptive Coefficient for Inference Trajectory Optimization in Flow and Diffusion

ICML 2025poster

Flow and diffusion models have demonstrated strong performance and training stability across various tasks but lack two critical properties of simulation-based methods: freedom of dimensionality and adaptability to different inference trajectories. To address this limitation, we propose the Multidim…

Cited by 0SourcePDFScholar
2025

PRESTO: Preimage-Informed Instruction Optimization for Prompting Black-Box LLMs

NeurIPS 2025poster

Large language models (LLMs) have achieved remarkable success across diverse domains, due to their strong instruction-following capabilities. This raised interest in optimizing instructions for black-box LLMs, whose internal parameters are inaccessible but popular for their strong performance and ea…

Cited by 0SourceScholar
2025

Representation Shift: Unifying Token Compression with FlashAttention

ICCV 2025poster

Transformers have demonstrated remarkable success across vision, language, and video. Yet, increasing task complexity has led to larger models and more tokens, raising the quadratic cost of self-attention and the overhead of GPU memory access. To reduce the computation cost of self-attention, prior…

2025

Super-Class Guided Transformer for Zero-Shot Attribute Classification

AAAI 2025technical

Attribute classification is crucial for identifying specific characteristics within image regions. Vision-Language Models (VLMs) have been effective in zero-shot tasks by leveraging their general knowledge from large-scale datasets. Recent studies demonstrate that transformer-based models with class…

2025

VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning

AAAI 2025technical

Despite the advancements of Video Large Language Models (VideoLLMs) in various tasks, they struggle with fine-grained temporal understanding, such as Dense Video Captioning (DVC). DVC is a complicated task of describing all events within a video while also temporally localizing them, which integrate…

2025

Visual Diversity and Region-aware Prompt Learning for Zero-shot HOI Detection

NeurIPS 2025poster

Zero-shot Human-Object Interaction detection aims to localize humans and objects in an image and recognize their interaction, even when specific verb-object pairs are unseen during training. Recent works have shown promising results using prompt learning with pretrained vision-language models such a…

Cited by 0SourcecodeScholar
2025

When Model Knowledge meets Diffusion Model: Diffusion-assisted Data-free Image Synthesis with Alignment of Domain and Class

ICML 2025poster

Open-source pre-trained models hold great potential for diverse applications, but their utility declines when their training data is unavailable. Data-Free Image Synthesis (DFIS) aims to generate images that approximate the learned data distribution of a pre-trained model without accessing the origi…

Cited by 0SourcePDFScholar
2024

Constant Acceleration Flow

NeurIPS 2024poster

Rectified flow and reflow procedures have significantly advanced fast generation by progressively straightening ordinary differential equation (ODE) flows under the assumption that image and noise pairs, known as coupling, can be approximated by straight trajectories with constant velocity. However,…

2024

DDMI: Domain-agnostic Latent Diffusion Models for Synthesizing High-Quality Implicit Neural Representations

ICLR 2024poster

Recent studies have introduced a new class of generative models for synthesizing implicit neural representations (INRs) that capture arbitrary continuous signals in various domains. These models opened the door for domain-agnostic generative models, but they often fail to achieve high-quality genera…

2024

Generative Subgraph Retrieval for Knowledge Graph–Grounded Dialog Generation

EMNLP 2024main

Knowledge graph–grounded dialog generation requires retrieving a dialog-relevant subgraph from the given knowledge base graph and integrating it with the dialog history. Previous works typically represent the graph using an external encoder, such as graph neural networks, and retrieve relevant tripl…

2024

Groupwise Query Specialization and Quality-Aware Multi-Assignment for Transformer-based Visual Relationship Detection

CVPR 2024poster

Visual Relationship Detection (VRD) has seen significant advancements with Transformer-based architectures recently. However we identify two key limitations in a conventional label assignment for training Transformer-based VRD models which is a process of mapping a ground-truth (GT) to a prediction.…

2024

LLaMo: Large Language Model-based Molecular Graph Assistant

NeurIPS 2024poster

Large Language Models (LLMs) have demonstrated remarkable generalization and instruction-following capabilities with instruction tuning. The advancements in LLMs and instruction tuning have led to the development of Large Vision-Language Models (LVLMs). However, the competency of the LLMs and instru…

2024

Multi-criteria Token Fusion with One-step-ahead Attention for Efficient Vision Transformers

CVPR 2024poster

Vision Transformer (ViT) has emerged as a prominent backbone for computer vision. For more efficient ViTs recent works lessen the quadratic cost of the self-attention layer by pruning or fusing the redundant tokens. However these works faced the speed-accuracy trade-off caused by the loss of informa…

2024

Retrieval-Augmented Open-Vocabulary Object Detection

CVPR 2024poster

Open-vocabulary object detection (OVD) has been studied with Vision-Language Models (VLMs) to detect novel objects beyond the pre-trained categories. Previous approaches improve the generalization ability to expand the knowledge of the detector using 'positive' pseudo-labels with additional 'class'…

2024

Stochastic Conditional Diffusion Models for Robust Semantic Image Synthesis

ICML 2024poster

Semantic image synthesis (SIS) is a task to generate realistic images corresponding to semantic maps (labels). However, in real-world applications, SIS often encounters noisy user inputs. To address this, we propose Stochastic Conditional Diffusion Model (SCDM), which is a robust conditional diffusi…

2024

Understanding Multi-compositional learning in Vision and Language models via Category Theory

ECCV 2024poster

"Pre-trained large language models (and multi-modal models) offer excellent performance across a wide range of tasks. Despite their effectiveness, we have limited knowledge of their internal knowledge representation. To get started, we use the classic problem of Compositional Zero-Shot Learning (CZS…

2024

vid-TLDR: Training Free Token Merging for Light-weight Video Transformer

CVPR 2024poster

Video Transformers have become the prevalent solution for various video downstream tasks with superior expressive power and flexibility. However these video transformers suffer from heavy computational costs induced by the massive number of tokens across the entire video frames which has been the ma…

2023

Advancing Bayesian Optimization via Learning Correlated Latent Space

NeurIPS 2023poster

Bayesian optimization is a powerful method for optimizing black-box functions with limited function evaluations. Recent works have shown that optimization in a latent space through deep generative models such as variational autoencoders leads to effective and efficient Bayesian optimization for stru…

2023

Large Language Models are Temporal and Causal Reasoners for Video Question Answering

EMNLP 2023long main

Large Language Models (LLMs) have shown remarkable performances on a wide range of natural language understanding and generation tasks. We observe that the LLMs provide effective priors in exploiting $\textit{linguistic shortcuts}$ for temporal and causal reasoning in Video Question Answering (Video…

Cited by 0SourcecodeScholar
2023

MELTR: Meta Loss Transformer for Learning To Fine-Tune Video Foundation Models

CVPR 2023poster

Foundation models have shown outstanding performance and generalization capabilities across domains. Since most studies on foundation models mainly focus on the pretraining phase, a naive strategy to minimize a single task-specific loss is adopted for fine-tuning. However, such fine-tuning methods d…

2023

NuTrea: Neural Tree Search for Context-guided Multi-hop KGQA

NeurIPS 2023poster

Multi-hop Knowledge Graph Question Answering (KGQA) is a task that involves retrieving nodes from a knowledge graph (KG) to answer natural language questions. Recent GNN-based approaches formulate this task as a KG path searching problem, where messages are sequentially propagated from the seed nod…

2023

Open-vocabulary Video Question Answering: A New Benchmark for Evaluating the Generalizability of Video Question Answering Models

ICCV 2023poster

Video Question Answering (VideoQA) is a challenging task that entails complex multi-modal reasoning. In contrast to multiple-choice VideoQA which aims to predict the answer given several options, the goal of open-ended VideoQA is to answer questions without restricting candidate answers. However, th…

Cited by 7PDFcodeScholar
2023

Read-only Prompt Optimization for Vision-Language Few-shot Learning

ICCV 2023poster

In recent years, prompt tuning has proven effective in adapting pre-trained vision-language models to down- stream tasks. These methods aim to adapt the pre-trained models by introducing learnable prompts while keeping pre- trained weights frozen. However, learnable prompts can affect the internal r…

Cited by 64PDFcodeScholar
2023

Robust Camera Pose Refinement for Multi-Resolution Hash Encoding

ICML 2023poster

Multi-resolution hash encoding has recently been proposed to reduce the computational cost of neural renderings, such as NeRF. This method requires accurate camera poses for the neural renderings of given scenes. However, contrary to previous methods jointly optimizing camera poses and 3D scenes, th…

Cited by 28SourcePDFScholar
2023

Self-Positioning Point-Based Transformer for Point Cloud Understanding

CVPR 2023poster

Transformers have shown superior performance on various computer vision tasks with their capabilities to capture long-range dependencies. Despite the success, it is challenging to directly apply Transformers on point clouds due to their quadratic cost in the number of points. In this paper, we prese…

2023

Semantic-Aware Implicit Template Learning via Part Deformation Consistency

ICCV 2023poster

Learning implicit templates as neural fields has recently shown impressive performance in unsupervised shape correspondence. Despite the success, we observe current approaches, which solely rely on geometric information, often learn suboptimal deformation across generic object shapes, which have hig…

Cited by 4PDFcodeScholar
2022

Consistency Learning via Decoding Path Augmentation for Transformers in Human Object Interaction Detection

CVPR 2022poster

Human-Object Interaction detection is a holistic visual recognition task that entails object detection as well as interaction classification. Previous works of HOI detection has been addressed by the various compositions of subset predictions, e.g., Image -> HO -> I, Image -> HI -> O. Recently, tran…

Cited by 31PDFcodeScholar
2022

Invertible Monotone Operators for Normalizing Flows

NeurIPS 2022accept

Normalizing flows model probability distributions by learning invertible transformations that transfer a simple distribution into complex distributions. Since the architecture of ResNet-based normalizing flows is more flexible than that of coupling-based models, ResNet-based normalizing flows have b…

2022

K-SALSA: K-Anonymous Synthetic Averaging of Retinal Images via Local Style Alignment

ECCV 2022poster

"The application of modern machine learning to retinal image analyses offers valuable insights into a broad range of human health conditions beyond ophthalmic diseases. Additionally, data sharing is key to fully realizing the potential of machine learning models by providing a rich and diverse colle…

2022

SageMix: Saliency-Guided Mixup for Point Clouds

NeurIPS 2022accept

Data augmentation is key to improving the generalization ability of deep learning models. Mixup is a simple and widely-used data augmentation technique that has proven effective in alleviating the problems of overfitting and data scarcity. Also, recent studies of saliency-aware Mixup in the image do…

2022

TokenMixup: Efficient Attention-guided Token-level Data Augmentation for Transformers

NeurIPS 2022accept

Mixup is a commonly adopted data augmentation technique for image classification. Recent advances in mixup methods primarily focus on mixing based on saliency. However, many saliency detectors require intense computation and are especially burdensome for parameter-heavy transformer models. To this e…

2022

Video-Text Representation Learning via Differentiable Weak Temporal Alignment

CVPR 2022poster

Learning generic joint representations for video and text by a supervised method requires a prohibitively substantial amount of manually annotated video datasets. As a practical alternative, a large-scale but uncurated and narrated video dataset, HowTo100M, has recently been introduced. But it is st…

Cited by 25PDFcodeScholar
2021

HOTR: End-to-End Human-Object Interaction Detection With Transformers

CVPR 2021poster

Human-Object Interaction (HOI) detection is a task of identifying "a set of interactions" in an image, which involves the i) localization of the subject (i.e., humans) and target (i.e., objects) of interaction, and ii) the classification of the interaction labels. Most existing methods have addresse…

Cited by 339PDFcodeScholar
2021

Metropolis-Hastings Data Augmentation for Graph Neural Networks

NeurIPS 2021poster

Graph Neural Networks (GNNs) often suffer from weak-generalization due to sparsely labeled data despite their promising results on various graph-based tasks. Data augmentation is a prevalent remedy to improve the generalization ability of models in many domains. However, due to the non-Euclidean nat…

Cited by 62SourcePDFScholar
2021

Neo-GNNs: Neighborhood Overlap-aware Graph Neural Networks for Link Prediction

NeurIPS 2021poster

Graph Neural Networks (GNNs) have been widely applied to various fields for learning over graph-structured data. They have shown significant improvements over traditional heuristic methods in various tasks such as node classification and graph classification. However, since GNNs heavily rely on smoo…

2021

Point Cloud Augmentation With Weighted Local Transformations

ICCV 2021poster

Despite the extensive usage of point clouds in 3D vision, relatively limited data are available for training deep neural networks. Although data augmentation is a standard approach to compensate for the scarcity of data, it has been less explored in the point cloud literature. In this paper, we prop…

Cited by 83PDFcodeScholar
2020

Robust Neural Networks inspired by Strong Stability Preserving Runge-Kutta methods

ECCV 2020poster

Deep neural networks have achieved state-of-the-art performance in a variety of fields. Recent works observe that a class of widely used neural networks can be viewed as the Euler method of numerical discretization. From the numerical discretization perspective, Strong Stability Preserving (SSP) met…

2020

Self-supervised Auxiliary Learning with Meta-paths for Heterogeneous Graphs

NeurIPS 2020poster

Graph neural networks have shown superior performance in a wide range of applications providing a powerful representation of graph-structured data. Recent works show that the representation can be further improved by auxiliary tasks. However, the auxiliary tasks for heterogeneous graphs, which cont…

2020

UnionDet: Union-Level Detector Towards Real-Time Human-Object Interaction Detection

ECCV 2020poster

Recent advances in deep neural networks have achieved significant progress in detecting individual objects from an image. However, object detection is not sufficient to fully understand a visual scene. Towards a deeper visual understanding, the interactions between objects, especially humans and obj…

Cited by 209SourcePDFScholar
2019

Mixed Effects Neural Networks (MeNets) With Applications to Gaze Estimation

CVPR 2019poster

There is much interest in computer vision to utilize commodity hardware for gaze estimation. A number of papers have shown that algorithms based on deep convolutional architectures are approaching accuracies where streaming data from mass-market devices can offer good gaze tracking performance, alth…

Cited by 126PDFcodeScholar
2019

Sampling-free Uncertainty Estimation in Gated Recurrent Units with Applications to Normative Modeling in Neuroimaging

UAI 2019poster

There has recently been a concerted effort to derive mechanisms in vision and machine learning systems to offer uncertainty estimates of the predictions they make. Clearly, there are enormous benefits to a system that is not only accurate but also has a sense for when it is not. Existing proposals c…

Cited by 8SourcePDFScholar
2018

Efficient Relative Attribute Learning using Graph Neural Networks

ECCV 2018poster

A sizable body of work on relative attributes provides compelling evidence that relating pairs of images along a continuum of strength pertaining to a visual attribute yields significant improvements in a wide variety of tasks in vision. In this paper, we show how emerging ideas in graph neural netw…

2018

Tensorize, Factorize and Regularize: Robust Visual Relationship Learning

CVPR 2018poster

Visual relationships provide higher-level information of objects and their relations in an image – this enables a semantic understanding of the scene and helps downstream applications. Given a set of localized objects in some training data, visual relationship detection seeks to detect the most like…

Cited by 74SourcePDFScholar
2017

Maximizing Subset Accuracy with Recurrent Neural Networks in Multi-label Classification

NeurIPS 2017spotlight

Multi-label classification is the task of predicting a set of labels for a given input instance. Classifier chains are a state-of-the-art method for tackling such problems, which essentially converts this problem into a sequential prediction problem, where the labels are first ordered in an arbitrar…

Cited by 232SourcePDFScholar
2017

Riemannian Nonlinear Mixed Effects Models: Analyzing Longitudinal Deformations in Neuroimaging

CVPR 2017poster

Statistical machine learning models that operate on manifold-valued data are being extensively studied in vision, motivated by applications in activity recognition, feature tracking and medical imaging. While non-parametric methods have been relatively well studied in the literature, efficient formu…

Cited by 34PDFScholar
2016

Latent Variable Graphical Model Selection Using Harmonic Analysis: Applications to the Human Connectome Project (HCP)

CVPR 2016spotlight

A major goal of imaging studies such as the (ongoing) Human Connectome Project (HCP) is to characterize the structural network map of the human brain and identify its associations with covariates such as genotype, risk factors, and so on that correspond to an individual. But the set of image derived…

Cited by 8PDFScholar
2015

Interpolation on the Manifold of K Component GMMs

ICCV 2015poster

Probability density functions (PDFs) are fundamental "objects" in mathematics with numerous applications in computer vision, machine learning and medical imaging. The feasibility of basic operations such as computing the distance between two PDFs and estimating a mean of a set of PDFs is a direct fu…

Cited by 6PDFScholar