← Search

Bumsub Ham

36 accepted papers

2026

Improving Visual Token Reduction via Rectifying Distortions for Efficient Multimodal LLM Inference

ICML 2026poster

Recent advancements in Multimodal Large Language Models (MLLMs) have achieved remarkable success in vision-language tasks, yet the quadratic computational complexity arising from the vast number of visual tokens creates significant memory and latency bottlenecks. While visual token reduction (VTR) s…

Cited by 0SourceScholar
2026

Jailbreaking on Text-to-Video Models via Scene Splitting Strategy

ICLR 2026poster

Along with the rapid advancement of numerous Text-to-Video (T2V) models, growing concerns have emerged regarding their safety risks. While recent studies have explored vulnerabilities in models like LLMs, VLMs, and Text-to-Image (T2I) models through jailbreak attacks, T2V models remain largely unexp…

Cited by 0SourcecodeScholar
2026

Relational Feature Caching for Accelerating Diffusion Transformers

ICLR 2026poster

Feature caching approaches accelerate diffusion transformers (DiTs) by storing the output features of computationally expensive modules at certain timesteps, and exploiting them for subsequent steps to reduce redundant computations. Recent forecasting-based caching approaches employ temporal extrapo…

Cited by 0SourceScholar
2025

AccuQuant: Simulating Multiple Denoising Steps for Quantizing Diffusion Models

NeurIPS 2025poster

We present in this paper a novel post-training quantization (PTQ) method, dubbed AccuQuant, for diffusion models. We show analytically and empirically that quantization errors for diffusion models are accumulated over denoising steps in a sampling process. To alleviate the error accumulation problem…

Cited by 0SourceScholar
2025

ELITE: Enhanced Language-Image Toxicity Evaluation for Safety

ICML 2025poster

Current Vision Language Models (VLMs) remain vulnerable to malicious prompts that induce harmful outputs. Existing safety benchmarks for VLMs primarily rely on automated evaluation methods, but these methods struggle to detect implicit harmful content or produce inaccurate evaluations. Therefore, we…

Cited by 0SourcePDFScholar
2025

Efficient Few-Shot Neural Architecture Search by Counting the Number of Nonlinear Functions

AAAI 2025technical

Neural architecture search (NAS) enables finding the best-performing architecture from a search space automatically. Most NAS methods exploit an over-parameterized network (i.e., a supernet) containing all possible architectures (i.e., subnets) in the search space. However, the subnets that share th…

Cited by 1SourcePDFScholar
2025

Maximizing the Position Embedding for Vision Transformers with Global Average Pooling

AAAI 2025technical

In vision transformers, position embedding (PE) plays a crucial role in capturing the order of tokens. However, in vision transformer structures, there is a limitation in the expressiveness of PE due to the structure where position embedding is simply added to the token embedding. A layer-wise metho…

2025

Subnet-Aware Dynamic Supernet Training for Neural Architecture Search

CVPR 2025poster

N-shot neural architecture search (NAS) exploits a supernet containing all candidate subnets for a given search space. The subnets are typically trained with a static training strategy (e.g., using the same learning rate (LR) scheduler and optimizer for all subnets). This, however, does not consider…

Cited by 0SourcePDFScholar
2023

ACLS: Adaptive and Conditional Label Smoothing for Network Calibration

ICCV 2023oral

We address the problem of network calibration adjusting miscalibrated confidences of deep neural networks. Many approaches to network calibration adopt a regularization-based method that exploits a regularization term to smooth the miscalibrated confidences. Although these approaches have shown the…

Cited by 26PDFScholar
2023

Camera-Driven Representation Learning for Unsupervised Domain Adaptive Person Re-identification

ICCV 2023poster

We present a novel unsupervised domain adaption method for person re-identification (reID) that generalizes a model trained on a labeled source domain to an unlabeled target domain. We introduce a camera-driven curriculum learning (CaCL) framework that leverages camera labels of person images to tra…

Cited by 40PDFScholar
2022

ALIFE: Adaptive Logit Regularizer and Feature Replay for Incremental Semantic Segmentation

NeurIPS 2022accept

We address the problem of incremental semantic segmentation (ISS) recognizing novel object/stuff categories continually without forgetting previous ones that have been learned. The catastrophic forgetting problem is particularly severe in ISS, since pixel-level ground-truth labels are available only…

Cited by 29SourcePDFScholar
2022

Bi-directional Contrastive Learning for Domain Adaptive Semantic Segmentation

ECCV 2022poster

"We present a novel unsupervised domain adaptation method for semantic segmentation that generalizes a model trained with source images and corresponding ground-truth labels to a target domain. A key to domain adaptive semantic segmentation is to learn domain-invariant and discriminative features wi…

Cited by 35SourcePDFScholar
2022

Decomposed Knowledge Distillation for Class-Incremental Semantic Segmentation

NeurIPS 2022accept

Class-incremental semantic segmentation (CISS) labels each pixel of an image with a corresponding object/stuff class continually. To this end, it is crucial to learn novel classes incrementally without forgetting previously learned knowledge. Current CISS methods typically use a knowledge distillati…

Cited by 38SourcePDFScholar
2022

OIMNet++: Prototypical Normalization and Localization-Aware Learning for Person Search

ECCV 2022poster

"We address the task of person search, that is, localizing and re-identifying query persons from a set of raw scene images. Recent approaches are typically built upon OIMNet, a pioneer work on person search, that learns joint person representations for performing both detection and person re-identif…

2021

Background-Aware Pooling and Noise-Aware Loss for Weakly-Supervised Semantic Segmentation

CVPR 2021poster

We address the problem of weakly-supervised semantic segmentation (WSSS) using bounding box annotations. Although object bounding boxes are good indicators to segment corresponding objects, they do not specify object boundaries, making it hard to train convolutional neural networks (CNNs) for semant…

Cited by 121PDFScholar
2021

HVPR: Hybrid Voxel-Point Representation for Single-Stage 3D Object Detection

CVPR 2021poster

We address the problem of 3D object detection, that is, estimating 3D object bounding boxes from point clouds. 3D object detection methods exploit either voxel-based or point-based features to represent 3D objects in a scene. Voxel-based features are efficient to extract, while they fail to preserve…

Cited by 169PDFcodeScholar
2021

Learning by Aligning: Visible-Infrared Person Re-Identification Using Cross-Modal Correspondences

ICCV 2021poster

We address the problem of visible-infrared person re-identification (VI-reID), that is, retrieving a set of person images, captured by visible or infrared cameras, in a cross-modal setting. Two main challenges in VI-reID are intra-class variations across person images, and cross-modal discrepancies…

Cited by 246PDFScholar
2021

Video-Based Person Re-Identification With Spatial and Temporal Memory Networks

ICCV 2021poster

Video-based person re-identification (reID) aims to retrieve person videos with the same identity as a query person across multiple cameras. Spatial and temporal distractors in person videos, such as background clutter and partial occlusions over frames, respectively, make this task much more challe…

Cited by 103PDFcodeScholar
2020

Learning with Privileged Information for Efficient Image Super-Resolution

ECCV 2020poster

Convolutional neural networks (CNNs) have allowed remarkable advances in single image super-resolution (SISR) over the last decade. Most SR methods based on CNNs have focused on achieving performance gains in terms of quality metrics, such as PSNR and SSIM, over classical approaches. They typically…

Cited by 162SourcePDFScholar
2017

FCSS: Fully Convolutional Self-Similarity for Dense Semantic Correspondence

CVPR 2017poster

We present a descriptor, called fully convolutional self-similarity (FCSS), for dense semantic correspondence. To robustly match points among different instances within the same object class, we formulate FCSS using local self-similarity (LSS) within a fully convolutional network. In contrast to exi…

Cited by 176PDFScholar
2017

SCNet: Learning Semantic Correspondence

ICCV 2017poster

This paper addresses the problem of establishing semantic correspondences between images depicting different instances of the same object or scene category. Previous approaches focus on either combining a spatial regularizer with hand-crafted features, or learning a correspondence model for appearan…

Cited by 159PDFcodeScholar
2016

Proposal Flow

CVPR 2016poster

Finding image correspondences remains a challenging problem in the presence of intra-class variations and large changes in scene layout. Semantic flow methods are designed to handle images depicting different instances of the same object or scene category. We introduce a novel approach to semantic…

Cited by 164PDFScholar
2015

DASC: Dense Adaptive Self-Correlation Descriptor for Multi-Modal and Multi-Spectral Correspondence

CVPR 2015poster

Establishing dense visual correspondence between multiple images is a fundamental task in many applications of computer vision and computational photography. Classical approaches, which aim to estimate dense stereo and optical flow fields for images adjacent in viewpoint or in time, have been dramat…

Cited by 119SourcePDFScholar