← Search

Changick Kim

39 accepted papers

2026

Jailbreak to Protect: Buffering Harmful Fine-Tuning via Temporary Jailbreaking LoRA in Large Language Models

ICML 2026spotlight

Fine-tuning-as-a-Service (FaaS) enables personalization of large language models (LLMs) but poses significant safety risks, as fine-tuning user-provided data degrades the model's safety-alignment. Prior works addressing this issue typically rely on explicit regularization, which leads to practical l…

Cited by 0SourceScholar
2025

Difficulty-aware Balancing Margin Loss for Long-tailed Recognition

AAAI 2025technical

When trained with severely imbalanced data, deep neural networks often struggle to accurately recognize classes with few samples. Previous studies in long-tailed recognition have attempted to rebalance biased learning using known sample distributions, primarily addressing different classification di…

2025

Diffusion Model Patching via Mixture-of-Prompts

AAAI 2025technical

We present Diffusion Model Patching (DMP), a simple method to boost the performance of pre-trained diffusion models that have already reached convergence, with a negligible increase in parameters. DMP inserts a small, learnable set of prompts into the model's input space while keeping the original m…

Cited by 0SourcePDFScholar
2025

Don’t Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models

ACL 2025finding

Large Vision Language Models (LVLMs) demonstrate strong capabilities in visual understanding and description, yet often suffer from hallucinations, attributing incorrect or misleading features to images. We observe that LVLMs disproportionately focus on a small subset of image tokens—termed blind to…

2025

Enhancing Robustness in Incremental Learning with Adversarial Training

AAAI 2025technical

Adversarial training is one of the most effective approaches against adversarial attacks. However, adversarial training has primarily been studied in scenarios where data for all classes is provided, with limited research conducted in the context of incremental learning where knowledge is introduced…

2025

Parameter Efficient Mamba Tuning via Projector-targeted Diagonal-centric Linear Transformation

CVPR 2025poster

Despite the growing interest in Mamba architecture as a potential replacement for Transformer architecture, parameter-efficient fine-tuning (PEFT) approaches for Mamba remain largely unexplored. In our study, we introduce two key insights-driven strategies for PEFT in Mamba architecture: (1) While s…

Cited by 0SourcePDFScholar
2025

SAFIRE: Segment Any Forged Image Region

AAAI 2025technical

Most techniques approach the problem of image forgery localization as a binary segmentation task, training neural networks to label original areas as 0 and forged areas as 1. In contrast, we tackle this issue from a more fundamental perspective by partitioning images according to their originating s…

2025

SplatFlow: Multi-View Rectified Flow Model for 3D Gaussian Splatting Synthesis

CVPR 2025poster

Text-based generation and editing of 3D scenes hold significant potential for streamlining content creation through intuitive user interactions. While recent advances leverage 3D Gaussian Splatting (3DGS) for high-fidelity and real-time rendering, existing methods are often specialized and task-focu…

2025

SteerX: Creating Any Camera-Free 3D and 4D Scenes with Geometric Steering

ICCV 2025poster

Recent progress in 3D/4D scene generation emphasizes the importance of physical alignment throughout video generation and scene reconstruction. However, existing methods improve the alignment separately at each stage, making it difficult to manage subtle misalignments arising from another stage. Her…

2025

VideoRFSplat: Direct Scene-Level Text-to-3D Gaussian Splatting Generation with Flexible Pose and Multi-View Joint Modeling

ICCV 2025poster

We propose VideoRFSplat, a direct text-to-3D model leveraging a video generation model to generate realistic 3D Gaussian Splatting (3DGS) for unbounded real-world scenes. To generate diverse camera poses and unbounded spatial extent of real-world scenes, while ensuring generalization to arbitrary te…

2024

Denoising Task Routing for Diffusion Models

ICLR 2024poster

Diffusion models generate highly realistic images by learning a multi-step denoising process, naturally embodying the principles of multi-task learning (MTL). Despite the inherent connection between diffusion models and MTL, there remains an unexplored area in designing neural architectures that exp…

2024

Flow-Assisted Motion Learning Network for Weakly-Supervised Group Activity Recognition

ECCV 2024poster

"Weakly-Supervised Group Activity Recognition (WSGAR) aims to understand the activity performed together by a group of individuals with the video-level label and without actor-level labels. We propose Flow-Assisted Motion Learning Network () for WSGAR, which consists of the motion-aware actor encode…

Cited by 1SourcePDFScholar
2024

HarmonyView: Harmonizing Consistency and Diversity in One-Image-to-3D

CVPR 2024poster

Recent progress in single-image 3D generation highlights the importance of multi-view coherency leveraging 3D priors from large-scale diffusion models pretrained on Internet-scale images. However the aspect of novel-view diversity remains underexplored within the research landscape due to the ambigu…

2024

Spatio-Temporal Proximity-Aware Dual-Path Model for Panoramic Activity Recognition

ECCV 2024poster

"Panoramic Activity Recognition (PAR) seeks to identify diverse human activities across different scales, from individual actions to social group and global activities in crowded panoramic scenes. PAR presents two major challenges: 1) recognizing the nuanced interactions among numerous individuals a…

Cited by 0SourcePDFScholar
2024

VideoMamba: Spatio-Temporal Selective State Space Model

ECCV 2024poster

"We introduce VideoMamba, a novel adaptation of the pure Mamba architecture, specifically designed for video recognition. Unlike transformers that rely on self-attention mechanisms leading to high computational costs by quadratic complexity, VideoMamba leverages Mamba’s linear complexity and selecti…

2023

Breaking Temporal Consistency: Generating Video Universal Adversarial Perturbations Using Image Models

ICCV 2023poster

As video analysis using deep learning models becomes more widespread, the vulnerability of such models to adversarial attacks is becoming a pressing concern. In particular, Universal Adversarial Perturbation (UAP) poses a significant threat, as a single perturbation can mislead deep learning model…

Cited by 6PDFScholar
2023

Introducing Competition To Boost the Transferability of Targeted Adversarial Examples Through Clean Feature Mixup

CVPR 2023poster

Deep neural networks are widely known to be susceptible to adversarial examples, which can cause incorrect predictions through subtle input modifications. These adversarial examples tend to be transferable between models, but targeted attacks still have lower attack success rates due to significant…

2023

Liveness Score-Based Regression Neural Networks for Face Anti-Spoofing

ICASSP 2023accepted

Previous anti-spoofing methods have used either pseudo maps or user-defined labels, and the performance of each approach depends on the accuracy of the third party networks generating pseudo maps and the way in which the users define the labels. In this paper, we propose a liveness score-based regre…

Cited by 0SourceScholar
2023

PG-RCNN: Semantic Surface Point Generation for 3D Object Detection

ICCV 2023poster

One of the main challenges in LiDAR-based 3D object detection is that the sensors often fail to capture the complete spatial information about the objects due to long distance and occlusion. Two-stage detectors with point cloud completion approaches tackle this problem by adding more points to the…

Cited by 43PDFcodeScholar
2023

ProtoFL: Unsupervised Federated Learning via Prototypical Distillation

ICCV 2023poster

Federated learning (FL) is a promising approach for enhancing data privacy preservation, particularly for authentication systems. However, limited round communications, scarce representation, and scalability pose significant challenges to its deployment, hindering its full potential. In this paper,…

Cited by 13PDFScholar
2023

Towards Good Practices for Missing Modality Robust Action Recognition

AAAI 2023technical

Standard multi-modal models assume the use of the same modalities in training and inference stages. However, in practice, the environment in which multi-modal models operate may not satisfy such assumption. As such, their performances degrade drastically if any modality is missing in the inference s…

2022

Improving the Transferability of Targeted Adversarial Examples Through Object-Based Diverse Input

CVPR 2022poster

The transferability of adversarial examples allows the deception on black-box models, and transfer-based targeted attacks have attracted a lot of interest due to their practical applicability. To maximize the transfer success rate, adversarial examples should avoid overfitting to the source model, a…

Cited by 82PDFcodeScholar
2021

DnD: Dense Depth Estimation in Crowded Dynamic Indoor Scenes

ICCV 2021poster

We present a novel approach for estimating depth from a monocular camera as it moves through complex and crowded indoor environments, e.g., a department store or a metro station. Our approach predicts absolute scale depth maps over the entire scene consisting of a static background and multiple movi…

Cited by 6PDFScholar
2021

Just a Few Points Are All You Need for Multi-View Stereo: A Novel Semi-Supervised Learning Method for Multi-View Stereo

ICCV 2021poster

While learning-based multi-view stereo (MVS) methods have recently shown successful performances in quality and efficiency, limited MVS data hampers generalization to unseen environments. A simple solution is to generate various large-scale MVS datasets, but generating dense ground truth for 3D stru…

Cited by 8PDFScholar
2021

Meta Batch-Instance Normalization for Generalizable Person Re-Identification

CVPR 2021poster

Although supervised person re-identification (Re-ID) methods have shown impressive performance, they suffer from a poor generalization capability on unseen domains. Therefore, generalizable Re-ID has recently attracted growing attention. Many existing methods have employed an instance normalization…

Cited by 185PDFcodeScholar
2020

Attract, Perturb, and Explore: Learning a Feature Alignment Network for Semi-supervised Domain Adaptation

ECCV 2020poster

Perturb, and Explore: Learning a Feature Alignment Network for Semi-supervised Domain Adaptation","Although unsupervised domain adaptation methods have been widely adopted across several computer vision tasks, it is more desirable if we can exploit a few labeled data from new domains encountered in…

Cited by 155SourcePDFScholar
2020

Hi-CMD: Hierarchical Cross-Modality Disentanglement for Visible-Infrared Person Re-Identification

CVPR 2020poster

Visible-infrared person re-identification (VI-ReID) is an important task in night-time surveillance applications, since visible cameras are difficult to capture valid appearance information under poor illumination conditions. Compared to traditional person re-identification that handles only the int…

Cited by 412PDFcodeScholar
2020

Learning to Discriminate Information for Online Action Detection

CVPR 2020poster

From a streaming video, online action detection aims to identify actions in the present. For this task, previous methods use recurrent networks to model the temporal sequence of current action frames. However, these methods overlook the fact that an input image sequence includes background and irrel…

Cited by 100PDFScholar
2019

Diversify and Match: A Domain Adaptive Representation Learning Paradigm for Object Detection

CVPR 2019poster

We introduce a novel unsupervised domain adaptation approach for object detection. We aim to alleviate the imperfect translation problem of pixel-level adaptations, and the source-biased discriminativity problem of feature-level adaptations simultaneously. Our approach is composed of two stages, i.e…

Cited by 384PDFScholar
2019

Self-Ensembling With GAN-Based Data Augmentation for Domain Adaptation in Semantic Segmentation

ICCV 2019poster

Deep learning-based semantic segmentation methods have an intrinsic limitation that training a model requires a large amount of data with pixel-level annotations. To address this challenging issue, many researchers give attention to unsupervised domain adaptation for semantic segmentation. Unsupervi…

Cited by 328PDFScholar
2019

Self-Training and Adversarial Background Regularization for Unsupervised Domain Adaptive One-Stage Object Detection

ICCV 2019oral

Deep learning-based object detectors have shown remarkable improvements. However, supervised learning-based methods perform poorly when the train data and the test data have different distributions. To address the issue, domain adaptation transfers knowledge from the label-sufficient domain (source…

Cited by 263PDFScholar