← Search

Hyunjung Shim

29 accepted papers

2026

Directional Textual Inversion for Personalized Text-to-Image Generation

ICLR 2026poster

Textual Inversion (TI) is an efficient approach to text‑to‑image personalization but often fails on complex prompts. We trace these failures to embedding norm inflation: learned tokens drift to out‑of‑distribution magnitudes, degrading prompt conditioning in pre‑norm Transformers. Empirically, we sh…

Cited by 2SourceScholar
2026

Interpretable Debiasing of Vision-Language Models for Social Fairness

CVPR 2026

The rapid advancement of Vision-Language models (VLMs) has raised growing concerns that their black-box reasoning processes could lead to unintended forms of social bias. Current debiasing approaches focus on mitigating surface-level bias signals through post-hoc learning or test-time algorithms, wh

Cited by 0SourceScholar
2026

Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via Perceptual Perturbation and Reward Modeling

ICML 2026poster

Recent multimodal large language models have demonstrated strong reasoning ability, yet their reliability as automated evaluators remains limited by a critical weakness: when visual evidence conflicts with textual cues, MLLM judges tend to reward plausible narratives over perceptually correct answer…

Cited by 0SourceScholar
2026

Real-Time Long Horizon Air Quality Forecasting via Group-Relative Policy Optimization

CVPR 2026

Accurate long horizon forecasting of particulate matter (PM) concentration fields is essential for operational public health decisions. However, achieving reliable forecasts remains challenging in regions with complex terrain and strong atmospheric dynamics such as East Asia. While foundation models

Cited by 0SourcecodeScholar
2026

Rethinking Direct Preference Optimization in Diffusion Models

AAAI 2026technical

Aligning text-to-image (T2I) diffusion models with human preferences has emerged as a critical research challenge. While Direct Preference Optimization (DPO) has established a foundation for preference learning in large language models (LLMs), its extension to diffusion models remains limited in ali

Cited by 0SourcePDFScholar
2026

SGSoft: Learning Fused Semantic-Geometric Features for 3D Shape Correspondence via Template-Guided Soft Signals

CVPR 2026

Learning dense correspondences across deformable 3D shapes remains a long-standing challenge due to structural variability, non-isometric deformation, and inconsistent topology. Existing methods typically trade off generalization, geometric fidelity, and efficiency. We address this by proposing SGSo

Cited by 0SourceScholar
2026

What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token Merging

ICLR 2026poster

State-of-the-art vision-language models (VLMs) suffer from a critical failure in understanding negation, often referred to as affirmative bias. This limitation is particularly severe in described object detection (DOD) tasks. To address this, we propose two primary contributions: (1) a new dataset p…

Cited by 0SourceScholar
2026

World in a Frame: Understanding Culture Mixing as a New Challenge for Vision-Language Models

CVPR 2026

In a globalized world, cultural elements from diverse origins frequently appear together within a single visual scene. We refer to these as culture mixing scenarios, yet how Large Vision-Language Models (LVLMs) perceive them remains underexplored. We investigate culture mixing as a critical challeng

Cited by 0SourceScholar
2025

3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation

EMNLP 2025

Vision-Language Models (VLMs) have shown remarkable performance on diverse visual and linguistic tasks, yet they remain fundamentally limited in their understanding of 3D spatial structures.We propose Geometric Distillation, a lightweight, annotation-free fine-tuning framework that injects human-ins

2025

Classifier-guided CLIP Distillation for Unsupervised Multi-label Classification

CVPR 2025poster

Multi-label classification is crucial for comprehensive image understanding, yet acquiring accurate annotations is challenging and costly. To address this, a recent study suggests exploiting unsupervised multi-label classification leveraging CLIP, a powerful vision-language model. Despite CLIP's pro…

2025

DGQ: Distribution-Aware Group Quantization for Text-to-Image Diffusion Models

ICLR 2025poster

Despite the widespread use of text-to-image diffusion models across various tasks, their computational and memory demands limit practical applications. To mitigate this issue, quantization of diffusion models has been explored. It reduces memory usage and computational costs by compressing weights…

2025

DreamCatalyst: Fast and High-Quality 3D Editing via Controlling Editability and Identity Preservation

ICLR 2025poster

Score distillation sampling (SDS) has emerged as an effective framework in text-driven 3D editing tasks, leveraging diffusion models for 3D-consistent editing. However, existing SDS-based 3D editing methods suffer from long training times and produce low-quality results. We identify that the root ca…

Cited by 13SourcePDFScholar
2025

Evaluating Image Hallucination in Text-to-Image Generation with Question-Answering

AAAI 2025technical

Despite the impressive success of text-to-image (TTI) models, existing studies overlook the issue of whether these models accurately convey factual information. In this paper, we focus on the problem of image hallucination, where images created by TTI models fail to faithfully depict factual content…

2025

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation

CVPR 2025poster

Open-Vocabulary Part Segmentation (OVPS) is an emerging field for recognizing fine-grained parts in unseen categories. We identify two primary challenges in OVPS: (1) the difficulty in aligning part-level image-text correspondence, and (2) the lack of structural understanding in segmenting object pa…

2025

I0T: Embedding Standardization Method Towards Zero Modality Gap

ACL 2025long

Contrastive Language-Image Pretraining (CLIP) enables zero-shot inference in downstream tasks such as image-text retrieval and classification. However, recent works extending CLIP suffer from the issue of *modality gap*, which arises when the image and text embeddings are projected to disparate mani…

2025

No Thing, Nothing: Highlighting Safety-Critical Classes for Robust LiDAR Semantic Segmentation in Adverse Weather

CVPR 2025poster

Existing domain generalization methods for LiDAR semantic segmentation under adverse weather struggle to accurately predict "things" categories compared to "stuff" categories. In typical driving scenes, "things" categories can be dynamic and associated with higher collision risks, making them crucia…

Cited by 0SourcePDFScholar
2024

Learning from Spatio-temporal Correlation for Semi-Supervised LiDAR Semantic Segmentation

IROS 2024

We address the challenges of the semi-supervised LiDAR segmentation (SSLS) problem, particularly in low-budget scenarios. The two main issues in low-budget SSLS are the poor-quality pseudo-labels for unlabeled data, and the performance drops due to the significant imbalance between ground-truth and

Cited by 1SourcecodeScholar
2024

Understanding Multi-Granularity for Open-Vocabulary Part Segmentation

NeurIPS 2024poster

Open-vocabulary part segmentation (OVPS) is an emerging research area focused on segmenting fine-grained entities using diverse and previously unseen vocabularies. Our study highlights the inherent complexities of part segmentation due to intricate boundaries and diverse granularity, reflecting the…

Cited by 2SourcePDFScholar
2024

Weakly Supervised Semantic Segmentation for Driving Scenes

AAAI 2024technical

State-of-the-art techniques in weakly-supervised semantic segmentation (WSSS) using image-level labels exhibit severe performance degradation on driving scene datasets such as Cityscapes. To address this challenge, we develop a new WSSS framework tailored to driving scene datasets. Based on extensiv…

2022

Commonality in Natural Images Rescues GANs: Pretraining GANs With Generic and Privacy-Free Synthetic Data

CVPR 2022poster

Transfer learning for GANs successfully improves generation performance under low-shot regimes. However, existing studies show that the pretrained model using a single benchmark dataset is not generalized to various target datasets. More importantly, the pretrained model can be vulnerable to copyrig…

Cited by 16PDFcodeScholar
2022

Logit Mixing Training for More Reliable and Accurate Prediction

IJCAI 2022poster

When a person solves the multi-choice problem, she considers not only what is the answer but also what is not the answer. Knowing what choice is not the answer and utilizing the relationships between choices, she can improve the prediction accuracy. Inspired by this human reasoning process, we propo…

Cited by 5SourcePDFScholar
2022

Threshold Matters in WSSS: Manipulating the Activation for the Robust and Accurate Segmentation Model Against Thresholds

CVPR 2022poster

Weakly-supervised semantic segmentation (WSSS) has recently gained much attention for its promise to train segmentation models only with image-level labels. Existing WSSS methods commonly argue that the sparse coverage of CAM incurs the performance bottleneck of WSSS. This paper provides analytical…

Cited by 96PDFcodeScholar
2021

Few-shot Font Generation with Localized Style Representations and Factorization

AAAI 2021technical

Automatic few-shot font generation is a practical and widely studied problem because manual designs are expensive and sensitive to the expertise of designers. Existing few-shot font generation methods aim to learn to disentangle the style and content element from a few reference glyphs, and mainly f…

2021

Multiple Heads Are Better Than One: Few-Shot Font Generation With Multiple Localized Experts

ICCV 2021poster

A few-shot font generation (FFG) method has to satisfy two objectives: the generated images should preserve the underlying global structure of the target character and present the diverse local reference style. Existing FFG methods aim to disentangle content and style either by extracting a universa…

Cited by 100PDFcodeScholar
2021

Railroad Is Not a Train: Saliency As Pseudo-Pixel Supervision for Weakly Supervised Semantic Segmentation

CVPR 2021poster

Existing studies in weakly-supervised semantic segmentation (WSSS) using image-level weak supervision have several limitations: sparse object coverage, inaccurate object boundaries, and co-occurring pixels from non-target objects. To overcome these challenges, we propose a novel framework, namely Ex…

Cited by 309PDFcodeScholar
2021

Rethinking the Truly Unsupervised Image-to-Image Translation

ICCV 2021poster

Every recent image-to-image translation model inherently requires either image-level (i.e. input-output pairs) or set-level (i.e. domain labels) supervision. However, even set-level supervision can be a severe bottleneck for data collection in practice. In this paper, we tackle image-to-image transl…

Cited by 120PDFcodeScholar
2020

Evaluating Weakly Supervised Object Localization Methods Right

CVPR 2020poster

Weakly-supervised object localization (WSOL) has gained popularity over the last years for its promise to train localization models with only image-level labels. Since the seminal WSOL work of class activation mapping (CAM), the field has focused on how to expand the attention regions to cover objec…

Cited by 237PDFcodeScholar
2018

Improved Training of Generative Adversarial Networks Using Representative Features

ICML 2018oral

Despite the success of generative adversarial networks (GANs) for image generation, the trade-off between visual quality and image diversity remains a significant issue. This paper achieves both aims simultaneously by improving the stability of training GANs. The key idea of the proposed approach is…