← Search

Junsung Park

10 accepted papers

2026

Contextualized Visual Personalization in Vision-Language Models

ICML 2026poster

Despite recent progress in vision-language models (VLMs), existing approaches often fail to generate personalized responses based on the user's specific experiences, as they lack the ability to associate visual inputs with a user’s accumulated visual-textual context. We newly formalize this challeng…

Cited by 0SourceScholar
2025

3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation

EMNLP 2025

Vision-Language Models (VLMs) have shown remarkable performance on diverse visual and linguistic tasks, yet they remain fundamentally limited in their understanding of 3D spatial structures.We propose Geometric Distillation, a lightweight, annotation-free fine-tuning framework that injects human-ins

2025

Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP

ICCV 2025poster

While CLIP has significantly advanced multimodal understanding by bridging vision and language, the inability to grasp negation -- such as failing to differentiate concepts like "parking" from "no parking" -- poses substantial challenges.By analyzing the data used in the public CLIP model's pre-trai…

Cited by 0SourcePDFScholar
2025

No Thing, Nothing: Highlighting Safety-Critical Classes for Robust LiDAR Semantic Segmentation in Adverse Weather

CVPR 2025poster

Existing domain generalization methods for LiDAR semantic segmentation under adverse weather struggle to accurately predict "things" categories compared to "stuff" categories. In typical driving scenes, "things" categories can be dynamic and associated with higher collision risks, making them crucia…

Cited by 0SourcePDFScholar
2025

Unleashing Multi-Hop Reasoning Potential in Large Language Models through Repetition of Misordered Context

NAACL 2025findings

Multi-hop reasoning, which requires multi-step reasoning based on the supporting documents within a given context, remains challenging for large language models (LLMs). LLMs often struggle to filter out irrelevant documents within the context, and their performance is sensitive to the absolute posit…

Cited by 0SourcePDFScholar
2024

DAFA: Distance-Aware Fair Adversarial Training

ICLR 2024poster

The disparity in accuracy between classes in standard training is amplified during adversarial training, a phenomenon termed the robust fairness problem. Existing methodologies aimed to enhance robust fairness by sacrificing the model's performance on easier classes in order to improve its performan…

2024

Entropy is not Enough for Test-Time Adaptation: From the Perspective of Disentangled Factors

ICLR 2024spotlight

Test-time adaptation (TTA) fine-tunes pre-trained deep neural networks for unseen test data. The primary challenge of TTA is limited access to the entire test dataset during online updates, causing error accumulation. To mitigate it, TTA methods have utilized the model output's entropy as a confiden…

2024

Interactive Text-to-Image Retrieval with Large Language Models: A Plug-and-Play Approach

ACL 2024long

In this paper, we primarily address the issue of dialogue-form context query within the interactive text-to-image retrieval task. Our methodology, PlugIR, actively utilizes the general instruction-following capability of LLMs in two ways. First, by reformulating the dialogue-form context, we elimina…

2024

Rethinking Data Augmentation for Robust LiDAR Semantic Segmentation in Adverse Weather

ECCV 2024oral

"Existing LiDAR semantic segmentation methods often struggle with performance declines in adverse weather conditions. Previous work has addressed this issue by simulating adverse weather or employing universal data augmentation during training. However, these methods lack a detailed analysis and und…

2023

PUCA: Patch-Unshuffle and Channel Attention for Enhanced Self-Supervised Image Denoising

NeurIPS 2023poster

Although supervised image denoising networks have shown remarkable performance on synthesized noisy images, they often fail in practice due to the difference between real and synthesized noise. Since clean-noisy image pairs from the real world are extremely costly to gather, self-supervised learning…

Cited by 17SourcePDFScholar