← Search

Sooel Son

6 accepted papers

2026

No Caption, No Problem: Caption-Free Membership Inference via Model-Fitted Embeddings

ICLR 2026poster

Latent diffusion models have achieved remarkable success in high-fidelity text-to-image generation, but their tendency to memorize training data raises critical privacy and intellectual property concerns. Membership inference attacks (MIAs) provide a principled way to audit such memorization by dete…

Cited by 0SourcecodeScholar
2026

SafeMoE: Safe Fine-Tuning for MoE LLMs by Aligning Harmful Input Routing

ICLR 2026poster

Recent large language models (LLMs) have increasingly adopted the Mixture-of-Experts (MoE) architecture for efficiency. MoE-based LLMs heavily depend on a superficial safety mechanism in which harmful inputs are routed safety-critical experts. However, our analysis reveals that routing decisions for…

Cited by 0SourcecodeScholar
2025

AdvPaint: Protecting Images from Inpainting Manipulation via Adversarial Attention Disruption

ICLR 2025poster

The outstanding capability of diffusion models in generating high-quality images poses significant threats when misused by adversaries. In particular, we assume malicious adversaries exploiting diffusion models for inpainting tasks, such as replacing a specific region with a celebrity. While existin…

2023

Effective Targeted Attacks for Adversarial Self-Supervised Learning

NeurIPS 2023poster

Recently, unsupervised adversarial training (AT) has been highlighted as a means of achieving robustness in models without any label information. Previous studies in unsupervised AT have mostly focused on implementing self-supervised learning (SSL) frameworks, which maximize the instance-wise classi…

Cited by 4SourcePDFScholar
2022

Learning to Generate Inversion-Resistant Model Explanations

NeurIPS 2022accept

The wide adoption of deep neural networks (DNNs) in mission-critical applications has spurred the need for interpretable models that provide explanations of the model's decisions. Unfortunately, previous studies have demonstrated that model explanations facilitate information leakage, rendering DNN…

Cited by 3SourcePDFScholar