← Search

Sungmin Han

4 accepted papers

2025

Contrast-CAT: Contrasting Activations for Enhanced Interpretability in Transformer-based Text Classifiers

UAI 2025

Transformers have profoundly influenced AI research, but explaining their decisions remains challenging – even for relatively simpler tasks such as classification – which hinders trust and safe deployment in real-world applications. Although activation-based attribution methods effectively explain t

2024

SwiftThief: Enhancing Query Efficiency of Model Stealing by Contrastive Learning

IJCAI 2024poster

Model-stealing attacks are emerging as a severe threat to AI-based services because an adversary can create models that duplicate the functionality of the black-box AI models inside the services with regular query-based access. To avoid detection or query costs, the model-stealing adversary must con…

Cited by 1SourcePDFScholar
2022

Libra-CAM: An Activation-Based Attribution Based on the Linear Approximation of Deep Neural Nets and Threshold Calibration

IJCAI 2022poster

Universal application of AI has increased the need to explain why an AI model makes a specific decision in a human-understandable form. Among many related works, the class activation map (CAM)-based methods have been successful recently, creating input attribution based on the weighted sum of activa…

2022

Model Stealing Defense against Exploiting Information Leak through the Interpretation of Deep Neural Nets

IJCAI 2022poster

Model stealing techniques allow adversaries to create attack models that mimic the functionality of black-box machine learning models, querying only class membership or probability outcomes. Recently, interpretable AI is getting increasing attention, to enhance our understanding of AI models, provid…

Cited by 12SourcePDFScholar