← Search

Sangmin Lee

35 accepted papers

2026

FedAMB: Adaptive Modality Balancing for Dominance-Robust Multimodal Federated Distillation

IJCAI 2026

Federated knowledge distillation (Fed-KD) exchanges distilled predictions on a shared proxy dataset, often reducing communication and accommodating heterogeneous client architectures. However, in multimodal federated learning (MFL), modality dominance can bias local optimization and contaminate the

Cited by 0Scholar
2026

MuCo: Multi-turn Contrastive Learning for Multimodal Embedding Model

CVPR 2026

Universal Multimodal embedding models built on Multimodal Large Language Models (MLLMs) have traditionally employed contrastive learning, which aligns representations of query-target pairs across different modalities. Yet, despite its empirical success, they are primarily built on a "single-turn" fo

Cited by 0SourcecodeScholar
2025

Derivative-Free Diffusion Manifold-Constrained Gradient for Unified XAI

CVPR 2025poster

Gradient-based methods are a prototypical family of "explainability for AI" (XAI) techniques, especially for image-based models. However, they (1) require white-box access to models, (2) are vulnerable to adversarial attacks, and (3) produce attributions that lie off the image manifold, leading to e…

2025

DynScene: Scalable Generation of Dynamic Robotic Manipulation Scenes for Embodied AI

CVPR 2025poster

Robotic manipulation in embodied AI critically depends on large-scale, high-quality datasets that reflect realistic object interactions and physical dynamics. However, existing data collection pipelines are often slow, expensive, and heavily reliant on manual efforts. We present DynScene, a diffusio…

Cited by 0SourcePDFScholar
2025

Gaze-LLE: Gaze Target Estimation via Large-Scale Learned Encoders

CVPR 2025highlight

We address the problem of gaze target estimation, which aims to predict where a person is looking in a scene. Predicting a person's gaze target requires reasoning both about the person's appearance and the contents of the scene. Prior works have developed increasingly complex, hand-crafted pipelines…

2025

IM-LUT: Interpolation Mixing Look-Up Tables for Image Super-Resolution

ICCV 2025poster

Super-resolution (SR) has been a pivotal task in image processing, aimed at enhancing image resolution across various applications. Recently, look-up table (LUT)-based approaches have attracted interest due to their efficiency and performance. However, these methods are typically designed for fixed…

2025

LAMA-UT: Language Agnostic Multilingual ASR Through Orthography Unification and Language-Specific Transliteration

AAAI 2025technical

Building a universal multilingual automatic speech recognition (ASR) model that performs equitably across languages has long been a challenge due to its inherent difficulties. To address this task we introduce a Language-Agnostic Multilingual ASR pipeline through orthography Unification and language…

2025

MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided Stylization

ICCV 2025poster

Speech-driven 3D facial animation aims to synthesize realistic facial motion sequences from given audio, matching the speaker's speaking style. However, previous works often require priors such as class labels of a speaker or additional 3D facial meshes at inference, which makes them fail to reflect…

Cited by 0SourcePDFScholar
2025

Object-aware Sound Source Localization via Audio-Visual Scene Understanding

CVPR 2025poster

Audio-visual sound source localization task aims to spatially localize sound-making objects within visual scenes by integrating visual and audio cues. However, existing methods struggle with accurately localizing sound-making objects in complex scenes, particularly when visually similar silent objec…

2025

Question-Aware Gaussian Experts for Audio-Visual Question Answering

CVPR 2025highlight

Audio-Visual Question Answering (AVQA) requires not only question-based multimodal reasoning but also precise temporal grounding to capture subtle dynamics for accurate prediction. However, existing methods mainly use question information implicitly, limiting focus on question-specific details. Furt…

2025

SocialGesture: Delving into Multi-person Gesture Understanding

CVPR 2025poster

Previous research in human gesture recognition has largely overlooked multi-person interactions, which are crucial for understanding the social context of naturally occurring gestures. This limitation in existing datasets presents a significant challenge in aligning human gestures with other modalit…

Cited by 0SourcePDFScholar
2025

Toward Human Deictic Gesture Target Estimation

NeurIPS 2025poster

Humans have a remarkable ability to use co-speech deictic gestures, such as pointing and showing, to enrich verbal communication and support social interaction. These gestures are so fundamental that infants begin to use them even before they acquire spoken language, which highlights their central r…

Cited by 0SourcecodeScholar
2025

Unleashing In-context Learning of Autoregressive Models for Few-shot Image Manipulation

CVPR 2025highlight

Text-guided image manipulation has experienced notable advancement in recent years. In order to mitigate linguistic ambiguity, few-shot learning with visual examples has been applied for instructions that are underrepresented in the training set, or difficult to describe purely in language. However,…

Cited by 3SourcePDFScholar
2025

Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection

AAAI 2025technical

The goal of video moment retrieval and highlight detection is to identify specific segments and highlights based on a given text query. With the rapid growth of video content and the overlap between these tasks, recent works have addressed both simultaneously. However, they still struggle to fully c…

2024

Defining Neural Network Architecture through Polytope Structures of Datasets

ICML 2024spotlight

Current theoretical and empirical research in neural networks suggests that complex datasets require large network architectures for thorough classification, yet the precise nature of this relationship remains unclear. This paper tackles this issue by defining upper and lower bounds for neural netwo…

Cited by 1SourcePDFScholar
2024

Learning to Visually Localize Sound Sources from Mixtures without Prior Source Knowledge

CVPR 2024poster

The goal of the multi-sound source localization task is to localize sound sources from the mixture individually. While recent multi-sound source localization methods have shown improved performance they face challenges due to their reliance on prior information about the number of objects to be sepa…

2024

Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations

CVPR 2024poster

Understanding social interactions involving both verbal and non-verbal cues is essential for effectively interpreting social situations. However most prior works on multimodal social cues focus predominantly on single-person behaviors or rely on holistic visual representations that are not aligned t…

2024

Self-supervised Debiasing Using Low Rank Regularization

CVPR 2024poster

Spurious correlations can cause strong biases in deep neural networks impairing generalization ability. While most existing debiasing methods require full supervision on either spurious attributes or target labels training a debiased model from a limited amount of both annotations is still an open q…

Cited by 4SourcePDFScholar
2023

Training Debiased Subnetworks With Contrastive Weight Pruning

CVPR 2023poster

Neural networks are often biased to spuriously correlated features that provide misleading statistical evidence that does not generalize. This raises an interesting question: "Does an optimal unbiased functional subnetwork exist in a severely biased network? If so, how to extract such subnetwork?" W…

2022

Audio-Visual Mismatch-Aware Video Retrieval via Association and Adjustment

ECCV 2022poster

"Retrieving desired videos using natural language queries has attracted increasing attention in research and industry fields as a huge number of videos appear on the internet. Some existing methods attempted to address this video retrieval problem by exploiting multi-modal information, especially au…

Cited by 8SourcePDFScholar
2022

OpenStreetMap-Based LiDAR Global Localization in Urban Environment Without a Prior LiDAR Map

RA-L 2022

Using publicly accessible maps, we propose a novel vehicle localization method that can be applied without using prior light detection and ranging (LiDAR) maps. Our method generates OSM descriptors by calculating the distances to buildings from a location in OpenStreetMap at a regular angle, and LiD

Cited by 58SourceScholar
2022

Weakly Paired Associative Learning for Sound and Image Representations via Bimodal Associative Memory

CVPR 2022poster

Data representation learning without labels has attracted increasing attention due to its nature that does not require human annotation. Recently, representation learning has been extended to bimodal data, especially sound and image which are closely related to basic human senses. Existing sound and…

Cited by 7PDFScholar
2021

Explaining Convolutional Neural Networks through Attribution-Based Input Sampling and Block-Wise Feature Aggregation

AAAI 2021technical

As an emerging field in Machine Learning, Explainable AI (XAI) has been offering remarkable performance in interpreting the decisions made by Convolutional Neural Networks (CNNs). To achieve visual explanations for CNNs, methods based on class activation mapping and randomized input sampling have ga…

Cited by 48SourcePDFScholar
2021

Towards a Better Understanding of VR Sickness: Physical Symptom Prediction for VR Contents

AAAI 2021technical

We address the black-box issue of VR sickness assessment (VRSA) by evaluating the level of physical symptoms of VR sickness. For the VR contents inducing the similar VR sickness level, the physical symptoms can vary depending on the characteristics of the contents. Most of existing VRSA methods focu…

Cited by 12SourcePDFScholar
2021

Two-Stage Textual Knowledge Distillation for End-to-End Spoken Language Understanding

ICASSP 2021accepted

End-to-end approaches open a new way for more accurate and efficient spoken language understanding (SLU) systems by alleviating the drawbacks of traditional pipeline systems. Previous works exploit textual information for an SLU model via pre-training with automatic speech recognition or finetuning…

Cited by 0SourceScholar
2021

Video Prediction Recalling Long-Term Motion Context via Memory Alignment Learning

CVPR 2021poster

Our work addresses long-term motion context issues for predicting future frames. To predict the future precisely, it is required to capture which long-term motion context (e.g., walking or running) the input motion (e.g., leg movement) belongs to. The bottlenecks arising when dealing with the long-t…

Cited by 146PDFcodeScholar
2021

Visual Comfort Aware-Reinforcement Learning for Depth Adjustment of Stereoscopic 3D Images

AAAI 2021technical

Depth adjustment aims to enhance the visual experience of stereoscopic 3D (S3D) images, which accompanied with improving visual comfort and depth perception. For a human expert, the depth adjustment procedure is a sequence of iterative decision making. The human expert iteratively adjusted the depth…

Cited by 9SourcePDFScholar
2020

SACA Net: Cybersickness Assessment of Individual Viewers for VR Content via Graph-based Symptom Relation Embedding

ECCV 2020poster

Recently, cybersickness assessment for VR content is required to deal with viewing safety issues. Assessing physical symptoms of individual viewers is challenging but important to provide detailed and personalized guides for viewing safety. In this paper, we propose a novel symptom-aware cybersickne…

Cited by 8SourcePDFScholar
2020

Structure Boundary Preserving Segmentation for Medical Image With Ambiguous Boundary

CVPR 2020poster

In this paper, we propose a novel image segmentation method to tackle two critical problems of medical image, which are (i) ambiguity of structure boundary in the medical image domain and (ii) uncertainty of the segmented region without specialized domain knowledge. To solve those two problems in au…

Cited by 169PDFScholar