← Search

Kwanghoon Sohn

49 accepted papers

2026

LEARNING WHAT TO HEAR: BOOSTING SOUND-SOURCE ASSOCIATION FOR ROBUST AUDIOVISUAL INSTANCE SEGMENTATION

ICASSP 2026poster

Audiovisual instance segmentation (AVIS) requires accurately localizing and tracking sounding objects throughout video sequences. Existing methods suffer from visual bias stemming from two fundamental issues: uniform additive fusion prevents queries from specializing to different sound sources, whil…

Cited by 0SourcePDFScholar
2025

Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video Representations

CVPR 2025poster

View-invariant representation learning from egocentric (first-person, ego) and exocentric (third-person, exo) videos is a promising approach toward generalizing video understanding systems across multiple viewpoints. However, this area has been underexplored due to the substantial differences in per…

2025

Faster Parameter-Efficient Tuning with Token Redundancy Reduction

CVPR 2025poster

Parameter-efficient tuning (PET) aims to transfer pre-trained foundation models to downstream tasks by learning a small number of parameters. Compared to traditional fine-tuning, which updates the entire model, PET significantly reduces storage and transfer costs for each task regardless of exponent…

2024

A Simple Framework for Generalization in Visual RL under Dynamic Scene Perturbations

NeurIPS 2024poster

In the rapidly evolving domain of vision-based deep reinforcement learning (RL), a pivotal challenge is to achieve generalization capability to dynamic environmental changes reflected in visual observations. Our work delves into the intricacies of this problem, identifying two key issues that appear…

Cited by 0SourcePDFScholar
2024

Diffusion-driven GAN Inversion for Multi-Modal Face Image Generation

CVPR 2024poster

We present a new multi-modal face image generation method that converts a text prompt and a visual input such as a semantic mask or scribble map into a photo-realistic face image. To do this we combine the strengths of Generative Adversarial networks (GANs) and diffusion models (DMs) by employing th…

2024

Improving Visual Recognition with Hyperbolical Visual Hierarchy Mapping

CVPR 2024poster

Visual scenes are naturally organized in a hierarchy where a coarse semantic is recursively comprised of several fine details. Exploring such a visual hierarchy is crucial to recognize the complex relations of visual elements leading to a comprehensive scene understanding. In this paper we propose a…

2023

Hierarchical Visual Primitive Experts for Compositional Zero-Shot Learning

ICCV 2023poster

Compositional zero-shot learning (CZSL) aims to recognize unseen compositions with prior knowledge of known primitives (attribute and object). Previous works for CZSL often suffer from grasping the contextuality between attribute and object, as well as the discriminability of visual features, and th…

Cited by 22PDFcodeScholar
2023

Knowing Where to Focus: Event-aware Transformer for Video Grounding

ICCV 2023poster

Recent DETR-based video grounding models have made the model directly predict moment timestamps without any hand-crafted components, such as a pre-defined proposal or non-maximum suppression, by learning moment queries. However, their input-agnostic moment queries inevitably overlook an intrinsic t…

Cited by 64PDFcodeScholar
2023

Local-Guided Global: Paired Similarity Representation for Visual Reinforcement Learning

CVPR 2023poster

Recent vision-based reinforcement learning (RL) methods have found extracting high-level features from raw pixels with self-supervised learning to be effective in learning policies. However, these methods focus on learning global representations of images, and disregard local spatial structures pres…

Cited by 10SourcePDFScholar
2023

PartMix: Regularization Strategy To Learn Part Discovery for Visible-Infrared Person Re-Identification

CVPR 2023poster

Modern data augmentation using a mixture-based technique can regularize the models from overfitting to the training data in various computer vision applications, but a proper data augmentation technique tailored for the part-based Visible-Infrared person Re-IDentification (VI-ReID) models remains un…

Cited by 93SourcePDFScholar
2023

Probabilistic Prompt Learning for Dense Prediction

CVPR 2023poster

Recent progress in deterministic prompt learning has become a promising alternative to various downstream vision tasks, enabling models to learn powerful visual representations with the help of pre-trained vision-language models. However, this approach results in limited performance for dense predic…

Cited by 23SourcePDFScholar
2023

Unsupervised Deep Asymmetric Stereo Matching With Spatially-Adaptive Self-Similarity

CVPR 2023poster

Unsupervised stereo matching has received a lot of attention since it enables the learning of disparity estimation without ground-truth data. However, most of the unsupervised stereo matching algorithms assume that the left and right images have consistent visual properties, i.e., symmetric, and eas…

Cited by 9SourcePDFScholar
2022

Meta-confidence estimation for stereo matching

ICRA 2022poster

We propose a novel framework to estimate the confidence of a disparity map taking into account, for the first time, the uncertainty affecting the confidence estimation process itself. Conversely to other tasks such as disparity estimation, the uncertainty of confidence directly hints that the confid…

Cited by 2SourceScholar
2022

Multi-Domain Unsupervised Image-to-Image Translation with Appearance Adaptive Convolution

ICASSP 2022accepted

Over the past few years, image-to-image (I2I) translation methods have been proposed to translate a given image into diverse outputs. Despite the impressive results, they mainly focus on the I2I translation between two domains, so the multi-domain I2I translation still remains a challenge. To addres…

Cited by 0SourceScholar
2022

Pin the Memory: Learning To Generalize Semantic Segmentation

CVPR 2022poster

The rise of deep neural networks has led to several breakthroughs for semantic segmentation. In spite of this, a model trained on source domain often fails to work properly in new challenging domains, that is directly concerned with the generalization capability of the model. In this paper, we prese…

Cited by 68PDFcodeScholar
2022

PointFix: Learning to Fix Domain Bias for Robust Online Stereo Adaptation

ECCV 2022poster

"Online stereo adaptation tackles the domain shift problem, caused by different environments between synthetic (training) and real (test) datasets, to promptly adapt stereo models in dynamic real-world applications such as autonomous driving. However, previous methods often fail to counteract partic…

2021

Adaptive Confidence Thresholding for Monocular Depth Estimation

ICCV 2021poster

Self-supervised monocular depth estimation has become an appealing solution to the lack of ground truth labels, but its reconstruction loss often produces over-smoothed results across object boundaries and is incapable of handling occlusion explicitly. In this paper, we propose a new approach to lev…

Cited by 35PDFcodeScholar
2021

Bridge To Answer: Structure-Aware Graph Interaction Network for Video Question Answering

CVPR 2021poster

This paper presents a novel method, termed Bridge to Answer, to infer correct answers for questions about a given video by leveraging adequate graph interactions of heterogeneous crossmodal graphs. To realize this, we learn question conditioned visual graphs by exploiting the relation between video…

Cited by 118PDFScholar
2021

CATs: Cost Aggregation Transformers for Visual Correspondence

NeurIPS 2021poster

We propose a novel cost aggregation network, called Cost Aggregation Transformers (CATs), to find dense correspondences between semantically similar images with additional challenges posed by large intra-class appearance and geometric variations. Cost aggregation is a highly important process in mat…

2021

Cross-Domain Grouping and Alignment for Domain Adaptive Semantic Segmentation

AAAI 2021technical

Existing techniques to adapt semantic segmentation networks across source and target domains within deep convolutional neural networks (CNNs) deal with all the samples from the two domains in a global or category-aware manner. They do not consider an inter-class variation within the target domain it…

2021

Deep Low-Contrast Image Enhancement using Structure Tensor Representation

AAAI 2021technical

We present a new deep learning framework for low-contrast image enhancement, which trains the network using the multi-exposure sequences rather than explicit ground-truth images. The purpose of our method is to enhance a low-contrast image so as to contain abundant details in various exposure levels…

Cited by 2SourcePDFScholar
2021

Learning Canonical 3D Object Representation for Fine-Grained Recognition

ICCV 2021poster

We propose a novel framework for fine-grained object recognition that learns to recover object variation in 3D space from a single image, trained on an image collection without using any ground-truth 3D annotation. We accomplish this by representing an object as a composition of 3D shape and its app…

Cited by 16PDFScholar
2021

Looking Into Your Speech: Learning Cross-Modal Affinity for Audio-Visual Speech Separation

CVPR 2021poster

In this paper, we address the problem of separating individual speech signals from videos using audio-visual neural processing. Most conventional approaches utilize frame-wise matching criteria to extract shared information between co-occurring audio and video. Thus, their performance heavily depend…

Cited by 56PDFScholar
2021

Mining Better Samples for Contrastive Learning of Temporal Correspondence

CVPR 2021poster

We present a novel framework for contrastive learning of pixel-level representation using only unlabeled video. Without the need of ground-truth annotation, our method is capable of collecting well-defined positive correspondences by measuring their confidences and well-defined negative ones by appr…

Cited by 34PDFScholar
2021

Stereo-augmented Depth Completion from a Single RGB-LiDAR image

ICRA 2021poster

Depth completion is an important task in computer vision and robotics applications, which aims at predicting accurate dense depth from a single RGB-LiDAR image. Convolutional neural networks (CNNs) have been widely used for depth completion to learn a mapping function from sparse to dense depth. How…

Cited by 8SourceScholar
2020

Cylindrical Convolutional Networks for Joint Object Detection and Viewpoint Estimation

CVPR 2020poster

Existing techniques to encode spatial invariance within deep convolutional neural networks only model 2D transformation fields. This does not account for the fact that objects in a 2D space are a projection of 3D ones, and thus they have limited ability to severe object viewpoint changes. To overcom…

Cited by 19PDFScholar
2019

Joint Learning of Semantic Alignment and Object Landmark Detection

ICCV 2019poster

Convolutional neural networks (CNNs) based approaches for semantic alignment and object landmark detection have improved their performance significantly. Current efforts for the two tasks focus on addressing the lack of massive training data through weakly- or unsupervised learning frameworks. In th…

Cited by 21PDFScholar
2019

LAF-Net: Locally Adaptive Fusion Networks for Stereo Confidence Estimation

CVPR 2019oral

We present a novel method that estimates confidence map of an initial disparity by making full use of tri-modal input, including matching cost, disparity, and color image through deep networks. The proposed network, termed as Locally Adaptive Fusion Networks (LAF-Net), learns locally-varying attent…

Cited by 65PDFScholar
2018

PARN: Pyramidal Affine Regression Networks for Dense Semantic Correspondence

ECCV 2018poster

This paper presents a deep architecture for dense semantic correspondence, called pyramidal affine regression networks (PARN), that estimates locally-varying affine transformation fields across images. To deal with intra-class appearance and shape variations that commonly exist among different insta…

Cited by 69SourcePDFScholar
2018

Recurrent Transformer Networks for Semantic Correspondence

NeurIPS 2018spotlight

We present recurrent transformer networks (RTNs) for obtaining dense correspondences between semantically similar images. Our networks accomplish this through an iterative process of estimating spatial transformations between the input images and using these transformations to generate aligned convo…

Cited by 115SourcePDFScholar
2018

Spatiotemporal Attention Based Deep Neural Networks for Emotion Recognition

ICASSP 2018accepted

We propose a spatiotemporal attention based deep neural networks for dimensional emotion recognition in facial videos. To learn the spatiotemporal attention that selectively focuses on emotional sailient parts within facial videos, we formulate the spatiotemporal encoder-decoder network using Convol…

Cited by 0SourceScholar
2017

Deeply Aggregated Alternating Minimization for Image Restoration

CVPR 2017spotlight

Regularization-based image restoration has remained an active research topic in image processing and computer vision. It often leverages a guidance signal captured in different fields as an additional cue. In this work, we present a general framework for image restoration, called deeply aggregated a…

Cited by 37PDFScholar
2017

FCSS: Fully Convolutional Self-Similarity for Dense Semantic Correspondence

CVPR 2017poster

We present a descriptor, called fully convolutional self-similarity (FCSS), for dense semantic correspondence. To robustly match points among different instances within the same object class, we formulate FCSS using local self-similarity (LSS) within a fully convolutional network. In contrast to exi…

Cited by 176PDFScholar
2015

DASC: Dense Adaptive Self-Correlation Descriptor for Multi-Modal and Multi-Spectral Correspondence

CVPR 2015poster

Establishing dense visual correspondence between multiple images is a fundamental task in many applications of computer vision and computational photography. Classical approaches, which aim to estimate dense stereo and optical flow fields for images adjacent in viewpoint or in time, have been dramat…

Cited by 119SourcePDFScholar