← Search

Dongbo Min

26 accepted papers

2026

FluoCLIP: Stain-Aware Focus Quality Assessment in Fluorescence Microscopy

CVPR 2026

Accurate focus quality assessment (FQA) in fluorescence microscopy is challenging due to stain-dependent optical variations that induce heterogeneous focus behavior across images. Existing methods, however, treat focus quality as a stain-agnostic problem, assuming a shared global ordering. We formul

Cited by 0SourcecodeScholar
2025

Hybrid-TTA: Continual Test-time Adaptation via Dynamic Domain Shift Detection

ICCV 2025poster

Continual Test Time Adaptation (CTTA) has emerged as a critical approach to bridge the domain gap between controlled training environments and real-world scenarios.Since it is important to balance the trade-off between adaptation and stabilization, many studies have tried to accomplish it by either…

Cited by 0SourcePDFScholar
2025

RobIA: Robust Instance-aware Continual Test-time Adaptation for Deep Stereo

NeurIPS 2025poster

Stereo Depth Estimation in real-world environments poses significant challenges due to dynamic domain shifts, sparse or unreliable supervision, and the high cost of acquiring dense ground-truth labels. While recent Test-Time Adaptation (TTA) methods offer promising solutions, most rely on static tar…

Cited by 0SourceScholar
2025

TADFormer: Task-Adaptive Dynamic TransFormer for Efficient Multi-Task Learning

CVPR 2025poster

Transfer learning paradigm has driven substantial advancements in various vision tasks. However, as state-of-the-art models continue to grow, classical full fine-tuning often becomes computationally impractical, particularly in multi-task learning (MTL) setup where training complexity increases prop…

Cited by 0SourcePDFScholar
2024

A Simple Framework for Generalization in Visual RL under Dynamic Scene Perturbations

NeurIPS 2024poster

In the rapidly evolving domain of vision-based deep reinforcement learning (RL), a pivotal challenge is to achieve generalization capability to dynamic environmental changes reflected in visual observations. Our work delves into the intricacies of this problem, identifying two key issues that appear…

Cited by 0SourcePDFScholar
2023

Environment Agnostic Representation for Visual Reinforcement Learning

ICCV 2023poster

Generalization capability of vision-based deep reinforcement learning (RL) is indispensable to deal with dynamic environment changes that exist in visual observations. The high-dimensional space of the visual input, however, imposes challenges in adapting an agent to unseen environments. In this wor…

Cited by 10PDFcodeScholar
2023

Local-Guided Global: Paired Similarity Representation for Visual Reinforcement Learning

CVPR 2023poster

Recent vision-based reinforcement learning (RL) methods have found extracting high-level features from raw pixels with self-supervised learning to be effective in learning policies. However, these methods focus on learning global representations of images, and disregard local spatial structures pres…

Cited by 10SourcePDFScholar
2022

Neural Matching Fields: Implicit Representation of Matching Fields for Visual Correspondence

NeurIPS 2022accept

Existing pipelines of semantic correspondence commonly include extracting high-level semantic features for the invariance against intra-class variations and background clutters. This architecture, however, inevitably results in a low-resolution matching field that additionally requires an ad-hoc int…

2022

Pin the Memory: Learning To Generalize Semantic Segmentation

CVPR 2022poster

The rise of deep neural networks has led to several breakthroughs for semantic segmentation. In spite of this, a model trained on source domain often fails to work properly in new challenging domains, that is directly concerned with the generalization capability of the model. In this paper, we prese…

Cited by 68PDFcodeScholar
2022

PointFix: Learning to Fix Domain Bias for Robust Online Stereo Adaptation

ECCV 2022poster

"Online stereo adaptation tackles the domain shift problem, caused by different environments between synthetic (training) and real (test) datasets, to promptly adapt stereo models in dynamic real-world applications such as autonomous driving. However, previous methods often fail to counteract partic…

2021

Adaptive Confidence Thresholding for Monocular Depth Estimation

ICCV 2021poster

Self-supervised monocular depth estimation has become an appealing solution to the lack of ground truth labels, but its reconstruction loss often produces over-smoothed results across object boundaries and is incapable of handling occlusion explicitly. In this paper, we propose a new approach to lev…

Cited by 35PDFcodeScholar
2021

Mining Better Samples for Contrastive Learning of Temporal Correspondence

CVPR 2021poster

We present a novel framework for contrastive learning of pixel-level representation using only unlabeled video. Without the need of ground-truth annotation, our method is capable of collecting well-defined positive correspondences by measuring their confidences and well-defined negative ones by appr…

Cited by 34PDFScholar
2019

Joint Learning of Semantic Alignment and Object Landmark Detection

ICCV 2019poster

Convolutional neural networks (CNNs) based approaches for semantic alignment and object landmark detection have improved their performance significantly. Current efforts for the two tasks focus on addressing the lack of massive training data through weakly- or unsupervised learning frameworks. In th…

Cited by 21PDFScholar
2019

LAF-Net: Locally Adaptive Fusion Networks for Stereo Confidence Estimation

CVPR 2019oral

We present a novel method that estimates confidence map of an initial disparity by making full use of tri-modal input, including matching cost, disparity, and color image through deep networks. The proposed network, termed as Locally Adaptive Fusion Networks (LAF-Net), learns locally-varying attent…

Cited by 65PDFScholar
2018

PARN: Pyramidal Affine Regression Networks for Dense Semantic Correspondence

ECCV 2018poster

This paper presents a deep architecture for dense semantic correspondence, called pyramidal affine regression networks (PARN), that estimates locally-varying affine transformation fields across images. To deal with intra-class appearance and shape variations that commonly exist among different insta…

Cited by 69SourcePDFScholar
2018

Recurrent Transformer Networks for Semantic Correspondence

NeurIPS 2018spotlight

We present recurrent transformer networks (RTNs) for obtaining dense correspondences between semantically similar images. Our networks accomplish this through an iterative process of estimating spatial transformations between the input images and using these transformations to generate aligned convo…

Cited by 115SourcePDFScholar
2017

Deeply Aggregated Alternating Minimization for Image Restoration

CVPR 2017spotlight

Regularization-based image restoration has remained an active research topic in image processing and computer vision. It often leverages a guidance signal captured in different fields as an additional cue. In this work, we present a general framework for image restoration, called deeply aggregated a…

Cited by 37PDFScholar
2017

FCSS: Fully Convolutional Self-Similarity for Dense Semantic Correspondence

CVPR 2017poster

We present a descriptor, called fully convolutional self-similarity (FCSS), for dense semantic correspondence. To robustly match points among different instances within the same object class, we formulate FCSS using local self-similarity (LSS) within a fully convolutional network. In contrast to exi…

Cited by 176PDFScholar
2015

DASC: Dense Adaptive Self-Correlation Descriptor for Multi-Modal and Multi-Spectral Correspondence

CVPR 2015poster

Establishing dense visual correspondence between multiple images is a fundamental task in many applications of computer vision and computational photography. Classical approaches, which aim to estimate dense stereo and optical flow fields for images adjacent in viewpoint or in time, have been dramat…

Cited by 119SourcePDFScholar
2015

SPM-BP: Sped-up PatchMatch Belief Propagation for Continuous MRFs

ICCV 2015oral

Markov random fields are widely used to model many computer vision problems that can be cast in an energy minimization framework composed of unary and pairwise potentials. While computationally tractable discrete optimizers such as Graph Cuts and belief propagation (BP) exist for multi-label discret…

Cited by 111PDFScholar