← Search

Masatoshi Okutomi

19 accepted papers

2025

RAM-W600: A Multi-Task Wrist Dataset and Benchmark for Rheumatoid Arthritis

NeurIPS 2025poster

Rheumatoid arthritis (RA) is a common autoimmune disease that has been the focus of research in computer-aided diagnosis (CAD) and disease monitoring. In clinical settings, conventional radiography (CR) is widely used for the screening and evaluation of RA due to its low cost and accessibility. The…

Cited by 0SourcecodeScholar
2024

CFDNet: A Generalizable Foggy Stereo Matching Network with Contrastive Feature Distillation

ICRA 2024poster

Stereo matching under foggy scenes remains a challenging task since the scattering effect degrades the visibility and results in less distinctive features for dense correspondence matching. While some previous learning-based methods integrated a physical scattering function for simultaneous stereo-m…

Cited by 1SourceScholar
2024

Reflection Removal Using Recurrent Polarization-to-Polarization Network

ICASSP 2024accepted

This paper addresses reflection removal, which is the task of separating reflection components from a captured image and deriving the image with only transmission components. Considering that the existence of the reflection changes the polarization state of a scene, some existing methods have exploi…

Cited by 0SourceScholar
2024

Self-Supervised Spatially Variant PSF Estimation for Aberration-Aware Depth-from-Defocus

ICASSP 2024accepted

In this paper, we address the task of aberration-aware depth-from- defocus (DfD), which takes account of spatially variant point spread functions (PSFs) of a real camera. To effectively obtain the spatially variant PSFs of a real camera without requiring any ground-truth PSFs, we propose a novel sel…

Cited by 0SourceScholar
2024

VSRD: Instance-Aware Volumetric Silhouette Rendering for Weakly Supervised 3D Object Detection

CVPR 2024poster

Monocular 3D object detection poses a significant challenge in 3D scene understanding due to its inherently ill-posed nature in monocular depth estimation. Existing methods heavily rely on supervised learning using abundant 3D labels typically obtained through expensive and labor-intensive annotatio…

2022

Self-Supervised Ego-Motion Estimation Based on Multi-Layer Fusion of RGB and Inferred Depth

ICRA 2022poster

In existing self-supervised depth and ego-motion estimation methods, ego-motion estimation is usually limited to only leveraging RGB information. Recently, several methods have been proposed to further improve the accuracy of self-supervised ego-motion estimation by fusing information from other mod…

Cited by 17SourcecodeScholar
2019

Automatic Labeled LiDAR Data Generation based on Precise Human Model

ICRA 2019poster

Following improvements in deep neural networks, state-of-the-art networks have been proposed for human recognition using point clouds captured by LiDAR. However, the performance of these networks strongly depends on the training data. An issue with collecting training data is labeling. Labeling by h…

Cited by 9SourceScholar
2019

Is This the Right Place? Geometric-Semantic Pose Verification for Indoor Visual Localization

ICCV 2019poster

Visual localization in large and complex indoor scenes, dominated by weakly textured rooms and repeating geometric patterns, is a challenging problem with high practical relevance for applications such as Augmented Reality and robotics. To handle the ambiguities arising in this scenario, a common st…

Cited by 59PDFScholar
2019

Pro-Cam SSfM: Projector-Camera System for Structure and Spectral Reflectance From Motion

ICCV 2019poster

In this paper, we propose a novel projector-camera system for practical and low-cost acquisition of a dense object 3D model with the spectral reflectance property. In our system, we use a standard RGB camera and leverage an off-the-shelf projector as active illumination for both the 3D reconstructio…

Cited by 21PDFScholar
2018

Benchmarking 6DOF Outdoor Visual Localization in Changing Conditions

CVPR 2018poster

Visual localization enables autonomous vehicles to navigate in their surroundings and augmented reality applications to link virtual to real worlds. Practical visual localization approaches need to be robust to a wide variety of viewing condition, including day-night changes, as well as weather and…

Cited by 780SourcePDFScholar
2018

InLoc: Indoor Visual Localization With Dense Matching and View Synthesis

CVPR 2018poster

We seek to predict the 6 degree-of-freedom (6DoF) pose of a query photograph with respect to a large indoor 3D map. The contributions of this work are three-fold. First, we develop a new large-scale visual localization method targeted for indoor environments. The method proceeds along three steps: (…

Cited by 582SourcePDFScholar
2018

Joint optimization for compressive video sensing and reconstruction under hardware constraints

ECCV 2018poster

Compressive video sensing is the process of encoding multiple sub-frames into a single frame with controlled sensor exposures and reconstructing the sub-frames from the single compressed frame. It is known that spatially and temporally random exposures provide the most balanced compression in terms…

Cited by 41SourcePDFScholar
2017

Are Large-Scale 3D Models Really Necessary for Accurate Visual Localization?

CVPR 2017poster

Accurate visual localization is a key technology for autonomous navigation. 3D structure-based methods employ 3D models of the scene to estimate the full 6DOF pose of a camera very accurately. However, constructing (and extending) large-scale 3D models is still a significant challenge. In contrast,…

Cited by 257PDFScholar
2016

Gradient-Domain Image Reconstruction Framework With Intensity-Range and Base-Structure Constraints

CVPR 2016poster

This paper presents a novel unified gradient-domain image reconstruction framework with intensity-range constraint and base-structure constraint. The existing method for manipulating base structures and detailed textures are classifiable into two major approaches: i) gradient-domain and ii) layer-de…

Cited by 67PDFScholar
2015

24/7 Place Recognition by View Synthesis

CVPR 2015poster

We address the problem of large-scale visual place recognition for situations where the scene undergoes a major change in appearance, for example, due to illumination (day/night), change of seasons, aging, or structural modifications over time such as buildings built or destroyed. Such situations re…

Cited by 742SourcePDFScholar