← Search

Naoto Yokoya

23 accepted papers

2026

GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote Sensing

CVPR 2026

Recent advances in multimodal large language models (MLLMs) have accelerated progress in domain-oriented AI, yet their development in geoscience and remote sensing (RS) remains constrained by distinctive challenges: wide-ranging disciplinary knowledge, heterogeneous sensor modalities, and a fragment

Cited by 0SourceScholar
2026

Is Pre-Training Applicable to the Decoder for Dense Prediction?

ICRA 2026poster

Encoder-decoder networks are commonly used model architectures for dense prediction tasks, where the encoder typically employs a model pre-trained on upstream tasks, while the decoder is often either randomly initialized or pre-trained on other tasks. In this paper, we introduce ×Net, a novel framew…

2026

LandCraft: Designing the Structured 3D Landscapes via Text Guidance

AAAI 2026technical

Modeling large-scale landscapes is a foundational yet time-consuming task in many 3D applications, typically requiring substantial expertise. Recently, Text-to-3D techniques have emerged as a promising, beginner-friendly prototyping approach for generating 3D content from textual input. However, ex

Cited by 0SourcePDFScholar
2026

MM-OVSeg: Multimodal Optical-SAR Fusion for Open-Vocabulary Segmentation in Remote Sensing

CVPR 2026

Open-vocabulary segmentation enables pixel-level recognition from an open set of textual categories, allowing generalization beyond fixed classes. Despite great potential in remote sensing, progress in this area remains largely limited to clear-sky optical data and struggles under cloudy or haze-con

Cited by 0SourcecodeScholar
2026

Proteo-R1: Thinking Foundation Models for De Novo Protein Binder Design

ICML 2026poster

Recent advances in generative diffusion and flow-matching models have revolutionized molecular design, enabling the creation of novel proteins, small molecules, and RNA sequences with unprecedented fidelity. Yet, these models remain intuitive rather than intelligent—they generate without reasoning. …

Cited by 0SourceScholar
2025

DisasterM3: A Remote Sensing Vision-Language Dataset for Disaster Damage Assessment and Response

NeurIPS 2025poster

Large vision-language models (VLMs) have made great achievements in Earth vision. However, complex disaster scenes with diverse disaster types, geographic regions, and satellite sensors have posed new challenges for VLM applications. To fill this gap, we curate the first remote sensing vision-langua…

Cited by 0SourcecodeScholar
2025

DynamicVL: Benchmarking Multimodal Large Language Models for Dynamic City Understanding

NeurIPS 2025poster

Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in visual understanding, but their application to long-term Earth observation analysis remains limited, primarily focusing on single-temporal or bi-temporal imagery. To address this gap, we introduce **DVL-Suite**, a…

Cited by 0SourcecodeScholar
2025

GaussianOcc: Fully Self-supervised and Efficient 3D Occupancy Estimation with Gaussian Splatting

ICCV 2025poster

We introduce GaussianOcc, a systematic method that investigates Gaussian splatting for fully self-supervised and efficient 3D occupancy estimation in surround views. First, traditional methods for self-supervised 3D occupancy estimation still require ground truth 6D poses from sensors during trainin…

2025

LR2Depth: Large-Region Aggregation at Low Resolution for Efficient Monocular Depth Estimation

IROS 2025

Monocular depth estimation (MDE) is crucial for various computer vision applications, but existing methods often struggle to balance inference speed and accuracy when processing large-region visual information. This paper introduces LR<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="

Cited by 0SourceScholar
2025

MP-HSIR: A Multi-Prompt Framework for Universal Hyperspectral Image Restoration

ICCV 2025poster

Hyperspectral images (HSIs) often suffer from diverse and unknown degradations during imaging, leading to severe spectral and spatial distortions. Existing HSI restoration methods typically rely on specific degradation assumptions, limiting their effectiveness in complex scenarios. In this paper, we…

2025

Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models

EMNLP 2025

Uncertainty quantification is essential for assessing the reliability and trustworthiness of modern AI systems. Among existing approaches, verbalized uncertainty, where models express their confidence through natural language, has emerged as a lightweight and interpretable solution in large language

Cited by 0SourcePDFScholar
2024

SynRS3D: A Synthetic Dataset for Global 3D Semantic Understanding from Monocular Remote Sensing Imagery

NeurIPS 2024spotlight

Global semantic 3D understanding from single-view high-resolution remote sensing (RS) imagery is crucial for Earth observation (EO). However, this task faces significant challenges due to the high costs of annotations and data collection, as well as geographically restricted data availability. To ad…

2022

ES6D: A Computation Efficient and Symmetry-Aware 6D Pose Regression Framework

CVPR 2022poster

In this paper, a computation efficient regression framework is presented for estimating the 6D pose of rigid objects from a single RGB-D image, which is applicable to handling symmetric objects. This framework is designed in a simple architecture that efficiently extracts point-wise features from RG…

Cited by 34PDFcodeScholar
2022

Learning Mutual Modulation for Self-Supervised Cross-Modal Super-Resolution

ECCV 2022poster

"Self-supervised cross-modal super-resolution (SR) can overcome the difficulty of acquiring paired training data, but is challenging because only low-resolution (LR) source and high-resolution (HR) guide images from different modalities are available. Existing methods utilize pseudo or weak supervis…

2022

Spectrum-Aware and Transferable Architecture Search for Hyperspectral Image Restoration

ECCV 2022poster

"Convolutional neural networks have been widely developed for hyperspectral image (HSI) restoration. However, making full use of the spatial-spectral information of HSIs still remains a challenge. In this work, we disentangle the 3D convolution into lightweight 2D spatial and spectral convolutions,…

Cited by 13SourcePDFScholar
2019

Non-Local Meets Global: An Integrated Paradigm for Hyperspectral Denoising

CVPR 2019oral

Non-local low-rank tensor approximation has been developed as a state-of-the-art method for hyperspectral image (HSI) denoising. Unfortunately, while their denoising performance benefits little from more spectral bands, the running time of these methods significantly increases. In this paper, we cla…

Cited by 189PDFcodeScholar
2019

Total-variation-regularized Tensor Ring Completion for Remote Sensing Image Reconstruction

ICASSP 2019accepted

In recent studies, tensor ring (TR) decomposition has shown to be effective in data compression and representation. However, the existing TR-based completion methods only exploit the global low-rank property of the visual data. When applying them to remote sensing (RS) image processing, the spatial…

Cited by 0SourceScholar
2018

Joint & Progressive Learning from High-Dimensional Data for Multi-Label Classification

ECCV 2018poster

Despite the fact that nonlinear subspace learning techniques (e.g. manifold learning) have successfully applied to data representation, there is still room for improvement in explainability (explicit mapping), generalization (out-of-samples), and cost-effectiveness (linearization). To this end, a no…

Cited by 39SourcePDFScholar
2017

A novel ensemble classifier of hyperspectral and LiDAR data using morphological features

ICASSP 2017accepted

Due to the benefits and limitation of different remote sensing sensors, fusion of the features from multiple sensors, such as hyperspectral and light detection and ranging (LiDAR) is an effective method for land cover mapping. In this paper, we propose a novel ensemble classifier to fuse hyperspectr…

Cited by 0SourceScholar