← Search

Noriaki Hirose

21 accepted papers

2026

Learning to Drive Anywhere With Model-Based Reannotation

RA-L 2026

Developing broadly generalizable visual navigation policies for robots is a significant challenge, primarily constrained by the availability of large-scale, diverse training data. While curated datasets collected by researchers offer high quality, their limited size restricts policy generalization.

Cited by 12SourcecodeScholar
2026

Learning to Drive Anywhere with Model-Based Reannotation

ICRA 2026poster

Developing broadly generalizable visual navigation policies for robots is a significant challenge, primarily constrained by the availability of large-scale, diverse training data. While curated datasets collected by researchers offer high quality, their limited size restricts policy generalization. …

2026

OmniVLA: An Omni-Modal Vision-Language-Action Model for Robot Navigation

ICRA 2026poster

Humans can flexibly interpret and compose different goal specifications, such as language instructions, spatial coordinates, or visual references, when navigating to a destination. In contrast, most existing robotic navigation policies are trained on a single modality, limiting their adaptability to…

2024

LeLaN: Learning A Language-Conditioned Navigation Policy from In-the-Wild Video

CoRL 2024poster

We present our method, LeLaN, which uses action-free egocentric data to learn robust language-conditioned object navigation. By leveraging the knowledge of large vision and language models and grounding this knowledge using pre-trained segmentation and depth estimation models, we can label in-the-wi…

Cited by 7SourceScholar
2024

SELFI: Autonomous Self-Improvement with RL for Vision-Based Navigation around People

CoRL 2024poster

Autonomous self-improving robots that interact and improve with experience are key to the real-world deployment of robotic systems. In this paper, we propose an online learning method, SELFI, that leverages online robot experience to rapidly fine-tune pre-trained control policies efficiently. SELFI…

Cited by 2SourceScholar
2023

ExAug: Robot-Conditioned Navigation Policies via Geometric Experience Augmentation

ICRA 2023poster

Machine learning techniques rely on large and diverse datasets for generalization. Computer vision, natural language processing, and other applications can often reuse public datasets to train many different models. However, due to differences in physical configurations, it is challenging to leverag…

Cited by 23SourceScholar
2023

GNM: A General Navigation Model to Drive Any Robot

ICRA 2023poster

Learning provides a powerful tool for vision-based navigation, but the capabilities of learning-based policies are constrained by limited training data. If we could combine data from all available sources, including multiple kinds of robots, we could train more powerful navigation models. In this pa…

Cited by 120SourcecodeScholar
2023

ViNT: A Foundation Model for Visual Navigation

CoRL 2023oral

General-purpose pre-trained models (``foundation models'') have enabled practitioners to produce generalizable solutions for individual machine learning problems with datasets that are significantly smaller than those required for learning from scratch. Such models are typically trained on large and…

Cited by 154SourceScholar
2022

Depth360: Self-supervised Learning for Monocular Depth Estimation using Learnable Camera Distortion Model

IROS 2022poster

Self-supervised monocular depth estimation has been widely investigated to estimate depth images and relative poses from RGB images. This framework is promising because the depth and pose networks can be trained from just time-sequence images without the need for the ground truth depth and poses. In…

Cited by 4SourceScholar
2022

Ex-DoF: Expansion of Action Degree-of-Freedom with Virtual Camera Rotation for Omnidirectional Image

ICRA 2022poster

Inter-robot transfer of training data is a little explored topic in learning- and vision-based robot control. Here we propose a transfer method from a robot with a lower Degree-of-Freedom (DoF) to one with a higher DoF utilizing the omnidirectional camera image. The virtual rotation of the robot cam…

Cited by 2SourceScholar
2022

Unsupervised Simultaneous Learning for Camera Re-Localization and Depth Estimation from Video

IROS 2022poster

We present an unsupervised simultaneous learning framework for the task of monocular camera re-localization and depth estimation from unlabeled video sequences. Monocular camera re-localization refers to the task of estimating the absolute camera pose from an instance image in a known environment, w…

Cited by 1SourceScholar
2021

PLG-IN: Pluggable Geometric Consistency Loss with Wasserstein Distance in Monocular Depth Estimation

ICRA 2021poster

We propose a novel objective for penalizing geometric inconsistencies and improving the depth and pose estimation performance of monocular camera images. Our objective is designed using the Wasserstein distance between two point clouds, estimated from images with different camera poses. The Wasserst…

Cited by 7SourceScholar
2021

Probabilistic Visual Navigation with Bidirectional Image Prediction

IROS 2021poster

Humans can robustly follow a visual trajectory defined by a sequence of images (i.e. a video) regardless of substantial changes in the environment or the presence of obstacles. We aim at endowing similar visual navigation capabilities to mobile robots solely equipped with a RGB fisheye camera. We pr…

Cited by 8SourceScholar
2019

Deep Visual MPC-Policy Learning for Navigation

RA-L 2019

Humans can routinely follow a trajectory defined by a list of images/landmarks. However, traditional robot navigation methods require accurate mapping of the environment, localization, and planning. Moreover, these methods are sensitive to subtle changes in the environment. In this letter, we propos

Cited by 114SourceScholar
2019

SoPhie: An Attentive GAN for Predicting Paths Compliant to Social and Physical Constraints

CVPR 2019poster

This paper addresses the problem of path prediction for multiple interacting agents in a scene, which is a crucial step for many autonomous platforms such as self-driving cars and social robots. We present SoPhie; an interpretable framework based on Generative Adversarial Network (GAN), which levera…

Cited by 1245PDFcodeScholar
2019

VUNet: Dynamic Scene View Synthesis for Traversability Estimation Using an RGB Camera

RA-L 2019

We present VUNet, a novel view(VU) synthesis method for mobile robots in dynamic environments, and its application to the estimation of future traversability. Our method predicts future images for given virtual robot velocity commands using only RGB images at previous and current time steps. The fut

Cited by 40SourceScholar
2018

GONet: A Semi-Supervised Deep Learning Approach For Traversability Estimation

IROS 2018poster

We present semi-supervised deep learning approaches for traversability estimation from fisheye images. Our method, GONet, and the proposed extensions leverage Generative Adversarial Networks (GANs) to effectively predict whether the area seen in the input image(s) is safe for a robot to traverse. Th…

Cited by 74SourceScholar
2015

Personal robot assisting transportation to support active human life

IROS 2015poster

Advanced countries are currently experiencing an aging society, and many people could benefit from a personal robot that can support a comfortable lifestyle. However, excessive and premature robot assistance may deteriorate the user's physical abilities and can accelerate their aging process. The au…

Cited by 17SourceScholar