← Search

Yutaka Satoh

12 accepted papers

2024

Formula-Supervised Visual-Geometric Pre-training

ECCV 2024poster

"Throughout the history of computer vision, while research has explored the integration of images (visual) and point clouds (geometric), many advancements in image and 3D object recognition have tended to process these modalities separately. We aim to bridge this divide by integrating images and poi…

2024

Subtle-Diff: A Dataset for Precise Recognition of Subtle Differences Among Visually Similar Objects

IROS 2024poster

Visual inspection robots used in factories and outdoor environments require the ability to accurately recognize visual differences between similar objects and further verbalize the recognition results to present the differences to humans. Despite the application of Large Language Models (LLMs) and m…

Cited by 0SourceScholar
2023

Question Generation for Uncertainty Elimination in Referring Expressions in 3D Environments

ICRA 2023poster

We introduce a new task of question generation to eliminate the uncertainty of referring expressions in 3D indoor environments (3D-REQ). Referring to an object using natural language is one of the most common occurrences in daily human conversations; therefore, instructing robots to identify a certa…

Cited by 2SourceScholar
2022

Can Vision Transformers Learn without Natural Images?

AAAI 2022technical

Is it possible to complete Vision Transformer (ViT) pre-training without natural images and human-annotated labels? This question has become increasingly relevant in recent months because while current ViT pre-training tends to rely heavily on a large number of natural images and human-annotated lab…

Cited by 38SourcePDFScholar
2021

Describing and Localizing Multiple Changes With Transformers

ICCV 2021poster

Existing change captioning studies have mainly focused on a single change. However, detecting and describing multiple changed parts in image pairs is essential for enhancing adaptability to complex scenarios. We solve the above issues from three aspects: (i) We propose a simulation-based multi-chang…

Cited by 65PDFScholar
2020

Joint Pedestrian Detection and Risk-level Prediction with Motion-Representation-by-Detection

ICRA 2020poster

The paper presents a pedestrian near-miss detector with temporal analysis that provides both pedestrian detection and risk-level predictions which are demonstrated on a self-collected database. Our work makes three primary contributions: (i) The framework of pedestrian near-miss detection is propose…

Cited by 6SourceScholar
2018

Anticipating Traffic Accidents With Adaptive Loss and Large-Scale Incident DB

CVPR 2018poster

In this paper, we propose a novel approach for traffic accident anticipation through (i) Adaptive Loss for Early Anticipation (AdaLEA) and (ii) a large-scale self-annotated incident database. The proposed AdaLEA allows us to gradually learn an earlier anticipation as training progresses. The loss fu…

Cited by 157SourcePDFScholar
2018

Can Spatiotemporal 3D CNNs Retrace the History of 2D CNNs and ImageNet?

CVPR 2018poster

The purpose of this study is to determine whether current video datasets have sufficient data for training very deep convolutional neural networks (CNNs) with spatio-temporal three-dimensional (3D) kernels. Recently, the performance levels of 3D CNNs in the field of action recognition have improved…

2018

Drive Video Analysis for the Detection of Traffic Near-Miss Incidents

ICRA 2018poster

Because of their recent introduction, self-driving cars and advanced driver assistance system (ADAS) equipped vehicles have had little opportunity to learn, the dangerous traffic (including near-miss incident) scenarios that provide normal drivers with strong motivation to drive safely. Accordingly,…

Cited by 49SourceScholar
2017

Illuminant-Camera Communication to Observe Moving Objects Under Strong External Light by Spread Spectrum Modulation

CVPR 2017spotlight

Many algorithms of computer vision use light sources to illuminate objects to actively create situation appropriate to extract their characteristics. For example, the shape and reflectance are measured by a projector-camera system, and some human-machine or VR systems use projectors and displays for…

Cited by 12PDFScholar