← Search

Kodai Nakashima

5 accepted papers

2023

SegRCDB: Semantic Segmentation via Formula-Driven Supervised Learning

ICCV 2023poster

Pre-training is a strong strategy for enhancing visual models to efficiently train them with a limited number of labeled images. In semantic segmentation, creating annotation masks requires an intensive amount of labor and time, and therefore, a large-scale pre-training dataset with semantic labels…

Cited by 13PDFcodeScholar
2022

Can Vision Transformers Learn without Natural Images?

AAAI 2022technical

Is it possible to complete Vision Transformer (ViT) pre-training without natural images and human-annotated labels? This question has become increasingly relevant in recent months because while current ViT pre-training tends to rely heavily on a large number of natural images and human-annotated lab…

Cited by 38SourcePDFScholar
2022

Replacing Labeled Real-Image Datasets With Auto-Generated Contours

CVPR 2022poster

In the present work, we show that the performance of formula-driven supervised learning (FDSL) can match or even exceed that of ImageNet-21k without the use of real images, human-, and self-supervision during the pre-training of Vision Transformers (ViTs). For example, ViT-Base pre-trained on ImageN…

Cited by 44PDFScholar
2021

Describing and Localizing Multiple Changes With Transformers

ICCV 2021poster

Existing change captioning studies have mainly focused on a single change. However, detecting and describing multiple changed parts in image pairs is essential for enhancing adaptability to complex scenarios. We solve the above issues from three aspects: (i) We propose a simulation-based multi-chang…

Cited by 65PDFScholar
2020

Joint Pedestrian Detection and Risk-level Prediction with Motion-Representation-by-Detection

ICRA 2020poster

The paper presents a pedestrian near-miss detector with temporal analysis that provides both pedestrian detection and risk-level predictions which are demonstrated on a self-collected database. Our work makes three primary contributions: (i) The framework of pedestrian near-miss detection is propose…

Cited by 6SourceScholar