← Search

Hirokatsu Kataoka

34 accepted papers

2026

3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Clouds

CVPR 2026

Despite recent progress in 3D self-supervised learning, collecting large-scale 3D scene scans remains expensive and labor-intensive. In this work, we investigate whether 3D representations can be learned from unlabeled videos recorded without any real 3D sensors. We present Laplacian-Aware Multi-lev

Cited by 0SourcecodeScholar
2026

CLIP-like Model as a Foundational Density Ratio Estimator

CVPR 2026

Density ratio estimation is a core concept in statistical machine learning because it provides a unified mechanism for tasks such as importance weighting, divergence estimation, and likelihood-free inference, but its potential in vision and language models has not been fully explored. Modern vision-

Cited by 0SourcecodeScholar
2026

HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics

CVPR 2026

Hand-object interaction (HOI) involves dynamics where human manipulations produce spatio-temporal effects on objects. However, existing semantic HOI benchmarks focus on either manipulation or effects at a coarse level, lacking fine-grained spatio-temporal reasoning to capture HOI dynamics. We introd

Cited by 0SourceScholar
2026

PowerCLIP: Powerset Alignment for Contrastive Pre-Training

CVPR 2026

Contrastive pre-training frameworks such as CLIP have demonstrated impressive zero-shot performance across a range of vision-language tasks. Recent studies have shown that aligning individual text tokens with specific image patches or regions enhances fine-grained compositional understanding. Howeve

Cited by 0SourcecodeScholar
2026

S3OD: Towards Generalizable Salient Object Detection with Synthetic Data

ICLR 2026poster

Salient object detection exemplifies data-bounded tasks where expensive pixel-precise annotations force separate model training for related subtasks like DIS and HR-SOD. We present a method that dramatically improves generalization through large-scale synthetic data generation and ambiguity-aware ar…

Cited by 0SourcecodeScholar
2025

AgroBench: Vision-Language Model Benchmark in Agriculture

ICCV 2025poster

Precise automated understanding of agricultural tasks such as disease identification is essential for the sustainable crop production. Recent advances in vision-language models (VLMs) are expected to further expand the range of agricultural tasks by facilitating human-model interaction through easy,…

2025

AnimalClue: Recognizing Animals by their Traces

ICCV 2025poster

Wildlife observation plays an important role in biodiversity conservation, necessitating robust methodologies for monitoring wildlife populations and interspecies interactions. Recent advances in computer vision have significantly contributed to automating fundamental wildlife observation tasks, suc…

2025

Approximate Domain Unlearning for Vision-Language Models

NeurIPS 2025spotlight

Pre-trained Vision-Language Models (VLMs) exhibit strong generalization capabilities, enabling them to recognize a wide range of objects across diverse domains without additional training. However, they often retain irrelevant information beyond the requirements of specific target downstream tasks,…

Cited by 0SourcecodeScholar
2025

Formula-Supervised Sound Event Detection: Pre-Training Without Real Data

ICASSP 2025accepted

In this paper, we propose a novel formula-driven supervised learning (FDSL) framework for pre-training an environmental sound analysis model by leveraging acoustic signals parametrically synthesized through formula-driven methods. Specifically, we outline detailed procedures and evaluate their effec…

Cited by 0SourceScholar
2024

Formula-Supervised Visual-Geometric Pre-training

ECCV 2024poster

"Throughout the history of computer vision, while research has explored the integration of images (visual) and point clouds (geometric), many advancements in image and 3D object recognition have tended to process these modalities separately. We aim to bridge this divide by integrating images and poi…

2024

Guided by the Way: The Role of On-the-route Objects and Scene Text in Enhancing Outdoor Navigation

ICRA 2024poster

In outdoor environments, Vision-and-Language Navigation (VLN) requires an agent to rely on multi-modal cues from real-world urban environments and natural language instructions. While existing outdoor VLN models predict actions using a combination of panorama and instruction features, this approach…

Cited by 1SourceScholar
2024

Pseudo-Outlier Synthesis Using Q-Gaussian Distributions for Out-of-Distribution Detection

ICASSP 2024accepted

Out-of-distribution (OOD) detection, which aims to determine whether an input is outside the training data distribution or not, is an indispensable task in many computer vision applications. In many of the previous studies on OOD detection for image classification, the class-conditional distribution…

Cited by 0SourceScholar
2024

Rethinking Image Super Resolution from Training Data Perspectives

ECCV 2024poster

"In this work, we investigate the understudied effect of the training data used for image super-resolution (SR). Most commonly, novel SR methods are developed and benchmarked on common training datasets such as DIV2K and DF2K. However, we investigate and rethink the training data from the perspectiv…

2024

Subtle-Diff: A Dataset for Precise Recognition of Subtle Differences Among Visually Similar Objects

IROS 2024poster

Visual inspection robots used in factories and outdoor environments require the ability to accurately recognize visual differences between similar objects and further verbalize the recognition results to present the differences to humans. Despite the application of Large Language Models (LLMs) and m…

Cited by 0SourceScholar
2024

Watermark-embedded Adversarial Examples for Copyright Protection against Diffusion Models

CVPR 2024poster

Diffusion Models (DMs) have shown remarkable capabilities in various image-generation tasks. However there are growing concerns that DMs could be used to imitate unauthorized creations and thus raise copyright issues. To address this issue we propose a novel framework that embeds personal watermarks…

Cited by 13SourcePDFScholar
2023

Graph Representation for Order-Aware Visual Transformation

CVPR 2023poster

This paper proposes a new visual reasoning formulation that aims at discovering changes between image pairs and their temporal orders. Recognizing scene dynamics and their chronological orders is a fundamental aspect of human cognition. The aforementioned abilities make it possible to follow step-by…

Cited by 4SourcePDFScholar
2023

Pre-training Vision Transformers with Very Limited Synthesized Images

ICCV 2023poster

Formula-driven supervised learning (FDSL) is a pre-training method that relies on synthetic images generated from mathematical formulae such as fractals. Prior work on FDSL has shown that pre-training vision transformers on such synthetic datasets can yield competitive accuracy on a wide range of d…

Cited by 10PDFcodeScholar
2023

Question Generation for Uncertainty Elimination in Referring Expressions in 3D Environments

ICRA 2023poster

We introduce a new task of question generation to eliminate the uncertainty of referring expressions in 3D indoor environments (3D-REQ). Referring to an object using natural language is one of the most common occurrences in daily human conversations; therefore, instructing robots to identify a certa…

Cited by 2SourceScholar
2023

SegRCDB: Semantic Segmentation via Formula-Driven Supervised Learning

ICCV 2023poster

Pre-training is a strong strategy for enhancing visual models to efficiently train them with a limited number of labeled images. In semantic segmentation, creating annotation masks requires an intensive amount of labor and time, and therefore, a large-scale pre-training dataset with semantic labels…

Cited by 13PDFcodeScholar
2023

Traffic Incident Database with Multiple Labels Including Various Perspective Environmental Information

IROS 2023poster

Traffic accident recognition is essential in developing automated driving and Advanced Driving Assistant System technologies. A large dataset of annotated traffic accidents is necessary to improve the accuracy of traffic accident recognition using deep learning models. Conventional traffic accident…

Cited by 0SourcecodeScholar
2023

Visual Atoms: Pre-Training Vision Transformers With Sinusoidal Waves

CVPR 2023poster

Formula-driven supervised learning (FDSL) has been shown to be an effective method for pre-training vision transformers, where ExFractalDB-21k was shown to exceed the pre-training effect of ImageNet-21k. These studies also indicate that contours mattered more than textures when pre-training vision t…

Cited by 27SourcePDFScholar
2022

Can Vision Transformers Learn without Natural Images?

AAAI 2022technical

Is it possible to complete Vision Transformer (ViT) pre-training without natural images and human-annotated labels? This question has become increasingly relevant in recent months because while current ViT pre-training tends to rely heavily on a large number of natural images and human-annotated lab…

Cited by 38SourcePDFScholar
2022

Neural Density-Distance Fields

ECCV 2022poster

"The success of neural fields for 3D vision tasks is now indisputable. Following this trend, several methods aiming for visual localization (e.g., SLAM) have been proposed to estimate distance or density fields using neural fields. However, it is difficult to achieve high localization performance by…

2022

Point Cloud Pre-Training With Natural 3D Structures

CVPR 2022poster

The construction of 3D point cloud datasets requires a great deal of human effort. Therefore, constructing a largescale 3D point clouds dataset is difficult. In order to remedy this issue, we propose a newly developed point cloud fractal database (PC-FractalDB), which is a novel family of formula-dr…

Cited by 46PDFcodeScholar
2022

Replacing Labeled Real-Image Datasets With Auto-Generated Contours

CVPR 2022poster

In the present work, we show that the performance of formula-driven supervised learning (FDSL) can match or even exceed that of ImageNet-21k without the use of real images, human-, and self-supervision during the pre-training of Vision Transformers (ViTs). For example, ViT-Base pre-trained on ImageN…

Cited by 44PDFScholar
2021

Describing and Localizing Multiple Changes With Transformers

ICCV 2021poster

Existing change captioning studies have mainly focused on a single change. However, detecting and describing multiple changed parts in image pairs is essential for enhancing adaptability to complex scenarios. We solve the above issues from three aspects: (i) We propose a simulation-based multi-chang…

Cited by 65PDFScholar
2021

MV-FractalDB: Formula-driven Supervised Learning for Multi-view Image Recognition

IROS 2021poster

The paper proposes a method for automatic multi-view dataset construction based on formula-driven supervised learning (FDSL). Although data collection and human annotation of 3D objects are labor-intensive, we automatically generate their training data and labels in the proposed multi-view dataset.…

Cited by 10SourceScholar
2020

Joint Pedestrian Detection and Risk-level Prediction with Motion-Representation-by-Detection

ICRA 2020poster

The paper presents a pedestrian near-miss detector with temporal analysis that provides both pedestrian detection and risk-level predictions which are demonstrated on a self-collected database. Our work makes three primary contributions: (i) The framework of pedestrian near-miss detection is propose…

Cited by 6SourceScholar
2018

Anticipating Traffic Accidents With Adaptive Loss and Large-Scale Incident DB

CVPR 2018poster

In this paper, we propose a novel approach for traffic accident anticipation through (i) Adaptive Loss for Early Anticipation (AdaLEA) and (ii) a large-scale self-annotated incident database. The proposed AdaLEA allows us to gradually learn an earlier anticipation as training progresses. The loss fu…

Cited by 157SourcePDFScholar
2018

Can Spatiotemporal 3D CNNs Retrace the History of 2D CNNs and ImageNet?

CVPR 2018poster

The purpose of this study is to determine whether current video datasets have sufficient data for training very deep convolutional neural networks (CNNs) with spatio-temporal three-dimensional (3D) kernels. Recently, the performance levels of 3D CNNs in the field of action recognition have improved…

2018

Drive Video Analysis for the Detection of Traffic Near-Miss Incidents

ICRA 2018poster

Because of their recent introduction, self-driving cars and advanced driver assistance system (ADAS) equipped vehicles have had little opportunity to learn, the dangerous traffic (including near-miss incident) scenarios that provide normal drivers with strong motivation to drive safely. Accordingly,…

Cited by 49SourceScholar