← Search

Kenji Iwata

8 accepted papers

2024

DailySTR: A Daily Human Activity Pattern Recognition Dataset for Spatio-temporal Reasoning

IROS 2024poster

Recognizing daily human activities is essential for domestic robots to assist humans effectively in indoor environments. These activities typically involve sequences of interactions between humans and objects across different locations and times within a household. Identifying these events and under…

Cited by 0SourceScholar
2024

Subtle-Diff: A Dataset for Precise Recognition of Subtle Differences Among Visually Similar Objects

IROS 2024poster

Visual inspection robots used in factories and outdoor environments require the ability to accurately recognize visual differences between similar objects and further verbalize the recognition results to present the differences to humans. Despite the application of Large Language Models (LLMs) and m…

Cited by 0SourceScholar
2024

The STVchrono Dataset: Towards Continuous Change Recognition in Time

CVPR 2024poster

Recognizing continuous changes offers valuable insights into past historical events supports current trend analysis and facilitates future planning. This knowledge is crucial for a variety of fields such as meteorology and agriculture environmental science urban planning and construction tourism and…

Cited by 8SourcePDFScholar
2023

Graph Representation for Order-Aware Visual Transformation

CVPR 2023poster

This paper proposes a new visual reasoning formulation that aims at discovering changes between image pairs and their temporal orders. Recognizing scene dynamics and their chronological orders is a fundamental aspect of human cognition. The aforementioned abilities make it possible to follow step-by…

Cited by 4SourcePDFScholar
2023

Question Generation for Uncertainty Elimination in Referring Expressions in 3D Environments

ICRA 2023poster

We introduce a new task of question generation to eliminate the uncertainty of referring expressions in 3D indoor environments (3D-REQ). Referring to an object using natural language is one of the most common occurrences in daily human conversations; therefore, instructing robots to identify a certa…

Cited by 2SourceScholar
2022

Can Vision Transformers Learn without Natural Images?

AAAI 2022technical

Is it possible to complete Vision Transformer (ViT) pre-training without natural images and human-annotated labels? This question has become increasingly relevant in recent months because while current ViT pre-training tends to rely heavily on a large number of natural images and human-annotated lab…

Cited by 38SourcePDFScholar
2021

Describing and Localizing Multiple Changes With Transformers

ICCV 2021poster

Existing change captioning studies have mainly focused on a single change. However, detecting and describing multiple changed parts in image pairs is essential for enhancing adaptability to complex scenarios. We solve the above issues from three aspects: (i) We propose a simulation-based multi-chang…

Cited by 65PDFScholar