← Search

Yasutomo Kawanishi

6 accepted papers

2026

Action Motifs: Self-Supervised Hierarchical Representation of Human Body Movements

CVPR 2026

Effective human behavior modeling requires a representation of the human body movement that capitalizes on its compositionality. We propose a hierarchical representation consisting of Action Atoms that capture the atomic joint movements and Action Motifs that are formed by their temporal composition

Cited by 0SourceScholar
2025

FROSS: Faster-Than-Real-Time Online 3D Semantic Scene Graph Generation from RGB-D Images

ICCV 2025poster

The ability to abstract complex 3D environments into simplified and structured representations is crucial across various domains. 3D semantic scene graphs (SSGs) achieve this by representing objects as nodes and their interrelationships as edges, facilitating high-level scene understanding. Existing…

2024

A Gaze-grounded Visual Question Answering Dataset for Clarifying Ambiguous Japanese Questions

COLING 2024main

Situated conversations, which refer to visual information as visual question answering (VQA), often contain ambiguities caused by reliance on directive information. This problem is exacerbated because some languages, such as Japanese, often omit subjective or objective terms. Such ambiguities in que…

2024

J-CRe3: A Japanese Conversation Dataset for Real-world Reference Resolution

COLING 2024main

Understanding expressions that refer to the physical world is crucial for such human-assisting systems in the real world, as robots that must perform actions that are expected by users. In real-world reference resolution, a system must ground the verbal information that appears in user interactions…

2023

DeePoint: Visual Pointing Recognition and Direction Estimation

ICCV 2023poster

In this paper, we realize automatic visual recognition and direction estimation of pointing. We introduce the first neural pointing understanding method based on two key contributions. The first is the introduction of a first-of-its-kind large-scale dataset for pointing recognition and direction est…

Cited by 5PDFcodeScholar
2023

Operative Action Captioning for Estimating System Actions

ICRA 2023poster

Human-assistive systems, such as robots, need to correctly understand the surrounding situation based on obser-vations and output the required support actions for humans. Language is one of the important channels to communicate with humans, and robots are required to have the ability to express thei…

Cited by 0SourceScholar