Causal Transformer-Based Online Action Recognition for High-Level Control of a Unitree Go1 Robot
Chaitanya Bandi, Kristof Kitz, Ulrike Thomas
Abstract
We present a new causal transformer system consisting of Spatial-Attention Tokenization (SAT) with MultiResolution Causal Temporal Mixing (MRCTM) to perform online skeleton-based action recognition during human–robot interaction. The novel architecture uses Spatial-Attention Tokenization (SAT) to generate soft tokens from human joint groups. MRCTM performs causal convolutions and selfattention operations to detect both detailed motion patterns and extended temporal relationships. We introduce GoHAR12 dataset as an evaluation tool as it contains 12 gesture and posture classes which are recorded in human-robot interaction (HRI) settings and directly translate to high-level commands for the Unitree Go1 quadruped. The proposed model reaches 98.4% accuracy on the GoHAR-12 dataset and it shows superior performance in distinguishing between actions that are quite similar in motion, maintains strong results on public benchmarks such as NTU-RGB+D and NW-UCLA. We demonstrate how causal transformer performs for reliable realtime skeleton-based control of the Unitree Go1 robot.