Interactive Robot Action Replanning using Multimodal LLM Trained from Human Demonstration Videos
Understanding human actions could allow robots to perform a large spectrum of complex manipulation tasks and make collaboration with humans easier. Recently, multimodal scene understanding using audio-visual Transformers has been used to generate robot action sequences from videos of human demonstra…