AXG-Reasoner: Error Detection and Explanation in Long Task Videos with Vision-Language Models
Virtual task assistants must recognize and explain users' mistakes to provide effective and corrective guidance. In this paper, we address the problem of error reasoning in long task videos, which is to detect and explain errors. Although recent Vision-Language Models (VLMs) demonstrate strong capab