← Search

Jiayao Ma

4 accepted papers

2025

Multimodal Large Language Models for Text-rich Image Understanding: A Comprehensive Review

ACL 2025finding

The recent emergence of Multi-modal Large Language Models (MLLMs) has introduced a new dimension to the Text-rich Image Understanding (TIU) field, with models demonstrating impressive and inspiring performance. However, their rapid evolution and widespread adoption have made it increasingly challeng…

Cited by 0SourcePDFScholar
2024

Advancing Virtual Reality Interaction: A Ring-Shaped Controller and Pose Tracking

ICRA 2024poster

Ensuring robust tracking of controllers’ movement is critical for human-robot interaction in virtual reality (VR) scenarios. This paper proposes a robust tracking algorithm based on a novel wearable ring-shaped controller equipped with an inertial measurement unit (IMU) and a light-emitting diode (L…

Cited by 0SourceScholar
2021

Hierarchical Temporal Multi-Instance Learning for Video-based Student Learning Engagement Assessment

IJCAI 2021poster

Video-based automatic assessment of a student's learning engagement on the fly can provide immense values for delivering personalized instructional services, a vehicle particularly important for massive online education. To train such an assessor, a major challenge lies in the collection of sufficie…

Cited by 10SourcePDFScholar