← Search

Mingjie Han

3 accepted papers

2026

Activating Visual Context and Commonsense Reasoning Through Masked Prediction in VLMs

AAAI 2026technical

Recent breakthroughs in reasoning models have markedly advanced the reasoning capabilities of large language models, particularly via training on tasks with verifiable rewards. Yet, a significant gap persists in their adaptation to real-world multimodal scenarios, most notably, vision-language tasks

Cited by 0SourcePDFScholar
2021

A Generative Model-Based Predictive Display for Robotic Teleoperation

ICRA 2021poster

We propose a new generative model-based predictive display for robotic teleoperation over high-latency communication links. Our method is capable of rendering photo-realistic images of the scene to the human operator in real time from RGB-D images acquired by the remote robot. A preliminary explorat…

Cited by 4SourceScholar
2021

Image-Based Joint State Estimation Pipeline for Sensorless Manipulators

IROS 2021poster

Motion planning is a largely solved problem for robot arms with joint state feedback, but remains an area of research for sensorless manipulators such as toy robot arms and heavy equipment such as excavators and cranes. A promising approach to this problem is deep learning, which employs a pre-train…

Cited by 2SourceScholar