← Search

Shao-Yuan Lo

9 accepted papers

2026

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning

ICML 2026poster

Theory of Mind (ToM) is a must-acquire skill for modern foundation model systems to operate effectively and safely in the real world. Recent works have explored honing ToM via post-training; however, we show that such progress is confounded by a pervasive “shortcut” issue: tasks can reach up to 99% …

Cited by 0SourceScholar
2025

Bridging Compressed Image Latents and Multimodal Large Language Models

ICLR 2025poster

This paper presents the first-ever study of adapting compressed image latents to suit the needs of downstream vision tasks that adopt Multimodal Large Language Models (MLLMs). MLLMs have extended the success of large language models to modalities (e.g. images) beyond text, but their billion scale hi…

Cited by 1SourcePDFScholar
2025

Filter Images First, Generate Instructions Later: Pre-Instruction Data Selection for Visual Instruction Tuning

CVPR 2025highlight

Visual instruction tuning (VIT) for large vision-language models (LVLMs) requires training on expansive datasets of image-instruction pairs, which can be costly. Recent efforts in VIT data selection aim to select a small subset of high-quality image-instruction pairs, reducing VIT runtime while main…

2025

Overcoming Multi-step Complexity in Multimodal Theory-of-Mind Reasoning: A Scalable Bayesian Planner

ICML 2025spotlight

Theory-of-mind (ToM) enables humans to infer mental states—such as beliefs, desires, and intentions—forming the foundation of social cognition. Existing computational ToM methods rely on structured workflows with ToM-specific priors or deep model fine-tuning but struggle with scalability in multimod…

Cited by 0SourcePDFScholar
2025

Towards Zero-Shot Anomaly Detection and Reasoning with Multimodal Large Language Models

CVPR 2025highlight

Zero-Shot Anomaly Detection (ZSAD) is an emerging AD paradigm. Unlike the traditional unsupervised AD setting that requires a large number of normal samples to train a model, ZSAD is more practical for handling data-restricted real-world scenarios. Recently, Multimodal Large Language Models (MLLMs)…

2024

Can't Make an Omelette Without Breaking Some Eggs: Plausible Action Anticipation Using Large Video-Language Models

CVPR 2024poster

We introduce PlausiVL a large video-language model for anticipating action sequences that are plausible in the real-world. While significant efforts have been made towards anticipating future actions prior approaches do not take into account the aspect of plausibility in an action sequence. To addre…

Cited by 18SourcePDFScholar
2024

Uncertainty-aware Action Decoupling Transformer for Action Anticipation

CVPR 2024highlight

Human action anticipation aims at predicting what people will do in the future based on past observations. In this paper we introduce Uncertainty-aware Action Decoupling Transformer (UADT) for action anticipation. Unlike existing methods that directly predict action in a verb-noun pair format we dec…

Cited by 11SourcePDFScholar
2023

Spatio-Temporal Pixel-Level Contrastive Learning-Based Source-Free Domain Adaptation for Video Semantic Segmentation

CVPR 2023poster

Unsupervised Domain Adaptation (UDA) of semantic segmentation transfers labeled source knowledge to an unlabeled target domain by relying on accessing both the source and target data. However, the access to source data is often restricted or infeasible in real-world scenarios. Under the source data…

2022

Learning Feature Decomposition for Domain Adaptive Monocular Depth Estimation

IROS 2022poster

Monocular depth estimation (MDE) has attracted intense study due to its low cost and critical functions for robotic tasks such as localization, mapping and obstacle detection. Supervised approaches have led to great success with the advance of deep learning, but they rely on large quantities of grou…

Cited by 16SourceScholar