← Search

Yihan Zhao

4 accepted papers

2026

From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models

ICML 2026oral

Latent actions serve as an intermediate representation that enables consistent modeling of vision-language-action (VLA) models across heterogeneous datasets. However, approaches to supervising VLAs with latent actions are fragmented and lack a systematic comparison. This work structures the study of…

Cited by 0SourceScholar
2025

Enhancing Multimodal Model Robustness Under Missing Modalities via Memory-Driven Prompt Learning

IJCAI 2025

Existing multimodal models typically assume the availability of all modalities, leading to significant performance degradation when certain modalities are missing. Recent methods have introduced prompt learning to adapt pretrained models to incomplete data, achieving remarkable performance when the

2024

Attention Shifting to Pursue Optimal Representation for Adapting Multi-granularity Tasks

IJCAI 2024poster

Object recognition in open environments, e.g., video surveillance, poses significant challenges due to the inclusion of unknown and multi-granularity tasks (MGT). However, recent methods exhibit limitations as they struggle to capture subtle differences between different parts within an object and a…

Cited by 1SourcePDFScholar
2024

UFDA: Universal Federated Domain Adaptation with Practical Assumptions

AAAI 2024technical

Conventional Federated Domain Adaptation (FDA) approaches usually demand an abundance of assumptions, which makes them significantly less feasible for real-world situations and introduces security hazards. This paper relaxes the assumptions from previous FDAs and studies a more practical scenario na…