← Search

Yawen Zhang

4 accepted papers

2026

OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation

CVPR 2026

Earth observation data presents a unique challenge: it is spatial like images, sequential like video or text, and highly multimodal. We present Helios: a multimodal, spatio-temporal foundation model that employs a novel self-supervised learning formulation, masking strategy, and loss all designed fo

Cited by 0SourcecodeScholar
2025

MoLA: MoE LoRA with Layer-wise Expert Allocation

NAACL 2025findings

Recent efforts to integrate low-rank adaptation (LoRA) with the Mixture-of-Experts (MoE) have managed to achieve performance comparable to full-parameter fine-tuning by tuning much fewer parameters. Despite promising results, research on improving the efficiency and expert analysis of LoRA with MoE…

2025

Non-Pharmacological Interventions: A Virtual Training Framework for Fine Motor Learning

ICASSP 2025accepted

In contemporary rehabilitation treatments, it is common to integrate advanced technologies such as virtual reality and human-computer interaction. However, there is insufficient research on their efficacy in enhancing the fine motor skills of children with developmental coordination disorder (DCD).…

Cited by 0SourceScholar
2024

Tackling Vision Language Tasks through Learning Inner Monologues

AAAI 2024technical

Visual language tasks such as Visual Question Answering (VQA) or Visual Entailment (VE) require AI models to comprehend and reason with both visual and textual content. Driven by the power of Large Language Models (LLMs), two prominent methods have emerged: (1) the hybrid integration between LLMs an…