← Search

Zhizhen Zhang

3 accepted papers

2026

MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent

CVPR 2026

Recent Vision-Language-Action (VLA) models reformulate vision-language models by tuning them with millions of robotic demonstrations. While they perform well when fine-tuned for a single embodiment or task family, extending them to multi-skill settings remains challenging: directly merging VLA exper

Cited by 0SourcecodeScholar
2026

StableI2I: Spotting Unintended Changes in Image-to-Image Transition

ICML 2026poster

In most real-world image-to-image (I2I) scenarios, existing evaluations primarily focus on instruction following and the perceptual quality or aesthetics of the generated images. However, they largely fail to assess whether the output image preserves the semantic correspondence and spatial structure…

Cited by 0SourceScholar
2025

Provable Ordering and Continuity in Vision-Language Pretraining for Generalizable Embodied Agents

NeurIPS 2025poster

Pre-training vision-language representations on human action videos has emerged as a promising approach to reduce reliance on large-scale expert demonstrations for training embodied agents. However, prior methods often employ time con- trastive learning based on goal-reaching heuristics, progressive…

Cited by 0SourcecodeScholar