← Search

Yanzhe Chen

5 accepted papers

2026

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation

ICML 2026poster

While Vision-Language-Action (VLA) models offer broad general capabilities, deploying them on specific hardware requires real-world adaptation to bridge the embodiment gap. Since robot demonstrations are costly, this adaptation must often occur under a strict data budget. In this work, we identify a…

Cited by 0SourceScholar
2026

UniAPO: Unified Multimodal Automated Prompt Optimization

AAAI 2026technical

Prompting is fundamental to unlocking the full potential of large language models. To automate and enhance this process, automatic prompt optimization (APO) has been developed, demonstrating effectiveness primarily in text-only input scenarios. However, extending existing APO methods to multimodal t

Cited by 0SourcePDFScholar
2025

MAI: A Multi-turn Aggregation-Iteration Model for Composed Image Retrieval

ICLR 2025poster

Multi-Turn Composed Image Retrieval (MTCIR) addresses a real-world scenario where users iteratively refine retrieval results by providing additional information until a target meeting all their requirements is found. Existing methods primarily achieve MTCIR through a "multiple single-turn" paradigm,…

Cited by 0SourcePDFScholar
2024

FashionERN: Enhance-and-Refine Network for Composed Fashion Image Retrieval

AAAI 2024technical

The goal of composed fashion image retrieval is to locate a target image based on a reference image and modified text. Recent methods utilize symmetric encoders (e.g., CLIP) pre-trained on large-scale non-fashion datasets. However, the input for this task exhibits an asymmetric nature, where the ref…

Cited by 5SourcePDFScholar