← Search

Yizhe Zhu

19 accepted papers

2026

Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis

ICML 2026poster

Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improvement. However, existing embodied benchmarks fail to provide actionable insights because they focus on task-level evaluation rather than discovering capability bottlenecks. To address t…

Cited by 0SourceScholar
2026

EquAct: An SE(3)-Equivariant Multi-Task Transformer for 3D Robotic Manipulation

ICLR 2026poster

Multi-task manipulation policy often builds on transformer's ability to jointly process language instructions and 3D observations in a shared embedding space. However, real-world tasks frequently require robots to generalize to novel 3D object poses. Policies based on shared embedding break geometri…

Cited by 0SourceScholar
2026

Generalizable Hierarchical Skill Learning via Object-Centric Representation

RA-L 2026

We present Generalizable Hierarchical Skill Learning (GSL), a novel framework for hierarchical policy learning that improves policy generalization and sample efficiency in robot manipulation. One core idea of GSL is to use object-centric skills as an interface that bridges the high-level vision-lang

Cited by 3SourceScholar
2025

Hierarchical Equivariant Policy via Frame Transfer

ICML 2025poster

Recent advances in hierarchical policy learning highlight the advantages of decomposing systems into high-level and low-level agents, enabling efficient long-horizon reasoning and precise fine-grained control. However, the interface between these hierarchy levels remains underexplored, and existing…

Cited by 2SourcePDFScholar
2024

MagicPose: Realistic Human Poses and Facial Expressions Retargeting with Identity-aware Diffusion

ICML 2024poster

In this work, we propose MagicPose, a diffusion-based model for 2D human pose and facial expression retargeting. Specifically, given a reference image, we aim to generate a person's new images by controlling the poses and facial expressions while keeping the identity unchanged. To this end, we propo…

2023

Shifted Diffusion for Text-to-Image Generation

CVPR 2023poster

We present Corgi, a novel method for text-to-image generation. Corgi is based on our proposed shifted diffusion model, which achieves better image embedding generation from input text. Different from the baseline diffusion model used in DALL-E 2, our method seamlessly encodes prior knowledge of the…

2021

TIME: Text and Image Mutual-Translation Adversarial Networks

AAAI 2021technical

Focusing on text-to-image (T2I) generation, we propose Text and Image Mutual-Translation Adversarial Networks (TIME), a lightweight but effective model that jointly learns a T2I generator G and an image captioning discriminator D under the Generative Adversarial Network framework. While previous met…

Cited by 39SourcePDFScholar
2021

Towards Faster and Stabilized GAN Training for High-fidelity Few-shot Image Synthesis

ICLR 2021poster

Training Generative Adversarial Networks (GAN) on high-fidelity images usually requires large-scale GPU-clusters and a vast number of training images. In this paper, we study the few-shot image synthesis task for GAN with minimum computing cost. We propose a light-weight GAN structure that gains sup…

2020

S3VAE: Self-Supervised Sequential VAE for Representation Disentanglement and Data Generation

CVPR 2020poster

We propose a sequential variational autoencoder to learn disentangled representations of sequential data (e.g., videos and audios) under self-supervision. Specifically, we exploit the benefits of some readily accessible supervision signals from input data itself or some off-the-shelf functional mode…

Cited by 136PDFScholar
2019

Learning Feature-to-Feature Translator by Alternating Back-Propagation for Generative Zero-Shot Learning

ICCV 2019poster

We investigate learning feature-to-feature translator networks by alternating back-propagation as a general-purpose solution to zero-shot learning (ZSL) problems. It is a generative model-based ZSL framework. In contrast to models based on generative adversarial networks (GAN) or variational autoenc…

Cited by 127PDFcodeScholar
2019

Semantic-Guided Multi-Attention Localization for Zero-Shot Learning

NeurIPS 2019poster

Zero-shot learning extends the conventional object classification to the unseen class recognition by introducing semantic representations of classes. Existing approaches predominantly focus on learning the proper mapping function for visual-semantic embedding, while neglecting the effect of learning…

Cited by 180SourcePDFScholar
2018

A Generative Adversarial Approach for Zero-Shot Learning From Noisy Texts

CVPR 2018poster

Most existing zero-shot learning methods consider the problem as a visual semantic embedding one. Given the demonstrated capability of Generative Adversarial Networks(GANs) to generate images, we instead leverage GANs to imagine unseen categories from text descriptions and hence recognize novel clas…

Cited by 499SourcePDFScholar
2017

Link the Head to the "Beak": Zero Shot Learning From Noisy Text Description at Part Precision

CVPR 2017poster

In this paper, we study learning visual classifiers from unstructured text description at part precision with no training images. We show that visual text terms can be encouraged to attend to its relevant parts, while image connections to non-visual text terms vanishes without any supervision. Thi…

Cited by 158PDFScholar