← Search

Zhenyu Zhao

7 accepted papers

2026

FrameOracle: Learning What to See and How Much to See in Videos

ICML 2026poster

Vision-language models (VLMs) advance video understanding but operate under tight computational budgets, making performance dependent on selecting a small, high-quality subset of frames. Existing frame sampling strategies, such as uniform or fixed-budget selection, fail to adapt to variations in con…

Cited by 3SourceScholar
2026

Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation

ICRA 2026poster

From loco-motion to dextrous manipulation, humanoid robots have made remarkable strides in demonstrating complex full-body capabilities. However, the majority of current robot learning datasets and benchmarks mainly focus on stationary robot arms, and the few existing humanoid datasets are either co…

2026

Ψ0Ψ0\Psi_0: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation

RSS 2026poster

We introduce Ψ₀ (Psi-Zero), an open foundation model to address challenging humanoid loco-manipulation tasks. While existing approaches often attempt to address this fundamental problem by co-training on large and diverse human and humanoid data, we argue that this strategy is suboptimal due to the …

Cited by 0SourceScholar
2025

LEDiT: Your Length-Extrapolatable Diffusion Transformer without Positional Encoding

NeurIPS 2025poster

Diffusion transformers (DiTs) struggle to generate images at resolutions higher than their training resolutions. The primary obstacle is that the explicit positional encodings (PE), such as RoPE, need extrapolating to unseen positions which degrades performance when the inference resolution differs…

Cited by 0SourcecodeScholar
2024

HiDiffusion: Unlocking Higher-Resolution Creativity and Efficiency in Pretrained Diffusion Models

ECCV 2024poster

"Diffusion models have become a mainstream approach for high-resolution image synthesis. However, directly generating higher-resolution images from pretrained diffusion models will encounter unreasonable object duplication and exponentially increase the generation time. In this paper, we discover th…

Cited by 5SourcePDFScholar
2020

Robust Machine Reading Comprehension by Learning Soft labels

COLING 2020main

Neural models have achieved great success on the task of machine reading comprehension (MRC), which are typically trained on hard labels. We argue that hard labels limit the model capability on generalization due to the label sparseness problem. In this paper, we propose a robust training method for…

2019

Learning Compact Partial Differential Equations for Color Images with Efficiency

ICASSP 2019accepted

Learning Partial Differential Equations (LPDEs) from training data for particular tasks has been successfully applied to many image processing problems. In this paper, we aim to learn compact Partial Differential Equations (LCPDEs) for color image tasks by proposing a more effective algorithm. The L…

Cited by 0SourceScholar