← Search

Wenyao Zhang

15 accepted papers

2026

Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining

ICLR 2026poster

Vision-language-action (VLA) models have shown great potential in building generalist robots, but still face a dilemma–misalignment of 2D image forecasting and 3D action prediction. Besides, such a vision-action entangled training manner limits model learning from large-scale, action-free web video…

Cited by 0SourcecodeScholar
2026

Goal-Driven Reward by Video Diffusion Models for Reinforcement Learning

CVPR 2026

Reinforcement Learning (RL) has achieved remarkable success in various domains, yet it often relies on carefully designed programmatic reward functions to guide agent behavior. Designing such reward functions can be challenging and may not generalize well across different tasks. To address this limi

Cited by 0SourceScholar
2026

Humanoid Generative Pre-Training for Zero-Shot Motion Tracking

CVPR 2026

We introduce Humanoid-GPT, a GPT-style Transformer with causal attention trained on a billion-scale motion corpus for whole-body control. Unlike prior shallow MLP trackers constrained by scarce data and an agility-generalization trade-off, Humanoid-GPT is pre-trained on a 2B-frame retargeted corpus

Cited by 0SourcecodeScholar
2026

OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models

ICLR 2026poster

Spatial reasoning is a key aspect of cognitive psychology and remains a bottleneck for current vision-language models (VLMs). While extensive research has aimed to evaluate or improve VLMs' understanding of basic spatial relations, such as distinguishing left from right, near from far, and object co…

Cited by 0SourcecodeScholar
2026

Reasoning in Space via Grounding in the World

ICLR 2026poster

In this paper, we claim that 3D visual grounding is the cornerstone of spatial reasoning and introduce the $\textit{Grounded-Spatial Reasoner (GS-Reasoner)}$ to explore the effective spatial representations that bridge the gap between them. Existing 3D LLMs suffer from the absence of a unified 3D re…

Cited by 0SourcecodeScholar
2025

DexVLG: Dexterous Vision-Language-Grasp Model at Scale

ICCV 2025poster

As large models gain traction, vision-language models are enabling robots to tackle increasingly complex tasks. However, limited by the difficulty of data collection, progress has mainly focused on controlling simple gripper end-effectors. There is little research on functional grasping with large m…

Cited by 0SourcePDFScholar
2025

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

NeurIPS 2025poster

Recent advances in vision-language-action (VLA) models have shown promise in integrating image generation with action prediction to improve generalization and reasoning in robot manipulation. However, existing methods are limited to challenging image-based forecasting, which suffers from redundant i…

Cited by 0SourcecodeScholar
2025

Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth Estimation

ICCV 2025poster

Current self-supervised monocular depth estimation (MDE) approaches encounter performance limitations due to insufficient semantic-spatial knowledge extraction. To address this challenge, we propose Hybrid-depth, a novel framework that systematically integrates foundation models (e.g., CLIP and DINO…

Cited by 0SourcePDFScholar
2025

SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation

NeurIPS 2025spotlight

While spatial reasoning has made progress in object localization relationships, it often overlooks object orientation—a key factor in 6-DoF fine-grained manipulation. Traditional pose representations rely on pre-defined frames or templates, limiting generalization and semantic grounding. In this pap…

Cited by 0SourceScholar
2024

Closed-Loop Unsupervised Representation Disentanglement with $\\beta$-VAE Distillation and Diffusion Probabilistic Feedback

ECCV 2024poster

"Representation disentanglement may help AI fundamentally understand the real world and thus benefit both discrimination and generation tasks. It currently has at least three unresolved core issues: (i) heavy reliance on label annotation and synthetic data — causing poor generalization on natural sc…

Cited by 7SourcePDFScholar
2024

Hierarchical Temporal Context Learning for Camera-based Semantic Scene Completion

ECCV 2024poster

"Camera-based 3D semantic scene completion (SSC) is pivotal for predicting complicated 3D layouts with limited 2D image observations. The existing mainstream solutions generally leverage temporal information by roughly stacking history frames to supplement the current frame, such straightforward tem…

2024

Scene Graph Disentanglement and Composition for Generalizable Complex Image Generation

NeurIPS 2024spotlight

There has been exciting progress in generating images from natural language or layout conditions. However, these methods struggle to faithfully reproduce complex scenes due to the insufficient modeling of multiple objects and their relationships. To address this issue, we leverage the scene graph, a…

Cited by 2SourcePDFScholar
2022

Local-Global Feature Aggregation for Light Field Image Super-Resolution

ICASSP 2022accepted

Deep convolutional neural networks (CNNs) have been widely explored in light field (LF) image super-resolution (SR) to achieve remarkable progress. However, most of the existing CNNs-based methods ignore the similarity of local neighbor views in the 4D LF data. Besides, due to the limitations of CNN…

Cited by 0SourceScholar
2022

Predict the Rover Mobility Over Soft Terrain Using Articulated Wheeled Bevameter

RA-L 2022

Robot mobility is critical for mission success, especially in soft or deformable terrains, where the complex wheel-soil interaction mechanics often leads to excessive wheel slip and sinkage, causing the eventual mission failure. To improve the rover performance, online mobility prediction using visi

Cited by 21SourceScholar