← Search

Bohan Zhou

10 accepted papers

2026

Learning Diverse Bimanual Dexterous Manipulation Skills from Human Demonstrations

AAAI 2026technical

Bimanual dexterous manipulation is a critical yet underexplored area in robotics. Its high-dimensional action space and inherent task complexity present significant challenges for policy learning, and the limited task diversity in existing benchmarks hinders general-purpose skill development. Existi

Cited by 0SourcePDFScholar
2026

Sobolev Gradient Ascent for Optimal Transport: Barycenter Optimization and Convergence Analysis

ICLR 2026poster

This paper introduces a new constraint-free concave dual formulation for the Wasserstein barycenter. Tailoring the vanilla dual gradient ascent algorithm to the Sobolev geometry, we derive a scalable Sobolev gradient ascent (SGA) algorithm to compute the barycenter for input distributions supported…

Cited by 0SourceScholar
2025

Cradle: Empowering Foundation Agents towards General Computer Control

ICML 2025poster

Despite their success in specific scenarios, existing foundation agents still struggle to generalize across various virtual scenarios, mainly due to the dramatically different encapsulations of environments with manually designed observation and action spaces. To handle this issue, we propose the Ge…

2025

MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion Generation

NeurIPS 2025poster

Egocentric hand-object motion generation is crucial for immersive AR/VR and robotic imitation but remains challenging due to unstable viewpoints, self-occlusions, perspective distortion, and noisy ego-motion. Existing methods rely on predefined 3D object priors, limiting generalization to novel obje…

Cited by 0SourcecodeScholar
2024

UniCode : Learning a Unified Codebook for Multimodal Large Language Models

ECCV 2024poster

"In this paper, we propose UniCode, a novel approach within the domain of multimodal large language models (MLLMs) that learns a unified codebook to efficiently tokenize visual, text, and potentially other types of signals. This innovation addresses a critical limitation in existing MLLMs: their rel…

Cited by 12SourcePDFScholar
2023

GFIE: A Dataset and Baseline for Gaze-Following From 2D to 3D in Indoor Environments

CVPR 2023poster

Gaze-following is a kind of research that requires locating where the person in the scene is looking automatically under the topic of gaze estimation. It is an important clue for understanding human intention, such as identifying objects or regions of interest to humans. However, a survey of dataset…

Cited by 10SourcePDFScholar
2023

Learning from Visual Observation via Offline Pretrained State-to-Go Transformer

NeurIPS 2023poster

Learning from visual observation (LfVO), aiming at recovering policies from only visual observation data, is promising yet a challenging problem. Existing LfVO approaches either only adopt inefficient online learning schemes or require additional task-specific information like goal states, making th…

Cited by 12SourcePDFScholar