← Search

Xiang Zhu

8 accepted papers

2026

Learning Generalizable Robot Policy with Human Demonstration Video As a Prompt

ICRA 2026poster

Recent robot learning methods commonly rely on imitation learning from massive robotic dataset collected with teleoperation. When facing a new task, such methods generally require collecting a set of new teleoperation data and finetuning the policy. Furthermore, the teleoperation data collection pip…

2025

A Multi-Sensor Fusion Approach for Rapid Orthoimage Generation in Large-Scale UAV Mapping

IROS 2025

Rapid generation of large-scale orthoimages from Unmanned Aerial Vehicles (UAVs) has been a long-standing focus of research in the field of aerial mapping. A multi-sensor UAV system, integrating the Global Positioning System (GPS), Inertial Measurement Unit (IMU), 4D millimeter-wave radar and camera

Cited by 3SourceScholar
2025

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

ICML 2025poster

Recent advancements in Vision-Language-Action (VLA) models have leveraged pre-trained Vision-Language Models (VLMs) to improve the generalization capabilities. VLMs, typically pre-trained on vision-language understanding tasks, provide rich semantic knowledge and reasoning abilities. However, prior…

Cited by 2SourcePDFScholar
2025

Where Does This Data Come From? Enhanced Source Inference Attacks in Federated Learning

IJCAI 2025

Federated learning (FL) enables collaborative model training without exposing raw data, offering a privacy-aware alternative to centralized learning. However, FL remains vulnerable to various privacy attacks that exploit shared model updates, including membership inference, property inference, and g

Cited by 0SourcePDFScholar
2024

Advancing Humanoid Locomotion: Mastering Challenging Terrains with Denoising World Model Learning

RSS 2024poster

Humanoid robots, with their human-like skeletal structure, are especially suited for tasks in human-centric environments. However, this structure is accompanied by additional challenges in locomotion controller design, especially in complex real-world environments. As a result, existing humanoid rob…

2024

Feature Re-Embedding: Towards Foundation Model-Level Performance in Computational Pathology

CVPR 2024poster

Multiple instance learning (MIL) is the most widely used framework in computational pathology encompassing sub-typing diagnosis prognosis and more. However the existing MIL paradigm typically requires an offline instance feature extractor such as a pre-trained ResNet or a foundation model. This appr…

2022

A Contact-Safe Reinforcement Learning Framework for Contact-Rich Robot Manipulation

IROS 2022poster

Reinforcement learning shows great potential to solve complex contact-rich robot manipulation tasks. However, the safety of using RL in the real world is a crucial problem, since unexpected dangerous collisions might happen when the RL policy is imperfect during training or in unseen scenarios. In t…

Cited by 7SourceScholar