← Search

Zhichao Wu

6 accepted papers

2026

Speedup Patch: Learning a Plug-and-Play Policy to Accelerate Embodied Manipulation

ICML 2026poster

While current embodied policies exhibit remarkable manipulation skills, their execution remains unsatisfactorily slow as they inherit the tardy pacing of human demonstrations. Existing acceleration methods typically require policy retraining or costly online interactions, limiting their scalability …

Cited by 0SourceScholar
2026

Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action Models

ICML 2026poster

Vision-Language-Action (VLA) models show strong generalization for robotic control, but finetuning them with reinforcement learning (RL) is constrained by the high cost and safety risks of real-world interaction. Training VLA models in interactive world models avoids these issues but introduces seve…

Cited by 0SourceScholar
2025

Apple detection method based on fusion of infrared thermal image and visible-light image

IROS 2025

To address the challenges posed by lighting variations and fruit occlusion in open orchard environments, which significantly affect the performance of apple-harvesting robots, this study proposes an apple detection method based on the fusion of infrared thermal images and visible-light images. First

Cited by 0SourceScholar
2025

FCConDubber: Fine And Coarse Grained Prosody Alignment For Expressive Video Dubbing via Contrastive Audio-Motion Pretraining

ICASSP 2025accepted

Automatic Video Dubbing (AVD) aims to synthesize speech that matches a character’s speaking style and emotion in silent video clips. However, existing approaches rely on attention mechanisms to learn cross-modal prosodic alignment implicitly, making it challenging to capture subtle prosodic variatio…

Cited by 0SourceScholar
2024

Continual Multi-Objective Reinforcement Learning via Reward Model Rehearsal

IJCAI 2024poster

Multi-objective reinforcement learning (MORL) approaches address real-world problems with multiple objectives by learning policies maximizing returns weighted by different user preferences. Typical methods assume the objectives remain unchanged throughout the agent's lifetime. However, in some real-…

Cited by 0SourcePDFScholar
2024

DCTTS: Discrete Diffusion Model with Contrastive Learning for Text-to-Speech Generation

ICASSP 2024accepted

In the Text-to-speech(TTS) task, the latent diffusion model has excellent fidelity and generalization, but its expensive resource consumption and slow inference speed have always been a challenging. To address this issue, this paper proposes the Discrete Diffusion Model with Contrastive Learning for…

Cited by 0SourceScholar