← Search

Zhiyuan Zhou

16 accepted papers

2026

CausalVAD: De-confounding End-to-End Autonomous Driving via Causal Intervention

CVPR 2026

Planning-oriented end-to-end driving models show great promise, yet they fundamentally learn statistical correlations instead of true causal relationships. This vulnerability leads to causal confusion, where models exploit dataset biases as shortcuts, critically harming their reliability and safety

Cited by 0SourceScholar
2026

Robust Fine-tuning of Vision-Language-Action Robot Policies via Parameter Merging

ICLR 2026poster

Generalist robot policies, trained on large and diverse datasets, have demonstrated the ability to generalize across a wide spectrum of behaviors, enabling a single policy to act in varied real-world environments. However, they still fall short on new tasks not covered in the training data. When fin…

Cited by 0SourceScholar
2026

Test-Time Adaptation without Source Data for Out-of-Domain Bioactivity Prediction

ICLR 2026poster

Accurate prediction of protein-ligand bioactivity is a cornerstone of modern drug discovery, yet current deep learning methods often struggle with out-of-domain (OOD) generalization. The existing methods rely on access to source data, making them impractical in scenarios where data cannot be accesse…

Cited by 0SourceScholar
2026

Topology-Optimized, Dual-Phase Gripper With Force Estimation for Underwater Operation

RA-L 2026

Addressing the critical trade-off between compliance and grasping force in underwater manipulation, this letter presents a novel topology optimization framework to automate the design of soft, variable-stiffness fingers from a single material. By employing multiple load cases within the optimization

Cited by 0SourceScholar
2026

Voices, Faces, and Feelings: Multi-modal Emotion-Cognition Captioning for Mental Health Understanding

AAAI 2026technical

Emotional and cognitive factors are essential for understanding mental health disorders. However, existing methods often treat multi-modal data as classification tasks, limiting interpretability especially for emotion and cognition. Although large language models (LLMs) offer opportunities for ment

Cited by 0SourcePDFScholar
2026

π∗0.6π0.6∗\pi^{*}_{0.6}: a VLA That Learns From Experience

RSS 2026poster

Vision–language–action (VLA) models offer a promising path toward general-purpose robots, but achieving the reliability and speed required for practical deployment remains challenging. We present a general-purpose method, RL with Experience and Corrections via Advantage-conditioned Policies (RECAP) …

Cited by 0SourceScholar
2025

AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World

CoRL 2025poster

Scalable and reproducible policy evaluation has been a long-standing challenge in robot learning: evaluations are critical to assess progress and build better policies, but evaluation in the real world, especially at a scale that would provide statistically reliable results, is costly in terms of hu…

Cited by 0SourcecodeScholar
2025

Behavioral Exploration: Learning to Explore via In-Context Adaptation

ICML 2025poster

Developing autonomous agents that quickly explore an environment and adapt their behavior online is a canonical challenge in robotics and machine learning. While humans are able to achieve such fast online exploration and adaptation, often acquiring new information and skills in only a handful of in…

Cited by 0SourcePDFScholar
2025

Compute-Optimal Scaling for Value-Based Deep RL

NeurIPS 2025poster

As models grow larger and training them becomes expensive, it becomes increasingly important to scale training recipes not just to larger models and more data, but to do so in a compute-optimal manner that extracts maximal performance per unit of compute. While such scaling has been well studied for…

Cited by 0SourcecodeScholar
2025

Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

ICLR 2025poster

The modern paradigm in machine learning involves pre-training on diverse data, followed by task-specific fine-tuning. In reinforcement learning (RL), this translates to learning via offline RL on a diverse historical dataset, followed by rapid online RL fine-tuning using interaction data. Most RL fi…

2025

Enhancing Bioactivity Prediction via Spatial Emptiness Representation of Protein-ligand Complex and Union of Multiple Pockets

NeurIPS 2025poster

Predicting the bioactivity of candidate ligands remains a central challenge in drug discovery. Ligands and endogenous substrates often compete for the same binding sites on target proteins, and the extent to which a ligand can modulate protein function depends not only on its binding but also on how…

Cited by 0SourceScholar
2025

LHQ-SVC: Lightweight and High Quality Singing Voice Conversion Modeling

ICASSP 2025accepted

Singing Voice Conversion (SVC) has emerged as a significant subfield of Voice Conversion (VC), enabling the transformation of one singer’s voice into another while preserving musical elements such as melody, rhythm, and timbre. Traditional SVC methods have limitations in terms of audio quality, data…

Cited by 0SourceScholar
2025

Learning Spatial-Aware Manipulation Ordering

NeurIPS 2025poster

Manipulation in cluttered environments is challenging due to spatial dependencies among objects, where an improper manipulation order can cause collisions or blocked access. Existing approaches often overlook these spatial relationships, limiting their flexibility and scalability. To address these l…

Cited by 0SourceScholar
2025

On Path to Multimodal Generalist: General-Level and General-Bench

ICML 2025oral

The Multimodal Large Language Model (MLLM) is currently experiencing rapid growth, driven by the advanced capabilities of language-based LLMs. Unlike their specialist predecessors, existing MLLMs are evolving towards a Multimodal Generalist paradigm. Initially limited to understanding multiple mod…

Cited by 0SourcePDFScholar
2024

Autonomous Improvement of Instruction Following Skills via Foundation Models

CoRL 2024poster

Intelligent robots capable of improving from autonomously collected experience have the potential to transform robot learning: instead of collecting costly teleoperated demonstration data, large-scale deployment of fleets of robots can quickly collect larger quantities of autonomous data useful for…

Cited by 12SourcecodeScholar