← Search

Hankz Hankui Zhuo

8 accepted papers

2025

TCCD: Tree-guided Continuous Causal Discovery via Collaborative MCTS-Parameter Optimization

IJCAI 2025

Learning causal relationships in directed acyclic graphs (DAGs) from multi-type event sequences is a challenging task, especially in large-scale telecommunication networks. Existing methods struggle with the exponentially growing search space and lack global exploration. Gradient-based approaches ar

2025

Transformer-based Reinforcement Learning for Net Ordering in Detailed Routing

IJCAI 2025

With feature size shrinking and design complexity increasing, detailed routing has become a crucial challenge in VLSI design. Although detailed routers have been proposed to judiciously handle hard-to-access pins and various design rules, their performances are sensitive to the order of nets to be r

2023

Gradient-Based Mixed Planning with Symbolic and Numeric Action Parameters (Extended Abstract)

IJCAI 2023poster

Dealing with planning problems with both logical relations and numeric changes in real-world dynamic environments is challenging. Existing numeric planning systems for the problem often discretize numeric variables or impose convex constraints on numeric variables, which harms the performance when s…

Cited by 0SourcePDFScholar
2023

Models as Agents: Optimizing Multi-Step Predictions of Interactive Local Models in Model-Based Multi-Agent Reinforcement Learning

AAAI 2023technical

Research in model-based reinforcement learning has made significant progress in recent years. Compared to single-agent settings, the exponential dimension growth of the joint state-action space in multi-agent systems dramatically increases the complexity of the environment dynamics, which makes it i…

2022

Creativity of AI: Automatic Symbolic Option Discovery for Facilitating Deep Reinforcement Learning

AAAI 2022technical

Despite of achieving great success in real life, Deep Reinforcement Learning (DRL) is still suffering from three critical issues, which are data efficiency, lack of the interpretability and transferability. Recent research shows that embedding symbolic knowledge into DRL is promising in addressing t…

Cited by 55SourcePDFScholar
2022

Plan To Predict: Learning an Uncertainty-Foreseeing Model For Model-Based Reinforcement Learning

NeurIPS 2022accept

In Model-based Reinforcement Learning (MBRL), model learning is critical since an inaccurate model can bias policy learning via generating misleading samples. However, learning an accurate model can be difficult since the policy is continually updated and the induced distribution over visited states…

2021

Coordinated Proximal Policy Optimization

NeurIPS 2021poster

We present Coordinated Proximal Policy Optimization (CoPPO), an algorithm that extends the original Proximal Policy Optimization (PPO) to the multi-agent setting. The key idea lies in the coordinated adaptation of step size during the policy update process among multiple agents. We prove the monoton…

2017

Plan explicability and predictability for robot task planning

ICRA 2017poster

Intelligent robots and machines are becoming pervasive in human populated environments. A desirable capability of these agents is to respond to goal-oriented commands by autonomously constructing task plans. However, such autonomy can add significant cognitive load and potentially introduce safety r…

Cited by 227SourceScholar
Hankz Hankui Zhuo — accepted AI-conference papers · AIConfPaper