← Search

Lipeng Wan

9 accepted papers

2026

Beyond Policy Training: Recursive Solution Search from Unannotated Videos

ICML 2026poster

Many real-world tasks are recorded as large collections of unannotated task executions, such as videos, which contain rich information about task progress but lack the supervision assumed by standard reinforcement learning (RL) pipelines. In many practical settings, the goal is not to train a reusab…

Cited by 0SourceScholar
2025

State Revisit and Re-explore: Bridging Sim-to-Real Gaps in Offline-and-Online Reinforcement Learning with An Imperfect Simulator

IJCAI 2025

In reinforcement learning (RL) based robot skill acquisition, a high-fidelity simulator is usually indispensable but unattainable since the real environment dynamics are difficult to model, which leads to severe sim-to-real gaps. Existing methods solve this problem by combining offline and online RL

Cited by 0SourcePDFScholar
2024

Grounded Answers for Multi-agent Decision-making Problem through Generative World Model

NeurIPS 2024poster

Recent progress in generative models has stimulated significant innovations in many fields, such as image generation and chatbots. Despite their success, these models often produce sketchy and misleading solutions for complex multi-agent decision-making problems because they miss the trial-and-error…

Cited by 0SourcePDFScholar
2024

Imagine, Initialize, and Explore: An Effective Exploration Method in Multi-Agent Reinforcement Learning

AAAI 2024technical

Effective exploration is crucial to discovering optimal strategies for multi-agent reinforcement learning (MARL) in complex coordination tasks. Existing methods mainly utilize intrinsic rewards to enable committed exploration or use role-based learning for decomposing joint action spaces instead of…

Cited by 3SourcePDFScholar
2023

Deep Hierarchical Communication Graph in Multi-Agent Reinforcement Learning

IJCAI 2023poster

Sharing intentions is crucial for efficient cooperation in communication-enabled multi-agent reinforcement learning. Recent work applies static or undirected graphs to determine the order of interaction. However, the static graph is not general for complex cooperative tasks, and the parallel message…

Cited by 7SourcePDFScholar
2023

MMRDN: Consistent Representation for Multi-View Manipulation Relationship Detection in Object-Stacked Scenes

ICRA 2023poster

Manipulation relationship detection (MRD) aims to guide the robot to grasp objects in the right order, which is important to ensure the safety and reliability of grasping in object stacked scenes. Previous works infer manipulation relationship by deep neural network trained with data collected from…

Cited by 2SourceScholar
2022

Greedy based Value Representation for Optimal Coordination in Multi-agent Reinforcement Learning

ICML 2022spotlight

Due to the representation limitation of the joint Q value function, multi-agent reinforcement learning methods with linear value decomposition (LVD) or monotonic value decomposition (MVD) suffer from relative overgeneralization. As a result, they can not ensure optimal consistency (i.e., the corresp…

Cited by 15SourcePDFScholar
2019

A Multi-task Convolutional Neural Network for Autonomous Robotic Grasping in Object Stacking Scenes

IROS 2019poster

Autonomous robotic grasping plays an important role in intelligent robotics. However, how to help the robot grasp specific objects in object stacking scenes is still an open problem, because there are two main challenges for autonomous robots: (1) it is a comprehensive task to know what and how to g…

Cited by 87SourceScholar