← Search

Yanwei Wang

7 accepted papers

2025

A Hierarchical Multi Robot Coverage Strategy for Large Maps With Reinforcement Learning and Dense Segmented Siamese Network

RA-L 2025

Complete coverage of multiple robots for a large map is an important collaborative planning task, which is widely used in disaster search and rescue, forest fire prevention, resource exploration, and other fields. It generally focuses on coverage completion with less robot (mostly drone) occupation

Cited by 3SourceScholar
2025

Inference-Time Policy Steering Through Human Interactions

ICRA 2025

Generative policies trained with human demonstrations can autonomously accomplish multimodal, longhorizon tasks. However, during inference, humans are often removed from the policy execution loop, limiting the ability to guide a pre-trained policy towards a specific sub-goal or trajectory shape amon

Cited by 37SourcecodeScholar
2025

Partner Modelling Emerges in Recurrent Agents (But Only When It Matters)

NeurIPS 2025poster

Humans are remarkably adept at collaboration, able to infer the strengths and weaknesses of new partners in order to work successfully towards shared goals. To build AI systems with this capability, we must first understand its building blocks: does such flexibility require explicit, dedicated mecha…

Cited by 0SourceScholar
2025

Versatile Demonstration Interface: Toward More Flexible Robot Demonstration Collection

IROS 2025

Previous methods for Learning from Demonstration leverage several approaches for a human to teach motions to a robot, including teleoperation, kinesthetic teaching, and natural demonstrations. However, little previous work has explored more general interfaces that allow for multiple demonstration ty

Cited by 3SourceScholar
2024

Grounding Language Plans in Demonstrations Through Counterfactual Perturbations

ICLR 2024spotlight

Grounding the common-sense reasoning of Large Language Models in physical domains remains a pivotal yet unsolved problem for embodied AI. Whereas prior works have focused on leveraging LLMs directly for planning in symbolic spaces, this work uses LLMs to guide the search of task structures and const…

2022

Temporal Logic Imitation: Learning Plan-Satisficing Motion Policies from Demonstrations

CoRL 2022oral

Learning from demonstration (LfD) has successfully solved tasks featuring a long time horizon. However, when the problem complexity also includes human-in-the-loop perturbations, state-of-the-art approaches do not guarantee the successful reproduction of a task. In this work, we identify the roots o…

Cited by 26SourceScholar