← Search

Yiyuan Pan

8 accepted papers

2026

Emergent Neural Automaton Policies: Learning Symbolic Structure from Visuomotor Trajectories

RSS 2026poster

Scaling robot learning to long-horizon tasks remains a formidable challenge. While end-to-end policies often lack the structural priors needed for effective long-term reasoning, traditional neuro-symbolic methods rely heavily on hand-crafted symbolic priors. To address the issue, we introduce ENAP (…

Cited by 0SourceScholar
2026

Predictive Local Planning with Multi-Step Reward and Q-Value Forecasting

ICRA 2026poster

Planning in dynamic environments often relies on explicit future observation prediction or value-based estimation, both of which can be brittle or hard to generalize in uncertain settings. We propose a novel model-based reinforcement learning framework that performs trajectory rollout and optimizati…

Cited by 0Scholar
2026

Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory

ICLR 2026poster

We introduce M3-Agent, a novel multimodal agent framework equipped with long-term memory. Like humans, M3-Agent can process real-time visual and auditory inputs to build and update episodic and semantic memories, gradually accumulating world knowledge. Its memory is organized in an entity-centric, m…

Cited by 0SourcecodeScholar
2025

FLAME: Learning to Navigate with Multimodal LLM in Urban Environments

AAAI 2025technical

Large Language Models (LLMs) have demonstrated potential in Vision-and-Language Navigation (VLN) tasks, yet current applications face challenges. While LLMs excel in general conversation scenarios, they struggle with specialized navigation tasks, yielding suboptimal performance compared to specializ…

2025

Planning from Imagination: Episodic Simulation and Episodic Memory for Vision-and-Language Navigation

AAAI 2025technical

Humans navigate unfamiliar environments using episodic simulation and episodic memory, which facilitate a deeper understanding of the complex relationships between environments and objects. Developing an imaginative memory system inspired by human mechanisms can enhance the navigation performance of…

Cited by 0SourcePDFScholar
2025

Seeing through Uncertainty: Robust Task-Oriented Optimization in Visual Navigation

NeurIPS 2025poster

Visual navigation is a fundamental problem in embodied AI, yet practical deployments demand long-horizon planning capabilities to address multi-objective tasks. A major bottleneck is data scarcity: policies learned from limited data often overfit and fail to generalize OOD. Existing neural network-b…

Cited by 0SourceScholar
2025

Wonder Wins Ways: Curiosity-Driven Exploration through Multi-Agent Contextual Calibration

NeurIPS 2025poster

Autonomous exploration in complex multi-agent reinforcement learning (MARL) with sparse rewards critically depends on providing agents with effective intrinsic motivation. While artificial curiosity offers a powerful self-supervised signal, it often confuses environmental stochasticity with meaningf…

Cited by 0SourceScholar
2021

CORAL: Colored structural representation for bi-modal place recognition

IROS 2021poster

Place recognition is indispensable for a drift-free localization system. Due to the variations of the environment, place recognition using single-modality has limitations. In this paper, we propose a bi-modal place recognition method, which can extract a compound global descriptor from the two modal…

Cited by 36SourceScholar