← Search

Haonan Luo

7 accepted papers

2026

Plan in Sandbox, Navigate in Open Worlds: Learning Physics-Grounded Abstracted Experience for Embodied Navigation

ICML 2026poster

Vision-Language Models (VLMs) have demonstrated exceptional general reasoning capabilities. However, their performance in embodied navigation remains hindered by a scarcity of aligned open-world vision and robot control data. Despite simulators providing a cost-effective alternative for data collect…

Cited by 0SourceScholar
2026

Plug-and-Play Label Map Diffusion for Universal Goal-Oriented Navigation

ICML 2026poster

In embodied vision, Goal-Oriented Navigation (GON) requires robots to locate a specific goal within an unexplored environment. The primary challenge of GON arises from the need to construct a Bird's-Eye-View (BEV) map to understand the environment while simultaneously localizing an unobserved goal. …

Cited by 0SourceScholar
2026

SparseSurf: Sparse-View 3D Gaussian Splatting for Surface Reconstruction

AAAI 2026technical

Recent advances in optimizing Gaussian Splatting for scene geometry have enabled efficient reconstruction of detailed surfaces from images. However, when input views are sparse, such optimization is prone to overfitting, leading to suboptimal reconstruction quality. Existing approaches address this

Cited by 0SourcePDFScholar
2025

A Continual Learning Approach for Embodied Question Answering with Generative Adversarial Imitation Learning

ICASSP 2025accepted

Embodied Question Answering (EQA) is a task in artificial intelligence where an intelligent agent is required to answer questions about its environment. For example, to answer a question such as "Is the TV on or off?", the agent must navigate to the room with the TV and answer with either "On." or "…

Cited by 0SourceScholar
2025

Enhancing Multi-Robot Semantic Navigation Through Multimodal Chain-of-Thought Score Collaboration

AAAI 2025technical

Understanding how humans cooperatively utilize semantic knowledge to explore unfamiliar environments and decide on navigation directions is critical for house service multi-robot systems. Previous methods primarily focused on single-robot centralized planning strategies, which severely limited explo…

2025

KLFormer: Karhunen-Loève Transform for Robust 3D Human Pose Estimation

ICASSP 2025accepted

In the scope of 3D human pose estimation, the task encompasses estimating the 3D positions of key skeletal points (i.e., wrists, elbows, and knees) from a 2D image or video sequence. This technology demonstrates widespread applicability across diverse domains, encompassing domains such as kinematic…

Cited by 0SourceScholar
2019

SegEQA: Video Segmentation Based Visual Attention for Embodied Question Answering

ICCV 2019poster

Embodied Question Answering (EQA) is a newly defined research area where an agent is required to answer the user's questions by exploring the real world environment. It has attracted increasing research interests due to its broad applications in automatic driving system, in-home robots, and personal…

Cited by 34PDFScholar