Lightweight Adaptive Topological Layout and Semantic Mapping in Vision-and-Language Navigation on Websites
Vision-and-Language navigation on websites requires agents to navigate target webpages and answer questions based on human instructions. Current web agents primarily leverage Large Language Models (LLMs) for semantic understanding and reasoning, but still suffer from limited navigation performance a