ActPlan-1K: Benchmarking the Procedural Planning Ability of Visual Language Models in Household Activities
Large language models(LLMs) have been adopted to process textual task description and accomplish procedural planning in embodied AI tasks because of their powerful reasoning ability. However, there is still lack of study on how vision language models(VLMs) behave when multi-modal task inputs are con…