← Search

Quanting Xie

3 accepted papers

2026

Affordance RAG: Hierarchical Multimodal Retrieval With Affordance-Aware Embodied Memory for Mobile Manipulation

RA-L 2026

In this study, we address the problem of openvocabulary mobile manipulation, where a robot is required to carry a wide range of objects to receptacles based on freeform natural language instructions. This task is challenging, as it involves understanding visual semantics and the affordance of manipu

Cited by 2SourceScholar
2026

Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation

ICRA 2026poster

In this study, we address the problem of open-vocabulary mobile manipulation, where a robot is required to carry a wide range of objects to receptacles based on free-form natural language instructions. This task is challenging, as it involves understanding visual semantics and the affordance of mani…

2026

MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning

CVPR 2026

Systems such as video chatbots and navigation robots often depend on streaming image captioning to interpret visual inputs. Existing approaches typically employ large multimodal language models (MLLMs) for this purpose, but their substantial computational cost hinders practical application.This limi

Cited by 0SourcecodeScholar