← Search

Xiaokang Ma

3 accepted papers

2026

EMKG: Embodied Memory Knowledge Graphs for Object-Goal Navigation in Dynamic Open Worlds

RA-L 2026

Object-Goal Navigation (OGN) in complex domestic environments remains challenging due to spatial memory and semantic uncertainties. To address this, we introduce EMKG, an embodied multimodal memory knowledge graph framework that enables open-world navigation. In contrast to conventional vision-langu

Cited by 0SourceScholar
2026

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective

ICLR 2026poster

Existing Multimodal Large Language Models (MLLMs) process a large number of visual tokens, leading to significant computational costs and inefficiency. Instruction-related visual token compression demonstrates strong task relevance, which aligns well with MLLMs’ ultimate goal of instruction followin…

Cited by 0SourceScholar
2025

Stimulating Imagination: Towards General-purpose "Something Something Placement"

IROS 2025

General-purpose object placement is a fundamental capability of an intelligent generalist robot: being capable of rearranging objects following precise human instructions even in novel environments. This work is dedicated to achieving general-purpose object placement with "something something" instr

Cited by 0SourceScholar