← Search

Shubo Zhang

6 accepted papers

2026

EMKG: Embodied Memory Knowledge Graphs for Object-Goal Navigation in Dynamic Open Worlds

RA-L 2026

Object-Goal Navigation (OGN) in complex domestic environments remains challenging due to spatial memory and semantic uncertainties. To address this, we introduce EMKG, an embodied multimodal memory knowledge graph framework that enables open-world navigation. In contrast to conventional vision-langu

Cited by 0SourceScholar
2026

MemClaw-RAG: Memory-Driven Navigation and Adaptive Locomotion for Wheeled-Legged Robots in Dynamic Environments

ICRA 2026poster

Object-Goal Navigation in dynamic environments remains challenging because many existing approaches rely primarily on reactive mapping and lack the ability to retain historical experience or establish structured memory associations. To address this limitation, we introduce MemClaw-RAG, an embodied m…

Cited by 0Scholar
2025

MemeReaCon: Probing Contextual Meme Understanding in Large Vision-Language Models

EMNLP 2025

Memes have emerged as a popular form of multimodal online communication, where their interpretation heavily depends on the specific context in which they appear. Current approaches predominantly focus on isolated meme analysis, either for harmful content detection or standalone interpretation, overl

Cited by 0SourcePDFScholar
2025

T2: An Adaptive Test-Time Scaling Strategy for Contextual Question Answering

EMNLP 2025

Recent advances in large language models have demonstrated remarkable performance on Contextual Question Answering (CQA). However, prior approaches typically employ elaborate reasoning strategies regardless of question complexity, leading to low adaptability. Recent efficient test-time scaling metho

2024

MUCH: A Multimodal Corpus Construction for Conversational Humor Recognition Based on Chinese Sitcom

COLING 2024main

Conversational humor is the key to capturing dialogue semantics and dialogue comprehension, which is usually generated in multiple modalities, such as linguistic rhetoric (textual modality), exaggerated facial expressions or movements (visual modality), and quirky intonation (acoustic modality). How…

Cited by 0SourcePDFScholar