← Search

Xiaojie Zhang

7 accepted papers

2026

MITA: A HIERARCHICAL MULTI-AGENT COLLABORATION FRAMEWORK WITH MEMORY-INTEGRATED AND TASK ALLOCATION

ICASSP 2026poster

Recent advances in large language models (LLMs) have substantially accelerated the development of embodied agents. LLM-based multi-agent systems mitigate the inefficiency of single agents in complex tasks. However, they still suffer from issues such as memory inconsistency and agent behavioral confl…

Cited by 0SourcePDFScholar
2026

Predicting What Matters: Robust Generalist Robot Policy Learning via Future Semantic Mask

ICML 2026poster

World models derived from large-scale video generative pre-training have emerged as a promising paradigm for generalist robot policy learning. However, standard approaches often focus on high-fidelity RGB video prediction, but this can result in overfitting to irrelevant factors, such as dynamic bac…

Cited by 0SourceScholar
2026

UniDoorManip: Learning Universal Door Manipulation Policy Over Large-Scale and Diverse Door Manipulation Environments

ICRA 2026poster

Learning a universal manipulation policy encompassing doors with diverse categories, geometries and mechanisms, is crucial for future embodied agents to effectively work in complex and broad real-world scenarios. Due to the limited datasets and unrealistic simulation environments, previous studies f…

2025

AdaManip: Adaptive Articulated Object Manipulation Environments and Policy Learning

ICLR 2025poster

Articulated object manipulation is a critical capability for robots to perform various tasks in real-world scenarios. Composed of multiple parts connected by joints, articulated objects are endowed with diverse functional mechanisms through complex relative motions. For example, a safe consists of a…

Cited by 4SourcePDFScholar
2025

Adaptive Articulated Object Manipulation On The Fly with Foundation Model Reasoning and Part Grounding

ICCV 2025poster

Articulated objects pose diverse manipulation challenges for robots. Since their internal structures are not directly observable, robots must adaptively explore and refine actions to generate successful manipulation trajectories. While existing works have attempted cross-category generalization in a…

Cited by 0SourcePDFScholar
2025

Unified Category-Level Object Detection and Pose Estimation from RGB Images using 3D Prototypes

ICCV 2025poster

Recognizing objects in images is a fundamental problem in computer vision. Although detecting objects in 2D images is common, many applications require determining their pose in 3D space. Traditional category-level methods rely on RGB-D inputs, which may not always be available, or employ two-stage…

2022

A Joint Learning Framework for Restaurant Survival Prediction and Explanation

EMNLP 2022main

The bloom of the Internet and the recent breakthroughs in deep learning techniques open a new door to AI for E-commence, with a trend of evolving from using a few financial factors such as liquidity and profitability to using more advanced AI techniques to process complex and multi-modal data. In th…