← Search

Xiyao Ma

5 accepted papers

2026

Disentangling for Transfer: Boosting Limited Modalities via Information-Theoretic Regularization and Cross-Modal Reconstruction

AAAI 2026technical

Missing critical modalities in medical imaging poses significant challenges for AI-driven diagnostic systems, particularly in scenarios where limited modalities must suffice for downstream tasks. Existing approaches often fail to fully leverage privileged features available only at training or addre

Cited by 0SourcePDFScholar
2025

MMPlanner: Zero-Shot Multimodal Procedural Planning with Chain-of-Thought Object State Reasoning

EMNLP 2025

Multimodal Procedural Planning (MPP) aims to generate step-by-step instructions that combine text and images, with the central challenge of preserving object-state consistency across modalities while producing informative plans. Existing approaches often leverage large language models (LLMs) to refi

Cited by 0SourcePDFScholar
2024

MEND: Meta Demonstration Distillation for Efficient and Effective In-Context Learning

ICLR 2024poster

Large Language models (LLMs) have demonstrated impressive in-context learning (ICL) capabilities, where a LLM makes predictions for a given test input together with a few input-output pairs (demonstrations). Nevertheless, the inclusion of demonstrations poses a challenge, leading to a quadratic inc…

2022

Contrastive Knowledge Graph Attention Network for Request-Based Recipe Recommendation

ICASSP 2022accepted

To improve daily customer experience, kitchen assistant becomes one of the enabled service in intelligent voice assistants, presenting personalized and relevant recipes to satisfy customer requests. Current solutions for recipe recommendation suffers from two limitations: First, user-recipe interact…

Cited by 0SourceScholar
2019

Adaptive Leader-Follower Formation Control and Obstacle Avoidance via Deep Reinforcement Learning

IROS 2019poster

We propose a deep reinforcement learning (DRL) methodology for the tracking, obstacle avoidance, and formation control of nonholonomic robots. By separating vision-based control into a perception module and a controller module, we can train a DRL agent without sophisticated physics or 3D modeling. I…

Cited by 29SourceScholar