← Search

Xu Luo

14 accepted papers

2026

InSpire: Vision-Language-Action Models with Intrinsic Spatial Reasoning

ICRA 2026poster

Leveraging pretrained Vision-Language Models (VLMs) to map language instruction and visual observations to raw low-level actions, Vision-Language-Action models (VLAs) hold great promise for achieving general-purpose robotic systems. Despite their advancements, existing VLAs tend to spuriously correl…

2026

MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents

ICLR 2026poster

The Model Context Protocol (MCP) standardizes how large language model (LLM) agents discover, describe, and call external tools. While MCP unlocks broad interoperability, it also enlarges the attack surface by making tools first-class, composable objects with natural-language metadata, and standardi…

Cited by 0SourcecodeScholar
2026

Policy Contrastive Decoding for Robotic Foundation Models

ICLR 2026poster

Generalist robot policies, or robotic foundation models, hold immense potential to enable flexible, general-purpose and dexterous robotic systems. Despite their advancements, our empirical experiments reveal that existing robot policies are prone to learning spurious correlations from pre-training t…

Cited by 0SourcecodeScholar
2025

Lumina-T2X: Scalable Flow-based Large Diffusion Transformer for Flexible Resolution Generation

ICLR 2025spotlight

Sora unveils the potential of scaling Diffusion Transformer (DiT) for generating photorealistic images and videos at arbitrary resolutions, aspect ratios, and durations, yet it still lacks sufficient implementation details. In this paper, we introduce the Lumina-T2X family -- a series of Flow-based…

2025

Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation

CoRL 2025poster

Generalist robot policies trained on large-scale datasets such as Open X-Embodiment (OXE) demonstrate strong performance across a wide range of tasks. However, they often struggle to generalize beyond the distribution of their training data. In this paper, we investigate the underlying cause of this…

Cited by 0SourceScholar
2025

Unlocking Smarter Device Control: Foresighted Planning with a World Model-Driven Code Execution Approach

EMNLP 2025

The automatic control of mobile devices is essential for efficiently performing complex tasks that involve multiple sequential steps. However, these tasks pose significant challenges due to the limited environmental information available at each step, primarily through visual observations. As a resu

2024

CoIN: A Benchmark of Continual Instruction Tuning for Multimodel Large Language Models

NeurIPS 2024poster

Instruction tuning demonstrates impressive performance in adapting Multimodal Large Language Models (MLLMs) to follow task instructions and improve generalization ability. By extending tuning across diverse tasks, MLLMs can further enhance their understanding of world knowledge and instruction inte…

Cited by 14SourcePDFScholar
2024

Lumina-Next : Making Lumina-T2X Stronger and Faster with Next-DiT

NeurIPS 2024poster

Lumina-T2X is a nascent family of Flow-based Large Diffusion Transformers (Flag-DiT) that establishes a unified framework for transforming noise into various modalities, such as images and videos, conditioned on text instructions. Despite its promising capabilities, Lumina-T2X still encounters chall…

2024

Unleashing the Potentials of Likelihood Composition for Multi-modal Language Models

EMNLP 2024finding

Model fusing has always been an important topic, especially in an era where large language models (LLM) and multi-modal language models (MLM) with different architectures, parameter sizes and training pipelines, are being created all the time. In this work, we propose a post-hoc framework, aiming at…

2023

A Closer Look at Few-shot Classification Again

ICML 2023poster

Few-shot classification consists of a training phase where a model is learned on a relatively large dataset and an adaptation phase where the learned model is adapted to previously-unseen tasks with limited labeled samples. In this paper, we empirically prove that the training algorithm and the adap…

2022

Alleviating the Sample Selection Bias in Few-shot Learning by Removing Projection to the Centroid

NeurIPS 2022accept

Few-shot learning (FSL) targets at generalization of vision models towards unseen tasks without sufficient annotations. Despite the emergence of a number of few-shot learning methods, the sample selection bias problem, i.e., the sensitivity to the limited amount of support data, has not been well un…

2021

Rectifying the Shortcut Learning of Background for Few-Shot Learning

NeurIPS 2021poster

The category gap between training and evaluation has been characterised as one of the main obstacles to the success of Few-Shot Learning (FSL). In this paper, we for the first time empirically identify image background, common in realistic images, as a shortcut knowledge helpful for in-class classif…