← Search

Xin Xiao

5 accepted papers

2026

RoboOmni: Actions Are Just Another Modality for Your Vision-Language Models

ICML 2026poster

Integrating Vision-Language Models (VLMs) into robotics has facilitated the development of generalizable Vision-Language Action (VLA) policies. However, unified discrete frameworks lag behind decoupled continuous designs due to limitations in action chunking and temporal modeling. To address this, w…

Cited by 0SourceScholar
2025

EEG Decoding and Visual Reconstruction via 3D Geometric with Nonstationarity Modelling

ICASSP 2025accepted

Electroencephalogram (EEG) signal processing has advanced in revealing the mechanisms of human visual perception, but existing methods often overlook two key EEG properties: (1) 3D geometric relationships between EEG electrodes, which reflects the ability to model the brain in stereoscopic terms; an…

Cited by 0SourceScholar
2024

Seeing the Image: Prioritizing Visual Correlation by Contrastive Alignment

NeurIPS 2024poster

Existing image-text modality alignment in Vision Language Models (VLMs) treats each text token equally in an autoregressive manner. Despite being simple and effective, this method results in sub-optimal cross-modal alignment by over-emphasizing the text tokens that are less correlated with or even c…

2024

Shifted Autoencoders for Point Annotation Restoration in Object Counting

ECCV 2024poster

"Object counting typically uses 2D point annotations. The complexity of object shapes and the subjectivity of annotators may lead to annotation inconsistency, potentially confusing counting model training. Some sophisticated noise-resistance counting methods have been proposed to alleviate this issu…

2024

World to Code: Multi-modal Data Generation via Self-Instructed Compositional Captioning and Filtering

EMNLP 2024main

Recent advances in Vision-Language Models (VLMs) and the scarcity of high-quality multi-modal alignment data have inspired numerous researches on synthetic VLM data generation. The conventional norm in VLM data construction uses a mixture of specialists in caption and OCR, or stronger VLM APIs and e…