← Search

Xiaohao Cai

5 accepted papers

2026

HiFi-Mamba: Dual-Stream ?-Laplacian Enhanced Mamba for High-Fidelity MRI Reconstruction

AAAI 2026technical

Reconstructing high-fidelity MR images from undersampled k-space data remains a challenging problem in MRI. While Mamba variants for vision tasks offer promising long-range modeling capabilities with linear-time complexity, their direct application to MRI reconstruction inherits two key limitations:

Cited by 0SourcePDFScholar
2026

MOGO: Residual Quantized Hierarchical Causal Transformer for Real-Time and Infinite-Length 3D Human Motion Generation

AAAI 2026technical

Recent advances in transformer-based text-to-motion generation have significantly improved motion quality. However, achieving both real-time performance and long-horizon scalability remains an open challenge. In this paper, we present MOGO (Motion Generation with One-pass), a novel autoregressive fr

Cited by 0SourcePDFScholar
2026

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently

ICML 2026poster

Transformers can acquire Chain-of-Thought (CoT) capabilities to solve complex reasoning tasks through fine-tuning. Reinforcement learning (RL) and supervised fine-tuning (SFT) are two primary approaches to this end. In this work, we examine them specifically for learning k-sparse Boolean functions w…

Cited by 0SourceScholar
2025

CALM: Culturally Self-Aware Language Models

NeurIPS 2025poster

Cultural awareness in language models is the capacity to understand and adapt to diverse cultural contexts. However, most existing approaches treat culture as static background knowledge, overlooking its dynamic and evolving nature. This limitation reduces their reliability in downstream tasks that…

Cited by 0SourceScholar
2025

Talk2Radar: Bridging Natural Language with 4D mmWave Radar for 3D Referring Expression Comprehension

ICRA 2025

Embodied perception is essential for intelligent vehicles and robots in interactive environmental understanding. However, these advancements primarily focus on vision, with limited attention given to using 3D modeling sensors, restricting a comprehensive understanding of objects in response to promp

Cited by 19SourcecodeScholar