← Search

Chengxuan Qian

6 accepted papers

2026

AutoDrive-R²: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving

ICLR 2026poster

Vision–Language–Action (VLA) models in autonomous driving systems have recently demonstrated transformative potential by integrating multimodal perception with decision-making capabilities. However, the interpretability and coherence of the decision process and the plausibility of action sequences r…

Cited by 0SourceScholar
2026

DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning

ICLR 2026poster

Multimodal representation learning aims to capture both shared and complementary semantic information across multiple modalities. However, the intrinsic heterogeneity of diverse modalities presents substantial challenges to achieve effective cross-modal collaboration and integration. To address this…

Cited by 0SourcecodeScholar
2026

Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools

ICLR 2026poster

Multimodal large language models (MLLMs) have demonstrated remarkable potential in bridging visual and textual reasoning, yet their reliance on text-centric priors often limits their ability to disentangle semantically similar actions in open-vocabulary scenarios. To address this, we propose Video-S…

Cited by 0SourceScholar
2026

fMRI-LM: Towards a Universal Foundation Model for Language-Aligned fMRI Understanding

CVPR 2026

Recent advances in multimodal large language models (LLMs) have enabled unified reasoning across images, audio, and video, but extending such capability to brain imaging remains largely unexplored. Bridging this gap is essential to link neural activity with semantic cognition and to develop cross-mo

Cited by 0SourcecodeScholar
2025

HAIF-GS: Hierarchical and Induced Flow-Guided Gaussian Splatting for Dynamic Scene

NeurIPS 2025poster

Reconstructing dynamic 3D scenes from monocular videos remains a fundamental challenge in 3D vision. While 3D Gaussian Splatting (3DGS) achieves real-time rendering in static settings, extending it to dynamic scenes is challenging due to the difficulty of learning structured and temporally consisten…

Cited by 0SourceScholar
2025

Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization

EMNLP 2025

The emergence of large Vision Language Models (VLMs) has broadened the scope and capabilities of single-modal Large Language Models (LLMs) by integrating visual modalities, thereby unlocking transformative cross-modal applications in a variety of real-world scenarios. Despite their impressive perfor