← Search

Xi Guo

4 accepted papers

2026

AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMs

AAAI 2026technical

Visual abductive reasoning (VAR) is a challenging task that requires AI systems to infer the most likely explanation for incomplete visual observations. While recent MLLMs develop strong general-purpose multimodal reasoning capabilities, they remain fall short in abductive inference, as compared to

Cited by 0SourcePDFScholar
2025

DriveScape: High-Resolution Driving Video Generation by Multi-View Feature Fusion

CVPR 2025poster

Recent advancements in generative models offer promising solutions for synthesizing realistic driving videos, aiding in training autonomous driving perception models. However, existing methods often struggle with high-resolution multi-view generation, mainly due to the significant memory and computa…

Cited by 0SourcePDFScholar
2025

InstaDrive: Instance-Aware Driving World Models for Realistic and Consistent Video Generation

ICCV 2025poster

Autonomous driving relies on robust models trained on high-quality, large-scale multi-view driving videos for tasks like perception and planning. While world models offer a cost-effective solution for generating realistic driving videos, they struggle to maintain instance-level temporal consistency…

Cited by 0SourcePDFScholar
2022

Learning Video Representations of Human Motion From Synthetic Data

CVPR 2022poster

In this paper, we take an early step towards video representation learning of human actions with the help of largescale synthetic videos, particularly for human motion representation enhancement. Specifically, we first introduce an automatic action-related video synthesis pipeline based on a photore…

Cited by 17PDFScholar