← Search

Jianhao Yuan

8 accepted papers

2026

Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference

ICLR 2026poster

Intuitive physics understanding in video diffusion models plays an essential role in building general-purpose physically plausible world simulators, yet accurately evaluating such capacity remains a challenging task due to the difficulty in disentangling physics correctness from visual appearance in…

Cited by 0SourcecodeScholar
2026

Inference-time Physics Alignment of Video Generative Models with Latent World Models

CVPR 2026

State-of-the-art video generative models produce promising visual content yet often violate basic physics principles, limiting their utility. While some attribute this deficiency to insufficient physics understanding from pre-training, we find that the shortfall in physics plausibility also stems fr

Cited by 0SourcecodeScholar
2025

SpatialBot: Precise Spatial Understanding with Vision Language Models

ICRA 2025

Vision Language Models (VLMs) have achieved impressive performance in 2D image understanding; however, they still struggle with spatial understanding, which is fundamental to embodied AI. In this paper, we propose SpatialBot, a model designed to enhance spatial understanding by utilizing both RGB an

Cited by 167SourcecodeScholar
2024

Hidden in Plain Sight: Evaluating Abstract Shape Recognition in Vision-Language Models

NeurIPS 2024poster

Despite the importance of shape perception in human vision, early neural image classifiers relied less on shape information for object recognition than other (often spurious) features. While recent research suggests that current large Vision-Language Models (VLMs) exhibit more reliance on shape, we…

2024

Not Just Pretty Pictures: Toward Interventional Data Augmentation Using Text-to-Image Generators

ICML 2024poster

Neural image classifiers are known to undergo severe performance degradation when exposed to inputs that are sampled from environmental conditions that differ from their training data. Given the recent progress in Text-to-Image (T2I) generation, a natural question is how modern T2I generators can be…

2024

RAG-Driver: Generalisable Driving Explanations with Retrieval-Augmented In-Context Multi-Modal Large Language Model Learning

RSS 2024poster

We need to trust robots that use often opaque AI methods. They need to explain themselves to us, and we need to trust their explanation. In this regard, explainability plays a critical role in trustworthy autonomous decision-making to foster transparency and acceptance among end users, especially in…

Cited by 83SourcePDFScholar
2024

Real-Fake: Effective Training Data Synthesis Through Distribution Matching

ICLR 2024poster

Synthetic training data has gained prominence in numerous learning tasks and scenarios, offering advantages such as dataset augmentation, generalization evaluation, and privacy preservation. Despite these benefits, the efficiency of synthetic data generated by current methodologies remains inferior…

2023

Off the Radar: Uncertainty-Aware Radar Place Recognition with Introspective Querying and Map Maintenance

IROS 2023poster

Localisation with Frequency-Modulated Continuous-Wave (FMCW) radar has gained increasing interest due to its inherent resistance to challenging environments. However, complex artefacts of the radar measurement process require appropriate uncertainty estimation - to ensure the safe and reliable appli…

Cited by 8SourceScholar