← Search

Jingcheng Ni

5 accepted papers

2026

Distilling Geometry Priors for 3D-Consistent Video Generation

ICML 2026poster

While recent video diffusion models (VDMs) produce visually impressive results, they fundamentally struggle to maintain 3D structural consistency, often resulting in object deformation or spatial drift. We hypothesize that these failures arise because standard denoising objectives lack explicit ince…

Cited by 0SourceScholar
2026

Iris: Integrating Language into Diffusion-based Monocular Depth Estimation

CVPR 2026

Conventional monocular depth estimators suffer from visual ambiguities and nuisances. We demonstrate that language can improve the fidelity of estimates by providing additional information through text as a condition, thereby reducing the solution space for depth estimates. This conditional distribu

Cited by 0SourceScholar
2025

MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction

CVPR 2025poster

World models that forecast environmental changes from actions are vital for autonomous driving models with strong generalization. The prevailing driving world model mainly build on pixel-level video prediction model. Although these models can produce high-fidelity video sequences with advanced diffu…

2025

UniMLVG: Unified Framework for Multi-view Long Video Generation with Comprehensive Control Capabilities for Autonomous Driving

ICCV 2025poster

The creation of diverse and realistic driving scenarios has become essential to enhance perception and planning capabilities of the autonomous driving system. However, generating long-duration, surround-view consistent driving videos remains a significant challenge. To address this, we present UniML…

2022

Motion Sensitive Contrastive Learning for Self-Supervised Video Representation

ECCV 2022poster

"Contrastive learning has shown great potential in video representation learning. However, existing approaches fail to sufficiently exploit short-term motion dynamics, which are crucial to various down-stream video understanding tasks. In this paper, we propose Motion Sensitive Contrastive Learning…

Cited by 20SourcePDFScholar