← Search

Jindong Jiang

8 accepted papers

2026

4DP-QA: Scalable QA for 4D Perception in Vision Language Models

CVPR 2026

Despite recent advances, Vision Language Models (VLMs) still struggle to grasp the dynamics of the world. We note that the ability to reason about a 4D scene, challenging in itself, is further complicated by two factors. First, VLMs observe motion indirectly via its projection onto 2D images. Second

Cited by 0SourceScholar
2025

Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models

NeurIPS 2025poster

We introduce Eagle2.5, a frontier vision-language model (VLM) for long-context multimodal learning. Our work addresses the challenges in long video comprehension and high-resolution image understanding, introducing a generalist framework for both tasks. The proposed training framework incorporates A…

Cited by 0SourceScholar
2024

Layout-Agnostic Scene Text Image Synthesis with Diffusion Models

CVPR 2024poster

While diffusion models have significantly advanced the quality of image generation their capability to accurately and coherently render text within these images remains a substantial challenge. Conventional diffusion-based methods for scene text generation are typically limited by their reliance on…

Cited by 5SourcePDFScholar
2020

Improving Generative Imagination in Object-Centric World Models

ICML 2020poster

The remarkable recent advances in object-centric generative world models raise a few questions. First, while many of the recent achievements are indispensable for making a general and versatile world model, it is quite unclear how these ingredients can be integrated into a unified framework. Second,…

Cited by 88SourcePDFScholar
2020

SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and Decomposition

ICLR 2020poster

The ability to decompose complex multi-object scenes into meaningful abstractions like objects is fundamental to achieve higher-level cognition. Previous approaches for unsupervised object-oriented scene representation learning are either based on spatial-attention or scene-mixture approaches and li…

Cited by 269SourceScholar