← Search

Jiawei He

24 accepted papers

2026

OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models

ICLR 2026poster

Spatial reasoning is a key aspect of cognitive psychology and remains a bottleneck for current vision-language models (VLMs). While extensive research has aimed to evaluate or improve VLMs' understanding of basic spatial relations, such as distinguishing left from right, near from far, and object co…

Cited by 0SourcecodeScholar
2026

Seeing Without Understanding: Disentangling Perception, Reasoning, and Simulation in VLM Gameplay

ICML 2026poster

While Vision-Language Models (VLMs) excel on static visual benchmarks, they consistently underperform in game-based reasoning environments. Existing evaluations conflate failures in perception, rule comprehension, and reasoning. We propose a two-stage diagnostic framework that decomposes VLM perform…

Cited by 0SourceScholar
2025

A Multi-Modal Fusion-Based 3D Multi-Object Tracking Framework With Joint Detection

RA-L 2025

In the classical tracking-by-detection (TBD) paradigm, detection and tracking are separately and sequentially conducted, and data association must be properly performed to achieve satisfactory tracking performance. In this letter, a new multi-object tracking framework is proposed, which integrates o

Cited by 18SourceScholar
2025

DexVLG: Dexterous Vision-Language-Grasp Model at Scale

ICCV 2025poster

As large models gain traction, vision-language models are enabling robots to tackle increasingly complex tasks. However, limited by the difficulty of data collection, progress has mainly focused on controlling simple gripper end-effectors. There is little research on functional grasping with large m…

Cited by 0SourcePDFScholar
2025

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

NeurIPS 2025poster

Recent advances in vision-language-action (VLA) models have shown promise in integrating image generation with action prediction to improve generalization and reasoning in robot manipulation. However, existing methods are limited to challenging image-based forecasting, which suffers from redundant i…

Cited by 0SourcecodeScholar
2025

End-to-End Driving with Online Trajectory Evaluation via BEV World Model

ICCV 2025poster

End-to-end autonomous driving has achieved remarkable progress by integrating perception, prediction, and planning into a fully differentiable framework. Yet, to fully realize its potential, an effective online trajectory evaluation is indispensable to ensure safety. By forecasting the future outcom…

2025

Enhancing End-to-End Autonomous Driving with Latent World Model

ICLR 2025poster

In autonomous driving, end-to-end planners directly utilize raw sensor data, enabling them to extract richer scene features and reduce information loss compared to traditional planners. This raises a crucial research question: how can we develop better scene feature representations to fully leverage…

2025

Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth Estimation

ICCV 2025poster

Current self-supervised monocular depth estimation (MDE) approaches encounter performance limitations due to insufficient semantic-spatial knowledge extraction. To address this challenge, we propose Hybrid-depth, a novel framework that systematically integrates foundation models (e.g., CLIP and DINO…

Cited by 0SourcePDFScholar
2025

SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation

NeurIPS 2025spotlight

While spatial reasoning has made progress in object localization relationships, it often overlooks object orientation—a key factor in 6-DoF fine-grained manipulation. Traditional pose representations rely on pre-defined frames or templates, limiting generalization and semantic grounding. In this pap…

Cited by 0SourceScholar
2024

AutoCast++: Enhancing World Event Prediction with Zero-shot Ranking-based Context Retrieval

ICLR 2024poster

Machine-based prediction of real-world events is garnering attention due to its potential for informed decision-making. Whereas traditional forecasting predominantly hinges on structured data like time-series, recent breakthroughs in language models enable predictions using unstructured text. In par…

2024

Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving

CVPR 2024poster

In autonomous driving predicting future events in advance and evaluating the foreseeable risks empowers autonomous vehicles to plan their actions enhancing safety and efficiency on the road. To this end we propose Drive-WM the first driving world model compatible with existing end-to-end planning mo…

2024

DrivingDojo Dataset: Advancing Interactive and Knowledge-Enriched Driving World Model

NeurIPS 2024poster

Driving world models have gained increasing attention due to their ability to model complex physical dynamics. However, their superb modeling capability is yet to be fully unleashed due to the limited video diversity in current driving datasets. We introduce DrivingDojo, the first dataset tailor-mad…

Cited by 7SourcePDFScholar
2024

OneTrack: Demystifying the Conflict Between Detection and Tracking in End-to-End 3D Trackers

ECCV 2024poster

"Existing end-to-end trackers for vision-based 3D perception suffer from performance degradation due to the conflict between detection and tracking tasks. In this work, we get to the bottom of this conflict, which was vaguely attributed to incompatible task-specific object features previously. We fi…

Cited by 2SourcePDFScholar
2023

3D Video Object Detection With Learnable Object-Centric Global Optimization

CVPR 2023poster

We explore long-term temporal visual correspondence-based optimization for 3D video object detection in this work. Visual correspondence refers to one-to-one mappings for pixels across multiple images. Correspondence-based optimization is the cornerstone for 3D scene reconstruction but is less studi…

2022

DeepFusionMOT: A 3D Multi-Object Tracking Framework Based on Camera-LiDAR Fusion With Deep Association

RA-L 2022

In the recent literature, on the one hand, many 3D multi-object tracking (MOT) works have focused on tracking accuracy and neglected computation speed, commonly by designing rather complex cost functions and feature extractors. On the other hand, some methods have focused too much on computation spe

Cited by 125SourcecodeScholar
2022

Densely Constrained Depth Estimator for Monocular 3D Object Detection

ECCV 2022poster

"Estimating accurate 3D locations of objects from monocular images is a challenging problem because of lacking depth. Previous work shows that utilizing the object’s keypoint projection constraints to estimate multiple depth candidates boosts the detection performance. However, the existing methods…

2021

Learnable Graph Matching: Incorporating Graph Partitioning With Deep Feature Learning for Multiple Object Tracking

CVPR 2021poster

Data association across frames is at the core of Multiple Object Tracking (MOT) task. This problem is usually solved by a traditional graph-based optimization or directly learned via deep learning. Despite their popularity, we find some points worth studying in current paradigm: 1) Existing methods…

Cited by 156PDFcodeScholar
2021

Variational Selective Autoencoder: Learning from Partially-Observed Heterogeneous Data

AISTATS 2021poster

Learning from heterogeneous data poses challenges such as combining data from various sources and of different types. Meanwhile, heterogeneous data are often associated with missingness in real-world applications due to heterogeneity and noise of input sources. In this work, we propose the variation…

Cited by 20SourcePDFScholar
2020

Piggyback GAN: Efficient Lifelong Learning for Image Conditioned Generation

ECCV 2020poster

Humans accumulate knowledge in a lifelong fashion. Modern deep neural networks, on the other hand, are susceptible to catastrophic forgetting: when adapted to perform new tasks, they often fail to preserve their performance on previously learned tasks. Given a sequence of tasks, a naive approach add…

Cited by 47SourcePDFScholar
2019

A Variational Auto-Encoder Model for Stochastic Point Processes

CVPR 2019poster

We propose a novel probabilistic generative model for action sequences. The model is termed the Action Point Process VAE (APP-VAE), a variational auto-encoder that can capture the distribution over the times and categories of action sequences. Modeling the variety of possible action sequences is a…

Cited by 70PDFScholar
2019

LayoutVAE: Stochastic Scene Layout Generation From a Label Set

ICCV 2019poster

Recently there is an increasing interest in scene generation within the research community. However, models used for generating scene layouts from textual description largely ignore plausible visual variations within the structure dictated by the text. We propose LayoutVAE, a variational autoencoder…

Cited by 185PDFScholar
2019

Lifelong GAN: Continual Learning for Conditional Image Generation

ICCV 2019poster

Lifelong learning is challenging for deep neural networks due to their susceptibility to catastrophic forgetting. Catastrophic forgetting occurs when a trained network is not able to maintain its ability to accomplish previously learned tasks when it is trained to perform new tasks. We study the pro…

Cited by 250PDFScholar
2019

Variational Autoencoders with Jointly Optimized Latent Dependency Structure

ICLR 2019poster

We propose a method for learning the dependency structure between latent variables in deep latent variable models. Our general modeling and inference framework combines the complementary strengths of deep generative models and probabilistic graphical models. In particular, we express the latent var…

Cited by 29SourcePDFScholar
2018

Probabilistic Video Generation using Holistic Attribute Control

ECCV 2018poster

Videos express highly structured spatio-temporal patterns of visual data. A video can be thought of as being governed by two factors: (i) temporally invariant (e.g., person identity), or slowly varying (e.g., activity), attribute-induced appearance, encoding the persistent content of each frame, and…

Cited by 88SourcePDFScholar