← Search

Desen Zhou

6 accepted papers

2026

VectorWorld: Efficient Streaming World Model via Diffusion Flow on Vector Graphs

ICML 2026spotlight

Closed-loop evaluation of autonomous-driving policies requires interactive simulation beyond log replay. However, existing generative world models often degrade in closed loop due to (i) history-free initialization that mismatches policy inputs, (ii) multi-step sampling latency that violates real-ti…

Cited by 0SourceScholar
2023

Weakly-supervised HOI Detection via Prior-guided Bi-level Representation Learning

ICLR 2023poster

Human object interaction (HOI) detection plays a crucial role in human-centric scene understanding and serves as a fundamental building block for many vision tasks. One generalizable and scalable strategy for HOI detection is to use weak supervision, learning from image-level annotations only. This…

Cited by 15SourcePDFScholar
2022

Action Quality Assessment with Temporal Parsing Transformer

ECCV 2022poster

"Action Quality Assessment(AQA) is important for action understanding and resolving the task poses unique challenges due to subtle visual differences. Existing state-of-the-art methods typically rely on the holistic video representations for score regression or ranking, which limits the generalizati…

Cited by 60SourcePDFScholar
2022

Human-Object Interaction Detection via Disentangled Transformer

CVPR 2022poster

Human-Object Interaction Detection tackles the problem of joint localization and classification of human object interactions. Existing HOI transformers either adopt a single decoder for triplet prediction, or utilize two parallel decoders to detect individual objects and interactions separately, and…

Cited by 77PDFScholar
2019

Pose-Aware Multi-Level Feature Network for Human Object Interaction Detection

ICCV 2019oral

Reasoning human object interactions is a core problem in human-centric scene understanding and detecting such relations poses a unique challenge to vision systems due to large variations in human-object configurations, multiple co-occurring relation instances and subtle visual difference between rel…

Cited by 276PDFcodeScholar
2016

Single-Image Crowd Counting via Multi-Column Convolutional Neural Network

CVPR 2016poster

This paper aims to develop a method that can accurately estimate the crowd count from an individual image with arbitrary crowd density and arbitrary perspective. To this end,we have proposed a simple but effective Multi-column Convolutional Neural Network (MCNN) architecture to map the image to its…

Cited by 2492PDFScholar