← Search

Ujjwal Upadhyay

2 accepted papers

2026

Time Blindness: Why Video-Language Models Can't See What Humans Can?

CVPR 2026

Recent advances in vision-language models (VLMs) have made impressive strides in understanding spatio-temporal relationships in videos. However, when spatial information is obscured, these models struggle to capture purely temporal patterns. We introduce SpookyBench, a benchmark where information is

Cited by 0SourcecodeScholar
2022

3D CoMPaT: Composition of Materials on Parts of 3D Things

ECCV 2022poster

"We present 3D CoMPaT, a richly annotated large-scale dataset of more than 7.19 million rendered compositions of Materials on Parts of 7262 unique 3D Models; 990 compositions per model on average. 3D CoMPaT covers 43 shape categories, 235 unique part names, and 167 unique material classes that can b…

Cited by 16SourcePDFScholar