← Search

Shijia Yang

3 accepted papers

2026

CaptionQA: Is Your Caption as Useful as the Image Itself?

CVPR 2026

Image captions serve as efficient surrogates for visual content in multimodal systems such as retrieval, recommendation, multi-step agentic inference pipelines. Yet current evaluation practices miss a fundamental question: Can captions stand-in for images in real downstream tasks? We propose a utili

Cited by 0SourcecodeScholar
2023

Time Will Tell: New Outlooks and A Baseline for Temporal Multi-View 3D Object Detection

ICLR 2023top-5%

While recent camera-only 3D detection methods leverage multiple timesteps, the limited history they use significantly hampers the extent to which temporal fusion can improve object perception. Observing that existing works' fusion of multi-frame images are instances of temporal stereo matching, we f…

2022

Image2Point: 3D Point-Cloud Understanding with 2D Image Pretrained Models

ECCV 2022poster

"3D point-clouds and 2D images are different visual representations of the physical world. While human vision can understand both representations, computer vision models designed for 2D image and 3D point-cloud understanding are quite different. Our paper explores the potential of transferring 2D mo…