← Search

Erez Yosef

2 accepted papers

2026

Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language Models

CVPR 2026

Segmenting long-form videos into semantically coherent scenes is a fundamental task in large-scale video understanding. Existing encoder-based methods are limited by visual-centric biases, classify each shot in isolation without leveraging sequential dependencies, and lack both narrative understandi

Cited by 0SourceScholar
2024

Mind The Edge: Refining Depth Edges in Sparsely-Supervised Monocular Depth Estimation

CVPR 2024poster

Monocular Depth Estimation (MDE) is a fundamental problem in computer vision with numerous applications. Recently LIDAR-supervised methods have achieved remarkable per-pixel depth accuracy in outdoor scenes. However significant errors are typically found in the proximity of depth discontinuities i.e…