← Search

Shunit Haviv Hakimi

2 accepted papers

2026

Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language Models

CVPR 2026

Segmenting long-form videos into semantically coherent scenes is a fundamental task in large-scale video understanding. Existing encoder-based methods are limited by visual-centric biases, classify each shot in isolation without leveraging sequential dependencies, and lack both narrative understandi

Cited by 0SourceScholar
2025

Group-Aware Reinforcement Learning for Output Diversity in Large Language Models

EMNLP 2025

Large Language Models (LLMs) often suffer from mode collapse, repeatedly generating the same few completions even when many valid answers exist, limiting their diversity across a wide range of tasks. We introduce Group-Aware Policy Optimization (GAPO) , a simple extension of the recent and popular G