← Search

Adam Botach

4 accepted papers

2026

Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language Models

CVPR 2026

Segmenting long-form videos into semantically coherent scenes is a fundamental task in large-scale video understanding. Existing encoder-based methods are limited by visual-centric biases, classify each shot in isolation without leveraging sequential dependencies, and lack both narrative understandi

Cited by 0SourceScholar
2025

Distilling the Knowledge in Data Pruning

ICML 2025poster

With the increasing size of datasets used for training neural networks, data pruning has gained traction in recent years. However, most current data pruning algorithms are limited in their ability to preserve accuracy compared to models trained on the full data, especially in high pruning regimes. I…

Cited by 5SourcePDFScholar
2025

Group-Aware Reinforcement Learning for Output Diversity in Large Language Models

EMNLP 2025

Large Language Models (LLMs) often suffer from mode collapse, repeatedly generating the same few completions even when many valid answers exist, limiting their diversity across a wide range of tasks. We introduce Group-Aware Policy Optimization (GAPO) , a simple extension of the recent and popular G

2022

End-to-End Referring Video Object Segmentation With Multimodal Transformers

CVPR 2022poster

The referring video object segmentation task (RVOS) involves segmentation of a text-referred object instance in the frames of a given video. Due to the complex nature of this multimodal task, which combines text reasoning, video understanding, instance segmentation and tracking, existing approaches…

Cited by 181PDFcodeScholar