← Search

Daniel Harari

2 accepted papers

2025

Seeing More with Less: Human-like Representations in Vision Models

CVPR 2025highlight

Large multimodal models (LMMs) typically process visual inputs with uniform resolution across the entire field of view, leading to inefficiencies when non-critical image regions are processed as precisely as key areas. Inspired by the human visual system's foveated approach, we apply a sampling meth…

Cited by 0SourcePDFScholar
2024

Why Not Use Your Textbook? Knowledge-Enhanced Procedure Planning of Instructional Videos

CVPR 2024poster

In this paper we explore the capability of an agent to construct a logical sequence of action steps thereby assembling a strategic procedural plan. This plan is crucial for navigating from an initial visual observation to a target visual outcome as depicted in real-life instructional videos. Existin…