← Search

A. Said Gurbuz

2 accepted papers

2026

Moving Beyond Sparse Grounding with Complete Screen Parsing Supervision

ICML 2026poster

Modern computer-use agents (CUA) must perceive a screen as a structured state, what elements are visible, where they are, and what text they contain, before they can reliably ground instructions and act. Yet, most available grounding datasets provide sparse supervision, with *insufficient* and *low-…

Cited by 0SourceScholar
2025

SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

ICCV 2025poster

We introduce SmolDocling, an ultra-compact vision-language model targeting end-to-end document conversion. Our model comprehensively processes entire pages by generating DocTags, a new universal markup format that captures all page elements in their full context with location. Unlike existing approa…