← Search

Victor Carbune

7 accepted papers

2025

ScreenQA: Large-Scale Question-Answer Pairs Over Mobile App Screenshots

NAACL 2025long

We introduce ScreenQA, a novel benchmarking dataset designed to advance screen content understanding through question answering. The existing screen datasets are focused either on low-level structural and component understanding, or on a much higher-level composite task such as navigation and task c…

2025

Self-play through Computational Runtimes improves Chart Reasoning

ACL 2025finding

Vision-language models (VLMs) achieve impressive zero-shot performance on multimodal reasoning tasks. Typically, best reported performance is achieved with a zero- or a few-shot prompt. We observe that asking the model to take other routes of solving the same task, such as through code generation, h…

Cited by 0SourcePDFScholar
2024

Chart-based Reasoning: Transferring Capabilities from LLMs to VLMs

NAACL 2024findings

Vision-language models (VLMs) are achieving increasingly strong performance on multimodal tasks. However, reasoning capabilities remain limited particularly for smaller VLMs, while those of large-language models (LLMs) have seen numerous improvements. We pro-pose a technique to transfer capabilities…

2024

LLMs cannot find reasoning errors, but can correct them given the error location

ACL 2024findings

While self-correction has shown promise in improving LLM outputs in terms of style and quality (e.g. Chen et al., 2023b; Madaan et al.,2023), recent attempts to self-correct logical or reasoning errors often cause correct answers to become incorrect, resulting in worse performances overall (Huang et…

2024

RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

ICML 2024poster

Reinforcement learning from human feedback (RLHF) has proven effective in aligning large language models (LLMs) with human preferences, but gathering high-quality preference labels is expensive. RL from AI Feedback (RLAIF), introduced in Bai et al. (2022b), offers a promising alternative that trains…

Cited by 99SourcePDFScholar
2024

ScreenAI: A Vision-Language Model for UI and Infographics Understanding

IJCAI 2024poster

Screen user interfaces (UIs) and infographics, sharing similar visual language and design principles, play important roles in human communication and human-machine interaction. We introduce ScreenAI, a vision-language model that specializes in UI and infographics understanding. Our model improves…

2021

Replacing Human Audio with Synthetic Audio for on-Device Unspoken Punctuation Prediction

ICASSP 2021accepted

We present a novel multi-modal unspoken punctuation prediction system for the English language which combines acoustic and text features. We demonstrate for the first time, that by relying exclusively on synthetic data generated using a prosody-aware text-to-speech system, we can outperform a model…

Cited by 0SourceScholar