← Search

Firoz Shaik

2 accepted papers

2026

PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks

ICML 2026poster

Creating and editing slides is a rich, multimodal activity that is ubiquitous in professional and educational settings, making it an ideal testbed for real-world computer-use agents. Microsoft PowerPoint is among the most widely adopted and feature-rich environments for presentation creation. We int…

Cited by 0SourceScholar
2025

A MISMATCHED Benchmark for Scientific Natural Language Inference

ACL 2025finding

Scientific Natural Language Inference (NLI) is the task of predicting the semantic relation between a pair of sentences extracted from research articles. Existing datasets for this task are derived from various computer science (CS) domains, whereas non-CS domains are completely ignored. In this pap…