← Search

Irene Pi

2 accepted papers

2026

Building a Precise Video Language with Human-AI Oversight

CVPR 2026

Video-language models (VLMs) learn to reason about the dynamic visual world through natural language. We introduce a suite of open datasets, benchmarks, and recipes for scalable oversight that enable precise video captioning. First, we define a structured specification for describing subjects, scene

Cited by 0SourcecodeScholar
2025

AgMMU: A Comprehensive Agricultural Multimodal Understanding Benchmark

NeurIPS 2025poster

We present **AgMMU**, a challenging real‑world benchmark for evaluating and advancing vision-language models (VLMs) in the knowledge‑intensive domain of agriculture. Unlike prior datasets that rely on crowdsourced prompts, AgMMU is distilled from 116,231 authentic dialogues between everyday growers…

Cited by 0SourceScholar