← Search

Ari Seff

7 accepted papers

2025

UINavBench: A Framework for Comprehensive Evaluation of Interactive Digital Agents

ICCV 2025poster

We build a comprehensive online evaluation benchmark for language-conditioned multi-step task execution on mobile interfaces. Our benchmark strives to evaluate the multi-step planning, reasoning, and visual grounding capabilities of agents, using mobile user interfaces as a concrete testbed. To buil…

Cited by 0SourcePDFScholar
2024

Improving Agent Behaviors with RL Fine-tuning for Autonomous Driving

ECCV 2024poster

"A major challenge in autonomous vehicle research is modeling agent behaviors, which has critical applications including constructing realistic and reliable simulations for off-board evaluation and forecasting traffic agents motion for onboard planning. While supervised learning has shown success in…

2023

MotionLM: Multi-Agent Motion Forecasting as Language Modeling

ICCV 2023poster

Reliable forecasting of the future behavior of road agents is a critical component to safe planning in autonomous vehicles. Here, we represent continuous trajectories as sequences of discrete motion tokens and cast multi-agent motion prediction as a language modeling task over this domain. Our model…

Cited by 104PDFScholar
2019

Discrete Object Generation with Reversible Inductive Construction

NeurIPS 2019poster

The success of generative modeling in continuous domains has led to a surge of interest in generating discrete data such as molecules, source code, and graphs. However, construction histories for these discrete objects are typically not unique and so generative models must reason about intractably l…

2015

DeepDriving: Learning Affordance for Direct Perception in Autonomous Driving

ICCV 2015poster

Today, there are two major paradigms for vision-based autonomous driving systems: mediated perception approaches that parse an entire scene to make a driving decision, and behavior reflex approaches that directly map an input image to a driving action by a regressor. In this paper, we propose a thir…

Cited by 2530PDFScholar
2015

Interleaved Text/Image Deep Mining on a Very Large-Scale Radiology Database

CVPR 2015poster

Despite tremendous progress in computer vision, effective learning on very large-scale (>100K patients) medical image databases has been vastly hindered. We present an interleaved text/image deep learning system to extract and mine the semantic interactions of radiology images and reports from a nat…