← Search

Alon Oved

3 accepted papers

2026

From Benchmarks to Business Impact: Deploying IBM Generalist Agent in Enterprise Production

AAAI 2026technical

Agents are rapidly advancing in automating digital work, but enterprises face a harder challenge: moving beyond prototypes to deployed systems that deliver measurable business value. This path is complicated by fragmented frameworks, slow development, and the absence of standardized evaluation pract

Cited by 0SourcePDFScholar
2026

ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents

ICLR 2026poster

Autonomous web agents solve complex browsing tasks, yet existing benchmarks measure only whether an agent finishes a task, ignoring whether it does so safely or in a way enterprises can trust. To integrate these agents into critical workflows, safety and trustworthiness (ST) are prerequisite conditi…

Cited by 0SourcecodeScholar
2025

SNAP: Semantic Stories for Next Activity Prediction

AAAI 2025technical

Predicting the next activity in an ongoing process is one of the most common tasks in the business process management (BPM) domain. It allows businesses to optimize resource allocation, enhance operational efficiency, and aid both in risk mitigation and strategic decision-making. Existing state-of-t…

Cited by 4SourcePDFScholar