IROS 20253 citations

Closing the intent-to-behavior gap via Fulfillment Priority Logic

Bassel El Mabsout, Abdelrahman AbdelGawad, Renato Mancuso

Abstract

Practitioners designing reinforcement learning policies face a fundamental challenge: translating intended behavioral objectives into representative reward functions. This challenge stems from behavioral intent requiring simultaneous achievement of multiple competing objectives, typically addressed through labor-intensive linear reward composition that yields brittle results. Consider the ubiquitous robotics scenario where performance maximization directly conflicts with energy conservation. Such competitive dynamics are resistant to simple linear reward combinations. In this paper, we present the concept of objective fulfillment upon which we build Fulfillment Priority Logic (FPL). FPL allows practitioners to define logical formulae representing their intentions and priorities within multi-objective reinforcement learning. Our novel Balanced Policy Gradient algorithm leverages FPL specifications to achieve up to 500% better sample efficiency compared to Soft Actor Critic. Notably, this work constitutes the first implementation of a non-linear utility scalarization design, intended explicitly for continuous control problems.

BibTeX
@inproceedings{iros2025_closingtheintent,
  title = {Closing the intent-to-behavior gap via Fulfillment Priority Logic},
  author = {Bassel El Mabsout and Abdelrahman AbdelGawad and Renato Mancuso},
  booktitle = {IROS 2025},
  year = {2025}
}
Closing the intent-to-behavior gap via Fulfillment Priority Logic · IROS 2025