2018
Negotiable Reinforcement Learning for Pareto Optimal Sequential Decision-Making
NeurIPS 2018poster
It is commonly believed that an agent making decisions on behalf of two or more principals who have different utility functions should adopt a Pareto optimal policy, i.e. a policy that cannot be improved upon for one principal without making sacrifices for another. Harsanyi's theorem shows that when…