← Search

Stuart Armstrong

2 accepted papers

2018

Occam's razor is insufficient to infer the preferences of irrational agents

NeurIPS 2018poster

Inverse reinforcement learning (IRL) attempts to infer human rewards or preferences from observed behavior. Since human planning systematically deviates from rationality, several approaches have been tried to account for specific human shortcomings. However, the general problem of inferring the rew…

Cited by 127SourcePDFScholar