← Search

Ondrej Bajgar

3 accepted papers

2024

Negative Human Rights as a Basis for Long-term AI Safety and Regulation (Abstract Reprint)

IJCAI 2024poster

If autonomous AI systems are to be reliably safe in novel situations, they will need to incorporate general principles guiding them to recognize and avoid harmful behaviours. Such principles may need to be supported by a binding system of regulation, which would need the underlying principles to be…

Cited by 0SourcePDFScholar
2024

Walking the Values in Bayesian Inverse Reinforcement Learning

UAI 2024poster

The goal of Bayesian inverse reinforcement learning (IRL) is recovering a posterior distribution over reward functions using a set of demonstrations from an expert optimizing for a reward unknown to the learner. The resulting posterior over rewards can then be used to synthesize an apprentice policy…

Cited by 1SourcePDFScholar