← Search

Zhuotong Chen

3 accepted papers

2026

Preference Optimization via Contrastive Divergence: Your Policy Is Secretly an NLL Estimator

AAAI 2026technical

Existing studies on preference optimization (PO) have been focused on constructing pairwise preference data following simple heuristics, such as maximizing the margin between chosen and rejected responses based on human (or AI) ratings. In this work, we develop a novel PO framework that provides th

Cited by 0SourcePDFScholar
2021

Bayesian Inference with Certifiable Adversarial Robustness

AISTATS 2021poster

We consider adversarial training of deep neural networks through the lens of Bayesian learning and present a principled framework for adversarial training of Bayesian Neural Networks (BNNs) with certifiable guarantees. We rely on techniques from constraint relaxation of non-convex optimisation probl…