← Search

Kaiyang Guo

4 accepted papers

2025

Proximalized Preference Optimization for Diverse Feedback Types: A Decomposed Perspective on DPO

NeurIPS 2025poster

Direct alignment methods typically train large language models (LLMs) by contrasting the likelihoods of preferred and dispreferred responses. While effective for matching relative preferences, these methods have been widely observed to depress the absolute likelihoods of example responses. Consequen…

Cited by 0SourceScholar
2024

Calibrated One Round Federated Learning with Bayesian Inference in the Predictive Space

AAAI 2024technical

Federated Learning (FL) involves training a model over a dataset distributed among clients, with the constraint that each client’s dataset is localized and possibly heterogeneous. In FL, small and noisy datasets are common, highlighting the need for well-calibrated models that represent the uncertai…

2022

Model-Based Offline Reinforcement Learning with Pessimism-Modulated Dynamics Belief

NeurIPS 2022accept

Model-based offline reinforcement learning (RL) aims to find highly rewarding policy, by leveraging a previously collected static dataset and a dynamics model. While the dynamics model learned through reuse of the static dataset, its generalization ability hopefully promotes policy learning if prope…

2022

Personalized Federated Learning via Variational Bayesian Inference

ICML 2022spotlight

Federated learning faces huge challenges from model overfitting due to the lack of data and statistical diversity among clients. To address these challenges, this paper proposes a novel personalized federated learning method via Bayesian variational inference named pFedBayes. To alleviate the overfi…

Cited by 122SourcePDFScholar