← Search

Lemin Kong

4 accepted papers

2025

Common Learning Constraints Alter Interpretations of Direct Preference Optimization

AISTATS 2025poster

Large language models in the past have typically relied on some form of reinforcement learning with human feedback (RLHF) to better align model responses with human preferences. However, because of oft-observed instabilities when implementing these RLHF pipelines, various reparameterization techniq…

Cited by 0SourceScholar
2025

Explicit Preference Optimization: No Need for an Implicit Reward Model

ICML 2025poster

The generated responses of large language models (LLMs) are often fine-tuned to human preferences through a process called reinforcement learning from human feedback (RLHF). As RLHF relies on a challenging training sequence, whereby a separate reward model is independently learned and then later ap…

2023

A Convergent Single-Loop Algorithm for Relaxation of Gromov-Wasserstein in Graph Data

ICLR 2023poster

In this work, we present the Bregman Alternating Projected Gradient (BAPG) method, a single-loop algorithm that offers an approximate solution to the Gromov-Wasserstein (GW) distance. We introduce a novel relaxation technique that balances accuracy and computational efficiency, albeit with some com…

Cited by 13SourcePDFScholar
2023

Outlier-Robust Gromov-Wasserstein for Graph Data

NeurIPS 2023spotlight

Gromov-Wasserstein (GW) distance is a powerful tool for comparing and aligning probability distributions supported on different metric spaces. Recently, GW has become the main modeling technique for aligning heterogeneous data for a wide range of graph learning tasks. However, the GW distance is kno…