← Search

Liu Leqi

17 accepted papers

2025

A Common Pitfall of Margin-based Language Model Alignment: Gradient Entanglement

ICLR 2025poster

Reinforcement Learning from Human Feedback (RLHF) has become the predominant approach for aligning language models (LMs) to be more helpful and less harmful. At its core, RLHF uses a margin-based loss for preference optimization, which specifies the ideal LM behavior only in terms of the difference…

2025

EquivaMap: Leveraging LLMs for Automatic Equivalence Checking of Optimization Formulations

ICML 2025poster

A fundamental problem in combinatorial optimization is identifying equivalent formulations. Despite the growing need for automated equivalence checks---driven, for example, by *optimization copilots*, which generate problem formulations from natural language descriptions---current approaches rely o…

2025

ExPO: Unlocking Hard Reasoning with Self-Explanation-Guided Reinforcement Learning

NeurIPS 2025poster

Recent advances in large language models have been driven by reinforcement learning (RL)-style post-training, which improves reasoning by optimizing model outputs based on reward or preference signals. GRPO-style approaches implement this by using self-generated samples labeled by an outcome-based v…

Cited by 0SourceScholar
2025

More of the Same: Persistent Representational Harms Under Increased Representation

NeurIPS 2025poster

To recognize and mitigate the harms of generative AI systems, it is crucial to consider whether and how different societal groups are represented by these systems. A critical gap emerges when naively measuring or improving *who* is represented, as this does not consider *how* people are represented.…

Cited by 0SourceScholar
2025

Prompting Fairness: Integrating Causality to Debias Large Language Models

ICLR 2025poster

Large language models (LLMs), despite their remarkable capabilities, are susceptible to generating biased and discriminatory responses. As LLMs increasingly influence high-stakes decision-making (e.g., hiring and healthcare), mitigating these biases becomes critical. In this work, we propose a causa…

Cited by 0SourcePDFScholar
2025

The Progress Illusion: Revisiting meta-evaluation standards of LLM evaluators

EMNLP 2025

LLM judges have gained popularity as an inexpensive and performant substitute for human evaluation. However, we observe that the meta-evaluation setting in which the reliability of these LLM evaluators is established is substantially different from their use in model development. To address this, we

2022

Action-Sufficient State Representation Learning for Control with Structural Constraints

ICML 2022spotlight

Perceived signals in real-world scenarios are usually high-dimensional and noisy, and finding and using their representation that contains essential and sufficient information required by downstream decision-making tasks will help improve computational efficiency and generalization ability in the ta…

Cited by 48SourcePDFScholar
2022

Modeling Attrition in Recommender Systems with Departing Bandits

AAAI 2022technical

Traditionally, when recommender systems are formalized as multi-armed bandits, the policy of the recommender system influences the rewards accrued, but not the length of interaction. However, in real-world systems, dissatisfied users may depart (and never come back). In this work, we propose a novel…

Cited by 17SourcePDFScholar
2022

Off-Policy Risk Assessment for Markov Decision Processes

AISTATS 2022poster

Addressing such diverse ends as mitigating safety risks, aligning agent behavior with human preferences, and improving the efficiency of learning, an emerging line of reinforcement learning research addresses the entire distribution of returns and various risk functionals that depend upon it. In the…

Cited by 8SourcePDFScholar
2022

Supervised Learning with General Risk Functionals

ICML 2022spotlight

Standard uniform convergence results bound the generalization gap of the expected loss over a hypothesis class. The emergence of risk-sensitive learning requires generalization guarantees for functionals of the loss distribution beyond the expectation. While prior works specialize in uniform converg…

Cited by 11SourcePDFScholar
2021

Off-Policy Risk Assessment in Contextual Bandits

NeurIPS 2021poster

Even when unable to run experiments, practitioners can evaluate prospective policies, using previously logged data. However, while the bandits literature has adopted a diverse set of objectives, most research on off-policy evaluation to date focuses on the expected reward. In this paper, we introduc…

Cited by 41SourcePDFScholar
2021

Rebounding Bandits for Modeling Satiation Effects

NeurIPS 2021poster

Psychological research shows that enjoyment of many goods is subject to satiation, with short-term satisfaction declining after repeated exposures to the same item. Nevertheless, proposed algorithms for powering recommender systems seldom model these dynamics, instead proceeding as though user prefe…

Cited by 30SourcePDFScholar
2019

Game Design for Eliciting Distinguishable Behavior

NeurIPS 2019poster

The ability to inferring latent psychological traits from human behavior is key to developing personalized human-interacting machine learning systems. Approaches to infer such traits range from surveys to manually-constructed experiments and games. However, these traditional games are limited becaus…

Cited by 2SourcePDFScholar
2018

The Sample Complexity of Semi-Supervised Learning with Nonparametric Mixture Models

NeurIPS 2018poster

We study the sample complexity of semi-supervised learning (SSL) and introduce new assumptions based on the mismatch between a mixture model learned from unlabeled data and the true mixture model induced by the (unknown) class conditional distributions. Under these assumptions, we establish an $\Ome…

Cited by 5SourcePDFScholar