← Search

Francesco Quinzan

5 accepted papers

2025

Doubly Robust Alignment for Large Language Models

NeurIPS 2025poster

This paper studies reinforcement learning from human feedback (RLHF) for aligning large language models with human preferences. While RLHF has demonstrated promising results, many algorithms are highly sensitive to misspecifications in the underlying preference model (e.g., the Bradley-Terry model),…

Cited by 0SourcecodeScholar
2024

Learning Decision Policies with Instrumental Variables through Double Machine Learning

ICML 2024poster

A common issue in learning decision-making policies in data-rich settings is spurious correlations in the offline dataset, which can be caused by hidden confounders. Instrumental variable (IV) regression, which utilises a key uncounfounded variable called the instrument, is a standard technique for…

2023

DRCFS: Doubly Robust Causal Feature Selection

ICML 2023poster

Knowing the features of a complex system that are highly relevant to a particular target variable is of fundamental interest in many areas of science. Existing approaches are often limited to linear settings, sometimes lack guarantees, and in most cases, do not scale to the problem at hand, in parti…

Cited by 11SourcePDFScholar
2023

Fast Feature Selection with Fairness Constraints

AISTATS 2023poster

We study the fundamental problem of selecting optimal features for model construction. This problem is computationally challenging on large datasets, even with the use of greedy algorithm variants. To address this challenge, we extend the adaptive query model, recently proposed for the greedy forwar…

Cited by 4SourcePDFScholar
2021

Adaptive Sampling for Fast Constrained Maximization of Submodular Functions

AISTATS 2021poster

Several large-scale machine learning tasks, such as data summarization, can be approached by maximizing functions that satisfy submodularity. These optimization problems often involve complex side constraints, imposed by the underlying application. In this paper, we develop an algorithm with poly-lo…

Cited by 3SourcePDFScholar