← Search

Kristjan Greenewald

35 accepted papers

2026

Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training

ICLR 2026poster

We revisit Group Relative Policy Optimization (GRPO) in both on-policy and off-policy optimization regimes. Our motivation comes from recent work on off-policy Proximal Policy Optimization (PPO), which improves training stability, sampling efficiency, and memory usage. In addition, a recent analysis…

Cited by 0SourceScholar
2025

Activated LoRA: Fine-tuned LLMs for Intrinsics

NeurIPS 2025poster

Low-Rank Adaptation (LoRA) has emerged as a highly efficient framework for finetuning the weights of large foundation models, and has become the go-to method for data-driven customization of LLMs. Despite the promise of highly customized behaviors and capabilities, switching between relevant LoRAs i…

Cited by 0SourcecodeScholar
2025

Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead

ICML 2025poster

Fine-tuning large language models (LLMs) with low-rank adaptations (LoRAs) has become common practice, often yielding numerous copies of the same LLM differing only in their LoRA updates. This paradigm presents challenges for systems that serve real-time responses to queries that each involve a diff…

Cited by 5SourcePDFScholar
2025

Know What You Don't Know: Uncertainty Calibration of Process Reward Models

NeurIPS 2025poster

Process reward models (PRMs) play a central role in guiding inference-time scaling algorithms for large language models (LLMs). However, we observe that even state-of-the-art PRMs can be poorly calibrated. Specifically, they tend to overestimate the success probability that a partial reasoning step…

Cited by 0SourceScholar
2025

Partially Observed Trajectory Inference using Optimal Transport and a Dynamics Prior

ICLR 2025poster

Trajectory inference seeks to recover the temporal dynamics of a population from snapshots of its (uncoupled) temporal marginals, i.e. where observed particles are \emph{not} tracked over time. Prior works addressed this challenging problem under a stochastic differential equation (SDE) model with a…

2024

Asymmetry in Low-Rank Adapters of Foundation Models

ICML 2024poster

Parameter-efficient fine-tuning optimizes large, pre-trained foundation models by updating a subset of parameters; in this class, Low-Rank Adaptation (LoRA) is particularly effective. Inspired by an effort to investigate the different roles of LoRA matrices during fine-tuning, this paper characteriz…

2024

Distributional Preference Alignment of LLMs via Optimal Transport

NeurIPS 2024poster

Current LLM alignment techniques use pairwise human preferences at a sample level, and as such, they do not imply an alignment on the distributional level. We propose in this paper Alignment via Optimal Transport (AOT), a novel method for distributional preference alignment of LLMs. AOT aligns LLMs…

Cited by 15SourcePDFScholar
2024

Multivariate Stochastic Dominance via Optimal Transport and Applications to Models Benchmarking

NeurIPS 2024poster

Stochastic dominance is an important concept in probability theory, econometrics and social choice theory for robustly modeling agents' preferences between random outcomes. While many works have been dedicated to the univariate case, little has been done in the multivariate scenario, wherein an age…

Cited by 1SourcePDFScholar
2024

Privacy without Noisy Gradients: Slicing Mechanism for Generative Model Training

NeurIPS 2024poster

Training generative models with differential privacy (DP) typically involves injecting noise into gradient updates or adapting the discriminator's training procedure. As a result, such approaches often struggle with hyper-parameter tuning and convergence. We consider the \emph{slicing privacy mech…

Cited by 0SourcePDFScholar
2024

Risk Aware Benchmarking of Large Language Models

ICML 2024poster

We propose a distributional framework for benchmarking socio-technical risks of foundation models with quantified statistical significance. Our approach hinges on a new statistical relative testing based on first and second order stochastic dominance of real random variables. We show that the second…

Cited by 1SourcePDFScholar
2024

Score Distillation via Reparametrized DDIM

NeurIPS 2024poster

While 2D diffusion models generate realistic, high-detail images, 3D shape generation methods like Score Distillation Sampling (SDS) built on these 2D diffusion models produce cartoon-like, over-smoothed shapes. To help explain this discrepancy, we show that the image guidance used in Score Distil…

2024

Slicing Mutual Information Generalization Bounds for Neural Networks

ICML 2024poster

The ability of machine learning (ML) algorithms to generalize well to unseen data has been studied through the lens of information theory, by bounding the generalization error with the input-output mutual information (MI), i.e., the MI between the training data and the learned hypothesis. Yet, these…

2024

Thermometer: Towards Universal Calibration for Large Language Models

ICML 2024poster

We consider the issue of calibration in large language models (LLM). Recent studies have found that common interventions such as instruction tuning often result in poorly calibrated LLMs. Although calibration is well-explored in traditional applications, calibrating LLMs is uniquely challenging. The…

2023

Identifiability Guarantees for Causal Disentanglement from Soft Interventions

NeurIPS 2023poster

Causal disentanglement aims to uncover a representation of data using latent variables that are interrelated through a causal model. Such a representation is identifiable if the latent model that explains the data is unique. In this paper, we focus on the scenario where unpaired observational and in…

2023

Learning Proximal Operators to Discover Multiple Optima

ICLR 2023poster

Finding multiple solutions of non-convex optimization problems is a ubiquitous yet challenging task. Most past algorithms either apply single-solution optimization methods from multiple random initial guesses or search in the vicinity of found solutions using ad hoc heuristics. We present an end-to-…

2023

Minimum-Entropy Coupling Approximation Guarantees Beyond the Majorization Barrier

AISTATS 2023poster

Given a set of discrete probability distributions, the minimum entropy coupling is the minimum entropy joint distribution that has the input distributions as its marginals. This has immediate relevance to tasks such as entropic causal inference for causal graph discovery and bounding mutual informat…

Cited by 14SourcePDFScholar
2023

Post-processing Private Synthetic Data for Improving Utility on Selected Measures

NeurIPS 2023poster

Existing private synthetic data generation algorithms are agnostic to downstream tasks. However, end users may have specific requirements that the synthetic data must satisfy. Failure to meet these requirements could significantly reduce the utility of the data for downstream use. We introduce a pos…

Cited by 10SourcePDFScholar
2022

$k$-Sliced Mutual Information: A Quantitative Study of Scalability with Dimension

NeurIPS 2022accept

Sliced mutual information (SMI) is defined as an average of mutual information (MI) terms between one-dimensional random projections of the random variables. It serves as a surrogate measure of dependence to classic MI that preserves many of its properties but is more scalable to high dimensions. Ho…

Cited by 19SourcePDFScholar
2022

Entropic Causal Inference: Graph Identifiability

ICML 2022spotlight

Entropic causal inference is a recent framework for learning the causal graph between two variables from observational data by finding the information-theoretically simplest structural explanation of the data, i.e., the model with smallest entropy. In our work, we first extend the causal graph ident…

Cited by 19SourcePDFScholar
2022

Log-Euclidean Signatures for Intrinsic Distances Between Unaligned Datasets

ICML 2022spotlight

The need for efficiently comparing and representing datasets with unknown alignment spans various fields, from model analysis and comparison in machine learning to trend discovery in collections of medical datasets. We use manifold learning to compare the intrinsic geometric structures of different…

2021

High-Dimensional Feature Selection for Sample Efficient Treatment Effect Estimation

AISTATS 2021poster

The estimation of causal treatment effects from observational data is a fundamental problem in causal inference. To avoid bias, the effect estimator must control for all confounders. Hence practitioners often collect data for as many covariates as possible to raise the chances of including the relev…

2021

Improving approximate optimal transport distances using quantization

UAI 2021poster

Optimal transport (OT) is a popular tool in machine learning to compare probability measures geometrically, but it comes with substantial computational burden. Linear programming algorithms for computing OT distances scale cubically in the size of the input, making OT impractical in the large-sample…

Cited by 12SourcePDFScholar
2021

Measuring Generalization with Optimal Transport

NeurIPS 2021spotlight

Understanding the generalization of deep neural networks is one of the most important tasks in deep learning. Although much progress has been made, theoretical error bounds still often behave disparately from empirical observations. In this work, we develop margin-based generalization bounds, where…

2020

Active Structure Learning of Causal DAGs via Directed Clique Trees

NeurIPS 2020poster

A growing body of work has begun to study intervention design for efficient structure learning of causal directed acyclic graphs (DAGs). A typical setting is a \emph{causally sufficient} setting, i.e. a system with no latent confounders, selection bias, or feedback, when the essential graph of the o…

2020

Asymptotic Guarantees for Generative Modeling Based on the Smooth Wasserstein Distance

NeurIPS 2020poster

Minimum distance estimation (MDE) gained recent attention as a formulation of (implicit) generative modeling. It considers minimizing, over model parameters, a statistical distance between the empirical data distribution and the model. This formulation lends itself well to theoretical analysis, but…

Cited by 25SourcePDFScholar
2020

Entropic Causal Inference: Identifiability and Finite Sample Results

NeurIPS 2020poster

Entropic causal inference is a framework for inferring the causal direction between two categorical variables from observational data. The central assumption is that the amount of unobserved randomness in the system is not too large. This unobserved randomness is measured by the entropy of the exoge…

Cited by 19SourcePDFScholar
2020

Gaussian-Smoothed Optimal Transport: Metric Structure and Statistical Efficiency

AISTATS 2020poster

Optimal transport (OT), and in particular the Wasserstein distance, has seen a surge of interest and applications in machine learning. However, empirical approximation under Wasserstein distances suffers from a severe curse of dimensionality, rendering them impractical in high dimensions. As a resul…

Cited by 43SourcePDFScholar
2019

Bayesian Nonparametric Federated Learning of Neural Networks

ICML 2019oral

In federated learning problems, data is scattered across different servers and exchanging or pooling it is often impractical or prohibited. We develop a Bayesian nonparametric framework for federated learning with neural networks. Each data server is assumed to provide local neural network weights,…

2019

Estimating Information Flow in Deep Neural Networks

ICML 2019oral

We study the estimation of the mutual information I(X;T_$\ell$) between the input X to a deep neural network (DNN) and the output vector T_$\ell$ of its $\ell$-th hidden layer (an “internal representation”). Focusing on feedforward networks with fixed weights and noisy internal representations, we d…

Cited by 181SourcePDFScholar
2019

Sample Efficient Active Learning of Causal Trees

NeurIPS 2019poster

We consider the problem of experimental design for learning causal graphs that have a tree structure. We propose an adaptive framework that determines the next intervention based on a Bayesian prior updated with the outcomes of previous experiments, focusing on the setting where observational data i…

Cited by 50SourcePDFScholar
2019

Statistical Model Aggregation via Parameter Matching

NeurIPS 2019poster

We consider the problem of aggregating models learned from sequestered, possibly heterogeneous datasets. Exploiting tools from Bayesian nonparametrics, we develop a general meta-modeling framework that learns shared global latent structures by identifying correspondences among local model parameteri…

2017

Time-dependent spatially varying graphical models, with application to brain fMRI data analysis

NeurIPS 2017poster

In this work, we present an additive model for space-time data that splits the data into a temporally correlated component and a spatially correlated component. We model the spatially correlated portion using a time-varying Gaussian graphical model. Under assumptions on the smoothness of changes in…

Cited by 18SourcePDFScholar