← Search

Anastasis Kratsios

10 accepted papers

2026

Inverse Entropic Optimal Transport Solves Semi-supervised Learning via Data Likelihood Maximization

ICML 2026poster

Learning conditional distributions $\pi^\star(\cdot|x)$ is a central problem in machine learning, which is typically approached via supervised methods with paired data $(x,y) \sim \pi^\star$. However, acquiring paired data samples is often challenging, especially in problems such as domain translati…

Cited by 0SourceScholar
2025

Filtered not Mixed: Filtering-Based Online Gating for Mixture of Large Language Models

ICLR 2025poster

We propose MoE-F — a formalized mechanism for combining N pre-trained expert Large Language Models (LLMs) in online time-series prediction tasks by adaptively forecasting the best weighting of LLM predictions at every time step. Our mechanism leverages the conditional information in each expert's ru…

2025

Neural Spacetimes for DAG Representation Learning

ICLR 2025poster

We propose a class of trainable deep learning-based geometries called Neural SpaceTimes (NSTs), which can universally represent nodes in weighted Directed Acyclic Graphs (DAGs) as events in a spacetime manifold. While most works in the literature focus on undirected graph representation learning or…

Cited by 1SourcePDFScholar
2024

A Comprehensive Analysis on the Learning Curve in Kernel Ridge Regression

NeurIPS 2024poster

This paper conducts a comprehensive study of the learning curves of kernel ridge regression (KRR) under minimal assumptions. Our contributions are three-fold: 1) we analyze the role of key properties of the kernel, such as its spectral eigen-decay, the characteristics of the eigenfunctions, and the…

Cited by 1SourcePDFScholar
2024

Characterizing Overfitting in Kernel Ridgeless Regression Through the Eigenspectrum

ICML 2024poster

We derive new bounds for the condition number of kernel matrices, which we then use to enhance existing non-asymptotic test error bounds for kernel ridgeless regression in the over-parameterized regime for a fixed input dimension. For kernels with polynomial spectral decay, we recover the bound from…

Cited by 13SourcePDFScholar
2024

Energy-Guided Continuous Entropic Barycenter Estimation for General Costs

NeurIPS 2024spotlight

Optimal transport (OT) barycenters are a mathematically grounded way of averaging probability distributions while capturing their geometric properties. In short, the barycenter task is to take the average of a collection of probability distributions w.r.t. given OT discrepancies. We propose a novel…

2024

Neural Snowflakes: Universal Latent Graph Inference via Trainable Latent Geometries

ICLR 2024poster

The inductive bias of a graph neural network (GNN) is largely encoded in its specified graph. Latent graph inference relies on latent geometric representations to dynamically rewire or infer a GNN's graph to maximize the GNN's predictive downstream performance, but it lacks solid theoretical foundat…

Cited by 4SourcePDFScholar
2023

A Theoretical Analysis of the Test Error of Finite-Rank Kernel Ridge Regression

NeurIPS 2023poster

Existing statistical learning guarantees for general kernel regressors often yield loose bounds when used with finite-rank kernels. Yet, finite-rank kernels naturally appear in a number of machine learning problems, e.g. when fine-tuning a pre-trained deep neural network's last layer to adapt it to…

Cited by 10SourcePDFScholar
2022

Universal Approximation Under Constraints is Possible with Transformers

ICLR 2022spotlight

Many practical problems need the output of a machine learning model to satisfy a set of constraints, $K$. Nevertheless, there is no known guarantee that classical neural network architectures can exactly encode constraints while simultaneously achieving universality. We provide a quantitative cons…

Cited by 37SourcePDFScholar