← Search

Ilja Kuzborskij

13 accepted papers

2026

DAG-Math: Graph-Guided Mathematical Reasoning in LLMs

ICLR 2026poster

Large Language Models (LLMs) demonstrate strong performance on mathematical problems when prompted with Chain-of-Thought (CoT), yet it remains unclear whether this success stems from search, rote procedures, or rule-consistent reasoning. To address this, we propose modeling CoT as a certain rule-bas…

Cited by 0SourcecodeScholar
2024

To Believe or Not to Believe Your LLM: Iterative Prompting for Estimating Epistemic Uncertainty

NeurIPS 2024poster

We explore uncertainty quantification in large language models (LLMs), with the goal to identify when uncertainty in responses given a query is large. We simultaneously consider both epistemic and aleatoric uncertainties, where the former comes from the lack of knowledge about the ground truth (such…

Cited by 11SourcePDFScholar
2023

Mixture Weight Estimation and Model Prediction in Multi-source Multi-target Domain Adaptation

NeurIPS 2023poster

We consider a problem of learning a model from multiple sources with the goal to perform well on a new target distribution. Such problem arises in learning with data collected from multiple sources (e.g. crowdsourcing) or learning in distributed systems, where the data can be highly heterogeneous.…

Cited by 3SourcePDFScholar
2021

Confident Off-Policy Evaluation and Selection through Self-Normalized Importance Weighting

AISTATS 2021poster

We consider off-policy evaluation in the contextual bandit setting for the purpose of obtaining a robust off-policy selection strategy, where the selection strategy is evaluated based on the value of the chosen policy in a set of proposal (target) policies. We propose a new method to compute a lower…

2021

On the Role of Optimization in Double Descent: A Least Squares Study

NeurIPS 2021poster

Empirically it has been observed that the performance of deep neural networks steadily improves with increased model size, contradicting the classical view on overfitting and generalization. Recently, the double descent phenomenon has been proposed to reconcile this observation with theory, suggesti…

Cited by 15SourcePDFScholar
2021

Stability & Generalisation of Gradient Descent for Shallow Neural Networks without the Neural Tangent Kernel

NeurIPS 2021poster

We revisit on-average algorithmic stability of Gradient Descent (GD) for training overparameterised shallow neural networks and prove new generalisation and excess risk bounds without the Neural Tangent Kernel (NTK) or Polyak-Łojasiewicz (PL) assumptions. In particular, we show ora…

Cited by 31SourcePDFScholar
2020

PAC-Bayes Analysis Beyond the Usual Bounds

NeurIPS 2020poster

We focus on a stochastic learning model where the learner observes a finite set of training examples and the output of the learning process is a data-dependent distribution over a space of hypotheses. The learned data-dependent distribution is then used to make randomized predictions, and the high-l…

Cited by 100SourcePDFScholar
2016

When Naive Bayes Nearest Neighbors Meet Convolutional Neural Networks

CVPR 2016poster

Since Convolutional Neural Networks (CNNs) have become the leading learning paradigm in visual recognition, Naive Bayes Nearest Neighbor (NBNN)-based classifiers have lost momentum in the community. This is because (1) such algorithms cannot use CNN activations as input features; (2) they cannot be…

Cited by 35PDFScholar