← Search

Ievgen Redko

21 accepted papers

2026

CauKer: Classification Time Series Foundation Models Can Be Pretrained on Synthetic Data

ICLR 2026oral

Time series foundation models (TSFMs) have recently gained significant attention due to their strong zero-shot capabilities and widespread real-world applications. Such models typically require a computationally costly pretraining on large-scale, carefully curated collections of real-world sequences…

Cited by 0SourcecodeScholar
2026

Mantis: Lightweight Foundation Model for Time Series Classification

ICML 2026poster

While foundation models have revolutionized various domains, their application to time series classification remains rather under-explored, with existing literature predominantly focused on forecasting. To bridge this gap, we introduce \textbf{Mantis}, a transformer-based foundation model pre-traine…

Cited by 0SourceScholar
2026

Optimal Self-Consistency for Efficient Reasoning with Large Language Models

ICML 2026poster

Self-consistency (SC) is a widely-used test-time inference technique for improving performance in chain-of-thought reasoning. It consists of generating multiple responses, or ``samples," from a large language model (LLM) and selecting the most frequent answer. This procedure can naturally be viewed …

Cited by 0SourceScholar
2026

Vision Transformer Finetuning Benefits from Non-Smooth Components

ICML 2026poster

The smoothness of the transformer architecture has been extensively studied in the context of generalization, training stability, and adversarial robustness. However, its role in transfer learning remains poorly understood. In this paper, we analyze the ability of vision transformer components to ad…

Cited by 1SourceScholar
2025

From Alexnet to Transformers: Measuring the Non-linearity of Deep Neural Networks with Affine Optimal Transport

CVPR 2025poster

In the last decade, we have witnessed the introduction of several novel deep neural network (DNN) architectures exhibiting ever-increasing performance across diverse tasks. Explaining the upward trend of their performance, however, remains difficult as different DNN architectures of comparable depth…

2025

Zero-shot Model-based Reinforcement Learning using Large Language Models

ICLR 2025poster

The emerging zero-shot capabilities of Large Language Models (LLMs) have led to their applications in areas extending well beyond natural language processing tasks. In reinforcement learning, while LLMs have been extensively used in text-based environments, their integration with continuous state s…

2024

Analysing Multi-Task Regression via Random Matrix Theory with Application to Time Series Forecasting

NeurIPS 2024spotlight

In this paper, we introduce a novel theoretical framework for multi-task regression, applying random matrix theory to provide precise performance estimations, under high-dimensional, non-Gaussian data distributions. We formulate a multi-task optimization problem as a regularization technique to enab…

Cited by 2SourcePDFScholar
2024

Breaking isometric ties and introducing priors in Gromov-Wasserstein distances

AISTATS 2024poster

Gromov-Wasserstein distance has many applications in machine learning due to its ability to compare measures across metric spaces and its invariance to isometric transformations. However, in certain applications, this invariant property can be too flexible, thus undesirable. Moreover, the Gromov-Was…

2024

Leveraging Ensemble Diversity for Robust Self-Training in the Presence of Sample Selection Bias

AISTATS 2024poster

Self-training is a well-known approach for semi-supervised learning. It consists of iteratively assigning pseudo-labels to unlabeled data for which the model is confident and treating them as labeled examples. For neural networks, \texttt{softmax} prediction probabilities are often used as a confide…

2024

SAMformer: Unlocking the Potential of Transformers in Time Series Forecasting with Sharpness-Aware Minimization and Channel-Wise Attention

ICML 2024oral

Transformer-based architectures achieved breakthrough performance in natural language processing and computer vision, yet they remain inferior to simpler linear baselines in multivariate long-term forecasting. To better understand this phenomenon, we start by studying a toy linear forecasting proble…

2023

Unbalanced CO-optimal Transport

AAAI 2023technical

Optimal transport (OT) compares probability distributions by computing a meaningful alignment between their samples. CO-optimal transport (COOT) takes this comparison further by inferring an alignment between features as well. While this approach leads to better alignments and generalizes both OT an…

Cited by 22SourcePDFScholar
2022

Improving Few-Shot Learning through Multi-task Representation Learning Theory

ECCV 2022poster

"In this paper, we consider the framework of multi-task representation (MTR) learning where the goal is to use source tasks to learn a representation that reduces the sample complexity of solving a target task. We start by reviewing recent advances in MTR theory and show that they can provide novel…

2021

All of the Fairness for Edge Prediction with Optimal Transport

AISTATS 2021poster

Machine learning and data mining algorithms have been increasingly used recently to support decision-making systems in many areas of high societal importance such as healthcare, education, or security. While being very efficient in their predictive abilities, the deployed algorithms sometimes tend t…

2021

Deep Neural Networks Are Congestion Games: From Loss Landscape to Wardrop Equilibrium and Beyond

AISTATS 2021poster

The theoretical analysis of deep neural networks (DNN) is arguably among the most challenging research directions in machine learning (ML) right now, as it requires from scientists to lay novel statistical learning foundations to explain their behaviour in practice. While some success has been achie…

Cited by 4SourcePDFScholar
2020

A Swiss Army Knife for Minimax Optimal Transport

ICML 2020poster

The Optimal transport (OT) problem and its associated Wasserstein distance have recently become a topic of great interest in the machine learning community. However, the underlying optimization problem is known to have two major restrictions: (i) it largely depends on the choice of the cost function…

2020

Margin-aware Adversarial Domain Adaptation with Optimal Transport

ICML 2020poster

In this paper, we propose a new theoretical analysis of unsupervised domain adaptation that relates notions of large margin separation, adversarial learning and optimal transport. This analysis generalizes previous work on the subject by providing a bound on the target margin violation rate, thus re…

2019

Optimal Transport for Multi-source Domain Adaptation under Target Shift

AISTATS 2019poster

In this paper, we tackle the problem of reducing discrepancies between multiple domains, i.e. multi-source domain adaptation, and consider it under the target shift assumption: in all domains we aim to solve a classification problem with the same output classes, but with different labels proportions…

2018

Revisiting $(\epsilon, \gamma, \tau)$-similarity learning for domain adaptation

NeurIPS 2018spotlight

Similarity learning is an active research area in machine learning that tackles the problem of finding a similarity function tailored to an observable data sample in order to achieve efficient classification. This learning scenario has been generally formalized by the means of a $(\epsilon, \gamma,…

Cited by 0SourcePDFScholar
2017

Co-clustering through Optimal Transport

ICML 2017poster

In this paper, we present a novel method for co-clustering, an unsupervised learning approach that aims at discovering homogeneous groups of data instances and features by grouping them simultaneously. The proposed method uses the entropy regularized optimal transport between empirical measures defi…

Cited by 59SourcePDFScholar