← Search

Dmitry Vetrov

31 accepted papers

2026

Can Stationary Distributions of Scale-Invariant Neural Networks Be Described by the Thermodynamics of an Ideal Gas?

IJCAI 2026

Understanding the training dynamics of deep neural networks remains a major open problem, with physics-inspired approaches offering promising insights. Building on this perspective, we develop a thermodynamic framework to describe the stationary distributions of stochastic gradient descent (SGD) wit

Cited by 0Scholar
2026

GeomMotif: A Benchmark for Arbitrary Geometric Preservation in Protein Generation

ICLR 2026poster

Motif scaffolding in protein design involves generating complete protein structures while preserving the 3D geometry of designated structural fragments, analogous to image outpainting in computer vision. Current benchmarks focus on functional motifs, leaving general geometric preservation capabiliti…

Cited by 0SourceScholar
2026

Guided Star-Shaped Masked Diffusion

ICML 2026poster

The performance of pre-trained masked diffusion models is often constrained by their sampling procedure, which makes decisions irreversible and struggles in low-step generation regimes. We introduce a novel sampling algorithm that works with pre-trained models and, after a lightweight fine-tuning of…

Cited by 0SourceScholar
2026

One-step Optimal Transport via Regularized Distribution Matching Distillation

ICML 2026poster

Unpaired domain translation remains a challenging task due to the need of finding a balance between faithfulness and realism. In this paper, we propose a method called Regularized Distribution Matching Distillation (RDMD) that combines the best properties of Optimal Transport (OT) and diffusion-base…

Cited by 0SourceScholar
2026

Smoothie: Smoothing Diffusion on Token Embeddings for Text Generation

ICML 2026poster

Diffusion models have achieved state-of-the-art performance in generating images, audio, and video, but their adaptation to text remains challenging due to its discrete nature. Prior approaches either apply Gaussian diffusion in continuous latent spaces, which inherits semantic structure but struggl…

Cited by 0SourcecodeScholar
2026

Streaming Generation of Co-Speech Gestures via Accelerated Rolling Diffusion

AAAI 2026technical

Generating co-speech gestures in real time requires both temporal coherence and efficient sampling. We introduce a novel framework for streaming gesture generation that extends Rolling Diffusion models with structured progressive noise scheduling, enabling seamless long-sequence motion synthesis whi

Cited by 0SourcePDFScholar
2025

Compressed and Smooth Latent Space for Text Diffusion Modeling

NeurIPS 2025poster

Autoregressive language models dominate modern text generation, yet their sequential nature introduces fundamental limitations: decoding is slow, and maintaining global coherence remains challenging. Diffusion models offer a promising alternative by enabling parallel generation and flexible control;…

Cited by 0SourcecodeScholar
2025

Diffusion on Language Model Encodings for Protein Sequence Generation

ICML 2025poster

Protein *sequence* design has seen significant advances through discrete diffusion and autoregressive approaches, yet the potential of continuous diffusion remains underexplored. Here, we present *DiMA*, a latent diffusion framework that operates on protein language model representations. Through sy…

2025

SDE Matching: Scalable and Simulation-Free Training of Latent Stochastic Differential Equations

ICML 2025poster

The Latent Stochastic Differential Equation (SDE) is a powerful tool for time series and sequence modeling. However, training Latent SDEs typically relies on adjoint sensitivity methods, which depend on simulation and backpropagation through approximate SDE solutions, which limit scalability. In thi…

Cited by 0SourcePDFScholar
2025

TEncDM: Understanding the Properties of the Diffusion Model in the Space of Language Model Encodings

AAAI 2025technical

This paper presents the Text Encoding Diffusion Model (TEncDM), a novel approach to diffusion modeling that operates in the space of pre-trained language model encodings. In contrast to traditionally used embeddings, encodings integrate contextual information. In our approach, we also employ a trans…

2024

HairFastGAN: Realistic and Robust Hair Transfer with a Fast Encoder-Based Approach

NeurIPS 2024poster

Our paper addresses the complex task of transferring a hairstyle from a reference image to an input photo for virtual hair try-on. This task is challenging due to the need to adapt to various photo poses, the sensitivity of hairstyles, and the lack of objective metrics. The current state of the art…

2024

Neural Flow Diffusion Models: Learnable Forward Process for Improved Diffusion Modelling

NeurIPS 2024poster

Conventional diffusion models typically relies on a fixed forward process, which implicitly defines complex marginal distributions over latent variables. This can often complicate the reverse process’ task in learning generative trajectories, and results in costly inference for diffusion models. To…

Cited by 12SourcePDFScholar
2024

The Devil is in the Details: StyleFeatureEditor for Detail-Rich StyleGAN Inversion and High Quality Image Editing

CVPR 2024poster

The task of manipulating real image attributes through StyleGAN inversion has been extensively researched. This process involves searching latent variables from a well-trained StyleGAN generator that can synthesize a real image modifying these latent variables and then synthesizing an image with the…

2024

Where Do Large Learning Rates Lead Us?

NeurIPS 2024poster

It is generally accepted that starting neural networks training with large learning rates (LRs) improves generalization. Following a line of research devoted to understanding this effect, we conduct an empirical study in a controlled setting focusing on two questions: 1) how large an initial LR is r…

2023

MARS: Masked Automatic Ranks Selection in Tensor Decompositions

AISTATS 2023poster

Tensor decomposition methods have proven effective in various applications, including compression and acceleration of neural networks. At the same time, the problem of determining optimal decomposition ranks, which present the crucial parameter controlling the compressionaccuracy trade-off, is still…

2023

StyleDomain: Efficient and Lightweight Parameterizations of StyleGAN for One-shot and Few-shot Domain Adaptation

ICCV 2023poster

Domain adaptation of GANs is a problem of fine-tuning GAN models pretrained on a large dataset (e.g. StyleGAN) to a specific domain with few samples (e.g. painting faces, sketches, etc.). While there are many methods that tackle this problem in different ways, there are still many important question…

Cited by 9PDFcodeScholar
2020

Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile Critics

ICML 2020poster

The overestimation bias is one of the major impediments to accurate off-policy learning. This paper investigates a novel way to alleviate the overestimation bias in a continuous control setting. Our method—Truncated Quantile Critics, TQC,—blends three ideas: distributional representation of a critic…

2020

Greedy Policy Search: A Simple Baseline for Learnable Test-Time Augmentation

UAI 2020poster

Test-time data augmentation—averaging the predictions of a machine learning model across multiple augmented samples of data—is a widely used technique that improves the predictive performance. While many advanced learnable data augmentation techniques have emerged in recent years, they are focused o…

2020

Pitfalls of In-Domain Uncertainty Estimation and Ensembling in Deep Learning

ICLR 2020poster

Uncertainty estimation and ensembling methods go hand-in-hand. Uncertainty estimation is one of the main benchmarks for assessment of ensembling performance. At the same time, deep learning ensembles have provided state-of-the-art results in uncertainty estimation. In this work, we focus on in-domai…

Cited by 412SourceScholar
2019

Subspace Inference for Bayesian Deep Learning

UAI 2019poster

Bayesian inference was once a gold standard for learning with neural networks, providing accurate full predictive distributions and well calibrated uncertainty. However, scaling Bayesian inference techniques to deep neural networks is challenging due to the high dimensionality of the parameter space…

2019

Variance Networks: When Expectation Does Not Meet Your Expectations

ICLR 2019poster

Ordinary stochastic neural networks mostly rely on the expected values of their weights to make predictions, whereas the induced noise is mostly used to capture the uncertainty, prevent overfitting and slightly boost the performance through test-time averaging. In this paper, we introduce variance l…

2017

Spatially Adaptive Computation Time for Residual Networks

CVPR 2017poster

This paper proposes a deep learning architecture based on Residual Network that dynamically adjusts the number of executed layers for the regions of the image. This architecture is end-to-end trainable, deterministic and problem-agnostic. It is therefore applicable without any modifications to a wid…

Cited by 431PDFcodeScholar
2016

Breaking Sticks and Ambiguities with Adaptive Skip-gram

AISTATS 2016poster

The recently proposed Skip-gram model is a powerful method for learning high-dimensional word representations that capture rich semantic relationships between words. However, Skip-gram as well as most prior work on learning word representations does not take into account word ambiguity and maintain…

2015

Inferring M-Best Diverse Labelings in a Single One

ICCV 2015poster

We consider the task of finding M-best diverse solutions in a graphical model. In a previous work by Batra et al. an algorithmic approach for finding such solutions was proposed, and its usefulness was shown in numerous applications. Contrary to previous work we propose a novel formulation of the pr…

Cited by 51PDFScholar