← Search

Igor Melnyk

15 accepted papers

2025

EpMAN: Episodic Memory AttentioN for Generalizing to Longer Contexts

ACL 2025long

Recent advances in Large Language Models (LLMs) have yielded impressive successes on many language tasks. However, efficient processing of long contexts using LLMs remains a significant challenge. We introduce **EpMAN** – a method for processing long contexts in an episodic memory module while holis…

2024

Distributional Preference Alignment of LLMs via Optimal Transport

NeurIPS 2024poster

Current LLM alignment techniques use pairwise human preferences at a sample level, and as such, they do not imply an alignment on the distributional level. We propose in this paper Alignment via Optimal Transport (AOT), a novel method for distributional preference alignment of LLMs. AOT aligns LLMs…

Cited by 15SourcePDFScholar
2024

Larimar: Large Language Models with Episodic Memory Control

ICML 2024poster

Efficient and accurate updating of knowledge stored in Large Language Models (LLMs) is one of the most pressing research challenges today. This paper presents Larimar - a novel, brain-inspired architecture for enhancing LLMs with a distributed episodic memory. Larimar's memory allows for dynamic, on…

2024

Risk Aware Benchmarking of Large Language Models

ICML 2024poster

We propose a distributional framework for benchmarking socio-technical risks of foundation models with quantified statistical significance. Our approach hinges on a new statistical relative testing based on first and second order stochastic dominance of real random variables. We show that the second…

Cited by 1SourcePDFScholar
2023

Reprogramming Pretrained Language Models for Antibody Sequence Infilling

ICML 2023poster

Antibodies comprise the most versatile class of binding molecules, with numerous applications in biomedicine. Computational design of antibodies involves generating novel and diverse sequences, while maintaining structural consistency. Unique to antibodies, designing the complementarity-determining…

2021

Fold2Seq: A Joint Sequence(1D)-Fold(3D) Embedding-based Generative Model for Protein Design

ICML 2021spotlight

Designing novel protein sequences for a desired 3D topological fold is a fundamental yet non-trivial task in protein engineering. Challenges exist due to the complex sequence–fold relationship, as well as the difficulties to capture the diversity of the sequences (therefore structures and functions)…

2021

ReGen: Reinforcement Learning for Text and Knowledge Base Generation using Pretrained Language Models

EMNLP 2021main

Automatic construction of relevant Knowledge Bases (KBs) from text, and generation of semantically meaningful text from KBs are both long-standing goals in Machine Learning. In this paper, we present ReGen, a bidirectional generation of text and graph leveraging Reinforcement Learning to improve per…

2021

Tabular Transformers for Modeling Multivariate Time Series

ICASSP 2021accepted

Tabular datasets are ubiquitous in data science applications. Given their importance, it seems natural to apply state-of-the-art deep learning algorithms in order to fully unlock their potential. Here we propose neural network models that represent tabular time series that can optionally leverage th…

Cited by 0SourceScholar
2020

Optimizing Mode Connectivity via Neuron Alignment

NeurIPS 2020poster

The loss landscapes of deep neural networks are not well understood due to their high nonconvexity. Empirically, the local minima of these loss functions can be connected by a learned curve in model space, along which the loss remains nearly constant; a feature known as mode connectivity. Yet, curre…

2019

Adversarial Semantic Alignment for Improved Image Captions

CVPR 2019poster

In this paper, we study image captioning as a conditional GAN training, proposing both a context-aware LSTM captioner and co-attentive discriminator, which enforces semantic alignment between images and captions. We empirically focus on the viability of two training methods: Self-critical Sequence T…

Cited by 46PDFScholar
2019

Estimating Information Flow in Deep Neural Networks

ICML 2019oral

We study the estimation of the mutual information I(X;T_$\ell$) between the input X to a deep neural network (DNN) and the output vector T_$\ell$ of its $\ell$-th hidden layer (an “internal representation”). Focusing on feedforward networks with fixed weights and noisy internal representations, we d…

Cited by 181SourcePDFScholar