← Search

Tom Sercu

15 accepted papers

2022

Learning inverse folding from millions of predicted structures

ICML 2022oral

We consider the problem of predicting a protein sequence from its backbone atom coordinates. Machine learning approaches to this problem to date have been limited by the number of available experimentally determined protein structures. We augment training data by nearly three orders of magnitude by…

2021

Language models enable zero-shot prediction of the effects of mutations on protein function

NeurIPS 2021poster

Modeling the effect of sequence variation on function is a fundamental problem for understanding and designing proteins. Since evolution encodes information about function into patterns in protein sequences, unsupervised models of variant effects can be learned from sequence data. The approach to da…

2021

Transformer protein language models are unsupervised structure learners

ICLR 2021poster

Unsupervised contact prediction is central to uncovering physical, structural, and functional constraints for protein structure determination and design. For decades, the predominant approach has been to infer evolutionary constraints from a set of related sequences. In the past year, protein langua…

Cited by 380SourcePDFScholar
2019

Adversarial Semantic Alignment for Improved Image Captions

CVPR 2019poster

In this paper, we study image captioning as a conditional GAN training, proposing both a context-aware LSTM captioner and co-attentive discriminator, which enforces semantic alignment between images and captions. We empirically focus on the viability of two training methods: Self-critical Sequence T…

Cited by 46PDFScholar
2019

Big-Little Net: An Efficient Multi-Scale Feature Representation for Visual and Speech Recognition

ICLR 2019poster

In this paper, we propose a novel Convolutional Neural Network (CNN) architecture for learning multi-scale feature representations with good tradeoffs between speed and accuracy. This is achieved by using a multi-branch network, which has different computational complexity at different branches with…

2019

Sobolev Descent

AISTATS 2019poster

We study a simplification of GAN training: the problem of transporting particles from a source to a target distribution. Starting from the Sobolev GAN critic, part of the gradient regularized GAN family, we show a strong relation with Optimal Transport (OT). Specifically with the less popular *dyna…

Cited by 48SourcePDFScholar
2019

Sobolev Independence Criterion

NeurIPS 2019poster

We propose the Sobolev Independence Criterion (SIC), an interpretable dependency measure between a high dimensional random variable X and a response variable Y. SIC decomposes to the sum of feature importance scores and hence can be used for nonlinear feature selection. SIC can be seen as a gradient…

2017

Fisher GAN

NeurIPS 2017poster

Generative Adversarial Networks (GANs) are powerful models for learning complex distributions. Stable training of GANs has been addressed in many recent works which explore different metrics between distributions. In this paper we introduce Fisher GAN that fits within the Integral Probability Metric…

2017

Knowledge distillation across ensembles of multilingual models for low-resource languages

ICASSP 2017accepted

This paper investigates the effectiveness of knowledge distillation in the context of multilingual models. We show that with knowledge distillation, Long Short-Term Memory(LSTM) models can be used to train standard feed-forward Deep Neural Network (DNN) models for a variety of low-resource languages…

Cited by 0SourceScholar
2017

Network architectures for multilingual speech representation learning

ICASSP 2017accepted

Multilingual (ML) representations play a key role in building speech recognition systems for low resource languages. The IARPA sponsored BABEL program focuses on building speech recognition (ASR) and keyword search (KWS) systems in over 24 languages with limited training data. The most common mechan…

Cited by 0SourceScholar
2016

Very deep multilingual convolutional neural networks for LVCSR

ICASSP 2016accepted

Convolutional neural networks (CNNs) are a standard component of many current state-of-the-art Large Vocabulary Continuous Speech Recognition (LVCSR) systems. However, CNNs in LVCSR have not kept pace with recent advances in other domains where deeper neural networks provide superior performance. In…

Cited by 227SourceScholar