← Search

Tom Arodz

3 accepted papers

2021

Shapeshifter: a Parameter-efficient Transformer using Factorized Reshaped Matrices

NeurIPS 2021poster

Language models employ a very large number of trainable parameters. Despite being highly overparameterized, these networks often achieve good out-of-sample test performance on the original task and easily fine-tune to related tasks. Recent observations involving, for example, intrinsic dimension of…

2020

Approximation Capabilities of Neural ODEs and Invertible Residual Networks

ICML 2020poster

Recent interest in invertible models and normalizing flows has resulted in new architectures that ensure invertibility of the network model. Neural ODEs and i-ResNets are two recent techniques for constructing models that are invertible, but it is unclear if they can be used to approximate any conti…

Cited by 116SourcePDFScholar
2020

word2ket: Space-efficient Word Embeddings inspired by Quantum Entanglement

ICLR 2020spotlight

Deep learning natural language processing models often use vector word embeddings, such as word2vec or GloVe, to represent words. A discrete sequence of words can be much more easily integrated with downstream neural layers if it is represented as a sequence of continuous vectors. Also, semantic re…

Cited by 45SourcecodeScholar