← Search

Dumitru Erhan

15 accepted papers

2023

Phenaki: Variable Length Video Generation from Open Domain Textual Descriptions

ICLR 2023poster

We present Phenaki, a model capable of realistic video synthesis given a sequence of textual prompts. Generating videos from text is particularly challenging due to the computational cost, limited quantities of high quality text-video data and variable length of videos. To address these issues, we i…

Cited by 440SourcePDFScholar
2023

StoryBench: A Multifaceted Benchmark for Continuous Story Visualization

NeurIPS 2023poster

Generating video stories from text prompts is a complex task. In addition to having high visual quality, videos need to realistically adhere to a sequence of text prompts whilst being consistent throughout the frames. Creating a benchmark for video generation requires data annotated over time, which…

2022

Information Prioritization through Empowerment in Visual Model-based RL

ICLR 2022poster

Model-based reinforcement learning (RL) algorithms designed for handling complex visual observations typically learn some sort of latent state representation, either explicitly or implicitly. Standard methods of this sort do not distinguish between functionally relevant aspects of the state and irre…

Cited by 33SourcePDFScholar
2020

Model Based Reinforcement Learning for Atari

ICLR 2020spotlight

Model-free reinforcement learning (RL) can be used to learn effective policies for complex tasks, such as Atari games, even from image observations. However, this typically requires very large amounts of interaction -- substantially more, in fact, than a human would need to learn the same games. How…

Cited by 1127SourcecodeScholar
2020

SurfelGAN: Synthesizing Realistic Sensor Data for Autonomous Driving

CVPR 2020oral

Autonomous driving system development is critically dependent on the ability to replay complex and diverse traffic scenarios in simulation. In such scenarios, the ability to accurately simulate the vehicle sensors such as cameras, lidar or radar is hugely helpful. However, current sensor simulators…

Cited by 125PDFScholar
2020

VideoFlow: A Conditional Flow-Based Model for Stochastic Video Generation

ICLR 2020poster

Generative models that can model and predict sequences of future events can, in principle, learn to capture complex real-world phenomena, such as physical interactions. However, a central challenge in video prediction is that the future is highly uncertain: a sequence of past observations of events…

Cited by 123SourcecodeScholar
2019

A Benchmark for Interpretability Methods in Deep Neural Networks

NeurIPS 2019poster

We propose an empirical measure of the approximate accuracy of feature importance estimates in deep neural networks. Our results across several large-scale image classification datasets show that many popular interpretability methods produce estimates of feature importance that are not better than a…

2019

High Fidelity Video Prediction with Large Stochastic Recurrent Neural Networks

NeurIPS 2019poster

Predicting future video frames is extremely challenging, as there are many factors of variation that make up the dynamics of how frames change through time. Previously proposed solutions require complex inductive biases inside network architectures with highly specialized computation, including segm…

Cited by 172SourcePDFScholar
2018

Hierarchical Long-term Video Prediction without Supervision

ICML 2018oral

Much of recent research has been devoted to video prediction and generation, yet most of the previous works have demonstrated only limited success in generating videos on short-term horizons. The hierarchical video prediction method by Villegas et al. (2017) is an example of a state-of-the-art metho…

Cited by 164SourcePDFScholar
2018

Learning how to explain neural networks: PatternNet and PatternAttribution

ICLR 2018poster

DeConvNet, Guided BackProp, LRP, were invented to better understand deep neural networks. We show that these methods do not produce the theoretically correct explanation for a linear model. Yet they are used on multi-layer networks with millions of parameters. This is a cause for concern since linea…

Cited by 433SourcePDFScholar
2018

Stochastic Variational Video Prediction

ICLR 2018poster

Predicting the future in real-world settings, particularly from raw sensory observations such as images, is exceptionally challenging. Real-world events can be stochastic and unpredictable, and the high dimensionality and complexity of natural images requires the predictive model to build an intrica…

Cited by 671SourcePDFScholar
2017

Unsupervised Pixel-Level Domain Adaptation With Generative Adversarial Networks

CVPR 2017oral

Collecting well-annotated image datasets to train modern machine learning algorithms is prohibitively expensive for many tasks. One appealing alternative is rendering synthetic data where ground-truth annotations are generated automatically. Unfortunately, models trained purely on rendered images fa…

Cited by 2021PDFScholar
2016

Domain Separation Networks

NeurIPS 2016poster

The cost of large scale data collection and annotation often makes the application of machine learning algorithms to new tasks or datasets prohibitively expensive. One approach circumventing this cost is training models on synthetic data where annotations are provided automatically. Despite their ap…

2015

Going Deeper With Convolutions

CVPR 2015poster

We propose a deep convolutional neural network architecture codenamed Inception that achieves the new state of the art for classification and detection in the ImageNet Large-Scale Visual Recognition Challenge 2014 (ILSVRC2014). The main hallmark of this architecture is the improved utilization of th…

Cited by 66966SourcePDFScholar