← Search

Bryan Seybold

7 accepted papers

2024

VideoPoet: A Large Language Model for Zero-Shot Video Generation

ICML 2024oral

We present VideoPoet, a language model capable of synthesizing high-quality video from a large variety of conditioning signals. VideoPoet employs a decoder-only transformer architecture that processes multimodal inputs -- including images, videos, text, and audio. The training protocol follows that…

Cited by 257SourcePDFScholar
2022

Learning Audio-Video Modalities from Image Captions

ECCV 2022poster

"There has been a recent explosion of large-scale image-text datasets, as images with alt-text captions can be easily obtained online. Obtaining large-scale, high quality data for video in the form of text-video and text-audio pairs however, is more challenging. To close this gap we propose a new vi…

Cited by 109SourcePDFScholar
2020

Collapsed Amortized Variational Inference for Switching Nonlinear Dynamical Systems

ICML 2020poster

We propose an efficient inference method for switching nonlinear dynamical systems. The key idea is to learn an inference network which can be used as a proposal distribution for the continuous latent variables, while performing exact marginalization of the discrete latent variables. This allows us…

Cited by 33SourcePDFScholar
2018

Instance Embedding Transfer to Unsupervised Video Object Segmentation

CVPR 2018poster

We propose a method for unsupervised video object segmentation by transferring the knowledge encapsulated in image-based instance embedding networks. The instance embedding network produces an embedding vector for each pixel that enables identifying all pixels belonging to the same object. Though tr…

Cited by 132SourcePDFScholar
2018

Rethinking the Faster R-CNN Architecture for Temporal Action Localization

CVPR 2018poster

We propose TAL-Net, an improved approach to temporal action localization in video that is inspired by the Faster R-CNN object detection framework. TAL-Net addresses three key shortcomings of existing approaches: (1) we improve receptive field alignment using a multi-scale architecture that can accom…

Cited by 846SourcePDFScholar
2018

Unsupervised Video Object Segmentation with Motion-based Bilateral Networks

ECCV 2018poster

In this work, we study the unsupervised video object segmentation problem where moving objects are segmented without prior knowledge of these objects. First, we propose a motion-based bilateral network to estimate the background based on the motion pattern of non-object regions. The bilateral networ…

Cited by 157SourcePDFScholar
2017

CNN architectures for large-scale audio classification

ICASSP 2017accepted

Convolutional Neural Networks (CNNs) have proven very effective in image classification and show promise for audio. We use various CNN architectures to classify the soundtracks of a dataset of 70M training videos (5.24 million hours) with 30,871 video-level labels. We examine fully connected Deep Ne…

Cited by 3037SourceScholar