← Search

Georg Heigold

8 accepted papers

2025

Massive Sound Embedding Benchmark (MSEB)

NeurIPS 2025poster

Audio is a critical component of multimodal perception, and any truly intelligent system must demonstrate a wide range of auditory capabilities. These capabilities include transcription, classification, retrieval, reasoning, segmentation, clustering, reranking, and reconstruction. Fundamentally, eac…

Cited by 0SourcecodeScholar
2023

Video OWL-ViT: Temporally-consistent Open-world Localization in Video

ICCV 2023poster

We present an architecture and a training recipe that adapts pretrained open-world image models to localization in videos. Understanding the open visual world (without being constrained by fixed label spaces) is crucial for many real-world vision tasks. Contrastive pre-training on large image-text d…

Cited by 17PDFScholar
2022

Conditional Object-Centric Learning from Video

ICLR 2022poster

Object-centric representations are a promising path toward more systematic generalization by providing flexible abstractions upon which compositional world models can be built. Recent work on simple 2D and 3D datasets has shown that models with object-centric inductive biases can learn to segment an…

2021

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

ICLR 2021oral

While the Transformer architecture has become the de-facto standard for natural language processing tasks, its applications to computer vision remain limited. In vision, attention is either applied in conjunction with convolutional networks, or used to replace certain components of convolutional net…

2021

ViViT: A Video Vision Transformer

ICCV 2021poster

We present pure-transformer based models for video classification, drawing upon the recent success of such models in image classification. Our model extracts spatio-temporal tokens from the input video, which are then encoded by a series of transformer layers. In order to handle the long sequences o…

Cited by 2888PDFcodeScholar
2020

Object-Centric Learning with Slot Attention

NeurIPS 2020spotlight

Learning object-centric representations of complex scenes is a promising step towards enabling efficient abstract reasoning from low-level perceptual features. Yet, most deep learning approaches learn distributed representations that do not capture the compositional properties of natural scenes. In…

2015

A Gaussian Mixture Model layer jointly optimized with discriminative features within a Deep Neural Network architecture

ICASSP 2015accepted

This article proposes and evaluates a Gaussian Mixture Model (GMM) represented as the last layer of a Deep Neural Network (DNN) architecture and jointly optimized with all previous layers using Asynchronous Stochastic Gradient Descent (ASGD). The resulting “Deep GMM” architecture was investigated wi…

Cited by 0SourceScholar