SIMONe: View-Invariant, Temporally-Abstracted Object Representations via Unsupervised Video Decomposition
To help agents reason about scenes in terms of their building blocks, we wish to extract the compositional structure of any given scene (in particular, the configuration and characteristics of objects comprising the scene). This problem is especially difficult when scene structure needs to be inferr…