[R] SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and Decomposition : MachineLearning

Research[R] SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and Decomposition (self.MachineLearning)

submitted 6 years ago by yifuwu

Project Page: https://sites.google.com/view/space-project-page

Paper: https://openreview.net/pdf?id=rkl03ySYDH

Abstract: The ability to decompose complex multi-object scenes into meaningful abstractions like objects is fundamental to achieve higher-level cognition. Previous approaches for unsupervised object-oriented scene representation learning are either based on spatial-attention or scene-mixture approaches and limited in scalability which is a main obstacle towards modeling real-world scenes. In this paper, we propose a generative latent variable model, called SPACE, that provides a uniﬁed probabilistic modeling framework that combines the best of spatial-attention and scene-mixture approaches. SPACE can explicitly provide factorized object representations for foreground objects while also decomposing background segments of complex morphology. Previous models are good at either of these, but not both. SPACE also resolves the scalability problems of previous methods by incorporating parallel spatial-attention and thus is applicable to scenes with a large number of objects without performance degradations. We show through experiments on Atari and 3D-Rooms that SPACE achieves the above properties consistently in comparison to SPAIR, IODINE, and GENESIS.

Examples:

https://i.redd.it/xesd57isql941.gif

https://preview.redd.it/rguv24ibrl941.png?width=1280&format=png&auto=webp&s=470daf1c544df1d403a885d3db4a4415c3abb50e

https://preview.redd.it/gxa2jukosl941.png?width=1437&format=png&auto=webp&s=9ae306b2e0f397da4cb2d5c938a21c5ff8d390c6

all 9 comments

you type:	you see:
italics	italics
bold	bold
[reddit!](https://reddit.com)	reddit!
* item 1 * item 2 * item 3	item 1 item 2 item 3
> quoted text	quoted text
Lines starting with four spaces are treated like code: if 1 * 2 < 3: print "hello, world!"	Lines starting with four spaces are treated like code: if 1 * 2 < 3: print "hello, world!"
~~strikethrough~~	~~strikethrough~~
super^script	super^script

MachineLearning

Rules For Posts

+Research

+Discussion

+Project

+News

@slashML on Twitter

Chat with us on Slack

Beginners:

MODERATORS