Enhancing variational generation through self-decomposition

Asperti, Andrea; Bugo, Laura; Filippini, Daniele

doi:10.1109/ACCESS.2022.3185654

Computer Science > Computer Vision and Pattern Recognition

arXiv:2202.02738 (cs)

[Submitted on 6 Feb 2022 (v1), last revised 14 Jul 2022 (this version, v2)]

Title:Enhancing variational generation through self-decomposition

Authors:Andrea Asperti, Laura Bugo, Daniele Filippini

View PDF

Abstract:In this article we introduce the notion of Split Variational Autoencoder (SVAE), whose output $\hat{x}$ is obtained as a weighted sum $\sigma \odot \hat{x_1} + (1-\sigma) \odot \hat{x_2}$ of two generated images $\hat{x_1},\hat{x_2}$, and $\sigma$ is a {\em learned} compositional map. The composing images $\hat{x_1},\hat{x_2}$, as well as the $\sigma$-map are automatically synthesized by the model. The network is trained as a usual Variational Autoencoder with a negative loglikelihood loss between training and reconstructed images. No additional loss is required for $\hat{x_1},\hat{x_2}$ or $\sigma$, neither any form of human tuning. The decomposition is nondeterministic, but follows two main schemes, that we may roughly categorize as either \say{syntactic} or \say{semantic}. In the first case, the map tends to exploit the strong correlation between adjacent pixels, splitting the image in two complementary high frequency sub-images. In the second case, the map typically focuses on the contours of objects, splitting the image in interesting variations of its content, with more marked and distinctive features. In this case, according to empirical observations, the Fréchet Inception Distance (FID) of $\hat{x_1}$ and $\hat{x_2}$ is usually lower (hence better) than that of $\hat{x}$, that clearly suffers from being the average of the former. In a sense, a SVAE forces the Variational Autoencoder to make choices, in contrast with its intrinsic tendency to {\em average} between alternatives with the aim to minimize the reconstruction loss towards a specific sample. According to the FID metric, our technique, tested on typical datasets such as Mnist, Cifar10 and CelebA, allows us to outperform all previous purely variational architectures (not relying on normalization flows).

Subjects:	Computer Vision and Pattern Recognition (cs.CV); Neural and Evolutionary Computing (cs.NE)
MSC classes:	68T07 Artificial neural networks and deep learning
ACM classes:	I.3.3
Cite as:	arXiv:2202.02738 [cs.CV]
	(or arXiv:2202.02738v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2202.02738
Journal reference:	IEEE Access, vol. 10, pp. 67510-67520, 2022
Related DOI:	https://doi.org/10.1109/ACCESS.2022.3185654

Submission history

From: Andrea Asperti [view email]
[v1] Sun, 6 Feb 2022 08:49:21 UTC (3,493 KB)
[v2] Thu, 14 Jul 2022 10:57:30 UTC (4,855 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Enhancing variational generation through self-decomposition

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Enhancing variational generation through self-decomposition

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators