Jet: A Modern Transformer-Based Normalizing Flow

Kolesnikov, Alexander; Pinto, André Susano; Tschannen, Michael

Computer Science > Computer Vision and Pattern Recognition

arXiv:2412.15129 (cs)

[Submitted on 19 Dec 2024]

Title:Jet: A Modern Transformer-Based Normalizing Flow

Authors:Alexander Kolesnikov, André Susano Pinto, Michael Tschannen

View PDF HTML (experimental)

Abstract:In the past, normalizing generative flows have emerged as a promising class of generative models for natural images. This type of model has many modeling advantages: the ability to efficiently compute log-likelihood of the input data, fast generation and simple overall structure. Normalizing flows remained a topic of active research but later fell out of favor, as visual quality of the samples was not competitive with other model classes, such as GANs, VQ-VAE-based approaches or diffusion models. In this paper we revisit the design of the coupling-based normalizing flow models by carefully ablating prior design choices and using computational blocks based on the Vision Transformer architecture, not convolutional neural networks. As a result, we achieve state-of-the-art quantitative and qualitative performance with a much simpler architecture. While the overall visual quality is still behind the current state-of-the-art models, we argue that strong normalizing flow models can help advancing research frontier by serving as building components of more powerful generative models.

Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as:	arXiv:2412.15129 [cs.CV]
	(or arXiv:2412.15129v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2412.15129

Submission history

From: Michael Tschannen [view email]
[v1] Thu, 19 Dec 2024 18:09:42 UTC (3,932 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Jet: A Modern Transformer-Based Normalizing Flow

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Jet: A Modern Transformer-Based Normalizing Flow

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators