RESEARCH PAPERYear 2023

Scalable Diffusion Models with Transformers

William Peebles; Saining Xie

Classification

View four quadrants
Major category
Components of WAMs
Architecture
Not applicable
Prediction paradigm
Not applicable
Source review status
Verified from primary sources

Category review. DiT explicitly defines the transformer denoising backbone and adaLN-Zero conditioning over VAE latent patches. This is a reusable neural backbone architecture subsequently instantiated by video/WAM models, rather than a task-specific training wrapper. Reading evidence

AT A GLANCE

Contribution

A contribution summary has not been added yet.

Abstract

An abstract has not been added yet.

Affiliations

Not listed in the collection.

BibTeX