RESEARCH PAPERYear 2021

Learning Transferable Visual Models From Natural Language Supervision

Alec Radford; Jong Wook Kim; Chris Hallacy; Aditya Ramesh; Gabriel Goh; Sandhini Agarwal; Girish Sastry; Amanda Askell; Pamela Mishkin; Jack Clark; Gretchen Krueger; Ilya Sutskever

Classification

View four quadrants
Major category
Components of WAMs
Architecture
Not applicable
Prediction paradigm
Not applicable
Source review status
Explicit in survey

Category review. CLIP learns transferable image/text encoders and aligned normalized embeddings. The manuscript explicitly uses these as global text conditioning and visual semantic representation components. Reading evidence

AT A GLANCE

Contribution

A contribution summary has not been added yet.

Abstract

An abstract has not been added yet.

Affiliations

Not listed in the collection.

BibTeX