RESEARCH PAPER

PaliGemma: A versatile 3B VLM for transfer

View all authors

Lucas Beyer; Andreas Steiner; André Susano Pinto; Alexander Kolesnikov; Xiao Wang; Daniel Salz; Maxim Neumann; Ibrahim Alabdulmohsin; Michael Tschannen; Emanuele Bugliarello; Thomas Unterthiner; Daniel Keysers; Skanda Koppula; Fangyu Liu; Adam Grycner; Alexey Gritsenko; Neil Houlsby; Manoj Kumar; Keran Rong; Julian Eisenschlos; Rishabh Kabra; Matthias Bauer; Matko Bošnjak; Xi Chen; Matthias Minderer; Paul Voigtlaender; Ioana Bica; Ivana Balazevic; Joan Puigcerver; Pinelopi Papalampidi; Olivier Henaff; Xi Xiong; Radu Soricut; Jeremiah Harmsen; Xiaohua Zhai

Classification

View four quadrants
Major category
Components of WAMs
Architecture
Not applicable
Prediction paradigm
Not applicable
Source review status
Verified from primary sources

Category review. PaliGemma is a general visual-language backbone connecting SigLIP image tokens to Gemma with prefix attention and task-adaptation training. The manuscript explicitly names it as a reusable VLM; no robot action or future-state prediction is required for this component role. Reading evidence

AT A GLANCE

Contribution

A contribution summary has not been added yet.

Abstract

An abstract has not been added yet.

Affiliations

Not listed in the collection.

BibTeX