RESEARCH PAPERYear 2025

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models

Weixin Liang; Lili Yu; Liang Luo; Srinivasan Iyer; Ning Dong; Chunting Zhou; Gargi Ghosh; Mike Lewis; Wen-tau Yih; Luke Zettlemoyer; Xi Victoria Lin

Classification

View four quadrants
Major category
Components of WAMs
Architecture
Not applicable
Prediction paradigm
Not applicable
Source review status
Verified from primary sources

Category review. Mixture-of-Transformers defines modality-specific projections/normalization/FFNs coupled through shared attention. This is a reusable multimodal backbone architecture, explicitly referenced by the manuscript for WAM information flow; it is not itself a robot control method. Reading evidence

AT A GLANCE

Contribution

A contribution summary has not been added yet.

Abstract

An abstract has not been added yet.

Affiliations

Not listed in the collection.

BibTeX