Revision History
Multimodal Model · 7 revisions
Sizes are character counts of the article source. The signed number is the change from the previous revision.
Recent edit summaries
Detailed summaries recorded by editors. Generic maintenance summaries are omitted here; every recorded revision remains below. Dates describe the edit, not necessarily the event it covers.
- Bring the multimodal model survey current through July 2026: add native multimodal pretraining versus the adapter pattern, early-versus-late fusion at pretraining scale with the conflicting industry and research definitions, unified understand-and-generate designs, audio and video as first-class modalities with streaming and long-video tokenisation, benchmarks answerable without the image, and safety exposure from image and audio inputs
Version 7 · Aug 1, 2026, 11:21 AM
- Correction: remove unsupported product, benchmark, architecture, and compute claims; rebuild the overview around sourced modality definitions, architectures, training, evaluation, failure modes, and reporting practice.
Version 6 · Jul 29, 2026, 03:15 AM
- Added 6 contextual internal links
Version 5 · Jul 23, 2026, 03:02 PM