Gemini Omni
Last edited
Fact-checked
In review queue
Sources
11 citations
Revision
v1 · 1,725 words
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Gemini Omni is a family of proprietary multimodal AI models from Google DeepMind for generating and editing media. Google announced the family at Google I/O on May 19, 2026, and released its first member, Gemini Omni Flash, on the same day. The initial product focuses on AI video generation: it accepts combinations of text, images, video, and, on some product surfaces, audio references, then produces short video with synchronized audio. Google describes Omni as a meeting point between the reasoning capabilities of the Gemini family and its generative media work.[1][2]
The released model's official name is Gemini Omni Flash. Its public Gemini Developer API identifier is gemini-omni-flash-preview. "Gemini Omni" names the broader family rather than a separate generally available model endpoint. Google has discussed future Omni output modalities and a higher-capability Pro variant, but as of July 24, 2026, it has not released an Omni Pro endpoint or announced a release date for one.[1][10]
Naming and release history
| Date | Development |
|---|---|
| May 19, 2026 | Google announced Gemini Omni and began rolling out Gemini Omni Flash in the Gemini app and Google Flow for Google AI Plus, Pro, and Ultra subscribers. It also started a free rollout through YouTube Shorts and YouTube Create.[1][2] |
| June 30, 2026 | The Gemini Developer API added gemini-omni-flash-preview in public preview through the Interactions API.[5] |
| July 16, 2026 | Google began rolling Gemini Omni into Google Vids for Google AI Pro and Ultra subscribers and Google Workspace business customers.[9] |
At launch, Google said Omni would eventually support outputs such as images and audio. The available Flash endpoint instead returns video, while some enterprise documentation lists text output alongside it. This distinction matters because "any input to any output" describes Google's planned family direction, not the complete output set of the June 2026 developer preview.[1][8]
Architecture and training
The Gemini Omni Flash model card identifies it as a transformer-based model with native support for text, vision, video, and audio inputs. Google says it trained the model on text, image, audio, and video data, with captions at several levels of detail for audio and video examples. Training videos went through semantic deduplication and filters for safety, compliance, and quality.[2]
Google trained the model on Tensor Processing Units using JAX and Pathways. It has not disclosed the parameter count, detailed network topology, training-compute budget, dataset size, or whether the production system routes requests among distinct media components. Google's launch description says Omni combines Gemini intelligence with generative media models, but the published model card does not describe a router or mixture-of-experts design.[1][2]
The model is proprietary and distributed through hosted Google products and APIs. Google has not released model weights or a technical paper that would allow independent inspection of its architecture or training corpus.
Inputs, outputs, and developer specifications
The Gemini Developer API and Google enterprise documentation expose different limits for the same preview identifier. These figures apply to their respective access surfaces rather than defining one universal configuration.[3][8]
| Property | Gemini Developer API | Google Enterprise Agent Platform |
|---|---|---|
| Model ID | gemini-omni-flash-preview | gemini-omni-flash-preview |
| Release stage | Public preview | Preview |
| Maximum input context | 1,048,576 tokens | 131,072 tokens |
| Documented output limit | Video lasting 3-10 seconds | 57,920 output tokens |
| Video format | 720p at 24 frames per second | 720p; 16:9 or 9:16 |
| Main supported tasks | Text-to-video, image-to-video, reference-to-video, and video editing | Text-to-video, reference-to-video, sound generation, and video editing |
The Developer API model page lists text, image, and video input, with videos up to 10 seconds for editing. Its detailed guide says direct audio-reference upload is not yet supported through that API, even though the model card describes audio as an architectural input modality and consumer launch material demonstrated voice references. Modality support therefore depends on the product and API version.[1][2][3][4]
Omni Flash can generate a clip from a text description, animate one or more still images, or alter an uploaded or previously generated clip. Developers can specify landscape 16:9 or portrait 9:16 output. The Interactions API also supports stateful editing: a follow-up request supplies the prior interaction ID, and the model creates a new video while applying the requested change. Each editing turn produces a new clip rather than modifying the original file in place.[4]
Google positions the model as having world knowledge and an intuitive grasp of physical relationships. Demonstrations include changes to materials, objects, actions, camera angles, and visual styles, as well as prompts tied to scientific or historical subjects. These examples show intended behavior, not a guarantee that a generated scene is physically or factually correct. Gemini Omni also differs from Google's Veo line in emphasis: Omni is built around multimodal references and conversational revision, while Veo remains a dedicated video-generation family.[1][7][10]
Product access and pricing
Consumer access began in the Gemini app, Google Flow, YouTube Shorts, and YouTube Create. The model card also lists Google Flow Music as a distribution channel. Google Vids later added text and image conditioned generation plus conversational editing for eligible Google AI and Workspace customers.[1][2][9]
Developer access is available through Google AI Studio and the paid tier of the Gemini Developer API. Standard API pricing lists input at $1.50 per million tokens for text, image, video, or audio. Output costs $9.00 per million text tokens and $17.50 per million video tokens. Google calculates 720p video at 5,792 output tokens per second, which makes the listed effective video price about $0.10 per second. There is no free API tier for this model.[6]
Google's enterprise platform lists the preview in the global region with fixed-quota consumption. It does not support pay-as-you-go, batch inference, or provisioned throughput on that surface. The enterprise documentation gives the preview a June 30, 2027 retirement date, although Google may update preview availability before then.[8]
Evaluation
Google reports human preference tests across four main video tasks. For video editing, human raters compared outputs on 504 examples and Google reported leading overall preference and instruction-following results against other video models. Its text-to-video test used 1,003 prompts from Meta's MovieGenBench. An image-to-video evaluation used 355 VBench image and text pairs, where Google reported a leading tie among Omni Flash, Grok Imagine Video, and Kling. A reference-to-video study used 468 multimodal examples and measured overall preference and speech adherence.[7]
These results are vendor-reported. The public model page describes sample counts and evaluation methods, but it does not expose all prompts, individual ratings, or enough detail to reproduce the comparisons independently. The initial May model card also said fuller evaluation results would accompany API access, so the later model page should be read as an update to that launch document.[2][7]
A separate 35-case test by K-Dense examined one-shot scientific and educational clips through an early-access API. Thirty-four generations succeeded. Reviewers rated 16 clips at least 4 out of 5 for scientific accuracy, four as mixed, and 14 as containing serious errors. The test found strong visual quality and prompt adherence but recurring problems with scientific labels, equations, mechanisms, and physical dynamics. It was a small, non-peer-reviewed evaluation without iterative prompt refinement, so it is evidence of specific failure modes rather than a general model ranking.[11]
Safety and provenance
Google says Gemini Omni Flash underwent automated and human evaluations during and after training, specialist human red teaming, automated red teaming, and pre-release ethics and safety reviews. The model card describes pre-training work on synthetic captioning and post-training production filters. Inputs and generated outputs in the API are subject to content-safety filters and Google's prohibited-use policy.[2][4]
Generated Omni videos carry an imperceptible SynthID watermark. Google also says content created or edited through the Gemini app, Flow, or YouTube includes C2PA Content Credentials. The two mechanisms serve different purposes: SynthID embeds a detectable signal in the media, while Content Credentials attach provenance metadata.[7]
Speech alteration receives additional restrictions. The model card says the underlying model can change a person's speech, but Google has restricted that capability while it continues safety testing. The public Developer API guide lists voice editing as unsupported. Google's personal-avatar workflows also require product-specific enrollment and likeness controls.[2][4][9]
Limitations
Google's model card identifies incomplete consistency across edits, complex motion, and accurate text rendering as known weaknesses. The Developer API has further operational limits: it does not support video extension, interpolation between first and last frames, reasoning across multiple video references, voice editing, YouTube videos as a media source, system instructions, temperature controls, top_p, stop sequences, or a separate negative-prompt field. Google says English is fully supported, while other languages have not been evaluated and may vary in quality.[2][4]
Some geographic restrictions apply. Editing uploaded videos is unavailable through the Developer API in the European Economic Area, Switzerland, and the United Kingdom, although editing a video generated by the model is supported there. Uploading or editing images containing minors is also restricted in those regions, and images of certain recognizable people may be rejected more broadly.[4]
Preview status is itself a limitation. Model behavior, quotas, interfaces, and documented ceilings may change before a stable release. The difference between the million-token Developer API context window and the smaller enterprise limit also means developers should verify the documentation for the exact surface they deploy on. Generated factual or educational media requires human review, especially where a visually plausible error could be mistaken for a correct explanation.[3][8][11]
References
- Google, "Introducing Gemini Omni," May 19, 2026. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni/ ↩
- Google DeepMind, "Gemini Omni Flash: Model Card," May 19, 2026. https://deepmind.google/models/model-cards/gemini-omni-flash/ ↩
- Google AI for Developers, "Gemini Omni Flash," updated June 30, 2026. https://ai.google.dev/gemini-api/docs/models/gemini-omni-flash ↩
- Google AI for Developers, "Generate and edit videos with Gemini Omni Flash," updated July 6, 2026. https://ai.google.dev/gemini-api/docs/omni ↩
- Google AI for Developers, "Gemini API release notes," June 30, 2026 entry. https://ai.google.dev/gemini-api/docs/changelog ↩
- Google AI for Developers, "Gemini Developer API pricing," updated July 9, 2026. https://ai.google.dev/gemini-api/docs/pricing ↩
- Google DeepMind, "Gemini Omni." https://deepmind.google/models/gemini-omni/ ↩
- Google Cloud, "Gemini Omni Flash Preview," accessed July 24, 2026. https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/omni-flash-preview ↩
- Google, "Create, edit and star in videos with two Google Vids updates," July 16, 2026. https://blog.google/products-and-platforms/products/workspace/gemini-omni-personal-avatars/ ↩
- TechCrunch, Rebecca Bellan, "Google's Gemini Omni turns images, audio, and text into video, and that's just the start," May 19, 2026. https://techcrunch.com/2026/05/19/googles-gemini-omni-turns-images-audio-and-text-into-video-and-thats-just-the-start/ ↩
- K-Dense, "Benchmarking Google's Omni Flash for Scientific Video," updated June 30, 2026. https://www.k-dense.ai/blog/benchmarking-google-omni-flash-scientific-video ↩
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.