Google introduces Gemini Omni - A new level of Multimodel GenAI Editing Reality
Ronni Holmvig Strøm · 2026-05-19
Google used I/O 2026 to introduce Gemini Omni, a new model family built to generate and edit media across modalities, beginning with video. The first release, Gemini Omni Flash, is rolling out to the Gemini app,
Google used I/O 2026 to introduce Gemini Omni, a new model family built to generate and edit media across modalities, beginning with video.
The first release, Gemini Omni Flash, is rolling out to the Gemini app, Google Flow, and YouTube Shorts, with developer and enterprise APIs expected in the coming weeks.
Google’s claim is that Omni can take images, audio, video, and text as input, then produce editable video through conversation.
AI video is then moving from studio tools into the interfaces where users already search, chat, make, and post.
Multimodal Generation Is Becoming An Interaction Model
The phrase multimodal generation has carried a faintly technical flavor for years, as if the central achievement were that a model could accept more than one kind of input. Omni suggests a different emphasis. The model is being presented less as a converter between formats and more as a conversational editing surface for media that can be changed in place.
Interaction changes in a significant way.. A user is no longer only writing a prompt and receiving a finished clip.
They can ask for a character to move, a background to change, a style to shift, an object to appear, or a scene to be re-lit, while the system preserves enough continuity that the conversation feels like revision rather than restart.
Google’s pitch is broad. “Anything from any input,” starting with video. We should treat that as product language, with all the compression and ambition that implies.
Still, even the narrower version is consequential. A video model that accepts prior media and conversational correction starts to behave like an editor that remembers the scene.
The Consumer Surface Is The Governance Event
AI video debates have focused on realism, duration, coherence, prompt adherence, and safety filters. Google is placing the first model inside Gemini, Flow, and YouTube Shorts.