Gemini Omni Flash is Google DeepMind's multimodal model that turns text and images into 720p video with native audio, then lets you refine each clip through natural conversation.
Powered by the Interactions API, edits build on each other turn by turn. Ask for a new camera angle or lighting change and the model keeps character, scene, and continuity intact.
Combines text, images and video in a single request; voice references work in Google's own apps, while audio uploads aren't yet supported in the API. The model decides how to weave a reference image or clip into the generated result.
Sound is produced alongside the visuals rather than dubbed on afterward, so dialogue, ambience, and sound design stay in sync with the scene.
An improved understanding of gravity, kinetic energy, and fluid dynamics yields more believable movement across the short clips it creates.
Gemini Omni Flash is the first model in Google's Gemini Omni family, unveiled at Google I/O on 19 May 2026 and opened to developers through the Gemini API and Google AI Studio on 30 June 2026. Unlike earlier pipelines that treat generation and editing as separate steps, it merges reasoning and creation into one architecture: you describe a scene, get a 720p clip with native audio, and then keep refining it through plain-language conversation without losing character, lighting, or scene continuity.
The model accepts text, images, and video as input (voice references work in Google's own apps, but the API doesn't yet accept audio uploads) and produces high-resolution video with sound as output. Multi-turn editing is handled by the Interactions API, and every clip carries an invisible SynthID watermark for provenance. Priced at $0.10 per second of output, the same as Veo 3.1 Fast, it now serves as the default video model in the Gemini app for AI Plus, Pro, and Ultra subscribers. Google made it generally available as gemini-omni-1.1-flash on 27 August 2026, adding video extension, first/last-frame interpolation and resolution control; the preview endpoint is retired as of 30 September 2026.
The Interactions API lets you treat a generated clip as a living draft. Follow-up prompts adjust one element at a time, and the model preserves continuity across turns so a series of edits reads as one coherent scene rather than disconnected regenerations.
Text, images, and video can be combined in a single request; voice references work in Google's own apps, while audio uploads aren't yet supported in the API. Provide a reference image for a character or a short clip to extend, and the model decides how to incorporate it based on your prompt.
Dialogue, ambience, music, and sound design are generated together with the picture. You can steer the soundtrack directly in the prompt, describing background music or specific sound-design cues alongside the visuals.
Google improved the model's intuition for physical forces such as gravity, kinetic energy, and fluid dynamics, producing motion and interactions that hold together more convincingly than prior fast video models.
Both 16:9 landscape (the default) and 9:16 portrait are supported, covering widescreen storytelling and vertical social formats from the same model.
Native output runs at 720p (the default) and 24 FPS in 3-to-10-second clips, billed in the Gemini API at $0.10 per second at 720p. A full ten-second clip lands around one dollar, making iterative editing affordable enough to actually iterate.
Every generated clip embeds Google's invisible SynthID watermark so AI-generated video can be identified downstream, with the safeguard enabled automatically rather than opt-in.
| Resolution | 360p, 720p (default); 1080p and 4K via upscaling — 720p on SharkFoto |
|---|---|
| Frame rate | 24 FPS |
| Clip length | 3–10 seconds (extendable to 40 s via the API); 10 s on SharkFoto |
| Aspect ratios | 16:9 (default), 9:16 |
| Audio | Native, generated with video |
| Watermark | SynthID (on by default) |
| Model ID | gemini-omni-1.1-flash |
|---|---|
| Family | Gemini Omni |
| Architecture | Transformer, native multimodal |
| Context window | ~1,048,576 tokens |
| Editing input | Video up to 10s |
| Language support | English fully supported |
| Price | $0.10 per second of output |
|---|---|
| 10s clip | ≈ $1.00 |
| APIs | Gemini API, Google AI Studio |
| Editing | Interactions API |
| Consumer app | Gemini app & Google Flow (AI Plus/Pro/Ultra); free in YouTube Shorts |
| Status | Generally available (Aug 27, 2026) |
Generate vertical 9:16 clips with matching audio for Reels, Shorts, and TikTok, then tweak the hook or pacing conversationally until the opening lands, all inside one session.
Spin up multiple 10-second variations of a product shot with different lighting, camera moves, or soundtracks to compare directions before committing to a full production.
Feed a reference image of a character or set and let the model animate it, refining camera angles and continuity turn by turn to previsualize a scene.
Because edits build on each other without regenerating from scratch, teams can direct a clip the way they would brief a person: change one thing, review, change the next.
The Gemini API and AI Studio access, plus a ~1M-token context, make it practical to wire conversational video generation into apps and test flows at $0.10 per second.
In the Gemini app, Omni Flash replaces Veo 3.1 as the default video model. Here is how the fast multimodal newcomer compares to Google's prior video generator on the facts Google has published.
| Feature | Veo 3.1 | Gemini Omni FlashNEW |
|---|---|---|
| Primary workflow | Prompt-to-video generation | Generate plus multi-turn conversational editing |
| Input modalities | Text, image | Text, image, video |
| Native audio | Yes | Yes |
| Iterative editing | Limited / separate workflow | Built in via Interactions API |
| Pricing (Fast tier) | $0.10 per second | $0.10 per second |
| Provenance | SynthID | SynthID |
Google's avatar feature (videos featuring your own likeness and voice) was available at launch, but some capabilities remain restricted: voice and speech editing isn't supported yet, audio references can't be uploaded through the API, and editing or extending uploaded videos is unavailable in the EEA, Switzerland and the UK. On SharkFoto, Gemini Omni Flash is currently offered for text-to-video only.
Each generation is a 3–10 second clip at 360p or 720p native resolution; through the Gemini API, output can also be upscaled to 1080p or 4K and a clip extended in 10-second steps, up to 40 seconds in total. On SharkFoto, Gemini Omni Flash currently produces fixed 10-second 720p clips. It is built for fast, iterative short-form work rather than long-form deliverables.
Google states English is fully supported while other languages have not been formally evaluated, so prompt reliability may vary outside English.
Clips render at 24 FPS in 16:9 or 9:16 only. Projects needing higher frame rates or other aspect ratios will need additional post-processing.
Describe a scene and get a 10-second 720p video with native audio in 16:9 or 9:16 — try Gemini Omni Flash in SharkFoto's Text to Video today.
Try Gemini Omni Flash now