Pricing Blog
Google · One model. Generate, edit, and iterate in a single conversation.

Gemini Omni Flash

Generate and edit video by talking to it

Gemini Omni Flash is Google DeepMind's multimodal model that turns text and images into 720p video with native audio, then lets you refine each clip through natural conversation.

720p
Resolution on SharkFoto
10s
Max clip length
24fps
Frame rate
$0.10
Per second output
What's new

Generate and edit video by talking to it

01

Conversational editing

Powered by the Interactions API, edits build on each other turn by turn. Ask for a new camera angle or lighting change and the model keeps character, scene, and continuity intact.

02

Truly multimodal input

Combines text, images and video in a single request; voice references work in Google's own apps, while audio uploads aren't yet supported in the API. The model decides how to weave a reference image or clip into the generated result.

03

Native audio, generated together

Sound is produced alongside the visuals rather than dubbed on afterward, so dialogue, ambience, and sound design stay in sync with the scene.

04

Physics-aware motion

An improved understanding of gravity, kinetic energy, and fluid dynamics yields more believable movement across the short clips it creates.

Overview

Model at a glance

Gemini Omni Flash is the first model in Google's Gemini Omni family, unveiled at Google I/O on 19 May 2026 and opened to developers through the Gemini API and Google AI Studio on 30 June 2026. Unlike earlier pipelines that treat generation and editing as separate steps, it merges reasoning and creation into one architecture: you describe a scene, get a 720p clip with native audio, and then keep refining it through plain-language conversation without losing character, lighting, or scene continuity.

The model accepts text, images, and video as input (voice references work in Google's own apps, but the API doesn't yet accept audio uploads) and produces high-resolution video with sound as output. Multi-turn editing is handled by the Interactions API, and every clip carries an invisible SynthID watermark for provenance. Priced at $0.10 per second of output, the same as Veo 3.1 Fast, it now serves as the default video model in the Gemini app for AI Plus, Pro, and Ultra subscribers. Google made it generally available as gemini-omni-1.1-flash on 27 August 2026, adding video extension, first/last-frame interpolation and resolution control; the preview endpoint is retired as of 30 September 2026.

Developer
Google DeepMind
Model ID
gemini-omni-1.1-flash
Modality
Multimodal to video
Resolution
360p/720p native, 1080p/4K upscaled; 24 FPS
Clip length
3–10 seconds
Audio
Native, synchronized
Price
$0.10 / second
Announced
Google I/O, 19 May 2026
Capabilities

Key features

Multi-turn conversational editing

The Interactions API lets you treat a generated clip as a living draft. Follow-up prompts adjust one element at a time, and the model preserves continuity across turns so a series of edits reads as one coherent scene rather than disconnected regenerations.

Multimodal input to video

Text, images, and video can be combined in a single request; voice references work in Google's own apps, while audio uploads aren't yet supported in the API. Provide a reference image for a character or a short clip to extend, and the model decides how to incorporate it based on your prompt.

Synchronized native audio

Dialogue, ambience, music, and sound design are generated together with the picture. You can steer the soundtrack directly in the prompt, describing background music or specific sound-design cues alongside the visuals.

Physics-aware generation

Google improved the model's intuition for physical forces such as gravity, kinetic energy, and fluid dynamics, producing motion and interactions that hold together more convincingly than prior fast video models.

Landscape and portrait output

Both 16:9 landscape (the default) and 9:16 portrait are supported, covering widescreen storytelling and vertical social formats from the same model.

Fast, cost-controlled clips

Native output runs at 720p (the default) and 24 FPS in 3-to-10-second clips, billed in the Gemini API at $0.10 per second at 720p. A full ten-second clip lands around one dollar, making iterative editing affordable enough to actually iterate.

SynthID provenance by default

Every generated clip embeds Google's invisible SynthID watermark so AI-generated video can be identified downstream, with the safeguard enabled automatically rather than opt-in.

Specs

Technical specifications

Output

Resolution360p, 720p (default); 1080p and 4K via upscaling — 720p on SharkFoto
Frame rate24 FPS
Clip length3–10 seconds (extendable to 40 s via the API); 10 s on SharkFoto
Aspect ratios16:9 (default), 9:16
AudioNative, generated with video
WatermarkSynthID (on by default)

Model

Model IDgemini-omni-1.1-flash
FamilyGemini Omni
ArchitectureTransformer, native multimodal
Context window~1,048,576 tokens
Editing inputVideo up to 10s
Language supportEnglish fully supported

Access & pricing

Price$0.10 per second of output
10s clip≈ $1.00
APIsGemini API, Google AI Studio
EditingInteractions API
Consumer appGemini app & Google Flow (AI Plus/Pro/Ultra); free in YouTube Shorts
StatusGenerally available (Aug 27, 2026)
In practice

Use cases

Short-form social video

Generate vertical 9:16 clips with matching audio for Reels, Shorts, and TikTok, then tweak the hook or pacing conversationally until the opening lands, all inside one session.

Ad and product concepting

Spin up multiple 10-second variations of a product shot with different lighting, camera moves, or soundtracks to compare directions before committing to a full production.

Storyboard-to-motion

Feed a reference image of a character or set and let the model animate it, refining camera angles and continuity turn by turn to previsualize a scene.

Iterative creative direction

Because edits build on each other without regenerating from scratch, teams can direct a clip the way they would brief a person: change one thing, review, change the next.

Rapid prototyping for developers

The Gemini API and AI Studio access, plus a ~1M-token context, make it practical to wire conversational video generation into apps and test flows at $0.10 per second.

Generational leap

Gemini Omni Flash vs. Veo 3.1

In the Gemini app, Omni Flash replaces Veo 3.1 as the default video model. Here is how the fast multimodal newcomer compares to Google's prior video generator on the facts Google has published.

FeatureVeo 3.1Gemini Omni FlashNEW
Primary workflowPrompt-to-video generationGenerate plus multi-turn conversational editing
Input modalitiesText, imageText, image, video
Native audioYesYes
Iterative editingLimited / separate workflowBuilt in via Interactions API
Pricing (Fast tier)$0.10 per second$0.10 per second
ProvenanceSynthIDSynthID
Honest look

Current limitations

Some capabilities still restricted

Google's avatar feature (videos featuring your own likeness and voice) was available at launch, but some capabilities remain restricted: voice and speech editing isn't supported yet, audio references can't be uploaded through the API, and editing or extending uploaded videos is unavailable in the EEA, Switzerland and the UK. On SharkFoto, Gemini Omni Flash is currently offered for text-to-video only.

Short native clips

Each generation is a 3–10 second clip at 360p or 720p native resolution; through the Gemini API, output can also be upscaled to 1080p or 4K and a clip extended in 10-second steps, up to 40 seconds in total. On SharkFoto, Gemini Omni Flash currently produces fixed 10-second 720p clips. It is built for fast, iterative short-form work rather than long-form deliverables.

English-first language support

Google states English is fully supported while other languages have not been formally evaluated, so prompt reliability may vary outside English.

Fixed frame rate and format set

Clips render at 24 FPS in 16:9 or 9:16 only. Projects needing higher frame rates or other aspect ratios will need additional post-processing.

FAQ

Frequently asked questions

Is Gemini Omni Flash a video model or an image model?
It is a video model. Gemini Omni Flash generates and edits high-resolution video with native audio, accepting text, images, and video as input. It is the first release in Google's Gemini Omni multimodal family.
How is it different from Veo 3.1?
The headline difference is conversational, multi-turn editing via the Interactions API: edits build on each other while preserving character and scene continuity. In the Gemini app, Omni Flash has replaced Veo 3.1 as the default video model, and both sit at $0.10 per second on their fast tier.
What resolution and length does it output?
In the Gemini API, each generation is a 3–10 second clip at 24 FPS in 16:9 or 9:16, with audio generated alongside the visuals. 720p is the default, with 360p drafts and upscaled 1080p or 4K output also available, and clips can be extended up to 40 seconds. On SharkFoto, Text to Video currently produces 10-second 720p clips with audio in 16:9 or 9:16.
How much does it cost?
Google prices Gemini Omni Flash at $0.10 per second of video output, so a full 10-second clip costs roughly one dollar. That matches Veo 3.1 Fast.
How do I access it right now?
Developers can use it through the Gemini API and Google AI Studio (model ID gemini-omni-1.1-flash), and it is available to Google AI Plus, Pro and Ultra subscribers in the Gemini app and Google Flow. On SharkFoto, you can use it today in Text to Video.
Are generated videos watermarked?
Yes. Every clip carries Google's invisible SynthID watermark by default so that AI-generated video can be identified after the fact.
Can I use Gemini Omni Flash on SharkFoto?
Yes. Gemini Omni Flash is available on SharkFoto in Text to Video, producing 10-second 720p clips with native audio in 16:9 or 9:16. Image to Video isn't available for this model on SharkFoto, and conversational editing, clip extension and other resolutions are offered only through Google's own apps and API.

Create with Gemini Omni Flash on SharkFoto

Describe a scene and get a 10-second 720p video with native audio in 16:9 or 9:16 — try Gemini Omni Flash in SharkFoto's Text to Video today.

Try Gemini Omni Flash now