VFX through conversation
Describe the transformation you want in everyday words and refine it without rebuilding the shot.
Official model guide and online generator
Combine Gemini reasoning with fast video generation, multimodal reference control, and conversational video editing. Use the Gemini Omni Flash AI video generator directly below to create from a prompt, image, frame pair, or supported references.

The complete MuseGen video workspace is embedded here and starts with Gemini Omni Flash selected. You can still compare variants or switch models without losing the rest of the workflow.
Generation requires an account and uses credits based on the selected model, variant, duration, resolution, and other settings. Failed tasks are refunded automatically.
Understand where Gemini Omni Flash fits before spending time on references, prompts, and final-resolution generations.
Gemini Omni Flash links multimodal understanding straight to video creation. It reasons over text, pictures and video before it generates or edits, which makes it particularly handy when your instruction depends on how several references relate or on understanding footage you already have.
For creative and editing teams, product marketers, teachers, social creators and fast-moving prototypers, the practical advantage is not a single headline benchmark. It is the way Gemini Omni Flash combines Text to video, multi-reference video and editing of existing footage with Text, several images and one source video. That combination determines whether the model can preserve a prepared visual direction or needs to invent most of the scene from language alone.
On this page, research and production live in one flow. Read the official specifications, study the source material, use the prompt framework, and then work in the embedded generator. The selected model defaults to Gemini Omni Flash, while the rest of MuseGen's upload, progress, history, reuse, and download experience stays available.
Use these official capabilities to plan the input, format, duration, and production tier before you generate.
Start with work where the model's strongest controls create a practical advantage rather than choosing only by maximum resolution.
Describe the transformation you want in everyday words and refine it without rebuilding the shot.
Feed in product and brand images so design details stay recognisable in a generated commercial moment.
Give an existing video a new look, location, visual joke or brand treatment for a short campaign.
Pull characters, places, props and art direction from different images into one request.
This official media was downloaded from the model developer's launch or product material and optimized for fast playback on this page.
Official source material: Footage released by Google for Gemini Omni Flash, downloaded and optimised for this guide.
Move from a creative idea to a configured Gemini Omni Flash generation without leaving this model page.
Write the subject, action, environment, camera, style, timing, and sound. If you have reference media, upload it and explain the role of each asset.
Keep Gemini Omni Flash selected, choose the appropriate variant, mode, duration, aspect ratio, resolution, and audio settings, then check the displayed credit cost.
Start the task, follow progress in the result panel, review the completed video, reuse its settings for another take, or download the finished file.
The defining capabilities that shape how Gemini Omni Flash handles direction, references, motion, sound, and delivery.
Gemini works out links, instructions, objects and surrounding context across every kind of input before it creates anything.
Change an existing clip in plain language — tweak an effect, swap visual elements or keep building on an idea.
Combine several images to steer subjects, products, locations, outfits, style and other parts of the scene.
A fixed video length suits social concepts, quick effects trials, product moments and fast iteration.
Transform footage you already have while keeping the motion or framing that made it worth using.
Get a full audiovisual result whenever sound, beat or effects belong to the change you asked for.
A strong Gemini Omni Flash prompt behaves like a compact production brief: it gives the model a subject, an ordered action, a camera plan, an art direction, and a soundtrack.
Subject + ordered action + environment + camera + lighting + visual style + timing + dialogue and sound + consistency constraints
“Images 1–4 fix the exact bottle shape, label, materials and colours. In the source video, swap the grey studio for a sunlit desert of pink sand, leave the camera path and the timing of the hands as they are, and add shimmering heat haze with a soft wind-chime sound rising underneath.”
Attaching references is not enough; say which subject, style, object, place or behaviour should come from each.
When editing, separate what must stay untouched from the specific elements you want swapped out.
Stick to one readable idea with a clear start, a change and a final payoff that suits the fixed length.
Choose a tier and format based on where you are in the creative process. Draft settings are for finding the shot; premium settings are for finishing a direction that already works.
| Capability | Gemini Omni Flash support |
|---|---|
| Variants | Preview |
| Generation modes | Text to video, multi-reference video and editing of existing footage |
| Inputs | Text, several images and one source video |
| Reference control | As many as 16 images plus one source video |
| Duration | 8 seconds |
| Resolution | 720p |
| Aspect ratios | 16:9 widescreen and 9:16 vertical |
| Audio | Native audio generation |
AI video is most reliable when the prompt gives each shot one readable visual idea. Review important details before publishing and treat the first generation as a directed take that can be refined.
Complex multi-character interaction, fast occlusion, readable text, logos, hands, and exact object counts can still vary between takes. Use clear references, simplify crowded action, and inspect continuity frame by frame.
Higher resolution does not replace art direction. Lock the story beat, composition, movement, and sound at an economical setting first; then move the strongest direction to the premium variant or resolution.
Compare a different balance of motion, references, audio, speed, duration, and resolution without leaving the MuseGen model library.
Create controlled cinematic shots with first-and-last-frame guidance, native audio, flexible duration, and true 4K output.
Open Kling v3Generate smooth, consistent video from text, a starting image, or up to nine reference images with native synchronized audio.
Open HappyHorse 1.1Generate 2K video with native stereo sound from text, frames, and up to fifteen reference images, clips, and audio tracks at once.
Open MiniMax H3Explore native multishot storytelling, synchronized audio and video, automatic duration, Diffusion Fidelity Rendering, and professional 4K HDR output from Lightricks' new open-weight foundation model.
Open LTX 2.5Gemini Omni Flash is a Google AI video generation model. Gemini Omni Flash links multimodal understanding straight to video creation. It reasons over text, pictures and video before it generates or edits, which makes it particularly handy when your instruction depends on how several references relate or on understanding footage you already have. MuseGen places the complete generator on this page so you can move from research to creation without opening a separate workspace.
Gemini Omni Flash supports Text to video, multi-reference video and editing of existing footage. That range lets you start with a written idea, guide the opening with an image, or use additional references when the composition, identity, or motion must be more controlled.
You can create 8 seconds video with output at 720p. Pick a lower resolution for quick creative exploration, then use the highest appropriate setting when you are ready to evaluate detail or deliver the shot.
Gemini Omni Flash supports Native audio generation. Write dialogue, ambience, music, and effects as deliberate parts of the prompt so the soundtrack supports the visible action and emotional rhythm of the scene.
The model accepts Text, several images and one source video. Its reference workflow supports As many as 16 images plus one source video. Give every uploaded asset a clear role in the prompt instead of expecting the model to infer which image controls identity, style, composition, or movement.
Gemini Omni Flash is a strong fit for creative and editing teams, product marketers, teachers, social creators and fast-moving prototypers. The best choice still depends on the shot: use this page's facts, features, examples, and prompt guide to decide whether its particular balance of control, speed, resolution, sound, and references matches the job.
A reliable prompt names the subject, action, location, camera, lighting, visual style, timing, and sound. Put events in chronological order, quote exact dialogue, and state what must remain consistent. When you upload references, identify each one explicitly.
Yes. The full Gemini Omni Flash generator is embedded directly below the hero on this page. Choose text, image, frames, or references as appropriate, configure the available controls, review the visible credit cost, and start the generation without leaving the model guide.
Model capabilities and media on this page were researched from the developer's official product pages, announcements, and documentation.
Open the complete Gemini Omni Flash AI video generator above, add your prompt or references, and turn the next shot on your list into a finished video.
Start generatingBrowse all models