Short scenes built on dialogue
Stage a brief exchange where voice, reaction, camera rhythm and location all pull together.
Official model guide plus a working online generator
Generate narrative video with native audio, cinematic camera language, frame control, and up to seven visual references. The Vidu Q3 AI video generator just below lets you create from a prompt, an image, a pair of frames or supported references.

The full MuseGen video workspace is built in here with Vidu Q3 already selected. You can still try other variants or change models and keep the rest of your setup.
You need an account to generate, and credits depend on the model, variant, length, resolution and other settings you choose. Credits for failed tasks come back automatically.
Know where Vidu Q3 fits before you invest time in references, prompts and renders at full resolution.
Vidu Q3 is made for short-form storytelling, not a silent visual snippet. Pro and Turbo handle text, image and frame-led generation with native audio, while the reference models let you carry characters, products and props from a set of pictures into one scene.
For short-film makers, ad teams, character designers, social studios, brand marketers and narrative designers, the real benefit is not one headline score. It is how Vidu Q3 pairs Text, image, first and last frame, or references to video with Text plus up to seven reference or frame images. That pairing decides whether the model can hold on to a visual direction you have prepared, or has to imagine most of the scene from words.
Here, research and production happen in the same place. Read the official specs, look through the source material, use the prompt structure, then create in the built-in generator. It opens on Vidu Q3, and everything else MuseGen offers — uploads, progress, history, reuse and downloads — is still there.
Plan your inputs, format, length and quality tier from these official capabilities before you spend a credit.
Begin with projects where this model's strongest controls give you a real edge, rather than choosing on top resolution alone.
Stage a brief exchange where voice, reaction, camera rhythm and location all pull together.
Feed in product, character, costume and location images to keep campaign elements intact in a brand-new scene.
Fit a setup, an action and a payoff into a tall or square clip with its own synchronized sound.
Travel between prepared key frames, looks, products or locations while you keep control of how it ends.
This official footage comes from the developer's own launch or product material and has been optimised to play quickly on this page.
From the official source: Native-audio sample from the official Vidu Q3 product page.
Take an idea from first prompt to a configured Vidu Q3 render without leaving the model page.
Write a prompt covering the subject, action, setting, camera, look, pacing and audio. Got references? Upload them and say what each one is there for.
Leave Vidu Q3 selected, pick the variant, mode, length, frame shape, resolution and audio options that suit the shot, then look at the credit cost displayed.
Start the job, watch it progress beside the generator, check the finished clip, then reuse the same settings for a fresh attempt or download the file.
The core capabilities that decide how Vidu Q3 deals with direction, references, movement, audio and final delivery.
Speech, room sound, effects and music are produced with the visuals and follow the rhythm of the scene.
Pick Pro for higher fidelity and stronger storytelling, or Turbo for quicker, cheaper exploration in the main input modes.
Blend as many as seven visual references while you direct characters, products, places, style and scene changes.
Set how the clip begins and ends for morphs, reveals, match cuts and carefully steered movement.
Make a tiny burst of motion, or give a fuller narrative moment enough room to unfold.
Direct framing, pacing, movement, acting and emphasis as part of the narrative from the start.
A good Vidu Q3 prompt reads like a short production brief: a subject, actions in order, a camera plan, an art direction and a soundtrack.
Subject + actions in order + setting + camera + lighting + look + timing + dialogue and sound + what must stay consistent
“Image 1 is the older man, Image 2 the young woman, Image 3 the train platform and Image 4 the costumes. Over twelve seconds they spot each other through the crowd, she says one line — “You came back” — and they both look at the old suitcase. Slow push-in, held-back expressions, station echo, a far-off whistle, no music.”
Say what the characters are after, how their faces shift and how close the camera sits to serve the beat.
Keep spoken lines apart from room tone, effects, score and silences so the intended sound stays clear.
State which image sets identity, costume, product, place, framing or style instead of weighting them equally.
Choose the tier and format that match your stage of the project. Draft settings help you find the shot; premium settings finish a direction that is already working.
| Feature | What Vidu Q3 offers |
|---|---|
| Tiers | Pro, Turbo, Reference or Reference Mix |
| Modes | Text, image, first and last frame, or references to video |
| Accepted inputs | Text plus up to seven reference or frame images |
| Reference support | A first frame, a last frame, or as many as seven images |
| Length | 1–16 seconds |
| Output resolution | 540p, 720p or 1080p |
| Frame shapes | 16:9, 9:16, 1:1, 4:3 and 3:4 |
| Sound | Native speech, music, background sound and effects in sync |
AI video behaves best when each prompt carries one clear visual idea. Check the important details before you publish, and treat a first render as a directed take you can improve.
Busy interactions between several characters, fast objects crossing the frame, legible lettering, brand marks, fingers and exact counts of things can still change from take to take. Lean on clear references, keep crowded action simple and check continuity frame by frame.
More pixels are no substitute for art direction. Nail the story beat, framing, motion and sound on an inexpensive setting first, then move the best version up to a premium tier or a higher resolution.
Weigh up a different mix of motion, references, audio, speed, length and resolution without leaving the MuseGen model library.
Generate polished 5-second 768p videos in around three seconds. H3 Max Turbo turns text, a starting image, or first and last frames into prompt-faithful video with native synchronized audio.
View H3 MaxTurn a slide deck, a PDF, a public webpage, or up to 20 references into one continuous 30-second shot at 1080P with native audio.
View Wan 3.0Create cinematic AI video with native dialogue, sound effects, first-and-last-frame control, and output up to 4K.
View Veo 3.1Direct multi-shot video with text, image, video, and audio references, native stereo sound, and precise creative control.
View Seedance 2.0Vidu Q3 is an AI video generation model from Vidu. Vidu Q3 is made for short-form storytelling, not a silent visual snippet. Pro and Turbo handle text, image and frame-led generation with native audio, while the reference models let you carry characters, products and props from a set of pictures into one scene. On MuseGen the full generator sits right here, letting you go from reading about the model to creating with it without switching workspaces.
Vidu Q3 offers Text, image, first and last frame, or references to video. That means you can begin from a written idea, fix the first frame with a picture, or bring in extra references when the layout, the identity or the motion needs tighter control.
Clips run 1–16 seconds, with output at 540p, 720p or 1080p. Use a lower resolution while you explore ideas, then switch to the highest sensible resolution once you are ready to judge fine detail or hand the shot over.
Vidu Q3 handles Native speech, music, background sound and effects in sync. Write speech, room tone, score and sound effects into the prompt on purpose so the soundtrack backs up the on-screen action and the emotional pace of the scene.
It takes Text plus up to seven reference or frame images. For references it supports A first frame, a last frame, or as many as seven images. Tell the prompt what each uploaded file is for rather than hoping the model guesses which image governs identity, style, composition or motion.
Vidu Q3 suits short-film makers, ad teams, character designers, social studios, brand marketers and narrative designers especially well. The right pick still comes down to the individual shot: use the official specs, features, examples and prompt tips on this page to judge whether its mix of control, pace, resolution, audio and references fits the job.
A dependable prompt covers the subject, the action, the place, the camera, the light, the look, the timing and the sound. List events in the order they happen, put exact dialogue in quotes, and say what has to stay consistent. If you upload references, name each one.
Yes. The complete Vidu Q3 generator is built into this page just below the hero. Pick text, an image, two frames or references as needed, set the available options, check the credit cost shown, and start generating without ever leaving this page.
Capabilities and media on this page were checked against the developer's official product pages, announcements and documentation.
Head back to the full Vidu Q3 AI video generator at the top, add a prompt or your references, and turn the next idea you have into a finished clip.
Generate nowSee every model