Gemini Omni Flash supports a practical multimodal workflow for text-to-video, image-to-video, reference-based generation, and instruction-based video editing.
Upload text, images, references, or a source video
Add a prompt, starting frame, multiple reference images, or an existing video so Gemini Omni Flash has clear visual context to follow.
Describe the video direction in natural language
Write direct instructions for subject action, camera movement, scene style, edits, effects, timing, and the final viewing format.
Choose format and generation settings
Select aspect ratio, duration, and output options for social clips, product demos, ads, concept previews, or video-editing drafts.
Generate, review, and refine the clip
Create a Gemini Omni Flash video, check motion and consistency, then revise the prompt or source media for a stronger final result.
Upload text, images, references, or a source video
Add a prompt, starting frame, multiple reference images, or an existing video so Gemini Omni Flash has clear visual context to follow.
Describe the video direction in natural language
Write direct instructions for subject action, camera movement, scene style, edits, effects, timing, and the final viewing format.
Choose format and generation settings
Select aspect ratio, duration, and output options for social clips, product demos, ads, concept previews, or video-editing drafts.
Generate, review, and refine the clip
Create a Gemini Omni Flash video, check motion and consistency, then revise the prompt or source media for a stronger final result.
Traditional AI video workflows can struggle when creators need one model to understand prompts, images, references, and uploaded footage together.

Gemini Omni Flash combines text prompts, images, reference photos, and uploaded videos into a guided workflow for faster generation and editing.

Gemini Omni Flash gives creators a flexible way to generate, guide, and edit AI video with text prompts, images, reference photos, and uploaded footage.