Google has released Gemini Omni Flash preview for developers, a multimodal video generation model accessible via the Interactions API. Key capabilities include text-to-video generation, image-to-video conversion, configurable aspect ratios (16x9 and 9x6) and durations, native audio generation alongside video, and multi-turn conversational editing using interaction IDs to preserve context across edits. A code walkthrough demonstrates generating videos from text prompts, creating images then converting them to video, merging multiple images into a single video, and editing existing videos with prompt-based instructions. Videos can be saved to Google Drive for reuse.

11m watch time
100 Impressions