Google Unleashes Gemini Omni 1.1 Flash for Advanced Video AI and Gemini 3.5 Transcribe GA
Google announces the general availability of Gemini Omni 1.1 Flash, offering advanced video generation with creative controls, and Gemini 3.5 Transcribe for high-accuracy, low-latency speech-to-text, empowering developers with powerful new AI capabilities.

In a significant move for the developer community, Google has announced the general availability (GA) of two powerful new additions to its Gemini AI ecosystem: Gemini Omni 1.1 Flash and Gemini 3.5 Transcribe. These releases, made available on August 27, 2026, introduce groundbreaking capabilities for generative video and high-accuracy speech-to-text, respectively, promising to revolutionize how developers build and interact with AI-powered applications.
The upgraded Gemini Omni 1.1 Flash model is designed to provide developers and professional content creators with unprecedented creative control over generative video, making it more suitable for real-world production through the Gemini API in Google AI Studio. Simultaneously, Gemini 3.5 Transcribe offers a suite of dedicated speech-to-text models that boast high accuracy and low latency, catering to a wide array of audio understanding needs. These advancements underscore Google's continuous commitment to pushing the boundaries of AI and providing robust tools for innovation.
1. Gemini Omni 1.1 Flash: Revolutionizing Generative Video for Developers
The general availability of Gemini Omni 1.1 Flash marks a pivotal moment for developers engaged in generative video workflows. This upgraded model, accessible via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, introduces a suite of features that enhance control, speed, and production readiness.
One of the most notable upgrades is Scene Extension, which allows developers to seamlessly continue an existing video from where it ends. Unlike previous models that only referenced the final second, Omni 1.1 can analyze up to 10 seconds of prior video context, enabling more coherent and natural continuations.
Another powerful capability is Interpolation (first + last frame). This feature empowers developers to generate a video transition between two specified images, enabling precise control over camera movements like orbits and zooms, or the creation of seamless looping clips. This level of control opens up new avenues for creative expression and efficiency in video production.
For faster iteration and cost-effectiveness, Google has introduced a 360p drafting mode. This allows developers to generate lightweight video previews up to 60% faster and at one-third the cost of the standard 720p resolution. Furthermore, the new resolution parameter in video_config now supports 360p, 720p (default), 1080p, and 4k outputs, with 1080p and 4K generated using upscaling, providing flexibility for various production needs.
Google emphasizes that these updates make generative video more controllable, faster to iterate on, and polished for real-world deployment, whether developers are building generative video workflows, creative tools, or media editing software.
2. Gemini 3.5 Transcribe: Precision Speech-to-Text for Enhanced Audio Understanding
Alongside the video generation advancements, Google has also made Gemini 3.5 Transcribe generally available. This release introduces two dedicated speech-to-text models based on Gemini's advanced audio understanding capabilities: gemini-3.5-transcribe and gemini-3.5-transcribe-live.
The gemini-3.5-transcribe model is designed for high-accuracy, low-latency non-streaming speech-to-text. It supports utterance-based language detection across over 85 languages, offering robust global applicability. Key features include speaker diarization, which identifies and separates different speakers in an audio recording, and word-level timestamps, providing precise timing for each spoken word. Developers can also leverage custom vocabulary biasing, allowing for the inclusion of up to 1,000 specific terms to improve recognition accuracy for specialized content.
For real-time applications, the gemini-3.5-transcribe-live model provides low-latency, bidirectional streaming speech-to-text over WebSockets using the Live API. This model supports interim and finalized transcription events, a 'Smart transcription mode,' and multiple Voice Activity Detection (VAD) strategies, making it ideal for live captioning, voice assistants, and interactive audio experiences.
These models signify a leap forward in audio processing for developers, offering tools that can be integrated into a wide range of applications requiring precise and efficient conversion of spoken language to text. From improving accessibility to enabling advanced analytics on audio content, Gemini 3.5 Transcribe provides a solid foundation for innovative solutions.
3. Impact on Developer Workflows and Creative Possibilities
The general availability of Gemini Omni 1.1 Flash and Gemini 3.5 Transcribe significantly expands the toolkit available to developers, particularly those in the media, entertainment, marketing, and accessibility sectors. For generative video, the enhanced controls mean less manual post-processing and more predictable, high-quality outputs directly from the AI model. Developers can now rapidly prototype and produce video content with specific creative directives, reducing iteration cycles and accelerating content creation pipelines. The ability to specify first and last frames, for instance, allows for precise storytelling and integration into existing video sequences, moving generative AI from mere novelty to a powerful production asset.
In the realm of audio, Gemini 3.5 Transcribe's precision and real-time capabilities open doors for more sophisticated applications. Developers can build more accurate voice command systems, enhance customer service interactions with detailed call transcriptions and speaker identification, or create dynamic content accessibility features. The multi-language support and custom vocabulary biasing are particularly valuable for global applications and niche industries, ensuring that the AI can accurately interpret diverse linguistic inputs and specialized terminology.
Collectively, these releases empower developers to build more intelligent, interactive, and media-rich applications. They lower the barrier to entry for complex AI tasks, allowing developers to focus on application logic and user experience rather than the intricate details of model training and optimization. The integration through the Gemini API and Google AI Studio streamlines the development process, making these advanced AI capabilities readily accessible and scalable for various projects.
4. Getting Started with the New Gemini APIs
For developers eager to explore these new capabilities, Google has provided comprehensive resources. Gemini Omni 1.1 Flash is available through Google AI Studio and the Gemini Enterprise Agent Platform, accompanied by developer documentation, a cookbook, and prompting guides to facilitate the integration of its video capabilities into applications. The existing gemini-omni-flash-preview endpoint will be deprecated on September 30, 2026, so developers currently using the preview version should plan their migration to the GA endpoint.
Similarly, for Gemini 3.5 Transcribe, developers can refer to the Audio transcription guide, the Live transcription guide, and the Gemini 3.5 Transcribe model page to get started. These resources provide detailed instructions and examples for utilizing both the non-streaming and streaming speech-to-text models effectively.
The availability of these models through the Gemini API ensures that developers can leverage Google's cutting-edge AI technology with familiar API paradigms, allowing for quick adoption and integration into existing and new projects. Google AI Plus, Pro, and Ultra subscribers can also access Omni 1.1 globally through Google Flow, with the scene extension feature available via the Gemini app.
Comparison Overview
| Feature/Model | Description/Capability | Developer Benefit |
|---|---|---|
| Gemini Omni 1.1 Flash (Video Extension) | Analyzes up to 10 seconds of previous video context for seamless continuations. | More coherent and natural video generation, reduced manual editing. |
| Gemini Omni 1.1 Flash (Interpolation) | Generates video transitions between two specified images (first + last frame). | Precise control over camera movements, seamless looping, enhanced creative direction. |
| Gemini Omni 1.1 Flash (360p Drafting Mode) | Generates lightweight video previews 60% faster at 1/3 the cost. | Accelerated iteration, cost-effective prototyping for generative video. |
| Gemini Omni 1.1 Flash (Resolution Control) | Supports 360p, 720p, 1080p, and 4K outputs (upscaling for higher resolutions). | Flexibility for various production qualities and target platforms. |
| Gemini 3.5 Transcribe (Accuracy & Latency) | High-accuracy, low-latency non-streaming speech-to-text. | Reliable and fast transcription for diverse audio content. |
| Gemini 3.5 Transcribe (Language & Diarization) | Utterance-based language detection (85+ languages) and speaker diarization. | Global application support, clear identification of multiple speakers in conversations. |
| Gemini 3.5 Transcribe (Custom Vocabulary) | Custom vocabulary biasing for up to 1,000 terms. | Improved accuracy for specialized terminology in niche domains. |
| Gemini 3.5 Transcribe Live (Streaming STT) | Low-latency, bidirectional streaming speech-to-text over WebSockets. | Enables real-time applications like live captioning and voice assistants. |
Frequently Asked Questions (FAQ)
Q: What is Gemini Omni 1.1 Flash?
Gemini Omni 1.1 Flash is Google's generally available, upgraded generative AI model focused on video creation. It offers new creative controls, faster iteration, and features like scene extension and interpolation, making generative video more suitable for real-world production.
Q: How does Scene Extension work in Gemini Omni 1.1 Flash?
Scene Extension allows the model to seamlessly continue an existing video. It can analyze up to 10 seconds of previous video context to generate coherent and natural continuations, a significant improvement over earlier models that only referenced the final second.
Q: What are the key features of Gemini 3.5 Transcribe?
Gemini 3.5 Transcribe offers high-accuracy, low-latency speech-to-text. Key features include utterance-based language detection across 85+ languages, speaker diarization, word-level timestamps, and custom vocabulary biasing (up to 1,000 terms). It also includes a live streaming version for real-time transcription.
Q: Where can developers access these new Gemini models?
Developers can access Gemini Omni 1.1 Flash and Gemini 3.5 Transcribe through the Gemini API in Google AI Studio. Gemini Omni 1.1 Flash is also available via the Gemini Enterprise Agent Platform. Developer documentation, cookbooks, and prompting guides are available to assist with integration.
Q: Is the preview version of Gemini Omni Flash still available?
No, the existing gemini-omni-flash-preview endpoint is scheduled for deprecation on September 30, 2026. Developers should migrate to the generally available Gemini Omni 1.1 Flash endpoint.
Try Our Developer Utilities
Simplify your engineering workflows with our free browser-native tools: