Gemini 3.8 Flash Is Here: Finding Moments for Opening Teasers and Video Chapters
Google released Gemini 3.8 Flash on September 2, 2026. The new AI model improves performance on coding and tasks involving multiple steps while keeping the same introductory pricing as its predecessor, Gemini 3.7 Flash.
What interests me this time is using it to find specific moments within video files. For example, it could suggest highlights for an opening teaser or potential YouTube chapters. Rather than simply comparing its answers with those from ChatGPT or Claude, I want to try bringing Gemini into the process of reviewing videos.
Gemini 3.8 Flash: Key Specifications
First, here is a summary of the basic specifications, based on the official model documentation and API pricing page.
| Item | Gemini 3.8 Flash |
|---|---|
| Supported inputs | Text, images, video, audio, and PDFs |
| Output | Text |
| Maximum input | 1,048,576 tokens (approximately 1.05 million) |
| Maximum output | 65,536 tokens (approximately 66,000) |
| Thinking levels | low, medium, high |
| API input price | $0.75 per million tokens |
| API output price | $3.75 per million tokens, including thinking tokens |
These are introductory prices for the standard API through December 31, 2026. From January 1, 2027, input and output will cost $1.50 and $7.50 per million tokens, respectively. These charges are separate from the Gemini app’s monthly subscription.
The model is available through the Gemini API, Google AI Studio, and other platforms. In the Gemini app, it is available to Google AI Pro and Ultra subscribers.
Finding Specific Moments in Long Videos
Gemini can analyze video files to describe or summarize their contents. It also supports extracting information from specific timestamps.
It now supports “Agentic video understanding,” which lets it locate relevant sections and examine them more closely. Rather than reading a set of frames sampled at fixed intervals just once, it navigates through the video based on the question, checking visual frames, audio, and transcripts.
This feature is not exclusive to Gemini 3.8 Flash. It was first announced on September 1 for models including 3.7 Flash. The current developer documentation also lists 3.8 Flash among the supported models.
Finding “the scene where this action happens” or “the footage that corresponds to this explanation” within a long video—that is what appeals to me. It makes it possible to search for moments that would be difficult to find just by reading a transcript.
This new processing mode is available through the API. According to the official announcement, its rollout to the Gemini app is still planned. The following are ways I would like to try using the officially documented capabilities.
Suggesting Opening Teaser Candidates from a Video File
One use I am particularly interested in is uploading a video file and asking Gemini to suggest moments for an opening teaser: a short introduction that shows highlights before the main content begins. In Japanese video production, this is called an “avant-title,” often shortened to “avan.”
Beyond memorable remarks, I would like it to notice moments when something changes during a demonstration or someone reacts to it. In other words, I want it to use both the dialogue and the visuals to find moments that make viewers want to keep watching.
Suggest three moments from this video that would work well in an opening teaser. Put the start and end timestamps, what is said and shown on screen, and your reason for choosing each moment in a table. Choose moments that spark interest in the main content, and avoid those that would be misleading when taken out of context.
With both candidates and reasons in hand, I can play back those sections and judge whether they convey the highlights of the video as a whole. If I can start by reviewing suggestions instead of searching from scratch, I can see this being useful in my day-to-day work.
Suggesting YouTube Chapter Candidates
Another use I would like to try is asking for YouTube chapter suggestions. Gemini could identify where topics or demonstrations change within a long video, then compile the start timestamps and short headings.
In a product video, that might be the transition from an overview to a hands-on demonstration. In a conversation, it might be the shift to a new question. I want to try having it suggest chapter boundaries based not only on the dialogue, but also on changes on screen.
Review the entire video and suggest YouTube chapters based on distinct topics or demonstrations. List each candidate as “start timestamp — short heading,” without dividing the video into too many small sections. After the list, add a note about any boundaries you are unsure of.
Starting with a video I already know would make it easier to check whether the headings match the content and whether any major topics are missing. I would want to verify the suggestions, including their timestamps, against the original video before using them.
Trying Gemini in a Different Role from ChatGPT and Claude
I can keep using the AI tools I know for everyday questions and writing, while asking Gemini to find scenes within videos. That is the kind of division of labor by workflow stage I have in mind.
I would start by uploading a video I am allowed to share and asking for opening teaser or chapter candidates. I still need to test how useful the results will be in practice. But if it reduces the time I spend searching through a long video wondering, “Where was that scene?”, that alone would give me a reason to use Gemini.
References
- Google: Official Gemini 3.8 Flash announcement
- Google AI for Developers: Gemini 3.8 Flash specifications
- Google AI for Developers: API pricing
- Google: Agentic video understanding announcement
- Google AI for Developers: Supported models and usage guide for video understanding
Specifications, pricing, and availability were checked against official information as of September 3, 2026.