谷歌更新 Gemini Omni 1.1 Flash:支持首尾帧控制、场景延展到 40 秒和 4K 升频 Google updates Gemini Omni 1.1 Flash with first/last-frame control, 40-second scene extension, and 4K upscaling
这次更新把视频生成从“单次出片”推进到“可对话式迭代”。Gemini Omni 1.1 Flash 支持首帧/尾帧补全、最多 10 秒前文感知、3 到 10 秒分段延展,并可叠加到 40 秒,同时提供 360p 草稿到 4K 成片的工作流。 This update moves video generation from one-shot output toward iterative control. Gemini Omni 1.1 Flash adds first/last-frame completion, up to 10 seconds of context for scene extension, 3- to 10-second incremental extension up to 40 seconds, and a draft-to-4K workflow.
核心变化
据源文,Gemini Omni 1.1 Flash 的重点不只是“能生成视频”,而是把视频制作往可指挥、可续写、可反复编辑的方向推进。谷歌强调它具备原生多模态能力、对话式编辑能力,以及继承自 Gemini 的世界知识。
- 场景延展(scene extension)现在可读取最多 10 秒前文,不再只看最后 1 秒。
- 每次延展可生成 3 到 10 秒,并可层层叠加到 40 秒。
- 系统会对输入视频末尾若干帧做调整,让接续更平滑。
具体怎么用
这版模型支持直接给首帧和尾帧,让中间视频补全,因此适合做环绕镜头、推拉变焦、无缝循环等效果。提示词里也支持 ``、``、``、`` 这类标签来指定不同素材的角色。
- 最多可接收 3 个视频参考片段。
- 每个参考片段最长 3 秒,主要用于保持人物外观一致。
- 源文说明视频参考中的音频会被忽略。
- 同时理解多个视频不是强项,可能降低效果。
编辑与迭代
编辑是有记忆的。传入 `previous_interaction_id` 后,模型会基于上一轮结果继续修改,并保留未被提及的内容,这意味着不必每次都重新上传旧视频。
成本和输出
源文把这套流程描述为“先低成本试错,再高质量出片”。草稿可先用 360p 跑,谷歌称这比 720p 快 60%,成本约为三分之一;最终可再升到 1080p 或 4K。
可用性与限制
源文称该能力已可通过 Gemini API、Google AI Studio 和 Gemini Enterprise Agent Platform 使用,也已出现在 Google Flow,并面向 AI Plus、Pro、Ultra 订阅用户;Gemini 应用中也能使用场景延展。
- 已点名的生产用户包括 Adobe、Figma Weave、GMI Cloud、Runway。
- 支持 SynthID 水印。
- 不支持 system instructions、temperature、top_p、stop sequences,也没有独立的负向提示词参数。
- 不支持语音编辑、音频参考,YouTube 链接不能作为素材来源。
What changed
The source frames Gemini Omni 1.1 Flash as a move beyond one-shot video generation toward controllable, iterative editing. Google emphasizes native multimodal input, conversational editing, and Gemini’s built-in world knowledge.
- Scene extension can now read up to 10 seconds of prior context, instead of only the last 1 second.
- Each extension pass can generate 3 to 10 seconds, and the result can be stacked up to 40 seconds.
- The system adjusts the last frames of the input video to make the transition smoother.
How it works
The model can take a first frame and a last frame and fill the middle, which supports orbit shots, push-ins, pull-outs, and seamless loops. The prompt format also supports tags such as ``, ``, ``, and `` to assign roles to different assets.
- Up to 3 video reference clips are supported.
- Each reference clip can be up to 3 seconds long and is mainly used to preserve subject identity.
- Audio in video references is ignored.
- The source notes that understanding multiple videos at once is not a strength and may hurt results.
Editing and iteration
Editing is stateful. When `previous_interaction_id` is provided, the model continues from the prior result and keeps details that were not explicitly changed, so you do not need to re-upload the old video every time.
Cost and output
The workflow is described as a low-cost draft pass followed by a higher-quality final pass. Drafts can run at 360p; Google says that is 60% faster than 720p and costs about one-third as much. Final output can then be upgraded to 1080p or 4K.
Availability and limits
According to the source, the feature is available through the Gemini API, Google AI Studio, and Gemini Enterprise Agent Platform, and it also appears in Google Flow for AI Plus, Pro, and Ultra subscribers. Scene extension is also available inside the Gemini app.
- Named production users include Adobe, Figma Weave, GMI Cloud, and Runway.
- All generated videos carry SynthID watermarks.
- There are no system instructions, temperature, top_p, or stop sequence controls, and no separate negative-prompt parameter.
- Voice editing, audio references, and YouTube links as input sources are not supported.
来源
- MarkTechPost · 08-29 22:30