Knowledge File / AI技能杠杆
趋势解读:9 demos of Gemini Omni and Gemini 3.5,评估 LLM Agent 表现
Google I/O 2026发布Gemini Omni和Gemini 3.5模型,前者支持多模态视频生成与对话式编辑,后者增强代理工作流性能。9个演示展示了自然语言编辑视频、保持角色一致性等能力。
SOURCE / AI技能杠杆
MIN / 9
ACCESS / 会员
POST / 2026-05-30 01:30:00
原贴
查看原文原文
With Gemini Omni, Gemini’s ability to reason meets the ability to create, while Gemini 3.5 is built to help you execute complex, agentic workflows. Your browser does not support the audio element. At Google I/O 2026 , we announced our latest models: Gemini Omni and the Gemini 3.5 family of models. Gemini Omni is our new model that can create anything from any input, starting with video. With Omni, you can combine images, audio, video and text as input and generate high-quality videos grounded in Gemini's real-world knowledge. You can also easily edit your videos through conversation. Then there’s Gemini 3.5, our latest family of models combining frontier intelligence with action. This represents a major leap forward in building more capable, intelligent agents. We’re kicking off the series by releasing 3.5 Flash. It delivers frontier performance for agents and coding, excelling at complex long-horizon tasks that deliver real-world utility. To give you a clearer understanding of Gemini Omni and Gemini 3.5 Flash, here are 9 demos of what they can help you do. Edit your videos through conversation. One capability that makes Omni special is that it gives you an easier way to edit video — with natural language. Every instruction builds on the last. Your characters stay consistent, the physics hold up and the scene remembers what came before. That means you can transform the world around you. Change specific things, or change everything. Your video becomes the starting point for something you never could have filmed yourself. Prompt: Make the sculpture out of bubbles. Reimagine the action. Take a video you shot and just ask Omni to change what’s happening. Edit the action, add in new characters or objects or transform a moment into something unexpected. Prompt: Dim the lights in the room. Put a black and white checkerboard room inside a glass sphere that floats tracking above the hand, inside it contains a recursive representation of the same hand holding the sphere, creating an infinite recursive of rooms. Camera slowly gets closer into the sphere, creating a video loop. Refine your videos across multiple turns. Change the environment, angle, style or even specific details, without ever losing the thread of your original scene. Scroll through the carousel to see how edits build on each other. Prompt: A video of a violinist playing a song.
中文翻译
随着Gemini Omni的发布,Gemini的推理能力与创造能力相结合,而Gemini 3.5旨在帮助你执行复杂的代理工作流。在2026年Google I/O大会上,我们宣布了最新模型:Gemini Omni和Gemini 3.5模型家族。
核心信息
Google I/O 2026发布Gemini Omni和Gemini 3.5模型,前者支持多模态视频生成与对话式编辑,后者增强代理工作流性能。9个演示展示了自然语言编辑视频、保持角色一致性等能力。
- Gemini Omni实现多模态输入到视频生成
- Gemini 3.5 Flash增强代理与编码能力
- 自然语言对话式编辑视频,保持一致性
- 支持多轮编辑,场景和角色记忆连贯
- 发布9个演示展示实际应用效果
试看内容
成为会员查看完整内容
你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。
详细解读
信息差价值
参考来源
成为会员查看完整内容