觉
AI觉醒星球
Awakening is here
Knowledge File / AI小生意项目库
2026-09-24 2 浏览 免费阅读

Google 新 Flash TTS 模型支持用文本描述从零设计 AI 声音

谷歌发布 Gemini 3.8 Flash TTS 与 Flash-Lite TTS 两款语音生成模型,支持 100 多种语言。前者可用文本描述从零设计音色,30 秒音频样本即可克隆声音;后者面向低成本规模化配音、音频内容和语音 agent。两款模型均支持逐句舞台指示、双人对白和非语言声音。

SOURCE / AI小生意项目库 MIN / 4 ACCESS / 免费阅读 POST / 2026-09-24 01:39:10

原贴

查看原文
作者:Matthias Bastian 来源站点:the-decoder.com 原贴时间:

原文

Google has released two new text-to-speech models, Gemini 3.8 Flash TTS and Flash-Lite TTS, which support more than 100 languages. Flash TTS can create new voices from text descriptions, and a voice cloning feature builds voice profiles from 30-second audio samples. Flash TTS is aimed at creative uses such as podcasts, audiobooks, and game characters, while Flash-Lite TTS is designed for low-cost speech generation at scale for dubbing, audio content, and voice agents. Both models support stage directions for each line, two-voice dialogue, and nonverbal sounds like laughter and sighs. They're rolling out through the Gemini API and Google AI Studio, and Gemini Enterprise API access will follow. Google is introducing Gemini 3.8 Flash TTS and Flash-Lite TTS, two new models for speech generation. Flash TTS can create new voices from text descriptions, and both models support more than 100 languages and let users add stage directions to individual lines of dialogue. Gemini 3.8 Flash TTS is designed for creative projects such as game characters, audiobooks, and podcasts, while Gemini 3.8 Flash-Lite TTS focuses on low-cost speech generation at scale for dubbing, audio content, and voice agents, according to Google. With Gemini 3.8 Flash TTS, users can design voices from scratch. According to Google, a text prompt can define a voice's role, accent, and vocal traits across a wide range of languages and dialects. For users who don't want to start from zero, Google offers a library of more than 2,000 preset voices, including regional variants such as Mexican Spanish, Quebec French, and Scottish English. Ad A voice cloning feature can build a voice profile from a 30-second audio sample. To use it, the person whose voice is being cloned has to record a spoken statement of consent, and the voice in that recording must match the sample. Every clip the Gemini audio models generate carries an inaudible SynthID watermark to help detect AI-generated speech, according to Google . Ad Google has also announced "Voice Remixing," a feature that will let users adjust the timbre, pitch, tempo, and accent of library voices, but it isn't available yet. Both models let users write directions for each line or have the model interpret script cues on its own. Google says the models can generate hours of audio with minimal "speaker drift," meaning the voice barely changes over time. Ad A two-voice mode generates dialogue from a single script while keeping the voices distinct, according to the company. Users can also script laughter, sighs, and sounds like "mhm" to place reactions and pauses exactly where they want them. In two of my own tests, I used a preset voice with a style prompt asking it to imitate an annoyed Berliner speaking English with a thick German accent. The style controls produced a convincing accent and intonation, but both tests had a high-pitched whine in the background at some points. In one of them, the voice also changed at the end of the clip. Ad Style prompt: A native German man from Berlin speaking English as a foreign language, with a thick, unmistakable German accent. He is clearly not a native English speaker: he pronounces English words the German way, applying German rhythm and intonation to English sentences. "Th" becomes "z" or "d" ("ze," "sink," "dat"), "w" becomes "v" ("vat," "vell"), and final consonants become harder ("goot," "bat" for "bad"). The "r" is guttural and throaty, never the English "r." Vowels are flat and short, lacking the softness found in American or British English.

中文翻译

谷歌发布了两款新的文本转语音模型,Gemini 3.8 Flash TTS 和 Flash-Lite TTS,支持超过 100 种语言。Flash TTS 可以通过文本描述创建新的声音,声音克隆功能可以从 30 秒的音频样本构建声音档案。Flash TTS 面向播客、有声书和游戏角色等创意用途,而 Flash-Lite TTS 则面向配音、音频内容和语音 agent 的低成本规模化语音生成。两款模型都支持为每一行添加舞台指示、双人对白,以及笑声和叹息等非语言声音。它们正通过 Gemini API 和 Google AI Studio 推出,Gemini Enterprise API 的接入将随后跟进。

核心信息

谷歌发布 Gemini 3.8 Flash TTS 与 Flash-Lite TTS 两款语音生成模型,支持 100 多种语言。前者可用文本描述从零设计音色,30 秒音频样本即可克隆声音;后者面向低成本规模化配音、音频内容和语音 agent。两款模型均支持逐句舞台指示、双人对白和非语言声音。

  • 谷歌发布 Gemini 3.8 Flash TTS 与 Flash-Lite TTS 两款语音生成模型,支持 100 多种语言。前者可用文本描述从零设计音色,30 秒音频样本即可克隆声音;后者面向低成本规模化配音、音频内容和语音 agent。两款模型均支持逐句舞台指示、双人对白和非语言声音。
  • 原贴提到:Google has released two new text-to-speech models, Gemini 3.8 Flash TTS
  • 来源:the-decoder.com

详细解读

这是什么信号

谷歌把文本转语音从“把字念出来”推进到“按描述造声音”。Gemini 3.8 Flash TTS 可以用一句文本提示定义声音的角色、口音和声线特征,覆盖多种语言和方言;不想从零设计的用户,还有 2000 多个预设音色可选,包括墨西哥西班牙语、魁北克法语、苏格兰英语这类地区变体。同时推出的 Flash-Lite TTS 走的是另一条线:低成本的规模化语音生成,面向配音、音频内容和语音 agent。

两款模型共有的能力同样值得注意:支持逐句舞台指示、由模型自行解读脚本提示、单脚本生成双人对白并保持音色区分、可以把笑声、叹息和“mhm”这类声音写进脚本精确摆放,官方称能生成数小时音频而“说话人漂移”极小。生成片段都会带上不可听的 SynthID 水印。谷歌还预告了尚未上线的“Voice Remixing”,允许调整预设音色的音色、音高、节奏和口音。

为什么重要

语音内容的生产瓶颈一直不在“能不能生成”,而在“能不能按指定的人设、情绪、节奏稳定生成”。文本描述造声、2000+ 预设音色、30 秒克隆这三件事叠在一起,等于把原本需要请配音、租棚、反复重录的环节压缩成提示词工程。双人对白和非语言声音的脚本化,直接把可用边界从朗读推到短剧、播客对谈、游戏 NPC 互动。

Flash-Lite 的定位同样关键:它把价格维度单独拆出来,意味着规模化语音场景——批量配音、语音 agent 的实时对话——会先在这条线上跑通成本模型。

对谁有价值

  • 播客、有声书、短视频创作者:用风格提示词固定一套“自己的声音”,批量产出。
  • 独立游戏开发者:用文本描述给角色定音,用双人对白模式做剧情对白。
  • 做语音 agent 的团队:Flash-Lite 瞄准的低成本规模化生成直接对应实时语音交互。
  • 出海与本地化团队:地区变体音色覆盖墨西哥西班牙语、魁北克法语、苏格兰英语等,适配区域市场。

可以怎么行动

  • 通过 Gemini API 或 Google AI Studio 先做小样本测试,用自己真实的脚本类型(对白、旁白、带情绪的独白)验证效果,而不是只看官方 demo。
  • 把“声音风格提示词”当作可复用资产沉淀:角色设定、口音、发音规则、节奏写进模板,团队内共享。
  • 对配音和语音 agent 场景,先评估 Flash-Lite 的成本与质量平衡点。
  • 涉及真人音色时,提前准备好克隆授权流程:被克隆者需录制同意声明,且声明中声音必须与样本一致。

风险与限制

作者自测给出了明确的反面证据:用预设音色加风格提示词模仿“带浓重德国口音的柏林人讲英语”,口音和语调有说服力,但两次测试都在某些段落出现背景高频啸叫,其中一次音色在片段结尾发生变化。这说明生成质量在长文本和复杂风格下仍不稳定,正式生产前必须逐段验收。

合规层面有两道硬约束:克隆需要同意录音且音色比对通过;所有生成片段带 SynthID 水印,可被检测。Voice Remixing 尚未上线,不能计入当前可用的能力范围。另外,模型实际表现与官方描述之间的差距,只有在自己业务语料上跑过才算数。

信息差价值

这条内容的真正价值,不只是“有人发布了一个新功能”,而是它揭示了 the-decoder.com 背后的产品方向、工作流变化或竞争信号。对 OPC 来说,这种信息可以转化成持续追踪的栏目选题。

如果把《Google 新 Flash TTS 模型支持用文本描述从零设计 AI 声音》放到你的内容系统里,它最大的价值在于帮助读者更快看懂“为什么值得关注”,而不是只看到一条碎片化动态。

参考来源

AI SUMMARY

这篇文章回答了什么

Google 新 Flash TTS 模型支持用文本描述从零设计 AI 声音主要讲什么?

谷歌发布 Gemini 3.8 Flash TTS 与 Flash-Lite TTS 两款语音生成模型,支持 100 多种语言。前者可用文本描述从零设计音色,30 秒音频样本即可克隆声音;后者面向低成本规模化配音、音频内容和语音 agent。两款模型均支持逐句舞台指示、双人对白和非语言声音。

这篇文章最值得关注的要点是什么?

谷歌发布 Gemini 3.8 Flash TTS 与 Flash-Lite TTS 两款语音生成模型,支持 100 多种语言。前者可用文本描述从零设计音色,30 秒音频样本即可克隆声音;后者面向低成本规模化配音、音频内容和语音 agen…;原贴提到:Google has released two new text-to-speech models, Gemini 3.8 Flash TTS;来源:the-decoder.com

这篇文章和哪些AI专题相关?

它适合放在AI副业、Agent工作流、AI工具专题里阅读。 关联原因:这篇内容命中「项目、小生意、变现」等主题信号。;这篇内容命中「Agent」等主题信号。;这篇内容命中「模型」等主题信号。

阅读这篇文章建议先理解哪些关键词?

建议先理解AI工具、工具、自动化、模型、Cursor这些关键词,再结合正文判断工具、机会或风险是否值得进入自己的工作流。

上一篇 Google Beam 扩展至新地区、新合作伙伴与新客户 下一篇 Claude Opus 5.5 与 GPT-6 Sol/Luna 发布,Simon Willison 详解新一轮价格战