觉
AI觉醒星球
Awakening is here
Knowledge File / AI小生意项目库
2026-07-09 3 浏览 免费阅读

ChatGPT 现在可以同时听和说,让 AI 对话更人性化

OpenAI 推出新版语音模型 GPT-Live,采用全双工架构,支持同时听和说,能处理打断、使用填充词,并将复杂任务委托给后台 GPT-5.5 处理。提供付费和免费两个版本,用户偏好率超 69%。

SOURCE / AI小生意项目库 MIN / 9 ACCESS / 免费阅读 POST / 2026-07-09 02:18:55

原贴

查看原文
作者:Matthias Bastian 来源站点:the-decoder.com 原贴时间:

原文

OpenAI has introduced GPT-Live, a new language model for live conversations in ChatGPT that can listen and speak simultaneously thanks to a full-duplex architecture. The model responds to interruptions, uses filler words like "mhmm" to keep the conversation flowing naturally, and is available in two versions: one for paying users and one for free users. For complex tasks that require web searches or logical reasoning, GPT-Live delegates the requests to the GPT-5.5 model in the background while maintaining the conversation, significantly improving answer quality. OpenAI introduces GPT-Live, a new generation of voice models with full-duplex architecture. The model can listen and speak at the same time and hands off complex tasks to GPT-5.5 in the background. With GPT-Live, OpenAI is shipping a new class of AI voice models designed to make talking to ChatGPT feel more like a real conversation. Two versions are rolling out worldwide right away. GPT-Live-1 is for paying users on Go, Plus, and Pro plans, while GPT-Live-1 mini is available on free accounts. Both work on iOS, Android, and ChatGPT.com . OpenAI plans to add API access soon, and developers can sign up through a form . Instead of the old rigid back-and-forth where users speak and the AI responds, GPT-Live uses a full-duplex architecture. The model can listen and talk at the same time, much like a real conversation between people. Earlier this year, Nvidia released a similar open-source model called PersonaPlex . Ad OpenAI says GPT-Live makes decisions multiple times per second about whether to speak, keep listening, pause, interrupt, or call a tool. The model can use filler phrases like "mhmm" or "got it" to signal it's following along. Users can interrupt, take a moment to think, or ask the model to slow down. Ad DEC_D_Incontent-1 Compared to the previous "Advanced Voice Mode," people preferred GPT-Live-1 in 75.7 percent of cases and GPT-Live-1 mini in 69.2 percent, OpenAI says. Beyond the new conversational features, the update adds visual cards that ChatGPT Voice can show during a conversation for things like weather, stock prices, or sports scores. OpenAI also revamped the nine available voices for GPT-Live. At launch, GPT-Live doesn't support Voice with video or screen sharing in ChatGPT, though OpenAI says those features are coming soon. The older Standard and Advanced Voice Mode with these features will stick around for now. Ad The biggest change is the split between live conversation management and actual reasoning. When a question needs a web search, reasoning, or agent-like capabilities, GPT-Live hands it off to a background model. Right now, that's GPT-5.5. While the background model works, GPT-Live keeps the conversation going. OpenAI says the architecture is built so GPT-Live stays connected to the latest frontier models at all times. Users can also pick a reasoning level, from "Instant" for quick answers to "Medium" and "High" for tasks where ChatGPT should spend more time thinking. Ad DEC_D_Incontent-2 This delegation fixes a major weakness of earlier live models, which answered questions based only on their own capabilities. That made them fall far behind frontier models and left them barely usable for serious work. Ad

中文翻译

OpenAI 推出了 GPT-Live,这是一种用于 ChatGPT 实时对话的新语言模型,由于全双工架构,它可以同时听和说。该模型能回应打断,使用“嗯”等填充词让对话自然流畅,并提供两个版本:一个面向付费用户,一个面向免费用户。对于需要网络搜索或逻辑推理的复杂任务,GPT-Live 在后台将请求委托给 GPT-5.5 模型,同时保持对话,显著提高了回答质量。OpenAI 推出 GPT-Live,一种采用全双工架构的新一代语音模型。该模型可以同时听和说,并将复杂任务交给后台的 GPT-5.5 处理。通过 GPT-Live,OpenAI 推出了一类新的 AI 语音模型,旨在让与 ChatGPT 的对话更像真实交流。两个版本立即在全球范围内推出。GPT-Live-1 面向 Go、Plus 和 Pro 计划的付费用户,而 GPT-Live-1 mini 在免费账户上可用。两者均适用于 iOS、Android 和 ChatGPT.com。OpenAI 计划很快提供 API 访问,开发者可以通过表格注册。与以前用户说话、AI 回答的僵硬来回模式不同,GPT-Live 采用全双工架构。该模型可以同时听和说,就像人与人之间的真实对话。今年早些时候,Nvidia 发布了类似的开源模型 PersonaPlex。Ad OpenAI 表示 GPT-Live 每秒多次决定是否说话、继续听、暂停、打断或调用工具。模型可以使用“嗯”或“知道了”等填充短语表示在跟随。用户可以打断、花点时间思考或让模型放慢语速。Ad DEC_D_Incontent-1 与之前的“高级语音模式”相比,OpenAI 表示用户在 75.7% 的情况下更喜欢 GPT-Live-1,在 69.2% 的情况下更喜欢 GPT-Live-1 mini。除了新的对话功能,更新还增加了视觉卡片,ChatGPT 语音可以在对话中显示天气、股票价格或体育比分等信息。OpenAI 还改版了 GPT-Live 的九种可用语音。发布时,GPT-Live 不支持视频或屏幕共享,但 OpenAI 表示这些功能即将推出。带有这些功能的旧版标准语音模式和高级语音模式将暂时保留。Ad 最大的变化是实时对话管理与实际推理之间的分离。当问题需要网络搜索、推理或类似代理的能力时,GPT-Live 将其交给后台模型。目前,那是 GPT-5.5。当后台模型工作时,GPT-Live 保持对话进行。OpenAI 表示该架构的设计使 GPT-Live 始终保持与最新前沿模型的连接。用户还可以选择推理级别,从快速回答的“即时”到需要更多思考时间的“中”和“高”。Ad DEC_D_Incontent-2 这种委托修复了早期实时模型的一个主要弱点,即仅根据自己的能力回答问题。这使得它们远远落后于前沿模型,并且几乎无法用于严肃工作。Ad

核心信息

OpenAI 推出新版语音模型 GPT-Live,采用全双工架构,支持同时听和说,能处理打断、使用填充词,并将复杂任务委托给后台 GPT-5.5 处理。提供付费和免费两个版本,用户偏好率超 69%。

  • OpenAI 推出新版语音模型 GPT-Live,采用全双工架构,支持同时听和说,能处理打断、使用填充词,并将复杂任务委托给后台 GPT-5.5 处理。提供付费和免费两个版本,用户偏好率超 69%。
  • 原贴提到:OpenAI has introduced GPT-Live, a new language model for live conversati
  • 来源:the-decoder.com

详细解读

这是什么信号:OpenAI 推出 GPT-Live,标志着 AI 语音交互从“轮流说话”迈向“实时对话”。全双工架构让 AI 能同时听和说,模拟人类交流中的打断、填充词等自然行为。这是 AI 交互范式的重大升级,尤其对实时语音助手的用户体验有质变影响。

为什么重要:传统语音助手(如 Siri、Alexa)需要等待用户说完才能响应,体验像“按键对讲机”。GPT-Live 通过每秒多次决策,实现类似真人对话的流畅感。它还引入“后台模型委托”机制,将复杂推理任务交给更强大的 GPT-5.5,解决了此前实时模型推理能力弱的问题,使语音助手不仅能闲聊,还能处理严肃工作。

对谁有价值:付费用户可获得完整版 GPT-Live-1,体验更真实;免费用户也能通过 mini 版感受全双工交互。开发者关注即将开放的 API,可构建语音驱动的应用(如客服、教育、陪伴)。企业可探索将 GPT-Live 集成到电话系统、智能硬件中,提升交互自然度。

可以怎么行动:1. 立即在 ChatGPT 应用中试用 GPT-Live(付费用户选择 GPT-Live-1,免费用户选择 GPT-Live-1 mini),体验打断、调整语速等交互。2. 开发者填写 OpenAI 表格申请 API 访问,提前规划语音应用场景。3. 关注 Nvidia 的 PersonaPlex 等开源替代方案,进行成本对比。

风险或限制:目前 GPT-Live 不支持视频和屏幕共享,且依赖后台模型(GPT-5.5)的推理能力,网络延迟可能影响体验。全双工语音伴随隐私风险(持续录音),需关注 OpenAI 的数据处理政策。另外,费用分层可能导致免费用户体验受限。

信息差价值

这条内容的真正价值,不只是“有人发布了一个新功能”,而是它揭示了 the-decoder.com 背后的产品方向、工作流变化或竞争信号。对 OPC 来说,这种信息可以转化成持续追踪的栏目选题。

如果把《ChatGPT 现在可以同时听和说,让 AI 对话更人性化》放到你的内容系统里,它最大的价值在于帮助读者更快看懂“为什么值得关注”,而不是只看到一条碎片化动态。

参考来源

AI SUMMARY

这篇文章回答了什么

ChatGPT 现在可以同时听和说,让 AI 对话更人性化主要讲什么?

OpenAI 推出新版语音模型 GPT-Live,采用全双工架构,支持同时听和说,能处理打断、使用填充词,并将复杂任务委托给后台 GPT-5.5 处理。提供付费和免费两个版本,用户偏好率超 69%。

这篇文章最值得关注的要点是什么?

OpenAI 推出新版语音模型 GPT-Live,采用全双工架构,支持同时听和说,能处理打断、使用填充词,并将复杂任务委托给后台 GPT-5.5 处理。提供付费和免费两个版本,用户偏好率超 69%。;原贴提到:OpenAI has introduced GPT-Live, a new language model for live conversati;来源:the-decoder.com

这篇文章和哪些AI专题相关?

它适合放在AI副业、AI工具专题里阅读。 关联原因:这篇内容命中「项目、小生意、变现」等主题信号。;这篇内容命中「自动化、模型」等主题信号。

阅读这篇文章建议先理解哪些关键词?

建议先理解AI工具、工具、自动化、模型、Cursor这些关键词,再结合正文判断工具、机会或风险是否值得进入自己的工作流。

上一篇 GitHub Copilot 在 Visual Studio Code 中的更新——2026年6月发布 下一篇 我们即将在API中推出GPT-Live-1和GPT-Live-1 mini