觉
AI觉醒星球
Awakening is here
Knowledge File / AI技能杠杆
2026-04-29 2 浏览 免费阅读

趋势解读:LLM 0.32a0 is a major backwards-compatible refactor,提升开发者接入体验

LLM Python 库发布 0.32a0 alpha 版本,进行向后兼容的重大重构,将模型输入改为消息序列,输出改为类型化流式部分,以适应当今多模态、工具调用等复杂场景。

SOURCE / AI技能杠杆 MIN / 4 ACCESS / 免费阅读 POST / 2026-04-29 19:01:47

原贴

查看原文
作者:Simon Willison 来源站点:simonwillison.net 原贴时间:
趋势解读:LLM 0.32a0 is a major backwards-compatible refactor,提升开发者接入体验

原文

I just released LLM 0.32a0 , an alpha release of my LLM Python library and CLI tool for accessing LLMs, with some consequential changes that I've been working towards for quite a while. Previous versions of LLM modeled the world in terms of prompts and responses. Send the model a text prompt, get back a text response. import llm model = llm . get_model ( "gpt-5.5" ) response = model . prompt ( "Capital of France?" ) print ( response . text ()) This made sense when I started working on the library back in April 2023. A lot has changed since then! LLM provides an abstraction over thousands of different models via its plugin system . The original abstraction - of text input that returns text output - was no longer able to represent everything I needed it to. Over time LLM itself has grown attachments to handle image, audio, and video input, then schemas for outputting structured JSON, then tools for executing tool calls. Meanwhile LLMs kept evolving, adding reasoning support and the ability to return images and all kinds of other interesting capabilities. LLM needs to evolve to better handle the diversity of input and output types that can be processed by today's frontier models. The 0.32a0 alpha has two key changes: model inputs can be represented as a sequence of messages, and model responses can be composed of a stream of differently typed parts. Prompts as a sequence of messages LLMs accept input as text, but ever since ChatGPT demonstrated the value of a two-way conversational interface, the most common way to prompt them has been to treat that input as a sequence of conversational turns. The first turn might look like this: user: Capital of France? assistant: (The model then gets to fill out the reply from the assistant.) But each subsequent turn needs to replay the entire conversation up to that point, as a sort of screenplay: user: Capital of France? assistant: Paris user: Germany? assistant: Most of the JSON APIs from the major vendors follow this pattern. Here's what the above looks like using the OpenAI chat completions API, which has been widely imitated by other providers: curl https://api.openai.com/v1/chat/completions \ -H " Authorization: Bearer $OPENAI_API_KEY " \ -H " Content-Type: application/json " \ -d ' { "model": "gpt-5.5", "messages": [ { "role": "user", "content": "Capital of France?" }, { "role": "assistant", "content": "Paris" }, { "role": "user", "content": "Germany?" } ] } ' Prior to 0.32, LLM modeled these as conversations: model = llm . get_model ( "gpt-5.5" ) conversation = model . conversation () r1 = conversation . prompt ( "Capital of France?" ) print ( r1 . text ()) # Outputs "Paris" r2 = conversation . prompt ( "Germany?" ) print ( r2 . text ()) # Outputs "Berlin" This worked if you were building a conversation with the model from scratch, but it didn't provide a way to feed in a previous conversation from the start. This made tasks like building an emulation of the OpenAI chat completions API much harder than they should have been. The llm CLI tool worked around this through a custom mechanism for persisting and inflating conversations using SQLite, but that never became a stable part of the LLM API - and there are many places you might want to use the Python library without committing to SQLite as the storage layer. The new alpha now supports this: import llm from llm import user , assistant model = llm . get_model ( "gpt-5.5" ) response = model . prompt ( messages = [ user ( "Capital of France?" ), assistant ( "Paris" ), user ( "Germany?" ), ]) print ( response . text ()) The llm.user() and llm.assistant() functions are new builder functions designed to be used within that messages=[] array. The previous prompt= option still works, but LLM upgrades it to a single-item messages array behind the scenes. You can also now reply to a response, as an alternative to building a conversation: response2 = response . reply ( "How about Hungary?" ) print ( response2 ) # Default __str__() calls .text() Streaming parts The other major new interface in the alpha concerns streaming results back from a prompt. Previously, LLM supported streaming like this: response = model . prompt ( "Generate an SVG of a pelican riding a bicycle" ) for chunk in response : print ( chunk , end = "" ) Or this async variant: import asyncio import llm model = llm . get_async_model ( "gpt-5.5" ) response = model . prompt ( "Generate an SVG of a pelican riding a bicycle" ) async def run (): async for chunk in response : print ( chunk , end = "" , flush = True ) asyncio . run ( run ()) Many of today's models return mixed types of content. A prompt run against Claude might return reasoning output, then text, then a JSON request for a tool call, then more text content. Some models can even execute tools on the server-side, for example OpenAI's code interpreter tool or Anthropic's web search . This means the results from the model can combine text, tool calls, tool outputs and other formats. Multi-modal output models are starting to emerge too, which can return images or even snippets of audio intermixed into that streaming response. The new LLM alpha models these as a stream of typed message parts. Here's what that looks like as a Python API consumer: import asyncio import llm model = llm . get_model ( "gpt-5.5" ) prompt = "invent 3 cool dogs, first talk about your motivations" def describe_dog ( name : str , bio : str ) -> str : """Record the name and biography of a hypothetical dog.""" return f" { name } : { bio } " def sync_example (): response = model . prompt ( prompt , tools = [ describe_dog ], ) for event in response . stream_events (): if event . type == "text" : print ( event . chunk , end = "" , flush = True ) elif event . type == "tool_call_name" : print ( f" \n Tool call: { event . chunk } (" , end = "" , flush = True ) elif event . type == "tool_call_args" : print ( event . chunk , end = "" , flush = True ) async def async_example (): model = llm . get_async_model ( "gpt-5.5" ) response = model . prompt ( prompt , tools = [ describe_dog ], ) async for event in response . astream_events (): if event . type == "text" : print ( event . chunk , end = "" , flush = True ) elif event . type == "tool_call_name" : print ( f" \n Tool call: { event . chunk } (" , end = "" , flush = True ) elif event . type == "tool_call_args" : print ( event . chunk , end = "" , flush = True ) sync_example () asyncio . run ( async_example ()) Sample output (from just the first sync example): My motivation: create three memorable dogs with distinct “cool” styles—one cinematic, one adventurous, and one charmingly chaotic—so each feels like they could star in their own story. Tool call: describe_dog({"name": "Nova Jetpaw", "bio": "A sleek silver-gray whippet who wears tiny aviator goggles and loves sprinting along moonlit beaches. Nova is fearless, elegant, and rumored to outrun drones just for fun."} Tool call: describe_dog({"name": "Mochi Thunderbark", "bio": "A fluffy corgi with a dramatic black-and-gold bandana and the confidence of a rock star. Mochi is short, loud, loyal, and leads a neighborhood 'security patrol' made entirely of squirrels."} Tool call: describe_dog({"name": "Atlas Snowfang", "bio": "A massive white husky with ice-blue eyes and a backpack full of trail snacks. Atlas is calm, heroic, and always knows the way home—even during blizzards, fog, or confusing camping trips."} At the end of the response you can call response.execute_tool_calls() to actually run the functions that were requested, or send a response.reply() to have those tools called and their return values sent back to the model: print ( response . reply ( "Tell me about the dogs" )) This new mechanism for streaming different token types means the CLI tool can now display "thinking" text in a different color from the text in the final response. The thinking text goes to stderr so it won't affect results that are piped into other tools. This example uses Claude Sonnet 4.6 (with an updated streaming event version of the llm-anthropic plugin) as Anthropic's models return their reasoning text as part of the response: llm -m claude-sonnet-4.6 ' Think about 3 cool dogs then describe them ' \ -o thinking_display 1 You can suppress the output of reasoning tokens using the new -R/--no-reasoning flag. Surprisingly that ended up being the only CLI-facing change in this release. A mechanism for serializing and deserializing responses As mentioned earlier, LLM has quite inflexible code at the moment for persisting conversations to SQLite. I've added a new mechanism in 0.32a0 that should provide Python API users a way to roll their own alternative: serializable = response . to_dict () # serializable is a JSON-style dictionary # store it anywhere you like, then inflate it: response = Response . from_dict ( serializable ) The dictionary this returns is actually a TypedDict defined in the new llm/serialization.py module. What's next? I'm releasing this as an alpha so I can upgrade various plugins and exercise the new design in real world environments for a few days. I expect the stable 0.32 release will be very similar to this alpha, unless alpha testing reveals some design flaw in the way I've put this all together. There's one remaining large task: I'd like to redesign the SQLite logging system to better capture the more finely grained details that are returned by this new abstraction. Ideally I'd like to model this as a graph, to best support situations like an OpenAI-style chat completions API where the same conversations are constantly extended and then repeated with every prompt. I want to be able to store those without duplicating them in the database. I'm undecided as to whether that should be a feature in 0.32 or I should hold it for 0.33. Tags: projects , python , ai , annotated-release-notes , generative-ai , llms , llm

中文翻译

我刚刚发布了 LLM 0.32a0,这是我用于访问 LLM 的 LLM Python 库和 CLI 工具的一个 alpha 版本,其中包含一些我长期以来一直在努力的重要更改。

核心信息

LLM Python 库发布 0.32a0 alpha 版本,进行向后兼容的重大重构,将模型输入改为消息序列,输出改为类型化流式部分,以适应当今多模态、工具调用等复杂场景。

  • 输入改为消息序列,支持多轮对话直接传入。
  • 输出拆分为类型化流式部分,含文本、工具等。
  • 向后兼容,现有代码无需修改。
  • 简化对话历史管理,无需外部存储。
  • 适配多模态模型,提升开发者体验。

详细解读

这是什么信号?LLM 0.32a0 将输入从单一文本提示改为消息序列,输出从单一文本改为类型化流式部分(文本、工具调用、图像等)。这是对 ChatGPT 以来对话接口和模型能力爆炸的回应,标志着 LLM 库从“文本进文本出”的抽象升级为“多模态消息进、流式部分出”的新范式。

为什么重要?向后兼容保证了现有用户的代码无需修改即可运行,同时新接口极大简化了多轮对话、工具链集成和多模态输出的处理。例如,开发者不再需要自己维护 SQLite 存储来恢复历史对话,直接传入 message 列表即可。这对基于 LLM 构建应用的工具链开发者是及时雨。

对谁有价值?Python 后端开发者、AI 应用工程师、RAG 系统构建者、需要调用多种模型的库使用者。尤其是那些希望统一 OpenAI、Anthropic 等多家 API 接口的团队,可以基于 LLM 构建一层抽象。

可以怎么行动?试用 alpha 版本,将现有项目中的 model.prompt() 调用迁移到 messages= 方式;利用新的 stream_events() 处理工具调用和混合输出;在测试环境中验证兼容性,并关注 0.32 正式版发布。

风险或限制。alpha 版本不稳定,API 可能在正式版前调整;stream_events() 接口尚未完全覆盖所有模型能力;对历史对话的完全兼容仍需社区测试。

信息差价值

信息差价值:多数开发者仍在使用旧式单一文本接口,而 LLM 库的这一重构提前揭示了业界主流 API 的演进方向——从 prompts/responses 转向多轮消息和类型化流式输出。掌握这一变化,可以预判未来一年内各大厂商 API 的升级路径,避免在旧接口上重复投资。

业务启发:如果你的产品依赖多个 LLM 供应商,可以基于 LLM 0.32a0 的新接口构建一个统一的“消息适配层”,将内部所有模型调用统一为 messages 格式。这不仅能降低切换成本,还能更容易地接入未来出现的新能力(如图像输出、工具执行)。对于 SaaS 工具,可借此打造“一次接入,多模型承载”的卖点。

可沉淀动作:立即 fork LLM 仓库并运行测试套件,评估新接口与现有代码的兼容性。开发一个内部“消息日志”中间件,自动记录 messages 和 stream_events,用于调试和合规审计。编写一份迁移指南,并在团队内进行技术分享,推动逐步过渡到新范式。

参考来源

AI SUMMARY

这篇文章回答了什么

趋势解读:LLM 0.32a0 is a major backwards-compatible refactor,提升开发者接入体验主要讲什么?

LLM Python 库发布 0.32a0 alpha 版本,进行向后兼容的重大重构,将模型输入改为消息序列,输出改为类型化流式部分,以适应当今多模态、工具调用等复杂场景。

这篇文章最值得关注的要点是什么?

LLM Python 库发布 0.32a0 alpha 版本,进行向后兼容的重大重构,将模型输入改为消息序列,输出改为类型化流式部分,以适应当今多模态、工具调用等复杂场景。;输入改为消息序列,支持多轮对话直接传入。;输出拆分为类型化流式部分,含文本、工具等。;向后兼容,现有代码无需修改。

这篇文章和哪些AI专题相关?

它适合放在AI工具、Agent工作流、AI超级个体专题里阅读。 关联原因:这篇内容命中「工具、自动化、模型」等主题信号。;这篇内容命中「Agent、工作流」等主题信号。;这篇内容命中「技能」等主题信号。

阅读这篇文章建议先理解哪些关键词?

建议先理解AI工具、工具、自动化、模型、Cursor这些关键词,再结合正文判断工具、机会或风险是否值得进入自己的工作流。

上一篇 版本更新:crewAIInc/crewAI 1.14.4a1,解读最新研究结论 下一篇 趋势解读:OpenAI models,Codex,and Managed Agents come to,解读最新 AI 进展
北竹游乐场 免费玩小游戏 免费玩