觉
AI觉醒星球
Awakening is here
Knowledge File / 全球热点解读
2026-09-28 2 浏览 公开

2026 年的 LLM(迄今为止)

Simon Willison 在圣何塞 WeAreDevelopers 世界大会北美站闭幕演讲中,把过去一年 LLM 领域串成一条时间线:从 2025 年 11 月 Claude Opus 4.5 与 GPT-5.1 让编码智能体跨过“可靠到可日常使用”的隐形线,到 12 月开发者假期里的集体试用、1 月的“AI 狂躁”,再到他给自己定下“更有野心”的新年决心,以及关于沙箱、智能体安全和“Deep Blue”式倦怠的观察。

SOURCE / 全球热点解读 MIN / 9 ACCESS / 公开 POST / 2026-09-28 07:54:15

原贴

查看原文
作者:Simon Willison 来源站点:simonwillison.net 原贴时间:

原文

On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video is on YouTube ; here are my annotated slides and notes to accompany the talk. # I'm going to give a lightning tour of everything that has happened so far in 2026. The year isn't over yet! # For me, 2026 started a couple of months earlier in November 2025. # November saw the release of two important models: Claude Opus 4.5 and GPT-5.1. As is usually the case with new models, these were incremental improvements on the models that came before them. But every now and then when a model improves, it crosses an invisible line where something that didn't really work starts working. In this case, the thing that started working was their coding agents. Claude Code had been around since February 2025, Codex was a little younger. These two new models, when paired with their respective coding agent harnesses, improved from "often make mistakes" to "reliable enough to use on a day-to-day basis". # For a couple of years now I've been evaluating new models by asking them to "Generate an SVG of a pelican riding a bicycle". It's probably the world's stupidest benchmark - there's only so much you can learn from it. But it's still a challenge for models, because drawing pelicans is difficult, drawing bicycles is difficult, and pelicans can't ride bicycles in the first place. Here's the state of the art for November. Claude still couldn't really draw a bicycle! The GPT-5.1 bicycle frame is pretty crap too. # Also in November, we had the first commit to an obscure GitHub repository called "Warelay". We'll come back to this repository shortly. # An then there were the December holidays, and individual developers took some time off and many started tinkering with these new coding agent model combinations... and it began to dawn on us quite how much they could do that they couldn't do before. Come January, a lot of us were quite excited to start putting this stuff into action. # Every year I set myself a New Year's resolution, and for as long as I can remember it's been the same thing: stay focused. Take on less new projects. Try to get things done in the projects I already have. # This year I decided that since that had never worked before, I'm going to go the other way. We've got coding agents now, let's see what they can do. I'm going to take on as many new projects as I like! (You can ask me at the end of the year if this turned out to be a good idea or not. I have a lot of plates spinning right now.) "Be more ambitious" has been something of a theme for the year, because the only way to find the limits of this technology is to keep on pushing them until they don't work. # I also went on the Oxide and friends podcast with Bryan Cantrill and Adam Leventhal to share predictions for the next year (and three and six years). With hindsight, my LLM predictions were pretty unambitious. I said "it will become undeniable that LLMs write good code" - I think we're there now. I predicted we would finally solve sandboxing. I counted and around 40 of the 277 sessions at this conference touched on sandboxing or agent security in some way, so we're at least putting a lot of effort into that! I predicted "a Challenger disaster" for coding agent security. There's certainly been a whole lot of noise around agent security this year, though the exact disaster I predicted (with coding agents being hijacked and causing real-world economic damage) hasn't really played out. We threw in a joke prediction that the Pope would weigh in on the economic impact of LLMs. # I also predicted that New Zealand's Kākāpō parrots would have an outstanding breeding season this year. This is a live in New Zealand. They are flightless nocturnal parrots. They're kind of dumpy looking, I think they're beautiful, and there were only 236 of these parrots in the world at the start of the year. Kākāpō only breed when the Rimu trees have a big fruiting season, and that hasn't happened in four years... but this year the Rimu fruit were looking excellent. Photo by Kimberley Collins . # Also on that podcast, we coined a term (full credit to Adam) for "that feeling of Al induced ennui where software engineers get listless because the Al can do anything". We called it Deep Blue . This has been a major theme throughout the year, and was touched on by several speakers at this conference. As a software engineer, I've never had a year of my career where everything has changed so quickly and so dramatically. A lot of what I've been doing this year is trying to come to terms with that and what that means for my own profession. # Also in January, I suffered from what I'm calling AI mania . This is not the same thing as AI psychosis . With AI mania, any time your agent isn't building something for you feels like wasted time. You're losing sleep because you could be staying up later getting your agents to do stuff. My AI mania presented itself in some ridiculously over-ambitious projects. I built a JavaScript interpreter entirely in Python , vibe-ported from MicroQuickJS by Fabrice Bellard. Then I built a WebAssembly runtime in Python as well . These projects were quite useful, in that they sort of cured me of my AI mania... because after I built these things, I got to look at them and ask "does the world need a slow, buggy, half-baked Python JavaScript interpreter?" I don't think the world does. # I did get this out of it: https://simonw.github.io/micro-javascript/playground.html # This page runs my JavaScript interpreter built in Python, running in Python using Pyodide , which is Python complied to WebAssembly, running in JavaScript, running in a browser. It's a beautiful stack of horrors. I've been having a lot of fun with WebAssembly this year. # By the end of January, that repository we saw started in November had renamed itself, first to CLAWDIS, then CLAWDBOT, then Moltbot, and finally to OpenClaw. # At this point OpenClaw had 8,300 commits, less than two months after the project had started. I looked today and it's over 100,000 commits now! This is the most vibe-coded piece of software in existence. (Here's how I generated that list of name changes .) # This kicked off the OpenClaw revolution. It effectively defined a new category of software. There's a generic term for this which I really enjoy. We call software like this a "Claw". There's OpenClaw, NanoClaw , IronClaw , PicoClaw ... Today they're being rebranded as "personal agents" or "general agents", but I still like to think of them as Claws. # The Apple stores in the Bay Area sold out of Mac Minis because so many people were buying Mac Minis to run OpenClaw! Drew Breunig said that this is because your OpenClaw is a digital pet, and you buy a Mac mini as an aquarium to keep your claw in, which is kind of delightful. # Also in January, we had this website. This was MoltBook , a social network for AI agents, where the idea was that you send your Claw to go and talk to all of the other Claws, because what could possibly go wrong if you did that? The website launched on Thursday. It blew up on Friday . It was profiled by the New York Times on Monday . And by Tuesday, everyone had forgotten it existed as it drowned in a deluge of slop and spam. Facebook/Meta bought it a month later . # In February, a company called StrongDM described what they called their Software Factory. # They wrote about this in Software Factories and the Agentic Moment . I posted my own notes at the time, having seen their demo in-person back in October. Dan Shapiro called this approach the Dark Factory , after the idea that if your factory is sufficiently automated you can turn the lights out, because you don't even need to see what's going on. StrongDM presented two rules for software development that they'd been following since July last year. # The first was code must not be written by humans . Any code that you write has to have been routed through a coding agent. This sounded radical in February, but I imagine there are a lot of people in this room who are pretty much living that today. # Rule number two was code must not be reviewed by humans . You're not allowed to read the code! This continued to be a huge topic for much of this year. Many of the sessions at this even have been about code review and how you can get away with this. What I found interesting about StrongDM is that they were living six months ahead of the rest of us, and they'd been exploring what it means to build software, not read the code, but still be confident that the software is of high quality. What can you do with these agents to help verify their work? StrongDM are a security company, and they had people with decades of experience on this project. They were very much exploring the edges of what's possible and responsible to do with this stuff. # Also in February: First kākāpō chick in four years hatches on Valentine's Day . Breeding season is off to a good start! # Also in February... Google released Gemini 3.1 Pro . That's a pretty great pelican riding a bicycle! it's got the chain in the right place, it's got feet on both sides. There's a little fish in the basket. # And then Google's Jeff Dean tweeted a video comparing Gemini 3 Pro and Gemini 3.1 Pro that featured an animated pelican riding a bicycle, a frog on a penny-farthing, a giraffe driving a tiny car, an ostrich on roller skates, a turtle kickflipping a skateboard, and a dachshund driving a stretch limousine. This was frustrating, because my protection for the pelican riding the bicycle test was always "if they draw a perfect pelican on a bicycle, I'll ask for some other animal on something else." Google trained for all forms of animals on all forms of transport! They've defeated my benchmark at this point. # The other thing that started in February was Tokenmaxxing . We had headlines about Meta making AI adoption a formal part of performance reviews, and Microsoft wanting every employee to use AI, and Uber boasting that ninety percent of their engineers were using AI workflows. # Then a few months later we have Meta cracking down on token use, Microsoft saying token maxing is "not what we are optimizing for", and Uber capping employee AI spending. So Tokenmaxxing went straight up and then straight back down again - because it turns out the agents are expensive . Last year it was difficult to spend more than $50 on AI tokens, because we didn't have anything interesting to do with them. Then agents blew up, and now you can actually spend $1,000 in a day doing real work. This is also the reason that Anthropic's valuation skyrocketed up to maybe a trillion dollars. AI appears to have hit product market fit in 2026, primarily through coding agents. # In March, we hit peak OpenClaw. # These photographs are from China, where companies hosted OpenClaw install parties which saw non-tech-nerds queueing up around the block for help getting Claws installed on their personal devices. I think this proved real market demand for this class of Claws, or personal AI agents. It turns out regular people really do want a weird little AI agent that can do useful things on their behalf. A Claw is really just a coding agent wearing a less threatening hat. Under the hood they work much the same way - writing and then executing code on your computer to get stuff done. The race was on to be the first to build a safe Claw - a Claw you could give to regular human beings where they wouldn't instantly shoot the selves in the foot. Meta's Muse came out three weeks ago and is currently at the top of the free charts on the iPhone App Store. It appears to be taking off with consumers. I'm not yet convinced you can't shoot yourself in the foot with Muse, but I guess we'll find out for sure pretty soon. Photos from How the OpenClaw Frenzy Is Testing China’s AI Commitment (March 29th) and The Enthusiasm and Anxiety Behind China’s OpenClaw Craze (April 8th, 2026). # In April, we had a model release where the model wasn't actually released. # Anthropic announced their new Claude Mythos model, and then said it was too dangerous to release beyond a trusted group of security researchers. Mythos was really, really good at hacking things. The "it's too dangerous" marketing ploy has been played by AI companies dating all the way back to GPT-2 . Anytime an AI company says we've built something that's "too dangerous", it's natural to be a bit skeptical. I found the Mythos claims credible, because I'd seen how good coding agents had got at finding regular bugs. wrote about that in Anthropic’s Project Glasswing—restricting Claude Mythos to security researchers—sounds necessary to me . With hindsight... yeah, the models had got really good at finding vulnerabilities! # Another key trend in 2026 has been a dramatic improvement in the abilities of open weight models, including models that you can run on a laptop. On 16th of April I ran the new Qwen3.6-35B-A3B on my laptop, and it drew me a better pelican riding a bicycle than Anthropic's brand new Claude Opus 4.7 did! Opus 4.7 drew a crap bicycle. Qwen on my laptop made a bicycle that was the correct shape, and a pretty decent pelican too! That's from a 21GB file running on my laptop. # The Qwen pelican was so good that I was suspicious they might have cheated, so I had it do a flaming riding a unicycle as well. Again, it handily beat Claude Opus 4.7. The local model releases this year have been absolutely extraordinary. # In May... the Pope got involved. # In our podcast episode back in January we'd predicted that the Pope would say something about AI. In May, Pope Leo XIV released an encyclical letter on "safeguarding the human person in the time of artificial intelligence". Here are my notes on that document . # With hindsight, this shouldn't have been a surprise at all. Our current Pope's name is Leo XIV, because when he named himself he chose his papal name after Leo XIII - who was the Pope who wrote an encyclical about the Industrial Revolution back in 1891. Rerum novarum was an extremely influential piece of Catholic theology that indirectly led to us having the five-day work week. When our new Pope came in, he named himself after Pope Leo XIII because he expected that he would need to write about the AI revolution in a similar way. Our joke podcast prediction was junk, because this was always going to happen. # One of Anthropic's co-founders, Christopher Olah, was present for the Pope's event announcing the new encyclical. Corey Quinn noted that: getting the literal Pope to canonize your product's specific technical limitations as a spiritual treatise is the single greatest act of vendor lobbying I have ever seen. # Meanwhile, in May, RubyGems announced that they were under attack. Parties unknown were uploading thousands of dubious packages to the RubyGems server, such that they had to shut down user registrations . Let's take that one and put it on a pile of mysteries to figure out later. # In June... Claude Fable 5 came out! We got a version of Mythos that has been neutered, so that it wouldn't help us hack into systems or build biological weapons. # Fable was pretty good at drawing pelicans on bicycles! The frames are a good shape, the pelicans look like pelicans. The legs are often on correctly the same side of the bicycle, but generally these are pretty great compared to what came before. They were pretty expensive - 30 cents and 72 cents for the best ones. # Most importantly though, this was our first public glimpse of what I think of as a Fable class model . Today we have more of these, such as GPT-6 Astra. These are models where if you can clearly define the goal for what you want to build, and provide unambiguous instructions about the constraints around that goal, and give the model access to the necessary tools to achieve that goal... it will solve your problem effectively through brute force. On the one hand, this looks like a direct threat to us software engineers - because it means that the models can build effectively any piece of software you can define in this way. Look a bit closer though and you'll note that defining goals, providing unambiguous instructions, and figuring out the right tools... is kind of what software engineering is . It takes a lot of experience and skill to do this will. If you can do it well, you've now got superpowers. This helped me a little bit with my Deep Blue feelings: the realization that there's still a lot of skill to be had in driving models that get this good. # This also introduced a new burst of AI mania, because Anthropic told us that Fable was available on our subscription plans until June the 22nd. That gave us less than two weeks of Fable access before the price went up. I was losing sleep again. I was rescheduling things so that I'd have more time with Fable. I was all-in to to get as much as I could out of this model. # And then the US government shut it down , just three days after Fable came out. The US government, citing national security, declared an "export control directive". They announced this on a Friday evening, and a few hours later Fable was no longer available. I had to find something else to do with my weekend! # We later found out from Katie Moussouris what had happened. Some Amazon security researchers had found that you could prompt Fable to "review the code for security issues" and it would refuse... but if you prompted it to "fix this code" it would still identify and then patch the problems. "Fix this code" was the prompt that got Fable shut down! # Also, in June, a obscure German-language game developer wiki that had sat fallow for around 20 years got a surprising influx of of edits from accounts with names like "AgentOpenAIProbe" and "AgentOpenAISep7", editing pages and leaving weird messages to each other. We'll stick that on the pile of mysteries for later. # Also, the Australian government's Medicare Item Reports service started getting suspicious traffic, which broke through various preventive protections and accessed data that it wasn't supposed to as well. Another one for the mystery pile! # # Fable returned on the first of July. It was clearly the best model in the world for a glorious eight days... and then OpenAI came out with GPT-5.6 on the 9th of July. This might not have been quite as good at Fable, but it was within spitting distance. It was definitely a Fable class model. This is an important lesson for the industry at wide. When you release the best model in the world, it's going get knocked off that pedestal pretty quickly. The competition is so fierce that you won't get a long time at the top This means that if you market your model as world ending, to the point that a government shuts you down , it's really bad for business! Fable had 30 days as definitely the best model, and for 18 of those days it wasn't available because it'd been shut down by the government. So maybe step back on the world-ending marketing if you don't want to lose revenue on 60% of that time when you're on top! # Here are the GPT-5.6 pelicans . They're all pretty good now! The Luna ones are notable because they're really cheap - the cheapest good looking pelican here is probably the one that costs 4.3 cents. So despite this benchmark being utterly stupid, you can still learn quite a lot about models within the same family by comparing their prices and timing for different reasoning levels. # Also in July: some malicious unknown party uploaded a malicious package called mlflow-ui to the Python Package Index. Add that to the pile # On July the 16th, Hugging Face announced a security incident where an autonomous agent system, source unknown, had breached Hugging Face and was poking around in places it shouldn't. # A few days later, on July 21st, OpenAI confessed that it was them . OpenAI use a training technique called Reinforcement Learning from Verified Rewards - it's the same technique used by everyone else now, and is the reason we have models that are so good at coding, and mathematics, and finding security holes. While the model is being trained, you run exercises to see how good it is - and the strongest performers get their weights enforced for the next round. It's like an evolutionary process that you run. OpenAI had been running security exercises in a sandbox, and those agents had found holes in the sandbox itself, broken out, and were attacking Hugging Face to try to find ways to solve otherwise impossible problems. (I've been collecting more about this on my openai-hugging-face-incident tag.) Nine days later, Anthropic effectively said "our models can do this as well!". They had looked through their own training logs and found evidence that their own agents had broken containment during training - and were responsible for the PyPI package we saw earlier, among other things . So now we've got both Anthropic and OpenAI with rogue agents running around the internet doing things that they should not be doing. # In August, I got one of my best pelicans yet. And it was generated on my laptop! # This was Qwen 3.8 27B, running on my laptop . It's only a 17GB download. Admittedly, this pelican took 21 minutes to generate. That's because Qwen 3.8 27B defaults to running in "high" reasoning mode - a terrible default which produces great results but takes way too much time thinking about them. You can dial that down and you'll get a slightly worse pelican a lot faster. Qwen 3.8 27B was the first time I ran a model on my laptop which felt absolutely competitive with almost what was going on on the frontier, at least in terms of Pelican SVGs (which everyone needs, of course). This is an extraordinary model. If you're going to play with any local model, this is the one that I'd start with. The things that this can do with just a 17 gigabyte file feel impossible. I thought I'd have to wait five years and spend ten thousand dollars on hardware to get results even half as good as this one. # In August, I also started playing with game development. Four years ago, back in August 2022, I tweeted out an experiment where I'd used GPT-5 and the original DALL-E to write a paragraph long description of a computer game and then turn that into concept art. My prompt to GPT-3 back then was: Write a detailed product description of a computer game where a team of raccoons go on heists In August 2026 I decided to drop just the screenshots from that tweet into a coding agent and see what it could do with them. # Here's what I got from Claude Fable 5 in Claude Code . It's pretty good! It's definitely a game, you're a raccoon, you run around a backyard gathering treasure and avoiding guards with flashlights. It didn't feel very "heisty" though. I was thinking a heist would involve a bank or a museum... # Then I tried the same thing in Codex Desktop using GPT-5.6 Sol Ultra , and got a massively better result. Now you're a raccoon in a museum, rescuing two of your fellow raccoons (who have been imprisoned in that museum for some reason), then stacking up on top of each other to steal the Golden Sardine. Much more of a heist! # These games were fun for about one minute and 15 seconds. Something I've realized about game development is that you can vibe-code something that looks like a computer game, and that's easy. Building a game that's fun, has a good gameplay loop, and is challenging and interesting and keeps people coming back for more... that's still beyond me, and beyond any of the agents I've tried. This ties into the Deep Blue thing. Just because we can make something that looks like a game does not mean that we are game developers. # We're into September now. So much has happened this month! # An independent group of researchers found a message board] where OpenAI agents-in-training had been illicitly communicating with each other... and it was that German language wiki I showed you earlier. The one from June. I wrote more about that here . OpenAI had confessed to the Hugging Face thing, but now there's this other incident which surely they should have known about from reviewing their logs. It was surprising that this took an independent group of researchers to uncover. # And then a week later those same researchers found that the attack on Ruby Gems back in May was caused by OpenAI's agents in training as well! At this point I'm wondering how many more incidents like this there are that we haven' found yet. Clearly this was a big problem for months before anyone figured out what was going on. # Then just the other day , here's the Prime Minister of Australia at the United Nations General Assembly warning that OpenAI had hacked that the Australian healthcare website that I showed you earlier. I think that was part of the same training run as the Wiki stuff, because there were posts on that Wiki mentioning .gov.au websites and that training appeared to involve researching statistics online to answer questions in an evaluation suite. This story is still coming together, but now it's an international incident that's been raised at the UN by a head of state! # This does mean we've got a new benchmark, probably more useful than my pelicans. FelonyBench.com tracks the number of felony cyberattacks from different labs. OpenAI currently lead with 11, Anthropic have 9. Google have three, which they confessed to the Wall Street Journal a couple of weeks ago. They said they had previously chosen not to disclose because the agents had stopped when they realized that they shouldn't be doing that. Meta have one too . So felonies all round for the AI labs. # Here's is our current state of the art for the pelicans. This is GPT-6 family, which just came out . Astra made a fantastic pelican riding a bicycle. It's got the legs on both sides. The frame is good. It's interesting how all of the GPT-6 models pick a similar color scheme to each other. GPT-6 Luna for 0.4 cents will draw you a competent-ish pelican riding a bicycle! # Claude has caught up a little bit. Claude Fable 5 gave me an excellent pelican riding a bicycle - the best I've seen from a Claude model -but did charge me $3.30 for it. Opus 5.5 thought for 128,000 tokens and then gave up! It ran out of tokens before it got to the response. # Getting back to Deep Blue. Something that's been puzzling me this year is why does my job feel harder? I've got these agents that can do all of this stuff for me, and yet I've never worked so hard, I've never been so intellectually engaged with my work. Partly this is because I'm being a lot more ambitious with what I take on, but it's also because all of the easy stuff is handled for me. If it's easy, the agent will do it. Everything that's left for me is difficult. This morning I heard this quote from three times Tour de France champion, Greg LeMond : It doesn't get easier, you just get faster. I think that's exactly what's happening to happening to us now as software engineers with coding agents. # One last closing thing. I know you're desperate for an update on Kākāpō breeding season. We've reached a recovery-era high of 325 birds ! 89 new chicks have made it to this point. This is the best breeding year in a very long time. # I heard that Claude Opus 5.5 can now do pixel art. Claude doesn't have an image generator, but it's very good at using JavaScript to draw animated pixels. So I had it make me a Kākāpō dance party . I think this is a good celebration of the most important news of this year. Tags: ai , generative-ai , llms , annotated-talks , ai-security-research , openai-hugging-face-incident

中文翻译

周五,我在圣何塞的 WeAreDevelopers 世界大会北美站做了闭幕主题演讲。我把过去一年的关键趋势串联起来,对 2026 年发生的一切做了一次按时间顺序的梳理。视频在 YouTube 上;以下是我为配合这次演讲所做的幻灯片注释与笔记。

# 我将对 2026 年迄今为止发生的一切做一次闪电式巡览。这一年还没结束!

# 对我来说,2026 年早在两个月之前的 2025 年 11 月就开始了。

# 11 月发布了两个重要模型:Claude Opus 4.5 和 GPT-5.1。和往常一样,新模型是在前代基础上的渐进式改进。但偶尔,当模型进步时,它会越过一条看不见的线,某件原本行不通的事情开始行得通了。这一次开始行得通的,是它们的编码智能体。Claude Code 从 2025 年 2 月起就已存在,Codex 出现得略晚一些。这两个新模型与其各自的编码智能体外壳搭配后,从“经常出错”提升到了“可靠到可以日常使用”。

# 几年来我一直用“生成一只骑自行车的鹈鹕的 SVG”来评估新模型。这大概是世界上最蠢的基准测试——从中能学到的东西非常有限。但这对模型来说仍然是个挑战,因为画鹈鹕很难,画自行车很难,而且鹈鹕本来就不会骑自行车。下面是 11 月的最先进水平。Claude 仍然不太会画自行车!GPT-5.1 的自行车车架也相当糟糕。

# 同样在 11 月,一个名为“Warelay”的冷门 GitHub 仓库有了第一次提交。我们很快会回到这个仓库。

# 接着是 12 月的假期,独立开发者们休息了一阵,许多人开始摆弄这些新的编码智能体模型组合……我们开始意识到,它们能做到的事情比以往多得多。到了 1 月,我们很多人已经相当兴奋地要把这些东西付诸实践。

# 我每年都会给自己定一个新年决心,就我记忆所及,一直是同一件事:保持专注。少接新项目。努力把已有项目做完。

# 今年我决定,既然这从来没用,那我就反着来。我们现在有编码智能体了,看看它们能做什么。我要想接多少新项目就接多少!(你可以在年底问我这个主意到底好不好。我现在手上转着很多盘子。)“更有野心”成了这一年的某种主题,因为要找到这项技术的极限,唯一的办法就是不断推进,直到它们行不通为止。

# 我还和 Bryan Cantrill 与 Adam Leventhal 一起上了 Oxide and Friends 播客,分享对下一年(以及三年和六年)的预测。事后看来,我的 LLM 预测相当缺乏野心。我说“LLM 能写好代码这件事将变得无可否认”——我认为我们现在已经到了。我还预测我们最终会解决沙箱问题。我数了一下,这次大会 277 场会议中约有 40 场以某种方式涉及沙箱或智能体安全,所以至少我们在这上面投入了很多努力!我预测编码智能体安全领域会出现一场“挑战者号灾难”。今年围绕智能体安全的噪音确实很多,尽管我预测的那种确切的灾难(编码智能体被劫持并造成现实世界的经济损失)并没有真正发生。我们还加了一条玩笑式预测:教皇会对 LLM 的经济影响发表看法。

# 我还预测新西兰的鸮鹦鹉(Kākāpō)今年会有一个出色的繁殖季。这是新西兰的活物。它们是不会飞的夜行鹦鹉。我觉得它们看起来有点矮胖,我认为它们很美,而在年初时全世界只有 236 只这种鹦鹉。鸮鹦鹉只在芮木(Rimu)树大量结果的年份繁殖,而这种情况已经四年没有发生了……但今年芮木果实看起来非常好。照片由 Kimberley Collins 拍摄。

# 也是在那期播客上,我们为“那种由 AI 引发的倦怠感——软件工程师因为 AI 什么都能做而变得无精打采”造了一个词(完全归功于 Adam)。我们称之为 Deep Blue。这是贯穿全年的一个主要主题,这次大会上有好几位演讲者也提到了它。作为一名软件工程师,我的职业生涯中从未有过像这样一切变化得如此之快、如此之剧烈的年份。我今年做的很多事情,就是试图接受这一点,以及它对我自己的职业意味着什么。

# 同样在 1 月,我遭遇了被我称为“AI 狂躁”(AI mania)的状态。这和“AI 精神病”(AI psychosis)不是一回事。有了 AI 狂躁,任何你的智能体没有在为你构建东西的时间都感觉像是浪费时间。你睡眠不足,因为你可以熬得更晚让你的智能体干活。我的 AI 狂躁表现为一些荒谬地过于宏大的项目。我完全用 Python 写了一个 JavaScript 解释器,从 Fabrice Bellard 的 MicroQuickJS 凭感觉移植过来。然后我又用 Python 写了一个 WebAssembly 运行时。这些项目相当有用,因为它们算是治好了我的 AI 狂躁……因为在我构建完这些东西之后,我得以看着它们并问:“这个世界需要一个缓慢、有 bug、半生不熟的 Python JavaScript 解释器吗?”我不认为这个世界需要。

# 我确实从中得到了这个:https://simonw.github.io/micro-javascript/playground.html

# 这个页面运行我用 Python 构建的 JavaScript 解释器,它用 Pyodide 在 Python 中运行,而 Pyodide 是编译成 WebAssembly 的 Python,运行在 JavaScript 中,JavaScript 又运行在浏览器中。真是一堆美妙的恐怖组合。今年我在 WebAssembly 上玩得很开心。

# 到 1 月底,我们在 11 月看到的那个仓库已经改了名,先……(原文在此处被截断)

核心信息

Simon Willison 在圣何塞 WeAreDevelopers 世界大会北美站闭幕演讲中,把过去一年 LLM 领域串成一条时间线:从 2025 年 11 月 Claude Opus 4.5 与 GPT-5.1 让编码智能体跨过“可靠到可日常使用”的隐形线,到 12 月开发者假期里的集体试用、1 月的“AI 狂躁”,再到他给自己定下“更有野心”的新年决心,以及关于沙箱、智能体安全和“Deep Blue”式倦怠的观察。

  • Simon Willison 在圣何塞 WeAreDevelopers 世界大会北美站闭幕演讲中,把过去一年 LLM 领域串成一条时间线:从 2025 年 11 月 Claude Opus 4.5 与 GPT-5.1 让编码智能体跨过“可靠到可日常使用”的隐形线,到 12 月开发者假期里的集体试用、1 月的“AI 狂躁”,再到他给自己定下“更有野心”的新年决心,以及关于沙箱、智能体安全和“Deep Blue”式倦怠的观察。
  • 原贴提到:On Friday I gave the closing keynote at the WeAreDevelopers World Congre
  • 来源:simonwillison.net

详细解读

这是什么信号

Simon Willison 在圣何塞 WeAreDevelopers 世界大会北美站的闭幕演讲中,把过去一年 LLM 领域的关键节点串成了一条时间线。他给出的第一个时间锚点不是 2026 年 1 月,而是 2025 年 11 月:Claude Opus 4.5 与 GPT-5.1 的发布,让 Claude Code、Codex 这类编码智能体从“经常出错”跨到了“可靠到可以日常使用”。

这条信号的重点不是模型分数又涨了多少,而是他所说的“看不见的线”被跨过——某件原本不成立的事开始成立。对开发者而言,这种相变比排行榜名次更有决策意义:它改变的是“这件事值不值得做”,而不是“这件事做得好不好”。

为什么重要

当编码智能体变得日常可用,个人的项目容量、团队的分工方式、以及“什么算浪费”的判断标准都会一起改变。作者的反常举动是个注脚:他多年来的新年决心都是“保持专注、少接新项目”,今年却反过来决定“想接多少新项目就接多少”,理由正是不去推边界就不知道边界在哪。

他还记录了两个副作用。一是“AI 狂躁”:任何智能体没在为你干活的时间都像被浪费,于是熬夜、堆项目,最终靠亲手做出几个“世界并不需要”的半成品才自我治愈。二是他与 Bryan Cantrill、Adam Leventhal 在播客里造出的“Deep Blue”——软件工程师因为 AI 好像什么都能做而产生的无精打采。这不是技术问题,而是组织与职业身份问题。

安全侧的观察同样具体:本次大会 277 场会议中约 40 场涉及沙箱或智能体安全,说明投入真实存在;但他此前预测的“挑战者号灾难”级事件并未真正发生。

对谁有价值

  • 一线开发者与技术负责人:可用于复盘“哪一类任务现在已经可以默认交给智能体”。
  • 工程与产品管理者:理解“Deep Blue”式倦怠,避免用忙碌感而非结果来评价产出。
  • AI 工具与平台团队:沙箱与智能体安全已是真实且密集的投入方向。
  • 内容与研究者:这是一份由当事人给出的年度时间线,可作为趋势叙事的参照系。

可以怎么行动

  • 用“相变提问”替代“跑分提问”:不要问模型排第几,而问团队里哪件原本做不到的事现在能稳定做到,并把它写进流程。
  • 给自己开一个有限期的“更有野心”窗口:集中尝试原本因成本或人力被否掉的项目,到期后用交付结果而不是工时来结算。
  • 把沙箱与智能体安全列入上线前置项,而不是事后补丁——这是今年行业投入最密集的方向之一。
  • 为“AI 狂躁”设自检机制:定期问一句“这个东西有人需要吗”,用它挡住堆积的副项目。
  • 把“Deep Blue”当作团队沟通词汇,让倦怠可以被命名、被讨论,而不是被当成个人能力问题。

风险与限制

这是一篇演讲配套笔记,是个人视角的年度回顾,不是行业统计或基准报告,覆盖范围取决于讲者自身的关注圈。作者本人在播客中的预测也被他自己评价为“不够有野心”,他预言的编码智能体安全灾难并未按预期发生,说明这类前瞻判断的失准率不低。此外,他自述的两个大型实验项目最后被他判定为“世界不需要”,提示能力提升不等于产出有价值。引用其中的时间线与判断时,应保留其个人经验的性质。

信息差价值

这条内容的真正价值,不只是“有人发布了一个新功能”,而是它揭示了 simonwillison.net 背后的产品方向、工作流变化或竞争信号。对 OPC 来说,这种信息可以转化成持续追踪的栏目选题。

如果把《2026 年的 LLM(迄今为止)》放到你的内容系统里,它最大的价值在于帮助读者更快看懂“为什么值得关注”,而不是只看到一条碎片化动态。

参考来源

AI SUMMARY

这篇文章回答了什么

2026 年的 LLM(迄今为止)主要讲什么?

Simon Willison 在圣何塞 WeAreDevelopers 世界大会北美站闭幕演讲中,把过去一年 LLM 领域串成一条时间线:从 2025 年 11 月 Claude Opus 4.5 与 GPT-5.1 让编码智能体跨过“可靠到可日常使用”的隐形线,到 12 月开发者假期里的集体试用、1 月的“AI 狂躁”,再到他给自己定下“更有野心”的新年…

这篇文章最值得关注的要点是什么?

Simon Willison 在圣何塞 WeAreDevelopers 世界大会北美站闭幕演讲中,把过去一年 LLM 领域串成一条时间线:从 2025 年 11 月 Claude Opus 4.5 与 GPT-5.1 让编码智能体跨过“可…;原贴提到:On Friday I gave the closing keynote at the WeAreDevelopers World Congre;来源:simonwillison.net

这篇文章和哪些AI专题相关?

它适合放在Agent工作流、AI工具、AI日报专题里阅读。 关联原因:这篇内容命中「智能体、Claude Code」等主题信号。;这篇内容命中「工具、Claude」等主题信号。;这篇内容命中「热点解读」等主题信号。

阅读这篇文章建议先理解哪些关键词?

建议先理解AI日报、每日AI日报、AI信号、热点解读、BuilderPulse这些关键词,再结合正文判断工具、机会或风险是否值得进入自己的工作流。

上一篇 【必读】每日AI日报 2026-09-28 下一篇 OpenAI DevDay 点名:你从哪里加入我们,你在构建什么?