原文
Friday's big release was Qwen 3.8 27B , an Apache 2 licensed 27B parameter vision-capable LLM from Alibaba's Qwen research lab. I've been looking forward to this one: 27B is an excellent size for running a model on a reasonably specced laptop, and its predecessor Qwen 3.6 27B was impressive. Qwen's self-reported benchmarks for this model are eye-opening. They show a boost from both Qwen 3.6 27B and the closed-weight Qwen 3.7-Plus, which was one of Qwen's strongest models of any size as recently as May this year . It will be interesting to hear what independent benchmarks have to say about the model. I've been running the model on two different machines: my 128GB M5 Max MacBook Pro, and an NVIDIA DGX Spark . On both machines I'm running LM Studio and their 17GB Q4_K_M quantized build . I also tried using llama-server directly on the Spark. The default of extra high results in spectacular over-thinking Qwen's documentation describes the model as defaulting to xhigh for the reasoning effort, and the LM Studio GGUF I've been trying preserves that default: Qwen3.8 comes with official support for reasoning_effort , which can be used to adjust reasoning depth and control cost: xhigh (default): for complex tasks demanding thorough analysis medium : balancing accuracy and speed low : efficient reasoning optimizing for speed and cost This is a hilarious default. It's absolutely not a good way to run the model, especially on consumer hardware. I've been finding the results extremely entertaining. I quickly ran into problems with LM Studio's default context limit of 8,192 tokens - Qwen was using them all up thinking about even the most mundane of problems. I loaded the model with the full 262,144 maximum context length and that problem went away. Here's the pelican riding a bicycle SVG I got from my first attempt with that increased context length. It took 21 minutes to generate, using 22,276 reasoning tokens to produce 3,223 tokens of output. You can read the reasoning trace here . This is by far the best pelican SVG I've been able to generate with a model that runs on a local machine - and this Qwen is pretty small, just a 17GB file on disk. There's a lot to like about this: The bicycle frame is the right shape It has legs on each side of the bike - that's very rare Good, clear pelican pouch The wings extend to touch the handlebars! The motion lines are behind, not in front It has a tasteful background - nice sun, clouds, hill, flowers and grass. Was that worth waiting 21 minutes for? Absolutely not. Here's that same prompt run with reasoning turned off - transcript here . This one produced 3,715 tokens and took 137s - just over two minutes. And just for completeness, I used OpenRouter to run the same prompt through the much larger Qwen 3.8 2.4T-A95B (released last week ) and got this snazzy animated SVG : Your browser does not support HTML5 video. I said Qwen at xhigh has a tendency to over-think things, but how bad really is it? I tried a much simpler prompt, again with that default extra high setting: draw an svg of a circle Qwen's reasoning trace started like this: The user is asking for an SVG drawing of a circle. Simple request — but I want it to be a carefully crafted piece. Let me make something that goes beyond just <circle> : a single self-contained SVG file with character — maybe a geometric "circle study," with subtle animation, layered rings, and a distinctive palette. Keep the scope right: they asked for an SVG of a circle. So the core is a circle. But I can add craft: concentric guide circles (like a compass/geometry drawing), tick marks, a soft gradient fill on the main circle, restrained ambient motion (a slowly rotating dashed ring, pulsing glow). Respect prefers-reduced-motion? That's more for CSS; SVG SMIL or CSS inside SVG will do. Palette options: deep teal ink on warm paper? Or bold vermilion circle on off-white with navy construction lines — Bauhaus/compass-drawing vibe. Let me go with a "geometric study" look: cool slate background, or bright paper white? Paper white is fine if it's not the cream-and-terracotta combo. [...] Several minutes later it produced this absolutely beautiful animated circle, which was entirely not what I had asked for! Your browser does not support HTML5 video. My strong recommendation: ignore that default. Run Qwen 3.8 27B on low or even no reasoning levels at first. It's a great model, but wow that default setting is a bad place to start. It's very good at bounding boxes A fun way to test a vision model is to see how well it can return bounding boxes around items in a photograph. I've seen previous Qwen models deal well with this, so I decided to put it to the test drawing bounding boxes around some pelicans. I've seen asking for 0-1000 scale produce good results in the past. I tried this: llm -a https://static.inaturalist.org/photos/714731804/large.jpg \ -m lmstudio/qwen/qwen3.8-27b \ ' Return JSON bounding boxes for the pelicans in this photo, 0-1000 scale for each dimension ' Here's the reasoning trace , which produced this: [ { "bbox_2d" : [ 195 , 290 , 370 , 780 ], "label" : " pelicans " }, { "bbox_2d" : [ 445 , 320 , 675 , 850 ], "label" : " pelicans " } ] This is such a good match . Here are those boxes rendered on top of the photo: Building a tool to label bounding boxes That visualization of the bounding boxes was taken using a new custom tool that I had Qwen 3.8 27B build for me, running offline on my laptop. I forgot to dial down the thinking effort so it was massively over-engineered , but it did manage to produce this full interface from this single prompt : [ {"bbox_2d": [195, 290, 370, 780], "label": "pelicans"}, {"bbox_2d": [445, 320, 675, 850], "label": "pelicans"} ] Build an HTML page which has an input box for accepting the URL to an image and a textarea for accepting the above style of JSON. It appends the image to the page, measures its width and height, then treats the coords in the bbox_2d as scaled from 0-1000 and scales them against the actual width and height, then it renders labelled boxes over the image. This screenshot shows one of the features I did not ask for - a demo scene, for if you don't have a photograph to test the tool with: Here's the relevant segment of the thinking trace, where it decided to draw its own pelicans purely because I had used the label "pelicans" in the example JSON I gave it in the prompt: Also a "load sample" that uses a known image? Can't depend on external images, but… the image URL input is user-provided; I could add a "try with sample" button [...] Hmm, I can draw a simple scene on canvas, export it as a data URL, and load it into the image — that's self-contained and demo-able! [...] But the user's coords are for an actual pelican image; a generated placeholder can still demo the scaling. Generate a 1000x1000 placeholder: gradient water + two blob-like "pelican" silhouettes placed at the given bboxes (using the same scale — cute: silhouettes at the exact 0-1000 positions, showing the boxes align). This makes for a fun, self-contained demo. Keep it simple: sky gradient, sun, water, two pelican-ish shapes (ellipse body, circle head, beak). Place at bbox centers. (I'm slightly nervous that models around the world might have a bias towards drawing pelicans at any chance they can get, brought on by nearly two years of exposure to my own stupid benchmark.) Is all that over-thinking necessary? Maybe it is, at least a bit. I tried with reasoning turned off and got this version , ( transcript here ), which nearly works but shows the boxes in the wrong place: So without reasoning it didn't quite one-shot a working tool. I'm sure it could get there with some follow-up prompts, but this is a good example of how reasoning can make a difference. Yes, it can drive coding agents One of the biggest questions around local models is whether or not they have enough horsepower to successfully run a coding agent loop. Coding agents require long context, strong code generation support and reliable tool-calling. On paper Qwen 3.8 27B has all three of these, so is it up to the task? My initial experiments with Pi have been very promising. I chose Pi because it has a shorter system prompt than most other options, making it a better fit for trying out smaller models. I configured Pi to use Qwen 3.8 27B running in LM Studio on the Spark (shared via tailscale serve ) by adding this to ~/.pi/agent/models.json : { "providers" : { "spark" : { "baseUrl" : " https://spark-18b3.tail68a31.ts.net/v1 " , "api" : " openai-responses " , "apiKey" : " dummy " , "models" : [ { "id" : " qwen3.8-27b " , "reasoning" : true } ] } } } Then ran pi --provider spark --model qwen3.8-27b in my ~/dev/datasette folder and prompted: how does auth work? After a sequence of reasoning and tool calls that accessed a bunch of different files it produced this reply , which is very solid. Just one problem: I wanted to share that transcript. So I pointed Pi and Qwen 3.8 27B at the JSONL transcript file in ~/.pi/agent/sessions/--Users-simon-Dropbox-dev-datasette-- and prompted: Write Python code to convert this jsonl to markdown And it built and tested this pi_jsonl_to_md.py , which did exactly what I needed. Here's that session transcript , published using the tool that it created. The quest for speed So far this is all looking very promising. We have a 17GB model that runs on high-end consumer hardware and can write code, drive tools, annotate images and generally do everything that I need from an LLM for getting real work done. There's one very significant catch: it feels slow - especially when it starts over-thinking, but even without that it's not particularly sprightly. I've been getting around 15-30 tokens a second from LM Studio. That's not terrible, but it's slow enough that it's going to be hard to win me away from hosted API models, which can return results a whole lot faster. Artificial Analysis track token speed and show OpenAI 5.6 Sol at 74 tokens/second and 5.6 Luna at an impressive 184/second. The good news is that the community have been exploring ways to speed things up since the model was first released two days ago. One of the most promising optimizations is baked into the model itself. Qwen supports Multi-Token Prediction , an architecture trick where a cheaper mechanism guesses several tokens ahead and the main model can then quickly verify if the guesses were correct. This can have quite a dramatic effect on inference performance. Based on this tweet from llama.cpp creator Georgi Gerganov I tried running the model with MTP like this on the Spark: llama serve \ -hf ggml-org/Qwen3.8-27B-GGUF:Q4_K_M \ -hfd ggml-org/Qwen3.8-27B-GGUF:Q4_0 \ --spec-default \ --spec-type draft-mtp \ --reasoning-preserve And sure enough, this gave me a significant boost. I had GPT-5.6 in Codex run a comparative benchmark on the Spark and the --spec-type draft-mtp server outperformed the LM Studio default GGUF by around 72%. I expect we'll see a whole lot more innovation around serving this model faster over the next few weeks. The MLX community likely have some tricks brewing as well. Some observations The fact that a 17GB file can do all of this stuff on my home machines is a miracle . Once again, I'm delighted and amazed at how much progress local models have made this year. A year ago this would have been competitive with the best and most expensive of the proprietary models - today it can run on a capable laptop. The only thing holding this back from being a daily driver is performance. It feels pretty slow on both the M5 Mac and the DGX Spark. That's the catch with these dense (non-Mixture-of-Experts) models - they require a whole lot of memory bandwidth to perform well, and neither of the machines I have access to are top performers in that regard. The most important thing about Qwen 3.8 27B is what it demonstrates . We can have an open weights general purpose model with a long context, effective tool calling, strong vision ability, and competent code generation, and we can fit the whole thing in just a 17GB file. The models at this size continue to get better at an impressive rate. We don't need to spend half a million dollars on datacenter-class hardware just to run a competent model. Tags: ai , generative-ai , local-llms , llms , qwen , pelican-riding-a-bicycle , llm-reasoning , llama-cpp , llm-release , coding-agents , lm-studio , ai-in-china , nvidia-spark , pi
中文翻译
周五的大发布是 Qwen 3.8 27B,这是阿里巴巴 Qwen 研究实验室推出的一款 Apache 2 许可的 270 亿参数视觉能力 LLM。我一直期待这个:27B 是在配置合理的笔记本电脑上运行模型的绝佳规模,其前身 Qwen 3.6 27B 令人印象深刻。Qwen 自我报告的基准测试令人大开眼界。它们显示出相比 Qwen 3.6 27B 和闭源的 Qwen 3.7-Plus 都有提升,后者在截至今年 5 月仍是 Qwen 最强模型之一。有趣的是,独立基准测试会怎么说。我已在两台不同的机器上运行该模型:我的 128GB M5 Max MacBook Pro 和 NVIDIA DGX Spark。在这两台机器上,我运行 LM Studio 及其 17GB Q4_K_M 量化构建。我还尝试了直接在 Spark 上使用 llama-server。
默认的高强度导致惊人的过度思考。Qwen 文档将模型默认的推理力度描述为 xhigh,而我尝试的 LM Studio GGUF 保留了这一默认值:Qwen3.8 官方支持 reasoning_effort,可用于调整推理深度和控制成本:xhigh(默认):适用于需要深入分析的复杂任务;medium:平衡准确性和速度;low:高效的推理,优化速度和成本。这是一个滑稽的默认值。这绝对不是运行模型的好方式,尤其是在消费级硬件上。我发现结果极其有趣。我很快就遇到了 LM Studio 默认上下文限制 8,192 个 token 的问题——Qwen 在思考最平凡的问题时也用光了它们。我加载了模型的最大上下文长度 262,144,问题就消失了。这是我第一次以增加上下文长度尝试得到的鹈鹕骑自行车 SVG。它花了 21 分钟生成,使用 22,276 个推理 token 产生 3,223 个输出 token。你可以在此处阅读推理痕迹。这是迄今为止我能够用本地机器上的模型生成的最好的鹈鹕 SVG——而这个 Qwen 相当小,磁盘上只有 17GB。有很多值得喜欢的地方:自行车架形状正确;自行车两侧都有腿——这非常罕见;鹈鹕的喉囊又大又清晰;翅膀伸展触碰到车把!运动线在后方,不在前方;它有一个雅致的背景——漂亮的太阳、云朵、小山、花朵和草地。等 21 分钟值得吗?绝对不值得。这是同样的提示在关闭推理时运行的结果——此处是记录。这个产生了 3,715 个 token,耗时 137 秒——刚过两分钟。为了完整起见,我使用 OpenRouter 通过更大得多的 Qwen 3.8 2.4T-A95B(上周发布)运行同样的提示,得到了这个时髦的动画 SVG:您的浏览器不支持 HTML5 视频。
我说 Qwen 在 xhigh 下倾向于过度思考,但到底有多糟呢?我尝试了一个更简单的提示,再次使用默认的高设置:画一个圆的 SVG。Qwen 的推理痕迹是这样开始的:用户要求画一个圆的 SVG。简单的请求——但我希望它是一件精心制作的作品。让我做一些超越 的东西:一个具有个性的自包含 SVG 文件——也许是一个几何"圆的研究",带有微妙的动画、分层环和独特的调色板。保持范围正确:他们要求一个圆的 SVG。所以核心是一个圆。但我可以添加工艺:同心引导圆(如同心圆/几何绘图)、刻度标记、主圆上的柔和渐变填充、克制的环境运动(缓慢旋转的虚线环、脉冲辉光)。尊重 prefers-reduced-motion?这更适合 CSS;SVG SMIL 或 SVG 内的 CSS 可以做到。调色板选项:暖纸上的深青色墨水?还是灰白色上的粗朱红圆配上海军蓝构造线——包豪斯/罗盘绘图的氛围。让我选择"几何研究"的外观:凉爽的板岩背景,还是明亮的纸白?如果这不是奶油色和赤陶色的组合,纸白就可以了。[...]几分钟后,它产生了这个绝对漂亮的动画圆,但这完全不是我要求的!您的浏览器不支持 HTML5 视频。我的强烈建议:忽略那个默认值。首先在 low 甚至 no reasoning 级别运行 Qwen 3.8 27B。它是一个伟大的模型,但哇,那个默认设置是一个糟糕的起点。
它非常擅长边界框。测试视觉模型的一个有趣方式是看它如何返回照片中项目的边界框。我之前见过 Qwen 模型处理得很好,所以我决定测试它在一些鹈鹕周围绘制边界框。我过去看到要求 0-1000 比例会产生好的结果。我尝试了这个:llm -a https://static.inaturalist.org/photos/714731804/large.jpg -m lmstudio/qwen/qwen3.8-27b '返回这张照片中鹈鹕的 JSON 边界框,每个维度 0-1000 比例' 这是推理痕迹,它产生了:[ { "bbox_2d" : [ 195 , 290 , 370 , 780 ], "label" : " pelicans " }, { "bbox_2d" : [ 445 , 320 , 675 , 850 ], "label" : " pelicans " } ] 这是一个非常好的匹配。下面是这些框渲染在照片上的效果:
构建一个标记边界框的工具。该边界框的可视化是使用一个新自定义工具拍摄的,该工具是我让 Qwen 3.8 27B 为我构建的,离线运行在我的笔记本电脑上。我忘了降低思考力度,所以它被大规模过度设计,但它确实成功地从单个提示中产生了这个完整的界面:[ {"bbox_2d": [195, 290, 370, 780], "label": "pelicans"}, {"bbox_2d": [445, 320, 675, 850], "label": "pelicans"} ] 构建一个 HTML 页面,其中有一个用于接受图像 URL 的输入框和一个用于接受上述 JSON 样式的文本框。它将图像附加到页面,测量其宽度和高度,然后将 bbox_2d 中的坐标视为从 0-1000 缩放,并...