AI觉醒星球
Awakening is here
Knowledge File / 全球热点解读
2026-06-14 10 浏览 公开

趋势解读:New AI model called "Count Anything" does exactly,提升开发者接入体验

清华大学等机构研究人员基于Meta SAM3开发的'Count Anything'模型,通过文本提示即可对多种图像中的物体进行计数和标记,在CLOC数据集训练后性能超越多数竞品,但仍难处理模糊术语和极密集场景。

SOURCE / 全球热点解读 MIN / 9 ACCESS / 公开 POST / 2026-06-14 01:00:19

原贴

查看原文
作者:Jonathan Kemper 来源站点:the-decoder.com 原贴时间:

原文

"Count Anything" counts and labels objects across a wide variety of image types, from satellite imagery and medical scans to everyday photos, using nothing more than a text prompt. The system builds on Meta's SAM3 and combines two approaches: it draws boxes around large objects and places points on small, dense targets, then merges the results without double counting. Trained on the custom-built CLOC dataset, the model outperforms many competitors in tests but still struggles with ambiguous terms and extremely dense scenes. Large language models can describe images, interpret charts, and pull text from photos. Multimodality is a given for modern AI systems. But one seemingly simple task remains surprisingly hard: reliably counting objects in an image. Getting those counts right has real consequences, whether it's a doctor reading a scan, a farmer estimating crop yields, or a city planner analyzing traffic. Until now, each of these tasks has required its own specialized system. That's where "Count Anything" comes in. The new AI model from researchers at Tsinghua University and other institutions aims to count objects across very different types of images, whether that's heads in crowds, cars in satellite photos, cells in medical scans, or bacterial colonies in the lab. Ad It's a familiar problem. A system that reliably counts heads in a crowd often chokes on tightly packed cells under a microscope or tiny vehicles seen from above. The researchers want a single model that takes text input, marks every counted object in the image, and handles wildly different image types. Ad DEC_D_Incontent-1 The key idea is combining two approaches that complement each other. One specializes in large, clearly visible objects and draws bounding boxes around them. The other handles small, densely packed objects by placing a dot on each detected target. Both predictions get merged at the end. A simple rule keeps the same object from being counted twice. When both counters flag the same target, only the prediction with higher confidence survives. Ad The system builds on a pretrained model from Meta called SAM3 that can process images and text together. Count Anything adds small adapter components on top for the counting task instead of retraining the whole model from scratch. For the model to learn this broadly, the researchers first had to build a matching dataset. Existing public datasets were typically built for a single purpose, like tumor cells or satellite images. The researchers merged them, cleaned up conflicting labels, and released the result as CLOC , which they say is the largest dataset for text-guided counting to date. Ad DEC_D_Incontent-2 It contains about 220,000 images, 619 categories, and 15 million labeled objects across six domains. Those include everyday photos, satellite and drone imagery, medical tissue samples, microscopic cell images, agricultural images like wheat ears, and bacterial culture photos. Ad

中文翻译

"Count Anything"仅使用文本提示即可对各种图像类型(从卫星图像和医学扫描到日常照片)中的对象进行计数和标记。该系统建立在Meta的SAM3之上,并结合了两种方法:它在大型物体周围绘制方框,并将点放置在小型、密集的目标上,然后合并结果,而不会重复计算。该模型在定制的CLOC数据集上进行训练,在测试中优于许多竞争对手,但仍然难以应对模糊的术语和极其密集的场景。

核心信息

清华大学等机构研究人员基于Meta SAM3开发的'Count Anything'模型,通过文本提示即可对多种图像中的物体进行计数和标记,在CLOC数据集训练后性能超越多数竞品,但仍难处理模糊术语和极密集场景。

  • Count Anything通过文本提示实现跨域物体计数。
  • 结合边界框和点标注,避免重复计数。
  • 在CLOC数据集训练,性能领先但难处理密集场景。
  • 基于Meta SAM3,添加适配器降低训练成本。
  • 有望统一医疗、农业、遥感等计数应用。

详细解读

这是什么信号
这是AI在视觉计数领域迈出的重要一步,解决了多模态模型长期以来难以可靠完成的目标计数任务。'Count Anything'通过结合边界框和点标注两种互补方法,实现了跨域通用计数,标志着AI从'识别'向'精确定量'的进化。

为什么重要
计数在医疗、农业、城市规划等场景中具有实际后果。此前各领域需独立开发专用系统,成本高且难以迁移。该模型首次用一种架构统一处理细胞计数、卫星图像中的车辆检测、农田作物估算等多样任务,大幅降低了开发门槛。

对谁有价值
1. 开发者:直接受益于模型的开源适配器设计,可低成本集成到现有工作流中,无需从头训练。2. 医疗影像分析团队:可快速自动化细胞计数,提升诊断效率。3. 农业科技公司:用于作物生长监测和产量预测。4. 城市规划者:从卫星图像中统计车辆或建筑密度。

可以怎么行动
1. 关注模型在Github上的开源进展,下载CLOC数据集进行领域内微调。2. 在低风险场景中试点替换现有专用计数系统,对比精度和速度。3. 记录模型在模糊术语和密集场景下的失败案例,为后续改进提供反馈。

风险或限制
1. 对模糊术语(如'杯子'可能指不同形态)和极密集场景(如细菌菌落重叠)性能不佳,需人工复核。2. 依赖Meta SAM3框架,可能受上游模型更新影响。3. CLOC数据集虽大,但覆盖领域有限,新场景泛化能力待验证。

信息差价值

信息差价值:多数从业者仍认为物体计数需专用模型,而本模型展示了通用化可能。这一认知差可转化为技术选型优势:开发者可提前储备基于SAM3的适配器能力,在竞争对手之前推出低成本计数解决方案。

业务启发:对于内容创作平台,可将该模型集成至图像分析工具,自动生成图片中物体数量描述(如'图片中有23辆汽车'),提升元数据丰富度。对于电商,可用于商品图片中同类物品计数(如'16个橘子'),辅助库存管理。

可沉淀动作:1. 建立'视觉计数Benchmark'数据库,跟踪各模型在多样化场景下的表现。2. 编写《Count Anything接入指南》,包含微调教程和常见失败案例。3. 探索将计数结果与LLM结合,生成结构化报告(如'作物密度分析报告')。

参考来源

上一篇 亚马逊首席执行官与美国官员会谈引发对 Anthropic 模型的整治 下一篇 趋势解读:Google Research's Gemini-SQL2 tops text-to-SQL benchmarks by a,提升开发者接入体验