觉
AI觉醒星球
Awakening is here
Knowledge File / AI小生意项目库
2026-07-16 8 浏览 免费阅读

解密扩散模型的创造力

研究表明,扩散模型能够生成新颖数据而非简单记忆训练集,其创造力源于神经网络学习分数函数时的自然平滑效应,导致模型在隐藏数据流形上的训练点之间进行插值。

SOURCE / AI小生意项目库 MIN / 9 ACCESS / 免费阅读 POST / 2026-07-16 02:06:27

原贴

查看原文
作者:Google Research Blog 来源站点:research.google 原贴时间:

原文

Zhengdao Chen, Research Scientist, Google Research We show that a diffusion model’s creativity (its ability to generate novel data, rather than just memorize its training set) is a mathematical consequence of neural networks learning a "smoothed" version of the score function, driving the model to interpolate between training data points along the hidden data manifold. Diffusion models are currently one of the most powerful types of tools for generative tasks that require complex and local structures, such as image generation and molecular discovery. They’ve shown an exciting capability to generalize beyond their training data and, in this sense, exhibit “creativity”. For instance, after being trained with datasets of actual images, they can transform random noise samples into novel, high-quality images. While this creative capability is impressive, it raises an intriguing question: where does it come from? Understanding the answer to this question is an important step towards demystifying the “black-box” nature of diffusion-based generative AI. To that end, in " On the Interpolation Effect of Score Smoothing in Diffusion Models ", presented at ICLR 2026 , we dive into the mathematics of diffusion models to answer this question. We show that a model’s creativity isn’t a random fluke. Instead, it is a consequence of how neural network training naturally "smooths" the transformation from noise back to the data during the generation process. Training a diffusion model begins with taking a set of real training data samples — like cat photos — and intentionally corrupting them with noise until they become completely unrecognizable. The model is then trained to reverse this corruption step-by-step so that it can reconstruct a realistic-looking image from pure noise, a process called denoising . If the model learns to perform this denoising process perfectly based only on its training samples, it should produce carbon copies of them during deployment time as well (a behavior known as memorization ). In this scenario, the model acts as a retrieval tool rather than as a creative engine capable of generating novel outputs. In practice, however, diffusion models usually do more than just memorize; they generalize to generate new data samples. To understand how diffusion models actually denoise data, imagine random noise as a cloud of gas particles scattered across a room, where a “force field” pulls each particle in a specific direction until they form a meaningful shape. In a diffusion model, the moving particles are individual data points undergoing denoising. The “force field” is the score function (SF), which is learned from the training data and dictates where the particles should flow at any given moment. If the model relies on a score function learned perfectly from the training data, the force field will drive the particles into positions that exactly replicate the training data points (i.e., memorization). The score function drives the denoising process which turns pure noise into meaningful data (e.g., images). We discovered that the creativity of diffusion models actually originates from the approximate nature of how neural networks typically learn: imperfect training due to regularization naturally leads to a slight blurring of the learned score function in a process called “score smoothing”. This, in turn, causes the denoising process to generate data that interpolates (in other words, fall in the space between) the training points, thus creating new and plausible data samples.

中文翻译

Zhengdao Chen, Research Scientist, Google Research 我们证明了扩散模型的创造力(即生成新颖数据而非仅记忆训练集的能力)是神经网络学习分数函数“平滑”版本的数学结果,这驱使模型沿着隐藏数据流形在训练数据点之间进行插值。当前,扩散模型是生成具有复杂局部结构的数据(如图像生成和分子发现)的最强大工具之一。它们展现出超越训练数据泛化的令人兴奋的能力,并在此意义上表现出“创造力”。例如,在通过真实图像数据集训练后,它们能将随机噪声样本转化为新颖的高质量图像。这种创造能力令人印象深刻,但引出一个有趣的问题:它从何而来?理解这个问题的答案是揭开基于扩散的生成式AI“黑箱”本质的重要一步。为此,在发表于ICLR 2026的论文《On the Interpolation Effect of Score Smoothing in Diffusion Models》中,我们深入扩散模型的数学以回答此问题。我们表明,模型的创造力并非随机偶然,而是神经网络训练在生成过程中自然“平滑”从噪声到数据转换的结果。训练扩散模型首先从一组真实训练数据样本(如猫照片)开始,故意用噪声破坏直到完全无法识别。然后训练模型逐步逆转此破坏过程,使其能从纯噪声重建逼真图像,这称为去噪。如果模型完全基于训练样本完美学习去噪过程,那么在部署时它应生成这些样本的精确副本(称为记忆化)。在此场景中,模型充当检索工具而非能生成新颖输出的创造性引擎。然而在实践中,扩散模型通常不仅记忆,还能泛化生成新数据样本。为了理解扩散模型实际如何去噪,想象随机噪声如同散布在房间中的气体粒子云,而“力场”将每个粒子拉向特定方向直到它们形成有意义形状。在扩散模型中,移动的粒子是经历去噪的单个数据点。“力场”是分数函数,它从训练数据学习并决定粒子在任何时刻的流向。如果模型依赖从训练数据完美学习的分数函数,力场会将粒子驱至精确复制训练数据点的位置(即记忆化)。分数函数驱动去噪过程,将纯噪声转化为有意义数据(如图像)。我们发现扩散模型的创造力实际上源于神经网络学习的近似本质:正则化导致的不完美训练自然使学习到的分数函数在称为“分数平滑”的过程中产生轻微模糊,进而导致去噪过程生成插值(即落在训练点之间空间)的数据,从而创造新颖且合理的数据样本。

核心信息

研究表明,扩散模型能够生成新颖数据而非简单记忆训练集,其创造力源于神经网络学习分数函数时的自然平滑效应,导致模型在隐藏数据流形上的训练点之间进行插值。

  • 研究表明,扩散模型能够生成新颖数据而非简单记忆训练集,其创造力源于神经网络学习分数函数时的自然平滑效应,导致模型在隐藏数据流形上的训练点之间进行插值。
  • 原贴提到:Zhengdao Chen, Research Scientist, Google Research We show that a diffus
  • 来源:research.google

详细解读

这是什么信号? Google Research 揭示扩散模型创造力的数学根源:并非偶然,而是神经网络训练中分数函数平滑的自然结果。该发现将生成模型的泛化能力从经验现象提升为可解释的理论,标志着对生成式AI“黑箱”本质的突破性理解。

为什么重要? 长期以来,扩散模型的创造性被视为经验奇迹,缺乏理论支撑。此研究首次将创造力归因于分数平滑导致的插值效应,解释了模型为何能生成超越训练集的多样样本。这对提升模型可靠性、可控性及设计更高效的生成算法具有基础性意义。

对谁有价值? 1) AI研究者:获得理解生成模型泛化机制的新框架,有助于改进模型结构、训练策略及避免过拟合。2) 应用开发者:可基于插值原理优化图像、分子等生成任务,例如通过调整平滑强度控制创新程度。3) 产品经理:理解创造力来源有助于评估模型潜力,预判生成内容的多样性与真实性边界。

可以怎么行动? 1) 在训练中引入正则化强度调节,平衡记忆与泛化以适配不同应用(如需要精准复现 vs 新颖创意)。2) 利用插值特性设计混合数据生成策略,如通过平滑路径生成对训练集的高质量变体。3) 开发者可测试分数平滑对特定领域(如医学影像)的影响,避免因过度平滑导致细节失真。

风险或限制 1) 研究基于理论推导及 ICLR 论文验证,但实际应用中的平滑效果可能受网络结构、优化算法等因素影响,需要更多实证。2) 过度平滑可能导致生成样本模糊或偏离真实分布,尤其在高细节任务(如人脸生成)中需谨慎控制。3) 该理论主要解释插值行为,未涵盖扩散模型的其他创造性来源(如随机采样路径的多样性),需综合理解。

信息差价值

这条内容的真正价值,不只是“有人发布了一个新功能”,而是它揭示了 research.google 背后的产品方向、工作流变化或竞争信号。对 OPC 来说,这种信息可以转化成持续追踪的栏目选题。

如果把《解密扩散模型的创造力》放到你的内容系统里,它最大的价值在于帮助读者更快看懂“为什么值得关注”,而不是只看到一条碎片化动态。

参考来源

AI SUMMARY

这篇文章回答了什么

解密扩散模型的创造力主要讲什么?

研究表明,扩散模型能够生成新颖数据而非简单记忆训练集,其创造力源于神经网络学习分数函数时的自然平滑效应,导致模型在隐藏数据流形上的训练点之间进行插值。

这篇文章最值得关注的要点是什么?

研究表明,扩散模型能够生成新颖数据而非简单记忆训练集,其创造力源于神经网络学习分数函数时的自然平滑效应,导致模型在隐藏数据流形上的训练点之间进行插值。;原贴提到:Zhengdao Chen, Research Scientist, Google Research We show that a diffus;来源:research.google

这篇文章和哪些AI专题相关?

它适合放在AI副业、AI工具、AI超级个体专题里阅读。 关联原因:这篇内容命中「项目、小生意、变现」等主题信号。;这篇内容命中「模型」等主题信号。;这篇内容命中「学习」等主题信号。

阅读这篇文章建议先理解哪些关键词?

建议先理解AI工具、工具、自动化、模型、Cursor这些关键词,再结合正文判断工具、机会或风险是否值得进入自己的工作流。

上一篇 OpenAI正在使用AI攻击自己的AI,效果比人类更好 下一篇 Thinking Machines 发布多模态模型 Inkling