觉
AI觉醒星球
Awakening is here
Knowledge File / AI技能杠杆
2026-09-13 2 浏览 免费阅读

两年大学研究发现:禁止课堂使用 AI 使学生表现更差

阿姆斯特丹自由大学法学教授 Thibault Schrepel 为期两年的实验显示,禁止使用 AI 的组连续两年排名最后;无指导 AI 组略高于禁 AI 组;结构化提示工程培训组第一年领先,但第二年优势几乎消失。研究者原本以为无指导使用 AI 弊大于利,最终承认“我错了”。

SOURCE / AI技能杠杆 MIN / 9 ACCESS / 免费阅读 POST / 2026-09-13 17:27:01

原贴

查看原文
作者:Jonathan Kemper 来源站点:the-decoder.com 原贴时间:

原文

A law professor spent two years testing how an AI ban, unguided AI suggestions, and structured training affect student performance. The group without AI finished last both years. "I was wrong," the researcher writes. He had assumed AI without guidance would do more harm than good. Thibault Schrepel of Vrije Universiteit Amsterdam randomly split students in his "Law of AI" course into three groups. The task was the same for everyone: working in small teams of four or five, they had 20 minutes to improve a provision of the EU AI Act. Grading covered substance, clarity, proportionality, and innovation. The first group couldn't use ChatGPT. The second received ChatGPT-generated revision suggestions embedded directly in the text and could keep using the tool but got no guidance on how. The third group got hands-on training in legal prompt engineering and checking AI suggestions for consistency and accuracy. All students took the same exams: a multiple-choice test and a take-home exam that required them to revise another AI Act provision. The experiment ran in 2024 with 66 students and was repeated in 2025 with 164 participants. The no-AI group mostly made minor wording tweaks, rephrasing for clarity and cutting redundancies. Only a few subgroups tried substantive legal improvements. After ten to fifteen minutes, many subgroups regularly ran dry, an effect Schrepel calls "idea exhaustion." Going without AI did force deeper discussion among group members, though. The second group accepted AI suggestions largely without question, often reasoning that they "sounded better," Schrepel writes. Some students replaced the terms "shall" and "individual" with AI-suggested alternatives without showing they understood the legal implications. Every subgroup kept at least one misleading or legally extraneous term from ChatGPT's output. Only the trained third group engaged in genuine back-and-forth with the AI, testing different phrasings and digging into the substantive questions behind the assignment. The trained group scored well above the other two in 2024, especially on the more demanding take-home exam. A year later, that gap had almost closed. All three groups performed at roughly the same level in 2025. Schrepel attributes this to growing chatbot familiarity. Many students already use these tools in daily life, so formal training delivered less of a boost in 2025 than it did a year earlier. Ethical use and legal responsibility still need to be taught, he adds. One finding held constant across both years: the no-AI group finished last. Banning AI produces worse average outcomes than allowing it, Schrepel concludes. Whether the advantage comes from structured training or just hands-on experience remains an open question. The second group's performance surprised him, Schrepel writes. He expected the uncritical errors from the classroom exercise to carry over into the exam. The opposite happened. The errors didn't repeat, and the group actually scored slightly higher than the AI-free group. The most likely explanation, he says, is that students learn to spot AI weaknesses through their own use, especially when accuracy has real consequences. That forced Schrepel to abandon his starting assumption. He had been convinced AI only helps with structured training and should otherwise stay out of the classroom. "I was wrong," he writes. Skipping AI instruction is a missed opportunity, but it doesn't cause the educational collapse some had predicted.

中文翻译

一位法学教授花了两年时间测试禁止使用 AI、无指导的 AI 建议和结构化培训如何影响学生表现。不使用 AI 的小组连续两年排名最后。“我错了,”研究者写道。他曾假设没有指导的 AI 弊大于利。

阿姆斯特丹自由大学的 Thibault Schrepel 将他“AI 法律”课程的学生随机分成三组。所有人的任务相同:以四到五人的小组合作,他们有 20 分钟时间改进《欧盟人工智能法案》的一项条款。评分涵盖实质内容、清晰度、比例性和创新性。第一组不能使用 ChatGPT。第二组收到直接嵌入文本中的 ChatGPT 生成的修改建议,可以继续使用该工具,但没有关于如何使用它的指导。第三组接受了法律提示工程以及检查 AI 建议一致性和准确性的实操培训。

所有学生参加相同的考试:一项多项选择题测试和一项带回家考试,后者要求他们修改《人工智能法案》的另一项条款。该实验于 2024 年进行,有 66 名学生参与,并于 2025 年重复,有 164 名参与者。

不使用 AI 的小组大多只做细微的措辞调整,为了清晰而改写并删减冗余。只有少数小组尝试了实质性法律改进。十到十五分钟后,许多小组经常思路枯竭,Schrepel 称之为“想法耗尽”。不过,不使用 AI 确实迫使小组成员进行更深入的讨论。

第二组基本上不加质疑地接受 AI 建议,常常理由是它们“听起来更好”,Schrepel 写道。一些学生用 AI 建议的替代词替换了“shall”和“individual”等术语,却没有表明他们理解其法律含义。每个小组都至少保留了 ChatGPT 输出中的一个误导性或法律上无关的术语。只有经过训练的第三组与 AI 进行了真正的来回互动,测试不同措辞并深入探究作业背后的实质性问题。

经过训练的组在 2024 年得分远高于另外两组,尤其是在要求更高的带回家考试中。一年后,这一差距几乎消失。2025 年,三个组的表现大致相同。Schrepel 将此归因于对聊天机器人熟悉度的增长。许多学生已经在日常生活中使用这些工具,因此正式培训在 2025 年带来的提升比一年前更少。他补充说,伦理使用和法律责任仍然需要教授。

一个发现两年都保持不变:不使用 AI 的组排名最后。Schrepel 总结说,禁止 AI 产生的平均结果比允许 AI 更差。优势是来自结构化培训还是仅仅来自实际使用,仍然是一个未决问题。

第二组的表现让他惊讶,Schrepel 写道。他原本预计课堂练习中的不加批判错误会延续到考试中。相反的情况发生了。错误没有重复出现,该组实际上得分略高于无 AI 组。他说,最可能的解释是,学生通过自己的使用学会识别 AI 的弱点,尤其是当准确性有真实后果时。这迫使 Schrepel 放弃他最初的假设。他曾确信 AI 只在结构化培训中有帮助,否则应该留在课堂之外。“我错了,”他写道。跳过 AI 教学是一个错失的机会,但它不会造成一些人预测的教育崩溃。

核心信息

阿姆斯特丹自由大学法学教授 Thibault Schrepel 为期两年的实验显示,禁止使用 AI 的组连续两年排名最后;无指导 AI 组略高于禁 AI 组;结构化提示工程培训组第一年领先,但第二年优势几乎消失。研究者原本以为无指导使用 AI 弊大于利,最终承认“我错了”。

  • 阿姆斯特丹自由大学法学教授 Thibault Schrepel 为期两年的实验显示,禁止使用 AI 的组连续两年排名最后;无指导 AI 组略高于禁 AI 组;结构化提示工程培训组第一年领先,但第二年优势几乎消失。研究者原本以为无指导使用 AI 弊大于利,最终承认“我错了”。
  • 原贴提到:A law professor spent two years testing how an AI ban, unguided AI sugge
  • 来源:the-decoder.com

详细解读

这是什么信号:这不是一项普通的“AI 能否提效”调查,而是一位法学教授在自己课程中连续两年做的对照实验。2024 年 66 人、2025 年 164 人,三组随机分配:禁止 AI、直接给 AI 建议但不培训、培训法律提示工程与 AI 核查。结果先打脸了“禁 AI 更安全”的直觉:无 AI 组两年都垫底;无指导 AI 组也没有像预期那样被误导拖垮,反而略高于无 AI 组;培训组第一年显著领先,但第二年优势几乎消失。

为什么重要:教育者和企业培训负责人常把“禁止使用 AI”当作防止作弊和质量下滑的最稳妥策略,但实验指向相反的平均结果:禁止 AI 不仅没有提升表现,还让学生更早陷入“想法耗尽”,只做措辞微调。更关键的是,随着工具普及,结构化提示培训的边际价值在下降——2025 年三组成绩趋同,说明“会用”正在变成基础能力,而不是差异化优势。不过,伦理使用与法律责任仍需正式教学,这一点原文明确保留。

对谁有价值:高校教师、课程设计者、企业 L&D 与知识型团队负责人,以及正在制定 AI 使用政策的组织。对个人学习者而言,信号是:不要等正式培训,先在有真实后果的任务中练习与 AI 互相校验。对管理者而言,与其一刀切禁用,不如设计任务、评分标准和核查机制,让 AI 使用暴露在可评估的流程里。

可以怎么行动:第一,把“禁止/允许”的二元政策改成分层任务:基础技能可无 AI 练,但复杂改进类任务应允许 AI,并要求提交修改理由和核查记录。第二,培训重点从“提示词模板”转向“识别 AI 弱点”:一致性、准确性、法律含义、误导性术语。第三,评分里加入对 AI 建议的批判性说明,避免学生无意识接受“听起来更好”的措辞。第四,定期复测,因为工具熟悉度变化会迅速侵蚀培训优势。

风险或限制:这是单一课程、单一教师、法律文本修改场景,样本为 66 人和 164 人,不能直接外推到所有学科或企业场景。三组表现趋同也可能来自学生日常使用增加,而非培训无效。原文没有给出具体分数、效应量或考试细节,因此不能把结论夸大为“AI 必然提升学习”。另外,无 AI 组的“想法耗尽”是课堂观察,不等于考试结果的全部解释。真正稳妥的做法是保留实验心态:允许 AI、教核查、用真实后果检验。

信息差价值

这条内容的真正价值,不只是“有人发布了一个新功能”,而是它揭示了 the-decoder.com 背后的产品方向、工作流变化或竞争信号。对 OPC 来说,这种信息可以转化成持续追踪的栏目选题。

如果把《两年大学研究发现:禁止课堂使用 AI 使学生表现更差》放到你的内容系统里,它最大的价值在于帮助读者更快看懂“为什么值得关注”,而不是只看到一条碎片化动态。

参考来源

AI SUMMARY

这篇文章回答了什么

两年大学研究发现:禁止课堂使用 AI 使学生表现更差主要讲什么?

阿姆斯特丹自由大学法学教授 Thibault Schrepel 为期两年的实验显示,禁止使用 AI 的组连续两年排名最后;无指导 AI 组略高于禁 AI 组;结构化提示工程培训组第一年领先,但第二年优势几乎消失。研究者原本以为无指导使用 AI 弊大于利,最终承认“我错了”。

这篇文章最值得关注的要点是什么?

阿姆斯特丹自由大学法学教授 Thibault Schrepel 为期两年的实验显示,禁止使用 AI 的组连续两年排名最后;无指导 AI 组略高于禁 AI 组;结构化提示工程培训组第一年领先,但第二年优势几乎消失。研究者原本以为无指导使用…;原贴提到:A law professor spent two years testing how an AI ban, unguided AI sugge;来源:the-decoder.com

这篇文章和哪些AI专题相关?

它适合放在Agent工作流、AI超级个体、AI工具专题里阅读。 关联原因:这篇内容命中「Agent、工作流」等主题信号。;这篇内容命中「技能、学习」等主题信号。;这篇内容来自该专题长期覆盖的栏目。

阅读这篇文章建议先理解哪些关键词?

建议先理解AI工具、工具、自动化、模型、Cursor这些关键词,再结合正文判断工具、机会或风险是否值得进入自己的工作流。

上一篇 AI导演DiDi_OK专访:用Seedance 2.5打造新作《糖果》,聊聊模型如何吞掉工作流 下一篇 Agent 长任务上下文工程解析:用预算控制、压缩、todo-state 和记忆对抗上下文溢出与目标丢失