AI觉醒星球
Awakening is here
Knowledge File / AI小生意项目库
2026-07-22 0 浏览 会员

AI系统帮助巴基斯坦法官清理大量积压案件,每投资1美元回报38.5美元

AI能让政府机构更高效吗?一项来自苏黎世联邦理工学院、帝国理工学院和新经济学院的研究提供了迄今最有力的实验证据。

SOURCE / AI小生意项目库 MIN / 9 ACCESS / 会员 POST / 2026-07-22 03:12:20

原贴

查看原文
作者:Matthias Bastian 来源站点:the-decoder.com 原贴时间:

原文

Can AI make government institutions more productive? A study from researchers at ETH Zurich, Imperial College London, and the New Economic School delivers the strongest experimental evidence yet. Pakistan has fewer than two judges per 100,000 residents, according to the authors. The EU has 22. England and Wales have 30. At the end of 2024, 2.26 million cases were pending, 82 percent of them in trial courts. Judges work with bare-bones tech and no support staff. Before the experiment, only about 25 percent had ever used a large language model like ChatGPT. Researchers from ETH Zurich, the New Economic School, and Imperial College London ran a large-scale field experiment with Pakistan's judiciary. The randomized trial covered 1,559 judges across 118 courts, roughly half of all Pakistani trial court judges. The tool was JudgeGPT, an AI assistant built on OpenAI's GPT-4 and designed for Pakistani trial courts. It uses retrieval augmented generation to search a database of 129,235 documents, including 128,292 court rulings and 943 Pakistani laws. When a judge enters a query, JudgeGPT picks the ten most relevant passages and generates a cited answer. The researchers split judges into three groups. One got JudgeGPT access plus targeted training: six 90-minute lectures over three weeks, taught by ETH Professor Elliott Ash after court hours. Judges learned which tasks suited the tool, where it fell short, and how to check its output. A second group got the same AI access but only a general seminar on technology and law. The control group attended that seminar with no JudgeGPT access. AI access alone did little. Judges with targeted training used JudgeGPT four times as much as those in the general seminar group. After 40 weeks, trained judges averaged nearly 60 logins and over 200 prompts. The comparison group averaged about 20 logins and fewer than 50 prompts. Districts with more trained judges resolved more cases. At moderate exposure levels, that meant about 1,848 extra cases per year per district, a 6.3 percent bump. Even districts in the bottom quartile still cleared about 616 more cases. Ruling quality held steady or improved. The appeal rate per 1,000 resolved cases fell slightly, and judges worked the same hours with no change in work-life balance. The researchers estimate savings of about $38.50 per dollar invested, based on what it would cost to hire enough extra judges to match the same output. Even conservative estimates put the return at "at least" $10 per dollar. A review of roughly 4,000 court judgments found more AI-flagged text, as expected. But readability, length, and the number of legal arguments held steady. An LLM-based quality check, validated by two Pakistani lawyers, showed a slight improvement. Rulings from trained judges were rated better in 59 percent of pairwise comparisons, up from 42 percent in the control group. The study found no evidence that AI use increased gender or religious bias in judicial language. The researchers reviewed anonymized chat logs from about 1,500 judges. Legal research, text editing, and text generation were the most common tasks. Around 60 percent of queries sought information about laws, procedures, or legal concepts. Trained judges used JudgeGPT more for editing and summarizing text, tasks where language models are more reliable. They asked fewer broad legal questions, where hallucination risk is higher. The researchers say the training steered judges toward limited support tasks while leaving the final decisions in their hands.

中文翻译

AI能让政府机构更高效吗?来自苏黎世联邦理工学院、帝国理工学院和新经济学院的研究人员提供了一项迄今最有力的实验证据。据作者称,巴基斯坦每10万居民中只有不到两名法官。欧盟有22名。英格兰和威尔士有30名。截至2024年底,有226万件案件待处理,其中82%在初审法院。法官使用的技术非常简陋,没有辅助人员。实验前,只有约25%的人曾使用过像ChatGPT这样的大型语言模型。来自苏黎世联邦理工学院、新经济学院和帝国理工学院的研究人员与巴基斯坦司法部门进行了一项大规模实地实验。这项随机试验覆盖了118个法院的1559名法官,约占巴基斯坦所有初审法院法官的一半。该工具是JudgeGPT,一个基于OpenAI的GPT-4构建的AI助手,专为巴基斯坦初审法院设计。它使用检索增强生成技术搜索包含129,235份文档的数据库,其中包括128,292份法院裁决和943部巴基斯坦法律。当法官输入查询时,JudgeGPT会选择十个最相关的段落并生成带引用的答案。研究人员将法官分为三组。一组获得了JudgeGPT访问权限以及有针对性的培训:三周内六次90分钟的讲座,由ETH教授Elliott Ash在法院下班后授课。法官们学习了哪些任务适合该工具,哪些地方有不足,以及如何检查其输出。第二组获得了相同的AI访问权限,但只参加了一个关于技术和法律的一般研讨会。对照组参加了该研讨会,但没有JudgeGPT访问权限。仅提供AI访问权限效果甚微。接受过针对性培训的法官使用JudgeGPT的次数是一般研讨组法官的四倍。40周后,接受培训的法官平均登录近60次,提示次数超过200次。比较组平均登录约20次,提示次数少于50次。拥有更多受过培训法官的地区解决了更多案件。在中等暴露水平下,这意味着每个地区每年约多解决1,848起案件,增加了6.3%。即使是排在最低四分位的地区,也仍多解决了约616起案件。裁决质量保持稳定或有所提高。每千件已决案件的上诉率略有下降,法官工作时间相同,工作与生活平衡没有变化。研究人员估计,每投资1美元可节省约38.5美元,这是基于雇佣足够多的额外法官以达到相同产出所需的成本。即使保守估计,回报也“至少”为每美元10美元。对大约4000份法院判决的审查发现,AI标记的文本有所增加,这在意料之中。但可读性、长度和法律论点的数量保持稳定。由两位巴基斯坦律师验证的基于LLM的质量检查显示略有改善。受过培训法官的裁决在59%的成对比较中得到更好评价,而对照组为42%。研究没有发现AI使用增加司法语言中性别或宗教偏见的证据。研究人员审查了约1500名法官的匿名聊天记录。法律研究、文本编辑和文本生成是最常见的任务。约60%的查询寻求有关法律、程序或法律概念的信息。受过培训的法官更频繁地使用JudgeGPT进行编辑和总结文本,这些任务中语言模型更可靠。他们提出的宽泛法律问题较少,而这类问题幻觉风险较高。研究人员表示,培训引导法官将AI用于有限的辅助任务,同时将最终决定权留给他们自己。

核心信息

AI能让政府机构更高效吗?一项来自苏黎世联邦理工学院、帝国理工学院和新经济学院的研究提供了迄今最有力的实验证据。

  • AI能让政府机构更高效吗?一项来自苏黎世联邦理工学院、帝国理工学院和新经济学院的研究提供了迄今最有力的实验证据。
  • 原贴提到:Can AI make government institutions more productive? A study from resear
  • 来源:the-decoder.com
试看内容

成为会员查看完整内容

你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。

详细解读 信息差价值 参考来源
成为会员查看完整内容
上一篇 物理AI仿真的现状:概述 下一篇 Claude Cowork通过屏幕录制和语音解说学习新技能