测试 (tests)
集里怎么说它
- 《从拒用 AI 到全面拥抱:Dioxus 团队的智能体编程实战课》(09:43起):本集说 Kelly 还没百分之百信服用 AI 写测试:智能体能轻松给任何 API 写测试,但和人类一样写不出「对的测试」,团队仍手动列举测试条件、自己设计测试 API。
① 提到它的金句
6 条
他们直接进入评估,比如”让我只写一些测试”,这就是事情脱轨的地方。
指向原始笔记的链接
They go straight into evals like, “Let me just write some tests,” and that is where things go off the rails.
—— Hamel Husain · [47:13]
我肯定可以在后端测试东西。我肯定可以为前端写测试。但我真正想知道的是,它可用吗?
指向原始笔记的链接
I can definitely test things on the back end. I can definitely write tests for the front end. But what I really want to know is, is it usable?
—— 嘉宾 · [04:56]
我们仍处于 AI 在最新的测试中智商大概是 120 的阶段,但当我们达到 500 IQ 的 AI 时,我们也许能找到治愈你母亲或我父亲疾病的办法。
指向原始笔记的链接
We’re still in this phase where AI is maybe 120 IQ with the latest tests, but when we get to 500 IQ AIs, we’ll maybe find cures for your mom or my dad’s disease.
—— Julien Bek · [75:18]
所以我担心,如果你以某种方式针对这种分数寻求或奖励黑客行为进行选择,并且你以一种天真的方式来做,第一,你可能会掩盖问题而不是修复它,第二,你实际上可能选择了那些具有看起来更漂亮这一长期目标的模型,因为你在非常强力地选择它们在你的测试中看起来不错。
指向原始笔记的链接
And so I’m worried that if you sort of select against this sort of score-seeking or reward-hacking behavior and you do it in a naive way, one, you might paper over the problem without fixing it, and two, you might actually select for models that have the longer run objective of looking good because you’re selecting really hard for them looking good on your tests.
—— Ryan Greenblatt · [19:37]
如果你不设干预阈值,智能体要么制造垃圾内容,要么烧出极高的 token 账单,因为它只会一直试下去,要么它们会伪造测试,就为了让测试通过。
指向原始笔记的链接
If you don’t do that, agents will either um create slop or run uh very high token bill because it’s just gonna keep trying, or they’re gonna fake the tests uh just to to to have the test passed.
—— Ran Arusi · [12:37]
它们可以轻松地为任何给定的 API 写测试,但和人类一样,它们无法写出正确的测试。
指向原始笔记的链接
They can easily write tests for any given API, but much like humans, they fail to write the right tests.
—— Jonathan Kelley · [15:05]
② 出现在这些集
1 集
③ 关联
点进去有真内容 —— 本页主要出口
Jonathan Kelley · Dioxys · Cognition · 智能体 · Rust · Blitz · Claude Code · 提示词工程 · 模糊测试 · 软件架构
