judge
集里怎么说它
- 《Portola:当AI变成即兴演员,不是助手》(34:29起):本集说他们用 LLM 做 judge 评估 Tolan 输出质量,但不能泛泛地问’怎么样’,必须把人类品味细化注入到单句层面——‘这是好的第一句吗?应该在这里问问题吗?‘需要大量人工标注来 brute force 编码品味。
① 提到它的金句
9 条
我认为有很多它表现不太好的情况,但我觉得如果你把 LLM 当裁判的用途从一个基准重新定义为一个异常检测器,那它其实是可以的。
指向原始笔记的链接
Well, I think there’s a lot of cases where it doesn’t work very well, but I think if you reframe the utility of an LLM as a judge from being a benchmark to being an anomaly detector, then I think it’s actually okay.
—— Ankur Goyal · [25:10]
现在这听起来很吸引人,但它是一个非常危险的指标,因为很多时候,错误,它们只发生在长尾上,并且不经常发生,所以如果你只有 10% 的时间有错误,那么你可以很容易地通过让评判者一直说它通过来达到 90% 的一致性。
指向原始笔记的链接
Now that sounds appealing, but it’s a very dangerous metric to use, because a lot of times, errors, they only happen on the long tail and they don’t happen as frequently, so if you only have the error 10% of the time, then you can easily have 90% agreement by just having a judge say it passes all the time.
—— Hamel Husain · [58:41]
这是产品需求文档应该是什么样子的最纯粹的意义,就是这个评估判断器,它确切地告诉你它应该是什么,并且它是自动的且不断运行的。
指向原始笔记的链接
This is the purest sense of what a product requirements document should be, is this eval judge that’s telling you exactly what it should be, and it’s automatic and running constantly.
—— Lenny · [61:35]
我实际上使用 GPT-5.5 来给所有这些东西打分,因为它实际上是最严厉的评委。
指向原始笔记的链接
I actually use GPT-5.5 to grade all these things because it’s actually the harshest judge.
—— 嘉宾 · [25:21]
如果你只围绕某个特定的 LLM 裁判去爬山,你很可能是在对它过拟合。
指向原始笔记的链接
And if you’re only hill climbing on a specific LLM judge, you’re probably overfitting on it.
—— Vishu · [07:35]
你可以把它理解为一种方式:把你的技能转化成、或者说基于它们创建出非常聚焦、高准确率、非常快、非常便宜的「LM as judge」工具。
指向原始笔记的链接
You can think of it as a way of turning your skills into or creating off of them very focused, high accuracy, very fast, very cheap LM as judge tools
—— Drew · [29:03]
那是因为它们全都是针对测试去训练的。所以我们有一个理念:真正重要的是现实。现实是唯一一个你真正能信得过的裁判。
指向原始笔记的链接
That’s because they all trained to the test. And so we had this philosophy that it’s really about reality. Reality is the only judge that you can actually trust.
—— Anastasios Angelopoulos · [03:35]
瓶颈最终转移到了验证这一侧,因此,我们没有办法大规模地衡量价值,并以同样的速度判断质量。
指向原始笔记的链接
the bottleneck ends up shifting to the verification side, and thus, we don’t have a way to measure value at scale and to judge quality at the same speed.
—— Maximillian Piras · [12:20]
顺便说一句,这比大多数 LLM-as-a-judge 方法——也就是让一个 LLM 来判断那是优质的人类作品还是 AI 生成的 slop——表现都要好。
指向原始笔记的链接
This performed better, by the way, than most LLM-as-a-judge methods of asking an LLM to judge if that is great human quality versus AI-generated slop.
—— Thais Castello Branco · [08:29]
② 出现在这些集
1 集
- 《Portola:当AI变成即兴演员,不是助手》 — 作为概念
③ 关联
点进去有真内容 —— 本页主要出口
Quintin · Elliot · Portola · Tolan · LLM · 提示词 · 记忆 · 响应时间 · hook · 即兴演员
