Vishu
① 他说过的话
6 条
但现在我们谈的是 Claygent 上数十亿次的运行,以及让 Sculptor 为你做端到端任务,或者说这些运行时间非常长的任务,评估就变成了不可协商的。
指向原始笔记的链接
But now that we’re talking about billions of runs on Claygent and having Sculptor do end-to-end tasks for you or these really long-running tasks, evals became non-negotiable.
—— Vishu · [02:11]
如果你有一套好的评估套件,你可以让 Claude、Codex 之类的进去替你修改提示词,而你确信你不会发布任何会毁掉生产环境的东西。
指向原始笔记的链接
If you have a good eval suite, you can let Claude or Codex or Devon kind of go in, make prompt changes for you, let BlinkchainEngine make prompt changes for you, and you know that you’re not shipping anything that is going to ruin production.
—— Vishu · [02:26]
那个智能体太吵了,它就像另一个你需要管理、保持更新、还要为它做 eval 的智能体,所以最终不值得。
指向原始笔记的链接
The agent was like too noisy and it was just like another agent that you had to manage and keep up to date and also have evals for and so it just ended up not being worth it.
—— Vishu · [05:49]
裁判漂移——所有这些模型和模型家族都有它们自己的内在偏好。
指向原始笔记的链接
JudgeDrift, all of these models and model families have their own internal biases.
—— Vishu · [07:30]
如果你只围绕某个特定的 LLM 裁判去爬山,你很可能是在对它过拟合。
指向原始笔记的链接
And if you’re only hill climbing on a specific LLM judge, you’re probably overfitting on it.
—— Vishu · [07:35]
同理,如果你只用一个很小的 eval 集去爬山,你的 prompt 很可能开始只是镜像那些 eval 示例。
指向原始笔记的链接
Same if you’re only hill climbing with like a small eval set, your prompt is probably going to start to mirror just those eval examples.
—— Vishu · [07:41]
② 出现在这些集
1 集
- 《Clay 的智能体矩阵:如何为数十亿次运行建评估》 — 作为主持
③ 他谈到的
点进去有真内容 —— 本页主要出口
Clay · Claygent · Sculptor · LangChain · 智能体 · 评估 · trace · LLM 当裁判 · harness · 数据湖
