VI

Vishu

精选演讲 主持
本站收录 1 集 · 6 条金句 · 关联 10

① 他说过的话

6 条

但现在我们谈的是 Claygent 上数十亿次的运行,以及让 Sculptor 为你做端到端任务,或者说这些运行时间非常长的任务,评估就变成了不可协商的。
But now that we’re talking about billions of runs on Claygent and having Sculptor do end-to-end tasks for you or these really long-running tasks, evals became non-negotiable.
—— Vishu · [02:11]

指向原始笔记的链接

如果你有一套好的评估套件,你可以让 Claude、Codex 之类的进去替你修改提示词,而你确信你不会发布任何会毁掉生产环境的东西。
If you have a good eval suite, you can let Claude or Codex or Devon kind of go in, make prompt changes for you, let BlinkchainEngine make prompt changes for you, and you know that you’re not shipping anything that is going to ruin production.
—— Vishu · [02:26]

指向原始笔记的链接

那个智能体太吵了,它就像另一个你需要管理、保持更新、还要为它做 eval 的智能体,所以最终不值得。
The agent was like too noisy and it was just like another agent that you had to manage and keep up to date and also have evals for and so it just ended up not being worth it.
—— Vishu · [05:49]

指向原始笔记的链接

裁判漂移——所有这些模型和模型家族都有它们自己的内在偏好。
JudgeDrift, all of these models and model families have their own internal biases.
—— Vishu · [07:30]

指向原始笔记的链接

如果你只围绕某个特定的 LLM 裁判去爬山,你很可能是在对它过拟合。
And if you’re only hill climbing on a specific LLM judge, you’re probably overfitting on it.
—— Vishu · [07:35]

指向原始笔记的链接

同理,如果你只用一个很小的 eval 集去爬山,你的 prompt 很可能开始只是镜像那些 eval 示例。
Same if you’re only hill climbing with like a small eval set, your prompt is probably going to start to mirror just those eval examples.
—— Vishu · [07:41]

指向原始笔记的链接

② 出现在这些集

1 集

③ 他谈到的

点进去有真内容 —— 本页主要出口

Clay · Claygent · Sculptor · LangChain · 智能体 · 评估 · trace · LLM 当裁判 · harness · 数据湖

④ 也在聊「智能体」的人