LL

LLM 当裁判 (LLM as a judge)

概念
本站收录 5 集 · 2 条金句 · 关联 10

集里怎么说它

① 提到它的金句

2 条

我认为有很多它表现不太好的情况,但我觉得如果你把 LLM 当裁判的用途从一个基准重新定义为一个异常检测器,那它其实是可以的。
Well, I think there’s a lot of cases where it doesn’t work very well, but I think if you reframe the utility of an LLM as a judge from being a benchmark to being an anomaly detector, then I think it’s actually okay.
—— Ankur Goyal · [25:10]

指向原始笔记的链接

顺便说一句,这比大多数 LLM-as-a-judge 方法——也就是让一个 LLM 来判断那是优质的人类作品还是 AI 生成的 slop——表现都要好。
This performed better, by the way, than most LLM-as-a-judge methods of asking an LLM to judge if that is great human quality versus AI-generated slop.
—— Thais Castello Branco · [08:29]

指向原始笔记的链接

② 出现在这些集

5 集

③ 关联

点进去有真内容 —— 本页主要出口

智能体 · 评估 · 多智能体 · trace · Claude · Codex · Corinne Riley · Lenny · GrokBot · Vishu