基准

基准测试 (benchmark)

概念
本站收录 10 集 · 4 条金句 · 关联 10

集里怎么说它

① 提到它的金句

4 条

我认为有很多它表现不太好的情况,但我觉得如果你把 LLM 当裁判的用途从一个基准重新定义为一个异常检测器,那它其实是可以的。
Well, I think there’s a lot of cases where it doesn’t work very well, but I think if you reframe the utility of an LLM as a judge from being a benchmark to being an anomaly detector, then I think it’s actually okay.
—— Ankur Goyal · [25:10]

指向原始笔记的链接

那些基准测试不代表你的工作负载。正如我们之前说的,AI 的边界是参差不齐的——某个模型可能在基准测试上做得很好,但这并不必然意味着它对你的工作负载也会做得很好。
They don’t represent your workloads. And as we said earlier, the boundary of AI is jagged, right? So although a certain model might do very well on a benchmark, it doesn’t necessarily translate into that it will do very well for your workloads as well.
—— Chetan Gupta · [38:19]

指向原始笔记的链接

不应该是某个自以为聪明的家伙随便搞出来的什么破基准测试。
There shouldn’t be some random f**ing benchmark that’s made by some dude that thinks they’re smart.*
—— Anastasios Angelopoulos · [06:28]

指向原始笔记的链接

当我们做基准测试时,它们在代码评审上是超越人类的。
When we benchmark them, it’s like they’re superhuman in code review.
—— Tibo Sottiaux · [47:53]

指向原始笔记的链接

② 出现在这些集

10 集

③ 关联

点进去有真内容 —— 本页主要出口

智能体 · Anthropic · OpenAI · 评估 · ChatGPT · Claude Code · 开源 · OpenRouter · NVIDIA · Codex