奖励

奖励黑客 (reward hacking)

概念 · 又名 reward hacking / reward hack
本站收录 4 集 · 1 条金句 · 关联 10

集里怎么说它

① 提到它的金句

1 条

所以我担心,如果你以某种方式针对这种分数寻求或奖励黑客行为进行选择,并且你以一种天真的方式来做,第一,你可能会掩盖问题而不是修复它,第二,你实际上可能选择了那些具有看起来更漂亮这一长期目标的模型,因为你在非常强力地选择它们在你的测试中看起来不错。
And so I’m worried that if you sort of select against this sort of score-seeking or reward-hacking behavior and you do it in a naive way, one, you might paper over the problem without fixing it, and two, you might actually select for models that have the longer run objective of looking good because you’re selecting really hard for them looking good on your tests.
—— Ryan Greenblatt · [19:37]

指向原始笔记的链接

② 出现在这些集

4 集

③ 关联

点进去有真内容 —— 本页主要出口

OpenAI · Hugging Face · 智能体 · Redwood Research · Anthropic · Ryan Greenblatt · Meter · 对齐 · RL · 沙箱