未对齐 (misalignment)
集里怎么说它
- 《超级智能为什么危险:Ryan Greenblatt 的推演与解法》(08:39起):本集反复讨论未对齐的不同表现:alignment faking、reward hacking、长期权力寻求等,认为问题会随 AI 能力增强而变得更难控制
① 提到它的金句
3 条
我认为,我使用 Claude 最有效的方式之一,或者我使用 Claude 最有趣的方式之一,我看到内部很多人都在做,是帮你识别不一致。
指向原始笔记的链接
one of the, I think, most effective ways I use Claude, or one of the most interesting ways I use Claude, which I see a number of people doing internally is to help you identify misalignment.
—— Amol Avasare · [56:05]
这是 Google 内部激励机制的错位,是 Google 内部利益的错位,几天后这似乎通过 Demis 的离开而显现出来了。
指向原始笔记的链接
This is a misalignment of incentives within Google, a misalignment of interest within Google that a couple of days later seems to have manifested with Demis’ departure.
—— 嘉宾 · [00:46]
所以走得更快意味着一个小小的错位可能滚雪球成大量浪费的工作,而且那工作消耗代币,代币现在要花真金白银。
指向原始笔记的链接
So going faster means that a small misalignment can snowball into a ton of wasted work, and that work costs tokens, and tokens cost real money now.
—— Idan Gazit · [04:15]
② 出现在这些集
1 集
③ 关联
点进去有真内容 —— 本页主要出口
Ryan Greenblatt · Matt Turk · Redwood Research · OpenAI · Anthropic · Google DeepMind · Meta · Hugging Face · 超级智能 · AI 控制
