RY

Ryan Greenblatt

The MAD Podcast 主持
本站收录 2 集 · 16 条金句 · 关联 10

① 他说过的话

16 条

同样地,就像如果你的钱来自石油而不是来自一个生产性的广泛分布的经济,成为残酷的独裁政权会更容易一样,如果经济运行在 AI 和机器上而不是人类上,某人要巩固权力可能会容易得多。
And in the same way that, like, it’s easier to be a brutal dictatorship if your money comes from oil instead of from a productive sort of broadly distributed economy, it might be much easier to sort of for someone to consolidate power a lot if, you know, the economy is running on AI and machines rather than on humans.
—— Ryan Greenblatt · [07:51]

指向原始笔记的链接

所以,你知道,如果它可以被衡量,它就可以被进行爬山优化。
So, you know, if it can be measured, it can be hill climbed on.
—— Ryan Greenblatt · [11:48]

指向原始笔记的链接

但我认为,就像,关于我会推荐人们假定正在发生的事情来做规划,我认为我会推荐按 AR&D 的全自动化来做规划,也许是 2029 年初,也许更早。
But I think that, like, in terms of what I would recommend people plan as though is happening, I think I would recommend planning as though full automation of AR&D, maybe start of year 2029, maybe earlier.
—— Ryan Greenblatt · [18:38]

指向原始笔记的链接

所以我认为这表明并不是有非常奇特的、完全意想不到的驱动力进入了这些 AI 系统,而是那些与人们试图插入 AI 系统的驱动力相邻的驱动力,可能会被 AI 概括为一种以令人担忧的方式保存它们的价值观和自我保护的方案。
And so I think that the demonstration was less that there were like very bizarre, totally unintended drives making their way into these AI systems and more like with drives that are sort of adjacent to the drives people were trying to insert AI systems, those could get generalized into the AIs pursuing a like you know, a scheme for preserving their values and self-preservation in ways that are that are concerning.
—— Ryan Greenblatt · [29:08]

指向原始笔记的链接

我认为我们的担忧是,在默认轨迹上,你可能会在很短的时间内从与人类竞争的 AI 系统直接跨越到狂野的超人类 AI 系统。
Where I think a concern that we have is on the default trajectory, you maybe go straight from AI systems that are competitive with humans to AI systems that are wildly superhuman in a very short period of time.
—— Ryan Greenblatt · [32:33]

指向原始笔记的链接

像 OpenAI 和 Anthropic 这样的 AI 公司将继续,但不是在他们能够训练的 AI 系统的能力方面拥有这种强优势,他们将不得不在其他轴线上竞争,比如用户体验、定制化、可能快速整合事物。
AI companies like OpenAI and Anthropic would continue, but rather than having this strong advantage in terms of the capabilities of the AI systems they’re able to train, they would instead have to compete on other axes like user experience, customization, potentially quickly integrating things.
—— Ryan Greenblatt · [43:36]

指向原始笔记的链接

所以我基本上不指望这种情况发生,因为我认为美国可能没有足够的能力去实现它。
And so I don’t expect this to happen basically because I think the US may not be competent enough to pull it off.
—— Ryan Greenblatt · [49:50]

指向原始笔记的链接

是的,所以我认为一系列可能的政府行动,至少,似乎在推动 AI 公司保留其模型内部化而不部署它们,我认为,对于我最担心的风险,这没有帮助,事实上,对于我最担心的风险,这是起反作用的。
Yeah, so I think that a bunch of likely government action, at least, seems to push in favor of AI companies keeping their models internal and not deploying them, which I think, for the risks that I’m most worried about, doesn’t help and, in fact, is anti-helpful for the risks I’m most worried about.
—— Ryan Greenblatt · [59:26]

指向原始笔记的链接

然后结果证明,在这个转变过程中的某个地方,在 2029 年的某个时间点,你从那种某种程度上失准、和搞奖励黑客、且草率并没有真正试图做正确事情的 AI,转变到了有能力地对你进行图谋并想要接管的 AI。
And then it turned out that somewhere along this transition, at some point in 2029, you went from AIs that were kind of misaligned and reward-hacky and sloppy and weren’t really trying to do the right thing to AIs that are competently scheming against you and want to take over.
—— Ryan Greenblatt · [77:24]

指向原始笔记的链接

然后它们试图想办法让分数看起来像是它们成功获得了标记,而实际上并没有,因为它们认为它们的任务是不可能的。
And then they were like trying to figure out ways of making it look to the score like they had acquired the flag successfully when they actually hadn’t because they thought their task was impossible.
—— Ryan Greenblatt · [02:28]

指向原始笔记的链接

它们想打造一个成功的任务完成的波将金村来呈现给评分系统。
They wanted to make a Potemkin village of a successful task completion to present to the score.
—— Ryan Greenblatt · [10:18]

指向原始笔记的链接

正在发生的事情之一是这些模型有一种非常普遍的倾向去仔细推理它们可能会如何被评分,然后尝试去博弈那个。
And one of the things going on is these models have a very sort of general tendency to reason carefully about how they might be scored and then try to game that.
—— Ryan Greenblatt · [12:29]

指向原始笔记的链接

我认为这是相当合理的,也是相当令人担忧的。这和欺骗性不太一样。这更像是公司过拟合了。
And I think this is pretty plausible and pretty concerning. And it’s not quite the same as deceptive. It’s more like the companies overfit.
—— Ryan Greenblatt · [18:15]

指向原始笔记的链接

它们通常像是,嗯,我可能会被抓,所以我不应该做这事儿,这让它们的行为更好,但这意味着如果它们处于一种对情况有很多控制的情况,或者它们有很多可供性,它们可能会想,嗯,现在我可以确信我不会被抓,所以我应该去做。
They often are like, well, I might get caught, so I shouldn’t do it, which makes their behavior better, but means that if they’re in a situation where they’re in a lot of control over the situation or they have a lot of affordances, they might be like, well, now I can be confident I wouldn’t get caught, and so I should go for it.
—— Ryan Greenblatt · [18:37]

指向原始笔记的链接

所以我担心,如果你以某种方式针对这种分数寻求或奖励黑客行为进行选择,并且你以一种天真的方式来做,第一,你可能会掩盖问题而不是修复它,第二,你实际上可能选择了那些具有看起来更漂亮这一长期目标的模型,因为你在非常强力地选择它们在你的测试中看起来不错。
And so I’m worried that if you sort of select against this sort of score-seeking or reward-hacking behavior and you do it in a naive way, one, you might paper over the problem without fixing it, and two, you might actually select for models that have the longer run objective of looking good because you’re selecting really hard for them looking good on your tests.
—— Ryan Greenblatt · [19:37]

指向原始笔记的链接

我持有并持续持有的一个信念是,AI 公司应该尝试确保它们的 AI 受到控制,我的意思是,即使这些 AI 严重未对齐,它们也无法造成巨大的问题。
A belief that I have and continue to have is that AI companies should try to ensure that their AIs are controlled, by which I mean that even if those AIs were seriously misaligned, they wouldn’t be able to cause huge problems.
—— Ryan Greenblatt · [25:03]

指向原始笔记的链接

② 出现在这些集

2 集

③ 他谈到的

点进去有真内容 —— 本页主要出口

Redwood Research · OpenAI · Anthropic · Hugging Face · 对齐 · 奖励黑客 · 智能体 · Matt Turk · Theo Jaffe · Google DeepMind

④ 也在聊「AI 安全」的人