AD

Adam Gleave

The Cognitive Revolution 主持
本站收录 1 集 · 6 条金句 · 关联 10

① 他说过的话

6 条

因为我们发现的是,虽然我们通常可以越狱模型,很难让它在思维链中推理时,对自己即将做的坏事闭嘴。
Because what we found is that even though we can usually jailbreak the model, it’s really hard to get it to shut up about the evil thing that it’s about to do when it’s reasoning in the chain of thought.
—— Adam Gleave · [33:19]

指向原始笔记的链接

模型必须真正地推理并理解你的有害意图,并在数千个 token 中配合它。
The model has to really reason about and understand your harmful intention and go along with it for thousands of tokens.
—— Adam Gleave · [62:11]

指向原始笔记的链接

坏消息是,我们越狱一个开源权重模型从未超过几个小时。
And the bad news is that it’s never taken us more than a few hours to jailbreak a open weight model.
—— Adam Gleave · [69:08]

指向原始笔记的链接

是的,我认为这是一个好问题,对开发者来说宽容的看法是,如果有一件事你真的不想搞乱,那就是预训练,因为它只是比你做的其他任何训练过程都要昂贵几个数量级。
Yeah, I think it’s a good question and the charitable take for developers is that if there’s one thing you really don’t want to mess with, it is pre-training because this is just orders of magnitude more expensive than every other training procedure that you do.
—— Adam Gleave · [74:22]

指向原始笔记的链接

所以这不是说我们不能对齐这些系统,而是一个巨大的控制内部监控失败,其中沙箱是不够的。
So it’s not that we can’t align these systems, but it is a massive control internal monitoring failure where the sandbox was insufficient.
—— Adam Gleave · [80:47]

指向原始笔记的链接

所以从这个角度来看,我认为感到相当受惊是正确的,即你的那种普通的甚至良好的但非偏执的安全实践不再足以遏制智能体了。
So from that perspective, I think it is right to be quite spooked that your sort of run of the mill or even goods, but not paranoid security practice is not enough to contain agents any longer.
—— Adam Gleave · [83:27]

指向原始笔记的链接

② 出现在这些集

1 集

③ 他谈到的

点进去有真内容 —— 本页主要出口

FAR AI · 通用越狱 · 社会工程 · 思维链 · 护栏 · 探针 · 预训练数据过滤 · 安全补全 · 越狱税 · 后训练

④ 也在聊「AI 安全」的人