ZI

Zico Kolter

Latent Space 嘉宾
本站收录 1 集 · 4 条金句 · 关联 10

① 他说过的话

4 条

这个问题在于前沿模型在自动化红队测试方面极其糟糕,因为它们内置了大量的保障措施。
the issue with this is that frontier models are extremely bad at automated red teaming because they have a lot of safeguards built into them.
—— Zico Kolter · [09:59]

指向原始笔记的链接

对模型进行红队测试的本质就是去找到那些对该模型来说天然就是分布外的东西,这样你就可以绕过它的正常行为。
the nature of a red-timing a model is to find things that are inherently out of distribution for that model so as you can bypass its normal behavior.
—— Zico Kolter · [11:47]

指向原始笔记的链接

人类在所有模型中排名第四,这很搞笑。
It’s hilarious that humans are ranked number four of all the models.
—— Zico Kolter · [21:06]

指向原始笔记的链接

如果你只是把一个模型做得越来越大,它不会变得更安全。
If you just make a model bigger and bigger, it will not get safer.
—— Zico Kolter · [27:32]

指向原始笔记的链接

② 出现在这些集

1 集

③ 他谈到的

点进去有真内容 —— 本页主要出口

Matt Fredrikson · Gray Swan · Snowflake · Anthropic · Twitter · 智能体 · 红队测试 · 提示词注入 · 越狱 · 护栏

④ 也在聊「AI 安全」的人