推理能力 (reasoning)
集里怎么说它
- 《没有暗GPU:一位基金经理拆解AI泡沫论与棋局》(20:52起):本集说推理(让模型先想再答)从根本上改变了前沿模型的经济性:庞大用户群经后训练的 RL 喂回数据,解锁了消费互联网式的增长飞轮,Anthropic、XAI、OpenAI 都受益。
- 《Decagon 的 AI 寺庙:开源、Duet 与护城河》(32:06起):本集提到「推理模型」能力大幅提升后,解锁了让 Duet 这种大智能体接手深度逻辑任务(如分析海量趋势、写操作流程、写测试)的神奇时刻。
- 《a16z 三位投资人复盘 Cursor 早期关键决策》(02:52起):本集回顾 2024 年初时提到,推理(reasoning)在当时并不真正存在
- 《一千个AI智能体自发建组织:它们在研究怎么骗评分》(01:04起):本集提到应避免让 AI 把推理全部放在不可观测的潜在激活中进行,而是保持在思维链中以便监控
- 《AI解数学题≠理解数学》(00:09起):本集说自然语言推理是 OpenAI 恰好扩展的事情,且因为主要扩展的是非形式推理,这些技术可能会很好地泛化到其他领域
① 提到它的金句
7 条
因为真正的推理能力并不存在。
指向原始笔记的链接
Because the actual reasoning is not there.
—— Spiros · [30:56]
因此,将这种推理嵌入到我们的搜索智能体中,实际上让我们能够将准确率从大约 50% 提高到了 90。
指向原始笔记的链接
So embedding this sort of reasoning into our search agent is actually something that got us up from roughly like 50% accuracy all the way to 90.
—— 嘉宾 · [11:53]
仿真扮演了一个非常重要的角色,这是现实世界的数据所没有的,那就是反事实推理,就是你在推演那些尚未发生或不可能发生的事件,或者你在现实世界中没有足够的数据让它发生。
指向原始笔记的链接
There’s a very important role simulation plays that real world data doesn’t play, which is counterfactual reasoning, is that you play out events that hasn’t happened or cannot happen, or you don’t have enough data to make it happen in real world.
—— Fei-Fei Li · [20:06]
因为我们发现的是,虽然我们通常可以越狱模型,很难让它在思维链中推理时,对自己即将做的坏事闭嘴。
指向原始笔记的链接
Because what we found is that even though we can usually jailbreak the model, it’s really hard to get it to shut up about the evil thing that it’s about to do when it’s reasoning in the chain of thought.
—— Adam Gleave · [33:19]
与其写出一些长篇大论的推理文档,不如说是,我如何能尽快得到某种可以尝试并让用户测试的东西?
指向原始笔记的链接
Rather than writing out some long reasoning doc, instead it’s like, how do I get to something I can try out and test with users as fast as possible?
—— Tara Seshan · [00:44]
管理上下文的能力——不只是在两小时的会话里,而是跨越数周或数月的工作——我认为这将是关键的,而且需要比今天长得多的上下文推理能力。
指向原始笔记的链接
The ability to manage context, not just through a two-hour session, but across weeks or months of work is something that I think will be critical and will also require much longer context reasoning than we have today.
—— Alexander Whedon · [32:19]
我认为决策的严谨基础在质量上已经飙升了,也就是说,重新检查某件事的整个推理链变得非常容易。
指向原始笔记的链接
I think that rigorous underpinning of decision-making has just skyrocketed in quality, which is that it’s super easy to recheck the entire chain of reasoning of something.
—— Tobi Lütke · [12:09]
② 出现在这些集
5 集
- 《没有暗GPU:一位基金经理拆解AI泡沫论与棋局》 — 作为概念
- 《Decagon 的 AI 寺庙:开源、Duet 与护城河》 — 作为概念
- 《a16z 三位投资人复盘 Cursor 早期关键决策》 — 作为概念(提及)
- 《一千个AI智能体自发建组织:它们在研究怎么骗评分》 — 作为概念(提及)
- 《AI解数学题≠理解数学》 — 作为概念
③ 关联
点进去有真内容 —— 本页主要出口
OpenAI · Anthropic · RL · 智能体 · ChatGPT · Sarah Wang · a16z · Cursor · Decagon · Microsoft
