Alexander Whedon
① 他说过的话
9 条
据我们所知,我们是唯一做过稳健的数百万 token 预训练的人。
指向原始笔记的链接
As far as we know, we are the only people to do robust multi-million token pre-training.
—— Alexander Whedon · [11:26]
研究在很大程度上表明,你在后训练阶段创造价值的能力,受限于你所做过的预训练。
指向原始笔记的链接
And so research has largely suggested that your ability to create value in the post-training phase is limited by the pre-training that you’ve done.
—— Alexander Whedon · [20:31]
我们在编码方面有意走得慢一些,因为我们已经意识到编码在多大程度上是一场预算游戏——你有多少数据预算?
指向原始笔记的链接
We are intentionally moving a little bit more slowly on the coding side because we’ve kind of realized like how much coding is like a budget game, like how much data budget do you have?
—— Alexander Whedon · [27:49]
一个是我们会有泛化能力更好的系统,因为每当你增加一个搜索引擎、或向量数据库(它就是一种搜索引擎)、或某种步骤间的路由逻辑,这种人工整理就真的限制了系统做很多不同事情的能力。
指向原始笔记的链接
One is we will have systems that can generalize better because every time you add a search engine or a vector database, which is a search engine, or some conditional logic to route between the steps, this human curation really limits the ability of that system to do a lot of different things.
—— Alexander Whedon · [30:24]
我们希望随着时间的推移,一百万 token 在智能水平、成本、延迟等方面,感觉起来像五万 token。
指向原始笔记的链接
We want a million tokens to feel like 50,000 tokens in terms of intelligence, cost, latency, et cetera, over time.
—— Alexander Whedon · [31:20]
管理上下文的能力——不只是在两小时的会话里,而是跨越数周或数月的工作——我认为这将是关键的,而且需要比今天长得多的上下文推理能力。
指向原始笔记的链接
The ability to manage context, not just through a two-hour session, but across weeks or months of work is something that I think will be critical and will also require much longer context reasoning than we have today.
—— Alexander Whedon · [32:19]
为此,我们真的希望人们觉得输入 token 是免费的——只管为你的问题考虑所需要的上下文就好。
指向原始笔记的链接
To make it, we really want people to feel like the input tokens are free, like just consider the context that you need to for your problem.
—— Alexander Whedon · [45:32]
我不认为只靠一个 LLM 就能做到。它们今天还不够,我也不认为它们有足够的创造力,但它们确实非常有帮助。
指向原始笔记的链接
I don’t think you can get there with just an LLM. They’re not enough today. I don’t think they’re creative enough either, but they’re definitely super helpful.
—— Alexander Whedon · [47:11]
我们想把算法范式的更替从每九年一次,缩短到每 12 个月一次。
指向原始笔记的链接
We want to move from shifting the algorithmic paradigm every nine years to doing it every 12 months.
—— Alexander Whedon · [49:15]
② 出现在这些集
1 集
③ 他谈到的
点进去有真内容 —— 本页主要出口
SubQuadratic · 稀疏注意力 · 上下文工程 · 智能体 · RAG · 预训练 · DeepSeek Sparse Attention · KVCache · Transformer · Opus 4.6
