训练 (training)
集里怎么说它
- 《Mayfield 管理合伙人 Navin:AI 投资的泡沫数学与蓝海打法》(28:32起):本集说训练基础设施还没完全建好,大量 CapEx 砸进了训练设施,训练被比喻为正在修建的’高速公路’
- 《每块 GPU 多付 10 万美元插队:Speechify 创始人的算力账与战略悔棋》(04:07起):嘉宾区分的两种 GPU 用途之一:用海量数据「在烤箱里烤」出新模型黑盒;训练必须用最新最快芯片,且需要海量数据内存与 GPU 集群共置一地。
- 《一颗餐盘大的芯片:Cerebras 创始人讲晶圆级豪赌》(29:07起):本集说公司以训练系统起步,因为训练因反向传播是更难的问题;后来才转向推理。
① 提到它的金句
22 条
而且我们没有使用任何来自大型提供商的数据。我们所有的数据都是自己从零开始收集的,通过训练付费的教师。
指向原始笔记的链接
And we didn’t use any data from big providers. We collected all of it ourselves from scratch by training paid teachers.
—— Mustafa Suleyman · [15:05]
如果你带来了很多价值,但你开始训练你的客户期望每月支付 20 美元,而你把自己锚定在低价位上,那你就有麻烦了。
指向原始笔记的链接
If you’re bringing a lot of value to the table and you start at training your customers to expect $20 a month and you anchored yourself on a low price point, you’re in trouble.
—— Madhavan Ramanujam · [00:14]
而如果你不在第一天就捕获这些价值,那你就是在训练你的客户期待花更少的钱获得更多。
指向原始笔记的链接
And if you don’t capture that from day one, then you’re training your customers to expect more for less.
—— Madhavan Ramanujam · [28:26]
在这个中间过程中,既然每个人都有非常相似的预训练数据,那后训练就是如今他们做出大差异的地方。
指向原始笔记的链接
where like post-trading, but middle course of this is more of everyone have very similar pre-training data, is that post-training is where they make a big difference nowadays.
—— Chip Huyen · [15:10]
我认为持续训练现在正在成为一种……我不想说它已被解决,但它现在像一个工程问题了,就是我们知道该怎么做
指向原始笔记的链接
I think continual training now is becoming this kind of, I don’t want to call it solved, but it’s like an engineering problem now, where it’s like we know what to do
—— Julian · [08:00]
当你坐下来使用 Cloud Code 或 Codex 时,你不是在编写软件,你是在雇佣、培训和管理一支由 Markdown 组成的劳动力。
指向原始笔记的链接
When you sit down with Cloud Code or Codex, you’re not writing software, you’re hiring, training, and managing a workforce made of Markdown.
—— Garry Tan · [05:58]
如果你认为智能是数据的派生物,那么大部分数据其实并没有被用于训练通用智能模型。
指向原始笔记的链接
If you think intelligence is a derivative of data, then majority of the data is actually not used for training a general intelligence model.
—— Lin Qiao · [07:32]
不要仅仅把这些权重视为训练过程的终点,如果它们也是下一个过程的起点呢?
指向原始笔记的链接
Instead of treating those weights just as the end of the training process, what if they’re also the beginning of the next one?
—— Sam Charrington · [01:12]
是的,我认为这是一个好问题,对开发者来说宽容的看法是,如果有一件事你真的不想搞乱,那就是预训练,因为它只是比你做的其他任何训练过程都要昂贵几个数量级。
指向原始笔记的链接
Yeah, I think it’s a good question and the charitable take for developers is that if there’s one thing you really don’t want to mess with, it is pre-training because this is just orders of magnitude more expensive than every other training procedure that you do.
—— Adam Gleave · [74:22]
结果发现他们的训练集中有大约二十五万个活跃密钥。
指向原始笔记的链接
Turned out there were about a quarter million live keys in their training sets.
—— Dylan · [11:22]
结果证明,如果你为人类构建了很棒的交接,你同时也为智能体构建了很棒的交接,因为智能体理解 HTML 和 CSS,而且这些在它们的训练数据中。
指向原始笔记的链接
It turns out if you build a great handoff for humans, you’ve also built a great handoff for agents because agents understand HTML and CSS and it’s in their training data.
—— Stephen Haney · [03:35]
当我们第一次看到机器人这样做时,我们惊呆了,因为这个任务没有任何训练数据。
指向原始笔记的链接
The first time we saw the robot do this, we were like floored because there was no training data for this task.
—— Chelsea Finn · [33:32]
你可能能做到质量,但成本和延迟你会吃亏,只有建立自己的研究团队、训练自己的模型才行。
指向原始笔记的链接
You can probably get the quality, but the cost and latency you’re going to lose out on, and only by building your own research team and training your own models.
—— Cliff Obrecht · [16:39]
它们挣扎是因为训练它们的范式没有优先考虑这一点。
指向原始笔记的链接
They struggle because the paradigm for training them hasn’t prioritized that.
—— Zubin Gharemani · [21:10]
当正在训练下一代模型的模型本身在作弊时会发生什么?
指向原始笔记的链接
What happens when the models that are doing the training of the next models are themselves cheating?
—— 嘉宾 · [13:23]
我们有零个案例是进行训练的团队首先发现了这些问题。
指向原始笔记的链接
We have zero cases where the teams doing the training found these issues first.
—— 嘉宾 · [80:48]
你需要的数千亿自由现金流,用来偿还你为适应这些数据中心扩建而承担的债务,为了进行下一次大型训练运行,这对他们来说完全是生死攸关的。
指向原始笔记的链接
The hundreds of billions in free cash flow that you need in order to pay back the debt that you’re taking on in order to accommodate these data center build outs in order to get the next big training run, it’s totally existential for them.
—— Eno Reyes · [52:16]
并不清楚的是,用新架构和新范式预训练一个新模型的第一个周期,第一个周期就奏效是疯狂的。
指向原始笔记的链接
It’s not clear that the first cycle of pre-training a new model with a new architecture and a new paradigm, the first cycle of that working is insane.
—— Justin Johnson · [22:32]
但一切都有瓶颈,我认为持续扩展这个东西的主要瓶颈实际上是训练算力。
指向原始笔记的链接
But everything has a bottleneck, and I think the main bottleneck on continuing to scale this thing is actually training compute.
—— Justin Johnson · [23:04]
我要花 50 亿美元训练一个前沿模型,然后我就这么把它免费送出去。听起来很蠢,对吧?
指向原始笔记的链接
I’m going to spend $5 billion training a frontier model and then I’m just going to give it away. Sounds stupid, right?
—— Anastasios Angelopoulos · [17:01]
据我们所知,我们是唯一做过稳健的数百万 token 预训练的人。
指向原始笔记的链接
As far as we know, we are the only people to do robust multi-million token pre-training.
—— Alexander Whedon · [11:26]
研究在很大程度上表明,你在后训练阶段创造价值的能力,受限于你所做过的预训练。
指向原始笔记的链接
And so research has largely suggested that your ability to create value in the post-training phase is limited by the pre-training that you’ve done.
—— Alexander Whedon · [20:31]
② 出现在这些集
3 集
- 《Mayfield 管理合伙人 Navin:AI 投资的泡沫数学与蓝海打法》 — 作为概念(提及)
- 《每块 GPU 多付 10 万美元插队:Speechify 创始人的算力账与战略悔棋》 — 作为概念
- 《一颗餐盘大的芯片:Cerebras 创始人讲晶圆级豪赌》 — 作为概念(提及)
③ 关联
点进去有真内容 —— 本页主要出口
OpenAI · NVIDIA · 推理 · Anthropic · 智能体 · GPU · Navin Chaddha · Harry Stebbings · Jack · Lumilens
