token
集里怎么说它
- 《不会写代码的人如何成为全职 vibe coder》(16:00起):本集将 token 描述为稀缺资源,用阿拉丁神灯三个愿望的比喻说明上下文记忆窗口有限;如果不给文件引用,agent 会把 80% 的 token 花在阅读代码上
- 《Claude Code 负责人:写代码已被解决,下一步是什么》(27:14起):本集说 Boris 建议 CTO 们先给工程师尽可能多的 token,小规模下 token 成本相对其他业务成本很低,Anthropic 内部已有工程师每月花数十万美元在 token 上
- 《DevOps 之父谈智能体开发:谁来管、怎么管、别踩什么坑》(31:33起):Daniel指出既然是智能体在做工作、消耗token及成本,公司突然开始关心软件开发生命周期的效率;Tamuz发现不感兴趣的开发者会把验证甩给智能体导致token消耗飙升
- 《「智能是数据的派生物」:Fireworks 创始人 Lin Kuo 的专用智能宣言》(00:12起):本集给出核心预测:未来三年 token 成本降 10 倍、驱动 100 倍使用量;并非所有 token 平等,应按任务建立评估 token 经济学的最佳实践——有的模型便宜 2 倍却要用 2 倍的量。
- 《Factory CEO Matan:早两年等于错,退款、路由器与软件工厂》(18:20起):本集从 token maxing 讲到成本理性:有银行每月花几十万美元问 Opus 天气;每个 CIO 都要回答’每多一个 token 放在哪’,进而变成’每多一美元投人力还是 token’。
- 《Heitor:用智能体重塑软件工程工作流的实操蓝图》(33:16起):本集作为成本度量单位,嘉宾提到一次重构花掉 2 亿个 token 才意识到必须停止全程用最贵模型,在 1400 人组织中每个工程师每月几千美元的 token 费用会引发领导层质疑
- 《Datadog 4000 人AI赋能实战:删掉上下文反而更好》(17:49起):在成本管理讨论中出现。Datadog 刻意不走配额限制路线,而是从系统角度去压缩输出、减少 token 浪费。
- 《AI 投资泡沫的崩盘剧本:为什么万亿美元建数据中心注定亏钱》(16:00起):本集把它说成:数据中心这个工厂生产的、史上贬值最快的商品,在恒定性能下每年跌价 70% 到 80%,持续至少四年没有停下来的迹象。
- 《给智能体建一个“人力资源部”:TrustWise 创始人谈运行时治理》(32:22起):本集说从生成式 AI 到智能体 AI,token 消耗量是两年前系统的 20 到 40 倍,一个输入可能触发 50 个动作,智能体可能陷入循环吃掉大量 token
- 《Ed Zitron:生成式 AI 是一场万亿级骗局》(11:47起):本集解释为约四分之三个单词,是 AI 公司的计费货币,像出租车的计价器;月费订阅隐藏了实际 token 消耗,200 美元月费可烧掉 14000 美元的 token。
- 《数据成了企业唯一的护城河:AI时代的数据基建怎么做》(27:18起):本集说我们不再处于 token 最大化(token maxing)的时代了,要为每一个 token 获取价值,因为它变得越来越贵
- 《NVIDIA 布局全栈、OpenAI 被迫上市与 AI 资本的”第五名效应”》(09:43起):本集说企业和个人已经’对 token 上瘾了’——‘我需要我的 10 个子智能体全天候 24 小时运行来做我的工作,否则我就辞职’;但 CFO 面临硬约束,token 账单会直接冲击 EPS
- 《最便宜的模型反而是最便宜的:Factory CTO 谈 AI 定价陷阱》(07:45起):本集说 token 本质上是智能,获得 token 要花钱,整个链条是把能量转化为智能、中间用美元交易;Factory 按项目和结果分配 token 而非按人头
- 《Token 都烧在哪了:Cursor 工程师教你把 AI 编程成本打下来》(02:07起):本集把它说成:模型输入输出的基本单位,平均一个词对应一到两个 token;分输入、输出、缓存写入、缓存读取四种,输出最贵,是计费的核心单位
- 《Ollama CEO:开源模型正吃掉企业 80-90% 的 token》(01:36起):本集贯穿的核心指标:token 消耗量、每 token 成本与每任务成本,Ollama 云端单用户用量年初以来涨 150 倍,Flash 模型会最先带来「无限 token」。
- 《开源权重不是威胁:Box CEO 聊 AI 的经济账》(15:57起):本集认为竞争会迫使 token 成本趋近基础设施成本——基础设施成本之上加 20-40% 而非 70-90%。
- 《邮箱里的 AI 助手:Plaid 前 CTO 谈如何在巨头围剿下赢》(41:57起):本集的核心经济变量:token 用得越多不算成功,token 最大化是危险思维;Town 反而会邮件提醒用户失控的 routine 在狂烧 token;免费内测期有用户五个月烧掉两万六千美元。
- 《斯坦福语言学家的代币经济学:你的 token 贬值了》(30:37起):本集讨论 AI 计费与价值的基本单位:其真实成本估计每 1 美元可能是 2 到 20 美元,且用途正从代码生成迁移到思考和解释
- 《鼠标力:为智能体时代找回「马力」这把尺子》(02:20起):本集说 token 只是系统的输出而非价值本身,必须被干净地追溯到成果(消灭多少 bug、关闭多少支持请求);行业内存在超支+使用不足的厄运循环。
- 《不到10人管7个SaaS:让智能体替你做营销的实操系统》(11:01起):本集说智能体工作流烧 token 极凶,有 YouTube 博主估算 200 美元订阅相当于每月 1.5 万美元的 API token 用量,公司等于在价格倾销。
① 提到它的金句
26 条
你可以尝试按 token 计费,然后你去找会计师事务所的 CEO,给他们这张按 token 计费的账单。他们看你的眼神就像你是个疯子。
指向原始笔记的链接
You could try to have a very metered approach or do per token and then you’re going to your accounting firm CEO and you’re giving them this bill per token. They look at you like you’re nuts.
—— 嘉宾 · [18:44]
通常,如果你使用最强大的模型,实际上反而更便宜,且消耗更少 token,因为它可以只用更少的纠错、更少的指导等就把同样的事情做得更快。
指向原始笔记的链接
Often, it’s actually cheaper and less token-intensive if you use the most capable model, because it can just do the same thing much faster with less correction, less hand holding, and so on.
—— Boris Cherny · [69:39]
我认为把 token 花费作为吹嘘的指标存在真正的危险,这就像人们吹嘘他们一天写了多少行代码一样。
指向原始笔记的链接
I think there’s a real danger in making the token spend the metric to boast about, which is the same as when people boast about how many lines of code they’ve written in a day.
—— Max Schoening · [36:08]
把你花在尝试另一个工具的时间看作拥有一个创新代币或一个创新预算,对吧?你没有世界上所有的时间。
指向原始笔记的链接
Think about the time that you take to invest in trying in another tool as having an innovation token or an innovation budget, right? You don’t have all the time in the world.
—— 嘉宾 · [08:46]
但下一层是关于,好,如果 token 并不是真正可互换的,你需要给它们分配不同的工作,比如这个 token 负责建议、那个负责执行,这个在「做梦」、那个在执行,诸如此类,你会想开始组合这些协同配合的编排式策略
指向原始笔记的链接
But the next one is about, okay, if tokens aren’t really fungible and you need to give them different jobs, like maybe this token is advising versus this token is executing, this token is dreaming versus this token is executing, so on and so forth, you want to start composing these kind of orchestrated strategies that go together
—— 嘉宾 · [11:34]
我确实认为 token 的成本会大幅下降,未来三年内成本降低 10 倍,而这 10 倍的成本降低将驱动 100 倍的使用量。
指向原始笔记的链接
I do think the cost of token will go down drastically, 10x cost reduction in the next three years, and this 10x cost reduction will drive 100x usage.
—— Lin Qiao · [00:12]
所以每一位 CIO 都需要回答:每多一个 token,我们该把它放在哪里?
指向原始笔记的链接
So every CIO is going to need to answer, for every incremental token, where do we put it?
—— Matan Grinberg · [26:09]
你可能不需要前沿的 token。你可能能够用前沿减一,或者你选择的开源权重 token 来应付。
指向原始笔记的链接
you may not need frontier tokens. You may be able to get by with frontier minus one or your open weight token of choice.
—— Sriram Krishnan · [06:41]
他们的那种 token 使用量每年增长十倍。
指向原始笔记的链接
their kind of token usage is growing like tenfold annually.
—— Ben Horowitz · [24:19]
大多数正在发生的 token 最大化并没有那么有价值,包括我自己过去几个月里做的很多。
指向原始笔记的链接
most token maxing a lot of token maxing that is happening isn’t that valuable, including a lot that I have done over time in the last few months.
—— 嘉宾 · [14:19]
至少我个人的决定是,我宁愿错在 token max 这一边,而不是不做的这边——反正你永远不会是对的。
指向原始笔记的链接
at least my personal decision is I’d rather be wrong on the token maxing side than the non right, and you’re never going to be right.
—— 嘉宾 · [15:44]
这里创造的价值太多了,以至于在那个场景下值得 token 最大化。
指向原始笔记的链接
There’s so much value created that it’s worth token maxing in that context.
—— 嘉宾 · [59:58]
如果你在 Opus 和 Haiku 上运行 terminal bench,Opus 的表现会好大约三倍,成本却是 Haiku 的十分之一,尽管 Haiku 每个 token 的价格要便宜得多。
指向原始笔记的链接
If you run terminal bench on Opus and Haiku, Opus will do about three times better at one-tenth the cost of Haiku, even though Haiku is significantly cheaper per token.
—— 嘉宾 · [15:26]
即使你是一个拥有十亿代币的人,即使你有 10 个终端没日没夜地运行 Fable,机会成本仍然存在,它就是一切。
指向原始笔记的链接
Even if you’re a token billionaire, even if you have 10 terminals running Fable night and day, then opportunity cost is still there, it’s everything.
—— Idan Gazit · [01:34]
当你说,嘿,清除暂存数据库,而它在你的笔记本电脑上发现了一个它可以使用的令牌,并且它认为它正在与暂存一起工作,但实际上它是生产环境,现在它刚刚删除了一切。
指向原始笔记的链接
when you say, hey, wipe the staging database and it finds a token on your laptop that it can use and it thinks it’s working with staging, but actually it’s production and now it just deleted everything.
—— Arjun Singh · [10:58]
这个产品本身,如果它真的存在,它会深埋在 Stripe 顶部的一个下拉菜单里,比如向下 11 层,因为它会被整个那种代币管理平台所吞并,对吧?
指向原始笔记的链接
This product itself, if it does exist, it’ll be deep in a dropdown menu on the top of Stripe, like 11 layers down, because it’ll be subsumed into the whole sort of token management platform, right?
—— Rory O’Driscoll · [26:20]
供应限制将在 2028 年左右缓解。我认为大型实验室按美元加权可能会获得未来 80% 的市场份额,因为这是我们历史上看到的大型在位者的情况。但我认为按 token 加权,60% 将是长尾和开源。
指向原始笔记的链接
Supply constraints will ease in 2028-ish. I think that the big labs will probably dollar weighted, get 80% of the market going forward because that’s historically what we’ve seen for large incumbents. But I think token weighted 60% will be long tail and open source.
—— Martin Casado · [16:39]
如果你看看一个每个 token 成本为双位数的美元的前沿模型,再看看一个每 token 成本为 10、11 美分的开源模型,虽然不是所有的 token 都生而平等,但这仍然是一个巨大的差异,足以带来大规模的采用。
指向原始笔记的链接
if you’re looking at a frontier model with double dollar digit cost per token, and you’re looking at an open source model that’s 10, 11 cents per token, while all tokens aren’t created equal, it’s still enough of a difference that there’s going to be a massive adoption.
—— Jerry Murdock · [14:06]
而且我们不再处于 token 最大化(token maxing)的时代了。我们试图实际上为我们拥有的每一个 token 获取价值,因为它变得越来越贵。
指向原始笔记的链接
And we’re not in the time of token maxing anymore. We’re trying to go to actually getting value for every token that we have, because it becomes more and more and more expensive.
—— Ofir Ehrlich · [27:18]
在法律领域,实际上你经常想要最大程度的智能,因为相比于应用到问题上的人类专业知识,token 支出或软件支出的比例真的很小。
指向原始笔记的链接
In law, you actually want the most amount of intelligence quite often because the fraction of token spend or software spend compared to human expertise applied to the problem is really tiny.
—— Max Junestrand · [33:16]
所以我们非常坚定地相信,下一个视点预测就是下一个 token 预测的等价物。
指向原始笔记的链接
So we do believe very strongly that next viewpoint prediction is the equivalent of next token prediction.
—— 嘉宾 · [43:15]
如果你不设干预阈值,智能体要么制造垃圾内容,要么烧出极高的 token 账单,因为它只会一直试下去,要么它们会伪造测试,就为了让测试通过。
指向原始笔记的链接
If you don’t do that, agents will either um create slop or run uh very high token bill because it’s just gonna keep trying, or they’re gonna fake the tests uh just to to to have the test passed.
—— Ran Arusi · [12:37]
据我们所知,我们是唯一做过稳健的数百万 token 预训练的人。
指向原始笔记的链接
As far as we know, we are the only people to do robust multi-million token pre-training.
—— Alexander Whedon · [11:26]
是的,根据我们在这里能想到的一切测量手段,你的 token 已经买不到它曾经能买到的东西了。
指向原始笔记的链接
Yes, your token is not buying you what it once did, according to everything we can think to measure here.
—— Chris Potts · [43:56]
人们把一整年的 token 预算挥霍一空——而且是在一个季度里就挥霍完了——或者他们在应付 token 排行榜之类的东西。
指向原始笔记的链接
People are blowing through their entire token budget for a year and they’re blowing through it in a quarter or they’re dealing with token leaderboards and such.
—— Maximillian Piras · [09:23]
我认为大型模型公司有一些大的护城河,因为它们最终提供 token。
指向原始笔记的链接
I think the large model companies have some big moat because they ultimately give the token.
—— Peter Steinberger · 来自原文
② 出现在这些集
20 集
- 《不会写代码的人如何成为全职 vibe coder》 — 作为概念
- 《Claude Code 负责人:写代码已被解决,下一步是什么》 — 作为概念(提及)
- 《DevOps 之父谈智能体开发:谁来管、怎么管、别踩什么坑》 — 作为概念
- 《「智能是数据的派生物」:Fireworks 创始人 Lin Kuo 的专用智能宣言》 — 作为概念
- 《Factory CEO Matan:早两年等于错,退款、路由器与软件工厂》 — 作为概念
- 《Heitor:用智能体重塑软件工程工作流的实操蓝图》 — 作为概念(提及)
- 《Datadog 4000 人AI赋能实战:删掉上下文反而更好》 — 作为概念(提及)
- 《AI 投资泡沫的崩盘剧本:为什么万亿美元建数据中心注定亏钱》 — 作为概念
- 《给智能体建一个“人力资源部”:TrustWise 创始人谈运行时治理》 — 作为概念
- 《Ed Zitron:生成式 AI 是一场万亿级骗局》 — 作为概念
- 《数据成了企业唯一的护城河:AI时代的数据基建怎么做》 — 作为概念
- 《NVIDIA 布局全栈、OpenAI 被迫上市与 AI 资本的”第五名效应”》 — 作为概念
- 《最便宜的模型反而是最便宜的:Factory CTO 谈 AI 定价陷阱》 — 作为概念(提及)
- 《Token 都烧在哪了:Cursor 工程师教你把 AI 编程成本打下来》 — 作为概念
- 《Ollama CEO:开源模型正吃掉企业 80-90% 的 token》 — 作为概念
- 《开源权重不是威胁:Box CEO 聊 AI 的经济账》 — 作为概念(提及)
- 《邮箱里的 AI 助手:Plaid 前 CTO 谈如何在巨头围剿下赢》 — 作为概念
- 《斯坦福语言学家的代币经济学:你的 token 贬值了》 — 作为概念
- 《鼠标力:为智能体时代找回「马力」这把尺子》 — 作为概念
- 《不到10人管7个SaaS:让智能体替你做营销的实操系统》 — 作为概念(提及)
③ 关联
点进去有真内容 —— 本页主要出口
智能体 · Anthropic · Cursor · OpenAI · 推理 · NVIDIA · Codex · 后训练 · OpenRouter · harness
