训练数据 (data)
集里怎么说它
- 《Scale AI 创始人 Alexandr Wang:AI 时代,最稀缺的不是智能而是愿景》(04:21起):本集说它是Scale最初的核心生意。嘉宾在MIT发现算力和代码只需联网按按钮就能拿到,唯独它没有获取途径。十年前VC觉得这门生意不性感,现在却大谈它是AI最大的商业机会。
① 提到它的金句
85 条
而且我们没有使用任何来自大型提供商的数据。我们所有的数据都是自己从零开始收集的,通过训练付费的教师。
指向原始笔记的链接
And we didn’t use any data from big providers. We collected all of it ourselves from scratch by training paid teachers.
—— Mustafa Suleyman · [15:05]
我们真正在构建的是一个数据平台,能够独特地把每一个产品使用单元连接到收入。
指向原始笔记的链接
What we’re really building is a data platform that can uniquely connect every unit of product usage to revenue.
—— Alvaro Morales · [29:15]
我们在行业内看到,通过算法数据和效率改进的结合,给定智能水平的成本降低了 10 倍。
指向原始笔记的链接
We’ve seen in the industry a 10X decrease in cost for a given amount of intelligence through a combination of algorithmic data and efficiency improvements.
—— Benjamin Mann · [58:33]
人类数据中唯一的护城河是对受众的访问。
指向原始笔记的链接
The only moat in human data is access to an audience.
—— Garrett Lord · [00:30]
如果你能够生产大量高质量的数据,你很可能会能够卖出你生产的任何东西。
指向原始笔记的链接
If you can produce high quality volumes of data, you most likely will be able to sell whatever you produce.
—— Garrett Lord · [48:11]
我认为合成数据在可验证的领域有一席之地,但我们从公司那里听到的说法始终是,合成数据不会占据主导地位。
指向原始笔记的链接
I think synthetic data has a role to play and in verifiable domains, but what we consistently hear from companies it’s like synthetic data is not going to dominate.
—— Garrett Lord · [62:00]
有一个短语,我和我的联合创始人会在我们之间非常早地讨论并且我们与我们合作的很多公司分享了它,那就是说你真正想要的是你想用数据诊断并用设计治疗。
指向原始笔记的链接
There’s one phrase that my co-founder and I would always discuss with us amongst ourselves very early on and which we shared with like a lot of the companies that we work with, which is what you really want is you want to diagnose with data and treat with design.
—— Julie Zhuo · [32:46]
在某个时刻,我们在互联网数据上实际上已经有点达到极限了。
指向原始笔记的链接
At some point, we are actually have kind of maxed out on the internet data.
—— Chip Huyen · [14:58]
在这个中间过程中,既然每个人都有非常相似的预训练数据,那后训练就是如今他们做出大差异的地方。
指向原始笔记的链接
where like post-trading, but middle course of this is more of everyone have very similar pre-training data, is that post-training is where they make a big difference nowadays.
—— Chip Huyen · [15:10]
‘不’是对‘是’的最好回答,因为‘不’是你可以使用的数据。
指向原始笔记的链接
No is the best answer to yes because no is data that you can use.
—— Lenny · [66:32]
在 AI 时代,单点解决方案没有足够的数据来发挥作用。
指向原始笔记的链接
Point solutions don’t have enough data in the age of AI to be useful.
—— Matt MacInnis · [78:36]
因为如果你一路尝试做数据驱动的决策,你要么不是在做差异化的产品,因为你在从其他东西获取数据,要么你只是在得到废话数据,对吧。
指向原始笔记的链接
Because if you try to do data-driven decisions all the way along, you’re either not doing a differentiated product because you’re taking data from another thing, or you’re just getting just bullshit data, right.
—— Tony Fadell · [09:23]
CDC 是最无聊的之一,但也是驱动现代社会的最基础的操作之一。但它是如此脆弱,以至于有一个冷笑话说它应该被称为持续数据损坏
指向原始笔记的链接
CDC is one of the most boring, but one of the most fundamental operations powering modern society. But it’s so brittle that a weak joke that it should be called continuous data corruption
—— Reynold Xin · [30:30]
让数据到位,然后在上面附加一些智能体。魔法就会出来。但没有正确的数据,你无法真正做到这一点。
指向原始笔记的链接
just get the data to be there and then slap some agent on top. Magic will come out. But without the right data, you can’t really do that.
—— Reynold Xin · [66:37]
把一个智能体叠加在一组数据孤岛化的任意工具之上,远比不上叠加在一个单一平台和唯一事实来源上——后者既拥有你所有的数据,也在同一个工具内执行你所有的操作。
指向原始笔记的链接
It is far more difficult to overlay an agent on top of this arbitrary set of tools with data silos than it is a single platform and source of truth that both has all of your data, but also takes all of your actions inside of the same tool.
—— Sam Blond · [37:47]
但对于我们担心的主要风险类别,比如提示词注入、数据窃取,风险远低于人类审核员的平均水平。
指向原始笔记的链接
But for the main categories of risks that we’re concerned about, like prompt injection, data exfiltration, the risks are far lower than the average human reviewer.
—— Cat Wu · [32:33]
所以我认为把数据中心可视化最好的方式,就是巨大的工厂,把电子变成 token。
指向原始笔记的链接
So I think the best way to visualize data centers is giant factories that are turning electrons into tokens.
—— Sachin Katti · [05:52]
所以无论我们在哪里建数据中心,我们都会做出硬性承诺,即我们不会从电网夺走电力。事实上,我们是在投资电网以产生新的电力,这样我们才能把电力用于数据中心。
指向原始笔记的链接
So whenever we build a data center anywhere, we make it a hard commitment that we are not taking power away from the grid. In fact, we are investing in the grid to generate new power so that we can consume it for data centers.
—— Sachin Katti · [08:44]
但我认为如果说设计和深刻的设计专业知识及思维因此而被挤出,那将是一个错误,仅仅是因为我们可以更快地编写代码,我们可以更快地进行数据分析。
指向原始笔记的链接
But I think it would be a mistake to say design and deep design expertise and thinking gets squeezed out, just because we can write code faster, we can do data analysis faster.
—— Elizabeth Stone · [21:32]
如果你认为智能是数据的派生物,那么大部分数据其实并没有被用于训练通用智能模型。
指向原始笔记的链接
If you think intelligence is a derivative of data, then majority of the data is actually not used for training a general intelligence model.
—— Lin Qiao · [07:32]
而且通常,爬坡的最终结果是:用你的数据解决你的独特问题,你会比通用模型更好。
指向原始笔记的链接
And often time, the end result of heel climbing is to solve your unique problem with your data, you are better than a general purpose model.
—— Lin Qiao · [15:21]
所以这里出现了一种类比:数据之于模型,正如模型之于 harness。
指向原始笔记的链接
So there’s a sort of analog that emerges where it’s what data is to a model, models are to a harness.
—— Matan Grinberg · [20:06]
那些数据在其他任何地方都不存在。Google Maps 里没有。它只存在于 DoorDash。
指向原始笔记的链接
That data doesn’t exist anywhere else. It doesn’t exist in Google Maps. It only exists at DoorDash.
—— 嘉宾 · [32:37]
彭博估计有超过 5000 亿美元的未偿 AI 数据中心债务,其中至少有 2000 亿美元由私人信贷持有,大约占未偿私人信贷贷款的 8%。
指向原始笔记的链接
Bloomberg estimates there’s over 200 billion of it held by private credit, making up roughly 8% of outstanding private credit loans.
—— Alex · [32:31]
为了极其明确一点,AI 数据中心算力收入的绝大部分取决于两家不盈利、不可持续的 AI 公司是否有能力继续每年筹集数百亿或数千亿美元。
指向原始笔记的链接
to be abundantly clear, the vast majority of AI data center compute revenue is contingent on the continued ability of two unprofitable, unsustainable AI companies to raise tens or hundreds of billions of dollars a year.
—— Alex · [35:38]
它们都是估值极高的公司,都在押注一个更大的关于机器人经济的承诺,也就是特斯拉和 Optimus 机器人无处不在,或者是太空中的数据中心。
指向原始笔记的链接
They’re both incredibly overvalued companies that are betting on a much larger promise of a robotic economy that Tesla and Optimus robots everywhere or data centers in space.
—— Ranjan Roy · [62:23]
用数据做诊断,用设计做解决。
指向原始笔记的链接
you diagnose with data and solve with design.
—— Jon Noronha · [23:19]
代码智能体非常善变、不可预测,它们经常犯错,然后当你真的很倒霉时,它们会发疯并试图销毁你的数据以及其余的一切。
指向原始笔记的链接
Coding agents are really fickle, unpredictable, they often make mistakes, and then when you’re really unlucky, they’ll go crazy and try and nuke your data and all the rest of it.
—— 嘉宾 · [17:00]
所以当它开始像是一 PB 级的数据,你有一千个微服务,那就是它变成的时候,与其说是一个 AI 问题,它变成了一个数据问题。
指向原始笔记的链接
So where it starts getting like a petabyte of data, you have a thousand microservices, that’s when it becomes, as much as it’s an AI problem, it becomes a data problem.
—— Anish · [27:42]
我不认为开源模型的进步会蚕食我们的核心业务,因为数据在模型性能的最前沿才是最有价值的。
指向原始笔记的链接
I wouldn’t say that open source model improvements cannibalize our core business because data is most valuable on the frontier of model performance.
—— Osvald Nitski · [04:21]
我们正在耗尽数据来训练大模型,但我们没有在使用旧模型的权重。所以为什么不使用权重,所有的知识,人们投入的所有算力,对吧?
指向原始笔记的链接
We’re running out of data to train the large models, but we’re not using the weights of older models. So why not using the weights, all the knowledge, all the compute that people invested, right?
—— Damian Borth · [30:39]
仿真扮演了一个非常重要的角色,这是现实世界的数据所没有的,那就是反事实推理,就是你在推演那些尚未发生或不可能发生的事件,或者你在现实世界中没有足够的数据让它发生。
指向原始笔记的链接
There’s a very important role simulation plays that real world data doesn’t play, which is counterfactual reasoning, is that you play out events that hasn’t happened or cannot happen, or you don’t have enough data to make it happen in real world.
—— Fei-Fei Li · [20:06]
事实上,Waymo 比重视现实世界数据更重视仿真。
指向原始笔记的链接
And actually, Waymo is more simulation-heavy than just real-world data-heavy.
—— Fei-Fei Li · [21:11]
这些模型在如此多的数据上训练,它们如此巨大,然而它们实际上真的不知道如何做其中的任何一项工作。
指向原始笔记的链接
The models are trained on so much data and they’re so large and yet they actually don’t really know how to do any of this work.
—— Frederick Rankin · [46:31]
我从根本上相信,对于未来的 AI 公司来说,你必须拥有有趣的数据策略。
指向原始笔记的链接
I fundamentally believe that for AI companies in the future, you have to have interesting data strategy.
—— Joon Sung Park · [57:50]
我的意思是,如果你真正理解数据,你应该能够非常好地压缩它。
指向原始笔记的链接
I mean, if you truly understand the data, you should be able to compress it really well.
—— Jeff Dean · [15:54]
作为一个产品经理,我在过去一年里花在与数据科学家交谈的时间比我职业生涯中任何时候都少,即使我可能花在数据和实际理解产品如何工作上的时间比我在职业生涯中任何时候都多 10 倍。
指向原始笔记的链接
as a product manager, I’ve spent less time in the last year talking to a data scientist than I ever have in my career, even though I’ve probably spent 10 times more time in data and understanding actually how the product’s working than I ever have in my career.
—— Tom Verrilli · [36:54]
我认为这就好比如果你经营一家汉堡店,一个拥有数百兆瓦的相当大的数据中心消耗的水量大约和一家麦当劳一样多。
指向原始笔记的链接
I think the analogy is like if you ran a burger shop, a fairly large data center with hundreds of megawatts would consume about the same amount of water as a McDonald’s.
—— 嘉宾 · [33:29]
我们实际上有白皮书,我们的数据在做模型的强化学习方面与真实数据一样好。
指向原始笔记的链接
And we actually have white papers where our data does as well at doing reinforcement learning on models as real data.
—— Ian · [03:20]
你要么需要做一件叫做 Safe Harbor 的事情,这是一个 HIPAA 流程,它基本上会破坏数据,因为它并不是真的那么有价值,或者你可以做这件事叫做专家判定。
指向原始笔记的链接
You either need to do something called Safe Harbor, which is a HIPAA process that essentially nukes the data, like it’s not really that valuable, or you can do this thing called expert determination.
—— Ian · [08:12]
结果证明,如果你为人类构建了很棒的交接,你同时也为智能体构建了很棒的交接,因为智能体理解 HTML 和 CSS,而且这些在它们的训练数据中。
指向原始笔记的链接
It turns out if you build a great handoff for humans, you’ve also built a great handoff for agents because agents understand HTML and CSS and it’s in their training data.
—— Stephen Haney · [03:35]
想象一下过去 20 年所有政府在互联网基础设施上花了多少钱,来路由那些完全不必要的数据包,而互联网本是被构建为点对点和端到端的。
指向原始笔记的链接
Imagine how much all the governments spent in internet infrastructure in the last 20 years to route packets data that is completely unnecessary when internet was built to be point-to-point and peer-to-peer.
—— Paolo Ardoino · [10:11]
一个如此智能、无所不知、可以扩展并分发到宇宙四角的 AI,不可能只待在地球上的数据中心里,不可能属于一个人。
指向原始笔记的链接
An AI that is so intelligent, that knows everything, that can scale and be distributed to the four corners of the universe cannot sit in a data center on Earth, cannot belong to one person.
—— Paolo Ardoino · [16:05]
所以如果你不理解你的 AI 是如何工作的,并不是你变得更智能,而是其他人利用你的数据变得更智能。
指向原始笔记的链接
So if you don’t understand how your AI works, it’s not you becoming more intelligent, it’s someone else becoming more intelligent with your data.
—— Paolo Ardoino · [33:15]
截至 2026 年第二季度,现在数据中心资金中超过 50% 是外部融资,这显然是资产负债表外和自有现金流之外的术语。
指向原始笔记的链接
As of the second quarter of 2026, This is now more than 50% of the funding for data centers is external financing, which is obviously the term of art for off-balance sheet and out of your own cash flows.
—— Paul Kedrosky · [03:58]
当我们第一次看到机器人这样做时,我们惊呆了,因为这个任务没有任何训练数据。
指向原始笔记的链接
The first time we saw the robot do this, we were like floored because there was no training data for this task.
—— Chelsea Finn · [33:32]
而有了元数据提示,当你添加低质量数据时性能实际上增加了,这表明当你包含这种提示时,它实际上能够从即使是低质量数据中获得更多收益。
指向原始笔记的链接
Whereas with the metadata prompting, the performance actually increases when you add that low quality data, suggesting that it’s actually able to get a lot more juice out of even low quality data when you include this kind of prompting.
—— Chelsea Finn · [36:35]
但我认为 AI 让我们能把知识编码到另一层。它不必完全编码在组织里;它可以编码在智能体能访问的数据中,这真正让决策民主化了。
指向原始笔记的链接
But I think AI gives us the ability to encode that knowledge at another layer. It doesn’t have to be encoded in the org entirely; it can be encoded in data that agents have access to, and that really democratizes decision-making.
—— Willem Avé · [10:39]
所以 Meta 保证如果它不坚持整整二十年,它会让债券持有人完整收回,但每个人都在押注如果 AI 需求爆发并且如果这些数据中心开始疯狂印钞,然后它们能够支付债券持有人而且还有富余并产生额外的现金,那么 Meta 就不需要承担任何责任。
指向原始笔记的链接
Meta is guaranteeing to make bondholders whole if it doesn’t stay for the entire two decades, but everyone is betting if AI demand explodes and if these data centers start printing cash and then they’re able to pay bondholders and then some and generate additional cash, then Meta’s not on the hook for anything.
—— Ranjan Roy · [11:30]
他们拥有所有的数据。他们拥有所有的智能。而且他们的模型正在被 OpenAI 和 Anthropic 击溃。
指向原始笔记的链接
They have all the data. They have all the intelligence. And like their models are getting trounced by OpenAI and by Anthropic.
—— Martin Casado · [52:31]
所以在人类历史上,我们从未创造过拥有那么多浮点运算量和那么多数据的单一数字制品。
指向原始笔记的链接
So like in the history of humanity, we’ve never created a single digital artifact that had that many flops and that much data in it.
—— Martin Casado · [56:20]
你得到的每一份数据都在某种程度上被人的解读、人的观点、偏见,以及他们自己的抱负等等所遮蔽。
指向原始笔记的链接
Every bit of data you got was also sort of clouded with human interpretation, human opinions, biases, sort of like their own ambitions.
—— Cliff Obrecht · [35:29]
我们在 Parallel 的观点是,人类点击数据是一个 bug,使用搜索进行工作的智能体应该依赖智能体反馈,而不是人类反馈。
指向原始笔记的链接
Our view at Parallel is that human click data is a bug and agent doing work with search should rely on agent feedback, not human feedback.
—— Parag · [00:00]
我们在 Parallel 的观点是,人类点击数据是一个缺陷,进行搜索工作的智能体应该依赖智能体反馈,而不是人类反馈。
指向原始笔记的链接
Our view at Parallel is that human click data is a bug and agent doing work with search should rely on agent feedback, not human feedback.
—— Parag · [00:00]
仍然有一些相当稀缺的事情你可以在内部解决。数据科学仍然稀缺。真正好的品味仍然稀缺。综合分析仍然稀缺。
指向原始笔记的链接
There are still some things that are quite scarce that you can work on internally. Data science is still scarce. Really good taste is still scarce. Synthesizing analytics is still scarce.
—— Elaina O’Mahoney · [12:02]
我们认为多智能体系统是一个分布式数据问题。
指向原始笔记的链接
We think of multi-agent systems as a distributed data problem.
—— James · [01:15]
所以我们要看到的关键点是:智能体不是计算,它们是数据。
指向原始笔记的链接
And so the key thing about agents that we see is that agents are not compute, they are data.
—— James · [05:52]
你知道,就在两天前,你看到 Google 从破产的 Spirit Airlines 那里买了一些东西。他们没有买飞机。他们买了数据。
指向原始笔记的链接
You know, just two days ago, you saw Google buy something from the bankrupt Spirit Airlines. They didn’t buy airplanes. They bought the data.
—— Ofir Ehrlich · [03:47]
所以它在一个组织内部创建了一整套行动者,不受组织规则的约束,也不一定在组织内部运行,但处理属于组织的敏感数据。
指向原始笔记的链接
So it creates a complete set of actors inside an organization, not bound by the rules of the organization, and not necessarily running within the premises of the organization, but handling sensitive data, which is the property of the organization.
—— Ofir Ehrlich · [21:27]
也许有一些是 AI 生成的数据。并且在 AI 生成的数据和最终提交的数据之间存在真正的差异。这是几乎其他人都没有的丰富信息。
指向原始笔记的链接
Maybe there’s some data that the AI generated. And there’s a real diff between the data that the AI generated and what was ultimately submitted. That’s rich information that almost no one else has.
—— Varun Shenoy · [12:02]
你需要的数千亿自由现金流,用来偿还你为适应这些数据中心扩建而承担的债务,为了进行下一次大型训练运行,这对他们来说完全是生死攸关的。
指向原始笔记的链接
The hundreds of billions in free cash flow that you need in order to pay back the debt that you’re taking on in order to accommodate these data center build outs in order to get the next big training run, it’s totally existential for them.
—— Eno Reyes · [52:16]
我会担心任何这样的记忆架构:它不依赖成本更低的嵌入器和重排序器,而是依赖 LLM 的多遍处理来帮你分类并缩小数据语料库。
指向原始笔记的链接
I would worry about any memory architecture that instead of relying on lower costs embedders and re-rankers is relying on multiple passes of the LLM to help you categorize and shrink the corpus of data
—— 嘉宾 · [66:35]
像糟糕的数据质量和糟糕的安全态势这类问题不会被 AI 解决,它们会被 AI 放大。
指向原始笔记的链接
Things like bad data quality and bad security posture don’t get solved by AI, they get amplified by AI.
—— 嘉宾 · [68:54]
当有人从一个智能体渠道过来时,智能体使用我们的目录数据时,转化率是它仅使用抓取数据时的 2 倍。
指向原始笔记的链接
And so when someone comes from an agentic channel, the conversion rate that when an agent is using our catalog data is 2x compared to if it’s just using scraped data.
—— Jess Hertz · [11:25]
我认为那是胡扯,因为无论是能源、液冷,还是所有能让数据中心变得更好的东西,这些都是可以在这里创造、在这里完成的工作岗位。
指向原始笔记的链接
I call BS on that because if you think about whether it’s around energy, liquid cooling, all of the things that make the data center better, those are all jobs that can be created and done here.
—— Rene Haas · [28:38]
这就是这一切的讽刺之处:今年一万亿美元的资本支出,以及未来五年即将到来的十万亿美元,是为了把数据移动几厘米和几毫米。这就是 AI。
指向原始笔记的链接
So this is the irony of it all, that the trillion dollars of CapEx this year and the 10 trillion over the next five years that are coming is to move data centimeters and millimeters. That’s AI.
—— Tony Kim · [07:08]
总有一天会变成芯片,但目前是数据。
指向原始笔记的链接
One day it’ll be chips, but for now it’s data.
—— 嘉宾 · [31:00]
唯一比没有数据更糟糕的就是非常糟糕的数据,所以我们不想要那样。
指向原始笔记的链接
The only thing worse than no data is really bad data, so we don’t want that.
—— Tyler Folkman · [21:59]
所以总的来说,存储数据是一个非常巨大的问题空间,我认为它永远不会下沉到模型层。
指向原始笔记的链接
So generally storing data is like a huge problem space that I don’t think will ever make its way down into the model layer.
—— Jeffrey Morgan · [17:21]
你越是需要多个模型来完成一项任务或一组任务,价值就越是积聚到那个能够理解任务、获取数据并处理工作流的层上
指向原始笔记的链接
the more that you need multiple models to do a task or a set of tasks, the more value accrues to the layer that can understand the task and get access to the data and handle the workflow
—— Aaron Levie · [26:22]
如果你让一群博士组成的数据中心去 FedEx 或达美乐披萨工作,他们会指数级地主导供应链和披萨吗?
指向原始笔记的链接
If you had a data center of PhDs working at FedEx or Domino’s Pizza, are they going to be exponentially dominating supply chain and pizzas?
—— Anish Acharya · [06:38]
我从根本上认为,在给定固定数据量的情况下,从数据中学习的最佳方式可能是拥有大量分布式的模型。
指向原始笔记的链接
I think fundamentally, it might be that given a fixed amount of data, the best way to learn from it is to have a lot of distributed models.
—— Lucas Kaiser · [04:43]
所以是的,Transformer 在你喂给它整个互联网时很强大,但我们知道存在一个更优秀的算法,它能从小得多的数据中学习,并在所学的东西上表现得令人惊叹。
指向原始笔记的链接
So yes, transformers are great when you feed them all of the internet, but we know there is an algorithm that’s even better and it can learn from much smaller data and be amazing at the things it’s learning.
—— Lucas Kaiser · [05:03]
而我说,把你 DevEx 团队的规模翻一倍。把你数据团队的规模翻一倍。投入平台建设。对人类有好处,对智能体也有好处,而这就是能让你跑起来的东西。
指向原始笔记的链接
And I say, double the size of your DevEx team. Double the size of your data team. Like, work on platform investments. Good for humans. Good for agents. And that’s what will let you run.
—— Claire Vaux · [18:04]
我认为五年内,你会信任你的智能体在不经你干预的情况下决定与别人分享什么数据。
指向原始笔记的链接
I think you’ll trust your agent to decide what data to share with other people without you intervening in five years.
—— Harry Stebbings · [13:03]
我们在编码方面有意走得慢一些,因为我们已经意识到编码在多大程度上是一场预算游戏——你有多少数据预算?
指向原始笔记的链接
We are intentionally moving a little bit more slowly on the coding side because we’ve kind of realized like how much coding is like a budget game, like how much data budget do you have?
—— Alexander Whedon · [27:49]
任何你展示给用户的东西,也需要作为数据提供给模型。
指向原始笔记的链接
anything that you show to the user also needs to be provided as data to the model.
—— Dustin Mihalik · [04:13]
你要在关注 UI 之前先关注数据
指向原始笔记的链接
you want to focus on the data before you focus on the UI
—— Dustin Mihalik · [13:39]
所有这些建立在对数据和学习直觉之上的分析工作,造就了我们如今拥有的模型——相比 2017 年的 Transformer,它就像一艘忒修斯之船。
指向原始笔记的链接
All of this analysis work built on intuitions about data and learning led to the model that we have now, which is like a ship of Theseus compared to the 2017 Transformer.
—— Chris Potts · [15:29]
如果研究实验室不拥有最好的产品,那才真是疯了,因为他们拿到了所有关于人们如何使用他们产品的数据。
指向原始笔记的链接
It would be really wild if the research labs didn’t have the best product, because they’re getting all of this data on how folks are using their products.
—— Demetrios Brinkmann · [38:34]
这其实与模型的强大程度无关。Databricks 的关键,甚至很多公司想在 AI 上取得成功的关键,全都在于你数据的上下文。
指向原始笔记的链接
It’s not really about the strength of the model. The key for Databricks and even the key for a lot of these companies to be successful at AI is all about the context of your data.
—— Ron Gabrisko · [19:57]
如果你让一数据中心博士生去 FedEx 或达美乐披萨工作,他们能在供应链和披萨上呈指数级称霸吗?
指向原始笔记的链接
If you had a data center of PhDs working at FedEx or Domino’s Pizza, are they going to be exponentially dominating supply chain and pizzas?
—— Anish Acharya · [07:18]
这是一个演示文稿产品,但它永远做不出惊艳的演示文稿。它缺上下文,缺数据结构。
指向原始笔记的链接
This is a presentations product and it’s never going to be capable of making amazing presentations. It was just missing the context. It was missing the data structure.
—— Keith Peiris · [04:14]
AI 以软件的速度前进,而数据中心以房地产的速度前进。
指向原始笔记的链接
AI is moving at the speed of software and data centers are moving at the speed of real estate.
—— Andrew Feldman · [19:51]
美国所有数据中心的用水量比加州杏仁种植者还少。
指向原始笔记的链接
All the data centers in the U.S. use less than California almond growers.
—— Andrew Feldman · [46:27]
② 出现在这些集
1 集
③ 关联
点进去有真内容 —— 本页主要出口
Alexandr Wang · Scale · Meta · MuseSpark · 开源模型 · 智能体 · 多智能体设置 · 前沿AI实验室 · 主观能动性 · Spark API
