斌叔OKmath
26-09-09 08:28 微博认证:橙旭园CEO 教育博主

据说前沿预训练是只有大实验室才能玩的游戏。我们还没有 10 万个芯片,所以只有一条路:算法效率。我们新的配方使用 50 倍更少的计算量就匹配了 DeepSeek V4 Pro 的预训练——这大约是 GPT3 所用 FLOPs 的一半,或者在 GB200 上大约 50 万美元。

We scaled this recipe 10x (~$4M) and exceeded all publicly available base models. Pretraining a model this capable would cost >$100M under DeepSeek V4 Pro’s recipe. Of course, we won't stop scaling there. To sanity check post-RL performance, we ran a short math RL run from base.

Our pretraining recipe is the multiplicative result of tens of changes across model architecture, optimizer, training objective, and data. And we learned to appreciate that fixing minor bugs is a compute multiplier too. We talk more about our research process in the blog.

Magic’s goal is to build the best model for SWE and autonomous AI R&D. To balance data trade-offs, we look at “knowledge evals” in addition to our standard out-of-sample suite. We make new knowledge evals for each model generation to avoid overfitting over time.

What’s next?
1) Scale RL with long context to teach agents test-time learning.
2) RL against the model’s own latent knowledge of its intent for stronger theoretical alignment properties.
3) Further improvements to pretraining.

And then, of course, release the thing!

We are likely the smallest team in the world training trillion parameter models. The impact a single person with strong judgement can have has never been higher. If you want to help build aligned superintelligence, consider joining:

http://t.cn/AX0melSN

发布于 北京