back to 2026-08-14
ᕼᑎ:49289844414 pts173 commentsAIworth reading

Accelerating GPT-5.6 Sol Ultrafast

Claude brief

HN 热门故事「Accelerating GPT-5.6 Sol Ultrafast」进入今日前列,值得先打开原文和讨论串判断它真正有价值的部分。

模型分析没有产出可用结构化结果;页面保留了 HN 热度、原文入口和讨论信号,避免用空泛总结替代一手材料。

它在 HN 上获得约 414 分和 173 条评论,说明这个话题至少触发了社区讨论;真正的判断仍要回到原文证据和评论区的分歧点。

这是一条降级分析:它不冒充完整解读,只把可验证的元数据、原始链接和 HN 讨论保留下来,方便稍后重新生成或人工阅读。

评论区已经提供了一些读者反应,但这里还没有形成完整综合。

它进入 HN 前列本身就是一个社区信号,但这还不是结论;更可靠的判断来自原文细节和评论区反例。

deep insight

这条记录目前缺少模型生成的深层解读。更好的阅读方式是先问:它的热度来自真正的新信息、可迁移的方法,还是只来自标题与时机。

可以先读原文第一屏和 HN 最高赞评论,再决定是否值得重新生成完整分析。

top comments

Right now, the "economic model" of AI is "who has the best model", or really weights.Honestly that'll go away eventually, just like operating systems eventually became free.Instead, it's going to come down to selling inference hardware. We'll likely see the "apple" model where a custom OS runs on their hardware, but we'll probably also see more things like Cerebras become commodity hardware instead of kilowatt-class datacenter only hardware.
I've been waiting so long for something amazing to come out of the OpenAI and Cerebras collaboration.> In our evaluations, GPT-5.6 Sol on Ultrafast mode answered all 2,500 HLE questions in 11 hours and 11 minutes. Claude Fable 5 needed 78 hours and 27 minutes, more than three days of continuous compute, to arrive at the same conclusions. In other words, Ultrafast worked through the frontier of human knowledge in a single working day, achieving comparable accuracy nearly 7× faster.This is actually insane.Hopefully the release ultrafast of Terra and Luna too. reply: I discovered yesterday that the “amazing thing that comes out of OpenAI” is Sol, due to its token efficiency.Dollar for tokens, Sol and Fable are the same price.However, Sol uses (literally:...
People underestimate the importance of speed on quality of thought, because people underestimate just how much quality is a result of simple iteration.When an LLM thinks, it typically just makes one pass. It outputs tokens from top to bottom, beginning to end, and then it's done. But when people think, especially strong thinkers, we typically iterate and revise our thoughts on the fly. We do numerous passes. We stop and restart, we reconsider, we review, we reevaluate. Sometimes we do this so quickly and automatically that we don't even realize we're doing it. I think a lot of what separates a highly intelligent or effective person from others has less to do with the quality of their first pass and more to do with just how many additional passes they're able to do in the same amount of time, and of course what kind of criteria they're habituated to consider during their review passes.Introspecting about this is difficult, but experimenting with LLMs is easy. First, simply ask an LLM to do something complex. For example, to come up with a new business idea, or to plan the next month of your life, etc.... reply:...
Unless I have read over it, besides the animation in the intelligence vs speed graph which only mentions internal data and not whether they truly reran the AA suite, there is no actually solid statement on the important aspect of performance.Neither the Cerebras or OpenAI post [0] outright state that this performs exactly the same as regular 5.6 Sol. I feel if this was 1:1 just Sol but much faster, they'd (rightfully) scream that off the rooftops. A line such as "this is the same performance, just faster, with no downsides" would go a long way in clarity and communication. Along with no pricing information, I'll hold out on further information.[0] https://openai.com/index/previewing-ultrafast/ reply: "delivering up to 750 output tokens per second and without any quality compromise" seems pretty definitive.
The corresponding OpenAI post https://openai.com/index/previewing-ultrafast/There is no pricing info, which could mean it's "if you have to ask..." territory or they are simply gauging interest before deciding reply: They're nearly certainly going to use it internally to speed up research that is serially bottlenecked. I would bet this is why they're interested in the Cerebras partnership more than everything else
> Compared with output speeds reported by Artificial Analysis GPT-5.6 Sol on Ultrafast mode runs 11x faster than Fable 5, and 5x faster than Opus 4.8 on Fast mode.Awesome work. I'm personally very excited for faster models/inference.I think speed is underrated to some degree in the current conversation. For a while, I was using Cursor's Composer quite a lot, even over frontier models, just because of how darn fast it was. reply: What do you need speed for? That's a genuine question, I feel like the limiting factor already is my creativity, attention span and budget. And I'm not even yet optimizing cost by batching things like review to slow local models over night, or schedule tasks to take full advantage of my subscriptions.
With this level of intelligence offered at this level of speed, new real-time applications become possible, such as providing expert advice during a phone call or court hearing. Current SOTA models are too slow in many cases to provide the kind of insights that we expect to receive from an intelligent human colleague, such as a sales coach or lawyer handing us a note or writing a Slack message during a difficult call. For these real-time applications, even a 10x increase in per-token cost would often be tolerable.
This does look pretty incredible, but don't forget that incredible token thoroughput can only necessarily solve certain bottlenecks. If your e2e tests take an hour, they'll still take an hour after Ultracode. If the agent runs a 10 minute typecheck after a change, that will still take 10 minutes. grep over a massive codebase is still just as slow, etc. I say this not to take away from this accomplishment but just to ensure everyone here keeps a clear head about what it means - 14x faster tokens does not mean it completes every task 14x faster.I suspect Humanity's Last Exam is without tool-calls, making it kind of the perfect benchmark to highlight how fast Ultrafast is, but not really the same as the everyday work you or I do.
Now I can blast through my weekly 20x pro codex credit in like an hour, great!The amount of usage you receive on Codex these days is dismal compared to what it was a few months ago, FYI.And they charge more for going faster.As a Codex customer, I am not impressed with their shenanigans over the past few months and I have resolved to master the art of Pi Coding Harness creation and loving it.Thanks for all the fish, Sam!
The omission of Mimo v2.5-Pro Ultraspeed, released in June, which can achieve 1000tok/s is an interesting flaw in the comparison graphs.It is a bit outdated (scores ± 40% lower), but smart enough for a lot of coding tasks, and can cost under 1/10th of Sol.https://mimo.mi.com/models/en-US/mimo-v2.5-pro-ultraspeed
> GPT-5.6 Sol on Ultrafast mode, delivering up to 750 output tokens per secondhttps://taalas.com/products/> delivering 17k tokens per second per user on Llama 3.1 8B model.Obviously this is a much smaller model, but I really can't wait for ASICs to take over the LLM space.Imagine running a model like Sol/Fable (even half the size with 60-70% of it's intelligence) on your own ASIC hardware.
Good news for Intel and AMD.Rught now on large scale codebases the bottleneck is both claude/codex inference, as well as time it takes to run tens of thousands of tests. We put those workloads on dedicated epyc 9005 build machines - but it still takes minutes per run. Those who can afford the fast tokens will be in the market for faster CPU that money can buy today.
This is really cool. Someone here commented about similarity between this and hardware advancements for AV encode/decode.I think it's only a matter of time before miniaturization can have a thumbnail sized user-replaceable accessory that contains the LLM built onto the hardware. I admit I don't know how any of that works, but would be amazing to experience. Fully local, fully offline, ultra fast local inference better than any personal computing product.
Wonder how many X usage this would consume when it becomes available for everyone. Fast mode already consumes 1.5x usage for 2.5x speed. Hopefully, this does not mean 8x usage for 14x speed, but rather something more reasonable such as 4x usage.
This is an amazing result. Can't wait until they release this to the general public, and I hope it's only a matter of time before other models are accelerated. I long for the day that regular consumers can run such models locally on specialized hardware.
Whoa. This looks both powerful and expensive.My prediction is that, this time next year, top developers outside ai labs will be spending 50k USD+ on inference.Within labs, I've heard spend is already far beyond this per developer.
I'd just like to point out that the largest model Cerebras has ever served is Kimi K2.6 which is 1T parameters, so that either means that theyve had a breakthrough on the hardware engineering side of things, or GPT-5.6 Sol is likely a lot smaller than people think.If it truly is only ~1-2T parameters, then this kinda kills 2 narratives for me.1. all the handwringing about open source catching up via Kimi K3 (3T params) is complete nonsense. All that matters imo for determining which labs are leading is intelligence per parameter. Anyone with a enough compute can train a giant model, but being able to squeeze capabilities into smaller models gives you a massive inference and training edge.2. Inference margins are clearly insane, and this explains why OpenAI was able to lower the price of Luna by 80%. Id guess that thing is probably 120b params based on the TPS they are serving it at.
I think token output speed is going to be one of the biggest fundamental shifts for AI this year. In my experience, models figure out problems after enough turns (or in agent swarms if its a lower tier model). Compressing that time horizon could take days of agentic coding into minutes. How ever will my monkey brain keep up?
Would gladly switch over to OpenAI and pay them 2x what I'm paying Claude if this becomes generally available
I haven't wrapped my head around what level of reasoning this involves. Is it equivalent to max?I didn't like Sol initially but it is growing on me the more I use it. Its personality is a bit flat and I caught it taking shortcuts a few times. But once I learned how to interact with it, I'm genuinely warming up to it. I find that it writes code that has fewer bugs even than Fable (although, to be fair I reach for Fable when the task is less well defined).If this has similar performance to Sol at max reasoning level, this would be a compelling reason to shift even more of my work (maybe the majority) to this model.
This is something I'm ready to pay for. Not more per token, but I will be happy to burn through 20x Pro subscription as fast as I consume my Plus weekly limit now, with 10x more tokens per unit of time. I've learned how to deal with and steer Sol medium quite efficiently, but at the same time I realize it's so slow for the small tasks it can do well, and still so unreliable for open-ended tasks.
I don’t know if this is that useful for coding. In some autonomous world, where no one check the code and the agent can just spend 10X more time checking its work and leading to better results, yes maybe it is useful.But if humans need to check its work, then 10X speed doesn’t really matter I guess.
In light of recent news, it is hard not to think about the 4.5-day hack on Hugging Face's systems happening ~10 times faster and be slightly concerned.
> allowing Sol Ultrafast to accelerate your most time-sensitive, mission-critical workCurious, what are some of the use cases?
I will happily pay 2x luna pricing for this speed with Luna :P
Why did OpenAI partner with Cerebras when they've already built their own chip, Jalapeño?
Their dinner plate chips are impressive.
Wow, that's even faster than diffusion LLMs but with the Fable-level quality! Congrats!