Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations:<p>1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months?<p>2) Distillation - also implausible for the reason above.<p>3) Benchmark hacking. AI companies have ways they can dial up performance artificially, and will reach for that to maintain the appearance of parity.<p>Other reasons?<p>Edit: Most replies are ignoring timing. It's the near-concurrent release of the same jump in capability that I find suspicious; not the fact that labs can catch up eventually.
> benchmark hacking<p>I think this is the main one. The benchmarks from this are heavily cherry-picked, and they also widely publicised their performance for 4.5 while downplaying the fact the benchmarks were "accidentally" in their training set
Its possible no AI lab has any unique edge, and success is a combination of (a) having access to GPUs (b) having access to large amounts of data (c) know about the handful of techniques to build an LLM, of which nearly all are likely open source and documented in papers.
So the cycle of growth is (a) and (b), get more GPUs and get more data and you have a better model.
Yea, this reads as LLMs are a pretty obvious technology to develop(for the highly intelligent researchers who are there). Also there's probably a lot of actual divergence in model capabilities and skills that concealed by the fairly narrow set of tests we run them against nowadays. Like wasn't Grok 4.20 super targeted at non-coding tasks.
GPUs might explain the remarkably concurrent timing. Data access doesn't really explain it unless all labs simultaneously got access to some treasure trove of data.
There is a widespread belief that the nature of intelligence is scalar, like how a person can have 100x more wealth than another person. If this were true, then we’d probably see breakaway RSI from a single lab.<p>But I think we’re discovering that intelligence is about universality, not magnitude. This is analogous to how building a universal Turing machine wasn’t merely a matter of building a calculator that could multiply higher numbers. The difference is that with calculators we consciously theorized about what universal computation would require, then we built one as a step change. Despite it having low memory and slow speeds, the first one built was as theoretically universal as any computer we have today, in terms of the surface of computations it can perform.<p>With intelligence, it’s turned out to be less discontinuous, which I believe has convinced people that intelligence is a never ending exponential rather than an S curve approaching a horizontal asymptote. I suspect the LLMs we have today are the same kind of thing we will have in 5-10 years, but in 5-10 years we’ll consider them to be fully universal. At that point we’ll still have improvements in tokens per second and volume of context window, but not in capability per token.
But in theory you can make an LLM A LOT faster than a human.<p>You can also run massive amount of LLMs in parallel.<p>There might be a limit to a normal LLM but not to theo everall system.
You're take basically lines up with Francois Chollet: <a href="https://arxiv.org/abs/1911.01547" rel="nofollow">https://arxiv.org/abs/1911.01547</a><p>intelligence is more like polishing a ball smooth than growing the ball to infinity.<p>For many tasks, it will be smooth enough.
But aren't today's frontier models already "fully universal"? To use your Turing machine analogy, I think we're past the calculator stage.
Why would you release a model if you are the current frontrunner? Only when a competitor pulls ahead, or comes close enough to actually get traffic, you prepare a new release.
I'm sure Grok 4.6 is not Fable level. Benchmarks are almost useless.<p>Having said that, Grok 4.6 (1.5T params) is without a doubt way smaller than Fable, maybe a Fable sized Grok would be Fable level?
4) There's nothing terribly special about Anthropic. No moat.
Agreed, but my suspicion is tied to the timing. Catching up eventually is to be expected. Having similar jumps in capability ready at the same time is odd.
There's also a bit of selection bias going on here because we forget about labs that don't have a jump and just focus on the ones that do. Notably Google is definitely not having that capability jump.
Maybe "readiness" is quite a flexible category? You're mid-training for your next model; a rival releases something; you clear the boards and release the model without completing the training run?
Gradual improvements in performance can look like jumps, when you go over critical thresholds.<p>Combustion engines improved gradually, each year. One year they got better than horses.
brand is their power, they'd be wise to not wreck it with dumb moves or PR statements (they already have some)
It's because Fable is just synthetic RL tasks + scale. The secret has been out for awhile now.
Does not explain timing
All frontier labs buy the same RL tasks from task producers.
keep in mind fable = mythos which as been "done" since february. so the gap is not 2 months, it's more like - techniques probably started "working" in late 2025, now are trickling down to 2nd tier labs 9 months later.
This is basically the answer, they generate A LOT of synthetic task rollouts in parallel, then use RL on the resulting reward signals to improve the model. Add scale to this and you have a Fable class model.
I'm sure the SF AI scene leaks like a sieve, and companies have a pretty good idea what each other is working on.
> It's the near-concurrent release of the same jump in capability that I find suspicious; not the fact that labs can catch up eventually.<p>When everyone's improvement (or at least, everyone's rate of increase in parameter count) is so rapid, "within 2 months" shouldn't be seen as "near-concurrent".
Yeah, I’m not convinced that there are any models as smart as Fable. Opus 5 definitely isn’t for all it has great benchmark scores. Fable displays judgement in a way I haven’t seen from any other model.
Any models available to us that is...
Yeah as models get better, valid benchmarks become more "trust me bro".
Possibility: They're all hitting the same plateau of what LLMs can do with their current architectures.<p>I'm not stating this as a fact, but it's a hypothesis I'm keeping in my mix.
Timing doesn't seem odd to me. It just seems like <a href="https://en.wikipedia.org/wiki/Multiple_discovery" rel="nofollow">https://en.wikipedia.org/wiki/Multiple_discovery</a> which I've noticed happen in many areas.
That's exactly what Anthropic said was going to happen!<p>Their big bet is that models are going to keep getting sharply better, not that they're going to quickly reach a plateau of quality that they can then defend.
what we're going through is the same thing as smartphones, the limiter is compute.<p>it used to be snapdragon came out HTC rushed out a janky phone everyone went omg htc is goat, then in the next few weeks and months others would impliment better versions and people would not notice those as much, finally sony would release a polished phone right as the next snapdragon cycle came.<p>eventually compute gains leveled off and apple won on taste.<p>nvidia/tpu is the new snapdragon. Anthropic and google both peaked on the first training run on a new tpu cycle.<p>you should expect amazing things within a few months of each other from everyone with access to chips and willingness to use them on a training run.<p>We haven't seen willingness from google to do that. So its currently xai,oai,anthropic, and probably soon meta.
I'm pretty sure both Anthropic and OpenAI haven't necessarily been secretive that they have internal models that are much more capable than commercially available ones.<p>It's probably a mix of all of that plus simply always keeping one in the chamber to 1up everyone else when the time is right.
> <i>1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months?</i><p>The assumed timeline (2 months) is slightly wrong because Fable (Latin) is essentially the same as Mythos (Greek) albeit with protections against cyber and biological misuse.<p>Mythos (Preview) was publicly announced in April 2026 [1] which means other labs have had 4 months to catch up, not 2 months.<p>Assuming everyone had access to Mythos from the start, your expression, similar to other folks would have been "Mythos-level intelligence" and not "Fable-level intelligence".<p>1: <a href="https://news.ycombinator.com/item?id=47679258">https://news.ycombinator.com/item?id=47679258</a>
No, it has happened to almost every other "sota" model before. There used to be a meme with a circular arrow going through Anthropic, OpenAI, Google as a hype circle. Now we can drop Google and add a couple of Chinese companies.<p>It's not an explanation of why it happens, I am just pointing Fable is not an exception, it has happened with almost every other model release by all these companies over the last 2-3 years.
Sometimes you just need to know that something is possible, not exactly how it is done.
I understand Mythos became internally available on the 24th of February.<p>Other labs catching up in half a year seems about right.
I think model level is more a function of the state of hardware. Once it exists and is available (and if a lab can afford it), then they can train their own 1T, 5T, coming up next 10T model.
This is a good candidate because it would also explain the timing. Most of the replies here do nothing to explain the timing I brought up.
this is exactly whats happening. Its funny having lived through this with snap dragons and phones.<p>Everyones hyped about the branded phone, but it was the chip that mattered and how fast you rushed a product out after you got it.<p>Sames true now, except size of training run is also a factor.
More compute is coming online at all times.
I'm solidly in the "they are benchmaxxing" camp. This became very apparent with GPT 5.6 Sol. It, too, was widely hailed to have near-Fable level intelligence. But I used it non-stop for a week and realized that they had mostly just dialed up the relentlessness meter to eleven, most likely via heavy RLHF.<p>Last week I gave it a small-sized auth ticket to work on, then stepped away. I came back later that afternoon and found that it had worked for 3+ hours and written 25,000+ lines of code. I skimmed over the code and it looked like a small fix followed by a massive number of additional checks around it, including static analysis tooling.<p>I gave it to another GPT 5.6 and said "check this code and see if it addresses the ticket". It looked at it and said that 98% of it was garbage and should be thrown away (its own words). I then gave it to Fable, which said it was massively over-engineered. Fable's theory was that the agent implemented the fix first, but then compacted and lost crucial context, forgot what the original task was about, and kept going. After many compaction cycles it was completely lost.<p>Some people complain that Opus 5 stops before finishing a task. But to me, that behavior is vastly preferable to what GPT 5.6 Sol does.
Yeah I found the timing on Sol especially curious since it came right on the heels of Fable. I've had mixed results with it - sometimes it seems great, other times it makes mistakes so stupid I cannot understand how it ever gets anything right.<p>Explaining it as a difference of effort would explain both.
> It's the near-concurrent release of the same jump in capability that I find suspicious; not the fact that labs can catch up eventually.<p>What are suspicious of? If the timing is similar maybe just everyone already are of similar capabilities and got there at a similar time?<p>> Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models?<p>It means Anthropic had no real moat and no real lead. Is that weird to you?
> other reasons<p>Maybe research is sufficiently public and simple to reproduce or the next steps of how to improve things are sufficiently obvious to the smart people working on frontier AI.
Maybe compute is the real moat (chinese possibly skip around it with distillation), xai is buildouts have been insanely fast (colossus 1 - 100,000 H100 GPUs brought online in 122 days lol) so maybe that explains them catching up<p>asked grok to give a compute estimate for each:
- SpaceX / xAI: ~1.4 GW (owned Colossus clusters)
- OpenAI: ~2–3 GW (mostly rented/cloud)
- Anthropic: ~1.5–2.5 GW (multi-cloud + xAI lease)<p>chatgpt estimates a lower:
- OpenAI: ~1.5M H100-eq ± ~0.8M
- Anthropic: ~1.4M H100-eq ± ~0.7M
- SpaceX/xAI: ~0.6M H100-eq ± ~0.3M<p>but it felt obligated to mention that "for single tightly interconnected NVIDIA training clusters, SpaceX/xAI has been unusually strong."
> <i>Maybe compute is the real moat (chinese possibly skip around it with distillation)</i><p>Makes no sense. At this point, all Western AI companies also engage in distillation. If distillation were such magic, they'd be insane not to.
we are in the process of transitioning from hype to commodity with llm tokens, moats are typically at the top of the stack or in the data warehouse