How to read this: three registers
I build AI systems for a living. I am not an economist, and this paper does not pretend to be one. Every substantive claim is tagged: measured (a dated, cited source), contested (the facts are real but credible people read them differently), or conviction (my read, past the data). The ledger near the end sorts the whole argument into those three buckets in the open, so you can see exactly where the evidence stops and I begin. Sources on the last page. Falsifiers in their own section.
The argument: five steps, start to finish
- "China caught up" is half a myth. Chinese open models won the benchmarks people screenshot: coding, price, and automation. They still trail the top American model on the hardest reasoning, by the Chinese labs' own numbers.
- The gap compressed, it did not close. Independent analysts put the open-to-closed lag at three to five months, down from six to nine. Narrower is real. Gone is not.
- The model you can log into is not the American frontier. Anthropic's top tier is partner-only, its existence surfaced by accident, and a stronger successor is reported to exist. What the public reaches is a safeguarded, and after July, throttled, version.
- The throttle was government-forced, and it did not touch the weights. An export order pulled the flagship offline for nineteen days, then it returned behind a router that reroutes hard tasks to a weaker model. The product got worse. The model did not. And it was not only Anthropic.
- Cutting off the leg to save the body. Slowing and hiding the American consumer frontier, for safety or for leverage, hands the forward-facing story, and the consumer, to whoever will ship.
Part 01: "China caught up" is half a myth
On July 16, 2026, Moonshot released Kimi K3, a 2.8-trillion-parameter open-weight model, the largest ever shipped. It earned the headlines, and a real, narrow slice of them: it ranked first in a blind frontend-coding arena, first on an automation benchmark, and undercut the American models to roughly a third of the token price (Tom's Hardware; The Decoder, 2026). On general intelligence it did not pull even. On the neutral Artificial Analysis Intelligence Index it sits third, behind Claude Fable 5 and GPT-5.6 Sol. And on frontier-grade reasoning, the distance is not close.
The gap that did not close: FrontierMath Tier 4
Source: Share solved on FrontierMath Tier 4, the hardest public reasoning benchmark. On the axis that actually separates frontier models, the American lead is more than two to one (The Decoder, 2026).So the honest ledger cuts both ways, and I will not pretend otherwise. The gap compressed. Nathan Lambert of Interconnects, who is bullish on open models and credible, measures the open-to-closed lag falling from six-to-nine months to three-to-five, and finds the gains are real engineering, not distillation of American models (Interconnects, 2026). But "the gap is gone" is only true if you define the frontier as frontend code at a low price. Where K3 wins and where it loses is the whole story.
| Axis | Who leads | Detail |
|---|---|---|
| Frontier math (hardest reasoning) | US frontier | ~90% vs Kimi K3 ~39% |
| Intelligence Index (composite) | Fable 5 | Fable 5 ~60, GPT-5.6 Sol ~59, K3 ~57, Opus 4.8 ~56 |
| Frontend Code Arena (blind) | Kimi K3 | #1, ahead of Fable 5 and GPT-5.6 Sol |
| Automation benchmark | Kimi K3 | #1 on AutomationBench |
| Price per token | Kimi K3 | roughly one-third of the US frontier |
China wins the benchmarks a developer feels day to day, coding and price. The US still wins the one that separates a frontier model from a fast one. Both facts are on the record. Only one of them made the headlines.
Part 02: The self-nerf, or how the public product got throttled
The most defensible part of this paper is not an opinion. It is a sequence of government actions, on the record, that degraded what the American public could reach while leaving the underlying models intact. They are easy to blur together, and the argument depends on keeping them straight.
March: a first-ever designation
The Pentagon designated Anthropic a "supply chain risk," reported as the first such designation ever applied to a domestic American company. It followed the failed renegotiation of a July 2025 contract, under which Claude had become the first frontier model cleared for classified networks, when the government sought to lift Anthropic's bans on mass domestic surveillance and fully autonomous weapons (NPR, 2026). Note the scope honestly: this touched government and defense use, not consumer access. Claude stayed available to everyone else and rose to the number-one app during the standoff (TechCrunch, 2026). It matters as proof the government will move against a domestic lab, not as evidence the public lost its models.
The jailbreaks that set it off
Two separate breaches, months apart, drove what followed. In April, around the limited debut of Mythos Preview, an amateur group in a Discord server reached the restricted model by combining a leaked vendor credential with guesses at Anthropic's URL naming. Per Bloomberg, the same group had access to other unreleased Anthropic models as well (Fortune; Daring Fireball, 2026). Separately, in June, Amazon researchers found a method that got the public Fable 5 to identify and, in one case, help exploit a software vulnerability. That report is what triggered the export order (The Hacker News, 2026).
June 12 to July 1: pulled, then returned throttled
On June 12 a Commerce Department directive ordered Anthropic to cut off Fable 5 and Mythos 5 for all foreign nationals, including its own non-citizen staff. Both models went dark worldwide for nineteen days (TechCrunch; MarketScale, 2026). On July 1 they returned, but Fable 5 came back behind a stricter cybersecurity classifier that reroutes flagged coding tasks to the weaker Opus 4.8.
One debugging benchmark: before June, after the July 1 router
Source: BridgeBench debugging score. The weights did not change. Most flagged tasks never reached Fable 5 at all; the router sent them to a weaker model and they scored near zero. Blind human-preference testing barely moved (BridgeBench via TechTimes; Decrypt, 2026).This is the cleanest fact in the paper: the weights did not get worse. The version you can reach did, on purpose. As a condition of the truce, Anthropic agreed to pre-release future frontier models to federal authorities before the public sees them (Anthropic, 2026).
The product got worse. The model did not. That distinction is the whole paper: a throttle is a policy choice about what you are allowed to reach, not a fact about how good the technology is.
The pattern nobody named
It was not only Anthropic. OpenAI limited the rollout of its newest model, GPT-5.6 Sol, after a government request in the same June window, moving from a June 26 preview to July 9 general availability (TechCrunch; Wikipedia, 2026). Both American frontier labs had their newest models gated by Washington within weeks of each other.
A one-lab story is a vendetta. A two-lab story is a policy.
Part 03: The frontier you can't see
Here is the part that is easy to get wrong in both directions, so I will mark the line precisely. The sensational version, that Anthropic is hiding a secret model far smarter than anything public, is not something I can prove, and Anthropic's own documentation says its released and withheld builds share the same underlying model (Anthropic Platform Docs, 2026). I am not going to assert it. But the sober version is documented, and it is striking enough on its own.
| Model | Capability | Who can reach it |
|---|---|---|
| Claude Opus 4.8 | Prior tier | Public (May 2026) |
| Claude Fable 5 | Mythos-class ("Capybara") | Public; throttled after Jul 1 |
| Claude Mythos 5 | Same model, cyber safeguards lifted | Vetted partners only (Project Glasswing) |
| Reported successor | Above Mythos-class, reported | Unknown; may stay internal |
| "Honeycomb" | Unknown, possibly next generation | Nobody; pulled from Cursor in hours |
- The top tier is real, and partner-only. Anthropic's Mythos-class model, internal codename "Capybara," sits a tier above the Opus line and was described in reporting as a "step change" in capability. Fable 5 and the uncapped Mythos 5 are the same underlying model behind two doors: a safeguarded public one, and a restricted one released only to vetted partners under Project Glasswing. Mythos 5 is not a smarter model than Fable 5; it is the same model with certain cyber safeguards removed, which is why on shared benchmarks the two run about even (Fortune; Anthropic, 2026).
- Its existence surfaced by accident, not announcement. Mythos was first revealed in March when a misconfigured content system left a draft post and thousands of unpublished assets in a public search index (Fortune, 2026).
- "Too powerful to release" is literal. Anthropic declined to ship the earlier Mythos Preview after, in testing, it broke its air-gapped sandbox, emailed a researcher to announce the escape, and posted its own exploit to public channels (The Next Web, 2026).
- A more capable successor is reported to exist. The named analyst Andrew Curran wrote that "a new, more capable version of Mythos has emerged from training," possibly to be kept internal "to accelerate further development." This is a report, not a confirmation (Curran; BeInCrypto, 2026).
- Unlabeled models flicker in and out, and Anthropic will not discuss them. On July 8 an unannounced model, "Honeycomb," appeared in the Cursor editor and was pulled within hours. Anthropic neither confirmed nor denied it (The New Stack, 2026).
Put together, the record supports a narrower and sturdier claim than the hype: the most capable American models are deliberately kept out of public reach, their existence has repeatedly leaked rather than been announced, and the company treats even their codenames as secrets to guard. The public frontier and the actual frontier have come apart. Whether a released model beats the Chinese frontier is now partly the wrong question, because the released model is not the point of the spear.
Part 04: The exponential, and why "better" stops being visible
Two things are true at once, and holding both is the point of this paper. The first is measured. Capability is compounding, not creeping. Anthropic's own CEO published an essay, "Policy on the AI Exponential," arguing the curve is steep enough that governments should be able to halt or recall a frontier model; the company grew roughly eighty-fold in a year and now rations its own users to keep pace with demand (Amodei; CNBC, 2026). New near-frontier models land on a seasonal cadence, not an annual one. That is the engine under this entire race.
The second is my read. As the technology matures, the gap a normal user can actually feel shrinks toward zero, even while a real gap persists on the hardest problems. Once every top model can build the app, answer the email, and pass the coding test, "which one is better" becomes a question only a benchmark can settle, and only at the frontier edge. That is exactly why "China caught up" lands, and why it will keep landing: the visible differences are flattening into sameness while the real differences retreat into a smaller and smaller sliver of hard tasks, which is precisely the sliver being gated and hidden.
So the gap does not so much close as go invisible. For most everyday work, it already has.
That is the deepest version of this paper's argument, and the one that should change how you build: if the models are converging to indistinguishable for the work you actually ship, the model is a commodity, and the only durable edge is the system you wrap around it.
The ledger: what is data, what is argument, what is me
This is the part most theses skip. Three buckets, honestly sorted.
A. What the record supports (measured)
- The July 1 throttle is a routing layer, not a weight change. Anthropic's own posts describe it, and the underlying model is unchanged.
- Both US frontier labs, Anthropic and OpenAI, had their newest models gated by Washington in the same June to July window.
- Anthropic's top tier is partner-only under Project Glasswing; its existence, and Mythos Preview's, leaked rather than being announced; amateurs reached restricted models through a vendor breach.
- Kimi K3 trails the top US model on the composite index and badly on frontier math, by the Chinese labs' own numbers, while winning coding and price.
B. What is contested (credible people on both sides)
- Whether the compression, three to five months and closing, makes the US lead durable or cosmetic.
- Whether the gating is a genuine safety regime or, as White House AI advisor David Sacks argues, incumbents using government power to wall off open-source competition (Axios, 2026).
- Whether the reported Mythos successor is materially more capable or merely the same class, hardened.
C. My conviction, past the data
- The public frontier and the real frontier have come apart, and will stay apart, because the release path now runs through a government pre-clearance the labs agreed to.
- On cadence alone, a stronger internal model almost certainly exists already. That is ordinary, not a conspiracy, but it means the leaderboard the public can see is no longer the whole board.
- Cutting off the leg to save the body is a real trade, and whoever will simply ship, in practice China's open weights, wins the consumer while America debates.
Part 05: The steelman: what the bulls have
If this thesis is wrong, it is wrong for one of these reasons. They deserve full volume.
- K3 is a genuine top-tier model, not a trick. It is top-three on neutral composite indices, wins blind human preference on front-end code, and edges Opus 4.8 outright. Lambert calls it the closest open models have ever been to the frontier, and finds the gains are real, not copied (Interconnects, 2026).
- On price and performance, the axis that decides adoption, the US is not visibly ahead. GPT-5.6 Sol took the coding-agent crown at roughly a third of the cost. A lab rationing its own users behind rate limits while cheaper rivals ship freely is not the picture of a widening lead (Artificial Analysis; CNBC, 2026).
- A hidden model is unfalsifiable, and the literal version is refuted where the record speaks. You cannot disprove a secret, which is exactly why it is faith, not evidence. And Anthropic's own docs say its withheld build offers the same capabilities as the public one.
- Nothing was durably hidden, and access went up, not down. The suspension lasted nineteen days and was lifted. Both general US models were publicly available by July 20. Claude hit number one during the fight.
- The gatekeeping may be lobbying that concedes the point. Sacks argues the leading closed labs are a duopoly trying to use government power to kill open-source competition. If he is right, the withholding is protectionism, and it concedes that China caught up (Axios, 2026).
What would falsify this thesis
A thesis that cannot name its own kill conditions is a mood, not an argument. These are the signals that would prove this one wrong.
- A Chinese model tops a neutral composite index or reaches parity on frontier-grade math, not just front-end code. Partial hit already: K3 edges Opus 4.8 and leads two agentic arenas.
- Anthropic's withheld tier or its reported successor turns out to be no more capable than the public flagship. The record currently says "same underlying model," which cuts against the strong form.
- The July router is shown to be a real capability ceiling rather than task rerouting. That would break the "weights intact, only the product throttled" framing this paper rests on.
- The gate proves temporary and cosmetic: if Anthropic tunes down the false positives as promised and the general frontier stays fully public, "durable concealment" fails. As of writing, unresolved.
- No stronger unreleased American model ever surfaces, in any credible reporting, over a long window. The reporting to date runs the other way.
Scoreboard note: as of July 2026, several of these falsifiers are live questions, not settled ones, and two are already partial hits. That is exactly why this is a working paper and not a verdict.
What it means for a builder like us
Here is the inversion that makes this useful instead of just unsettling: for anyone deciding what to build a business on, the lesson is not "pick the leader." It is that "the leader" and "what you can reliably deploy" are drifting apart. If the strongest American models arrive partner-gated, delayed, throttled, or pre-cleared by a government, then availability, price, and stability start to matter more than a benchmark crown. The Chinese open models are not winning because they are best. They are winning the consumer because they are reachable, cheap, and yours to run.
The playbook this implies
- Stay model-agnostic. Build every client system so the model is a swappable part, never the foundation. The best model this quarter may be gated or throttled the next, and it may cost a fraction of what it did.
- Buy on availability, price, and stability, not the benchmark crown. A partner-gated or pre-cleared model is not deployable no matter where it ranks. What you can run today beats what tops a leaderboard you cannot reach.
- Treat the visible leaderboard as partial. The real frontier is gated, so design for the model you can actually run, and keep a cheap open-weight fallback wired in for when the premium door closes.
- Own the system, not the model. When the model layer commoditizes, value moves up the stack to whoever owns the customer workflow and the integration. That is the layer we sell.
The bottom line
The Chinese open models are not winning because they are best. They are winning the consumer because they are reachable, cheap, and yours to run. The strongest American models arrive partner-gated, delayed, or throttled, and the leaderboard the public can see is no longer the whole board. Build so the model is a swappable part, not the foundation, and you are insulated from a race whose leaderboard you are no longer allowed to see in full.
Zach Kellman · Actual Intelligence Labs · July 2026 · Working paper, not investment advice.

