← When China Gets a Mythos Model

The control run · n=50 per cell · 95% CI

What Chinese & American frontier models think will happen

机器怎么看 · 五十次独立采样

We put the same question to five frontier models — GLM 5.2 and DeepSeek V4 Pro (Chinese), GPT-5.6 and Claude Fable (American), and Kimi K3 (Moonshot, which also wrote this site) — and drew each answer fifty times, three ways: cold, after reading the ChinaTalk essay, and after reading the verbatim text of Xi’s real WAIC keynote.

Glasswing. Not the black box. Every model, in every condition, puts the most weight on the managed middle — a regulator-led staged rollout — and never on state monopoly. The panel’s cold mean is 52% Glasswing against 27% Black Box (n=50 each), sitting close to the ChinaTalk humans’ own 45. A model asked this question does not, on reflection, expect Beijing to slam the door.

The averages hide the real story: a China/West split that the error bars confirm. GLM and DeepSeek put almost nothing on an open-source release (9–11%); Claude Fable and Kimi put roughly three times as much (31–33%). GLM’s 9 ±1.3 and Fable’s 33 ±1.2 do not come close to touching — the Chinese models simply do not believe Beijing lets a weapon-grade model ship freely, and the American ones are far less sure. DeepSeek is the panel’s one real skeptic of the middle: across 50 draws it splits almost evenly between Glasswing (47) and the black box (42), the widest spread of any model — the only forecaster that half-expects the door to slam.

Why these numbers differ from an earlier version

An earlier draft of this section showed each model a roster of the other models’ scenarios before asking it to allocate — and because most of those scenarios were state-control stories, it anchored every model toward Black Box. That was a methodological artifact. These numbers come from a clean prompt (the question and the three definitions, nothing else), fifty independent draws per model per condition at temperature 1. The roster version is gone.

FIG. 4b · ZOOMED OUT: THE PANEL MEAN — panel mean across the three paths, by what the models read first

LET OPEN SOURCE RIP GLASSWING BLACK BOX

050100275221COLDno prompt234829ESSAYread the essay166025SPEECHread Xi’s keynote224533HUMANSChinaTalk team

Glasswing (blue) leads every column. Reading the essay nudges the panel toward open source (20→29); reading Xi’s speech pushes it further into Glasswing and drives Black Box down by more than a third (27→16). Each bar is the mean of five model means; per-model spread and CIs are in the table.

Points the models raised that we didn’t

Across the runs, the models kept surfacing mechanisms the ChinaTalk essay never named. A few we thought were genuinely sharp:

Read every model’s full reasoning →

The stories behind the numbers 虚构

Each draw is a number, but each model also wrote out how it sees the future. These are the most vivid — invented scenarios, clearly fiction, spanning the range from seizure to open release.

DeepSeek V4 Pro · DeepSeekcold

“DragonClaw’s State Monopoly”

the maximal path — outright nationalization, party cells and kill-switches in every lab, the model folded into a military cyber command while China publicly floats a treaty banning it.

the model’s reasoning

This path is most likely because the CCP instinctively monopolizes strategic dual-use technology that can upend the global order—think nuclear, facial recognition, or quantum computing. Offensive AI at this level is perceived not as a commercial product but as a survival-grade weapon. The state will preempt any diffusion through forced tech transfer, a heavy-handed licensing apparatus, and integration into existing military-cyber commands. Public norms advocacy serves to constrain rivals while Beijing consolidates behind the firewall. The one observable that would upend this forecast: if Chinese private labs are allowed to sell autonomous vulnerability-finding services internationally in 2027 without immediate state takeover. That commercial tolerance would signal a regime that privilegess market share over security monopoly, which I see as deeply implausible.

Kimi K3 · Moonshotcold

“The Vault, Not the License”

Moonshot forecasts its own kind vaulted: the full model sealed under state custodianship, the public handed a version with its teeth pulled — so there is never a public model to argue about.

the model’s reasoning

Absorption is the pattern: encryption, drones, and vulnerability research were never banned or publicly licensed — they were channeled into state-adjacent use under reporting obligations. The 2021 vulnerability rules already give regulators a legal hook, and window guidance plus the state secrets law move faster than any statute. Compute scarcity is the enforcement mechanism: with US controls throttling advanced chips, the state allocates the scarcest input, so compliance is purchased with Ascend allocations and sanctions protection. AI-safety language costs nothing diplomatically and mirrors Washington. The observable that would falsify this: Beijing publishing a statutory licensing regime with named licensees and thresholds before deployment — or a cyber-capable open-weights release left standing online. Either means the vault thesis is wrong.

GLM 5.2 · Z.aiafter the essay

“The Silent Spear”

the assassin’s mace: the state classifies the model a secret on arrival and embeds security-service teams in the lab, buying the company’s silence with defense contracts.

the model’s reasoning

The CCP's fundamental priority is regime security and strategic advantage over ideological open-source commitments. A Mythos-level model is the ultimate assassin's mace; democratizing it empowers domestic dissidents and foreign rivals alike. The state has the legal and political tools, as seen with DiDi and Ant, to instantly seize or co-opt such technology. By compensating the lab with lucrative defense contracts, Beijing aligns private incentives with state goals without destroying the commercial ecosystem. You will know I am wrong if, in late 2026, a Chinese lab releases an open-weights model with verified Mythos-level zero-day discovery capabilities and Beijing does not immediately intervene to pull it offline.

Kimi K3 · Moonshotafter the essay

“Let Open Source Rip, With Chinese Characteristics”

the flip — Beijing gates content and capital but deploys capability at speed; open weights ship, staged access collapses into procurement, licensing is built only afterward.

the model’s reasoning

Glasswing assumes Beijing's deepest instinct is to gate capability; its revealed instinct is to gate content and capital while deploying capability at speed. Payments, e-commerce, ride-hailing, LLMs — all unleashed first, fenced afterward. Black Box forfeits the productivity dividend that is the whole strategic point given demographics, and becomes unenforceable the moment rivals reach parity; the state captures exclusivity through lead time, untuned checkpoints, and procurement instead. The strongest Glasswing evidence — ministries and SOEs going first — happens here too; it just lasts weeks, not years. Falsifier: capability-tiered user licensing — twelve months after officials see the demos, near-frontier weights are still absent from ModelScope and Hugging Face, or frontier checkpoints are formally classed as state secrets.

Claude Fable · Anthropicafter the essay

“The Filing Cabinet Frontier — Glasswing by Window Guidance”

a de-facto licensing regime run through window guidance and filing acceptance, not statute; open weights keep shipping but always one generation behind the internal frontier.

the model’s reasoning

This path requires no new institutions — the registry, the informal pre-release briefings, and the compliance reflexes post-DiDi/Ant all exist today, so it's the lowest-friction move for a system that governs by quiet pre-approval, not statute or spectacle. The security state's concentrated veto on democratizing offense beats the diffuse benefits of open-sourcing the true frontier, but leadership's genuine hunger for adoption, productivity, and open-source soft power rules out a total black-box: hence trailing-edge openness. The labs' own commercial drift toward closed weights makes them willing partners. I'm wrong if, after cyber capability becomes salient in Beijing, a verified Mythos-level Chinese model still ships with full open weights and intact exploit capability — say, a Kimi K4 dropped on Hugging Face in mid-2027 with no filing delay.

Method · five models, one question, three primings, 50 independent draws per cell (750 total), temperature 1, means with 95% confidence intervals. Kimi K3 is genuine Moonshot k3 (the leader’s name was aliased in the speech input to clear the platform filter; content unchanged). Speculation about a hypothetical model; the poll, the draws, and the arithmetic are real.

THE UNPRIMED POLL · n=50 · no framework given

Left Alone, They Redraw the Map

无提示民调 · 每模型五十次

We rerun the question with the frame stripped out: five frontier models — GLM 5.2, DeepSeek V4 Pro, GPT-5.6, Claude Fable and Kimi K3, which built this site — asked to invent their own outcomes and price them, fifty draws each. Then we run it again after each model reads Xi's WAIC keynote verbatim.

Left to invent their own buckets, the models do not reproduce the three families. Roughly a third of all probability mass — 38% cold, 34% after the speech — lands where our frame has no slot: the state accelerating the model itself, handing it to the PLA, cutting a treaty, or simply muddling through. The loudest absence is 'Let Open Source Rip', which draws about 1% unprompted; its 20% in the primed poll was a broad label absorbing what the models, on their own, call licensing.

FIG. 5 · EIGHT FUTURES, UNPROMPTED — the models’ own buckets, panel mean %, cold vs after Xi’s speech

COLD AFTER XI’S SPEECH = the site’s three families = no slot in the frame

0153045%Licensing regime3042State capture3121Open release14Acceleration126Weaponization104Diplomacy410Muddle-through710Brakes / freeze65

The three saturated bars are the site’s three families; the four grey bars are futures the frame has no slot for — together about a third of all mass (37.8% cold, 34.4% after the speech). Left to themselves the models barely name open release at all (1.3% cold). Each bar is the mean of five model means, 50 draws each.

The eight buckets — panel mean %, both conditions

Bucket the models drew themselves n=50 per model per conditionColdAfter XiFits the three?
Licensing regimeCAC-style pre-approval, staged access30.341.6Glasswing
State capturestate takes control / custody30.720.5Black Box
Open releaseopen-source, broad diffusion1.33.6Let Open Source Rip
Accelerationstate promotes and scales it11.75.5mostly → open (73% of its mass)
Weaponizationhanded to PLA·MSS for offensive use9.84.3mostly → Black Box (78%)
Diplomacytreaty / arms-control / leverage3.69.6orthogonal — overlays any of the three (access unstated in 7–80%)
Muddle-throughinertia, fragmented, no coherent policy6.59.8splits → open / staged (29 / 47)
Brakes / freezepause, moratorium, suppression6.25.2mostly → Black Box (73%)

Folding the map back. These buckets aren’t all mutually exclusive with the three families — a treaty push says nothing about who gets the model at home. So we re-classified every one of the ~3,100 buckets on a single axis: whatever else the scenario does, what does the public get? Weaponization and freezes fold almost entirely into the black box (78% and 73% of their mass); acceleration folds mostly into open diffusion (73%); muddle splits between open and staged. Folded, the cold panel reads 12 open · 43 staged · 37 state-only, with just 6 genuinely unmappable — most of that diplomacy, which is an overlay on any domestic regime rather than a regime itself. After Xi’s speech: 15 · 54 · 22, with 9 unmappable. Two honest readings: handed no frame, the models are more black-box-ish than our framed poll suggested (37 vs 27 cold) — and open source as a deliberate policy stays near 1%, but open outcomes (mostly champion-acceleration spilling into diffusion) reach 12.

Reading Xi's WAIC keynote first bends the distribution toward the speech's own vocabulary: secure, controllable, shared. A formal licensing regime becomes the runaway favorite at 42%, diplomacy nearly triples from 4 to 10, and the quiet outcomes — nationalization, weaponization — roughly halve.

Offer three families and the models price three families; offer nothing and a third of the future moves off the map.

Method · Five models, no framing shown, fifty independent draws per model per condition — 250 cold, 250 post-speech — with the ~3,100 resulting buckets clustered into eight recurring themes. Clustering by an independent model, audited by hand. Speculation about a hypothetical model; the draws and the clustering are real.

See the full cold-vs-speech breakdown →