AI Brief, 15 September 2026: a 753-billion-parameter model whose config file is someone else's
Shanghai Artificial Intelligence Laboratory has put an agentic model called Atria Dawn Preview on
Hugging Face under an MIT licence, and its own model card says in the opening paragraph that it is
"Built on the 744B-parameter MoE GLM-5.2 foundation model" — Z.ai's model, released in June. That
disclosure is welcome. What it understates is how little was changed around it. The checkpoint's
config.json and GLM-5.2's agree on all fifty-three shared fields, including every structural one,
and differ on exactly one: the version of the transformers library that wrote the file.
This is the second day running that a Shanghai AI Laboratory release has turned out to be another Chinese lab's foundation model with post-training on top. Yesterday's edition covered Intern-S2-397B, whose config matched Alibaba's Qwen3.5-397B-A17B except for 3,072 vocabulary slots. Atria Dawn Preview does not differ even by that much. The difference between the two cases is disclosure: the Intern-S2 card named no base model, and this one names its base in its first sentence.
The benchmark table is the part that does not hold up. Four of its sixteen benchmarks also appear on Z.ai's own GLM-5.3 card, so nineteen cells report the same rival model on the same named benchmark in both places. Three of the nineteen agree. All three are on CyberGym; on Terminal-Bench 2.1, AutomationBench and GDPval not one figure matches. Z.ai's card spends a page of footnotes on how each number was produced — harness, temperature, task count, timeout. The Atria card carries no footnote, names no harness, and gives no version for a benchmark whose version Z.ai states. A table mixing figures copied from a rival's card with figures apparently re-run in-house, without marking which is which, is harder to trust than one openly all self-run.
Elsewhere on a modest day, Microsoft published a code of conduct for its models that turns out to govern nothing yet, and Tencent released a sparse-attention module that bolts onto frozen Qwen3 weights.
- Atria Dawn Preview: 753.3B parameters in bf16, MIT licence, weights uploaded 11 September with no announcement anywhere, and no independent evaluation since.
- Its config matches Z.ai's GLM-5.2 on every structural field; the only differing key is
transformers_version. - Of nineteen benchmark cells reporting the same rival on the same benchmark in both labs' cards, three agree, and all three are CyberGym.
- Microsoft AI's Code of Conduct says its current models "are not yet trained on this document."
- The Clay Mathematics Institute acknowledged the Navier-Stokes claim on 11 September without naming OpenAI, and left the problem's status marked Active.
A model that is GLM-5.2 with different weights in the same holes
The repository history dates the release, because nothing else does: there is no announcement post on
any Atria or Shanghai AI Laboratory surface, no Techmeme entry, and no Hacker News submission. The
first weight upload to
internlm/Atria-Dawn-Preview landed at
2026-09-11T10:50:50Z and the model card a day later. Then the apparatus arrived in one burst on
14 September: a Discord server at 06:42Z, a 23-page technical report compiled at 15:28Z, and
arXiv:2609.15818 submitted at 16:22Z under roughly 144 authors.
Neither the model card nor the README links that paper, and the card's citation block is still
commented out around the words "Add the official citation here."
Both repositories declare architectures: ["GlmMoeDsaForCausalLM"], and every number underneath is
the same: 78 layers, 256 routed experts with 8 active per token and 1 shared, hidden size 6144, 64
attention heads, the DSA sparse-attention indexer with 32 heads and index_topk 2048, a
154,880-token vocabulary and max_position_embeddings of 1,048,576. The weights match too: Atria
ships 353 safetensors shards totalling 1,506,667,381,320 bytes;
GLM-5.2 ships 282 shards totalling 1,506,667,387,408. The
6,088-byte gap across 71 extra files is about 86 bytes per file, the size of a safetensors header
rather than of a model. Halving either gives 753.3 billion bf16 parameters, against the 744B the
card claims, the difference being the embedding and multi-token-prediction tensors the headline count
excludes. Identical shapes are what post-training produces, since fine-tuning changes values and not
shapes; the narrow point is that nothing architectural was added.
Both cards report Terminal-Bench 2.1 under that exact name, and on all four models appearing in both, the numbers differ:
The pattern does not point one way. Three rivals score lower on the Atria card than on their own vendor's page and one scores higher, which rules out a simple thumb on the scale and leaves the duller explanation: different harnesses and settings, unstated. AutomationBench is the same story across five models, with a 9.9-point swing on Qwen 3.8 Max. CyberGym is the exception showing the rest is not a naming artefact: three of its five cells match Z.ai's figures to the decimal.
The table has also moved since it went up. A commit at 2026-09-14T04:07:20Z titled "Update benchmark values in README.md" rewrote five rows of an already-public table, including two of the model's own scores — AutomationBench from 54.5 to 53.8, τ³-Bench Banking from 40.5 to 41.2 — and rescaled the whole GDPval row from percentage-like figures to the Elo-like ones now shown.
Read on its own terms, the table describes a model good at finding things and unremarkable at building them. It leads its comparison set on AutomationBench (53.8), BFCL v4 (77.0), CyberGym (86.5), DeepSearchQA (96.0), BrowseComp (92.5) and SkillsBench (66.4) — search, browsing, tool calls and vulnerability work. It places sixth of seven on GDPval (1583) and SWE-bench Pro (59.6 against Opus 5's 74.7), and last of the six reporting Terminal-Bench 2.1 (78.3 against Opus 5's 90.2). For a card that leads with "code implementation," the coding rows are its weakest.
One last oddity: the card credits the Shanghai Artificial Intelligence Laboratory and the weights sit under its InternLM account, but the technical report names no institution, giving two Fudan University addresses for correspondence, and the Atria website runs on infrastructure belonging to a Shanghai company that describes itself as born out of Fudan's NLP laboratory. No source states what the relationship is.
Microsoft published a code of conduct for models it has not trained yet
Microsoft AI published a Code of Conduct for MAI Models at 13:00 UTC on 14 September, with an announcement opening a six-week public consultation. The coverage framed it as Microsoft's position on whether models have an inner life. The document is narrower and stranger than that.
It is a draft, and not merely one that might be revised. It governs nothing: "we are not using it to train our models today," says the announcement, and Appendix B states that current models "are not yet trained on this document." A revised version is promised late in the year, to guide development in 2027 and beyond. The document calls itself "descriptive and aspirational" and "not a guarantee of present-day performance." Nearly every rule is future tense about model behaviour, making it a specification for models that do not exist yet.
On the question the headlines took up there is one substantive paragraph, titled "AI is Artificial." A model "is not conscious and should not be designed to imitate consciousness." Then the sentence worth quoting: "We reject the pursuit of legal personhood, or the idea that models might deserve welfare, or be entitled to rights." The reasoning is instrumental rather than metaphysical — imitating consciousness-like states "increases the challenge of containment, control, and alignment" — while the document concedes in the same breath that the science is "far from settled." The phrases "inner life" and "seemingly conscious" appear nowhere in it; they come from the coverage, which also supplies the contrast with Anthropic. The document mentions no other lab.
The part that matters for the monitoring argument is section 2.4. Models "will not tamper with chain of thoughts or code, or misrepresent or conceal their reasoning or action traces," and "do not communicate in neuralese or any form beyond simple human understanding," because "if humans can't understand it, humans can't oversee it." That lands on the problem OpenAI's chief scientist described on 7 September, when he wrote that chain-of-thought monitoring is getting less reliable. Microsoft's answer is to forbid the failure mode by policy.
Two things the coverage missed: a three-layer chain of command in which the Code of Conduct outranks the operator, who outranks the user; and that neither page carries a byline, the attribution to Mustafa Suleyman tracing to a quote he gave the paywalled Information rather than to the artifact.
A gate that exists only to carry gradient
Tencent published
Simple Attention Sparsification at
08:40 UTC on 14 September: AttnGates checkpoints for Qwen3-4B, 8B and 14B at 33.0M, 33.0M and 42.0M
parameters, three days after the paper went up. These are
"router-only checkpoints, not standalone language models" — the backbone was frozen throughout and is
not included.
The gates rank blocks of the KV cache so attention can skip most of them: keys are pooled into 64-token blocks, a gate network of hidden size 128 scores each block against the current query, and only the top-scoring blocks are attended. The interesting part is how that ranking gets trained. Top-K selection is piecewise constant — nudge a score slightly and either nothing changes or the selection jumps — so the gradient of the loss with respect to the score is zero almost everywhere. Prior work sidesteps this by not training on the task at all, distilling the dense attention map instead. The paper's objection is that this optimises the wrong thing: the router learns where the dense model attends, not which blocks matter most under a limited budget.
SAS keeps hard Top-K in the forward pass but adds a continuous gate inside the softmax, in log space:
Here
The arithmetic pays for precision about what is saved. At a 128k context, 64-token blocks and the default 2,048-token budget mean a decode step attends 32 of 2,048 blocks — 1.6% — so attention compute alone should fall about 64×. The measured decode speedups are 4.6× at 256k and 5.6× at 512k on Qwen3-4B at batch 1. The gap between 64 and 5.6 is the mechanism eating its own gains: Top-K selection grows from 21% of the decode step at 8k to 90% at 512k, because ranking every candidate block is itself work that scales with context. Prefill is unchanged, and the KV cache does not shrink — every token's keys and values stay resident, since any block may be selected next step. That is the opposite trade from DeepSeek-V4.1-Flash, covered here on 11 September, which cut the cache to 890 bytes per token. All figures are the authors' own. The budget is a serving decision:
# The gate checkpoint is the served path; the frozen Qwen3 backbone is fetched
# separately via config.json. This env var is the compute/accuracy knob —
# 1024, 2048 or 4096, no retraining needed.
export SGLANG_SEER_TOKEN_BUDGET=2048
python -m sglang.launch_server \
--model-path /path/to/Qwen3-4B-AttnGates \
--attention-backend seer_attn --trust-remote-code
Three caveats. The seer_attn backend is not in released SGLang — the current 0.5.19 wheel's list of
accepted attention backends does not contain it, so these weights run only on the authors' fork. The
gates were trained on OpenR1-Math-220k at 32,768 tokens, a narrow distribution for a mechanism sold
on long context. And accuracy does not fully recover: at the 4,096-token budget AIME25 still trails
dense attention by 2.9 to 6.4 points across the three sizes, and at 2,048 the 4B gap is 10.0. The
large reported wins are over the previous sparse method, not over dense attention.
Also notable
- Clay acknowledged the Navier-Stokes claim, and did less than reported. The 9 September edition
asked whether any mathematician outside OpenAI would read the 166-page argument. The Clay
Mathematics Institute published a statement
on 11 September, three days before the coverage that surfaced it, saying it "shares in the
excitement" as the community "contemplate the announcement that the Navier-Stokes problem has
apparently been settled." It does not name OpenAI, does not say a review has begun, and the
problem's own page still reads Active, so reports describing the proof as under formal review
overstate it. Clay's rules require publication in a refereed outlet plus two further years; the
argument is still a self-hosted PDF with no preprint or DOI, so that clock has not started, and
OpenAI's own
formalization.yamlstill records its review status as "self-assessed." - Twenty-seven probes across eleven labs' models. Ant Group's
inclusionAIpublished SingProbe checkpoints on 14 September for Llama, Gemma, ten Qwen models, GLM, DeepSeek, MiniMax, Step, Hunyuan and gpt-oss backbones, all Apache-2.0. Each taps three hidden layers and emits a per-token score for intent, unsafety and hallucination risk — 2.23M parameters on Qwen3-0.6B, 8.13M on Qwen3.5-397B-A17B, under 0.5% decode overhead by Ant's own measurement. Like SAS, it ships against a fork. - Rubin's first third-party inference numbers. SemiAnalysis published agentic-inference benchmarks for Vera Rubin NVL72 late on 14 September, claiming up to 7× Blackwell's token throughput per megawatt on a 1.6-trillion-parameter DeepSeek model, against the 3× Nvidia gave at GTC. The runs are on pre-release software and unaudited.
- A code-review cost comparison from a vendor with a stake in it. Entelligence ran GPT-5.6 Luna against GPT-6 Astra on 50 public benchmark pull requests on 14 September: 69 verified bugs for $0.20 against 92 for $5.66, but Luna returned 24 false positives in 93 findings to Astra's 4 in 96, and caught 9 of 24 security bugs to Astra's 19. Data is public; the company sells cost-based routing.
What to watch
- Whether anyone evaluates Atria Dawn Preview independently. Its base model is a ready-made control: scoring Atria and GLM-5.2 on one harness would settle in a single run what the post-training bought.
- Whether Microsoft's consultation changes anything measurable. The six weeks end around 26 October. The test for the next version is whether it names a threshold, a test for reasoning legibility, or an external reviewer — all three are absent from this one.
- Whether the Navier-Stokes argument reaches a journal. Nothing in Clay's process can begin until it does, and no named mathematician has yet published an assessment of the proof's correctness.
- Independent numbers for Tencent's Hy4 preview remain absent from Artificial Analysis for a tenth consecutive issue, though cwe-bench has carried a third-party Hy4 score since 1 September.