AI Brief, 3 October 2026: a superhuman result that cost under eight thousand dollars
The most consequential result of the past few days was published on 30 September and did not appear in this brief at the time, which is worth correcting now. In Nature, a group from Carnegie Mellon, NYU, Stanford and MIT reported that an agent called Ataraxos beat Pim Niemeijer, the most decorated Stratego player in the game's history, by 15 wins to 1 with 4 draws over a 20-game series. Stratego is the imperfect-information benchmark that resisted the methods that solved chess and Go: each side hides the identity of all 40 of its pieces, so a player cannot even enumerate the position they are in.
The margin is not the story. The cost is. DeepMind's DeepNash, the previous serious attempt, trained on 1,024 tensor-processing-unit nodes for two to three months, which the Ataraxos authors price at roughly $3,000,000 to $4,500,000 at 2025 rates; it never beat top humans, losing to most of the strongest players it faced at the 2023 world championship, Niemeijer among them. Ataraxos trained on 16 H100s for one week, plus four more for four days for a second network: under $8,000. It consumed about 160 million self-play games against DeepNash's 5.5 billion. A result that was previously out of reach of an industrial research budget was reached on a cluster a well-funded lab would consider a rounding error, by six academics.
Separately, and with more immediate consequences for anyone reading this, arXiv began rate-limiting every author on 1 October: a maximum of two submissions per calendar month and three active submissions at any time, across all categories, with rejected submissions counting against the total. arXiv received 40,363 submissions this September against 20,569 in September 2024, and names AI-assisted writing as the driver.
Otherwise the window was quiet in the way Saturdays are: no arXiv announcement batch, and none of the dozen labs that ship weights published anything in the last 24 hours.
- Ataraxos beat Pim Niemeijer 15-1-4, an 85% effective win rate, and went 38-2 across 40 games against world-championship attendees in a separate demonstration.
- Training cost under $8,000 against an estimated $3m-$4.5m for DeepNash, and about 160 million self-play games against 5.5 billion.
- DeepMind declined a head-to-head, telling the authors DeepNash's code no longer runs.
- arXiv now allows two submissions per author per calendar month, three active at once, all categories.
- llama.cpp merged
/v1/systemoneon 2 October with official GGUF conversions for five decision models — a different endpoint name from the/v1/decisionsSGLang shipped 32 hours earlier. - An independently built cyber-range benchmark puts a local Qwen3.8 27B at 28.1% pass@1, and 0% on binary exploitation.
Stratego fell to a learning rate schedule
Ataraxos (Sokota, Vinitsky, Hu, Fan, Kolter and Farina, Nature 658, 55-59, published 30 September, submitted 20 November 2025) is built from three ordinary-sounding parts: a policy-value network trained by self-play, a belief network trained to predict the opponent's hidden pieces, and a search procedure applied at test time.
The belief network is what makes search possible at all. In a perfect-information game you search from the position; in Stratego there is no single position to search from. So Ataraxos samples plausible completions of the hidden state, evaluates candidate moves under each one, and then performs one policy-improvement step, applied only to the decision at hand.
Because that step has the same form as a training update, the search inherits the improvement guarantees of the learning algorithm rather than needing its own theory. That is why the same code produced a superhuman Barrage Stratego agent, a new state of the art on Hanabi, and wins over PerfectDou and DouZero at dou dizhu, spanning adversarial, cooperative and team games.
The training innovation is smaller to state and is the part that actually bought the compute saving. Policy updates of this kind take the general form
where
Work the compute arithmetic and the gap is larger than "orders of magnitude" suggests. Against DeepNash's $3m-$4.5m, Ataraxos's sub-$8,000 run is 375 to 563 times cheaper. On data it used about 160 million games against 5.5 billion, a factor of 34, and roughly 50 billion training examples against 5 to 10 trillion, a factor of 100 to 200. The cost ratio is larger than the data ratio, which says the saving came from hardware efficiency and a shorter schedule as much as from sample efficiency.
The comparison with DeepNash cannot be settled directly. The authors asked DeepMind for a head-to-head and offered to build whatever infrastructure it needed; DeepMind replied that DeepNash's code is no longer functional. A four-year-old result from a major lab, published in Science, is now unrunnable, so the headline claim of this paper rests on a cost estimate rather than a match. Ataraxos's own code is public at github.com/AtaraxosAI and the 20 evaluation games are posted, so the newer side of the comparison is at least reproducible.
arXiv is rationing submissions, and says why
arXiv changed its rate-limit policy on 1 October: two submissions per author per calendar month, three active submissions at any time, across every category. A rejected submission still counts, because the policy is rationing moderator attention rather than archive space.
The numbers behind it are the interesting part, because they are a measurement of what generative models have done to scientific publishing volume.
Submissions have doubled in two years. cs.AI alone is up more than sixfold over the same period. September's
40,363 submissions generated close to 9,000 support tickets for staff and volunteer moderators. arXiv's
stated diagnosis is specific rather than general hand-wringing: an increase in "thin papers of narrow scope",
in "salami" papers where one piece of work is split into several submissions, and in dense, AI-written
manuscripts. Its policy already permits AI as a research tool when disclosed; the complaint is that a small
number of authors are consuming a disproportionate share of moderator time and delaying everyone else's
papers by days or weeks.
The decision models reach llama.cpp, under a third name
Yesterday's edition reported SGLang shipping a native /v1/decisions endpoint at 01:09 UTC on 2 October and
called it the clearest sign that typed-decision models had become infrastructure. Eight hours later the point
was made again, more strongly, and with a different name on it.
llama.cpp merged PR #29818 on 2 October, from maintainer
ngxson, adding a /v1/systemone server endpoint with runtime support for five models: laya, julia-1,
lev, openjev (including vision) and kev. Pre-converted weights are published by the project itself
rather than by the vendors, at ggml-org/OpenJev-GGUF, ggml-org/lev-GGUF, ggml-org/Laya-GGUF,
ggml-org/Julia-1-GGUF and ggml-org/Kev-4B-GGUF, so serving one is a single command:
llama-server -hf ggml-org/Laya-GGUF
# then POST to /v1/systemone instead of /v1/chat/completions
The detail that makes it cheap is that these models are wrappers around existing embedding architectures,
BERT and Qwen, so the change is largely confined to a server_decision_context plus a small decision-head
addition in libllama. That is also the most informative thing anyone has said
about the class this week: a decision model is an embedding model with a head and a calling convention, and
the engine that supports the most hardware could add five of them without restructuring anything.
Two things follow. First, Laya — the Apache-2.0 release from Convai Innovations that reached the top of
Hugging Face trending in September — now has official GGUF conversions in the most widely deployed local
inference engine, which is a larger distribution event than any vendor announcement this week. Second, the
field now has two incompatible endpoint names shipped 32 hours apart: SGLang's /v1/decisions and
/v1/score, and llama.cpp's /v1/systemone, which is the name TypeSafe AI's proprietary Jev API uses. A
class that was one startup's API eighteen days ago now has two open implementations that cannot be swapped by
changing a base URL.
Also notable
- An independently built cyber-range benchmark published numbers on 2 October that are worth more than the vendor evaluations around them. RangeBench v2 gives each model a shell in an isolated Docker target and requires the exact flag: 19 tasks, 6 models, 544 scored attempts, run 27 September. MiMo 2.6 Flash leads at 73.7% pass@1 and a local Qwen3.8 27B manages 28.1%, with zero on binary exploitation. Read the coverage column before the ranking: pass@1 is averaged over different task subsets per model (Luna 16 of 19, LongCat 18 of 19), and Luna's 90.9% pass@3 covers only 11 of the 19 tasks, so the leaderboard ordering is not like-for-like.
- FLUX 3 Image (Black Forest Labs, 1 October) replaces prompt engineering with layout. The canvas is a
0-to-1000 grid whatever the aspect ratio, and each element is a row in a JSON table carrying an id, a
description and a box written
[y_min, x_min, y_max, x_max]— y first, which will catch people out. It renders natively at 4K (a published sample is 5456x3072) and accepts up to ten reference images. It is not open weights: the page offers a commercial weights licence on request. The release carries no benchmark numbers of any kind, which is now the second major image model in a fortnight to ship without one. - DeepSeek Harness went into public preview worldwide as an open-source desktop application for macOS on Apple silicon and Windows, built on a plugin architecture the company calls Cordis. It took 385 points on Hacker News. The interesting part is not the app but the framing: everything, including the agent loop and subagents, is a plugin.
- ds4, Salvatore Sanfilippo's MIT-licensed C inference engine, reached the Hacker News front page on 2 October with 189 points. Its Qwen3.8 Flash Next build is a good illustration of what asymmetric quantisation buys: 41.73 GiB of resident weights using IQ2_XXS for the gate and up expert projections and padded Q2_K for the down projections, inside a single 137.10 GiB file whose remaining 95.37 GiB of bf16 n-gram tables are read from SSD rather than loaded. That is what puts the model on a 64 GB Mac. The project is not new, and its traction is long-standing rather than a release event.
- Ling 3.1 Flash (inclusionAI) became callable on OpenRouter at 14:07 UTC on 2 October: a 560-billion parameter mixture of experts with 25 billion active and a 262,144-token context, served by a single provider at promotional zero pricing. No weights and no vendor announcement, so the API listing is the only artifact.
- llama.cpp also merged a CUDA change fusing shared experts into the MMVQ path, with the contributor's own measurements showing 2.3% to 5.1% higher decode throughput. Self-reported, single machine, unreplicated.
What to watch
- Whether anyone reproduces Ataraxos. The code and the evaluation games are public and the training run is under $8,000, which puts replication inside a university budget for the first time in this line of work. The comparison that matters most, Ataraxos against DeepNash, is permanently unavailable.
- Whether arXiv publishes an appeals or exemption route. Two submissions a month across all categories is a hard constraint on large collaborations, and the policy is explicitly a stopgap. The enforcement detail that will decide its real effect is whether withdrawn or replaced submissions count.
- Whether
/v1/systemoneor/v1/decisionswins. Two serving engines shipped incompatible names for the same capability in 32 hours. The vendors have not aligned on one, and the second, broader decision-readout pull request against llama.cpp has been open since 23 September. - Whether Cloudflare publishes an ECE and a Brier score for Clef, asked here yesterday, and whether Clef or Intern-Decision is submitted to a sealed board at all — asked on the 25th, 27th, 29th, 30th, 1st and 2nd. Neither has moved.