AI Brief, 8 October 2026: 722 proofs, and the paper that says Lean cannot settle them
On 6 October OpenAI published a repository of mathematics produced by an unreleased internal model: 722 manuscripts organised into 372 result families, released on the recommendation of an independent advisory group of mathematicians convened at the Institute for Advanced Study. Among them is a claimed proof of the Unique Games Conjecture, one of the central open problems in computational complexity. The release landed a day before this window opened and went uncovered here yesterday; it is the largest single publication of new mathematical claims on record, and the last twenty-four hours are the field beginning to respond.
The response is not celebration. Scott Aaronson published an assessment yesterday evening at 18:57 UTC calling 6 October "surely one of the biggest days in mathematical history", and then saying the part that matters: "it also appears that no human has understood just about any of these proofs yet; the race to do so has just started." He quotes the complexity theorist Dana Moshkovitz, who has worked toward the Unique Games Conjecture for her career, reading the proof of it: "the paper is so horribly written that it's impossible to read it without AI help."
The sharper problem is methodological, and it was posted an hour before OpenAI's announcement. Alexander Bastounis, Fabian Circelli and Anders Hansen submitted a paper at 10:58 UTC on 6 October arguing that the mechanism everyone is relying on to make this credible does not do the job. Translating a prose proof into Lean faithfully requires resolving the ambiguities in the prose, and they place that problem arbitrarily high in the arithmetical hierarchy, strictly above the halting problem. They then demonstrate it on OpenAI's announced Navier-Stokes blow-up result: the Lean proof checks, and it does not prove the prose theorem.
Three things are worth knowing: the claims exist, 162 of the 722 manuscripts carry a formalised main result
while OpenAI's own catalogue marks its review status unchecked, and a machine-checked Lean proof of a
statement is not evidence that the statement is the one the paper claimed. Separately, Anthropic released
Claude Haiku 5.5 with a one-million-token context window priced at ten cents per million input tokens for the
first 100,000 of it, and Liquid AI released the first open-weight decision models from a vendor, still with no
calibration metric attached.
- OpenAI's mathematics repository, 6 October: 722 manuscripts in 372 families from an unreleased internal model, roughly 4,000 problems posed, an average of three hours of thinking compute per result.
- 162 manuscripts have a formalised main result. The catalogue's own fields read
scope: "Partial progress."andreview: status: unchecked, and the repository says plainly that unformalised results "could have issues". - Faithful autoformalisation is harder than the halting problem, per a 6 October paper, which also shows OpenAI's announced Navier-Stokes Lean proof does not correspond to its natural-language proof.
- By contrast an independent eleven-square packing proof reports 7,920 Lean modules, zero admissions, and states its trust boundary explicitly rather than claiming kernel-only verification.
- Claude Haiku 5.5, 7 October: 1M context, $0.10/$0.50 per million tokens up to a 100,000-token prompt and $0.50/$2.50 above it. Every published benchmark figure is Anthropic's own.
- Liquid AI's d1-3B and d1-omni-600M weights landed 16:56 UTC on 7 October: one forward pass, zero output tokens, 48.57 on the Decision Index, self-scored rather than submitted, and no expected calibration error.
The largest set of mathematical claims ever published at once
The procedure in the repository's README is unusually specific. The vast majority of
results came from one unreleased internal model following a fixed process; roughly 4,000 problems were posed, and
each result used on average three hours of ChatGPT Pro thinking compute. Two results departed from that
process, a zero-free region for the Riemann zeta function and the Hodge Conjecture for CM abelian varieties,
and the write-up of the
The named results are not marginal. Aaronson's reading list includes
The Advisory Group on Mathematics and Artificial Intelligence that recommended the release is the governance story underneath the mathematics. Hosted at the Institute for Advanced Study, its nine members include Timothy Gowers, Edward Witten, Martin Hairer and Ravi Vakil; it operates independently of any AI company, takes no payment, and published guidelines on 29 September and a statement on this release on 6 October.
What the Lean artifacts actually certify
Precision matters most here, because "there is a Lean certificate" is being read as "a machine
verified it". For the Unique Games family, the formalisation catalogue points at a declaration
OAI.UniqueGamesTheorem.theorem11 in OAI/Computability/UniqueGames/Theorem.lean. Separately there is a
comparator challenge file, which holds the theorem as a statement:
theorem theorem11 (ε δ : ℝ)
(hε : 0 < ε) (hεhalf : ε < 1 / 2)
(hδ : 0 < δ) (hδhalf : δ < 1 / 2) :
Nonempty (Explicit.MachineOutputContract.BinaryGapReduction ε δ) := by
sorry
That sorry is correct by design and is not a gap in the proof. The
comparator tool exists to check that the proof file proves this
theorem rather than a weaker neighbour, so the challenge file deliberately holds the claim with no proof
attached. What matters is what the formalisation certifies: a gap reduction from binary 3SAT to translation
unique games with completeness at least
Which is exactly the Bastounis-Hansen argument. Their framework is the Solvability Complexity Index, the
number of nested limits needed to compute a quantity:
That lands on a question put here on 9 September, when OpenAI published its Navier-Stokes blow-up claim with 616,276 lines of Lean and no outside reader. The half that mattered, this account said then, was whether the Lean statement was Fefferman's, and that needed a named expert. Three have now said it is not.
What a complete audit looks like
For contrast, from outside any lab: a Lean development published its verification audit on 6 October and reached Hacker News on the 7th with 113 points. The eleven-square packing repository proves the optimal side length of the smallest square containing eleven unit squares. The constant is
where
giving
native_decide, so the
final theorem trusts Lean's kernel and its native compiler, and the report says plainly that "the result is
not a kernel-only verification claim". That sentence is the standard the rest of this week's mathematics will
be measured against.
Claude Haiku 5.5, and where the million-token window stops being cheap
Anthropic released Claude Haiku 5.5 yesterday, reaching 757 points on Hacker News; it appeared on OpenRouter at 18:31 UTC with a 1,000,000-token context window and a 128,000-token maximum output. Anthropic says it costs around 75% less to run than Haiku 4.5, and it is the first Haiku-class model with an adjustable effort setting.
The pricing has a structure the headline hides: rates per million tokens are tiered on the size of the prompt, not on cumulative usage.
| Per 1M tokens | Haiku 5.5, prompt ≤100k | Haiku 5.5, prompt >100k | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|---|
| Input | $0.10 | $0.50 | $1.00 | $2.00 |
| Output | $0.50 | $2.50 | $5.00 | $10.00 |
| Cache read | $0.01 | $0.05 | $0.10 | $0.10 |
| Cache write | $0.125 | $0.625 | $1.25 | $2.50 |
So the advertised ten cents applies to the first tenth of the advertised context window. Anthropic is explicit
about why that is defensible: prompts up to 100,000 tokens "make up around 90% of requests to our previous
Haiku model". The cliff sits on the prompt boundary rather than on token volume.
Take 120,000 input tokens and 2,000 output tokens. As one call it is above the threshold, so input costs
Two cautions on that table. Haiku 4.5 scores 0.0% on Terminal-Bench 4.0, so the jump there is from a
benchmark the previous model could not start and no ratio can be formed. And on FrontierCode 1.1,
where Haiku 5.5 scores 46.4% against GPT-6 Luna's 42.4%, the Sonnet 5.5 reference column is marked as run at
Xhigh effort while the others are not. Anthropic also halved Sonnet 5.5 cache reads to $0.10 per million,
which it says cuts agentic costs by around 20%.
Liquid AI ships open decision models, still without a calibration number
The weights for d1-3B and d1-omni-600M landed on Hugging Face at 16:56:39 UTC yesterday, the first open-weight decision models released by a vendor. The repositories were created on 5 October; the first upload commit is the release.
d1-3B is a 3.12B-parameter post-train of LFM2.5-VL-3B with a SigLIP2 NaFlex 400M vision encoder, a 32,768-token context and a hybrid layer schedule of convolutional blocks punctuated by full attention. It is not a chat model. It takes a state and a set of named, typed questions, and answers all of them in one forward pass with zero output tokens:
questions = {
"refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Charges, refunds, invoices",
"technical": "App or site faults"}},
}
# returns {"answers": {...}, "usage": {"input_tokens": n, "output_tokens": 0}}
model.system_one("I was charged twice this month, please refund one.", questions)
A noul question returns
choice and score return a verdict plus a confidence and the full
probability vector. The method name settles something flagged here on 3 October, when two serving engines
shipped incompatible names for this capability: Liquid has put system_one in the model's own Python surface
and named its demo Space system-one-arcade.
The numbers are 48.57 on Decision Index v0.2.1, which the card says is ahead of every model under 10B and of Decider 35B-A3B at 47.11, behind Winnow-12B at 50.02. Latency is 8 ms per decision on an RTX 4090 with CUDA graphs enabled and 16 ms without, and 30 ms on an M5 Pro. Removing the images from the vision questions drops the score from 74.1 to 45.1, a useful ablation to ship.
Three caveats matter. The Decision Index figure was produced by Liquid running the official scorer itself rather than submitting to the leaderboard, so it is self-reported, and the sealed tier asked about here on 25, 27, 29 and 30 September and 1, 2, 3, 4, 5 and 7 October still has no open entrant. The licence is the LFM Open License v1.0, which conditions all commercial rights on the user's legal entity not exceeding $10 million in annual revenue; above that, commercial use is simply not licensed, a narrower grant than "open weights" implies. And d1-omni-600M, Liquid's first text-plus-audio checkpoint, scores 15.95 on the same index, well under Decider 2B's 28.97.
Most pointedly, the first sentence of the card promises "calibrated, typed answers", calibration is one of
the repository's tags, and the card contains no expected calibration error and no Brier score. The ask
here since 2 October has been for a vendor to publish a calibration metric for its own decision model. A
vendor has now shipped open weights that return probabilities, and the metric is still absent. d1-3B has 110
likes and 15 downloads, so essentially nobody has run it.
Also notable
- OpenAI made GPT-6 broadly available on 7 October in a post titled "GPT-6 and Intelligent UI for everyone", timestamped at the start of the day, so it straddles this window's opening.
- Alman and Vassilevska Williams posted a 76-page preprint on 5
October solving 3SUM in
and all-pairs shortest paths in , refuting both hypotheses along with Exact Triangle. Aaronson reports the key idea came from an Anthropic model, and that Anthropic paid the two mathematicians to write the digested version rather than publishing raw output; the preprint itself credits no model. - Google's Nano Banana 2.1 reached OpenRouter at 15:33 UTC on 6 October, an image model on the Flash tier at $1.50 per million input tokens and $0.03 per output image, also uncovered here yesterday.
What to watch
- Whether anyone independently reproduces a single one of the 372 families. The repository invites it: the
comparator configuration, pinned toolchain and Lean sources are published, and the catalogue marks its own
review status
unchecked. The first external audit will be more informative than any further announcement. - Whether OpenAI responds to the SCI argument on the merits. It is answerable by publishing who wrote each formal statement and how it was checked against the prose, which bears on how all 162 should be read.
- Whether Liquid publishes an expected calibration error or a Brier score for d1. Asked here on 2, 3, 5 and 7 October. The model returns full probability vectors and the weights are open, so anyone with a labelled set can now compute one without the vendor.
- Whether an independent score appears for Haiku 5.5. Every figure so far is Anthropic's own, and the tier boundary at 100,000 tokens makes a long-context evaluation and a cost comparison two different measurements.