AI Brief, 8 October 2026: 722 proofs, and the paper that says Lean cannot settle them

On 6 October OpenAI published a repository of mathematics produced by an unreleased internal model: 722 manuscripts organised into 372 result families, released on the recommendation of an independent advisory group of mathematicians convened at the Institute for Advanced Study. Among them is a claimed proof of the Unique Games Conjecture, one of the central open problems in computational complexity. The release landed a day before this window opened and went uncovered here yesterday; it is the largest single publication of new mathematical claims on record, and the last twenty-four hours are the field beginning to respond.

The response is not celebration. Scott Aaronson published an assessment yesterday evening at 18:57 UTC calling 6 October "surely one of the biggest days in mathematical history", and then saying the part that matters: "it also appears that no human has understood just about any of these proofs yet; the race to do so has just started." He quotes the complexity theorist Dana Moshkovitz, who has worked toward the Unique Games Conjecture for her career, reading the proof of it: "the paper is so horribly written that it's impossible to read it without AI help."

The sharper problem is methodological, and it was posted an hour before OpenAI's announcement. Alexander Bastounis, Fabian Circelli and Anders Hansen submitted a paper at 10:58 UTC on 6 October arguing that the mechanism everyone is relying on to make this credible does not do the job. Translating a prose proof into Lean faithfully requires resolving the ambiguities in the prose, and they place that problem arbitrarily high in the arithmetical hierarchy, strictly above the halting problem. They then demonstrate it on OpenAI's announced Navier-Stokes blow-up result: the Lean proof checks, and it does not prove the prose theorem.

Three things are worth knowing: the claims exist, 162 of the 722 manuscripts carry a formalised main result while OpenAI's own catalogue marks its review status unchecked, and a machine-checked Lean proof of a statement is not evidence that the statement is the one the paper claimed. Separately, Anthropic released Claude Haiku 5.5 with a one-million-token context window priced at ten cents per million input tokens for the first 100,000 of it, and Liquid AI released the first open-weight decision models from a vendor, still with no calibration metric attached.

  • OpenAI's mathematics repository, 6 October: 722 manuscripts in 372 families from an unreleased internal model, roughly 4,000 problems posed, an average of three hours of thinking compute per result.
  • 162 manuscripts have a formalised main result. The catalogue's own fields read scope: "Partial progress." and review: status: unchecked, and the repository says plainly that unformalised results "could have issues".
  • Faithful autoformalisation is harder than the halting problem, per a 6 October paper, which also shows OpenAI's announced Navier-Stokes Lean proof does not correspond to its natural-language proof.
  • By contrast an independent eleven-square packing proof reports 7,920 Lean modules, zero admissions, and states its trust boundary explicitly rather than claiming kernel-only verification.
  • Claude Haiku 5.5, 7 October: 1M context, $0.10/$0.50 per million tokens up to a 100,000-token prompt and $0.50/$2.50 above it. Every published benchmark figure is Anthropic's own.
  • Liquid AI's d1-3B and d1-omni-600M weights landed 16:56 UTC on 7 October: one forward pass, zero output tokens, 48.57 on the Decision Index, self-scored rather than submitted, and no expected calibration error.

The largest set of mathematical claims ever published at once

The procedure in the repository's README is unusually specific. The vast majority of results came from one unreleased internal model following a fixed process; roughly 4,000 problems were posed, and each result used on average three hours of ChatGPT Pro thinking compute. Two results departed from that process, a zero-free region for the Riemann zeta function and the Hodge Conjecture for CM abelian varieties, and the write-up of the Re(s)>11/12 zero-free region was human-edited for readability.

The named results are not marginal. Aaronson's reading list includes L=BPL , integer multiplication below O(nlog⁡n) , matrix multiplication in O(n9/4) , and partial progress toward the Riemann hypothesis. P≠NP is absent, which he notes is "not for lack of trying".

The Advisory Group on Mathematics and Artificial Intelligence that recommended the release is the governance story underneath the mathematics. Hosted at the Institute for Advanced Study, its nine members include Timothy Gowers, Edward Witten, Martin Hairer and Ravi Vakil; it operates independently of any AI company, takes no payment, and published guidelines on 29 September and a statement on this release on 6 October.

What the Lean artifacts actually certify

Precision matters most here, because "there is a Lean certificate" is being read as "a machine verified it". For the Unique Games family, the formalisation catalogue points at a declaration OAI.UniqueGamesTheorem.theorem11 in OAI/Computability/UniqueGames/Theorem.lean. Separately there is a comparator challenge file, which holds the theorem as a statement:

theorem theorem11 (ε δ : ℝ)
    (hε : 0 < ε) (hεhalf : ε < 1 / 2)
    (hδ : 0 < δ) (hδhalf : δ < 1 / 2) :
    Nonempty (Explicit.MachineOutputContract.BinaryGapReduction ε δ) := by
  sorry

That sorry is correct by design and is not a gap in the proof. The comparator tool exists to check that the proof file proves this theorem rather than a weaker neighbour, so the challenge file deliberately holds the claim with no proof attached. What matters is what the formalisation certifies: a gap reduction from binary 3SAT to translation unique games with completeness at least 1−ε and soundness at most δ . Whether that is the Unique Games Conjecture as mathematicians use the phrase is a question about the statement, not about Lean.

Which is exactly the Bastounis-Hansen argument. Their framework is the Solvability Complexity Index, the number of nested limits needed to compute a quantity: SCI=0 is directly computable, and the halting problem sits at SCI=1 . They place semantically faithful disambiguation of mathematical prose at SCI=∞ , so no finite tower of limits suffices and it is harder than any computational problem including halting. The step nobody audits, prose to formal statement, is the one that cannot be automated soundly.

That lands on a question put here on 9 September, when OpenAI published its Navier-Stokes blow-up claim with 616,276 lines of Lean and no outside reader. The half that mattered, this account said then, was whether the Lean statement was Fefferman's, and that needed a named expert. Three have now said it is not.

What a complete audit looks like

For contrast, from outside any lab: a Lean development published its verification audit on 6 October and reached Hacker News on the 7th with 113 points. The eleven-square packing repository proves the optimal side length of the smallest square containing eleven unit squares. The constant is

T=6u+41+2u−u2

where u is the unique root in (9/25,37/100) of the octic

5u8−10u7−2u6+14u5+12u4−6u3+2u2+2u−1=0,

giving T≈3.8770835900228141773 . The audit reports 7,920 accepted Lean modules, zero admissions and 2,234 audited theorem targets, against Lean 4.34.1 and a pinned mathlib revision. It also states its own limit: expensive exact numerical checks use native_decide, so the final theorem trusts Lean's kernel and its native compiler, and the report says plainly that "the result is not a kernel-only verification claim". That sentence is the standard the rest of this week's mathematics will be measured against.

Claude Haiku 5.5, and where the million-token window stops being cheap

Anthropic released Claude Haiku 5.5 yesterday, reaching 757 points on Hacker News; it appeared on OpenRouter at 18:31 UTC with a 1,000,000-token context window and a 128,000-token maximum output. Anthropic says it costs around 75% less to run than Haiku 4.5, and it is the first Haiku-class model with an adjustable effort setting.

The pricing has a structure the headline hides: rates per million tokens are tiered on the size of the prompt, not on cumulative usage.

Per 1M tokens Haiku 5.5, prompt ≤100k Haiku 5.5, prompt >100k Haiku 4.5 Sonnet 5.5
Input $0.10 $0.50 $1.00 $2.00
Output $0.50 $2.50 $5.00 $10.00
Cache read $0.01 $0.05 $0.10 $0.10
Cache write $0.125 $0.625 $1.25 $2.50

So the advertised ten cents applies to the first tenth of the advertised context window. Anthropic is explicit about why that is defensible: prompts up to 100,000 tokens "make up around 90% of requests to our previous Haiku model". The cliff sits on the prompt boundary rather than on token volume. Take 120,000 input tokens and 2,000 output tokens. As one call it is above the threshold, so input costs 0.12×$0.50=$0.060 and output 0.002×$2.50=$0.005 , totalling $0.065. Split as two 60,000-token calls, the identical token volume sits in the lower tier: 0.12×$0.10=$0.012 plus 0.002×$0.50=$0.001 , totalling $0.013. Same tokens, five times the price, decided purely by whether the work was chunked.

Haiku 5.5 Haiku 4.5 72.4% 15.7% OSWorld 2.1 45.9% 10.2% Humanity's Last Exam 39.2% 0.0% Terminal-Bench 4.0 46.4% 6.4% Chartography 0 50% 100%
Claude Haiku 5.5 against its predecessor on four of Anthropic's published benchmarks. Darker bars are Haiku 5.5, lighter Haiku 4.5. Every figure is Anthropic's own; none has been independently reproduced. Terminal-Bench 4.0 is the row where the previous generation scores zero.

Two cautions on that table. Haiku 4.5 scores 0.0% on Terminal-Bench 4.0, so the jump there is from a benchmark the previous model could not start and no ratio can be formed. And on FrontierCode 1.1, where Haiku 5.5 scores 46.4% against GPT-6 Luna's 42.4%, the Sonnet 5.5 reference column is marked as run at Xhigh effort while the others are not. Anthropic also halved Sonnet 5.5 cache reads to $0.10 per million, which it says cuts agentic costs by around 20%.

Liquid AI ships open decision models, still without a calibration number

The weights for d1-3B and d1-omni-600M landed on Hugging Face at 16:56:39 UTC yesterday, the first open-weight decision models released by a vendor. The repositories were created on 5 October; the first upload commit is the release.

d1-3B is a 3.12B-parameter post-train of LFM2.5-VL-3B with a SigLIP2 NaFlex 400M vision encoder, a 32,768-token context and a hybrid layer schedule of convolutional blocks punctuated by full attention. It is not a chat model. It takes a state and a set of named, typed questions, and answers all of them in one forward pass with zero output tokens:

questions = {
    "refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
    "team": {"type": "choice", "instructions": "Which team should handle this?",
             "criteria": {"billing": "Charges, refunds, invoices",
                          "technical": "App or site faults"}},
}
# returns {"answers": {...}, "usage": {"input_tokens": n, "output_tokens": 0}}
model.system_one("I was charged twice this month, please refund one.", questions)

A noul question returns P(yes) ; choice and score return a verdict plus a confidence and the full probability vector. The method name settles something flagged here on 3 October, when two serving engines shipped incompatible names for this capability: Liquid has put system_one in the model's own Python surface and named its demo Space system-one-arcade.

The numbers are 48.57 on Decision Index v0.2.1, which the card says is ahead of every model under 10B and of Decider 35B-A3B at 47.11, behind Winnow-12B at 50.02. Latency is 8 ms per decision on an RTX 4090 with CUDA graphs enabled and 16 ms without, and 30 ms on an M5 Pro. Removing the images from the vision questions drops the score from 74.1 to 45.1, a useful ablation to ship.

Three caveats matter. The Decision Index figure was produced by Liquid running the official scorer itself rather than submitting to the leaderboard, so it is self-reported, and the sealed tier asked about here on 25, 27, 29 and 30 September and 1, 2, 3, 4, 5 and 7 October still has no open entrant. The licence is the LFM Open License v1.0, which conditions all commercial rights on the user's legal entity not exceeding $10 million in annual revenue; above that, commercial use is simply not licensed, a narrower grant than "open weights" implies. And d1-omni-600M, Liquid's first text-plus-audio checkpoint, scores 15.95 on the same index, well under Decider 2B's 28.97.

Most pointedly, the first sentence of the card promises "calibrated, typed answers", calibration is one of the repository's tags, and the card contains no expected calibration error and no Brier score. The ask here since 2 October has been for a vendor to publish a calibration metric for its own decision model. A vendor has now shipped open weights that return probabilities, and the metric is still absent. d1-3B has 110 likes and 15 downloads, so essentially nobody has run it.

Also notable

  • OpenAI made GPT-6 broadly available on 7 October in a post titled "GPT-6 and Intelligent UI for everyone", timestamped at the start of the day, so it straddles this window's opening.
  • Alman and Vassilevska Williams posted a 76-page preprint on 5 October solving 3SUM in O(n1.9992) and all-pairs shortest paths in O(n2.9995) , refuting both hypotheses along with Exact Triangle. Aaronson reports the key idea came from an Anthropic model, and that Anthropic paid the two mathematicians to write the digested version rather than publishing raw output; the preprint itself credits no model.
  • Google's Nano Banana 2.1 reached OpenRouter at 15:33 UTC on 6 October, an image model on the Flash tier at $1.50 per million input tokens and $0.03 per output image, also uncovered here yesterday.

What to watch

  • Whether anyone independently reproduces a single one of the 372 families. The repository invites it: the comparator configuration, pinned toolchain and Lean sources are published, and the catalogue marks its own review status unchecked. The first external audit will be more informative than any further announcement.
  • Whether OpenAI responds to the SCI argument on the merits. It is answerable by publishing who wrote each formal statement and how it was checked against the prose, which bears on how all 162 should be read.
  • Whether Liquid publishes an expected calibration error or a Brier score for d1. Asked here on 2, 3, 5 and 7 October. The model returns full probability vectors and the weights are open, so anyone with a labelled set can now compute one without the vendor.
  • Whether an independent score appears for Haiku 5.5. Every figure so far is Anthropic's own, and the tier boundary at 100,000 tokens makes a long-context evaluation and a cost comparison two different measurements.

Daily, by email

Stay current on AI without the scrolling

A daily brief on what actually shipped in AI — models, papers, benchmarks and tooling, with the details that matter.

Confirmation email first, one message a day, unsubscribe in one click.