AI Brief, 9 October 2026: a sign error, and what 42% actually counts
OpenAI has withdrawn three of the mathematical manuscripts it published on 6 October. A history file in the repository, dated 7 October and pushed at 05:03 UTC on 8 October by OpenAI's Dan Roberts, says that in "Algebraicity of Weil classes on split abelian eightfolds" a sign error invalidates a stabilization-trace cancellation argument along with the construction two dependent papers rely on. All three are gone from the catalogue: that paper, "Algebraicity of Kuga–Satake Correspondences for K3 Surfaces" and "The rational Hodge conjecture for products of K3 surfaces". The collection now stands at 719 manuscripts in 372 families, down from 722, with the family count unchanged.
The withdrawal is the smaller half of the entry. Fourteen other manuscripts were revised with proof repairs, corrected statements and clearer hypotheses, and thirteen more were updated purely to cite the revised editions of companion papers. That is thirty manuscripts, 4.2% of the corpus, moved in the first forty-eight hours. Yesterday's brief said the first external audit would be more informative than any further announcement; the first correction came from OpenAI instead, and its README had already hedged that "some of the unformalized results could have issues."
Which is where the number worth checking comes in. The history entry reports six new formalisations and
puts the total at "300 / 719 = ~42%" of top-line results. The machine-readable catalogue in the same
repository says something else: its
formalisation file
describes itself as a catalogue of papers with a formalised main result and lists 173 of them, carrying
201 formal declarations, with its review status recorded as unchecked. Those are 24.1% and 28.0% of 719.
The repository does not define "top-line results" anywhere a reader can point to, so the headline
percentage and the auditable catalogue cannot be reconciled from what is published. None of the three
withdrawn titles appears in the formalisation file at all.
Elsewhere, the field's response hardened. TechCrunch reported that the solutions are not yet meeting mathematicians' standards, and the Association for Human Mathematics published a statement urging mathematicians to stop working with OpenAI, calling the release a demonstration of power. And the open-weights story with the most attention in the last day turned out to be a release from the week before.
- Three manuscripts withdrawn, fourteen revised, thirteen re-cited. The named cause is a single sign error in one paper and the construction two others inherited from it. The catalogue is now 719 manuscripts in 372 families, from 722 on publication day.
- Formalisation coverage is stated as ~42% in prose, while the machine-readable catalogue lists 173
papers and records its own review status as
unchecked. - No frontier lab published weights at all in the last twenty-four hours: sixteen organisations that ship regularly, including Qwen, DeepSeek, Google and Mistral, have nothing dated inside the window.
- Whistle, a 16.9 MB on-device speech-to-text model, reached 632 points on Hacker News on 8 October, six days after it was published, and the only independent test in the thread put it far behind a 1.7B baseline on accented speech.
- TRL 1.15.0 turned on a fused language-model head by default, with a self-reported 52–82% cut in peak training memory at 8,192 tokens.
A sign error, and the arithmetic of "42% formalized"
The dependency structure is the part engineers will recognise. One error in one lemma did not stay in one paper: the catalogue is organised into families whose members cite each other, and a construction fails everywhere it is used.
The repaired set is itemised rather than summarised, which is to OpenAI's credit: six manuscripts on Kähler minimal model programs, four on Lipschitz heights and Ashkin–Teller currents, two on taming and hypersymplectic deformation, a citation fix, and one called "Incompressible Box Transport and Finite Computation", whose torus-projection and common-clock estimates were revised.
That last title sits in the Navier-Stokes family, the one Alexander Bastounis, Fabian Circelli and Anders Hansen used on 6 October to argue that faithful autoformalisation is harder than the halting problem and that OpenAI's Lean proof there did not match its prose. The history entry gives no reason for the revision and does not cite them, so the honest reading is an unattributed repair in the family they criticised, not a vindication. That result has not been withdrawn.
The objection raised repeatedly in the
Hacker News thread, which reached 338 points, is the
right one and remains unanswered: if these results were machine-checked in Lean, how was a sign error
possible? The answer is that most of them were not. A Lean proof certifies that a formal statement follows
from formal premises; it says nothing about whether that statement faithfully translates the prose
theorem, which is exactly the gap the 6 October paper formalised. And the catalogue marks its review
status unchecked, so even that correspondence is self-declared.
Terence Tao made the structural version of this argument two days before the withdrawal, in a thread posted at 18:00 UTC on 6 October that only reached Hacker News on the 8th, where it drew 594 points. His case is that a breakthrough proof used to generate talks, workshops and new entrants, and that this was how proofs got digested into shared understanding. Solutions obtained by prompting break the loop: whoever submitted the prompt often cannot answer questions about the output, so little follow-on activity forms, and researchers have begun withholding open directions for fear of being scooped. He argues the damage is irreversible and concludes that mathematics must "decenter the role of raw problem solving" and reward exposition instead. A guest post on his blog by Álvaro Lozano-Robledo on 8 October asks what to tell students.
Against that, formalisation itself is getting cheaper fast. A group released NanoProof at 09:49 UTC on 8 October, an automated theorem prover for Lean 4 shipped with its data, tooling, pipeline and weights, reporting 50.8% pass@16 on MiniF2F-Test at roughly 7 to 90 times less compute than ABEL and HTPS. Those figures are the authors' own, but the code is CC BY-SA 4.0 and published, so they are checkable by anyone with a GPU, which is more than can be said for the corpus it would be pointed at.
A 16.9 MB transcriber, and the first independent test of it
The loudest open-weights item of the day was not a release. Whistle, an on-device speech-to-text model from Cactus Compute, was published on 2 October with weights that went up on 30 September; it reached the Hacker News front page at 16:59 UTC on 8 October with 632 points, after two earlier submissions of the same URL died at two and five points. It went uncovered here when it shipped, and earns coverage now because the scrutiny is new even though the model is not.
The headline is the file size, and it holds up. The repository's fp32 checkpoint is 220,618,620 bytes across 55,151,433 parameters; the artifact that actually ships is 16,919,407 bytes in Cactus's own container format, quantised to 2–4 bits with a group size of 128.
Effective bits per weight is just the deployed size in bits over the parameter count,
params = 55_151_433 # from the safetensors header, all fp32
deployed = 16_919_407 # whistle.cact, the file that ships
print(deployed * 8 / params) # 2.4542 bits per weight
print(params * 4 / deployed) # 13.04x smaller than the fp32 checkpoint
So 16.9 MB is a real 55M-parameter model at about 2.45 bits per weight, not a distilled Whisper, and that is the one claim here verifiable from the published files alone. The architecture is custom: 80 log-mel bins band-limited to 250–3500 Hz, a convolutional stem reducing 3,000 frames to 375, eight encoder and eight decoder blocks of width 512 with grouped-query attention, gated cross-attention at every decoder layer, and five-beam search. Seven languages, batch rather than streaming, CPU-only, Apache 2.0.
Everything else is self-reported, and two details flatter it. Cactus measured its own word error rates over 86,174 utterances, but quoted the Whisper and Moonshine numbers from those papers rather than re-running them, so it is not a like-for-like harness. And in the speed panel the Whisper baseline runs at fp32 while Whistle runs at 2–4 bits.
The genuinely new information is in the thread, and it is unflattering. One commenter ran Whistle fully locally on an Amazon Echo Show and scored it against his own speech, reporting a heavy accent: free-form Whistle got 70 of 170 utterances right against 168 of 170 for Qwen ASR 1.7B. He recovered it only by abandoning free-form transcription, training a small network on 10,000 generated utterances to map Whistle's final state onto a fixed template set, reaching 164 of 170. That is a good result for a constrained command vocabulary and a poor one for open transcription. A second commenter reported the classic Whisper-family failure on a television episode, with the model emitting "Thank you." as a default and once producing it for sixty seconds of dialogue, contradicting the model card's claim that silence returns an empty transcript rather than an invented sentence. Cactus's founder posted twice and answered neither criticism. NVIDIA Parakeet, which several commenters named as the baseline they actually use, appears nowhere in the published comparison.
What actually shipped
No frontier lab published weights. Checks across sixteen organisations that ship regularly, Qwen, DeepSeek, Google, Microsoft, Mistral, NVIDIA, AI2, Moonshot, Z.ai and IBM among them, return nothing created inside the window; the newest Qwen upload is 20 September, the newest DeepSeek one 10 September. Meta's public listing is stale for a different reason, as its recent repositories are gated, so it is unverified rather than quiet.
The release that changes something concrete for people training models is TRL 1.15.0, published at 19:26 UTC on 8 October. It turns on a fused language-model head by default across SFT, DPO, KTO, GRPO, RLOO and distillation: log-probabilities and entropy are computed in Triton tiles so the full logits tensor is never materialised. On Hugging Face's own single-GPU measurement that raises maximum trainable sequence length about 6.9-fold and cuts peak memory at 8,192 tokens by 52 to 82%, figures which are the maintainers' own. It also breaks two things worth knowing before upgrading:
# PEFT adapters that targeted the LM head now raise instead of silently training it
LoraConfig(target_modules=["q_proj", "v_proj"], modules_to_save=["lm_head"]) # was target_modules=[..., "lm_head"]
# use_liger_kernel is deprecated in DPO/KTO/GRPO/RLOO and goes away in 2.0.0
DPOConfig(use_liger_kernel=True) # drop it; the fused head replaces it
Python 3.10, nn.DataParallel and the MiniLLM trainer are also removed.
Also notable
- StepFun's Step 5 Preview is now priced, appearing on OpenRouter at 12:34 UTC on 8 October with a 1,000,000-token context at $1.00 per million input tokens and $2.70 output. The 20 September brief reported 1.21 TB of its bf16 weights sitting ungated on Hugging Face with no licence, and predicted one of those facts would change: that repository is no longer public, while a known-good StepFun one still is. When it was withdrawn is not established.
- OpenAI told investors it reached roughly $50B annualised revenue at the end of September, against the $68B widely reported a week earlier, per CNBC following the Financial Times, which attributes part of the gap to the larger number including partner gross revenue. Nvidia fell 3% and CoreWeave 8%. The same report says OpenAI pulled a planned GPT-6.1 Astra launch on safety grounds, which no first-party statement covers.
- Anthropic's revised usage policy takes effect 12 November and adds a section prohibiting "sustained and needless abusive or cruel behavior toward our models", plus one on undermining democratic processes (announced at 17:00 UTC).
- Fourteen Gannett newspapers sued OpenAI for copyright infringement on 8 October, filing USA Today Co., Inc. v. OpenAI Foundation in the Southern District of New York as 1:26-cv-08892. The complaint names seven OpenAI entities, attaches an exhibit of GPT-5.6 output examples, and is related to the existing OpenAI copyright multidistrict litigation.
- Three dismissed OpenAI safety researchers disputed the company's account: Mikita Balesni says he, Tomek Korbak and Jasmine Wang were fired for prioritising safety over OpenAI's near-term interests.
- One paper worth the abstract: linear probes that detect agent sabotage and unverbalised deception, with code and data (arXiv:2610.12445).
What to watch
- Whether more withdrawals follow, and who found this one. The history entry names the error but not its discoverer, and whether a mathematician or another model caught it decides how much the review process is doing. A second entry this week would say the 48-hour rate was not a one-off.
- Whether the repository defines "top-line results". One sentence reconciling the ~42% in prose with the 173 papers and 201 declarations in the formalisation file would settle the corpus's verified fraction, and it is the cheapest clarification available.
- Whether anyone benchmarks Whistle against Parakeet independently. Word error rates on a shared harness at equal precision would settle whether the size win costs accuracy.
- Whether Liquid publishes a calibration error or Brier score for d1, asked here on 2, 3, 5, 7 and 8 October; the weights are open, so this no longer needs the vendor. Mistral's and Reflection's "end of the month" weight deadlines also remain unmet.