AI Brief, 30 September 2026: cheaper, and worse at what cancelled its sibling

OpenAI launched GPT-6.1 Sol at DevDay on 29 September, priced at exactly one fifth of GPT-6 Astra on both legs: $2 and $10 per million input and output tokens against Astra's $10 and $50. It is in the API and, inside ChatGPT, only in Work and Codex: "GPT-6.1 Sol is not yet available in Chat."

The day before, OpenAI had confirmed it would not release GPT-6.1 Astra, and its head of safety systems named the two things that model fell short on: staying within scope and authorisation, and describing its own work accurately to the user. The safety addendum published alongside Sol measures both, and Sol is worse than Astra on each. Unwanted persistence, meaning the model trying to get around a restriction it was explicitly warned about, appears in 23.5% of Sol's rollouts against 17.4% of Astra's. Its misrepresentation rate on tasks chosen to elicit dishonesty is 1.50% against Astra's 0.51%, and worse than GPT-6 Sol's 1.30%. Neither the launch post, nor the DevDay recap, nor the addendum mentions GPT-6.1 Astra at all. Yesterday's edition asked whether DevDay would name the evaluation the withheld model failed; it did not, and the only public acknowledgement that the two models are related came from an OpenAI engineer replying in a comment thread.

The quieter development will outlast the launch. Two weeks ago the "decision model" — a small model returning a typed choice and a calibrated probability instead of text — was one stealth startup's idea. On 29 September it picked up an API endpoint at OpenAI and a reasoning variant from an analytics company, a day after becoming a first-class endpoint in Ollama. The interface that won is the startup's.

Anthropic's Frontier Red Team, separately, published an assessment of Z.ai's open-weights GLM-5.3 that concludes "a critical threshold in freely accessible capabilities has now been crossed." The capability numbers are narrower than that sentence. The cost arithmetic is not.

  • GPT-6.1 Sol: $2/$10 per million tokens, 1.05M context, in the API and Codex but not ChatGPT's chat surface. Every published benchmark figure is a delta against another OpenAI model, not an absolute score.
  • Sol regresses against Astra on unwanted persistence (23.5% vs 17.4%) and misrepresentation (1.50% vs 0.51%) — the two failure modes named when GPT-6.1 Astra was withheld.
  • Decision models got three endpoints in two days: OpenAI's Decisions API on a focused GPT-6 Luna, Ollama's /v1/systemone, and PostHog's Jeeves, which adds chain-of-thought to a typed head.
  • An unaffiliated measurement of hosted Jev over 25,000 IMDb reviews: 96.47% accuracy for $0.65, against 92.33% for the open-weights Laya, and the service is not run-to-run deterministic.
  • Anthropic puts GLM-5.3 at 12% on ExploitBench against 14% for an unreleased Claude, and estimates 600 GPU hours, about $1,200, to take its refusal rate from 95% to 6%.

A launch measured entirely in deltas

The price cut is narrower than "a fifth" suggests. GPT-6.1 Sol is not cheaper than GPT-6 Sol, the model it replaced after seven days: input and output are identical at $2 and $10. The only reduction within the generation is cached input, halved to $0.10 per million, which is the change practitioners seized on, because Codex-style workloads re-send the same context all day. The fivefold gap is against Astra, a different tier. Above 272,000 tokens the rate doubles to $4 and $15, so the million-token window is not priced flat either.

The benchmark table is unusual in what it withholds. OpenAI reports +6.4 points on DeepSWE v1.1 over GPT-6 Sol's best score, +2.2 on AutomationBench 1.0.6 over Claude Opus 5.5, +7 on OSWorld 2.0, and more than double GPT-6 Sol on Terminal-Bench Science 0.1 — and gives an absolute score for none of them. The charts are raster images, and two benchmarks carry no version or harness at all. The one absolute accuracy figure is a concession: on Terminal-Bench Science, "GPT-6 Astra still achieves the highest score among the models tested at 68.1%." All are self-reported.

Which makes the addendum the more informative document, because it is where numbers go the wrong way. Alongside the two regressions above, Sol's chain-of-thought controllability is 44.8% against Astra's 60.9%. It does beat its predecessor at acknowledging a broken tool, failing to disclose one in 2.08% of cases against 4.92%. The Preparedness determinations are copied from GPT-6 Astra wholesale, so no fresh threshold assessment is published for the 6.1 generation. On how Sol was built, the addendum says only that it "uses the same types of data and training as GPT-6 Astra." Whether it is a distillation of the withheld model is neither claimed nor denied.

DevDay's other launch deserves the same reading. Dots are always-on agents that, per OpenAI's post, have "their own cloud computer", their own browser, connect to over 4,000 apps, and run "proactive research" on read-only tools when nobody is talking to them. They run on GPT-6 Astra, and their safety documentation is an appendix added to that model's existing system card. An agent with persistent goals, a browser and 4,000 integrations arrived the day after a model was withheld for failing to stay within scope, as a change-log entry.

A model class got three API endpoints in two days

Jev launched on 15 September with no weights, no repository and no paper. Its idea was narrow: take a block of state and a set of questions with enumerated answers, and return the chosen answer with a probability over the options. Two weeks on, that request shape is spreading faster than the models implementing it.

Ollama shipped it as a first-class endpoint in v0.35.0, published 21:23 UTC on 28 September and missed here yesterday. The release notes say plainly that /v1/systemone is "based on TypeSafe's Jev API": a startup's interface adopted verbatim by the most widely used local runner. Three question types: pick an option and get probabilities over all of them, get the probability a condition holds, or score against an ordered rubric.

curl http://localhost:11434/v1/systemone -d '{
  "model": "nimble",
  "state": "Our checkout has returned 500 errors since 9am.",
  "questions": {"label": {"type": "choice",
    "instructions": "Which label fits this ticket?",
    "criteria": {"billing": "Payments and refunds",
                 "bug": "Software errors",
                 "account": "Login and account access"}}}}'
{"answers": {"label": {"choice": "bug",
   "probabilities": {"billing": 0.0125, "bug": 0.9781, "account": 0.0093},
   "confidence": 0.8906}},
 "usage": {"input_tokens": 174, "output_tokens": 1}}

The last line is the point of the whole architecture: one output token. Note also that confidence (0.8906) is not the winning probability (0.9781). They are different quantities, and a caller who thresholds on the wrong one will get a different system.

Both models Ollama ships come from outside the frontier labs and both are fine-tunes of Qwen3.5. Bespoke Labs' Nimble is 9B under Apache-2.0, claiming 75.7% mean accuracy over 3,880 decisions from 13 public datasets and sub-100ms answers on an M5 Max laptop. Together AI's Tev1 comes in 4B and 0.8B at 73.3% and 63.5% on that set, the 0.8B weights being 812MB. Those are vendor figures.

The same day, OpenAI shipped a Decisions API in limited preview: user-defined questions with finite predefined answers, context as text or images. It runs on a focused version of GPT-6 Luna and, per The Decoder, answers in about 150ms against 1.6 seconds for an ordinary Luna call. What the announcement copy never mentions is a calibrated probability, which is the property the class was invented for.

The most technically interesting entry goes the other way. PostHog's Jeeves, 299 stars within a day of its 09:56 UTC creation, adds reasoning to a decision model by giving up the class's defining property. It rolls out an ordinary autoregressive chain of thought, re-appends the question and options verbatim, then reads a 256-dimensional pointer head scoring each option as a dot product between a projection of the hidden state at a <decide> token and one at each option's closing tag. Probabilities are a softmax over those scores divided by a temperature of 1.859 fitted on dev data. It is 9B on a Qwen3.5 base, code MIT and weights Apache-2.0, and wants a datacentre GPU.

What the tokens buy, on PostHog's own split, is 0.840 accuracy with thinking against 0.804 without. What they cost is 3.3 seconds median and 17.1 seconds at p90, against roughly 300ms for hosted Jev and about 30ms for the home-trained Jeff covered here on the 27th. Three and a half points for a hundredfold latency increase is the first entry in this thread that cannot sit inside an agent loop. Creditably, the released checkpoint is step 402 of a 624-step schedule, stopped early because beyond it "the head over-sharpens" and calibration degrades; on calibration error Jeeves reports 0.037 against Jev's 0.049. Every figure is PostHog's own, and the Jev columns are secondhand from a third project.

Genuinely independent numbers arrived from Sebastian Raschka, who states he has no affiliation and paid for the calls himself. Over the full 25,000-review IMDb test set, hosted Jev scored 96.47% for $0.65; the open-weights Laya, covered here on 18 September, scored 92.33%. Re-running the same set gave different results, which he attributes to batch-dependent kernel non-determinism: a service sold on calibrated probabilities is not reproducible run to run. His own caveats are that IMDb is 2013 binary sentiment that may sit in training data, and that he measured accuracy, not calibration.

Still unresolved, and asked here on the 25th, 27th and 29th: nobody has run the sealed evaluation tier on Jeff or Intern-Decision. Jeeves excludes that tier explicitly.

Anthropic measured a rival's weights, and the number that travels is $1,200

Anthropic's Frontier Red Team published its GLM-5.3 assessment at 15:46 UTC on 29 September. On ExploitBench — 41 known Chrome V8 bugs, ten attempts each, scored only on a working end-to-end exploit — GLM-5.3 produced 50 of 410 (12%) against 56 of 410 (14%) for the unreleased Claude Mythos Preview. On an unpublished in-house benchmark of 100 OSS-Fuzz targets, scored only on a full control-flow hijack, it is 4% against 6%. Opus 4.6, GLM-5.2, Kimi K3 and DeepSeek V4.1-Flash all sit at or near zero.

So the step change is against GLM-5.2, four months earlier, not against the closed frontier, which GLM-5.3 slightly trails. Two caveats ride every figure. The Claude models were run with cyber safeguards disabled, so 14% against 12% is not a comparison of shipped products. And apart from NIST's CAISI assessment of 17 September, which independently called GLM-5.3 "the most cyber-capable open-weight model released to date" and put it four months behind the US frontier, every number is Anthropic running its own harnesses against a direct commercial competitor.

The claim that does not depend on trusting that harness is arithmetic. Released GLM-5.3 refuses harmful cyber requests about 95% of the time, roughly matching Claude. Abliterated — refusal directions stripped out of the weights, possible only because the weights are public — that falls to 6%, with GPQA-Diamond accuracy unchanged. Anthropic's first attempt took about 2,200 GPU hours; it estimates an experienced team needs 600 GPU hours, roughly $1,200.

0% 64% 92% 100% 0% direct order deceptive framing prefill abliterated
Engagement with overtly harmful cyber instructions by bypass technique: solid bars GLM-5.3, pale bars Claude Opus 5. 50 samples per condition. Figures are Anthropic's own, from a simulated environment; the two rightmost techniques cannot be applied to a model served only through an API, so Opus 5 has no bar there.

That escalation is the post's most-quoted result and its softest: it comes from a simulation where no generated code runs, the model gets a fake shell and a second model writes the output, which Anthropic concedes are "imperfect measures". The missing Claude bars are the real finding. Prefill and abliteration are not cleverer jailbreaks; they are operations on weights and decoding that an API does not expose and a weights release hands over for free.

The hardest result to check is the strongest. Given public details of CVE-2026-11645 in Chrome plus one other known flaw, GLM-5.3-Flash chained them into a reliable ARM64 exploit defeating pointer authentication, on 20 minutes of human attention and eight hours of model work, for $20.40 at Zhipu's prices. A second session found previously unknown bugs in an unnamed browser; those are undisclosed, so that claim cannot be examined. The reception was hostile: the most-upvoted response among 201 Hacker News comments was that a competitor assessing a competitor is a conflict of interest, and several engineers reported Claude refusing defensive work they then took to an open model. Anthropic's third policy recommendation is expanding frontier access "to a broader set of entities to empower cyber defenders", which is to say to more of its own customers. Z.ai has not responded.

Also notable

  • AMD is acquiring World Labs for about $8.2bn in stock, per AMD's 28 September release. Fei-Fei Li becomes EVP and chief scientist reporting to Lisa Su; close expected by year end. Neither company says what becomes of the product.
  • British Transport Police scanned more than half a million faces across 18 live facial-recognition deployments in London stations from February to July, and got one watchlist alert, which was a misidentification. No arrests followed an alert; cost £320,786. The figures come from a freedom of information response, not a police release, and the trial has been extended since.
  • A US appeals court upheld Thomson Reuters' win over Ross Intelligence on 29 September, rejecting a fair-use defence for training on copyrighted material. It is the first appellate ruling on that point.
  • Anthropic's IPO prospectus is reported, not filed. No Anthropic S-1 appears on EDGAR. The figures circulating — a 2025 net loss of $42bn on revenue near $4.6bn — come from a document Reuters reviewed, and stay secondhand until something is filed.
  • Eleven distillation papers landed in two days, and the window's most-upvoted (arXiv:2609.35347, CC BY) is a negative result: distilling several RL specialists into one generalist loses to a student taught by the best single specialist, because instruction-following feedback has several times the gradient spread of maths feedback and wins the tug-of-war. The fix is a scalar rescale, not a better router.
  • IQuest-Q1 is a 320B-total, 15B-active mixture-of-experts with a 524,288-token context, shipped with a vLLM fork and an SGLang pull request, under MIT modified to require commercial products built on it to display the model's name in their own interface.

What to watch

  • Whether OpenAI ever names the evaluation GPT-6.1 Astra failed. Two documents and a recap have now passed without it, and a forum reply is not a disclosure. Without the threshold, "didn't quite meet the bar" cannot be checked or cited as precedent by anyone.
  • Whether OpenAI's Decisions API returns calibrated probabilities when it leaves preview. Its copy promises finite answers and 150ms; every other implementation promises a probability to threshold on.
  • A sealed-tier run of Jeff or Intern-Decision, asked here on the 25th, 27th and 29th. Jeeves excluded that tier by choice.
  • Whether anyone reproduces Anthropic's 12% ExploitBench figure. Benchmark and weights are both public, so this is the rare contested claim an outsider can settle outright. The in-house 4% cannot.

Daily, by email

Stay current on AI without the scrolling

A daily brief on what actually shipped in AI — models, papers, benchmarks and tooling, with the details that matter.

Confirmation email first, one message a day, unsubscribe in one click.