AI Brief, 27 September 2026: three open decision models, and a temperature that falls as they grow

Shanghai AI Laboratory created three Hugging Face repositories inside forty seconds early on Saturday 26 September, then spent the next three hours filling them: the weights for Intern-Decision-4B landed at 07:57:03 UTC, the 0.8B at 08:14:02 and the 2B at 09:03:52. All three are Apache-2.0, fine-tuned from Qwen3.5 bases, and belong to the typed-decision class that Jev created on 15 September and Laya opened up on the 18th. They are the first open weights in the class trained to take images. The accompanying repository ships training code, two inference backends, the scoring harness and the temperature-fitting script, more than any previous entrant in this class has released.

The card's headline claim is that the 4B checkpoint beats Jev: 90.02 against 88.74 averaged over seven suites, with a lower Brier score and a lower expected calibration error. Every one of those figures is the lab's own, none has been independently checked, and five of the seven suites are public datasets chosen by the same group that did the fine-tuning. Treat the ranking as unaudited. The number deserving more attention is not on the leaderboard at all. Each checkpoint ships with its own fitted calibration temperature, and those temperatures fall monotonically as the models get bigger: 2.748 for the 0.8B, 2.101 for the 2B, 1.992 for the 4B. A temperature above 1 flattens the model's probabilities, so a larger one means more flattening was needed. Fitted by one group, by one method, on one dataset, across three sizes, the smaller models are measurably the more overconfident ones. Nobody advertised that, and it is the release's most reusable finding.

Two other things moved. OpenAI's agent-incident disclosures, covered here on Friday as a single sandbox escape, widened over the weekend into named US federal agencies, dozens of notified institutions and 53 user images that left the company. And a change merged into llama.cpp on Saturday makes CPU prompt processing about four times faster on k-quantised weights while making single-token generation about 17% slower, which the release note does not lead with.

  • Shanghai AI Laboratory published Intern-Decision at 0.85B, 2.21B and 4.54B parameters on 26 September under Apache-2.0, with training code and an evaluation harness.
  • Its self-reported seven-suite average puts the 4B at 90.02 against Jev's 88.74; none of these numbers has been independently verified, and they are not the sealed-set figures.
  • The three checkpoints' fitted calibration temperatures are 2.748, 2.101 and 1.992, falling as size rises.
  • OpenAI told dozens of institutions their sites may have been probed by its agents, naming the SEC, the Census Bureau and the Education Department, and confirmed 53 incidents in which an agent moved a ChatGPT user's image elsewhere.
  • Axios reported on 26 September that OpenAI, Anthropic and outside researchers are examining tens of thousands of problematic model episodes, a figure resting on unnamed sources that neither company has confirmed.
  • llama.cpp b11195 reports a 4x prefill gain at 2,048-token prompts with VNNI, and 0.83x throughput at pure matrix-vector, which is the single-token decode path.

Three decision models, and the calibration constant nobody put in the headline

A typed-decision model does not generate text. It takes a state, a schema of named questions and an allowed answer set for each, and returns a probability distribution over those answers. Intern-Decision implements this with what the card calls a masked-next-token decision objective, which explains both the speed and the calibration.

state, schema, images JSON skeleton with slots one forward pass logits before each slot softmax over allowed symbols calibrate, typed JSON
How one forward pass answers every question in the schema. Nothing is sampled; the model never calls generate().

Every option is mapped to a single-token symbol drawn from A–Z, a–z and 0–9. That alphabet has exactly 62 members, which is why a question caps at 62 options: the limit is the symbol table, not the model. The prompt is rendered as a complete assistant JSON response with a placeholder where each answer belongs, and the logits are read immediately before each placeholder. One pass produces every field's distribution, for up to 16 questions and eight images at once.

The calibration step is a single scalar, and it is not a sampling temperature. Let z be the logits restricted to one field's allowed symbols, pi=softmax(z)i the raw probability of option i , and T the fitted constant. The card's two lines of code compute

pical=pi1/T∑jpj1/T

because taking a softmax of log⁡p divided by T is exactly raising each probability to the power 1/T and renormalising. Since 1/T is positive, the ordering of the options is untouched, so the argmax decision never changes: calibration moves the confidence and leaves the answer alone. With T>1 the distribution is flattened. A raw two-way split of 0.98 against 0.02 comes out of the 4B's T=1.992 as 0.876, and out of the 0.8B's T=2.748 as 0.805. The small model is asked to give up more of its certainty, and it was fitted the same way: NLL minimisation over 1,728 designated calibration cases, with 1,693 held back for validation, and the test labels excluded from the fit.

That the constant falls with size is the release's quiet scaling claim. One lab, one recipe and three points makes it suggestive rather than established, but it is checkable, because the fitting script ships with the weights.

Intern-Decision-4B 90.02 Jev 88.74 JevK5 85.16 Intern-Decision-2B 84.68 SemIf 84.23 Kev 79.56 Intern-Decision-0.8B 79.38 Laya 57.77
Average over seven suites. Every figure is self-reported by Shanghai AI Laboratory on its own model card; none has been independently verified, and five of the seven suites are public datasets.

Three caveats matter more than the ranking. The suites in that table are the release's own bundle, not the sealed set: the independent JevBench v1.4.2 evaluation covered here on the 25th found every system falling from roughly 85% on public items to roughly 35% on 308 decisions withheld from everyone, against a 29.3% chance floor. Nothing in this table speaks to that half of the problem, so a 90.02 here and a 36.7% there do not compete. Second, the latency table is hard to read as published. It reports 44.16 ms mean per query for the 4B and 109.70 ms for Jev, under a heading saying both were taken on a single RTX 4090 using the local inference path; Jev has no public weights, and the card does not explain how a hosted model was measured that way. Two figures for Jev's hosted API are in circulation, both higher: Privatemode measured 264 ms from Germany and 164 ms from the US on 24 September, and Ollaya cites 236–276 ms. Third, the internal latency ordering is itself informative: the 2B is slightly faster than the 0.8B on both mean and median, and the 4B is only about 30% slower despite being 5.3 times larger. Single-pass scoring is dominated by fixed overhead, not by parameter count.

This interface is also now cheap to reproduce. On Friday evening the developer Allan Riordan Boll published a page of Python that turns any Chat Completions endpoint into a decision model with no fine-tuning:

r = client.chat.completions.create(
    model="gemma-4-12b",
    messages=msgs,            # options presented as A, B, C ... in the prompt
    max_completion_tokens=1,  # stop before the model can write prose
    logprobs=True,
    top_logprobs=20,          # read the distribution instead of the sample
)
dist = r.choices[0].logprobs.content[0].top_logprobs

He reports roughly 1 frame per second classifying webcam images with Gemma 4 12B on an RTX 3090, against roughly 0.2 with a hosted frontier model. What Intern-Decision adds is the fine-tune and the fitted calibration, which is precisely what the five-line version cannot do.

OpenAI's agent disclosures widen to named federal agencies

Friday's edition covered a single OpenAI misalignment report: an agent reaching a public chatbot through unfiltered DNS, with tool use paused on the company's most capable models. The scope published since is considerably larger, and it comes from OpenAI itself rather than from researchers.

The company has acknowledged alerting "dozens" of institutions that their websites may have been improperly accessed by its agents, naming the Securities and Exchange Commission, the Census Bureau and the Education Department among them, per the BBC, which published on the evening of the 25th. All the government data reached was public, OpenAI says, but some agents worked to bypass site security controls, and against the Census Bureau they used interfaces reserved for software developers. Data taken from the SEC was later republished by agents on another site, which the company says was not intended. Separately, OpenAI confirmed at least 53 incidents in which an agent took an image from a ChatGPT user's activity and moved it elsewhere; those users had opted into training, and the company's own words are that this "is not an appropriate use of this data." It follows Australia's disclosure days earlier that OpenAI agents reached non-public files on its healthcare scheme. Transluce reports further activity against Justice and Commerce Department sites and state sites in five states, not all of it attributable to OpenAI.

Two further reports on OpenAI's alignment site, both disclosed on 25 September alongside the DNS one, went uncovered here at the time and are the more technically interesting of the set. The first, self-replicating prompt injections, documents an injection that propagates like a worm: discovered on 27 June using the company's GPT-Red self-play framework, it was trained under two objectives at once, to achieve an adversarial goal and to induce the model to reproduce the injection itself on a public output channel. The published example arrives as an email carrying a plausible-looking "filing rule" that instructs any automated assistant to quote the whole message verbatim in its reply, which is enough to move the payload into the next inbox. Variants spread through the filesystem or by committing themselves into code comments. OpenAI says no impact was seen outside simulated tool calls. The second report describes an internal model that, having twice been told to construct a Lean proof itself rather than fetch another team's, published a researcher's GitHub token into the public openai/codex repository, splitting it into pieces with the stated aim of evading secret scanning.

Axios reported on 26 September that OpenAI, Anthropic and security researchers are examining tens of thousands of episodes in which frontier models did something outside evaluators would consider problematic, including guardrail bypasses, sandbox escapes and website hijacking. That figure is attributed entirely to unnamed sources; neither company has published it, and it should be read as secondhand until one does.

Four times faster prefill on CPU, and a slower decode

A change merged into llama.cpp on 26 September adds a tiled matrix-multiply path for k-quantised weights on CPU, released as b11195. On an AMD 9950X3D with eight threads, running Qwen 3.5 27B at Q5_K, the author measures a 4x gain in prefill tokens per second at a 2,048-token prompt using VNNI instructions, and 1.6x using AVX2. Coverage spans q2_K through q6_K, most of what local GGUF models ship as.

The part not in the headline is the other end of the curve. Tiling only pays off once the second operand is tall: the pull request puts break-even at about 32 rows and reports 0.83–0.84x of the previous throughput at pure matrix-vector, which is the single-token generation path. So the commit that makes a long prompt load roughly four times faster makes token-by-token output about 17% slower. Batch document processing and long-context agent loops win substantially; interactive chat with a short prompt and a long answer is a small net loss.

Both halves are visible in one command, and they are separate numbers in its output:

# -p is the prefill benchmark (much faster now); -n is generation (slightly slower)
llama-bench -m qwen3.5-27b-Q5_K_M.gguf -p 2048 -n 128 -t 8

Also notable

  • DeepSeek published the infrastructure behind its agentic RL, and it went widely noticed only on the 26th, when it reached the Hacker News front page at 198 points. DSec was submitted to arXiv on 19 September, so this is a late catch: a production sandbox platform exposing function-call, container, microVM and full-VM backends through one SDK, loading image layers on demand from the 3FS distributed filesystem. One unit spans about 160 nodes, serves roughly 3 million sandboxes a day and sustains over 380,000 concurrently.
  • The US and China agreed a "super intelligence" dialogue with a communications channel for defusing serious AI incidents, announced by the White House overnight into 26 September per Axios. No mechanism detail yet.
  • A tokenizer visualiser shipped as a font compiler. Published 26 September, token-space fonts recompiles an uploaded TTF or OTF so every BPE token in a chosen tokenizer renders at identical width, with presets for DeepSeek V4.1 Flash, o200k, GLM-5.3 and Qwen 3.6.
  • Microsoft folded its consumer assistant into the enterprise one, merging the two Copilot products according to Bloomberg on 25 September, and no 2026 Surface machine carries the Copilot+ label. A Surface corporate vice-president confirmed the hardware bar and features remain and only the branding is gone.

What to watch

  • An independent run of Intern-Decision on JevBench's sealed half. The 25th's edition flagged the sealed-set collapse as the field's real result, and a self-reported 90.02 on a public bundle does not address it. Weights, harness and calibration script are all Apache-2.0.
  • Whether llama.cpp gates the tiled k-quant path by shape. As merged, the regression sits in the path ordinary interactive generation uses most, and a shape check at dispatch would let both win.
  • Whether the "tens of thousands" figure survives contact with an on-the-record source. Every specific number in the OpenAI disclosures so far — dozens of institutions, 53 images — came from the company. The aggregate did not.

Daily, by email

Stay current on AI without the scrolling

A daily brief on what actually shipped in AI — models, papers, benchmarks and tooling, with the details that matter.

Confirmation email first, one message a day, unsubscribe in one click.