AI Brief, 2 October 2026: three vendors shipped a decision model, and the sealed scores went negative

The typed-decision model stopped being a product and became infrastructure. Cloudflare published Clef and clef-flash under Apache-2.0, weights landing 30 September and the announcement following at 15:34 UTC on 1 October; Amazon's Strands Decider 2B went up on 30 September, also Apache-2.0; Perplexity's pplx-decider-v1-27b at 21:24 UTC on 1 October with no announcement at all. Then at 01:09 UTC this morning SGLang shipped v0.5.21 with a native /v1/decisions endpoint. Fifteen days ago this class was one startup's proprietary API.

Almost nobody has run any of it. Clef carries 385 likes against 18 downloads, all-time; clef-flash 132 likes and the same 18; Amazon's adapter 14 likes and none. The Cloudflare post took 462 points on Hacker News, where a moderator edited its title from "open-source" to "open weights" after readers noted that no training data or code was published. Those measure attention, not use.

The substance is calibration, because that is the entire claim of the class: a decision model should return a probability you can threshold on rather than a sentence you parse. Cloudflare publishes its own 73-model leaderboard carrying expected calibration error and Brier score for every entry. Clef and clef-flash are the only two of the 73 with both fields empty, and the suite's one calibration-named benchmark is unscored for both. The three scalars that actually set Clef's confidence sit in the checkpoint at exactly ln 10, exactly ln 10, and exactly −2, bit-identical across the 27-billion and the 9-billion model, weighting string matching roughly eight times the head that reads the state.

And the question asked here on the 25th, 27th, 29th, 30th and 1st has an answer. JevBench now publishes per-system sealed-tier scores for all 106 ranked systems, open weights included. The open entrants do not merely fall on the held-back items; they go negative.

  • Clef (27.36B backbone plus a 128M head, Apache-2.0, multimodal) answers Jev's own /v1/systemone request body — a drop-in replacement for the endpoint it competes with.
  • Every Clef benchmark number is Cloudflare's own, and its board scores Clef on 36 of 38 index benchmarks against 38 for every rival, the omissions including the suite's hardest.
  • Three independent reproductions measure latency 4x to 17x worse than the claims.
  • JevBench v1.5.4: jeff scores 12.71 open and −3.56 sealed; OpenDecision 9.47 and −5.26. Neither Clef nor Intern-Decision is on that board at all.
  • NVIDIA published a 15.2-billion-parameter multimodal model as 480 raw training-checkpoint shards with no config, no tokenizer and no loader.

Clef, and the six bytes that outrank the reasoning head

Clef pairs a 27,356,728,560-parameter backbone with a separate 128,056,324-parameter decision head: 64 layers, hidden size 5120, 24 attention heads over 4 key-value heads, and three linear-attention layers to every full one, 16 full and 48 linear. The card names Qwen3.8-27B as the base while the config declares the qwen3_5 architecture class. A 27-layer vision tower makes it the first in this class to take images. clef-flash is the same recipe at 9,409,813,744 parameters.

Context length is quoted three ways: the blog says 64k, the docs 65,536, the backbone config allows 262,144, and the shipped inference code defaults max_length to 16,384. That is a caller-settable argument, so not strictly a contradiction, but the advertised figure appears in no downloadable artifact.

The head is four transformer layers of width 1024 with 16 heads, cross-attending to the backbone's hidden states. It takes a JSON state, optional images, and up to 64 questions typed noul, choice or score, returning one logit per allowed option, softmaxed per question. For an ordered score field it reports the expectation over levels and the top probability as confidence — precisely the ordinal behaviour yesterday's edition covered a paper measuring. There is no post-hoc recalibration anywhere in the inference path.

Each option's logit is the sum of two paths:

ℓ(o)=esp⟨u^o,a^q⟩+σ(g)[esjcos⁡(fq,vo)+r(fq,vo)]

ℓ(o) is the logit for option o of question q . The first term is the lexical prior: u^o is the normalised embedding of the option's own text and a^q the normalised question anchor, so their inner product measures how much the option string resembles the question string, and it never consults the state. The second is the joint path: fq is the field vector the head produces after cross-attending over the backbone's representation of the state, vo the routed option vector, and r a small MLP over [fq,vo,fq⊙vo,|fq−vo|] . σ is the logistic function, and sp , sj and g are three learnable scalars that the published code initialises to zero.

The values in the checkpoint are not zero:

import math
# from joint_head.safetensors; byte-identical in clef and clef-flash
s_p = s_j = 2.296875   # bf16 round-trip of ln(10), hex 0x4013
g = -2.0               # hex 0xc000

math.exp(s_p)          # 9.943  weight on the lexical match
1 / (1 + math.exp(-g)) # 0.119  gate on everything the head computed
0.119 * math.exp(s_j)  # 1.185  effective weight on the reasoning cosine

So the lexical path carries 9.94 and the reasoning cosine an effective 1.18, a ratio of 8.4. Work it through. Take two options where A's text resembles the question slightly more than B's, by 0.12 of cosine, while the head that read the state prefers B strongly, by 0.85. The prior gives A 9.94×0.12=1.19 ; the joint path gives B 1.185×0.85=1.01 . A wins by 0.18, which over two options is 54.6% to 45.4%. The model reports 54.6% confidence in the answer its reasoning head argued against. And because cosine is bounded, the joint term's whole available swing is 1.185×2=2.37 , so any lexical gap wider than 2.37/9.94=0.238 cosine cannot be overturned by that term at all. Only the residual MLP can, and it is gated to 0.119.

Whether those scalars were trained to round numbers or written in by hand cannot be settled from outside the company. What is checkable: two separately post-trained checkpoints carry byte-identical values, both are exact constants, and the code initialises them elsewhere. Cloudflare's post does not mention them.

Nor is the surrounding evidence strong. Every figure in the release is Cloudflare's own — its board marks Clef and clef-flash self-reported while all 71 other entries are community, and the two ran on an H200 against an upstream board run on an RTX PRO 6000. Its own data file records panel_coverage: 36 for both its models and 38 for every rival, the omissions being iSarcasmEval and HLE, the suite's hardest benchmark and one in the joint-highest-weighted area. The blog claims the Clef models "beat the decision models on latency", contradicted by the table printed above it: Clef's 238.6 ms p95 is worse than three of the five columns. The board also notes that rival latency includes a network round trip and is "not comparable" to Clef's local figures, a caveat neither the blog nor the card carries. And of the four models tabulated against, three rank 24th, 29th and 63rd on that same board, while the two systems sitting between Clef and clef-flash go unmentioned.

Three readers posted their own measurements within hours, and all three disagree sharply. One reports clef-flash at 661 ms median against a claimed 38.8 ms; another, publishing a harness, puts hosted Clef near 850 ms p50 against a claimed 209.3 ms and measures Jev about five times faster than Cloudflare's figure for it; a third reports Clef "2-3x slower and worse" than Jev on a production moderation pipeline. None of the three is replicated either, but unaudited vendor numbers contradicted by three unaudited user numbers is a weaker evidentiary position than the launch suggests.

Amazon's entrant takes the opposite approach to evidence and deserves credit for it. It is a LoRA adapter on Qwen3.5-2B-Base under Apache-2.0, and the repository ships the whole evaluation: results.jsonl, per-bin expected calibration error, a SHA-256 manifest, training stage logs, and the harness commit it ran against, 1bcc55eb. The headline is 72.29% on public items, 167 of 231, at strict schema validity of 1.0 and a median 85 ms on one H100. Two caveats the artifacts supply themselves: the tiers are 48/48 easy, 65/72 standard and 54/111 hard, so the aggregate rides on the easy half; and the runs at context widths 3072 and 4096 are genuinely different — different checkpoint hashes, different GPUs, mean Brier 0.3488 against 0.3477 — yet return the identical 167 correct and the identical count in all three tiers. The width change moved the confidences and not one answer.

The sealed half finally has numbers, and they are negative

JevBench is now at v1.5.4 and runs 720 sealed decisions per system against 904 open, the sealed half weighted at 50% of the Intelligence axis, across all 106 ranked systems rather than the API entrants only. The run asked for here on the 25th, 27th, 29th, 30th and 1st exists. The board publishes no last-updated date, so when it landed is not establishable.

jeff, the MIT-licensed 400M model covered here on the 29th, sits at rank 84 with 12.71 open and −3.56 sealed, a gap of 16.27. OpenDecision, an Apache-2.0 build on ModernBERT-large, is rank 87 with 9.47 and −5.26. The board's first-place system, Cygnet, built on a frozen gemma-4-12B-it, gains 2.7 points when the items are hidden. A negative score is worse than the floor: these models are not failing to generalise, they are anti-correlated with the hidden answers while looking mid-table on the half they could have trained against.

0 12.71 -3.56 jeff 9.47 -5.26 OpenDecision left: open · right: sealed
Open and sealed Intelligence for two open-weights decision models on JevBench v1.5.4, measured by the board operator rather than the developers. Both go below zero on the 720 held-back items.

Neither Clef nor Intern-Decision appears in JevBench's 112-system roster, nor in its not_measured list. The week's two most prominent open decision models are absent from the only board that seals items.

NVIDIA published fifteen billion parameters that nothing can load

PixelUMM went up at 21:40 UTC on 1 October: an encoder-free unified multimodal model, 15,199,672,064 parameters, built on Qwen3-8B with code derived from ByteDance's BAGEL. Rather than take embeddings from a pretrained vision encoder it splits images into raw 16-by-16 pixel patches, sharing one transformer representation between understanding and generation.

The repository contains 491 files. Four hundred and eighty are .distcp shards — PyTorch Distributed Checkpoint fragments, in four directories of 96, 128, 128 and 128 — plus four .metadata files and four __SAVE_COMPLETE markers. There is no config.json, no model.safetensors.index.json, no tokenizer_config.json and no generation_config.json; no inference code and no paper. The card links a LICENSE and a THIRD_PARTY_LICENSES.md that are both absent, while stating the source is Apache-2.0 and the checkpoint is under NVIDIA's One-Way Noncommercial License, research only. What has been published is the raw output of a training job, four times over, with the loader left behind.

Also notable

  • SGLang v0.5.21 (01:09 UTC, 2 October): 779 merged pull requests from 227 contributors, prefill and decode instances switching roles live without a restart, the prefix cache moved to a Rust core by default, and new /v1/decisions and /v1/score endpoints. A serving engine adding a first-class decision route is the clearest sign yet that this class is now plumbing.
  • Pi 1.0 (19:14 UTC, 1 October, MIT, 111,326 GitHub stars, 920 Hacker News points) added codemode: the model emits JavaScript run in a QuickJS WASM sandbox with no filesystem or network, and non-LLM classifier models become first-class, Jev among them. Separately, Figma's remote MCP server allowlists the client_name sent during OAuth and its catalog of 24 clients omits Pi — so Earendil shipped an oauth.clientName setting whose documented example sends "clientName": "Claude Code". An allowlist keyed on a self-reported string is a convention, not a control.
  • Hierarchical Continuous Diffusion Language Models (arXiv:2610.02193, 17:59 UTC on 1 October, UIUC with Amazon) couples a continuous latent to a discrete token scaffold in one denoising process. Its own table is admirably unflattering: it loses in-distribution Sudoku to a reproduced baseline, 94.21 to 94.65, and posts 75.5 generative perplexity on LM1B against an autoregressive transformer's 66.7 at comparable size. No code or weights.
  • Context Language Models (arXiv:2609.37725) took 127 Hacker News points yesterday but was submitted 29 September, so it is late here: it hands the agent its context as a file to edit with Bash, self-reporting 59.4% on BrowseComp-Plus.
  • Project Suncatcher's prototype satellite is in orbit, confirmed by Google at 23:30 UTC on 1 October — built with Planet, launched on Transporter-18, and meant to measure TPU radiation, thermal and stress tolerance in orbit. The 25th asked whether Transporter-18 would fly and whether Google would then publish telemetry rather than another video. Half of that is answered.
  • Artificial Analysis measured Grok 4.7 at Low effort, Intelligence Index 42.2. Its index remains v4.3.2, and its methodology page now carries a version history from v1.0 — but still no statement on whether scores compare across versions, asked here for a twenty-eighth consecutive issue, as is a number for Tencent's Hy4 preview. Gemini 4 Argon still has no OpenRouter row and one effort tier.
  • Policy: an FTC spokesperson confirmed on 30 September that the agency is investigating OpenAI, Anthropic and other AI companies over consumer risk, with information requests also going to METR; the probe opened over the summer, the FTC has published no document on it, and this account follows CBS News, which credits the New York Post with breaking it. Separately, California's attorney general served OpenAI an investigative subpoena over the Hugging Face incident, and a judge dismissed the Chegg and Penske antitrust suits over Google's AI Overviews (The Verge).

What to watch

  • Whether Cloudflare publishes an ECE and a Brier score for its own models. Its board has the columns and fills them for 71 of 73 entries. A decision model with no calibration number is a classifier with a confidence field, and this is the one gap the vendor can close in an afternoon.
  • Whether the latency figures survive a fourth party. Three measurements 4x to 17x worse than the claims are the only non-vendor evidence so far, and one of them ships a public harness.
  • Whether Clef or Intern-Decision is submitted to a sealed board at all. The tier missing for five issues now exists and runs on open weights, and the week's two highest-profile open releases are not in it. Absence is a choice now, not a gap in the evaluation.
  • Whether anyone reproduces Amazon's identical tier counts at two context widths, and whether NVIDIA ships a loader for PixelUMM. Both are settleable from outside: Amazon shipped every artifact including the pinned harness commit, and PixelUMM's architecture is untestable until it loads.

Daily, by email

Stay current on AI without the scrolling

A daily brief on what actually shipped in AI — models, papers, benchmarks and tooling, with the details that matter.

Confirmation email first, one message a day, unsubscribe in one click.