AI Brief, 4 October 2026: a million tokens, by taking the positions out

Aleph Alpha published Kolibri on 3 October and timed it for the Day of German Reunification, describing it as a "sovereign open-weight model". The weights landed on Hugging Face at 06:13 UTC, the vLLM plugin needed to serve them reached PyPI an hour later at 07:12, and the announcement followed the same morning. It is a mixture-of-experts model with 78,103,074,560 total parameters and 3,457,573,120 active per token, trained from scratch on German and English only, and released under Apache 2.0. One token touches 4.4% of it.

The reason to read the configuration rather than the benchmark table is an attention decision that almost nothing else ships. Forty of Kolibri's fifty layers are sliding-window layers with a window of 512 preceding tokens plus the current one. The other ten, every fifth layer, attend across the whole context. That pattern is becoming common. What is not common is where the positional encoding goes: rotary embeddings are applied only in the sliding-window layers, so the ten layers that actually see the full sequence have no notion of position at all. Nothing in them is indexed to a sequence length, so there is nothing to rescale when the context grows. Aleph Alpha trained to 262,144 tokens and says the context can then be pushed further "without any position scaling, in principle to arbitrary lengths"; it reports validating quality to 1,048,576.

That is a real claim with a published price, and the company prints the price itself. On RULER, the retrieval benchmark that sweeps context length, Kolibri Base scores 86.9 at 4k, 76.3 at 32k and 67.9 at 128k, against 96.3, 93.7 and 89.9 for Qwen3.5 35B-A3B Base at the same lengths and a comparable active-parameter count. From 8k to 128k it is behind every mixture-of-experts baseline in its own table; at 4k only Gemma 4 26B-A4B Base, on 84.7, is lower. At 1M tokens it scores 63.2, where Qwen3.5 scores 57.5 and Nemotron 3 Nano 58.5. The architecture buys a flatter curve, not a higher one, and it buys it by starting lower.

All of those figures are Aleph Alpha's own and no third party has run the model: it carries no OpenRouter listing and no independent index score. What is unusual is that the numbers it published are the ones that complicate its case.

  • Kolibri: 78B total, 3.46B active, 384 experts per layer with 6 firing, 50 layers, 1M context, Apache 2.0, German and English. Weights 3 October 06:13 UTC.
  • RoPE only in the local layers. The ten global layers use no positional encoding, which makes extension beyond the trained 262,144 tokens a configuration override rather than a retraining job.
  • The KV cache is the point. Interleaving 40 capped layers with 10 uncapped cuts the cache at a million tokens from 50 GiB to 10 GiB.
  • Self-reported RULER has Kolibri behind every comparable open MoE baseline from 8k to 128k, and first at 1M. Nobody outside the company has measured it.
  • 23.6 trillion training tokens and 950 MWh, both disclosed, which few labs do.
  • The licence grants the weights under Apache 2.0 but explicitly reserves "model architecture, parameter settings or any training method".

The attention pattern, and what it saves

The interleave is a block of five layers, four local and one global, repeated ten times.

4 × sliding window 513 tokens 1 × full attention whole context RoPE, base 10,000 cache capped at 513 tokens no positional encoding cache grows with the sequence × 10
One repeat of Kolibri's 50-layer stack. Four sliding-window layers see 513 tokens and carry rotary position embeddings; the fifth sees the whole context and carries none. The block repeats ten times.

The configuration gives 4 key-value heads at a head dimension of 128 against 48 query heads, and the recommended serving command uses an FP8 KV cache, so one token costs

2×nkv×dhead×1 byte=2×4×128=1024 bytes

per layer, where nkv is the number of key-value heads, dhead the per-head dimension and the leading 2 the key and value tensors. Exactly 1 KiB, which makes the rest arithmetic a reader can do in their head. The ten global layers grow with the sequence; the forty local layers pin at 513 tokens and stop:

10×n×1 KiB⏟global+40×513×1 KiB⏟local, constant

At the trained length of 262,144 tokens that is 2.50 GiB plus a fixed 20.04 MiB. At 1,048,576 tokens it is 10.00 GiB plus the same 20.04 MiB. Had all fifty layers attended fully, the same million-token context would need 50.00 GiB. The saving is a flat factor of five, the ratio of fifty layers to ten, and it decides usability: the weights occupy about 78 GB in FP8, so on the two-H100 node Aleph Alpha lists as its minimum, 10 GiB of cache for a million tokens leaves room to serve and 50 GiB does not.

This is the second million-token architecture in a week whose real subject is the cache. Naive-N0.5-Flash, covered here on the 28th, went further and kept no full-attention layers at all, pairing a 128-token sliding window with nine sparse layers selecting the top 2,048 tokens of the history, for about 24 GiB at a million tokens. Kolibri keeps ten uncapped layers and lands at 10 GiB. The two bracket the same question, how much global attention a long context needs, from opposite sides, and Kolibri is the one that published a RULER curve against named baselines.

Serving past the trained window is an opt-in. The shipped config.json sets max_position_embeddings to 262,144; reaching a million requires overriding it:

vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 --tool-call-parser kolibri1 \
  --max-model-len 1048576 \
  --hf-overrides '{"max_position_embeddings": 1048576}'

The experts are unusually small and numerous: 384 routed experts per layer with a hidden size of 512 against the model's 2,560, plus one shared expert, top-6 sigmoid token-choice routing, and norm_topk_prob set to false, so the six gate values are used unnormalised. Training ran 20 trillion tokens of pre-training, 3.44 trillion of mid-training at 65,536 tokens and 201 billion in a long-context phase at 262,144, on 768 GPUs at a constant learning rate after warmup, with Nesterov Muon on the two-dimensional backbone parameters and AdamW on the embeddings, router and head. The card puts the energy at 950 MWh excluding fine-tuning: 0.14 joules per training token across 23.6 trillion. Most labs publish neither figure, and the EU general-purpose AI Code of Practice, which Aleph Alpha has signed, is a likelier reason than candour.

What the tables concede

100 50 86.9 96.3 4k 76.3 93.7 32k 67.9 89.9 128k 63.2 57.5 1M Kolibri Base Qwen3.5 35B-A3B Base
RULER, the mean of four retrieval tasks, at four context lengths. Kolibri Base against the strongest comparable open model in Aleph Alpha's own table. All figures are self-reported by Aleph Alpha; Qwen3.5 is served with static YaRN beyond its trained window.

The post-trained model does better in the comparison Aleph Alpha leads with. Its unweighted average is 75.5 in English and 70.8 in German, the highest of the seven mixture-of-experts models at roughly 3B active parameters in its table, ahead of Qwen3.5 35B-A3B at 74.7 and 69.8; it loses to dense Qwen3.8 27B at 80.2 and 79.9. German reasoning is where the margins are wide: 90.0 on AIME 2026 in German, the best at its active size by 3.3 points, and 38.1 on τ³-bench banking where the next-best mixture-of-experts entry manages 16.0, a gap large enough to want a second party to check.

The same tables are candid about where it is weak, which an aggregator will not carry. On RGB Closed-Book it scores 51.0, the lowest of all fourteen models listed, where most score between 73 and 93, and on RGB Fact-Check error correction 34.0 against 53 to 90 for everyone else. Its multi-turn tool calling is mid-table at best, 39.8 on BFCL v3 and 47.5 on v4, which sits awkwardly beside the τ³-bench result. And in the base-model table it scores 16.7 on Wahl-O-Mat, the German voting-advice questionnaire, where every other model scores between 43.2 and 61.1 and its own 30B predecessor scores 52.3. The card does not define how that benchmark is scored, so the number is not a quality judgement either way, but it is the release's largest outlier and it sits next to a section saying the company "actively reduced" political bias through data filtering and alignment training. A model built for German public administration scoring lowest in its class on German party positions is interesting, and the company volunteered it.

The licence is narrower than "Apache 2.0" alone suggests. The weights and the serving plugin both carry Apache 2.0, but the card adds that the grant "especially does not extend to underlying code, model architecture, parameter settings or any training method", with all rights to those retained. The reservation covers the method rather than the artifact, and it is the second time in three days that an open release has needed its openness qualified in the small print, after the Cloudflare post whose Hacker News title a moderator edited from "open-source" to "open weights".

A second safety rupture at OpenAI in three days

The OpenAI agent disclosures and unlifted tool-use pause tracked here since the 26th are the technical half of a story that now has a staffing half. David Robinson, who spent three and a half years in OpenAI's safety organisation and helped write its Preparedness Framework, published a signed essay in The Atlantic at 11:00 UTC on 3 October explaining why he left. The argument is narrower and more interesting than the headline quote. He does not say the company is indifferent to risk; he says the method is wrong. OpenAI's stated approach is iterative deployment, which is to ship, observe what breaks and harden. His position is that iterative deployment only works while the failures are survivable, and that frontier systems have left that regime, so assurance has to happen before release the way it does in aviation and nuclear power. The essay is not publicly retrievable, and the summaries of its argument here follow the wire coverage of it rather than the text, so no passage of it is quoted.

This is one person's account of a culture, under his own name, with no corroborating documents and no response from the company: testimony, not an audit. What lifts it above testimony is the sequencing. Two days earlier OpenAI said it had parted ways with three individuals over the handling of sensitive company information, three safety researchers who, as the Wall Street Journal reported, had shared material with an outside AI-safety organisation. Two ruptures in the same team inside three days carry a different weight than either alone. The first says the company polices what its safety staff may tell outsiders; the second says a long-tenured insider concluded a magazine was the remaining channel.

Also notable

  • A vendor in the decision-model class finally published its own ECE and Brier score, and this brief missed it. The ask here for six consecutive issues has been for a vendor to put a calibration number on its own model, since a calibrated probability is the whole claim of the class, and Cloudflare's board carries those columns for 71 community entries but not for its own two. One vendor now has. IFM's K2-Type-0.9B went up at 19:30 UTC on 2 October, nine hours before yesterday's edition, and reports ECE 0.065 and a Brier score of 0.328 alongside 76.2% on the JevBench public set, with p50 latency of 27 ms per decision. It is Apache-2.0 and 1.08B parameters stored. Every figure is the vendor's own, and the card concedes both that the official JevBench score also uses sealed items on which every listed system scores far below its public accuracy, and that the calibration is fitted, so the probabilities can drift off that distribution. A disclosed ECE with a stated limitation is still the first real evidence the class has offered.
  • Two more entrants arrived without one. Sber's FRIDA-Decisions (2 October, MIT) is the first in the class built on an encoder rather than a decoder with a head, and reports 0.893 against Jev's 0.897 over 735 items of a Russian benchmark of its own authorship, at a paired McNemar p of 0.84, which is parity rather than a win. Its card has no calibration section. vllm-sr/Decision-2.0-Vega-27B, the largest open entry at a claimed 29.37B parameters, ships a config.json containing "calibration": null and no benchmark table at all.
  • Nebius is reported to be buying Inferize, a GPU-utilisation startup, for $100m to $150m, on unnamed sources with nothing filed behind either figure. Unconfirmed.
  • No arXiv batch falls in this window. 4 October is a Sunday and arXiv announces nothing at the weekend, so the newest listing in cs.LG, cs.CL and cs.AI is still Friday 2 October's, and Hugging Face published no Daily Papers page for 3 or 4 October: a calendar artifact, not a quiet field. Serving-engine releases are not covered in this edition.

What to watch

  • Whether anyone outside Aleph Alpha runs Kolibri, and at what context length. It has no OpenRouter listing and no independent index score. The claim needing a second party is not the headline average but the RULER curve: a reproduction at 4k and at 1M would settle whether the positional swap costs what the company's own table says it costs.
  • Whether the Wahl-O-Mat result gets an explanation. 16.7 against a field of 43.2 to 61.1, from a model built for German public administration, is either a scoring artifact or a direct consequence of the de-biasing the card describes. The card does not say which.
  • Whether a sealed-tier JevBench number appears for any open decision model. Asked here on the 25th, 27th, 29th, 30th, 1st, 2nd and 3rd. The public-set calibration figures now exist; the sealed tier, where the 2nd found open entrants scoring below zero, still has no open entrant in it. Clef still carries no ECE and no Brier on Cloudflare's own 73-row board.
  • Whether OpenAI answers the substance. Robinson's claim is about method, not conduct, and it is checkable: either pre-deployment assurance of the kind he describes is being built, or iterative deployment remains the policy.

Daily, by email

Stay current on AI without the scrolling

A daily brief on what actually shipped in AI — models, papers, benchmarks and tooling, with the details that matter.

Confirmation email first, one message a day, unsubscribe in one click.