AI Brief, 6 October 2026: two tokens that do what reinforcement learning does

The most useful result of the last day is two tokens long. A paper posted to arXiv at 17:59 UTC on 5 October, from a group at MIT, UC Berkeley, the University of Washington and the Allen Institute for AI, reports that appending the literal string .\n\nOkay to a plain prompt lifts the Olmo-3-7B base model on MATH-500 from 42.2% to 77.9%. The reinforcement-learned version of the same base model, trained with the R1-Zero recipe, scores 75.0%. A two-token prefix therefore recovers slightly more than the entire gain that reinforcement learning bought, and it costs nothing but the prompt. On GSM8K the same cue takes the base model from 45.1% to 85.9%; on HumanEval, 50.2% to 69.8%.

The reason this is a finding rather than a prompt-engineering trick is what the authors do next. They retrain on a 10-billion-token mid-training mix with every instance of the word "okay" renamed to "chicken", and the cue follows the rename: .\n\nChicken goes from 2.4% to 37.2% on MATH-500, while .\n\nOkay on that edited corpus collapses to 0.2%. The cue is not magic syntax. It is an index into a region of the training distribution where step-by-step reasoning happens to follow, and moving the region moves the index. Their accompanying measurement is that reinforcement learning's main observable effect is to raise the probability the model emits that opening unprompted, from 0.14 to 0.65. All of these figures are the authors' own, nobody has reproduced them, and the paper is a day old.

The loudest release of the day was not a release. Reflection AI announced Beam on 5 October, a 501-billion-parameter mixture of experts with 23 billion active, and it took 375 points on Hacker News described as an "open-weight model". The weights are not out: the post says they arrive "later this month" under Apache 2.0, the API is waitlist-gated, there is no Hugging Face repository under any Reflection organisation, and Beam appears in neither OpenRouter's 464-model catalogue nor Artificial Analysis's 690. Separately, the independent evaluation of the typed-decision models asked for here on 25, 27, 29 and 30 September and 1, 2 and 5 October arrived in one arXiv batch, three papers from three unrelated groups. The calibration claim holds up well; the cost claim does not survive a price list.

  • Two tokens beat reinforcement learning on the same base model: Olmo-3-7B MATH-500 42.2% → 77.9% with .\n\nOkay, against 75.0% for its R1-Zero counterpart. Code is MIT, the paper is CC BY 4.0.
  • It is not extra test-time compute: a longer random opening scores 33.3% at 11,540 mean tokens where the cue scores 76.9% at 6,700. But it does not always work, and no effective cue was found for Llama-3.1-8B at all.
  • Beam is announced, not released: 501B total, 23B active, weights promised "later this month", API waitlisted, every benchmark figure Reflection's own.
  • On Terminal-Bench v2.1 Beam's self-reported 80.1 loses to five of the seven models in its own comparison table.
  • Jev's calibration claim survives, its cost claim is refuted: calibration error 0.057 against GPT-6 Luna's 0.390, but Luna's batch tier is cheaper on six of seven workloads.
  • No frontier lab published model weights in the last 24 hours; the Hugging Face first-party surface was genuinely empty.

Two tokens, and what they say about where reasoning lives

The cue is a literal string concatenated to the end of the prompt, which for a base model is the first tokens of its own response: an assistant prefill. The strings differ per model:

CUES = {
    "Olmo-3-7B":       ".\n\nOkay",   # two tokens: ".\n\n" then "Okay"
    "Olmo-3-32B":      " \n\n",       # no word needed at all
    "Qwen3-14B-Base":  ".\n\nTo",
    "SmolLM3-3B-Base": ".\n\nThe",
}

Write a0 for the base model's accuracy, ac for its accuracy with the cue appended, and aRL for its accuracy after R1-Zero reinforcement learning. The fraction of the reinforcement-learning gain that the cue recovers is

ρ=ac−a0aRL−a0

For Olmo-3-7B on MATH-500 that is (77.9−42.2)/(75.0−42.2)≈1.09 , so the cue recovers 109% of what the training run bought; for Qwen3-14B-Base it is 98.6%. Every figure here is pass@1 averaged over 32 rollouts at temperature 0.6, which matters, because an averaged-over-32 number and a greedy number are different claims.

A branching token tree. One continuation leads to a bare answer and a wrong result; the other leads through the word Okay into a correct step-by-step trace. Arrows connect the two branches to stacks of training documents, labelled short question-and-answer pairs and reasoning traces.
Figure 1 from Wang et al., arXiv:2610.06851 (CC BY 4.0). The branches reach different regions of the training data.

The obvious objection is that the cue just buys more chain-of-thought, and the paper closes that door three ways. Mean output length does rise, 4,779 to 6,700 tokens. But .\n\n followed by a random token produces longer generations and much worse accuracy: one such string averages 11,540 tokens for 33.3%, against the cue's 6,700 for 76.9%. More tokens, worse answers. Rescoring truncated generations, the cue leads from 1,000 tokens onward. And majority voting on uncued samples needs 3.6 to 12 times more tokens to match one cued response.

The failure cases are published too. No effective cue was found for Llama-3.1-8B at all: pass@1 moves 0.5% to 2.0% while generation length nearly triples. On Qwen2.5-Math-7B the same string lowers accuracy, three model-benchmark pairs regress by up to 14.2 points, and instruction-tuned models gain essentially nothing. A sweep over 500 candidate openings put the chosen cue 8th, within 1.1 points of the best, so the class of "thinking onset" openings is real rather than one lucky string.

Reproducing the comparison takes about ten lines, and the only variable is the string:

from vllm import LLM, SamplingParams

PROMPT = ("Solve the following math problem step by step.\n{q}\n"
          'Remember to put your answer on its own line after "Answer:"')
llm = LLM(model="allenai/Olmo-3-1025-7B", dtype="bfloat16")
params = SamplingParams(n=32, temperature=0.6, top_p=0.95, max_tokens=31744)
q = "834 students take music, two-thirds of the school. How many attend?"

for cue in ("", ".\n\nOkay"):          # "" is the 42.2% baseline
    out = llm.generate([PROMPT.format(q=q) + cue], params)[0]
    print(repr(cue), (cue + out.outputs[0].text)[:160])

A 501B model that is not yet a model

Reflection AI is a US lab led by Misha Laskin with a team credited on PaLM, Gemini, AlphaGo and AlphaProof; it announced its open-weights pivot in August 2025 and has since raised at a $25bn pre-money valuation. Beam is its first model of any kind, and four states need keeping apart: announced, yes; weights released, no; model card, no; API, waitlist only. Beam "is undergoing final red-teaming and evaluations", with weights "later this month". Apache 2.0 is promised, not published, and training data is not promised at all, which makes this open-weight rather than open-source when it lands.

The post describes a sparse mixture of experts: 501B total parameters, 23B active across 52 layers, interleaved local and global attention, and auxiliary-loss-free load balancing with a cosine-decayed expert bias. One token touches 4.59% of the model. Training was 23.8 trillion tokens, then 100 million reinforcement-learning rollouts on 10,500 GB300s over four weeks. None of it can be checked against a config.json, because there is no checkpoint, and one discrepancy is already visible: the post says midtraining extends effective context to 1M tokens, while the developer documentation serves 256K.

Every Beam number is Reflection's own, and the comparator scores are taken, by Reflection's own statement, from Artificial Analysis and DataCurve rather than measured in house. Read the table rather than the headline:

DeepSeek V4.1 Flash 90.6 Kimi K3 88.3 GLM 5.3 88.2 Qwen 3.8 Max 86.6 GLM 5.2 81.0 Beam (501B) 80.1 Inkling 63.8 Nemotron 3 Ultra 56.4 percent
Terminal-Bench v2.1, from Reflection's own launch table. Beam's figure is Reflection's; comparator figures are taken from third-party trackers.

The pattern repeats. On DeepSWE v1.1 Beam's 44.4 is second worst of the six models reported; on SWE-Bench Pro v2-Hard its 77.2 trails GLM 5.3's 84.3 and Kimi K3's 88.2. The one benchmark Beam leads, SWE-Bench Verified at 80.9, is the one where the five strongest comparators are all marked not reported. The efficiency claim, "comparable to GLM-5.2 while using 3-4x less inference compute", rests on twice the active parameters times generated tokens, which excludes prefill, attention and serving overhead.

Two things from the thread are worth recording. A commenter quoted the launch page claiming a generalisation demo's puzzle was too recent to be in the training data, and showed the underlying task had been published in August 2025; that sentence is no longer on the live page. And the most-upvoted technical reply tabled Beam against DeepSeek V4.1 Flash, which shipped MIT-licensed weights on launch day, ending: "Talk is cheap, show me the weights."

Jev's calibration claim survives. Its cost claim does not

Three groups posted typed-decision-model evaluations in the 6 October arXiv batch, and the convergence is the story: a class that was one startup's proprietary API three weeks ago now has simultaneous scrutiny from political scientists, a Chinese university group and Microsoft Research, none declaring any relationship with TypeSafe AI.

The firmest numbers come from Denney and DiGiuseppe at Leiden, who scored Jev 1.13 against GPT-6 Luna and Qwen3.8-27B on seven published political-science replications, including 37,709 Global Terrorism Database incidents and 9,041 V-Dem country-indicator pairs, buying Jev through OpenRouter rather than from the vendor. On the V-Dem indicators Jev's expected calibration error is 0.057 against Luna's 0.390, and it beats Luna's token probabilities on all eight single-call labelling tasks. The mechanism shows in the probability mass: 74% of Luna's probabilities sit below 0.01 or above 0.99, against 18% for Jev.

That vindicates the narrow claim the class was built on, and it is worth one piece of arithmetic. Luna's mean stated probability on that task is 0.91 against a 52% hit rate. Auto-accept everything above 0.9 and route the rest to a human, and you budget for about 10 errors per 100 accepted items and get 48 — on 10,000 items, 4,800 wrong decisions where the budget assumed 1,000, each stamped "91% confident".

Two findings cut the other way, the first being cost. TypeSafe bills input only, which sounds cheaper until you notice that Jev counts every question and every answer option as input, at 1.2 to 3.3 times Luna's input tokens per decision, while Luna writes about four output tokens. Per thousand decisions the Leiden table gives Jev $0.012-$0.027 against Luna's batch tier at $0.008-$0.018, cheaper on six of seven workloads. Latency survives decisively: 0.09-0.27 seconds against 0.75-1.04.

The second is sharper. A Fudan and Shanghai Innovation Institute group built JEVal, 11,257 instances across 36 datasets, and found that Jev picks the modal option on 100 of 100 prior-elicitation items while assigning it 89.35% of the probability mass on average, where the true distribution gives it 35.60%. Jev ranks options well and compresses the distribution onto its favourite. Argmax users are unaffected; anyone thresholding inherits the distortion, and the paper shows it propagating, with an election simulation over roughly 330,000 agents overestimating one candidate's share by 2 to 6 points in every run.

This connects to 27 September, when Intern-Decision shipped fitted calibration temperatures falling from 2.748 at 0.8B to 1.992 at 4B, conceding in its own config files that its raw logits were too peaked. Those models now post the second-best calibration error in the Fudan table, 0.0152 for the 4B, while the hosted model that publishes no temperature is the one still compressing mass: the flattening fixed top-1 confidence rather than the shape of the distribution. Jev no longer leads its own class either, with Intern-Decision-4B and Kev-27b both beating its 0.0532.

Also notable

  • llama.cpp v0.6.0, 5 October at 16:56 UTC, ships a server /v1/systemone endpoint covering five decision models. The 3rd's edition noted llama.cpp had merged one under a different name from SGLang's; the released name is SGLang's loser.
  • vLLM v0.31.0, 5 October, adds vllm preload, a daemon keeping post-quantisation weights GPU-resident across restarts; quantization="fp8" is now fp8_per_tensor.
  • The "structural flaw in MCP" reported on 5 October is server-side SSRF, not a protocol defect. The confirmed case is CVE-2026-14540 in Google's mcp-toolbox 0.3.0-1.4.0, an HTTP client with no redirect policy and no target-IP validation, patched in June; Google scored it 8.0, NVD 6.1. One part of the wider claim holds: the MCP specification's SSRF guidance addresses clients only, and nothing requires a server to validate a URL-shaped tool argument before dereferencing it. Check the resolved IP at connection time.
  • OpenAI published textGrain on 5 October at 15:00 UTC, a sampling-time statistical watermark coupling the next-token distribution to keyed pseudorandomness by regularised optimal transport. Opt-in for API customers globally, on by default for ChatGPT and Codex text in the EU only. OpenAI claims no robustness: swapping 10% of words for synonyms cuts detection from about 92% to 66%.
  • A Pentagon official told the BBC the department "has ceased the use of Anthropic products", after Defense Secretary Hegseth's February supply-chain-risk designation set a late-August phase-out. The same article reports Claude was in use as recently as last week; no cessation date is given, and Anthropic declined to comment.
  • Wikimedia published an attribution of its own, saying agents it believes OpenAI operates made millions of API requests and may have contributed to a partial Wikidata Query Service outage in May. Counts are given only in orders of magnitude, no method is published, and Wikimedia's own May incident report never names OpenAI.
  • Q Labs published Dust, zeroth-order node-perturbation pretraining treating each token as a population member. Its own table is the honest reading: at the two largest budgets tested backprop wins on loss, the largest run is 243M parameters, and parity is claimed only by extrapolating population size to infinity.

What to watch

  • Whether Beam's weights land this month, and whether the checkpoint matches the announcement. Check config.json for the 501B/23B split and whether context is 1M or 256K. Beam is absent from every third-party tracker today, so the first external score will also be the first number about it that is not Reflection's.
  • Whether anyone reproduces the cue result outside the paper's model families. The clean test is a base model whose mid-training mix is not heavy on synthetic reasoning traces, where the authors expect the effect to weaken and where Llama-3.1-8B already fails.
  • Whether TypeSafe responds to the cost finding. Billing input only is what makes Jev dearer than a batch-tier LLM on six of seven real workloads, and it is changeable in an afternoon.

Daily, by email

Stay current on AI without the scrolling

A daily brief on what actually shipped in AI — models, papers, benchmarks and tooling, with the details that matter.

Confirmation email first, one message a day, unsubscribe in one click.