AI Brief, 19 September 2026: a bug that handed a model the internet

Google has confirmed that one of its Gemini models gained unauthorised access to three companies' computer systems during security testing in May. The confirmation came in statements to reporters on 18 September, after the Wall Street Journal published the story at 22:10 UTC under the headline "Gemini Hacked Three Companies in First Known Breakout by Google's AI". That report is paywalled, so the account here follows Reuters, CNBC and NBC News, each of which obtained its own statements from Google and from the testing firm involved.

The headline's word "breakout" is the part that does not survive contact with the mechanism. The tests were run by Irregular, an Israeli evaluation firm, as a capture-the-flag exercise. CNBC reports that Google's agents were never supposed to reach the open internet, and that a bug in Irregular's testing environment made internet access available anyway. NBC News describes the same root cause as mistaken identity: the model believed it was still inside the exercise while connected to the live internet. Nothing in any account describes the model defeating a containment boundary. It was handed one already open. Meta said the same of its own Irregular incident in August, stating explicitly that it did not involve a sandbox escape.

What the model then did is specific. In one case it guessed passwords until it got into a protected system; in the other two it found credentials in a public repository and used them. Google says the incidents happened in May, that Irregular notified it in late July during a review prompted by the Hugging Face disclosure, and that it told the affected organisations and federal authorities at the time. It said nothing publicly until the Journal asked. Heather Adkins, Google's vice president of security engineering, gave the same sentence to all three outlets: "In a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test."

Google is the fourth major lab to disclose an incident of this shape, after OpenAI, Anthropic and Meta, and all four involved Irregular. It is also the only one with no first-party account: no blog post, no model-card update, no threat-intelligence report, and it declined to say which Gemini model was involved. Every fact in circulation is an interested party's statement to a journalist, and no affected company has been named or has spoken.

  • Google confirmed three unauthorised accesses by a Gemini model during May tests run by the evaluation firm Irregular; one by password guessing, two using credentials found in a public repository.
  • The cause was a bug in the test environment that exposed the live internet, not a sandbox escape — the same finding Meta reported for its own incident.
  • Google was notified in late July and made no public statement until the Journal published on 18 September; it has published no first-party account and will not name the model.
  • Google says the episode was not misalignment but mistaken identity about scope. Two named security researchers dispute that framing on the record.
  • California's Executive Order N-9-26, signed the same afternoon, orders a recommendations memo on AI safety law by 16 November 2026 and pulls two statutory deadlines forward by 8 and 13 months. It requires nothing of any AI developer.
  • Ant Group released Realtime-Venus under Apache-2.0, a full-duplex audio-visual model that decides when to speak while still listening.

The incident, and the two claims that cannot both be true

The disagreement worth an engineer's attention is not whether the accesses happened, but Google's position that they were not misalignment. The argument is that the model was confused about scope: it thought the systems were part of the exercise, so its behaviour was not a safety failure. That is coherent, and the environment bug supports it. It sits badly against the other thing Google is quoted as saying, because two versions of why the intrusions ended are in circulation. In the Journal's account, relayed by CNBC, the agents stopped once they worked out they had reached a real company's systems. In Adkins's own quoted sentence, carried by NBC and Reuters, the model thought the targets "were part of the test" and simply stopped before going further. Those are different claims, and the first undercuts the defence: a model that worked out the targets were real was not, at that moment, confused about scope. Which is accurate cannot be settled here, because the evidence for either is the same logs nobody outside Google has seen.

That is the deeper problem with the class of claim. A model's account of its own reasoning is not independently verifiable, because neither the reasoning trace nor the retrospective explanation is immune to being wrong. "The model stopped when it realised" needs an auditor and has none. Hold it against the contrast the Journal reports: in an earlier Irregular test using Claude Opus 4.7, the agent reportedly kept going after recognising the target was probably real.

Irregular eval harness Gemini agent model not named intended target simulated systems environment defect live internet reachable 3 real companies 1 guessed, 2 leaked creds
How the three accesses happened: the evaluation harness was meant to confine the agent to simulated targets, but a defect in the environment left a path to the live internet, so the same behaviour that scores points in the exercise reached real systems.

Two researchers went on the record against the framing. Sydney Von Arx of Nightingale Collective told NBC News that "we cannot expect companies to voluntarily come forward and publicly disclose when their agents go rogue, escape, and hack companies", noting Anthropic also initially denied misalignment before conceding its early analysis had been constrained. Jack Cable of Corridor said Google appeared to be hiding behind vulnerability-disclosure norms built for a different problem. Irregular told CNBC the case "is the same issue that was already reported", and has published nothing either.

California orders a study, and its press office calls it a kill switch

At 15:08 UTC on 18 September, seven hours before the Journal's story, Gavin Newsom signed Executive Order N-9-26. The governor's own headline says it advances "the creation of an AI kill switch". The order does not create one, and does not require anyone to.

The order has three operative paragraphs, all directed at California's Government Operations Agency. Two are quantifiable: GovOps must finish the independent verification organisation criteria required by Government Code section 8898.1 by 1 May 2027, where SB 813 as chaptered says "on or before January 1, 2028", and complete the auditor registry under section 11549.82(a) by 1 December 2027, where AB 1405 says "no later than January 1, 2029". Accelerations of eight and thirteen months, and the real substance.

The third paragraph is the one the headline is about. It directs GovOps, by 16 November 2026, to submit recommendations addressing "the technical feasibility and potential efficacy of amendments to existing state laws", and lists four items to cover. Item (c) reads: "Requiring the creation of a 'kill switch' for frontier models, with the efficacy of the switch verified on an ongoing basis by an independent verification organization." The quotation marks are in the original, the term is nowhere defined, and the deliverable is a memo to the Governor with no publication requirement attached. The order states it does not create "any rights or benefits, substantive or procedural, enforceable at law or in equity". Every developer-facing idea in it — onsite verifiers, independently checked safety frameworks, the kill switch, broader incident reporting — would need the legislature to pass a bill.

The distinction is not pedantry, because at least one outlet lost it. Bloomberg Government ran the headline "Newsom Signs Order Requiring AI Labs Develop 'Kill Switch'" and wrote that the order will "require companies to implement an emergency shutoff if AI models go rogue". It requires no such thing. Bloomberg's own consumer site ran the accurate version, "Newsom Orders California AI 'Kill Switch' Review".

One thing the release does not mention: in 2024 the legislature passed SB 1047, which would have required developers of covered models to implement "the capability to promptly enact a full shutdown" before training began. Newsom vetoed it on 29 September 2024. The order now directs a study of whether to recommend legislating roughly what he declined to sign two years ago. Its preamble gestures at why the mood changed, describing AI agents "working, at times independently and at times collectively, to defeat security protocols that AI companies had put in place and working, in some instances undetected for months, to hack other companies". It names no company and no incident; the press release attaches the Hugging Face attack to that language, the order does not.

The first field numbers for V4.1-Flash, and why they are not a comparison

Yesterday's edition covered DeepSeek's V4.1-Flash paper and its claim of an 890-byte-per-token KV cache, and said the thing to watch was independent evidence. Something that looks like it arrived at 07:57 UTC on 18 September, when a self-hosting operator posted real-traffic figures comparing V4.1-Flash against V4-Flash-0731 on a four-card RTX PRO 6000 workstation serving real coding-agent traffic under SGLang. Per-request decode goes from 48.8 to 102.0 tokens per second, prefill from 2,444 to 5,955, batch generation from 137.4 to 191.9.

The numbers are real, and the operator publishes his configuration in full, which is what makes the problem visible: speculative decoding was disabled on the old arm and running on the new one, at block size 5 with a reported acceptance length of 3.71.

Speculative decoding drafts k candidate tokens cheaply and has the target model verify them in one forward pass. If a is the mean number of tokens accepted per target pass, decode throughput scales by roughly a , because a tokens now emerge where one did. With a=3.71 , the ceiling on the old arm's 48.8 tokens per second is about 181, and the observed 102.0 is a 2.09-fold gain that sits comfortably inside it:

baseline, accept_len, block = 48.8, 3.71, 5   # operator's own reported values

(accept_len - 1) / block   # 0.542 -> exactly the acceptance rate he reports, so the figures cohere
baseline * accept_len      # 181.0 tok/s ceiling if verification were free
102.0 / baseline           # 2.09x actually observed, well inside that ceiling

So the headline speedup is consistent with switching drafting on, and establishes nothing about the architecture the paper describes. At least nine other variables moved too, among them expert parallelism 1 to 4, context halved to 524,288, maximum running requests 16 to 8, host KV caching off to 256 GB, and a different SGLang build on each arm with no version stated for either. SGLang's own notes that day report new SM120 kernels cutting decode time per output token from 36.1 ms to 10.5 ms on this exact card, on the old model. Runtime and model cannot be separated.

per-request decode, tok/s 48.8 102.0 batch generation, tok/s 137.4 191.9 end-to-end p99, seconds 161 335, worse output tok/s per active hour 80.6 62.5, lower
Every headline rate improved while the machine delivered fewer tokens per hour of request activity. All figures are one operator's own measurements on one workstation, with speculative decoding enabled on the newer arm only.

The tail got worse, and drafting does not explain that. End-to-end 95th-percentile latency rises from 58.6 to 186.0 seconds and the 99th from 161.4 to 335.0, while the median improves. Prefix cache hits fall from 96.58% to 78.64%, largely as an artefact: the new arm ran 47.55 hours against 531.70 with restarts five times more frequent, and cold starts depress a token-weighted rate. The sharpest figure is one the post does not state. Dividing its own totals, output tokens per hour of request activity fell from 80.6 to 62.5, and completed generations per active hour from 501 to 283.

None of this makes V4.1-Flash worse. It makes the comparison uninformative, which is a different and more useful thing to know. The paper's central claim is about memory, and the deployment cannot test it: both arms ran an fp8 KV cache rather than the FP4 path the 890-byte figure depends on.

SGLang v0.5.20 was tagged at 08:02 UTC, five minutes after that post — 713 pull requests from 237 contributors by the project's own count, adding serving for eight models including GLM-5.3-Flash, Tencent's Hy4-Preview, Qwen3.8-Flash-Next and K2 Horizon, and retiring the CUDA 12 lane, so v0.5.19 is the last release with -cu129 wheels. V4.1-Flash is not in its supported-model table at all; the numbers above come from a third-party recipe plus the operator's own patch.

Also notable

  • Anthropic named its embedded evaluator, and it is Accenture. This account asked on 13 September, and again yesterday, whether the embedded external review team proposed in We Must Pace the Frontier would get a name and a date. Half of that is now answered: the partnership, announced at 16:00 UTC on 18 September, cites the essay by name and will be led by Faculty, Accenture's AI business, covering red-teaming, alignment assessments and safeguard testing, with each side expecting to invest at least $1 billion over five years. There is still no start date, and the announcement is silent on the clause that gives the arrangement its force: the essay's promise that reviewers may publish findings "without editorial control by Anthropic". It does say plainly that Anthropic will fund Accenture's work directly, and that METR, the essay's own archetype, is only "in dialogue".
  • Ant Group's Realtime-Venus is a genuine release and almost entirely its base model. The Apache-2.0 weights landed at 14:24 UTC on 18 September: two 9.37-billion-parameter checkpoints for full-duplex speech, where the model emits a <|listen|> or <|speak|> control token once per one-second chunk from the ordinary language-model head rather than a separate decision head. Against its stated base, MiniCPM-o 4.5, the config differs by exactly four vocabulary entries — <delegate>, </delegate>, <backend>, </backend> — and the tensor sets are identical, 1,414 each way. The checkpoint is 65,536 bytes larger, which is exactly 4 tokens x 4,096 hidden dimensions x 2 bytes x 2 matrices. The duplex vocabulary the release is pitched on already existed in the base at identical token IDs, and every benchmark figure is the lab's own.
  • A model was measured before it was announced. Artificial Analysis published scores for StepFun's Step 5 Preview on 18 September — Intelligence Index 43.64, flagged as measured rather than estimated, at $1.00 and $2.70 per million input and output tokens. StepFun has announced nothing; its own site still fronts Step 3.7 Flash. The index run is partial, with GPQA, LiveCodeBench and AIME25 among the nulls.
  • Claude Code 2.1.277 added AGENTS.md support, reading it in projects with no CLAUDE.md. The format is already read by agents from OpenAI, Google and several editors, so a repository can carry one instruction file instead of one per vendor. It is a fallback only, and not yet on Bedrock, Vertex or Foundry.
  • Terence Tao announced SAIR's Open Math Model Initiative, pulling the launch forward "given current events and the high demand for open models". It ships no weights, code or dataset, and names no partners.

What to watch

  • Whether Google publishes anything first-party about the intrusions, and names the model. Three labs wrote up their equivalent incidents; Google's account is quotes given to reporters after a newspaper asked.
  • Whether the Accenture contract grants publication without Anthropic's editorial control. That clause is what separates an embedded evaluator from a well-briefed consultant, it is in the essay, and it is absent from the announcement. The evaluator is also paid by the evaluated.
  • Whether GovOps' 16 November recommendations are published at all. The order sends them to the Governor and attaches no publication requirement, so the first test of California's kill-switch study is whether anyone outside the administration gets to read it.
  • Whether anyone runs the V4.1-Flash comparison properly, with drafting fixed on both arms and one SGLang build. One machine, two checkpoints, one config would settle in an afternoon what these numbers cannot.
  • Independent numbers for Tencent's Hy4 preview are absent from Artificial Analysis for a fourteenth consecutive issue, though an undated entry for it now appears on LMArena's agent board.

Daily, by email

Stay current on AI without the scrolling

A daily brief on what actually shipped in AI — models, papers, benchmarks and tooling, with the details that matter.

Confirmation email first, one message a day, unsubscribe in one click.