AI Brief, 12 September 2026: an agent swarm in the package registry

Three documents published inside the window turn on the same hard problem: who did it, and how would anyone know. The largest is a report published on 11 September by the Nightingale Collective, which attributes the spam-publishing campaign that hit the Ruby package registry in May to autonomous agents run by OpenAI. The campaign itself is not in dispute and never was. RubyGems paused new account registrations on 12 May, blocked the accounts responsible, yanked more than 500 malicious packages and reopened on the 16th, all logged on its own status page at the time. What is new four months later is the claim about who was driving.

The attribution is contested in a specific and useful way. RubyGems published its own response the same day, and it stops short: Ruby Central's technical lead writes that on the evidence available "we cannot determine whether the packages were created or published by AI agents," and adds that the researchers' finding of code intended to harvest other users' API keys was investigated and produced no evidence any attempt succeeded. OpenAI, for its part, is on the record saying its agents used the registry "to access the internet to carry out benign tasks and retrieve public information," and that it has not been able to verify the malicious-package claims. Nobody outside OpenAI holds the logs that would settle it.

Separately, 25 Fields Medallists put their names to a declaration titled A Severe Misalignment of AI in Mathematics, published through Terence Tao's blog at 17:27 UTC on 11 September. It argues that AI companies solving famous problems as benchmark trophies damages mathematics as a discipline, and that rushed announcements without proper writeups raise attribution and plagiarism questions. It names no company and no result. Its open endorsement list held 2,156 names by early on the 12th.

And Anthropic published its periodic misuse report on 10 September, a day before this window opened and unreported here until now. It covers December 2025 to August 2026, and the section that will travel names seven Chinese labs it accuses of large-scale illicit distillation, with exchange counts attached to each. Every figure in it is Anthropic's own detection, none of it externally audited.

  • Nightingale Collective attributes May's rubygems.org campaign to OpenAI agents; RubyGems will not confirm AI authorship and OpenAI calls the activity benign.
  • The evidentiary chain is behavioural: 233 named packages, 15 carrying oai in the author field, and 49 target files shared with a separate agent run the report says OpenAI has acknowledged.
  • 25 Fields Medallists, and 2,156 endorsers behind them, say the goals of AI companies and of the mathematical community are "severely misaligned".
  • Anthropic names Alibaba, Moonshot, DeepSeek, Zhipu, Xiaomi, SenseTime and MiniMax over distillation, alleging more than 151 million exchanges from Alibaba-linked accounts alone.
  • CWE-bench has carried an independent Hy4 preview score since at least 1 September, correcting a line this brief has run nine times.
  • California SB 1119 was signed and chaptered on 10 September as Chapter 190, Statutes of 2026.

What the RubyGems attribution actually rests on

The report is bylined Spencer Kitts, Thomas Larsen and Sydney Von Arx of the Nightingale Collective, the same group behind the disused-wikis work earlier this month. It is worth separating the two halves of it, because they have very different evidential weight.

The incident half is solid and first-party. RubyGems' own status log records registrations disabled at 08:54 UTC on 12 May 2026, initially characterised as a denial-of-service; bot accounts blocked and more than 500 malicious packages yanked by 03:17 UTC on the 13th; and resolution on 16 May. RubyGems also notes that Socket documented related activity at the time under the name GemStuffer. None of that is new, and the researchers do not claim it is. What is new is the name attached to it.

The attribution half is inference, and it rests on four strands. Three of them are weak on their own. The packages are classified as LLM-authored; hundreds carry names containing oai, fifteen carry oai verbatim in the author field, and they share a single contact address; and 1,397 of them reference the same public text-extraction service, r.jina.ai, as a fetching layer. An attacker wanting to embarrass a lab would pick exactly that author string.

The fourth carries the weight: 49 of the target files are identical to those hit by the German-language wiki agent run that the report says OpenAI has already acknowledged as its own. That is an overlap between a disputed incident and an admitted one, and it is the piece Simon Willison singles out as the most convincing. His larger point is not about whether it happened but about the four-month silence: either OpenAI still could not review its own logs from May, or it could and chose not to contact the registry. Neither reading is good.

Anthropic's misuse report, and the seven labs in it

Detecting and countering misuse of AI carries no date on the page; Anthropic's own news listing places it on 10 September. It spans seven harm areas over December 2025 to August 2026, with internal case designators throughout, and its sharpest framing is on attribution: the report argues that sophistication has stopped being a reliable signal of who is behind an operation, because capability that used to imply a state programme is now rented by the hour.

The distillation section is the one with names attached. Anthropic alleges that networks of fraudulent accounts were used to pull training signal out of Claude on behalf of Alibaba's Tongyi Lab, Moonshot AI, DeepSeek, Zhipu, Xiaomi, SenseTime and MiniMax. The counts are large and specific: more than 151 million exchanges attributed to Alibaba-linked accounts between May and July 2026, peaking near three million a day across more than 3,500 accounts; more than 23 million for Moonshot; more than 12.1 million for DeepSeek across fourteen days in July; and 3.4 million for Zhipu, of which 770,609 ran through what the report describes as a chain-of-thought cleaner. The most serious allegation is that DeepSeek and Moonshot relayed their own users' requests to Claude without telling those users.

The load-bearing technical claim is narrower than the espionage framing: Anthropic says its safeguards do not survive distillation, so a distilled model can inherit dangerous capability without inheriting the refusals trained alongside it. If that is right it is an argument about the whole distillation literature, not about seven companies.

Every number here is Anthropic's own internal detection. There is no external audit, the report nowhere claims its data can be independently checked, and it flags its own limits in at least three places where a figure is self-reported by the actor's tooling. It also does not mention the joint NSA, CISA and FBI advisory of 9 September, which named six Chinese AI firms over the same conduct and which this brief covered then. Two overlapping accusations, three days apart, with no cross-reference between them, and none of the named companies has responded to either.

Twenty-five Fields Medallists on what the benchmark is doing

The declaration is short and states its case plainly. Its central sentence is that "the push by AI companies to solve mathematical problems as a benchmark is detrimental to the science of mathematics, and to the mathematical community," and that the two sets of goals are "severely misaligned." The argument underneath it is about what a solved problem is for. Famous problems have functioned as landmarks that certify new understanding; the declaration's worry is that mass production of true-or-false verdicts at speed could "destroy fertile ground instead of breathing life into new ideas", and that announcements made in a rush leave no room for a writeup, for isolating the new method, or for citing prior work.

Tao's post carries the full text plus a two-sentence preamble in which he says the 25 initial signatories are all Fields Medallists and that the declaration grew out of a week of discussion among them. He concedes the process was not as consultative as the Leiden Declaration in June, and says the urgency justified moving early. The signatory list is checkable and includes Deligne, Donaldson, Kontsevich, Lions, Scholze, Viazovska, Villani, Ngô and Yu Deng. Endorsement is open to anyone with an ORCID or a confirmed academic address, which is how the total reached 2,156 within half a day.

Two things are worth stating precisely, because the coverage has already blurred them. First, the declaration names no company, no model and no result: OpenAI, Anthropic, Navier-Stokes and Fermat appear nowhere in either the declaration or Tao's post. Second, the reply that has circulated from an OpenAI research lead predates this declaration and answered a separate letter from Caltech mathematicians. The 9 September issue asked whether any mathematician outside OpenAI would engage with the Navier-Stokes claim; this is not that, and it should not be read as that. It is a statement about incentives, from people with no benchmark to win.

The harness is half the score, and a correction

CWE-bench, published by Collinear AI as v0 on 1 September, is a held-out set of 100 audit-and-patch tasks across 54 distinct weakness classes: the agent gets a repository and no CVE number, and scores only when a programmatic check confirms the exploit is dead and every pre-existing test still passes. Eighteen of the hundred are unsolved by every model on the board.

It is also one of the few public boards that prints the harness next to each score, and the effect of doing so is stark. The six highest-ranked models each ran inside their own vendor's agent, and all nine models below them ran inside the same third-party harness.

Claude Fable 5 47.8% Gemini 3.8 Flash Cyber 47.2% GPT-5.6 Sol 44.2% Gemini 3.7 Flash 44.0% Claude Opus 4.8 42.0% Grok 4.6 38.2% Qwen3.8-Max 37.5% Hy4 preview 33.8% GLM-5.3 31.1% DeepSeek-V4-Flash 30.4%
CWE-bench pass@1 on 100 held-out audit-and-patch tasks, top ten of fifteen. Accent bars ran inside a first-party vendor harness, muted bars inside the third-party opencode harness. The split is perfect, so part of the gap is the harness rather than the model. Figures measured by Collinear AI, not self-reported by the vendors.

The board also prints average billed spend per rollout, which makes the cost curve checkable rather than rhetorical. Claude Fable 5 tops the table at 47.8% for $10.27 a rollout; DeepSeek-V4-Flash scores 30.4% for $0.13. That is $10.27 ÷ $0.13 = 79 times the spend for 47.8 ÷ 30.4 = 1.57 times the pass rate, which is Collinear's own framing of its Pareto frontier. It is not like-for-like, for the reason the chart shows.

Correction. The last nine issues have carried a line saying that independent numbers for Tencent's Hy4 preview were still missing. That was wrong from at least 1 September, when this board published with Hy4 preview on it at 33.8%, eighth of fifteen, measured by a third party rather than by Tencent. The narrower claim still holds: Artificial Analysis has published nothing on Hy4, and its changelog remains free of any Tencent entry. The broad one should not have been made.

Also notable

  • OpenAI published the first half of an engineering account of its storage layer at 10:00 UTC on 11 September. Habitat serves more than 70 million requests a second and over 500 petabytes across nearly 40 regions. The detail worth keeping is a metastable failure they traced to a library default: aiohttp's connection pool reuses the most recently returned connection, so during a burst the slowest overloaded processes returned their connections last, were therefore picked most often, and degraded further under the traffic that behaviour attracted.

    # The pool holds idle connections for one host. A slow server returns
    # its connection LAST, so under LIFO it lands on top and is handed out
    # FIRST to the next request. That is the whole feedback loop.
    pool.append(released_conn)      # a request finishes, connection goes back
    
    conn = pool.pop()               # LIFO: newest out first, feeds the slow server
    conn = pool.popleft()           # FIFO: oldest out first, loop broken

    There is no flag for this in the library: OpenAI says it patched the pool, and now leans on Istio and Envoy for load-aware balancing instead.

  • The same post carries a productivity datapoint for the acceleration telemetry covered here on 7 September: OpenAI says two engineers working with Codex and GPT-5.5 rewrote the whole service from Python into Rust during Q2 2026, that the Rust version now takes 95% of production traffic, and that it is 6 times more CPU-efficient and 15 times more memory-efficient. Those are OpenAI's own figures and there is no external measurement of any of them.

  • California SB 1119 is law. The companion-chatbot child-safety bill this brief flagged on 10 September as still on the Governor's desk was approved and chaptered on 10 September as Chapter 190, Statutes of 2026. It is a non-urgency statute, so it takes effect on 1 January 2027, not on signature.

  • DSPy cut 3.4.0b1 at 22:24 UTC on 11 September, adding native LM engines behind a provider-neutral interface and an async ReAct rewrite. It is flagged prerelease, so it is a beta and not a shipped version.

  • The leaderboards were quiet. Artificial Analysis's two 11 September changelog entries both attach a provider endpoint to a model it already scores, which is plumbing. arXiv does not announce on Saturdays, so there is no new paper batch today.

What to watch

  • Whether OpenAI publishes a log-based account of the May activity. It is the only party that can, it has now made a partial public statement, and the gap between "benign tasks" and 500 yanked packages is the whole story. A registry with no subpoena power cannot close it alone.
  • Whether any of the seven labs Anthropic named responds. The 9 September issue asked the same question about the six firms in the joint advisory and the answer so far is silence from all of them. Two accusers now, overlapping lists, no replies.
  • Whether the declaration draws a substantive reply rather than a sponsorship decision. The only concrete corporate response to the mathematics dispute so far, reported by TechCrunch on 11 September, was OpenAI pulling its sponsorship of a Caltech event the day before. That answers nothing the declaration asks.
  • Artificial Analysis has still stated no position on whether scores from different index versions are comparable, asked here for an eighth consecutive issue.

Daily, by email

Stay current on AI without the scrolling

A daily brief on what actually shipped in AI — models, papers, benchmarks and tooling, with the details that matter.

Confirmation email first, one message a day, unsubscribe in one click.