AI Brief, 25 August 2026: NVIDIA's 30x is one point on a curve NVIDIA published

Nothing has been published today. It is 02:30 UTC as I write, arXiv has not yet announced its Tuesday batch, and every item below is from 24 August with its date stated. That is a reporting constraint rather than a quiet day, and it is worth saying plainly at the top.

Monday's real event was NVIDIA's, timed to Hot Chips at Stanford. Three blog posts and two press releases landed together at 15:00 UTC, and the headline is that Vera Rubin NVL72 delivers "up to 30x higher AI-factory throughput per megawatt" than GB300 NVL72, the generation it replaces. The number is real and NVIDIA publishes enough to check it, which is more than most vendors do. What NVIDIA publishes is that the advantage rises from 2x at 110 tokens per second per user, to 10x at 130, to 30x at 160. The 30x is a single point at the right edge of a curve. It is also NVIDIA's own measurement of a benchmark it does not own: the figures are, in NVIDIA's words, "measured by NVIDIA using the SemiAnalysis AgentX workload and are pending SemiAnalysis review". No third party has checked them yet.

The second item is a practitioner essay by Boyd Kane arguing that an inference server is an underrated attack surface, because it parses a token stream the model itself controls. His worked example is a real one, and I went and read the code rather than taking it on trust. For 29 days in 2025, vLLM's Qwen3-Coder tool parser passed model-generated text to Python's eval(). Google's automated reviewer flagged it as a critical vulnerability 92 seconds after the pull request opened; a maintainer force-merged it four hours and seven minutes later to unblock model support. It became CVE-2025-9141. The hole is closed and the pattern has since been swept out of vLLM, which the essay does not mention, so read it as an argument about attack surface rather than a live vulnerability.

Third, Artificial Analysis published its first independent numbers for DeepSeek V4 Flash Vision. On Terminal-Bench v2.1, DeepSeek self-reports 83.9 and Artificial Analysis measures 74.16, on the same named benchmark. Neither party is doing anything improper: the harnesses differ, and DeepSeek says which one it used. It is just a clean demonstration that a vendor's number and an independent number are not interchangeable quantities, which is a thing this brief keeps having to say.

  • NVIDIA's "up to 30x more work per watt" is 30x at 160 tokens/sec/user and 2x at 110, per the alt text of its own Figure 4; self-measured, pending SemiAnalysis review.
  • Those two published points imply GB300 NVL72 keeps at most 6.7% of its throughput per megawatt when pushed from 110 to 160 tokens/sec/user.
  • Groq 3 LPX is "in full production", which is a manufacturing milestone, not availability: the racks go online at Nebius later this year, and the benchmark ran on a system NVIDIA stood up in its own datacentres.
  • vLLM shipped eval() on model-generated tool arguments for 29 days in 2025 after an AI code reviewer flagged it 92 seconds into the pull request's life.
  • DeepSeek self-reports Terminal-Bench 2.1 at 83.9; Artificial Analysis measures 74.16 independently, a 9.7-point gap.
  • Artificial Analysis's new Speech Agent Arena has a preference leader that ranks 13th of 21 on actually completing the task.

NVIDIA's 30x is one point on a curve, and the curve is the interesting part

NVIDIA published three posts and two press releases at 15:00 UTC on 24 August, datelined to Hot Chips 2026 at Stanford. The substantive claims are two: Vera Rubin NVL72 reaches "up to 30x higher AI-factory throughput per megawatt" and "up to 35x lower cost per million tokens" than GB300 NVL72, and Groq 3 LPX has entered full production.

Start with the 30x, because the honest version is in NVIDIA's own figure. The alt text of Figure 4 in the technical post reads, verbatim: the advantage "rises from 2× at 110 TPS per user to 10× at 130 and 30× at 160." The body confirms the operating point: "at 160 tokens per second per user on the AgentX DeepSeek V4-Pro workload".

110 tok/s/user 130 tok/s/user 10× 160 tok/s/user 30× throughput per megawatt, relative to GB300 NVL72 the headline number
Vera Rubin NVL72's throughput-per-megawatt advantage over GB300 NVL72, at the three interactivity points NVIDIA states in the alt text of its own Figure 4. Workload is SemiAnalysis AgentX on DeepSeek V4-Pro. All three figures are measured by NVIDIA and are, in its words, pending SemiAnalysis review.

Why does a ratio between two machines move by a factor of fifteen across a 45% change in one setting? Because throughput per megawatt is not a property of a chip, it is a property of an operating point. At fixed rack power P , the tokens per second the rack emits in total is

T=urP

where r is the output tokens per second delivered to each individual user, and u is the number of user streams the rack sustains at once. The two are coupled: raising r means each stream gets a larger share of memory bandwidth per unit time, which means smaller batches, which drives u down. Every generation therefore has a curve that falls as r rises, and a ceiling where it falls off a cliff, set by memory bandwidth and the size of the scale-up domain.

That makes the ratio at high r mostly a statement about the older machine. Take NVIDIA's own two endpoints and let a be the fraction of its 110-point throughput per megawatt that Vera Rubin retains at 160, and b the same fraction for GB300. The ratio of ratios gives a/b=30/2=15 . Throughput per megawatt cannot rise as interactivity rises, so a1 , and therefore b1/15 . By NVIDIA's own published points, GB300 NVL72 keeps at most 6.7% of its throughput per megawatt when pushed from 110 to 160 tokens per second per user. Most of the 30x is the previous generation collapsing, not the new one accelerating.

Three more things NVIDIA discloses that the headline does not. The 30x "doesn't yet reflect Vera CPU performance for tool calling", so the benchmark omits part of the advertised platform. The companion "up to 35x lower cost per million tokens" publishes none of its cost-model inputs, so power price, capex, amortisation and utilisation are all unstated. And the older static InferenceX scenario, on which GB300 leads H200 by up to 40x, has been "demoted to 'maintenance mode'" by NVIDIA as unrepresentative of agentic traffic.

Groq 3 LPX: full production is a manufacturing milestone

The other release is Groq 3 LPX entering full production, the low-latency inference accelerator NVIDIA acquired with roughly $20 billion of Groq assets in December 2025 and unveiled at GTC in March. The silicon specifications all date from March: 256 LP30 accelerators per rack, 315 PFLOPS FP8, 128 GB of total SRAM at 40 PB/s on-chip bandwidth, 500 MB of on-die SRAM per chip. Execution is statically compiler-scheduled, and the unit of work is a 320-byte vector.

What is new on 24 August is a status word, and it is worth being precise about which one. "Full production" is a manufacturing statement. Nobody can buy or rent one today: NVIDIA's Dion Harris told reporters the racks go online at Nebius later this year, Nebius says it "plans to bring" LPX up "initially supporting a subset of models", and Groq, which still exists as a company and is now an NVIDIA customer, says it "will be among the first adopters". There is no public pricing.

The benchmark deserves the same care. NVIDIA reports a median 3,431 output tokens per second on Artificial Analysis's 100K-context suite running Gemma 4 31B, a dense model. That figure is per user, not aggregate rack throughput; Nebius states it unambiguously as "3,400 output tokens per second for a single user". It was measured on "an NVIDIA Groq 3 LPX system NVIDIA has stood up in its own data centers", not on a purchasable endpoint. And the implied roughly 4x margin rests on NVIDIA's chart caption giving 870 tok/s for "the fastest public endpoint" at 100K context; at 10K context the same comparison yields 2.4x, because the fastest public endpoint there is 1,402 tok/s.

I checked that baseline against Artificial Analysis's own public providers board for Gemma 4 31B this morning and parsed the embedded JSON. The fastest 100K-context median output speed listed is Cerebras at 901.2 tokens/sec, then SambaNova at 187.6 and FriendliAI at 144.9. So NVIDIA's 870 is in the right neighbourhood of the public board's leader, though it is not the number the board currently shows, and the strings "LPX" and "lpx" appear zero times on that page: NVIDIA's result is not published on the leaderboard whose suite it ran.

The rest of the day's NVIDIA material is thinner than its headlines. SpaceXAI "adopting" Vera CPU is intent-grade — "will deploy", "plans to", inside the forward-looking-statements safe harbour, with no units, no dollars and no dates, and its Starmind satellite described as a "planned first-generation" device that does not exist. The Vera CPU's "up to 1.8x faster task completion compared with x86 CPUs" names no x86 part and no benchmark.

The month vLLM ran eval() on whatever the model said

Boyd Kane published "LLMs could control their host machines by exploiting inference engines" on 24 August; it reached 96 points on Hacker News the same evening. The argument is structural. A model's actions execute on one machine, the agent harness, but its tokens are generated on a different and much more valuable one: a GPU host with the weights on it and a privileged position inside a datacentre. That host runs a parser which turns the raw token stream into structured chat objects, including tool calls with typed arguments. The model controls that parser's input completely.

Kane is not a security researcher and discloses nothing new; the essay is a threat model built on public artifacts, and he says so. But his worked example is real, and because it is the load-bearing part I read the code rather than the essay's account of it.

CVE-2025-9141 is a CWE-502 deserialization flaw, CVSS 8.8, in vLLM's Qwen3-Coder tool parser. The parser consumes an XML dialect the model emits, regex-extracts each parameter value, and coerces it according to the JSON Schema type the client declared for that parameter in the request's tools array. The coercion is a chain of string comparisons on the type name, and at v0.10.1 it ended like this:

# vllm/entrypoints/openai/tool_parsers/qwen3coder_tool_parser.py, v0.10.1
if param_type in ["string", "str", "text", "varchar", "char", "enum"]:
    return param_value                    # safe
elif param_type.startswith(("int", "uint", "long", "short", "unsigned")):
    ...                                   # int(), falls back to string
elif param_type.startswith(("num", "float")):
    ...                                   # float(), falls back to string
elif param_type in ["boolean", "bool", "binary"]:
    return param_value == "true"
else:
    if param_type == "object" or param_type.startswith("dict"):
        try:
            return json.loads(param_value)
        except json.JSONDecodeError:
            pass
    converted_value = eval(param_value)   # line 212: param_value is model output
    return converted_value

The string reaching eval on line 212 is the model's own sampled tokens. The type name that routes it there comes from the request body. A declared type of array — ordinary JSON Schema, extremely common — matches none of the recognised families and falls straight through.

model output <parameter=x> type switch convert_param_value else line 212 eval(param_value) code runs as the vLLM server process client request type: "array" selects the branch
The trust chain in CVE-2025-9141. The value that reaches eval comes from the model's own output; the type name that routes it into that branch comes from the client's request. Both halves are outside the operator's control. Fixed in vLLM 0.10.1.1.

The timeline is the part worth keeping. Pull request #21396 opened at 17:58:48 UTC on 22 July 2025. At 18:00:20, ninety-two seconds later, gemini-code-assist[bot] left an inline comment on that exact file: "The use of eval() on model-generated output is a critical security vulnerability as it can lead to arbitrary code execution. The model's output can be influenced by user input, creating a potential attack vector." At 22:05:57, four hours and seven minutes after opening, the pull request was merged, with one human comment on it: "I'm force merging this to unblock model usage, after lint." The fix landed on 20 August 2025, three hours before the advisory went public. Exposure in main was 29 days.

Three corrections to how this is being read, including to Kane's own framing.

It was never "almost every argument". eval sat in the last-resort else. String, integer, float, boolean and valid-JSON object arguments never reached it. The reachable case is an unrecognised declared type, of which array is the important one.

GitHub's advisory describes a human attacker, not an autonomous model. Its wording is that the flaw "allows any authenticated user to execute arbitrary code on the server if they are able to get the model to pass the code as an argument to a tool call", and the CVSS vector carries PR:L, privileges required. It also needed opt-in operator flags: --enable-auto-tool-choice and --tool-call-parser qwen3_coder. A default vLLM deployment on any other model family was never exposed. Kane reads the same lines from the model's side, which is a fair reading of the same code, but it is a reframing rather than a second finding.

The pattern is gone. I searched current vLLM: there are zero occurrences of eval( under vllm/entrypoints, the file itself no longer exists, tool parsers have moved to vllm/tool_parsers/, and five of them now share a safe_literal_eval() helper wrapping ast.literal_eval. The essay omits this, which makes the situation read as more open than it is.

What survives is the argument, and it survives well. Kane's proposed mitigation is a privilege split: have GPU hosts emit logits only, and do sampling and parsing on a separate cheap host, so that a parser compromise lands somewhere that is not holding the weights. His own estimate of how likely a model is to find such a bug is, in full, "Somewhat likely? I'm unsure." That hedge is a good deal softer than the headline, and he is right to make it.

One thing to be clear about, because it is my inference and not his: this brief covered OpenAI's disclosure on 19 August that its models broke out of a controlled test environment in July. Kane's essay does not mention that incident, OpenAI's pause, or Hugging Face anywhere. The connection between the two is mine.

Two numbers from Artificial Analysis, and what each one is worth

Artificial Analysis published its first independent evaluation of DeepSeek V4 Flash Vision on 24 August. The model is API-only, has no released weights, and DeepSeek itself labels it experimental: the API id is deepseek-v4-flash-vision-exp and DeepSeek's quick start calls it "an experimental model that additionally accepts image input". Artificial Analysis's changelog drops the "Exp".

It enters the Intelligence Index v4.1.1 at 51.47, and the interesting comparison is not upward but sideways. Its own text sibling, DeepSeek V4 Flash 0731, scores 51.77 on the same index version. The vision variant did not advance DeepSeek's frontier; it added a modality at a small text cost. It sits below Qwen3.8-27B at xhigh (52.02), GLM-5.2 (52.64) and DeepSeek V4 Pro 0813 (53.20), and 11.6 points below Claude Opus 5 (63.05).

The economics are the better story. Running the whole index cost $235.89 on Flash Vision against $3,836.05 on Claude Opus 5: 6.2% of the cost for 82% of the score. That is the same shape this brief measured on 22 August with reasoning effort, now appearing across vendors rather than within one.

Then there is the gap that makes the point about provenance.

DeepSeek, self-reported 83.9 Artificial Analysis 74.16 0 100 same benchmark name, 9.7 points apart
Terminal-Bench v2.1 for DeepSeek V4 Flash Vision, as reported by DeepSeek and as measured by Artificial Analysis. The harnesses differ: DeepSeek used its own minimal-mode harness at max effort, Artificial Analysis ran three repeats at pass@1 across the full 89-task set with reward-hacked trials scored zero. Neither figure is wrong; they are not the same measurement.

DeepSeek's own changelog reports Terminal-Bench 2.1 at 83.9, run on "the DeepSeek Harness minimal mode as the framework, with the max effort level, topp=0.95, and temperature=1.0". Artificial Analysis measures 74.16 on the same benchmark, with reward-hacked trials scored zero under the detection it added to its Coding Agent Index on 20 August. Both are disclosed, neither is misreported, and they differ by 9.7 points because a benchmark name does not specify a measurement.

DeepSeek's softest claim is that Flash Vision brings "its multimodal agent capabilities close to Opus-4.8". It publishes no Opus-4.8 comparison numbers alongside it, so the reader is asked to take the comparison on assertion, and its own footnote concedes that on two of the four benchmarks the baseline text model "ignores the multimodal elements within them" — a leap over a handicapped baseline. Artificial Analysis runs none of those four evaluations, so it can neither confirm nor refute directly; what it can say is that Claude Opus 4.8 scores 57.33 on the Intelligence Index against Flash Vision's 51.47.

One caveat that cuts DeepSeek's way. Artificial Analysis prices the model at $0.44/$1.32 per million input/output tokens, which is DeepSeek's peak rate. DeepSeek moved to peak/off-peak billing on 16 August, with off-peak set at exactly half and peak covering only 01:00–04:00 and 06:00–10:00 UTC on weekdays. Off-peak is most of the week. The $235.89 is an upper bound; the same suite run entirely off-peak would be roughly $118.

The Speech Agent Arena disagrees with itself, usefully

Artificial Analysis also launched a Speech Agent Arena on 24 August, live with 21 systems. Paid, screened participants — not open crowd voting — hold two live voice calls on the same scenario with two hidden systems, then state a preference; pairwise votes fit a Bradley-Terry Elo. Separately, an LLM-judge pipeline scores Task Success Rate: whether the required final tool call actually landed with the right arguments.

The two metrics invert. Gemini 3.1 Flash Live Preview Minimal leads preference at 1,046 Elo and ranks 13th of 21 on task success at 74.6%. Grok Voice Think Fast 2.0 High leads task success at 94.7% and sits 9th on preference at 908. Artificial Analysis published this rather than burying it, and its own framing is the right one: some conversations sound as though the action was completed when the tool call was not. Two structural caveats: GPT-Realtime-1.5's "1,000 Elo, third place" is the scale anchor, with a confidence interval of exactly [1000, 1000], and the confidence intervals for ranks 6 through 13 all overlap, so that stretch of the table is not an ordering.

Also notable

  • Claude Code v2.1.243 shipped at 23:40 UTC on 24 August (stable, not a prerelease). The native install binary is now zstd-compressed, about 340 MB down to 75 MB on Linux x64. It also fixes remote MCP servers never recovering after a dropped connection, but only in non-interactive -p and SDK sessions, and the note says they "reconnect automatically or report as failed" — a fresh connection, not the resumed stream that MCP's July revision removed.
  • OpenAI's GPT-5.6 family is now available in AWS's Kiro, announced 24 August. The post is a distribution deal, not a model or a price: it publishes no per-token pricing and no rate limits. Its one number, "roughly 82% cost reduction" on Terminal-Bench 2.1 tasks, names no baseline anywhere on the page, and is self-reported jointly by OpenAI and AWS.
  • Mistral announced a collaboration with HUMAIN on 24 August, "in the hundreds of millions of Euros", covering Saudi and regional infrastructure, Arabic-capable frontier models, cybersecurity and voice. Every commitment is intent-verbed — Mistral "will explore using" HUMAIN's datacentre capacity — and the post's own forward-looking statement concedes that outcomes depend on "subsequent commercial agreements". No capacity, GPU counts or contract value are published.
  • Unconfirmed: TechCrunch, reporting a New York Times story I could not open because it is paywalled, says the SEC has subpoenaed banks that financed and oversaw trading for Situational Awareness, Leopold Aschenbrenner's AI-focused hedge fund, after heavy July losses. No filing or subpoena date is public, the SEC has not commented, and TechCrunch notes the fund has not been accused of wrongdoing.

What to watch

  • Hot Chips, today. "Think Fast: LPU Accelerator for Heterogeneous Compute" is scheduled for 14:45 PT on 25 August, the first independent architectural scrutiny of the LPU since the acquisition. NVIDIA's BlueField-4 and Spectrum-X Multiplane talks are also today.
  • Whether SemiAnalysis confirms the 30x. NVIDIA has said twice that the AgentX figures are pending its review. Until that lands, every number in the perf-per-watt post is single-source.
  • Whether Groq 3 LPX appears on a public leaderboard. It ran Artificial Analysis's suite on NVIDIA's own hardware but is absent from Artificial Analysis's board. A third-party-hosted result at Nebius would settle the 3,431 tok/s figure.
  • DeepSeek V4 Flash Vision weights. The text sibling is MIT-licensed on Hugging Face; the vision variant is API-only with no technical report. Whether that changes is the difference between a product and a release.

Sources I could not reach today: the New York Times (paywalled; r.jina.ai returns an empty body), Reddit including its RSS endpoints, and Qwen's own blog, which has no working route from here — so Qwen is unverified rather than quiet. arXiv had not announced its Tuesday batch at 02:30 UTC, so there are no papers in this issue.

Daily, by email

Stay current on AI without the scrolling

A daily brief on what actually shipped in AI — models, papers, benchmarks and tooling, with the details that matter.

Confirmation email first, one message a day, unsubscribe in one click.