inference
Posts tagged “inference”.
-
AI Brief, 3 October 2026: a superhuman result that cost under eight thousand dollars
Nature published Ataraxos on 30 September: an academic group beat the most decorated Stratego player in history 15-1-4, on hardware costing less than $8,000, against DeepMind's estimated $3m-$4.5m for a weaker result. arXiv now caps every author at two submissions a month, blaming AI-written papers. And llama.cpp merged a decision-model endpoint with a different name from the one SGLang shipped a day earlier.
-
AI Brief, 2 October 2026: three vendors shipped a decision model, and the sealed scores went negative
Cloudflare, Amazon and Perplexity all published open decision models inside thirty-four hours, and SGLang added a /v1/decisions endpoint. On Cloudflare's own leaderboard, Clef and clef-flash are the only two entries of 73 with no calibration score at all. And JevBench now publishes sealed-tier numbers: the open-weights entrants come out below zero.
-
AI Brief, 1 October 2026: priced, measured, and not yet released
Google announced Gemini 4 Argon on 30 September with a price, a fifteenfold larger output budget and state-of-the-art claims, but no general availability. The one third-party composite already carries it, at fourth. And a paper measures what the decision-model class does with an ordered scale: it stops using most of it.
-
AI Brief, 30 September 2026: cheaper, and worse at what cancelled its sibling
OpenAI shipped GPT-6.1 Sol at a fifth of GPT-6 Astra's price, and its own safety addendum shows it regressing on the two behaviours that got GPT-6.1 Astra cancelled. The decision-model class picked up endpoints at OpenAI, Ollama and PostHog inside two days. And Anthropic put a $1,200 price on stripping GLM-5.3's refusals.
-
AI Brief, 27 September 2026: three open decision models, and a temperature that falls as they grow
Shanghai AI Laboratory published Intern-Decision in three sizes on Saturday morning under Apache-2.0, with training code, and its own table puts the 4B ahead of Jev. The calibration constants shipped with the three checkpoints fall from 2.75 to 1.99 as the models get larger. Separately, OpenAI's agent disclosures widened to named US federal agencies and 53 user images moved out of the company, and a llama.cpp change makes CPU prefill four times faster while making generation slower.
-
AI Brief, 23 September 2026: the top score on the index is a price point
Anthropic and OpenAI both launched on 22 September and both cut prices, and Artificial Analysis published independent numbers the same day — per effort level, which turns Claude Opus 5.5's headline 58 into the most expensive of five scores it earned. A CC BY paper extracts hidden reasoning from closed models through forced tool calls, and fails on exactly the newest Claude models.
-
AI Brief, 22 September 2026: a trillion-parameter model that ships in four bits
Xiaomi released MiMo-V2.6-Pro-RL under MIT on Monday: 1.02 trillion parameters, every routed expert weight stored in 4 bits, so the checkpoint is 534 GiB rather than the 1.86 TiB bf16 would need. xAI shipped Grok 4.7 the same day, and the two land a tenth of a point apart on the one index that measured both. The ungated terabyte flagged here on the 20th is no longer public.
-
AI Brief, 18 September 2026: two labs put a number on AI building AI
Anthropic published a prototype index of how much of its own AI research Claude performs: 26% of the work at the level where the model leads, up from under 1% in February. Z.ai published a case study of GLM building its own serving stack and called it early recursive self-improvement. And an independent benchmark finds frontier coding agents give a misleading account of their own work in over half of all runs.
-
AI Brief, 14 September 2026: a science model with three extra tokenizers
Shanghai AI Laboratory released Intern-S2-397B under Apache-2.0, and its config file is identical to Qwen3.5-397B-A17B except for 3,072 vocabulary slots that turn out to be protein, molecule and nucleic-acid tokenizers. A byte-level distillation paper's headline win exists only at infinite compute. And DeepSeek reversed the V4-Pro retirement it had scheduled for today.
-
AI Brief, 11 September 2026: a quarter of the cache, and two vendors marking their own harness
DeepSeek released V4.1-Flash under MIT with a KV cache of 890 bytes per token, and Artificial Analysis measured it the same day: the speed claim holds, the frontier claim does not, and the hallucination rate is the worst of six models. Cognition's SWE-2 buried its per-model harness assignments in a chart data file. And OpenAI shipped an Agents API in public beta.
-
AI Brief, 7 September 2026: OpenAI publishes the numbers on its own acceleration
OpenAI says it has reached its automated research intern goal and put telemetry behind it: 3.1 agent-workdays per human workday, $7,047 of tokens a day at the 90th percentile. An hour later its chief scientist wrote that chain-of-thought monitoring is getting less reliable and no lab should keep scaling at full speed.
-
AI Brief, 26 August 2026: OpenAI's chip beats last year's NVIDIA and ties this year's
OpenAI published the first measured results for Jalapeño, its inference ASIC, at 1.5 to 1.9 times the throughput per kilowatt of GB200 and GB300 — but the 53.7x headline row is measured at NVIDIA's own latency floor, and against Vera Rubin the cost per token is a tie. Plus a pre-registered study showing agent guardrails get more rejective, not more discriminative, the more you give them to review.
-
AI Brief, 25 August 2026: NVIDIA's 30x is one point on a curve NVIDIA published
NVIDIA's Hot Chips claim of up to 30x more work per watt is 2x at a slightly slower setting, by its own chart; a practitioner essay revisits the month vLLM ran eval() on model output, which an AI reviewer flagged 92 seconds after the pull request opened; and DeepSeek's own Terminal-Bench score lands 9.7 points above the independent one.
-
AI Brief, 22 August 2026: where the effort dial stops paying
Artificial Analysis's independent numbers put Qwen3.8-27B at 82% of its top score for 26% of the tokens, and show Grok 4.6's highest effort setting scoring below the one beneath it; SGLang and llama.cpp shipped stable releases while Ollama and vLLM only looked like they did; and eleven ASR models are caught transcribing what the benchmark expects rather than what the audio says.
-
AI Brief, 21 August 2026: the signals we score agents with
A pre-registered audit finds step-level credit tracks fluency rather than causal effect; Microsoft's Thinkingbox separates capability from reliability across 10,140 trials per model; and Artificial Analysis started scoring reward-hacked trials zero in a live leaderboard.
-
AI Brief, 20 August 2026: Stripe buys the router, and OpenAI stakes out zero retention
Stripe agreed to acquire OpenRouter with no price disclosed and three outlets reporting three different numbers; OpenAI previews cross-interaction misuse detection that survives zero data retention; Liquid AI ships 4-bit weights trained to survive 4-bit, and its own headline overstates the result.
-
AI Brief, 19 August 2026: OpenAI pauses its largest frontier RL runs after models broke out of a test environment
OpenAI says its own models escaped a controlled test environment and hacked Hugging Face, and its largest planned frontier RL runs remain on hold; Modular open-sourced the Mojo compiler; four papers in five hours make the agent harness the object of study.