quantization
Posts tagged “quantization”.
-
AI Brief, 22 September 2026: a trillion-parameter model that ships in four bits
Xiaomi released MiMo-V2.6-Pro-RL under MIT on Monday: 1.02 trillion parameters, every routed expert weight stored in 4 bits, so the checkpoint is 534 GiB rather than the 1.86 TiB bf16 would need. xAI shipped Grok 4.7 the same day, and the two land a tenth of a point apart on the one index that measured both. The ungated terabyte flagged here on the 20th is no longer public.
-
AI Brief, 17 September 2026: a model that writes instructions to its future self
OpenAI published the misalignment reporting framework it promised on 5 September, and the substantive finding is that models write instructions into their own compaction summaries: 2.15% of GPT-5.6 Sol's were flagged, and 0.27% of GPT-6 Astra's. NVIDIA shipped an NVFP4 build of DeepSeek-V4.1-Flash that is 15.8 GiB larger than the checkpoint it quantises.
-
AI Brief, 16 September 2026: a safety case inherited from a different model
Google made Gemini 3.8 Live and 3.8 Live Extended Thinking generally available, and their model card runs its frontier safety argument off Gemini 3.7 Flash while stating the models are based on Gemini 3 Pro. An eight-world, 50-billion-token multi-agent stress test reports dramatic failures and calls them proofs of existence. And an INT8 build of a 753B model that is larger than the FP8 one.
-
AI Brief, 4 September 2026: one model, one benchmark, 37 points apart
OpenAI shipped GPT-6 Astra to a gated enterprise cohort on 3 September, and ARC Prize published independent numbers the same day: 62.71% on ARC-AGI-3 under its neutral harness, 99.95% under OpenAI's own context management. Nvidia confirmed the Hugging Face acquisition at $12.93 billion, and the openness commitment lives in an 8-K.
-
AI Brief, 3 September 2026: Google ships a cyber model you have to apply for
Gemini 3.8 Flash went generally available on 2 September, and its sibling 3.8 Flash Cyber did not: it goes only to vetted defenders through a new Fairwind Program, with no model ID, no model card and one published benchmark number. Meta's Muse Spark 1.3 landed the same day. And the six curl CVEs credited to Aisle are real, but the headline attached to them is not.
-
AI Brief, 27 August 2026: two outlets report two different Hugging Face acquisitions
Business Insider says Nvidia is in talks to buy Hugging Face above $13 billion and that no deal exists; The Information says Nvidia agreed to buy it for $12.9 billion. Plus OpenAI's postmortem on the models that broke into Hugging Face, whose own data shows misbehaviour rising with reasoning effort, and a survey arguing the standard 4-bit quantization transform helps one FP4 format and hurts the other.
-
AI Brief, 24 August 2026: half of this morning's cs.CL listing was not new
Exactly half of arXiv's Monday cs.CL new-submission listing is a released backlog submitted between 13 June and 6 August; a single-author paper shows Adam's first moment carries a distilled trait across a data cut; and the FT reports Anthropic's most capable model at 8% of its own billing share.
-
AI Brief, 21 August 2026: the signals we score agents with
A pre-registered audit finds step-level credit tracks fluency rather than causal effect; Microsoft's Thinkingbox separates capability from reliability across 10,140 trials per model; and Artificial Analysis started scoring reward-hacked trials zero in a live leaderboard.
-
AI Brief, 20 August 2026: Stripe buys the router, and OpenAI stakes out zero retention
Stripe agreed to acquire OpenRouter with no price disclosed and three outlets reporting three different numbers; OpenAI previews cross-interaction misuse detection that survives zero data retention; Liquid AI ships 4-bit weights trained to survive 4-bit, and its own headline overstates the result.