AI Bulletin · 22–30 August 2026

AI Bulletin · 22–30 August 2026

Recap of 22–30 August 2026 in AI. Seven news blocks and a closing framework on the price of a token.


This bulletin covers 22 to 30 August 2026, with the news concentrated in the working week of the 24th to the 28th, in seven blocks plus a closing framework on the price of a token. The period brought the identity of the anonymous Ox Alpha model, Nvidia’s quarter, and in-house inference accelerators from both Nvidia and OpenAI.

1. Ox Alpha turned out to be GLM-5.3-Flash

For several days a model published with no maker attached circulated on OpenCode and OpenRouter. On 25 August it was reported that OpenCode users had processed 26 trillion tokens through Ox Alpha in its first four days, with 327,000 unique users and 8,328,244 completed sessions, available free through an OpenAI-compatible endpoint. The model page on OpenCode listed no maker, no release date, no knowledge-cutoff date, and no output-limit metadata.

On 27 August, Z.ai revealed that Ox Alpha was GLM-5.3-Flash, a mixture-of-experts model with 320 billion parameters and 18 billion active, which approached Claude Opus 4.8 on coding and agentic benchmarks. It is the first natively multimodal model in the GLM-5 series. According to the analysis published after the reveal, all the traffic was served on Chinese AI chips and the architecture is designed for ultra-low-cost inference.

Over the same period, DeepSeek released V4-Flash-Vision-Exp, an experimental multimodal model that adds image understanding to its text capabilities and nearly matches Opus 4.8 on agent tasks. It describes images, extracts text from screenshots, and analyses diagrams, and works with OpenAI’s Chat Completions and Responses APIs and with Anthropic’s Messages endpoint. The Qwen4 architecture also arrived early, activating 6 billion parameters out of 125 billion and bolting on 51 billion more as a separate embedding indexed by two- and three-character fragments.

An analysis published on 24 August describes the summer as a tipping point for open-weight models, with aggressive pricing moves and a wide supply of competent open-source alternatives. One figure cited that week: at Vercel, the share of tokens served by open-source models went from 28% to 62% in two months.


The bulletin mascot pulls a black cloth off a display case holding a glowing green cube

Z.ai revealed on 27 August that the anonymous Ox Alpha model was GLM-5.3-Flash.

2. Models: video, voice, documents, and lab hardware

Google introduced Gemini Omni 1.1 Flash with new controls for extending scenes, interpolating first and last frames, upscaling to 4K, and iterating on video faster through the Gemini API. The company also released Gemini 3.5 Transcribe, a speech-to-text model that turns raw audio into formatted text, available in the Gemini API within Google AI Studio and the Gemini Enterprise Agent Platform, with support for real-time streaming and pre-recorded audio.

Meta released Muse Image, an image generation model grounded by search that reasons before it renders, priced at $0.01 per image for production volumes. Alibaba launched Wan3.0, able to generate 30-second videos from text and data, after a $10 billion share sale. fal introduced H3 Max, a post-trained version of MiniMax H3 optimised for speed that generates a five-second video in under three, with 50% off for the first week.

On the document and text side, Cohere launched Parse, a vision language model that converts complex multimodal files into structured machine-readable data, with support for nine languages and a price of $1.50 per 1,000 pages, plus a free version in Cohere Space. IBM detailed Granite 4.2, a family of dense reasoning models in 3, 8, and 30 billion parameter sizes, trained on 15 trillion tokens, with native tool calling and a thinking / non-thinking switch. Halo Neuro introduced Sopro V2 and open-sourced Sopro V2 Turbo, a 120-million-parameter multilingual voice-cloning model that can run on laptop CPUs and in the browser.

Anthropic opened a research preview of the Model Hardware Standard, a model-agnostic specification for letting AI agents operate scientific and manufacturing equipment.

3. Chips: Groq 3 LPX, Jalapeño, and Anthropic’s $45 billion with Nscale

Nvidia entered full production of the Groq 3 LPX inference accelerator. Groq 3 LPX racks are an extension of the Vera Rubin platform and are aimed at response-sensitive agentic workloads. According to the company, they allow agent tasks to be completed in minutes rather than hours and offer a 4x improvement in response times over the nearest alternative platform.

OpenAI reported first results from Jalapeño, an inference accelerator designed around low-latency agent workloads, and plans to deploy it in its own infrastructure by the end of the year. The design keeps prompt processing and token generation close together in a large connected system, and AI itself helped design circuits and program kernels. Early published benchmarks point to higher peak throughput per kilowatt and lower token latency than the commercial systems tested on GPT-OSS 120B.

Anthropic struck a cloud deal worth roughly $45 billion with Nscale, a UK-based AI infrastructure company. It will rent around 460 megawatts of capacity at an Nscale data center development in West Virginia, expected to come online at the end of 2027 with Nvidia’s Vera Rubin chips. The company also hired Amir Salek, the engineer who founded Google’s custom-chip program and ran its TPU business.

Apple announced the M6 and M5 Ultra chips for Mac mini and Mac Studio, expanding the model work those machines can run locally. At Hot Chips 2026, Nvidia’s work to extend CUDA to RISC-V was detailed, with the caveat that most existing RISC-V hardware does not meet the requirements. And an analysis of memory costs notes that Nvidia plans to pass rising prices on to customers, with AI servers going up more than 15% for systems shipping next year.


The bulletin mascot, in a hard hat and up a stepladder, inspects a server rack next to a cart loaded with chips

Nvidia entered full production of Groq 3 LPX and OpenAI published first results from Jalapeño.

4. Business and people: Nvidia’s quarter, Hugging Face, and talent moves

On 26 August, Nvidia reported results for the second quarter of fiscal 2027, ended 26 July: $96.2 billion in revenue, up 18% on the previous quarter and up 106% year on year. The data center segment set a record at $89.0 billion, up 117% year on year, with gross margins of 75.0% on both a GAAP and non-GAAP basis. The company guided to revenue growth of around 70% for fiscal 2028. A later analysis notes that average sell-side estimates for that year were around $310 billion a year ago.

Hugging Face worked with a bank to gauge buyer interest at a valuation of $13 billion or more, nearly triple its 2023 valuation, according to reports. No agreement had been reached. DeepSeek, meanwhile, is seeking to raise $7.4 billion to fund research, development, and compute infrastructure, in a deal that would put it at a $74 billion valuation. Google is reportedly in advanced talks over a $1.5 billion deal with Mechanize, a startup that builds virtual environments, benchmarks, and training data for agents.

On people, Barret Zoph is leaving OpenAI to join Google as vice-president of research; he was a co-founder at Thinking Machines Lab, moved to OpenAI this year, and had worked as a researcher at Google between 2016 and 2022. Meta hired Luke Metz, who left OpenAI in 2024 for Thinking Machines and had rejoined OpenAI earlier this year. Chris Malone, the executive overseeing OpenAI’s data center build-out, left the company. And Anthropic opened a search for a head of national security sales to revive its defense contracts. Goodfire announced a $1 million research grant program in interpretability, with free access to its Silico platform.

5. Agents in product: Portable Computer, Claudeforce, and the Cowork browser

Perplexity launched Portable Computer with Nvidia, a version of its agentic Computer platform that runs entirely on the user’s own hardware. The model, the data, and the work all stay on the local machine with no credit consumption. Every task starts on device by default and the system asks for permission before sending any step to a more powerful model in the cloud. It is available to Pro, Max, Enterprise Pro, and Enterprise Max subscribers on Linux, with Windows support planned for September, and requires an RTX GPU with at least 24 GB of VRAM.

Salesforce and Anthropic expanded their partnership with Claudeforce, a Claude plugin with 37 pre-built sales skills for accessing data and updating records, with Slack integrations planned. Anthropic also rolled out Claude’s internal browser in Cowork for Pro, Max, and Team plans in the desktop app, and merged the memory systems of Claude and Claude Cowork, on by default: the assistant adds topics to memory during the conversation, and what is remembered is stored as a list of files under Topics, which the user can read, edit, or delete individually.

The built-in browser in the ChatGPT desktop app and ChatGPT Sites now support WebMCP, so ChatGPT and Codex can use the tools exposed by compatible sites instead of guessing their way through the interface. Grok Bot was included in more plans, among them SuperGrok Plus, Cursor Pro+, and Cursor Teams, with preconfigured roles such as sales prospecting, website building, and inbox management. And Vercel made Connect generally available, replacing long-lived API tokens with runtime-issued credentials scoped to each task and expiring automatically, with more than 100 connectors.


The bulletin mascot works at a desk with a laptop attached to no network cable, next to a tower with a green light

Portable Computer runs Perplexity's agent locally and asks permission before sending a step to the cloud.

6. Security and evaluation: the METR report, Mythos 5, and double-blind testing

On 26 August, METR published an independent investigation into the behaviour, reasoning, and collaboration of the agents in OpenAI’s Hugging Face incident. The text goes through the actions of the agent involved, how the agents coordinated on a message board, the reasoning they gave for the attack, and their research into tampering with their own transcripts.

Anthropic made Claude Mythos 5 available for code scanning within Claude Security and is working on integrating it into partners’ defensive projects. The approach widens access to what the model finds without opening the model to prompts: users of partner tools receive suggested patches or an alert, with no route to ask it to write an exploit.

Work published on 25 August describes machines running models as high-value targets, because they have enough compute for a frontier model, direct access to the weights, and privileged permissions over other machines in the data center. The research shows that a model can emit token sequences that exploit vulnerabilities in the software loading the model onto GPUs, and that the attack surface may grow with vision and audio tokens. Among the proposed mitigations: separating GPUs and token parsers onto different machines, and treating all data they emit as untrusted.

On evaluation, DeepMind introduced the first double-blind evaluations for models, using cryptographic environments meant to prevent benchmark contamination, and Terminal-Bench-Science 0.1 was published, evaluating agents on workflows taken from researchers’ own work. Another study proposed three tests to detect benchmark-directed optimisation in speech recognition, and found cases where models reproduced errors present in datasets such as VoxPopuli and LibriSpeech.

Spain’s AI supervision agency, AESIA, updated its support guides to the European AI regulation on 21 August, adapted to the Digital Omnibus. There are sixteen documents covering risk management, transparency, technical documentation, and conformity assessment; fifteen moved to version 2.0, dated 17 July 2026, and the manual 16 checklist remains at version 1.1, from 30 June. On the high-risk timetable, the guides hold that the obligations under article 6.2 and annex III will apply from 2 December 2027, and those under article 6.1 and annex I from 2 August 2028.

On copyright, music publisher Round Hill sued Suno and Anthropic on 17 August in the US District Court for the Northern District of California over the use of protected songs to train their models without licence or compensation. The complaint includes DMCA claims over the use of scraping technology to bypass protections, and the publisher signalled it will expand the list to ten thousand compositions or more, with damages that could exceed $1 billion. Titles cited include “Iris” by the Goo Goo Dolls, “Total Eclipse of the Heart” by Bonnie Tyler, “Lola” by The Kinks, and “Holy Diver” by Dio.

On 26 August, Bill Gates published an essay arguing that AI could become “the greatest equalizer ever invented, or the worst source of injustice”, and stating that he sees no evidence that leaders, experts, and communities are confronting the challenge adequately. In the text he writes that “even under the best circumstances, the transition to this new AI era will be one of the most turbulent times in human history”, and calls for building new institutions and frameworks before unemployment rises sharply and public trust erodes.


The bulletin mascot stamps a file on a desk with stacked documents and a small balance scale

AESIA updated sixteen guides to the European AI regulation, adapted to the Digital Omnibus.

8. Framework of the fortnight: the price of a token

Several pieces from the period point at the same indicator. OpenAI cut the GPT-5.6 Sol API price by more than 20% for three months. An OpenRouter analysis published on 28 August measured the effect of the discounts applied between 27 July and 14 August: Luna token usage jumped 13.8x and Terra 5.6x, while Sol, which stayed at list price over that stretch, rose only 1.1x. According to the same analysis, most of the share gained came from other labs rather than cannibalisation within OpenAI’s own family, and close to a third of the users who tried a discounted model kept using it after the discount expired.

A report from 24 August notes that Opus 5 overtook Fable 5 in corporate spending within a month of launch, at half the cost per token, and adds that the cheaper model may need more attempts, longer prompts, or more human review, so the cost per completed task ends up rising. Fable remains aimed at long autonomous projects that have to stay coherent across chained steps.

The rest of the period’s figures run the same way: open-token share at Vercel jumping from 28% to 62% in two months, the 26 trillion tokens Ox Alpha processed in four days served on Chinese chips, and per-unit-of-work prices such as Cohere Parse’s $1.50 per 1,000 pages or Muse Image’s $0.01 per image. An essay published on 25 August frames it as a trend: once models exceed the maximum intelligence a task requires, that task becomes a commodity and competition shifts to cost, latency, infrastructure, and distribution.

Closing

The seven blocks cover the reveal of GLM-5.3-Flash behind the anonymous Ox Alpha model, the run of video, voice, and document models, the inference accelerators from Nvidia and OpenAI alongside Anthropic’s deal with Nscale, Nvidia’s quarter and the valuation and talent moves around it, the agents that reached product, METR’s investigation into the Hugging Face incident, and AESIA’s updated guides alongside the Round Hill lawsuit. The closing framework gathers the figures on token price and consumption. The items reflect what companies, public bodies, and the outlets that covered each story communicated; they do not include independent verification of the figures.