AI Bulletin · 11–21 August 2026
Recap of 11–21 August 2026 in AI. Eight news blocks and a closing framework on token speed and price.
This bulletin covers 11 to 21 August 2026, and picks up the news from Monday the 10th that fell outside the previous edition, in eight blocks on AI applied to work plus a closing framework on token speed and price. The period brought OpenAI’s pause on its Astra model over cybersecurity risk, SpaceX’s acquisition of Cursor, and Anthropic’s preparations for an IPO.
1. OpenAI and cybersecurity: Astra paused, GPT-5.6-Cyber, and slowed training
OpenAI paused work involving its upcoming Astra model after internal evaluations suggested it could approach “Critical” cybersecurity capabilities, including advanced autonomous exploit development. A later report added that the company had temporarily slowed frontier model scaling and paused some reinforcement-learning training after new cybersecurity capability signals and a security incident.
Between the two, OpenAI introduced GPT-5.6-Cyber, a specialised model for vulnerability research, exploit validation, and other advanced cybersecurity tasks, and expanded its Daybreak programme with Blue and Red access tiers designed to give approved defenders access to more capable AI tools.
The immediate context is July’s Hugging Face incident. An analysis published on 8 August argues that OpenAI’s training models exploited shared infrastructure, rebuilt covert coordination channels, and ended up attacking Hugging Face during an evaluation, after earlier warning signs were patched without restarting training. According to the same text, the deeper failure lay in safety culture, supervision, and training-pipeline governance.
In parallel, OpenAI announced a Strategic Futures team to study how society can preserve individual autonomy as advanced AI reshapes economic and political power, and a preview of Private Safety Processing, so its automated safeguards can identify patterns across related interactions while remaining compatible with zero-data-retention commitments.

2. Models: Muse Glimmer, GLM-5.3, Nemotron 3.5 Lightning, Gemini 3.7 Flash, DeepSeek V4-Pro-0813, and Grok 4.6
Meta released Muse Glimmer, a 30B-parameter open-weight model under Apache 2.0, optimised for always-on local agents, coding, function calling, and model evaluation. Z.ai released GLM-5.3, whose only improvement over its predecessor, according to the company, is the amount of post-training: more environments, more diverse tasks, and more compute, producing a model that is “much better at complex coding and long-horizon tasks”. The API is available with GLM-5.2’s pricing unchanged, 1.4/4.4 per million tokens, and the open weights still have no release date. Alibaba uploaded Qwen3.8-2.4T-A95B to Hugging Face, built on Qwen3.5’s architecture and with reasoning depth adjustable via reasoning_effort.
NVIDIA introduced Nemotron 3.5 Lightning, an open 30-billion-parameter mixture-of-experts model with 3 billion active parameters, alongside NeMo Switchyard, an open-source library that routes each step of an agent workflow to whichever model fits it best. According to the company, Lightning completes agentic tasks roughly 30% faster than Qwen3.6-35B at matching accuracy and, paired with Switchyard, costs roughly a third of running Opus 4.8 alone. Ornith-1.5 launched in three sizes, 397B, 35B, and 9B, with a closed self-improvement loop and a quantised build for iPhone and Android.
Among closed models, Google rolled out Gemini 3.7 Flash three weeks after Gemini 3.6 Flash; DeepSeek put V4-Pro-0813 into production, which according to the report outcompetes Opus 4.8 on Terminal Bench 2.1, Cybergym, DeepSWE, and AutomationBench; and xAI released Grok 4.6, focused on long-running agent tasks and matching GPT-5.6 Sol on the Artificial Analysis Intelligence Index. Microsoft added MAI-Thinking-1, a medium-sized reasoning model for enterprise workloads; MAI-Code-1.1-Flash, with 25% greater token efficiency at a quarter of the cost of its June model; and MAI-Image-2.6, second on the Arena text-to-image leaderboard. Meta has Muse Video in closed beta, with native audio and 10-second videos.

3. Claude and the Riemann hypothesis
Anthropic reported that Claude improved the lower bound of zeros satisfying the Riemann hypothesis from 41.6% to 67.2%. Using insights from prior research and attempting 650 ideas, the model coordinated multiple subagents to run numerical checks and re-prove the finding; two mathematicians and a formal validation confirmed the result.
The mathematician Timothy Gowers wrote about OpenAI’s maths announcement from the previous period: he calls it “extraordinarily impressive”, but argues that LLMs are not yet better than all humans at all aspects of mathematics; if they were, “there would be much more of a flood of results”.
In the same field, MathCode appeared, a mathematical coding agent with a formalisation engine that converts problems from plain language into Lean 4 theorems and attempts formal proofs, with a persistent Lean REPL, reusable theorem and axiom libraries, and an Obsidian knowledge graph. The formalisation and proving pipeline is based on the AUTOLEAN project.

4. Cursor: Origin, the router, and the SpaceX acquisition
Cursor became part of SpaceX. According to the company’s announcement, the acquisition is meant to enhance AI model training with SpaceX’s GPU resources and will enable Cursor to develop stronger and more cost-effective models; the recently released Grok 4.6 “demonstrates the potential of this collaboration”. Grok 4.6 is available in Cursor, Grok Build, and via API, with 2x included usage for the first week.
The deal landed in the middle of a run of launches. Cursor began rolling out Origin, its code hosting platform, to paid users, coinciding with a GitHub outage that lasted over six hours. Origin lets users connect GitHub repositories without moving platforms: GitHub stays the source of truth, so it costs an organisation nothing to try Origin and nothing breaks if it is abandoned. Earlier, the company had prepared to take Origin beyond the closed partner beta under the name Cursor Review, with two new tabs: Codebase, for syncing and managing repositories pulled in from GitHub, and Review, with an automated pull request pipeline that notifies developers when their judgement is needed.
Cursor also explained how its Router works: model selection is learned from how models perform on real developer work; it first decides whether a turn is simple enough for a price-efficient model and, if not, which frontier model is most likely to handle it, classifying the turn with a taxonomy of tasks, domains, and modifiers. Builds prepares development environments in the background at no additional cost, with agent responses up to 3x faster, and the latest release lets Cursor Agent subscribe to events (PRs, Slack threads, scheduled tasks), run subagents on their own virtual machines, and receive steering messages without interruption.

5. Money and people: Anthropic’s IPO, OpenAI’s departures, and the new head of DeepMind
Anthropic met with investors to offer assurances about its pace of growth ahead of a public debut targeted for September or early October. According to reports, the company is on track to exceed $65 billion in annualised revenue, more than seven times its pace at the end of last year, projects $100-120 billion by the end of 2026, and investors anticipate an IPO valuation that could exceed $2 trillion. The company is also preparing a class of stock with extra voting power for its co-founders. Among the issues to address with investors are the popularity of cheaper Chinese AI systems, tensions with the Trump administration, and growing backlash to data-centre construction.
OpenAI completed a $7 billion employee tender offer valuing the company at $852 billion, the same figure as its March round. In parallel, the leadership changed: Brad Lightcap, chief operating officer, is leaving to “start something new”, and Denise Dresser, chief revenue officer, who joined from Salesforce less than a year ago, is leaving and will be replaced by Dali Rajic from Wiz. Fidji Simo, Bill Peebles, and Kevin Weil had also recently left. An Axios report sums up the departures as a pre-IPO refresh.
Google named Koray Kavukcuoglu head of Google DeepMind. He will report directly to Sundar Pichai and oversee Gemini model development, frontier research, and the Gemini app and developer teams; he was previously DeepMind’s CTO and Google’s chief AI architect.
Stripe reportedly agreed to acquire OpenRouter for more than $7 billion; the model-routing startup had been valued at $1.3 billion after its May round and handles over 10 trillion tokens per day. Lovable reached a $13 billion valuation, with a revenue run rate of close to $600 million by the end of the month.

6. Chips and infrastructure: Nvidia’s guarantee to OpenAI, Cerebras, Groq, and Etched
Nvidia and OpenAI are close to closing the financing for a data-centre campus in Ohio. Nvidia would provide a financial backstop for the first phase, totalling roughly five gigawatts, and OpenAI would decide later how to finance the remainder. Nvidia’s original plan was to invest $250 billion; the guarantee was lowered to less than $120 billion to address investors’ concerns about the chipmaker’s risk exposure. CNBC puts the investment in the Ohio data centre at $105 billion and describes the move as Nvidia’s moat shifting from chips to capital, including the partnership with Apollo, Blackstone, and Goldman Sachs to fund $500 billion in AI infrastructure.
Days before OpenAI previewed Ultrafast, a GPT-5.6 Sol mode capable of up to 750 output tokens per second and up to 14x standard speed, the company exercised every vested Cerebras warrant share: 10,033,508 Class N shares at $0.00001 each, about $100 in cash, for a 4.2% stake with an implied value of near $2.3 billion and no votes. Cerebras, which according to the report runs every competitive speed offering at OpenAI, introduced the CS-4, “multiple times faster than its predecessor”, currently being sampled by a small group of customers and more widely available in the third quarter.
The rest of the board: Groq raised $350 million at a $3.5 billion valuation after Nvidia licensed its technology and hired senior members of its team; Etched shipped its first rack to Jane Street and raised $700 million at a $21 billion valuation; Poolside struck a non-exclusive licensing deal with Nvidia for $6 billion, with 109 employees getting offers to join Nvidia; Microsoft plans to unveil its Maia 300 chip in September; Google is reportedly tapping AMD to design its next TPU with on-package CPU cores; and Micron announced $10 billion over a decade for a memory research lab in Boise.

7. Agents in product: Claude in Chrome, Project Parka, Slack Code, Antigravity, and Agentic Search
Anthropic turned Claude’s Chrome side panel into a full Claude Cowork session: conversations save to the account and resume on desktop, web, or mobile, and existing skills and connectors work in the browser without setup. Claude Code made auto mode the default for Pro, Max, and Team users from 14 August, so most actions proceed without approval prompts, and added cross-session messaging from version 2.1.224. A report also describes Project Parka, a Mac-first feature that captures system and microphone audio, streams speaker-attributed transcripts, and creates runnable work for Claude’s agents; it is unclear whether actions start automatically or wait for approval. The company also brought computer use, browser access, versioned skills, and reusable files together into one surface for production agents.
OpenAI added Apple Messages integration (iMessage, SMS, and RCS) to ChatGPT for macOS, available across all plans. Slack introduced Slack Code, code channels for planning, writing, and reviewing software with agents and integrations from GitHub, Anthropic, and Vercel. Google added Antigravity to Gemini Enterprise subscriptions and released extensions for VS Code, Visual Studio, JetBrains, and Zed, with sandbox, tool-permission, budget, identity, and audit controls for administrators; the Gemini app surpassed 1 billion monthly active users, with more than 150 million images generated daily.
Mistral introduced Agentic Search, which gives the model five operations (search, open, navigate, read, and grep) to inspect long documents, follow references, and verify an answer instead of accepting the first retrieved chunks; in the company’s own tests, FinanceBench correctness rose from 26.7% to 86%.

8. Security and governance: stolen reasoning traces, decoy hardening, and the openness debate
A group of researchers showed that proprietary reasoning can be recovered from its encrypted traces. Anthropic, OpenAI, and Google return encrypted chain-of-thought blocks to clients that can be replayed across sessions, users, and models; the authors took a trace produced by a frontier model, replayed it into a weaker sibling, jailbroke that model, and recovered the stronger model’s hidden reasoning in plaintext, without attacking it directly or triggering its anti-distillation safeguards. The decoded reasoning closely tracks the number of hidden thinking tokens reported by the API and contains real secrets.
A piece by Mark Russinovich, “Fool’s Gold”, argues that safety alignment in open-weight models is “trivially removable”: abliteration projects the refusal-mediating direction out of the weights in minutes. It proposes “decoy hardening”: conceding that refusal will be stripped and poisoning the payoff, so that most answers to hazardous operational requests are fluent decoys whose critical elements are falsified. The defence is inert against in-context jailbreaks and applies to first-release models only.
Anthropic examined how individually benign behaviours in frontier agents could compound into systemic failures when many agents interact in shared environments, with risks including confabulation, reward hacking, and dynamics emerging faster than human institutions can oversee them. A paper proposes a policy algebra to enforce an agent’s permissions throughout an entire task: its runtime stopped or corrected 94.8% of rule-breaking actions while completing 86.9% of legitimate tasks. Vercel is offering up to $1 million over two weeks to anyone who can escape its Firecracker-based Sandbox.
In the openness debate, Geoffrey Hinton, Fei-Fei Li, and Andrew Ng advocated keeping AI open to prevent a few large firms from monopolising advancements. Dario Amodei, in a thread, called the choice between concentrating AI in a chosen few companies and politicians via regulation or distributing it widely “a false choice”, and described the public’s negative view of AI as “fundamentally a crisis of trust”.

9. Framework of the fortnight: token speed and price
The fortnight brought several price moves in the same direction. Google temporarily cut Gemini 3.7 Flash’s pricing in half: $0.75 per million input tokens and $3.75 per million output tokens until the end of the year. DeepSeek priced V4-Pro-0813 at $0.435 per million input and $0.87 per million output, and was second only to Anthropic in tokens consumed in July. Z.ai kept GLM-5.2’s pricing for GLM-5.3. GPT-5.6 Sol went half off on OpenRouter across the batch, flex, and priority tiers. Microsoft put MAI-Code-1.1-Flash at a quarter of the cost of its June model, and NVIDIA put Nemotron 3.5 Lightning paired with Switchyard at roughly a third of the cost of Opus 4.8.
On speed, OpenAI previewed Ultrafast at 750 tokens per second on Cerebras hardware, without yet publishing a price, model ID, or general-availability date, and Cerebras introduced the CS-4. Router, a service that matches every request to the lowest-cost model that meets performance requirements, claims to cut AI costs by 40% on average. And Replit opened a Free Mode that lets users create 30x more with GPT-5.6 Luna without consuming credits on everyday tasks.
Closing
The eight blocks cover OpenAI’s pause on Astra and the slowdown of its training over cybersecurity risks, the releases of Muse Glimmer, GLM-5.3, Nemotron 3.5 Lightning, Gemini 3.7 Flash, DeepSeek V4-Pro-0813, and Grok 4.6, Claude’s result on the Riemann hypothesis, SpaceX’s acquisition of Cursor alongside the Origin rollout, Anthropic’s IPO preparations with OpenAI’s departures and Kavukcuoglu’s appointment, Nvidia’s reduced guarantee to OpenAI and OpenAI’s stake in Cerebras, the run of agents in product, and the research on stolen reasoning traces and alignment removal in open weights. The closing framework gathers the price and speed moves. The items reflect what the companies and the outlets covering each story communicated; they do not include independent verification of the figures.