AI News Roundup — Week of August 24, 2026
This week in AI: OpenAI's live-video GPT-5.5 Omni, an 80% price cut on ChatGPT's free model, a wave of agent-first infrastructure, and $700M into AI inference chips.

AI News Roundup — Week of August 24, 2026
The week's theme was scale and plumbing: a new multimodal flagship, another round of price cuts, and a rush of infrastructure built specifically for AI agents. Here is what mattered for the tools people actually use.
New models
- OpenAI released GPT-5.5 Omni, a flagship that can interpret and respond to live video and audio streams at the same time — enabling things like real-time coaching during a physical task or joining a multi-person conversation with awareness of non-verbal cues. It pushes ChatGPT further into always-on, real-time assistance rather than turn-by-turn chat.
- GLM (Z.ai) shipped GLM-5.2 Turbo, a faster, cheaper variant of its open-weight coding and agentic line — the latest sign that Chinese labs keep matching frontier coding quality at a fraction of the cost.
- A stealth model nicknamed "OX Alpha" reportedly outran GPT-5.6 on several coding benchmarks and saw rapid production adoption within a day, with its maker undisclosed (reported). Treat anonymous benchmark-toppers with healthy skepticism until the source and evaluation are public.
Pricing & product updates
- The price war kept escalating. OpenAI cut its GPT-5.6 Luna tier by roughly 80% to about $0.20 per 1M input tokens and made it the default free model in ChatGPT. If you build on these APIs, "cheapest capable model" is now a monthly moving target — worth re-checking with our LLM cost calculator.
- Usage hit eye-watering scale: ChatGPT reportedly reached about 1 billion weekly active users, and Gemini crossed roughly 1 billion monthly active users. Distribution, not just model quality, is increasingly the moat.
- Agent infrastructure was everywhere. Cloudflare launched Kitesurf, a browser runtime built for AI agents that reportedly uses 3–7× less CPU and memory than Chromium; AWS made web search generally available on Bedrock AgentCore; and Google Cloud consolidated Vertex AI and Agentspace into a single Gemini Enterprise Agent Platform for building and governing agents.
- On the enterprise side, Anthropic reported growing uptake of a compliance-tuned Claude variant ("Claude 3.1 Guardian") in financial services, with checks aligned to frameworks like MiFID II — a sign of AI moving into regulated, audit-heavy workflows.
Money
- Capital kept pouring into the plumbing beneath AI rather than the models themselves. Etched raised about $700 million (roughly $1.9 billion total), led by Jane Street with Kleiner Perkins, Sequoia, Andreessen Horowitz, Tiger Global, and Bain Capital Ventures, to build specialized AI inference chips (reported). With frontier model quality converging, investors are betting on who can serve those models fastest and cheapest.
- Anthropic reported second-quarter 2026 revenue of about $10.9 billion (up ~130%) and its first operating profit of roughly $559 million — well ahead of its own timeline (reported).
- Elsewhere, Velaura AI raised a $110 million Series A at a valuation above $1 billion, and Singapore's Graas raised $17 million alongside acquiring Trustana to expand its AI ecommerce platform.
Tool shake-ups worth knowing
- The "browser for AI agents" is fast becoming its own category. Cloudflare's Kitesurf arrives just weeks after OpenAI retired its standalone Atlas browser and Perplexity made its Comet browser free — if you are building agents, expect the runtime layer to consolidate quickly.
- Stealth launches like "OX Alpha" show a strong model can grab developer mindshare overnight; benchmark leaderboards are moving faster than vendors can publish methodology, so verify claims before you migrate.
- xAI's Grok Voice TF 2.0 reportedly went live this week, pushing Grok further into real-time voice — another front in the multimodal race that GPT-5.5 Omni also opened.
Figures here are third-party and fast-moving; verify current pricing and availability directly with each vendor before relying on them.