AI News Roundup — Week of September 14, 2026
DeepSeek V4.1-Flash undercuts frontier pricing, Cognition raises $2B and ships SWE-2, OpenAI opens full-duplex voice to developers, and Sora’s API dies September 24.

AI News Roundup — Week of September 14, 2026
AI News Roundup — Week of September 14, 2026
After a month dominated by frontier flagship launches, this week was about the layer underneath: cheaper inference, voice, and agent tooling. Two of the largest stories came from the same company two days apart.
New models
- DeepSeek released V4.1-Flash on September 10. It is a sparse mixture-of-experts model and the first DeepSeek build on the company's Causal Encoder-Decoder architecture, activating roughly 8B parameters on input and 16B on output from a 552B-parameter backbone. It takes a context of 1,048,576 tokens, emits up to 384,000, and handles images natively. DeepSeek says V4.1-Flash beats its own V4 Pro on quality, cost and speed; third-party coverage broadly agrees it is competitive with frontier models at a fraction of the price, though leaderboard positions move weekly. Available on the DeepSeek API.
- Cognition released SWE-2 on September 10, reported to be post-trained from Kimi K3 and the first model in its SWE line with selectable reasoning-effort levels. Cognition reports 50.0% on FrontierCode 1.1 Main against a reported 50.9% for Claude Fable 5.1 and 53.3% for GPT-6 Astra, plus 92.8% on Terminal-Bench 2.1. On Terminal-Bench 4 it reports 27.3%, roughly half of either rival — worth reading before treating the headline number as parity. It shipped the same day in Devin Desktop and CLI.
- OpenAI brought GPT-Live-1 to the API on September 10. The model listens and speaks at the same time instead of chaining speech-to-text, a language model and text-to-speech, removing the turn-taking pause that makes most voice agents feel stilted. It ships with 12 voices, and OpenAI reports a 30% gain on its own full-duplex benchmark over GPT-Realtime-2.1. Deeper reasoning is delegated to whatever model you pair it with.
- OpenAI released ChatGPT Images 2.5 on September 8, adding a Sketch input that lets you draw a reference directly in the conversation, plus templates and prompt sharing. OpenAI reports up to 50% lower latency than Images 2.0 and better preservation of unedited regions across multi-turn edits. Two API models landed alongside it: GPT-Image-2.5 Flare as the default, and GPT-Image-2.5 Sunburst for slower, higher-control production work. Rolling out across ChatGPT tiers, ChatGPT Work and Codex.
Pricing & product updates
- V4.1-Flash is the week's most aggressive price. Off-peak rates are reported at $0.003 per million input tokens on a cache hit, $0.15 per million on a cache miss, and $0.60 per million output tokens, with peak rates at double those figures. For a well-cached workload that can run off-peak, the gap against US frontier pricing is now wide enough to be worth re-running your cost model over.
- DeepSeek began routing deepseek-v4-pro API requests to V4.1-Flash at 04:00 UTC on September 14, billed at Flash rates, and says this continues until V4.1-Pro ships. Anyone pinned to the V4 Pro alias is now on a different model without changing a line of code — re-run your evals.
- GPT-Live-1 is priced at $0.05 per minute for the voice layer, separate from whatever reasoning model sits behind it. Custom voices require contacting sales.
- Cognition says SWE-2 costs about 64% less than Claude Fable 5.1 at its FrontierCode comparison point, and roughly a quarter of GPT-6 Astra. That is a vendor comparison at a vendor-chosen operating point; price it against your own traffic before switching.
- Gemini Advanced added direct code-repository upload — one folder per conversation, up to 1,000 files and 100MB.
- GitHub Copilot promotional flex credits expired at the start of the month, reportedly cutting included monthly credits for Business and Enterprise seats with no change in seat price. Reported figures vary across sources; check your own billing page rather than a comparison table.
Money
- Cognition raised over $2 billion in a Series E at a $48 billion valuation on September 8, led by Andreessen Horowitz and Accel with Founders Fund, General Catalyst and Avenir participating. The company reports run-rate revenue near $900 million, up from a reported $492 million at its May round — a valuation that roughly doubled in four months. SWE-2 landed two days later.
- Gimlet Labs raised $300 million at a $3 billion valuation, led by Andreessen Horowitz with Arm and Microsoft's M12 joining, six months after an $80 million Series A. The company builds a multi-silicon inference cloud that splits phases of a single inference across different accelerators.
- Outside the US, Shenzhen-based Kinetix AI disclosed more than RMB 500 million across Angel+ financings for embodied AI, and South Korea's AIDIN Robotics closed a reported KRW 16 billion round involving HD Hyundai Robotics and Samsung Venture Investment.
- Sam Altman told Fortune that OpenAI will not go public in 2026, reportedly calling it an ill-advised moment given open safety questions — worth noting as a correction to persistent IPO-timing chatter.
Tool shake-ups worth knowing
- OpenAI's Sora API is removed on September 24 — ten days out. Every sora-2 and sora-2-pro alias stops working, and OpenAI's deprecation table lists no recommended replacement. The consumer app closed in April. If you are still calling it, this is the last practical week to migrate.
- Google Assistant is being removed from mobile, with Gemini taking over as the Assistant experience on Android; removal began September 4 across phones, tablets, Wear OS and Android Auto. Anything built against Assistant behaviour needs revisiting.
- DeepSeek V4 Pro is effectively retired for now, folded into V4.1-Flash until a V4.1-Pro exists. Treat the V4 Pro name as an alias, not a fixed target.
We are updating affected Toolglade listings as these changes take effect. Pricing, benchmark and valuation figures above come from vendor announcements and third-party reporting, both of which move fast — verify against the vendor's own pricing page before budgeting against them.