StepFun Opens Step 5 Preview API to Business Teams
StepFun launched Step 5 Preview on September 20, opening API access to a 600-billion-parameter sparse Mixture-of-Experts model the same day as the announcement. The model activates 27 billion parameters per inference call and supports a one-million-token context window, placing it directly in competition with frontier models from OpenAI and Anthropic for production AI workloads.
Shanghai AI Lab Ships Atria Dawn: A 744B Agentic Model Anyone Can Deploy
Shanghai AI Laboratory released Atria Dawn Preview, a 744B-parameter mixture-of-experts model built for agentic, multi-step workflows. The model is available under an MIT licence, supports 1 million tokens of context, and outperforms Claude Opus 5 and GPT-5.6 Sol on several key benchmarks. Any organisation with the GPU infrastructure to host it can deploy it with no licensing negotiation.
OpenAI Launches GPT-6 Astra: Million-Token Context and 2x Faster Computer Use for Enterprise
OpenAI released GPT-6 Astra on September 3, 2026, its most capable model to date, featuring a 1.05 million token context window, 2x faster computer use, and phased rollout to enterprise customers. API pricing starts at $10 per million input tokens and $50 per million output tokens, with enterprise access disabled by default until an administrator enables it.
OpenAI's GPT-6 Astra Is Now Live for Business Users
OpenAI released GPT-6 Astra on September 3, 2026, and is rolling it out to ChatGPT Business and Enterprise users this week. The model can autonomously fill out forms, update CRM records, organise calendars, and conduct web research with roughly twice the speed of its predecessor. Enterprise admins can enable it per workspace, with access included in existing plan allowances.
Google Gemini 3.8 Flash Triples Agent Task Completions at Entry Price
Google released Gemini 3.8 Flash on 2 September 2026, delivering more than three times the task completions of its predecessor and scoring 54.9 percent on the HLE-Verified multi-step reasoning benchmark. The model is available immediately via Google AI Studio and the Gemini API at $0.75 per million input tokens, with that introductory rate expiring on 31 December 2026. A companion model, Gemini 3.8 Flash Cyber, is being rolled out to trusted security practitioners for autonomous vulnerability detection.
Anthropic Keeps Claude Sonnet 5 at Its Launch Price
Anthropic announced on 10 August 2026 that the introductory price for Claude Sonnet 5, $2 per million input tokens and $10 per million output tokens, is now its permanent standard price. A 50 per cent increase to $3/$15 per million tokens that was scheduled to take effect on 1 September 2026 will not occur. This is a direct cost saving for any business using Claude via API or building Claude-powered workflows.
DeepSeek V4-Pro Launches Adaptive Reasoning and Off-Peak Pricing for Enterprise Operators
DeepSeek released V4-Pro on August 13, 2026, introducing three-tier adaptive reasoning that lets operators dial compute effort up or down per task, and a peak/off-peak pricing model that cuts API costs in half during off-peak windows. The model is backward compatible with existing DeepSeek endpoints and natively supports the OpenAI Responses API, making it a drop-in upgrade for teams already using OpenAI-compatible tooling.
Meta Open-Sources Muse Glimmer: AI Agents on Your Own Hardware
Meta released Muse Glimmer on August 10, 2026, a 30-billion-parameter AI model that runs on a single consumer GPU under an Apache 2.0 open-source licence. The model is built for autonomous agentic tasks including coding, file management, and tool use, and operates entirely on local hardware without sending data to the cloud. Businesses with a capable workstation or high-end Mac can now deploy a powerful AI agent without ongoing cloud API costs or data-sharing agreements.
OpenAI Refreshes GPT-5.6 Sol with 68% Fewer Factual Errors and Unlimited Free Access
OpenAI rolled out a significant update to GPT-5.6 Sol on August 6, 2026, reducing factual errors by 68% compared to the previous default model according to internal testing. Simultaneously, free ChatGPT users gained unlimited text conversations with GPT-5.6 Luna, removing message caps for the first time. The dual move signals OpenAI competing on both quality and access as the AI market matures.
OpenAI Cuts GPT-5.6 Luna by 80% as AI Cost War Accelerates
On 30 July 2026, OpenAI reduced the price of its GPT-5.6 Luna model by 80%, dropping API costs from $1 to $0.20 per million input tokens and from $6 to $1.20 per million output tokens. The GPT-5.6 Terra model was cut by 20% at the same time, while the flagship Sol model held unchanged. The move arrives three weeks after the GPT-5.6 family launched, and signals that competitive pressure from global AI providers is now driving costs down faster than many businesses anticipated.
Anthropic's Opus 5: Near-Flagship Performance at Half the Cost
Anthropic released Claude Opus 5 on July 24, 2026, delivering near-flagship performance at roughly half the API cost of its previous top model, Claude Fable 5. The model introduces built-in effort toggles that let businesses dial cost up or down by task complexity, and it outperforms Fable 5 on coding and knowledge benchmarks while carrying a fresher training data cutoff of May 2026. Claude Max subscribers get access immediately with no additional charge.
Kimi K3: China's Open-Source AI Just Hit Frontier Level
Moonshot AI released Kimi K3 on 16 July 2026, a 2.8-trillion-parameter open-weights model that rivals the best proprietary models from OpenAI and Anthropic. It is the largest open-source AI model ever built, and its performance gap with closed frontier models is now smaller than at any point in AI history. Open weights are scheduled for public release on 27 July 2026, giving any organisation the ability to download, customise, and self-host a near-frontier AI system.
Google Gemini 3.5 Pro Launches With 2-Million Token Context
Google DeepMind released Gemini 3.5 Pro on 17 July 2026, the company's most capable model to date. The model ships a 2-million-token context window, double the current frontier, alongside a new Deep Think extended reasoning mode. It is available via the Gemini API and Vertex AI, with Deep Think gated behind the $250 per month Ultra subscription.
Mira Murati's Thinking Machines Releases Its First AI Model
Thinking Machines Lab, founded by former OpenAI CTO Mira Murati, released its first AI model on 15 July 2026. Named Inkling, it is an open-weight mixture-of-experts system trained natively on text, image, audio, and video. Unlike most frontier releases, Inkling is explicitly designed as a customisation starting point rather than a finished product, and organisations can download and modify it directly.
SpaceXAI Launches Grok 4.5: Opus-Class AI for Coding and Knowledge Work
SpaceXAI, the AI division formed after SpaceX absorbed xAI following its public listing as SPCX, released Grok 4.5 on July 8, 2026. The model is priced at $2 per million input tokens and $6 per million output tokens, available today inside Cursor for all plans and through the SpaceXAI console. It targets coding, agentic workflows, and knowledge work, with SpaceXAI positioning it as an Opus-class model at a fraction of comparable frontier costs.
Meta Launches Its First Paid AI Model at a Quarter of Rival Prices
Meta launched Muse Spark 1.1 on 9 July 2026, its first ever paid commercial AI model, ending the company's long-standing practice of releasing frontier AI only as free open-source software. The model is priced at $1.25 per million input tokens and $4.25 per million output tokens, significantly undercutting comparable tiers from OpenAI and Anthropic. Access is currently limited to a US-only public preview via the new Meta Model API, with a free consumer version available globally through the Meta AI app.
OpenAI Releases GPT-5.6: Three-Tier Model Family Now Public
OpenAI publicly launched its GPT-5.6 model family on 9 July 2026, offering three tiers named Sol, Terra, and Luna at different price points. The release followed approval from the US Department of Commerce after additional safety testing, and brings a clear tiered pricing structure ranging from $1 to $5 per million input tokens. Terra, the mid-tier option, matches the performance of the previous generation GPT-5.5 while costing roughly half as much.
Anthropic's Claude Sonnet 5 Is Now the Default AI for Every Free User
Anthropic released Claude Sonnet 5 on June 30, 2026 and made it the default model for all Claude Free and Pro users from July 1. It is the most agentic Sonnet ever built, benchmarks close to the flagship Opus 4.8 on key tasks, and carries introductory API pricing of $2 per million input tokens through August 31. For businesses already using Claude in any capacity, the model they are running changed without any action required on their part.
MiniMax M3 Exceeds GPT-5.5 and Gemini Benchmarks at One-Tenth the Price
Shanghai-based MiniMax launched M3 on June 1, a model that independently eclipses GPT-5.5 and Gemini 3.1 Pro on key performance benchmarks while costing between 5 and 10 percent as much. The release confirms a structural shift in the AI market: frontier-grade capability is no longer the exclusive domain of Western providers or high-cost API contracts.
Anthropic Ships Claude Opus 4.8 With Sharper Judgement and Dynamic Workflows
Anthropic released Claude Opus 4.8 on 28 May 2026 with sharper judgement, stronger coding performance, and a new Dynamic Workflows feature that orchestrates up to 1,000 parallel subagents in a single session. Pricing for the standard model is unchanged from Opus 4.7, while Fast mode is now 2.5 times faster and three times cheaper. The release lands less than two months after Opus 4.7 and reframes what a single agent run can accomplish.
OpenAI Launches GPT-5.5: First Fully Retrained Base Model Since GPT-4.5
OpenAI released GPT-5.5 on April 23, 2026, its first fully retrained base model since GPT-4.5. The model is designed to complete complex multi-step tasks with minimal human direction, operates across email, spreadsheets, calendars, and other applications, and matches GPT-5.4 latency while using significantly fewer tokens in Codex deployments.
OpenAI Launches GPT-5.5 with Stronger Agentic and Computer-Use Capabilities
OpenAI released GPT-5.5 on April 23, 2026, with significant advances in agentic coding, computer use, and long-horizon task execution. Available to Plus, Pro, Business, and Enterprise users, it carries a 1 million-token context window and is priced at $5 per million input tokens in the API. OpenAI describes it as its smartest and most intuitive model to date.
Anthropic Releases Claude Opus 4.7 with Stronger Agent and Vision Capabilities
Anthropic released Claude Opus 4.7 on April 16, 2026, its most capable commercial model to date. The release delivers significant gains in software engineering, vision, and long-running agent workflows at unchanged pricing of $5 per million input tokens and $25 per million output tokens. It is positioned just below the restricted Mythos Preview model.
DeepSeek V4 Achieves Near-Frontier Performance at $5.2M Training Cost
DeepSeek released V4, a one-trillion-parameter Mixture-of-Experts open-weights model achieving near-frontier performance for an estimated $5.2 million training cost. At $0.28 per million input tokens versus $2+ for Western flagships, it is reshaping cost assumptions for enterprise AI procurement.
Google Gemini 3.1 Pro Leads 13 of 16 Major Benchmarks at One-Third of GPT-5.4 Cost
Google Gemini 3.1 Pro leads 13 of 16 major benchmarks on the Artificial Analysis Intelligence Index and ties GPT-5.4 Pro on the overall index, while costing approximately one-third of the API price. This puts direct pressure on OpenAI enterprise pricing across cost-conscious buyer segments.
OpenAI GPT-5.4 Fully Deployed Across All Surfaces With Native Computer-Use
GPT-5.4 is now fully deployed across ChatGPT, Codex, and the OpenAI API, completing a rollout that began in March. The model introduces native computer-use capabilities, enabling agents to interact directly with desktop applications and browsers without custom integrations.
Meta Launches Muse Spark, Its First Proprietary Model From Superintelligence Labs
Meta released Muse Spark, the first model from its new Superintelligence Labs, marking a sharp pivot from open-source Llama to proprietary AI. The multimodal reasoning model uses 'thought compression' to achieve frontier performance at a fraction of the compute cost, processing text and images natively. Meta AI app downloads jumped 87% on launch day.
Google Launches Gemini 3.1 Flash-Lite at $0.25 Per Million Tokens
Google has released Gemini 3.1 Flash-Lite, its most cost-efficient AI model to date, priced at $0.25 per million input tokens, one-eighth the cost of Gemini 3.1 Pro. The model delivers 2.5 times faster responses and 45% higher output speeds than its predecessor, while supporting a one-million-token context window and multimodal inputs including text, images, audio, video, and PDFs. For operators running high-volume AI workflows, the pricing shift opens use cases that were previously too expensive to sustain.
OpenAI's GPT-5.4 Surpasses Humans at Autonomous Desktop Tasks
OpenAI launched GPT-5.4 on 5 March 2026, the company's first general-purpose model with native computer-use capabilities. The model scored 75% on the OSWorld-V benchmark, outperforming the human baseline of 72.4%, and 83% on the GDPVal benchmark for economically valuable knowledge work. It marks the clearest shift yet from AI as a conversational tool to AI as an autonomous digital coworker capable of executing multi-step tasks across software environments.
GPT-5.4 Turns ChatGPT into an Autonomous Digital Coworker
OpenAI released GPT-5.4 and GPT-5.4 Pro across ChatGPT, the API, and Codex on 17 March 2026. The model features a 1-million-token context window and can autonomously execute multi-step workflows across documents, spreadsheets, and software environments. A new Skills feature lets teams build and share reusable automations, marking a practical shift from AI as a chat assistant to AI as an autonomous digital coworker.
Meta's Llama 4 Brings Frontier AI to Self-Hosted Deployments
Meta's Llama 4 family delivers frontier-class AI capability at roughly one-ninth the per-token cost of GPT-4o, with full self-hosting support for organisations that cannot send data to third-party cloud providers. Scout and Maverick are available across AWS, Azure, and Snowflake, with dedicated deployment guides for regulated industries including finance, healthcare, and defence.
GPT-5.4 Beats the Human Baseline on Real Desktop Work
OpenAI's GPT-5.4 has become the first general-purpose AI model to score above the human baseline on OSWorld-V, a benchmark that simulates real desktop productivity tasks. Released on 5 March 2026, the model introduces native computer-use capabilities, a 1-million-token context window, and autonomous multi-step workflow execution across software environments. It is available through ChatGPT, the API, and Codex, with enterprise-grade security controls for business accounts.
GPT-5.4 Can Now Control Your Computer Autonomously
OpenAI released GPT-5.4 on 5 March 2026, the first general-use AI model with native computer-use capabilities. The model surpasses the human benchmark for real-world computer tasks and embeds directly into Excel and Google Sheets, bringing autonomous workflow execution to everyday business tools.
GPT-5.4 Launches with Native Computer Use and 1M Token Context
OpenAI launched GPT-5.4 on 5 March 2026, its most capable general-purpose frontier model to date. The release combines native computer-use capabilities with a 1-million-token context window and 33% fewer factual errors than its predecessor, and is available immediately to API developers and ChatGPT paid subscribers.
OpenAI GPT-5.4 Launches with 1M Token Context Window
OpenAI launched GPT-5.4 in three variants (Standard, Thinking, Pro) with a 1.05M-token context window and 33% fewer factual errors than GPT-5.2. API pricing starts at $2.50 per million input tokens. The extended context window allows entire contracts, codebases, or customer histories to be processed in a single API call.