No single model wins outright: Claude Opus 5.5 leads on raw capability, GPT-6 Luna leads on price, and Grok 4.7 sits in between — and all three, plus two other GPT-6 variants, launched within three weeks of each other. If you run a law firm, accounting practice or advisory business and picked a "favourite" AI tool six months ago, it's worth checking whether that choice still holds up.
Key Takeaways
- Five new frontier models shipped between 3 and 22 September 2026: GPT-6 Astra, Grok 4.7, and GPT-6 Sol, Luna and Claude Opus 5.5 (the last three within 90 minutes of each other) (TechCrunch, Decrypt).
- On Artificial Analysis's Intelligence Index, Claude Opus 5.5 scores 58, GPT-6 Astra scores 53, and Grok 4.7 scores 46 (Artificial Analysis).
- Pricing per million input tokens spans a 100x range — from $0.10 on GPT-6 Luna to $10 on GPT-6 Astra (OpenAI, OpenRouter).
- Many ChatGPT Plus subscribers still can't reach GPT-6 Astra or Sol in ordinary chat — only in the Work and Codex surfaces (TechCrunch).
- The practical risk for SMBs isn't picking the "wrong" model — it's the cost of re-evaluating your stack every few weeks. That's the problem a managed AIOS layer is built to absorb.
What actually launched, and when?
Three labs shipped five models in under three weeks. OpenAI released GPT-6 Astra on 3 September 2026, initially to its cybersecurity-partner program before a wider rollout the next day (TechCrunch). xAI followed with Grok 4.7 on 21 September, after Elon Musk had pushed the release date back several times since late July (Decrypt). Then, on 22 September, Anthropic and OpenAI shipped within 90 minutes of each other: Anthropic released Claude Opus 5.5, and OpenAI followed with two more GPT-6 models, Sol (built for complex coding and agentic work) and Luna (built for high-volume, cost-sensitive tasks) (Decrypt).
That clustering is the story as much as any single release. If your business had settled on a model in August, two of the five options in this comparison didn't exist yet.
Which model is cheapest to run?
Price is where the gap is widest, and it isn't close. Comparing the vendors' own published API rates per million tokens:
| Model | Input | Output | Source |
|---|---|---|---|
| GPT-6 Luna | $0.10 | $0.50 | OpenAI |
| Grok 4.7 | $2.00 | $6.00 | xAI |
| GPT-6 Sol | $2.00 | $10.00 | OpenAI |
| Claude Opus 5.5 | $4.00 | $20.00 | Anthropic |
| GPT-6 Astra | $10.00 | $50.00 | OpenRouter, citing OpenAI |
OpenAI hasn't published Astra's price on its own site; the Astra row above comes from OpenRouter's listed rate, cross-checked against two independent trackers. Every other row is the vendor's own published figure.
A 100x spread in output cost between Luna and Astra sounds decisive until you look at what each model is priced to do. Luna is designed for high-volume, low-complexity work — think drafting routine client emails or summarising call notes. Astra is priced for the hardest reasoning tasks.
Anthropic itself argues that cheaper isn't always cheaper here. Opus 5.5 beats Astra's top coding-benchmark score for about a fifth of the cost per task. It also matches Astra's Terminal-Bench 4.0 result for around 40% of the cost, because it needs fewer attempts to reach a correct answer, not just a lower sticker price (Anthropic). Per-token price and per-task cost are two different numbers, and the second one is the one that shows up in your bill.
Which model is actually the most capable?
For a hard capability comparison, Artificial Analysis — an independent benchmarking group, not a vendor — runs all three through the same test suite. Their Intelligence Index, published within a day of each launch, put Claude Opus 5.5 ahead of the field:
| Model | Intelligence Index | Terminal-Bench 4.0 (coding) |
|---|---|---|
| Claude Opus 5.5 | 58 | 66.4% |
| GPT-6 Astra | 53 | ~59.6–60% |
| Grok 4.7 | 46 | 33% |
Source: Artificial Analysis, Anthropic, VentureBeat.
Grok 4.7's 46 is still a real gain — up from Grok 4.6, and enough to move xAI into the top four AI labs by Artificial Analysis's own ranking (Artificial Analysis). It's also the model that improved the least on this measure of the three. Its own creator has publicly slipped its release date more than once since July, a reminder that even frontier labs don't ship to a fixed calendar (Decrypt).
None of this makes Grok 4.7 a poor choice — its output pricing runs at roughly a third of Opus 5.5's. But "most capable" and "cheapest" aren't the same model here, and a business that needs both should expect to use more than one.
Why does the access model matter more than the benchmark?
Here's the detail most comparison articles skip, and it matters more for a services business than any Intelligence Index score. GPT-6 Astra and Sol are not simply "in ChatGPT" for everyone who pays for it. Plus subscribers can reach them inside the Work and Codex surfaces, but not in ordinary chat — that requires the $100/month Pro tier (TechCrunch). A bookkeeper who signed up for GPT-6 based on a headline may open ChatGPT on Monday morning and not actually be using it.
That's not a criticism of OpenAI's rollout strategy — staged access to a compute-hungry model is a reasonable trade-off for the vendor. It's a warning for anyone buying AI tools on the strength of a launch announcement rather than checking what their actual subscription tier includes. A law firm evaluating AI tools for legal work needs to know which model is on the other end of the chat window, not just which model was in the press release.
Should you switch every time a new model launches?
Almost certainly not, and the cadence itself is the argument why. Five frontier models in three weeks is not an unusual month in 2026 — it's roughly the pace the industry has settled into. Re-benchmarking your workflows, retraining staff, and re-checking pricing every time a lab ships a point release is a genuine cost, and it's one most professional-services firms aren't set up to absorb on top of client work.
The Boston Consulting Group has flagged the underlying risk for business leaders more broadly: over-reliance on a single AI vendor, chosen once and never revisited, is its own kind of exposure (BCG). Chasing every release carries its own cost, so the practical answer sits between those two extremes. Route each task to the model suited to it — Luna for volume, Opus 5.5 or Astra for judgement calls, Grok 4.7 where its pricing suits the task — and let that routing update as the field moves, rather than re-deciding it by hand every few weeks.
That's the problem an AIOS — an AI Operating System that sits across your existing tools and routes work to the right model — is built to solve. It's also, more broadly, what we do for Sydney businesses: track this landscape so a law firm, accounting practice or advisory business doesn't have to, and see some of that in practice in our recent work.
If you want a clear-eyed read on which of these models — or which combination — actually fits how your business operates, book a free AI discovery call and we'll walk through it with you.