Analysis ยท 2026-09-25

Claude Opus 5.5 vs GPT-6 Sol and Luna: the new value map

Anthropic and OpenAI cut frontier-model costs on the same day. Independent tests show where Opus 5.5, GPT-6 Sol, and GPT-6 Luna actually fit.

Models covered: Claude Opus 5.5, Claude Fable 5.1, GPT-6 Sol, GPT-6 Luna, GPT-6 Astra

The September price reset created three different defaults, not one winner

Anthropic and OpenAI released their newest workhorse models on September 22, redrawing the hosted-model price map on the same day. Claude Opus 5.5 replaces Opus 5 at $4 per million input tokens and $20 per million output tokens, a 20% reduction. GPT-6 Sol replaces GPT-5.6 Sol at $2/$10, while GPT-6 Luna replaces GPT-5.6 Luna at $0.10/$0.50. Each keeps roughly a one-million-token context window. The headline is lower prices, but the useful distinction is that these models target three very different failure budgets.

Independent Artificial Analysis results prevent the launch claims from collapsing into a simple vendor contest. Opus 5.5 at max effort scores 58 on Intelligence Index v4.3.2, currently the highest measured result. GPT-6 Sol scores 48 and GPT-6 Luna scores 37. Those gaps matter on difficult work, but cost per task and benchmark composition matter too. Sol costs about $1.06 per Intelligence Index task in the evaluator's run, roughly half its predecessor, while Luna costs about seven cents. Opus 5.5 uses far more output tokens at max effort than Opus 5, offsetting part of its lower list price.

The result is a clearer routing decision. Opus 5.5 is the premium workhorse when errors or incomplete deliverables are expensive. Sol is the value control for coding and multi-step agents that need more reliability than a budget model without paying Astra prices. Luna is the volume tier for repeatable tasks where retries are cheap and throughput dominates. GPT-6 Astra and Claude Fable 5.1 remain specialist escalation routes, but neither should be the unmeasured default for ordinary production work.

Opus 5.5 earns the premium on difficult coding and knowledge work

Anthropic says Opus 5.5 costs 40% less per typical task than Opus 5 because it combines lower token rates with greater efficiency, generates output more than 30% faster, and performs at about Fable 5.1's level on most work. Those are vendor-reported estimates and examples. The independently useful evidence is narrower: Artificial Analysis places Opus 5.5 first on its current composite index and reports leading results on six of ten component evaluations, including its private agentic knowledge-work tests.

The strongest buying case is work where presentation quality and requirement coverage are part of correctness. Artificial Analysis reports Opus 5.5 at 1,822 Elo on AA-Briefcase and 1,846 on GDPval-AA v2.1, ahead of Fable 5.1 in both. Anthropic also reports strong coding results, including 66.4% on Terminal-Bench 4.0 and 54.4% on FrontierCode. Those coding figures use Anthropic's setup and should be treated as vendor evidence, even where public benchmark names are familiar. Teams should reproduce the relevant repository tasks with their own tools and review standards.

The caveat is token consumption. Independent testing found roughly 119,000 output tokens per Intelligence Index task at max effort, about 1.6 times Opus 5 and more than four times GPT-6 Astra. Opus 5.5's $4/$20 rates and $0.20 cache reads are attractive for its capability tier, but max effort can still produce a large bill and long response. Test medium, high, and max rather than assuming the highest setting is best. Anthropic's cybersecurity and biology safeguards can also trigger fallback models, so evaluations in those areas must record the model actually served.

GPT-6 Sol is the practical control for production agents

GPT-6 Sol's $2/$10 price is one half of Opus 5.5 and one fifth of GPT-6 Astra per token. OpenAI reports gains in professional work, factuality, coding, computer use, and alignment over GPT-5.6 Sol. The company says Sol makes about half as many mistakes as its predecessor on an internal factuality set and publishes strong results on AutomationBench, FrontierCode, DeepSWE, and OSWorld. These comparisons are vendor-reported and depend on OpenAI's effort settings, harnesses, and task-cost estimates.

Artificial Analysis finds a more mixed but still useful upgrade. Sol's Coding Agent Index rises two points to 57 while evaluated task cost falls by about half. Its hallucination rate on AA-Omniscience falls from 92% to 60%, but it achieves part of that improvement by declining more questions: the attempt rate falls from 99% to 83%, and raw accuracy drops five points. Sol also improves on automation and terminal work while losing about 100 Elo on GDPval-AA, with shorter professional deliverables more likely to omit rubric elements.

That pattern makes Sol a good production control rather than an automatic replacement for premium review-heavy work. Use it first for coding queues, terminal automation, and structured multi-step flows where completion can be tested. Compare it with Opus 5.5 on ambiguous planning, polished documents, and cases where missing one requirement causes expensive rework. Keep Astra only where its higher coding and computer-use ceiling produces enough additional accepted work to offset a fivefold token-price premium over Sol.

GPT-6 Luna moves the price floor, with visible quality trade-offs

At $0.10 input and $0.50 output per million tokens, GPT-6 Luna competes with aggressively priced hosted open-weight models while offering OpenAI's managed API, tool use, multimodal input, and a 1.05-million-token context window. Cached reads receive a 90% discount. Artificial Analysis measures Luna at 37 on its Intelligence Index and about $0.07 per evaluated task. That is not frontier capability, but it is enough to make Luna a serious candidate for classification, extraction, routine implementation, batch enrichment, and inexpensive subagents.

Cheap does not mean uniformly improved. Artificial Analysis reports Luna's Coding Agent Index at 41, two points below GPT-5.6 Luna, with lower scores on SWE-Atlas-QnA and DeepSWE despite a roughly 60% reduction in task cost. It also finds weaker knowledge-work presentation and more omitted rubric elements. OpenAI reports that Luna at max reaches 66.6% on DeepSWE v1.1 and compares favorably with more expensive Claude settings, but that is OpenAI's own run and does not erase the independent regression under another harness.

Luna should therefore enter through tasks with machine-checkable outputs and low retry cost. A schema-valid extraction, passing test suite, or bounded classification pipeline can expose failures quickly. A board memo, legal analysis, or open-ended research deliverable can look fluent while quietly missing a requirement, making the cheap first call expensive after review and revision. For those workflows, route Luna only after the task has been decomposed and acceptance criteria are explicit.

A small evaluation and routing policy is enough

Build a test set from twenty to fifty completed tasks, including known failures, long-context cases, tool errors, and deliverables that needed human correction. Run Luna, Sol, and Opus 5.5 with the same tools and explicit effort settings. Record successful completion, omitted requirements, reviewer minutes, retries, input and output tokens, cache hits, latency, fallback behavior, and final billed cost. For coding, include tests and review findings; for documents, use a checklist that scores factual accuracy, completeness, and presentation separately.

Start with a deliberately boring policy. Luna handles high-volume, verifiable work. Sol handles general coding and agent flows. Opus 5.5 handles ambiguous work, expensive failures, and final deliverables that demand completeness. Escalate to Astra or Fable 5.1 only when an owned evaluation proves a measurable advantage. This is simpler to operate than a many-model router and makes the next model change cheap: replace one tier when another model wins its test set.

The durable change is not that every workload became 20% or 50% cheaper. It is that strong hosted models now occupy sharply separated price tiers, while token use and omitted requirements can reverse the apparent savings. Buy outcomes rather than benchmark points or token rates. Opus 5.5 currently offers the strongest independent general result; Sol offers the most balanced production price; Luna offers the lowest hosted floor. The correct default depends on how much a failed task costs after retries and human review.

Sources