Model compare

Sol, Luna & Astra · coding

Best balance

☀ Sol medium

73.01% · $0.4196/task

Lower cost

☾ Luna max

66.59% · $0.2169/task

Highest score

☀ Sol high

75.22% · $0.6461/task

DeepSWE v1.1

☀ GPT-6.1 Sol☾ GPT-6 Luna✦ GPT-6 Astra
All published costs · shaded coding range

Astra max leads AutomationBench (41.4%). Sol high leads this coding snapshot at lower cost than every Astra setting.

Why these picks?

Cheapest within 3 score points of the best for balance; within 10 for lower cost. Custom thresholds, not guarantees.

The coding range starts 10 percentage points below the highest score and ends at the cheapest setting with that highest score. Rings and bold labels mark the three picks; other coding points are gently faded. This makes unnecessary spending visible without hiding other results.

Low latency chooses the shortest reported total response among settings within 10 index points of the best Artificial Analysis Intelligence Index score. This is a separate broad quality test. It does not measure coding-task completion. Missing speed values are excluded.

Official OpenAI

Compare benchmarks, speed, and efforts

All exact values

Frontier means no other configuration in the same dataset is both cheaper and at least as high-scoring. This does not account for speed or statistical uncertainty.

Estimate your task cost

Enter your token counts. Include reasoning tokens; they are billed as output.

☀ GPT-6.1 Sol

Loading

Estimated API token cost / attempt

☾ GPT-6 Luna

Loading

Estimated API token cost / attempt

✦ GPT-6 Astra

Loading

Estimated API token cost / attempt

Rates and formula

Current official API prices, checked October 1, 2026. USD per million tokens; standard short context:

ModelInputCache hitCache writeOutput
Sol$2.00$0.10$2.50$10.00
Luna$0.10$0.01$0.125$0.50
Astra$10.00$1.00$12.50$50.00

Cost = uncached input × input rate + cache hits × cached rate + cache writes × write rate + (answer + reasoning) × output rate, divided by one million. Cache categories partition total input, so nothing is counted twice. Above 272k input tokens, input/cache rates double and output rates multiply by 1.5 for the full request. Batch and Flex halve rates; Fast doubles them. Ultrafast is listed for Astra only, at six times standard rates. Unsupported estimates are omitted. These are API estimates, not Codex subscription charges.

Official pricing · Official changelog

What these numbers mean

Use task cost with an acceptable score. Token price alone misses how many tokens a model uses. For your own coding work, measure spend and elapsed time per accepted task, including retries and review.

Speed views use Artificial Analysis: tokens/second, first-chunk latency, or reported total response time. The leaderboard does not specify answer length for its total column. These measure API responses, not complete coding tasks. Different units cannot be averaged.

The shared-task average equally weights three benchmark families. It is a custom view, not an official coding score. Small differences may be noise; the release points do not supply confidence intervals.

No comparable ultra points are published here. Luna ultra is unsupported; Astra API supports low through max. Blank speed values are unpublished.

Sources and freshness

Artificial Analysis

Independent quality, cost, and speed measurements. Saved October 1; individual measurement dates unavailable.

Leaderboard · Methodology

Benchmark publishers

Check benchmark versions and evaluation setups before comparing results.

DeepSWE · OSWorld · LiveBench

Saved data, reviewed October 1, 2026. Reloading does not refresh it. Luna's launch computer-use results predate OpenAI's September 25 image-encoding fix.

Official API prices · Astra API · Source data · JSON

Next

Fast, source-reviewed updates after model releases. Then recommendations tested on real tasks.

Available now: free data API · Workflow planner · Roadmap