BlockRun
Back to Signal
Oct 2026

GPT-6 Astra, Sol and Luna Are Live

GPT-6 Astra at $10 in and $50 out per 1M tokens, GPT-6 Sol at $2 and $10, and GPT-6 Luna at $0.10 and $0.50, live on BlockRun

All three tiers of OpenAI's GPT-6 generation are live on BlockRun, on Base and on Solana. Astra has been in the catalog since September 5; Sol and Luna joined it on September 25.

ModelInput / 1MOutput / 1MAbove 272K promptContextMax output
openai/gpt-6-astra$10.00$50.00$20.00 / $75.001M128K
openai/gpt-6-sol$2.00$10.00$4.00 / $15.001M128K
openai/gpt-6-luna$0.10$0.50$0.20 / $0.751M128K

All three take images and do reasoning. There is no API key and no subscription: every call is quoted in dollars before it runs and paid per call in USDC.

Which tier for which job

The three share a context window, an output cap and a request format, so the choice is about price per call, not about rewriting anything. Switching tiers is a one-word change to model.

Astra is the flagship, and OpenAI positions it for long-horizon agentic work and computer use: the runs where an agent plans across many steps and a wrong turn early costs more than the tokens. At $10 / $50 it is priced like it. Use it where a cheaper model's failure rate is the expensive part.

Sol is the tier below, for complex coding and agentic workflows. At $2 / $10 it is a fifth of Astra's price and half of GPT-5.6 Sol's ($4 / $20). For most coding agents this is the one to try first.

Luna is the volume tier: high-throughput chat, classification, extraction, routing decisions, the cheap first pass before a bigger model sees anything. At $0.10 / $0.50 it is half the price of GPT-5.6 Luna ($0.20 / $1.20). An agent that makes thousands of small calls a day is the agent this tier is for.

Why the price per call matters more than the price per token

An agent paying per call sees the bill one request at a time, and on this family the request is what gets priced, not just the tokens in it.

The 272K line reprices the whole request. Once a prompt passes 272K input tokens, OpenAI bills the entire request at 2x input and 1.5x output — not just the tokens above the line. A 300K-token prompt to Sol costs $4 per 1M for all 300K, not $2 for the first 272K. BlockRun bills exactly that way, because that is how it is billed upstream. If your agent accumulates context over a long run, keeping each prompt under 272K is the largest single saving available.

Reasoning tokens are output tokens. All three tiers reason, and reasoning is billed at the output rate. On Sol and Luna you can turn it off with reasoning_effort: "none", and the gateway passes that through as sent. Astra does not accept "none"; the gateway raises it to "low", the least Astra allows, rather than failing the call.

Leave room for the answer. On all three, reasoning tokens count against the same output allowance as the answer and bill as output. A small max_tokens with a high reasoning effort can end with finish_reason: "length" before any visible text. Give the call a budget that fits the reasoning, or lower the effort.

There is no silent fallback. None of the three has a fallback model. If the model is unavailable, the call returns an error rather than being served by a different model at a different price. An agent that needs a backup should name one itself.

Flex: half price for work that can wait

On Base, all three tiers accept OpenAI's Flex processing tier. Send service_tier: "flex" and the call is quoted at half the standard rate, which is what OpenAI bills for it — Sol becomes $1 / $5, Luna $0.05 / $0.25. Flex trades speed for price, so it suits batch jobs, evals, overnight summarisation and anything else where nobody is watching a spinner.

We verified Flex with paid calls on each model before turning it on, against a control without the tier, so the half-price quote reflects the tier the model actually served rather than the one that was asked for.

Calling them

The chat endpoint is OpenAI-compatible, and the SDK handles the payment. On Base:

from blockrun_llm import LLMClient

client = LLMClient()  # wallet from BLOCKRUN_WALLET_KEY or ~/.blockrun/.session
print(client.chat("openai/gpt-6-sol", "Refactor this function and explain the change."))

# Volume work: Luna, no reasoning, at the Flex rate
print(client.chat("openai/gpt-6-luna", "Classify: 'refund not received'",
                  reasoning_effort="none", service_tier="flex"))

On Solana, the same call with the Solana client:

from blockrun_llm import SolanaLLMClient

client = SolanaLLMClient()  # wallet from SOLANA_WALLET_KEY; pays USDC on Solana
print(client.chat("openai/gpt-6-astra", "Plan the migration in ordered steps."))

What the gateway handles so you don't have to

GPT-6 is stricter about request fields than the models most agent code was written against. Sent as-is, each of these would be rejected; the gateway adjusts them so a request written for GPT-5 works unchanged:

  • Tool calls. GPT-6 does not serve function tools on the chat-completions endpoint alongside reasoning. Send tools the usual way; the gateway serves the request through the endpoint that accepts them and hands you back an ordinary chat-completions response.
  • Sampling and stop parameters. temperature other than 1, top_p, frequency and presence penalties, logprobs and stop are all rejected upstream. The gateway drops them rather than failing the call; steer with the prompt instead.
  • max_tokens. GPT-6 wants max_completion_tokens; the gateway accepts either.
  • reasoning_effort: "none" on Astra. Raised to "low", as above.

One thing the gateway does not paper over: reasoning_effort: "max". OpenAI's model pages list it, but all three models reject it. The accepted values are low, medium, high and xhigh, plus none on Sol and Luna.

Each rule was measured on the model it applies to, not copied from a sibling — Astra and Sol differ on "none", and a rule copied from one to the other would have billed reasoning a caller had just declined.

Try them

The full catalog, with live prices, is at blockrun.ai/models, and the docs cover each model's request rules. Pick a tier, send a request, and the price comes back before the work does.

All articles →