GPT-6 Astra, Sol and Luna Are Live

All three tiers of OpenAI's GPT-6 generation are live on BlockRun, on Base and on Solana. Astra has been in the catalog since September 5; Sol and Luna joined it on September 25.
| Model | Input / 1M | Output / 1M | Above 272K prompt | Context | Max output |
|---|---|---|---|---|---|
openai/gpt-6-astra | $10.00 | $50.00 | $20.00 / $75.00 | 1M | 128K |
openai/gpt-6-sol | $2.00 | $10.00 | $4.00 / $15.00 | 1M | 128K |
openai/gpt-6-luna | $0.10 | $0.50 | $0.20 / $0.75 | 1M | 128K |
All three take images and do reasoning. There is no API key and no subscription: every call is quoted in dollars before it runs and paid per call in USDC.
Which tier for which job
The three share a context window, an output cap and a request format, so the
choice is about price per call, not about rewriting anything. Switching tiers
is a one-word change to model.
Astra is the flagship, and OpenAI positions it for long-horizon agentic work and computer use: the runs where an agent plans across many steps and a wrong turn early costs more than the tokens. At $10 / $50 it is priced like it. Use it where a cheaper model's failure rate is the expensive part.
Sol is the tier below, for complex coding and agentic workflows. At $2 / $10 it is a fifth of Astra's price and half of GPT-5.6 Sol's ($4 / $20). For most coding agents this is the one to try first.
Luna is the volume tier: high-throughput chat, classification, extraction, routing decisions, the cheap first pass before a bigger model sees anything. At $0.10 / $0.50 it is half the price of GPT-5.6 Luna ($0.20 / $1.20). An agent that makes thousands of small calls a day is the agent this tier is for.
Why the price per call matters more than the price per token
An agent paying per call sees the bill one request at a time, and on this family the request is what gets priced, not just the tokens in it.
The 272K line reprices the whole request. Once a prompt passes 272K input tokens, OpenAI bills the entire request at 2x input and 1.5x output — not just the tokens above the line. A 300K-token prompt to Sol costs $4 per 1M for all 300K, not $2 for the first 272K. BlockRun bills exactly that way, because that is how it is billed upstream. If your agent accumulates context over a long run, keeping each prompt under 272K is the largest single saving available.
Reasoning tokens are output tokens. All three tiers reason, and reasoning is
billed at the output rate. On Sol and Luna you can turn it off with
reasoning_effort: "none", and the gateway passes that through as sent. Astra
does not accept "none"; the gateway raises it to "low", the least Astra
allows, rather than failing the call.
Leave room for the answer. On all three, reasoning tokens count against
the same output allowance as the answer and bill as output. A small max_tokens
with a high reasoning effort can end with finish_reason: "length" before any
visible text. Give the call a budget that fits the reasoning, or lower the effort.
There is no silent fallback. None of the three has a fallback model. If the model is unavailable, the call returns an error rather than being served by a different model at a different price. An agent that needs a backup should name one itself.
Flex: half price for work that can wait
On Base, all three tiers accept OpenAI's Flex processing tier. Send
service_tier: "flex" and the call is quoted at half the standard rate, which
is what OpenAI bills for it — Sol becomes $1 / $5, Luna $0.05 / $0.25. Flex trades
speed for price, so it suits batch jobs, evals, overnight summarisation and
anything else where nobody is watching a spinner.
We verified Flex with paid calls on each model before turning it on, against a control without the tier, so the half-price quote reflects the tier the model actually served rather than the one that was asked for.
Calling them
The chat endpoint is OpenAI-compatible, and the SDK handles the payment. On Base:
from blockrun_llm import LLMClient
client = LLMClient() # wallet from BLOCKRUN_WALLET_KEY or ~/.blockrun/.session
print(client.chat("openai/gpt-6-sol", "Refactor this function and explain the change."))
# Volume work: Luna, no reasoning, at the Flex rate
print(client.chat("openai/gpt-6-luna", "Classify: 'refund not received'",
reasoning_effort="none", service_tier="flex"))
On Solana, the same call with the Solana client:
from blockrun_llm import SolanaLLMClient
client = SolanaLLMClient() # wallet from SOLANA_WALLET_KEY; pays USDC on Solana
print(client.chat("openai/gpt-6-astra", "Plan the migration in ordered steps."))
What the gateway handles so you don't have to
GPT-6 is stricter about request fields than the models most agent code was written against. Sent as-is, each of these would be rejected; the gateway adjusts them so a request written for GPT-5 works unchanged:
- Tool calls. GPT-6 does not serve function tools on the chat-completions
endpoint alongside reasoning. Send
toolsthe usual way; the gateway serves the request through the endpoint that accepts them and hands you back an ordinary chat-completions response. - Sampling and stop parameters.
temperatureother than 1,top_p, frequency and presence penalties,logprobsandstopare all rejected upstream. The gateway drops them rather than failing the call; steer with the prompt instead. max_tokens. GPT-6 wantsmax_completion_tokens; the gateway accepts either.reasoning_effort: "none"on Astra. Raised to"low", as above.
One thing the gateway does not paper over: reasoning_effort: "max". OpenAI's
model pages list it, but all three models reject it. The accepted values are
low, medium, high and xhigh, plus none on Sol and Luna.
Each rule was measured on the model it applies to, not copied from a sibling —
Astra and Sol differ on "none", and a rule copied from one to the other would
have billed reasoning a caller had just declined.
Try them
The full catalog, with live prices, is at blockrun.ai/models, and the docs cover each model's request rules. Pick a tier, send a request, and the price comes back before the work does.





