BlockRun
Get started
Back to Pricing

GPT-OSS 120B

nvidia/gpt-oss-120b

nvidia

OpenAI's GPT-OSS 120B hosted free by NVIDIA. Hidden from public catalog over privacy (NVIDIA's free tier may use prompts for service improvement); still callable by direct ID for legacy integrations.

Code Examples

from blockrun_llm import LLMClient

client = LLMClient()  # Uses BLOCKRUN_WALLET_KEY (never sent to server)
response = client.chat("nvidia/gpt-oss-120b", "Hello!")

Pricing

InputFree / 1M tokens
OutputFree / 1M tokens
Context128K tokens
Max Output16K tokens

Free model — no payment or wallet required.

Payment

Network
Base
Currency
USDC
Protocol
x402

Pay per request with USDC on Base. No subscription required.

Try It

Send a message to try GPT-OSS 120B

Connect your wallet to enable payments

About NVIDIA GPT-OSS 120B

OpenAI's GPT-OSS 120B hosted free by NVIDIA. Hidden from public catalog over privacy (NVIDIA's free tier may use prompts for service improvement); still callable by direct ID for legacy integrations. It is built by NVIDIA and served through BlockRun's OpenAI-compatible API, which means you can call it without an account, an API key, or a subscription. Requests are paid for individually, in USDC, at the moment they are made.

What it costs

GPT-OSS 120B is free to call on BlockRun. No payment header is required, no wallet needs funding, and no API key is issued — requests are rate limited per IP rather than billed. It is a practical way to develop against the API surface before moving production traffic onto a paid model.

Specifications

Context window
128,000 tokens
Maximum output
16,384 tokens
API compatibility
OpenAI-compatible
Payment
Free — no payment required
Categories
chat, reasoning, coding

Calling it from your code

Pass nvidia/gpt-oss-120b as the model field. Because the endpoint mirrors the OpenAI chat completions schema, any existing OpenAI client works by changing the base URL — streaming, tool use, and multi-turn messages all behave the same way.

The first request comes back as an HTTP 402 carrying a signed price quote. Your client signs that quote with a wallet holding USDC and retries; the second request returns the completion, and the payment settles on-chain. Client libraries handle this handshake for you, so in practice it is a single call. See the documentation for request and response shapes, or browse the full catalog of 71 chat models to compare alternatives.