GPT-OSS 120B
nvidia/gpt-oss-120b
OpenAI's GPT-OSS 120B hosted free by NVIDIA. Hidden from public catalog over privacy (NVIDIA's free tier may use prompts for service improvement); still callable by direct ID for legacy integrations.
Code Examples
from blockrun_llm import LLMClient
client = LLMClient() # Uses BLOCKRUN_WALLET_KEY (never sent to server)
response = client.chat("nvidia/gpt-oss-120b", "Hello!")Pricing
Free model — no payment or wallet required.
Payment
Pay per request with USDC on Base. No subscription required.
Try It
Send a message to try GPT-OSS 120B
Connect your wallet to enable payments
About NVIDIA GPT-OSS 120B
OpenAI's GPT-OSS 120B hosted free by NVIDIA. Hidden from public catalog over privacy (NVIDIA's free tier may use prompts for service improvement); still callable by direct ID for legacy integrations. It is built by NVIDIA and served through BlockRun's OpenAI-compatible API, which means you can call it without an account, an API key, or a subscription. Requests are paid for individually, in USDC, at the moment they are made.
What it costs
GPT-OSS 120B is free to call on BlockRun. No payment header is required, no wallet needs funding, and no API key is issued — requests are rate limited per IP rather than billed. It is a practical way to develop against the API surface before moving production traffic onto a paid model.
Specifications
- Context window
- 128,000 tokens
- Maximum output
- 16,384 tokens
- API compatibility
- OpenAI-compatible
- Payment
- Free — no payment required
- Categories
- chat, reasoning, coding
Calling it from your code
Pass nvidia/gpt-oss-120b as the model field. Because the endpoint mirrors the OpenAI chat completions schema, any existing OpenAI client works by changing the base URL — streaming, tool use, and multi-turn messages all behave the same way.
The first request comes back as an HTTP 402 carrying a signed price quote. Your client signs that quote with a wallet holding USDC and retries; the second request returns the completion, and the payment settles on-chain. Client libraries handle this handshake for you, so in practice it is a single call. See the documentation for request and response shapes, or browse the full catalog of 71 chat models to compare alternatives.