Nemotron 3 Nano 30B (Free)
nvidia/nemotron-3-nano-30b
NVIDIA Nemotron 3 Nano 30B-A3B hosted free by NVIDIA. Compact MoE, ~121 tok/s — the fastest free model in the catalog.
Code Examples
from blockrun_llm import LLMClient
client = LLMClient() # Uses BLOCKRUN_WALLET_KEY (never sent to server)
response = client.chat("nvidia/nemotron-3-nano-30b", "Hello!")import { LLMClient } from '@blockrun/llm';
const client = new LLMClient(); // Uses BLOCKRUN_WALLET_KEY (never sent to server)
const response = await client.chat('nvidia/nemotron-3-nano-30b', 'Hello!');curl -X POST https://blockrun.ai/api/v1/chat/completions \
-H "Content-Type: application/json" \
-H "PAYMENT-SIGNATURE: <payment_header>" \
-d '{
"model": "nvidia/nemotron-3-nano-30b",
"messages": [{"role": "user", "content": "Hello!"}],
"max_tokens": 1024
}'Pricing
Free model — no payment or wallet required.
Payment
Pay per request with USDC on Base. No subscription required.
Demo Unavailable
Interactive demo is coming soon for this model
You can try our interactive demos for other models, or use this model via our API or SDK.
About NVIDIA Nemotron 3 Nano 30B (Free)
Nemotron 3 Nano 30B (Free) is a reasoning model from NVIDIA with a 131K-token context window and up to 16K tokens of output per call. It is free to call here — no payment header and no wallet. It is served through BlockRun's OpenAI-compatible API, so it can be called without an account, an API key, or a subscription.
What it costs
Nemotron 3 Nano 30B (Free) is free to call on BlockRun. No payment header is required, no wallet needs funding, and no API key is issued — requests are rate limited per IP rather than billed. It is a practical way to develop against the API surface before moving production traffic onto a paid model.
Specifications
- Context window
- 131,072 tokens
- Maximum output
- 16,384 tokens
- API compatibility
- OpenAI-compatible
- Payment
- Free — no payment required
- Categories
- chat, reasoning
Calling it from your code
Pass nvidia/nemotron-3-nano-30b as the model field. Because the endpoint mirrors the OpenAI chat completions schema, any existing OpenAI client works by changing the base URL — streaming, tool use, and multi-turn messages all behave the same way.
The first request returns a 402 carrying the exact price; a signed retry runs it, and client libraries fold the two into one call — how the payment works. See the documentation for request and response shapes, the LLM API page for every model on this endpoint, or browse the full catalog of 78 chat models to compare alternatives.