69 models
GPT-5.6 Sol
openaiOpenAI flagship tier — deepest reasoning for complex coding, agentic workflows, and long-horizon problems. 1M context
GPT-5.6 Terra
openaiBalanced GPT-5.6 tier — everyday coding, reasoning, and agentic tasks at half the flagship price. 1M context
GPT-5.6 Luna
openaiCost-efficient GPT-5.6 tier for high-volume, latency-sensitive chat and lightweight agentic workflows. 1M context
GPT-5.6 Sol Pro
openaiHighest-capability GPT-5.6 — Sol with pro reasoning mode for the hardest problems and long-running agentic work. 1M context
GPT-5.6 Terra Pro
openaiGPT-5.6 Terra with pro reasoning mode — deeper responses on complex tasks at half the standard Terra rate. 1M context
GPT-5.6 Luna Pro
openaiGPT-5.6 Luna with pro reasoning mode — budget tier with deeper reasoning for high-volume workloads. 1M context
GPT-5.5
openaiFirst fully retrained base since GPT-4.5. 1M context, 128K output, native agent + computer use
GPT-5.5 Pro
openaiPremium GPT-5.5 with maximum compute for the hardest problems
ChatGPT Instant (GPT-5.5)
openaiChatGPT's default model — the rolling `chat-latest` alias, currently GPT-5.5 Instant. Tuned for speed and concision, same price as GPT-5.5
GPT-5.4
openaiMost capable and efficient frontier model with 1M context, native computer use, and thinking mode
GPT-5.4 Pro
openaiPremium GPT-5.4 with maximum compute for the hardest problems
GPT-5.3
openaiHigh intelligence with medium speed. Multimodal with vision, function calling, and structured outputs
GPT-5.2
openaiFrontier model with 400K context and adaptive reasoning
GPT-5.4 Mini
openaiStrongest mini model for coding, computer use, and subagents with GPT-5.4 capabilities
GPT-5 Mini
openaiCost-optimized reasoning and chat
GPT-5.4 Nano
openaiFastest and most affordable GPT-5.4 model for high-throughput tasks
GPT-5.2 Pro
openaiUses more compute for consistently better answers
GPT-5.3 Codex
openaiIndustry-leading agentic coding model. 400K context, reasoning, tool use, and complex execution
GPT-4.1
openaiLatest GPT-4 generation model
GPT-4.1 Mini
openaiFast and affordable GPT-4.1 model
GPT-4.1 Nano
openaiUltra-fast and cost-effective GPT-4.1
GPT-4o
openaiMultimodal model with vision and audio
GPT-4o Mini
openaiFast and affordable GPT-4o model
o1
openaiAdvanced reasoning model for complex tasks
o3
openaiLatest reasoning model with improved performance
o3-mini
openaiEfficient reasoning model for STEM tasks
o4-mini
openaiLatest generation efficient reasoning model
Claude Haiku 4.5
anthropicFastest and most efficient Claude, near-frontier intelligence
Claude Sonnet 5
anthropicNewest Sonnet — near-Opus coding/agentic quality at Sonnet cost. 1M context, 128k output, adaptive thinking, vision
Claude Sonnet 4.6
anthropicBest balance of intelligence, speed, and cost
Claude Sonnet 4.5
anthropicSonnet 4.5 — strong coding and agentic performance, vision
Claude Opus 4.5
anthropicLatest Anthropic flagship with enhanced reasoning and creativity
Claude Opus 4.7
anthropicPowerful Claude Opus for complex reasoning and agentic coding. 1M context, 128k output, adaptive thinking
Claude Fable 5
anthropicAnthropic's most capable model — Mythos-class tier above Opus, for the most demanding reasoning and long-horizon agentic work. 1M context, 128K output, always-on thinking
Claude Opus 4.8
anthropicMost capable Claude 4-series Opus for complex reasoning and agentic coding. 1M context, 128k output, adaptive thinking
Claude Opus 5
anthropicNewest Opus — step-change over Opus 4.8 for deep reasoning and agentic coding at the same price. 1M context, 128k output, adaptive thinking
Gemini 3.1 Pro
googleLatest Gemini with improved thinking, token efficiency, and agentic capabilities. Optimized for software engineering (requires new SDK)
Gemini 3 Flash Preview
googleFrontier-class performance with Pro-level intelligence at Flash speed and pricing. Includes thinking mode (requires new SDK)
Gemini 3.6 Flash
googleNewest-generation Flash with built-in thinking mode — frontier-class quality at Flash speed
Gemini 3.5 Flash
googleLatest-generation Flash with built-in thinking mode — frontier-class quality at Flash speed
Gemini 2.5 Pro
googleState-of-the-art for reasoning, coding, and mathematics
Gemini 2.5 Flash
googleFast and efficient Gemini model with vision support
Gemini 3.5 Flash Lite
googleLatest Flash Lite — ultra-fast, lightweight Gemini with thinking mode for high-throughput tasks
Gemini 3.1 Flash Lite
googleUltra-fast and lightweight Gemini 3.1 model with thinking mode for high-throughput tasks
Gemini 2.5 Flash Lite
googleMost economical Gemini model - ultra-fast and lightweight (requires new SDK)
DeepSeek V4 Pro
deepseekDeepSeek V4 flagship — 1.6T MoE / 49B active, 1M context. Strongest open-weight reasoner. Thinking mode default.
DeepSeek V4 Flash Chat
deepseekPaid V4 Flash in non-thinking mode (1.6T-class quality at $0.14 in / $0.28 out). Production-grade reliability and 5MB request bodies.
DeepSeek V4 Flash Reasoner
deepseekPaid V4 Flash in thinking mode for reasoning tasks. Same upstream as deepseek/deepseek-chat but with thinking enabled by default.
Kimi K3
moonshotMoonshot's flagship — a 2.8-trillion-parameter open MoE with 1M context, image + text input, returning reasoning_content. Live-verified 2026-07-17 (chat, tools, vision).
GLM-5.3
zaiZ.AI's flagship — 1M-token context with always-on reasoning, strong at long-horizon coding. Verified live on Z.AI.
GLM-5.2
zaiZ.AI GLM-5.2 — 1M-token context, strong open-source long-horizon coding. Verified live on Z.AI.
GLM-5.1
zaiZ.AI flagship — #1 open source on SWE-Bench Pro, 8-hour autonomous execution. 200K context
GLM-5
zaiZ.AI's foundation model with 200K context. Strong reasoning and agentic capabilities
GLM-5 Turbo
zaiOptimized GLM-5 variant with faster inference
Grok 4.3
xaixAI's Grok 4.3 reasoning model. 1M context, vision-capable, tuned for agentic workflows and instruction-following.
Grok Build 0.1
xaixAI's fast agentic coding model, trained for interactive software-engineering workflows. 256K context, text + image input.
Grok 4.5
xaixAI's flagship Grok 4.5 — their most intelligent and fastest model. 500K context, vision-capable, chain-of-thought reasoning. Supports Live Search (+$0.025/source)
MiniMax M2.7
minimaxMiniMax's flagship reasoning model with recursive self-improvement. Great value for complex tasks (~60 tps)
MiniMax M3
minimaxMiniMax's M3 flagship — 1M context, strong reasoning + coding.
Qwen3.7 Max
qwenAlibaba's Qwen flagship — the Max tier. 1M context, strong reasoning, coding, and agentic tool use. Live-verified 2026-07-20.
Qwen3.7 Plus
qwenAlibaba's balanced Qwen tier — 1M context with reasoning, coding, and agentic tool use at a fraction of the Max price
Qwen3.7 Flash
qwenAlibaba's fastest Qwen tier — 1M context reasoning for high-volume, latency-sensitive workloads
Tencent Hy3
tencentTencent's Hy3 — fast, inexpensive reasoning at 262K context. One of the most-used open models of 2026.
Xiaomi MiMo-V2.5 Pro
xiaomiXiaomi's MiMo-V2.5 Pro — 1M context reasoning model, priced well below the frontier tier.
Nemotron 3 Nano Omni (Free)
nvidiaFreeNVIDIA's multimodal reasoning Nemotron Nano Omni hosted free by NVIDIA. 31B / 3.2B active MoE. Accepts text, images, video, audio. ChartQA 90.3, DocVQA 95.6, MMMU 70.8 — the only vision-capable free model in our catalog
Mistral Nemotron (Free)
nvidiaFreeMistral × NVIDIA Nemotron instruction model hosted free by NVIDIA. Fast (~0.2s), strong instruction following.
StepFun Step 3.7 Flash (Free)
nvidiaFreeStepFun Step 3.7 Flash hosted free by NVIDIA. Fast lightweight reasoning, 131K context.
Nemotron Nano 9B v2 (Free)
nvidiaFreeNVIDIA Nemotron Nano 9B v2 hosted free by NVIDIA. Compact + fast (~0.7s), good for high-volume light tasks.
Nemotron Nano 12B v2 VL (Free)
nvidiaFreeNVIDIA Nemotron Nano 12B v2 Vision-Language hosted free by NVIDIA. Accepts images; compact + fast.