Infer

74 models

Claude Fable 5

Anthropic's high-capability Fable 5 model for advanced reasoning, repository-scale coding, long-context analysis, and agentic workflows.

Anthropic·1M context·700 Credits / 1M in·3,500 Credits / 1M out·70 Credits / 1M cache read·875 Credits / 1M cache write

Claude Haiku 4.5

Default Claude Haiku route for fast background work, summaries, and lightweight agent loops.

Anthropic·1M context·70 Credits / 1M in·350 Credits / 1M out·7 Credits / 1M cache read·87.5 Credits / 1M cache write

Claude Opus 4.6

Default Claude Opus route for complex planning and high-quality reasoning.

Anthropic·1M context·350 Credits / 1M in·1,750 Credits / 1M out·35 Credits / 1M cache read·437.5 Credits / 1M cache write

Claude Opus 4.7

Claude Opus 4.7 route for advanced reasoning, coding, and long-context agent work.

Anthropic·1M context·350 Credits / 1M in·1,750 Credits / 1M out·35 Credits / 1M cache read·437.5 Credits / 1M cache write

Claude Opus 4.8

Claude Opus 4.8 route for high-quality reasoning, repository-scale coding, and agentic workflows.

Anthropic·1M context·350 Credits / 1M in·1,750 Credits / 1M out·35 Credits / 1M cache read·437.5 Credits / 1M cache write

Claude Opus 5

Default Claude Opus 5 route for the strongest planning, coding, and long-running agent work.

Anthropic·1M context·350 Credits / 1M in·1,750 Credits / 1M out·35 Credits / 1M cache read·437.5 Credits / 1M cache write

Claude Sonnet 4.6

Default Claude Sonnet route for balanced coding, analysis, and agent traffic.

Anthropic·1M context·210 Credits / 1M in·1,050 Credits / 1M out·21 Credits / 1M cache read·262.5 Credits / 1M cache write

DeepSeek Chat V3 0324

DeepSeek V3 chat snapshot for dependable coding, math, and general reasoning.

DeepSeek·164K context·20 Credits / 1M in·77 Credits / 1M out·13.5 Credits / 1M cache read·20 Credits / 1M cache write

DeepSeek Chat V3.1

DeepSeek V3.1 chat model for balanced quality, cost, and coding support.

DeepSeek·33K context·15 Credits / 1M in·75 Credits / 1M out·15 Credits / 1M cache read·15 Credits / 1M cache write

DeepSeek V3.1 Terminus

DeepSeek V3.1 Terminus model for stronger chat and technical reasoning workloads.

DeepSeek·164K context·21 Credits / 1M in·79 Credits / 1M out·13 Credits / 1M cache read·21 Credits / 1M cache write

DeepSeek V3.2

DeepSeek V3.2 served through Qianfan International for high-ROI chat, coding, math, and reasoning workloads.

DeepSeek·144K context·35 Credits / 1M in·105 Credits / 1M out·30 Credits / 1M cache read

DeepSeek V3.2 Exp

Experimental DeepSeek V3.2 model for previewing newer reasoning and chat behavior.

DeepSeek·164K context·27 Credits / 1M in·41 Credits / 1M out·27 Credits / 1M cache read·27 Credits / 1M cache write

DeepSeek V4 Flash

Lower-latency DeepSeek V4 route for cost-sensitive chat and structured generation.

DeepSeek·1M context·Official time-window pricing (UTC):Off-peak:22 Credits / 1M in·66 Credits / 1M out·0.7 Credits / 1M cache read·22 Credits / 1M cache write·Peak:44 Credits / 1M in·132 Credits / 1M out·1.4 Credits / 1M cache read·44 Credits / 1M cache write

DeepSeek V4 Pro

Infer DeepSeek Pro route for coding, math, and complex analysis with official DeepSeek primary routing and provider fallbacks.

DeepSeek·1M context·Official time-window pricing (UTC):Off-peak:66 Credits / 1M in·198 Credits / 1M out·2.2 Credits / 1M cache read·66 Credits / 1M cache write·Peak:132 Credits / 1M in·396 Credits / 1M out·4.4 Credits / 1M cache read·132 Credits / 1M cache write

ERNIE 5.0

Baidu's flagship Qianfan International model for text chat plus verified vision input with image URLs, Base64 images, multi-turn image conversations, and detail controls.

Baidu·128K context·98 Credits / 1M in·392 Credits / 1M out

Gemini 2.5 Flash

Discounted Gemini Flash route for fast multimodal-capable assistant traffic.

Google·1M context·7.5 Credits / 1M in·62.55 Credits / 1M out·1.875 Credits / 1M cache read

Gemini 2.5 Flash Image

Gemini image-capable route for visual prompts and image-token billing.

Google·1M context·7.5 Credits / 1M in·750 Credits / 1M out

Gemini 3 Pro Image Preview

Gemini 3 Pro image route for higher-fidelity visual generation and multimodal prompts.

Google·1M context·80 Credits / 1M in·4,800 Credits / 1M out

Gemini 3.1 Flash Image Preview

Gemini 3.1 Flash image route for low-latency visual generation and image-token billing.

Google·1M context·20 Credits / 1M in·600 Credits / 1M out

Gemma 3 12B IT

Small Gemma 3 instruction model for very low-cost chat and classification tasks.

Google·131K context·4 Credits / 1M in·13 Credits / 1M out·4 Credits / 1M cache read·4 Credits / 1M cache write

Gemma 3 27B IT

Instruction-tuned Gemma 3 model for lightweight assistants and text generation.

Google·131K context·8 Credits / 1M in·16 Credits / 1M out·8 Credits / 1M cache read·8 Credits / 1M cache write

Gemma 4 26B A4B IT

Compact instruction-tuned Gemma model for fast and low-cost production use.

Google·262K context·13 Credits / 1M in·40 Credits / 1M out·13 Credits / 1M cache read·13 Credits / 1M cache write

Gemma 4 31B IT

Instruction-tuned Gemma model for economical chat, coding support, and analysis.

Google·262K context·13 Credits / 1M in·38 Credits / 1M out·13 Credits / 1M cache read·13 Credits / 1M cache write

GLM-4.5 Air

Efficient GLM model for low-cost bilingual chat, summarization, and extraction.

Zhipu·131K context·13 Credits / 1M in·85 Credits / 1M out·2.5 Credits / 1M cache read·13 Credits / 1M cache write

GLM-4.6

Earlier GLM generation with dependable bilingual chat and instruction following.

Zhipu·205K context·39 Credits / 1M in·190 Credits / 1M out·39 Credits / 1M cache read·39 Credits / 1M cache write

GLM-4.7

Zhipu GLM model for balanced reasoning, coding support, and Chinese-English tasks.

Zhipu·203K context·40 Credits / 1M in·175 Credits / 1M out·8 Credits / 1M cache read·40 Credits / 1M cache write

GLM-4.7 Flash

Fast GLM flash model for inexpensive chat, classification, and structured output.

Zhipu·203K context·6 Credits / 1M in·40 Credits / 1M out·1 Credits / 1M cache read·6 Credits / 1M cache write

GLM-5

Zhipu AI's latest foundation model with strong bilingual capabilities and advanced reasoning for enterprise applications.

Zhipu·200K context·55 Credits / 1M in·176 Credits / 1M out·11 Credits / 1M cache read·55 Credits / 1M cache write

GLM-5.1

Newer Zhipu flagship with stronger bilingual reasoning and enterprise-oriented instruction following.

Zhipu·200K context·77 Credits / 1M in·242 Credits / 1M out·14.3 Credits / 1M cache read·77 Credits / 1M cache write

GLM-5.2

Z.ai flagship route for long-context bilingual reasoning and agent workloads.

Zhipu·1M context·98 Credits / 1M in·308 Credits / 1M out·18.2 Credits / 1M cache read·98 Credits / 1M cache write

GPT Image 2

OpenAI image generation route for prompt-driven image creation and image-token billing.

OpenAI·1M context·200 Credits / 1M in·1,200 Credits / 1M out

GPT OSS 120B

Large open-weight GPT OSS model for economical reasoning, coding, and general chat.

OpenAI·131K context·3.9 Credits / 1M in·19 Credits / 1M out·3.9 Credits / 1M cache read·3.9 Credits / 1M cache write

GPT OSS 20B

Smaller open-weight GPT OSS model for lightweight chat and high-volume routing.

OpenAI·131K context·3 Credits / 1M in·14 Credits / 1M out·3 Credits / 1M cache read·3 Credits / 1M cache write

GPT-5.3 Codex

Default Codex-oriented GPT route for developer agents and code tasks.

OpenAI·400K context·140 Credits / 1M in·1,120 Credits / 1M out·14 Credits / 1M cache read·175 Credits / 1M cache write

GPT-5.4

Default GPT-5.4 route for cost-sensitive production agents.

OpenAI·1M context·200 Credits / 1M in·1,200 Credits / 1M out·20 Credits / 1M cache read·250 Credits / 1M cache write

GPT-5.5

Premium GPT-5.5 route for high-capability agent planning, coding, and analysis.

OpenAI·1M context·400 Credits / 1M in·2,400 Credits / 1M out·40 Credits / 1M cache read·500 Credits / 1M cache write

GPT-5.6 Luna

Default GPT-5.6 Luna route for cost-sensitive agent traffic and high-volume reasoning workloads.

OpenAI·1M context·Tiered pricing:0-272K:14 Credits / 1M in·84 Credits / 1M out·1.4 Credits / 1M cache read·17.5 Credits / 1M cache write·272K-1M:28 Credits / 1M in·126 Credits / 1M out·2.8 Credits / 1M cache read·35 Credits / 1M cache write

GPT-5.6 Sol

High-capability GPT-5.6 route for complex coding, review, and agent planning.

OpenAI·1M context·350 Credits / 1M in·2,100 Credits / 1M out·35 Credits / 1M cache read·437.5 Credits / 1M cache write

GPT-5.6 Terra

Default GPT-5.6 Terra route for stronger reasoning, coding, and agentic workloads.

OpenAI·1M context·Tiered pricing:0-272K:140 Credits / 1M in·840 Credits / 1M out·14 Credits / 1M cache read·175 Credits / 1M cache write·272K-1M:280 Credits / 1M in·1,260 Credits / 1M out·28 Credits / 1M cache read·350 Credits / 1M cache write

Grok 4.1 Fast

Default Grok route for low-latency chat and production traffic.

xAI·2M context·13 Credits / 1M in·32.5 Credits / 1M out·3.25 Credits / 1M cache read

Kimi K2.5

Moonshot AI's powerful model with strong long-context understanding and reasoning. Excels at document analysis and complex tasks.

Moonshot·256K context·42 Credits / 1M in·210 Credits / 1M out·7 Credits / 1M cache read·42 Credits / 1M cache write

Kimi K2.6

Updated Moonshot model with stronger long-context reasoning and higher output quality for research and document-heavy workflows.

Moonshot·256K context·66.5 Credits / 1M in·280 Credits / 1M out·11.2 Credits / 1M cache read·66.5 Credits / 1M cache write

Kimi K3

Moonshot's flagship long-horizon coding and knowledge-work model with a 1M-token context window and always-on reasoning.

Moonshot·1.0M context·210 Credits / 1M in·1,050 Credits / 1M out·21 Credits / 1M cache read·210 Credits / 1M cache write

Llama 3.1 70B Instruct

Larger Llama instruct model for stronger open-weight reasoning and generation.

Meta·131K context·40 Credits / 1M in·40 Credits / 1M out·40 Credits / 1M cache read·40 Credits / 1M cache write

Llama 3.1 8B Instruct

Small Llama instruct model for inexpensive chat, classification, and extraction.

Meta·16K context·2 Credits / 1M in·5 Credits / 1M out·2 Credits / 1M cache read·2 Credits / 1M cache write

Llama 3.3 70B Instruct

Updated 70B Llama instruct model for general-purpose chat and reasoning.

Meta·66K context·10 Credits / 1M in·32 Credits / 1M out·10 Credits / 1M cache read·10 Credits / 1M cache write

Llama 4 Maverick

Llama 4 model for long-context open-weight chat and multimodal-adjacent workflows.

Meta·1M context·15 Credits / 1M in·60 Credits / 1M out·15 Credits / 1M cache read·15 Credits / 1M cache write

MiMo V2 Flash

Xiaomi MiMo flash model for high-throughput, latency-sensitive production traffic.

Xiaomi·262K context·9 Credits / 1M in·29 Credits / 1M out·4.5 Credits / 1M cache read·9 Credits / 1M cache write

MiMo V2.5

General-purpose MiMo model with 1M context and tiered pricing for longer prompts.

Xiaomi·1M context·Tiered pricing:0–256K:40 Credits / 1M in·200 Credits / 1M out·8 Credits / 1M cache read·40 Credits / 1M cache write·256K–1M:80 Credits / 1M in·400 Credits / 1M out·16 Credits / 1M cache read·80 Credits / 1M cache write

MiMo V2.5 Pro

Updated MiMo pro model for broad long-context reasoning and production assistants.

Xiaomi·1M context·Tiered pricing:0–256K:100 Credits / 1M in·300 Credits / 1M out·20 Credits / 1M cache read·100 Credits / 1M cache write·256K–1M:200 Credits / 1M in·600 Credits / 1M out·40 Credits / 1M cache read·200 Credits / 1M cache write

MiniMax M2.5

MiniMax's versatile model with balanced performance across reasoning, creative writing, and multilingual tasks.

MiniMax·200K context·21 Credits / 1M in·84 Credits / 1M out·4.2 Credits / 1M cache read·26.25 Credits / 1M cache write

MiniMax M2.7

Refreshed MiniMax model with broader context and balanced performance for multilingual chat, reasoning, and creative work.

MiniMax·192K context·21 Credits / 1M in·84 Credits / 1M out·4.2 Credits / 1M cache read·26.25 Credits / 1M cache write

Mistral Nemo

Efficient Mistral model for broad multilingual chat and economical generation.

Mistral·131K context·2 Credits / 1M in·4 Credits / 1M out·2 Credits / 1M cache read·2 Credits / 1M cache write

Mistral Small 3.2 24B Instruct

Instruction-tuned Mistral Small model for compact reasoning and production assistants.

Mistral·128K context·7.5 Credits / 1M in·20 Credits / 1M out·7.5 Credits / 1M cache read·7.5 Credits / 1M cache write

Nemotron 3 Nano 30B A3B

Compact NVIDIA Nemotron model for efficient chat and instruction following.

NVIDIA·256K context·5 Credits / 1M in·20 Credits / 1M out·5 Credits / 1M cache read·5 Credits / 1M cache write

Nemotron 3 Super 120B A12B

Larger NVIDIA Nemotron model for stronger reasoning and production assistants.

NVIDIA·262K context·9 Credits / 1M in·45 Credits / 1M out·9 Credits / 1M cache read·9 Credits / 1M cache write

Qwen 3 235B A22B 2507

Large Qwen mixture model for economical reasoning and long-context generation.

Alibaba·262K context·10 Credits / 1M in·60 Credits / 1M out·10 Credits / 1M cache read·10 Credits / 1M cache write

Qwen 3 30B A3B Instruct 2507

Compact Qwen mixture model for fast instruction following and practical chat.

Alibaba·262K context·9 Credits / 1M in·30 Credits / 1M out·9 Credits / 1M cache read·9 Credits / 1M cache write

Qwen 3 32B

Mid-sized Qwen 3 model for affordable reasoning, chat, and content generation.

Alibaba·41K context·8 Credits / 1M in·24 Credits / 1M out·4 Credits / 1M cache read·8 Credits / 1M cache write

Qwen 3 Coder

Qwen 3 coder model for software engineering, code repair, and technical analysis.

Alibaba·262K context·22 Credits / 1M in·180 Credits / 1M out·22 Credits / 1M cache read·22 Credits / 1M cache write

Qwen 3 Coder Next

Qwen coding model for code generation, refactoring, and agentic developer workflows.

Alibaba·262K context·14 Credits / 1M in·80 Credits / 1M out·9 Credits / 1M cache read·14 Credits / 1M cache write

Qwen 3 Coder Plus

Specialized coding model from Alibaba with strong code generation, debugging, and analysis capabilities.

Alibaba·1M context·Tiered pricing:0–32K:40.18 Credits / 1M in·160.58 Credits / 1M out·4.018 Credits / 1M cache read·50.225 Credits / 1M cache write·32K–128K:60.27 Credits / 1M in·240.87 Credits / 1M out·6.027 Credits / 1M cache read·75.3375 Credits / 1M cache write·128K–256K:100.38 Credits / 1M in·401.45 Credits / 1M out·10.038 Credits / 1M cache read·125.475 Credits / 1M cache write·256K–1M:200.76 Credits / 1M in·2,006.97 Credits / 1M out·20.076 Credits / 1M cache read·250.95 Credits / 1M cache write

Qwen 3 Next 80B A3B Instruct

Qwen Next mixture model for efficient reasoning, chat, and instruction following.

Alibaba·262K context·9 Credits / 1M in·110 Credits / 1M out·9 Credits / 1M cache read·9 Credits / 1M cache write

Qwen 3 VL 235B A22B Instruct

Qwen VL model for multimodal-oriented prompts and strong text instruction following.

Alibaba·262K context·20 Credits / 1M in·88 Credits / 1M out·11 Credits / 1M cache read·20 Credits / 1M cache write

Qwen 3.5 27B

Mid-sized Qwen 3.5 model for general chat, analysis, and structured output.

Alibaba·262K context·30 Credits / 1M in·240 Credits / 1M out·30 Credits / 1M cache read·30 Credits / 1M cache write

Qwen 3.5 35B A3B

Efficient Qwen mixture model balancing quality, cost, and multilingual coverage.

Alibaba·262K context·25 Credits / 1M in·200 Credits / 1M out·25 Credits / 1M cache read·25 Credits / 1M cache write

Qwen 3.5 397B A17B

Large Qwen 3.5 model for high-quality multilingual reasoning and synthesis.

Alibaba·262K context·39 Credits / 1M in·234 Credits / 1M out·19.5 Credits / 1M cache read·39 Credits / 1M cache write

Qwen 3.5 9B

Small Qwen 3.5 model for low-latency chat, routing, and simple transformations.

Alibaba·262K context·10 Credits / 1M in·15 Credits / 1M out·10 Credits / 1M cache read·10 Credits / 1M cache write

Qwen 3.5 Flash

Fast and cost-effective Qwen model optimized for high-throughput tasks. Supports up to 1M context with tiered pricing.

Alibaba·1M context·Tiered pricing:0–128K:2.03 Credits / 1M in·20.09 Credits / 1M out·0.203 Credits / 1M cache read·2.5375 Credits / 1M cache write·128K–256K:8.05 Credits / 1M in·80.29 Credits / 1M out·0.805 Credits / 1M cache read·10.0625 Credits / 1M cache write·256K–1M:12.04 Credits / 1M in·120.4 Credits / 1M out·1.204 Credits / 1M cache read·15.05 Credits / 1M cache write

Qwen 3.5 Flash 02-23

Qwen 3.5 Flash snapshot for fast, low-cost multilingual assistant traffic.

Alibaba·1M context·10 Credits / 1M in·40 Credits / 1M out·1 Credits / 1M cache read·12.5 Credits / 1M cache write

Qwen 3.5 Plus

Alibaba's flagship language model with strong multilingual and reasoning capabilities. Supports up to 1M context with tiered pricing.

Alibaba·1M context·Tiered pricing:0–128K:8.05 Credits / 1M in·48.16 Credits / 1M out·0.805 Credits / 1M cache read·10.0625 Credits / 1M cache write·128K–256K:20.09 Credits / 1M in·120.4 Credits / 1M out·2.009 Credits / 1M cache read·25.1125 Credits / 1M cache write·256K–1M:40.11 Credits / 1M in·240.8 Credits / 1M out·4.011 Credits / 1M cache read·50.1375 Credits / 1M cache write

Qwen 3.5 Plus 02-15

Qwen 3.5 Plus snapshot with 1M context and tiered long-prompt pricing.

Alibaba·1M context·Tiered pricing:0–256K:40 Credits / 1M in·240 Credits / 1M out·4 Credits / 1M cache read·50 Credits / 1M cache write·256K–1M:50 Credits / 1M in·300 Credits / 1M out·5 Credits / 1M cache read·62.5 Credits / 1M cache write

Qwen 3.6 Plus

Latest generation Qwen model with improved reasoning and instruction following capabilities.

Alibaba·1M context·Tiered pricing:0–256K:19.32 Credits / 1M in·115.57 Credits / 1M out·1.932 Credits / 1M cache read·24.15 Credits / 1M cache write·256K–1M:77.07 Credits / 1M in·462.14 Credits / 1M out·7.707 Credits / 1M cache read·96.3375 Credits / 1M cache write

Step 3.5 Flash

StepFun flash model for quick multilingual chat, extraction, and structured generation.

StepFun·262K context·10 Credits / 1M in·30 Credits / 1M out·10 Credits / 1M cache read·10 Credits / 1M cache write