Models
19 models from 8 providers. One API.
Compare context windows, capabilities and per-token rates. To use a model, pass its ID as the model parameter.
Model catalog
Showing all 19 models
Prices in USD per 1 million tokens
OpenAI
GPT-6 Astra
gpt-6-astraOpenAI's largest model, for complex reasoning and long multi-step work.
- Class
- Most capable
- Context
- 1M tokens
- Max output
- 128K tokens
- Input
- Text, Image
- Input
- $10.00
- Cached input
- $1.00
- Output
- $50.00
- Tool calling
- Reasoning
- Image input
- Prompt caching
- Structured output
OpenAI
GPT-6.1 Sol
gpt-6.1-solGeneral-purpose model balancing capability, speed and cost.
- Class
- Balanced
- Context
- 1M tokens
- Max output
- 128K tokens
- Input
- Text, Image
- Input
- $2.00
- Cached input
- $0.20
- Output
- $10.00
- Tool calling
- Reasoning
- Image input
- Prompt caching
- Structured output
OpenAI
GPT-6 Luna
gpt-6-lunaSmall, fast model for classification, extraction and high-volume tasks.
- Class
- Fast, low cost
- Context
- 400K tokens
- Max output
- 64K tokens
- Input
- Text, Image
- Input
- $0.10
- Cached input
- $0.01
- Output
- $0.50
- Tool calling
- Image input
- Prompt caching
- Structured output
Anthropic
Claude Fable 5.1
claude-fable-5-1Anthropic's most capable model, for demanding reasoning and long-horizon agents.
- Class
- Most capable
- Context
- 1M tokens
- Max output
- 128K tokens
- Input
- Text, Image
- Input
- $10.00
- Cached input
- $0.25
- Output
- $50.00
- Tool calling
- Reasoning
- Image input
- Prompt caching
- Structured output
Anthropic
Claude Opus 5.5
claude-opus-5-5High-capability model for coding, agents and complex analysis.
- Class
- Most capable
- Context
- 1M tokens
- Max output
- 128K tokens
- Input
- Text, Image
- Input
- $4.00
- Cached input
- $0.20
- Output
- $20.00
- Tool calling
- Reasoning
- Image input
- Prompt caching
- Structured output
Anthropic
Claude Sonnet 5.5
claude-sonnet-5-5Fast, capable model for everyday coding, agent and production workloads.
- Class
- Balanced
- Context
- 1M tokens
- Max output
- 128K tokens
- Input
- Text, Image
- Input
- $2.00
- Cached input
- $0.20
- Output
- $10.00
- Tool calling
- Reasoning
- Image input
- Prompt caching
- Structured output
Anthropic
Claude Haiku 5.5
claude-haiku-5-5Anthropic's fastest model, for latency-sensitive and high-volume work.
- Class
- Fast, low cost
- Context
- 1M tokens
- Max output
- 128K tokens
- Input
- Text, Image
- Input
- $0.10
- Cached input
- $0.01
- Output
- $0.50
- Tool calling
- Reasoning
- Image input
- Prompt caching
- Structured output
Google
Gemini 3.1 Pro
gemini-3.1-proGoogle's most capable model, with long context and multimodal input.
- Class
- Most capable
- Context
- 1M tokens
- Max output
- 66K tokens
- Input
- Text, Image, Audio, Video
- Input
- $2.00
- Cached input
- $0.20
- Output
- $12.00
- Tool calling
- Reasoning
- Image input
- Prompt caching
- Structured output
Google
Gemini 3.8 Flash
gemini-3.8-flashFast multimodal model for production workloads at moderate cost.
- Class
- Balanced
- Context
- 1M tokens
- Max output
- 66K tokens
- Input
- Text, Image, Audio, Video
- Input
- $0.75
- Cached input
- $0.075
- Output
- $3.75
- Tool calling
- Reasoning
- Image input
- Prompt caching
- Structured output
Google
Gemini 3.5 Flash-Lite
gemini-3.5-flash-liteLowest-cost Gemini model, for simple tasks at very high volume.
- Class
- Fast, low cost
- Context
- 1M tokens
- Max output
- 66K tokens
- Input
- Text, Image
- Input
- $0.30
- Cached input
- $0.03
- Output
- $2.50
- Tool calling
- Image input
- Prompt caching
- Structured output
DeepSeek
DeepSeek V4 Pro
deepseek-v4-proOpen-weight reasoning model, strong on code and mathematics.
- Class
- Most capable
- Context
- 256K tokens
- Max output
- 64K tokens
- Input
- Text
- Input
- $1.74
- Cached input
- $0.145
- Output
- $3.48
- Tool calling
- Reasoning
- Prompt caching
- Structured output
DeepSeek
DeepSeek V4.1 Flash
deepseek-v4.1-flashFast, low-cost open-weight model for general text tasks.
- Class
- Fast, low cost
- Context
- 256K tokens
- Max output
- 64K tokens
- Input
- Text
- Input
- $0.30
- Cached input
- $0.03
- Output
- $1.20
- Tool calling
- Prompt caching
- Structured output
Z.ai
GLM-5.2
glm-5.2Z.ai's flagship model, tuned for coding and agentic tool use.
- Class
- Most capable
- Context
- 200K tokens
- Max output
- 128K tokens
- Input
- Text
- Input
- $1.40
- Cached input
- $0.26
- Output
- $4.40
- Tool calling
- Reasoning
- Prompt caching
- Structured output
Z.ai
GLM-5.3 Flash
glm-5.3-flashLightweight GLM model for fast, inexpensive text generation.
- Class
- Fast, low cost
- Context
- 200K tokens
- Max output
- 32K tokens
- Input
- Text
- Input
- $0.15
- Cached input
- —
- Output
- $0.50
- Tool calling
- Structured output
Moonshot AI
Kimi K3
kimi-k3Moonshot AI's largest model, built for long-context agentic work.
- Class
- Most capable
- Context
- 256K tokens
- Max output
- 64K tokens
- Input
- Text, Image
- Input
- $3.00
- Cached input
- $0.30
- Output
- $15.00
- Tool calling
- Reasoning
- Image input
- Prompt caching
- Structured output
Moonshot AI
Kimi K2.6
kimi-k2.6Open-weight mixture-of-experts model for coding and tool use.
- Class
- Balanced
- Context
- 256K tokens
- Max output
- 64K tokens
- Input
- Text
- Input
- $0.95
- Cached input
- $0.16
- Output
- $4.00
- Tool calling
- Prompt caching
- Structured output
MiniMax
MiniMax M3
minimax-m3Long-context open-weight model for agents and document work.
- Class
- Balanced
- Context
- 1M tokens
- Max output
- 128K tokens
- Input
- Text
- Input
- $0.60
- Cached input
- $0.06
- Output
- $2.40
- Tool calling
- Reasoning
- Prompt caching
- Structured output
Qwen
Qwen3.8 Max
qwen3.8-maxThe largest Qwen model, for complex reasoning and multilingual work.
- Class
- Most capable
- Context
- 262K tokens
- Max output
- 66K tokens
- Input
- Text, Image
- Input
- $2.00
- Cached input
- $0.40
- Output
- $6.00
- Tool calling
- Reasoning
- Image input
- Prompt caching
- Structured output
Qwen
Qwen3.8 Flash
qwen3.8-flashFast, inexpensive Qwen model with a long context window.
- Class
- Fast, low cost
- Context
- 1M tokens
- Max output
- 33K tokens
- Input
- Text
- Input
- $0.15
- Cached input
- $0.03
- Output
- $0.47
- Tool calling
- Prompt caching
- Structured output
No models match these filters.
Model availability and rates can change. Applications should read the current list fromGET /v1/models. Model names are trademarks of their respective owners.
One account. One balance. One API.
Create an account, add balance and make your first request with the client you already use.