Pricing
Pay per token, from one prepaid balance.
Add balance once and use it with any model. Every request is charged at that model’s published rate, and your usage history shows exactly what it cost.
How billing works
Add balance
Top up your account in the dashboard. Your balance is shared by every model and every API key.
Make requests
Each request is priced from its input, cached input and output tokens at the model’s published rate.
See every charge
The cost is deducted from your balance and recorded in your usage history, by model.
Rate card
Rates for all 19 models.
Input, cached input and output tokens are priced separately, because they cost different amounts to serve. All prices are in US dollars per million tokens.
| Model | Input | Cached input | Output |
|---|---|---|---|
| OpenAI | |||
| GPT-6 Astragpt-6-astra | $10.00 | $1.00 | $50.00 |
| GPT-6.1 Solgpt-6.1-sol | $2.00 | $0.20 | $10.00 |
| GPT-6 Lunagpt-6-luna | $0.10 | $0.01 | $0.50 |
| Anthropic | |||
| Claude Fable 5.1claude-fable-5-1 | $10.00 | $0.25 | $50.00 |
| Claude Opus 5.5claude-opus-5-5 | $4.00 | $0.20 | $20.00 |
| Claude Sonnet 5.5claude-sonnet-5-5 | $2.00 | $0.20 | $10.00 |
| Claude Haiku 5.5claude-haiku-5-5 | $0.10 | $0.01 | $0.50 |
| Gemini 3.1 Progemini-3.1-pro | $2.00 | $0.20 | $12.00 |
| Gemini 3.8 Flashgemini-3.8-flash | $0.75 | $0.075 | $3.75 |
| Gemini 3.5 Flash-Litegemini-3.5-flash-lite | $0.30 | $0.03 | $2.50 |
| DeepSeek | |||
| DeepSeek V4 Prodeepseek-v4-pro | $1.74 | $0.145 | $3.48 |
| DeepSeek V4.1 Flashdeepseek-v4.1-flash | $0.30 | $0.03 | $1.20 |
| Z.ai | |||
| GLM-5.2glm-5.2 | $1.40 | $0.26 | $4.40 |
| GLM-5.3 Flashglm-5.3-flash | $0.15 | — | $0.50 |
| Moonshot AI | |||
| Kimi K3kimi-k3 | $3.00 | $0.30 | $15.00 |
| Kimi K2.6kimi-k2.6 | $0.95 | $0.16 | $4.00 |
| MiniMax | |||
| MiniMax M3minimax-m3 | $0.60 | $0.06 | $2.40 |
| Qwen | |||
| Qwen3.8 Maxqwen3.8-max | $2.00 | $0.40 | $6.00 |
| Qwen3.8 Flashqwen3.8-flash | $0.15 | $0.03 | $0.47 |
USD per 1 million tokens. A dash means cached input is not offered for that model.
Estimate
What would a month cost?
Describe your usage and compare the same workload across every model. Estimates use published rates; your actual bill depends on real token counts.
Estimated monthly costclaude-sonnet-5-5
- Input20M × $2.00
- $40.00
- Cached input20M × $0.20
- $4.00
- Output8M × $10.00
- $80.00
Total$124.00
The same workload on every model
- GPT-6 Luna OpenAI$6.20
- Claude Haiku 5.5 Anthropic$6.20
- Qwen3.8 Flash Qwen$7.36
- GLM-5.3 Flash Z.ai$10.00
- DeepSeek V4.1 Flash DeepSeek$16.20
- Gemini 3.5 Flash-Lite Google$26.60
- MiniMax M3 MiniMax$32.40
- Gemini 3.8 Flash Google$46.50
- Kimi K2.6 Moonshot AI$54.20
- DeepSeek V4 Pro DeepSeek$65.54
- GLM-5.2 Z.ai$68.40
- Qwen3.8 Max Qwen$96.00
- GPT-6.1 Sol OpenAI$124.00
- Claude Sonnet 5.5 Anthropic$124.00
- Gemini 3.1 Pro Google$140.00
- Kimi K3 Moonshot AI$186.00
- Claude Opus 5.5 Anthropic$244.00
- Claude Fable 5.1 Anthropic$605.00
- GPT-6 Astra OpenAI$620.00
Billing questions
The details.
When am I charged?
When you add balance. After that, each request deducts its cost from your balance as it completes.
What is cached input?
Many applications send the same long prefix on every request: instructions, a policy document, a codebase summary. On models that support prompt caching, tokens served from cache are billed at the lower cached input rate. Models without caching show a dash in the rate card and bill all input at the standard rate.
Are reasoning tokens billed?
Yes. Models that reason before answering produce tokens while they think, and those are billed as output tokens. The usage object in each response reports them.
Can rates change?
Rates can change, for example when a model provider changes its own pricing. The rate applied to a request is the one published at the time the request is made.
What happens when my balance runs out?
Requests are declined with a clear error instead of running up a bill. Add balance in the dashboard and you can continue straight away.
One account. One balance. One API.
Create an account, add balance and make your first request with the client you already use.