Models

19 models from 8 providers. One API.

Compare context windows, capabilities and per-token rates. To use a model, pass its ID as the model parameter.

Model catalog

Showing all 19 models

Prices in USD per 1 million tokens

  1. OpenAI

    GPT-6 Astra

    gpt-6-astra

    OpenAI's largest model, for complex reasoning and long multi-step work.

    Class
    Most capable
    Context
    1M tokens
    Max output
    128K tokens
    Input
    Text, Image
    Input
    $10.00
    Cached input
    $1.00
    Output
    $50.00
    • Tool calling
    • Reasoning
    • Image input
    • Prompt caching
    • Structured output
  2. OpenAI

    GPT-6.1 Sol

    gpt-6.1-sol

    General-purpose model balancing capability, speed and cost.

    Class
    Balanced
    Context
    1M tokens
    Max output
    128K tokens
    Input
    Text, Image
    Input
    $2.00
    Cached input
    $0.20
    Output
    $10.00
    • Tool calling
    • Reasoning
    • Image input
    • Prompt caching
    • Structured output
  3. OpenAI

    GPT-6 Luna

    gpt-6-luna

    Small, fast model for classification, extraction and high-volume tasks.

    Class
    Fast, low cost
    Context
    400K tokens
    Max output
    64K tokens
    Input
    Text, Image
    Input
    $0.10
    Cached input
    $0.01
    Output
    $0.50
    • Tool calling
    • Image input
    • Prompt caching
    • Structured output
  4. Anthropic

    Claude Fable 5.1

    claude-fable-5-1

    Anthropic's most capable model, for demanding reasoning and long-horizon agents.

    Class
    Most capable
    Context
    1M tokens
    Max output
    128K tokens
    Input
    Text, Image
    Input
    $10.00
    Cached input
    $0.25
    Output
    $50.00
    • Tool calling
    • Reasoning
    • Image input
    • Prompt caching
    • Structured output
  5. Anthropic

    Claude Opus 5.5

    claude-opus-5-5

    High-capability model for coding, agents and complex analysis.

    Class
    Most capable
    Context
    1M tokens
    Max output
    128K tokens
    Input
    Text, Image
    Input
    $4.00
    Cached input
    $0.20
    Output
    $20.00
    • Tool calling
    • Reasoning
    • Image input
    • Prompt caching
    • Structured output
  6. Anthropic

    Claude Sonnet 5.5

    claude-sonnet-5-5

    Fast, capable model for everyday coding, agent and production workloads.

    Class
    Balanced
    Context
    1M tokens
    Max output
    128K tokens
    Input
    Text, Image
    Input
    $2.00
    Cached input
    $0.20
    Output
    $10.00
    • Tool calling
    • Reasoning
    • Image input
    • Prompt caching
    • Structured output
  7. Anthropic

    Claude Haiku 5.5

    claude-haiku-5-5

    Anthropic's fastest model, for latency-sensitive and high-volume work.

    Class
    Fast, low cost
    Context
    1M tokens
    Max output
    128K tokens
    Input
    Text, Image
    Input
    $0.10
    Cached input
    $0.01
    Output
    $0.50
    • Tool calling
    • Reasoning
    • Image input
    • Prompt caching
    • Structured output
  8. Google

    Gemini 3.1 Pro

    gemini-3.1-pro

    Google's most capable model, with long context and multimodal input.

    Class
    Most capable
    Context
    1M tokens
    Max output
    66K tokens
    Input
    Text, Image, Audio, Video
    Input
    $2.00
    Cached input
    $0.20
    Output
    $12.00
    • Tool calling
    • Reasoning
    • Image input
    • Prompt caching
    • Structured output
  9. Google

    Gemini 3.8 Flash

    gemini-3.8-flash

    Fast multimodal model for production workloads at moderate cost.

    Class
    Balanced
    Context
    1M tokens
    Max output
    66K tokens
    Input
    Text, Image, Audio, Video
    Input
    $0.75
    Cached input
    $0.075
    Output
    $3.75
    • Tool calling
    • Reasoning
    • Image input
    • Prompt caching
    • Structured output
  10. Google

    Gemini 3.5 Flash-Lite

    gemini-3.5-flash-lite

    Lowest-cost Gemini model, for simple tasks at very high volume.

    Class
    Fast, low cost
    Context
    1M tokens
    Max output
    66K tokens
    Input
    Text, Image
    Input
    $0.30
    Cached input
    $0.03
    Output
    $2.50
    • Tool calling
    • Image input
    • Prompt caching
    • Structured output
  11. DeepSeek

    DeepSeek V4 Pro

    deepseek-v4-pro

    Open-weight reasoning model, strong on code and mathematics.

    Class
    Most capable
    Context
    256K tokens
    Max output
    64K tokens
    Input
    Text
    Input
    $1.74
    Cached input
    $0.145
    Output
    $3.48
    • Tool calling
    • Reasoning
    • Prompt caching
    • Structured output
  12. DeepSeek

    DeepSeek V4.1 Flash

    deepseek-v4.1-flash

    Fast, low-cost open-weight model for general text tasks.

    Class
    Fast, low cost
    Context
    256K tokens
    Max output
    64K tokens
    Input
    Text
    Input
    $0.30
    Cached input
    $0.03
    Output
    $1.20
    • Tool calling
    • Prompt caching
    • Structured output
  13. Z.ai

    GLM-5.2

    glm-5.2

    Z.ai's flagship model, tuned for coding and agentic tool use.

    Class
    Most capable
    Context
    200K tokens
    Max output
    128K tokens
    Input
    Text
    Input
    $1.40
    Cached input
    $0.26
    Output
    $4.40
    • Tool calling
    • Reasoning
    • Prompt caching
    • Structured output
  14. Z.ai

    GLM-5.3 Flash

    glm-5.3-flash

    Lightweight GLM model for fast, inexpensive text generation.

    Class
    Fast, low cost
    Context
    200K tokens
    Max output
    32K tokens
    Input
    Text
    Input
    $0.15
    Cached input
    —
    Output
    $0.50
    • Tool calling
    • Structured output
  15. Moonshot AI

    Kimi K3

    kimi-k3

    Moonshot AI's largest model, built for long-context agentic work.

    Class
    Most capable
    Context
    256K tokens
    Max output
    64K tokens
    Input
    Text, Image
    Input
    $3.00
    Cached input
    $0.30
    Output
    $15.00
    • Tool calling
    • Reasoning
    • Image input
    • Prompt caching
    • Structured output
  16. Moonshot AI

    Kimi K2.6

    kimi-k2.6

    Open-weight mixture-of-experts model for coding and tool use.

    Class
    Balanced
    Context
    256K tokens
    Max output
    64K tokens
    Input
    Text
    Input
    $0.95
    Cached input
    $0.16
    Output
    $4.00
    • Tool calling
    • Prompt caching
    • Structured output
  17. MiniMax

    MiniMax M3

    minimax-m3

    Long-context open-weight model for agents and document work.

    Class
    Balanced
    Context
    1M tokens
    Max output
    128K tokens
    Input
    Text
    Input
    $0.60
    Cached input
    $0.06
    Output
    $2.40
    • Tool calling
    • Reasoning
    • Prompt caching
    • Structured output
  18. Qwen

    Qwen3.8 Max

    qwen3.8-max

    The largest Qwen model, for complex reasoning and multilingual work.

    Class
    Most capable
    Context
    262K tokens
    Max output
    66K tokens
    Input
    Text, Image
    Input
    $2.00
    Cached input
    $0.40
    Output
    $6.00
    • Tool calling
    • Reasoning
    • Image input
    • Prompt caching
    • Structured output
  19. Qwen

    Qwen3.8 Flash

    qwen3.8-flash

    Fast, inexpensive Qwen model with a long context window.

    Class
    Fast, low cost
    Context
    1M tokens
    Max output
    33K tokens
    Input
    Text
    Input
    $0.15
    Cached input
    $0.03
    Output
    $0.47
    • Tool calling
    • Prompt caching
    • Structured output

Model availability and rates can change. Applications should read the current list fromGET /v1/models. Model names are trademarks of their respective owners.

One account. One balance. One API.

Create an account, add balance and make your first request with the client you already use.