Home / Models

Model catalogue

Claude & GPT at 66% of official rates · DeepSeek & Kimi K3 at 88% · GLM from 70% · Qwen at official rates. Discounts are a policy, not a promo — verify live models with GET /v1/models.

Model catalogue

Every major model on one key

Claude for reasoning, DeepSeek for code, Qwen for value, Kimi and GLM for multilingual workflows. Claude & GPT billed at 66% of official rates; other major models at competitive discounts. Live list per key via GET /v1/models.

Opus 5 · 1M ctx
Anthropic · Reasoning

Near-frontier intelligence at roughly half the price, 1M context.

Live
Fable 5 · 1M ctx
Anthropic · Frontier

Anthropic's frontier model, 1M context, for the hardest reasoning work.

Live
Sonnet 5 · 1M ctx
Anthropic · Most agentic

The most agentic Sonnet, 1M context, balanced cost and capability.

Live
Haiku 4.5 · 200K
Anthropic · Fastest

Fastest and cheapest Claude, SWE-bench 73.3%.

Live
GPT-5.5 · Flagship
OpenAI · Flagship

Flagship reasoning model, Responses and Chat both supported.

Live
Codex · Coding line
OpenAI · Coding

Dedicated line for code generation, works with Codex CLI.

Live
Kimi K3
Moonshot · Long context

First 3T-class open model: 2.8T params, native vision, 1M context.

Live
GLM-5.3
Zhipu · Coding flagship

All gains from post-training, coding +50% over 5.2, strongest open.

Live
GLM-5.3 Flash
Zhipu · Fast & cheap

Fast, low-cost variant of GLM-5.3 for high volume.

Live
GLM-5.2
Zhipu · MIT flagship

MIT-licensed flagship with true 1M context.

Live
V4 Pro
DeepSeek · Flagship

V4-Pro (1.6T/49B active) rivals top closed models, native 1M context.

Live
V4 Flash
DeepSeek · Efficient

V4-Flash (284B/13B) efficient and affordable, native 1M context.

Live
3.8 Max
Alibaba · Flagship

2.4T MoE open flagship, can autonomously code 10+ day projects.

Live
3.8 Flash
Alibaba · Long context

6B active params per token with 1M context.

Live
3-VL Plus · Vision
Alibaba · Vision

Vision understanding with structured image and video analysis.

Live

How usage is billed

Usage-based, metered per call. No subscription, credits never expire.

Input tokens
Prompt and context sent to the model. Billed at each model's published rate — Claude & GPT at 66% of official rates, other major models at official rates.
Output tokens 5× input
Generated text. Most models bill output at 5× the input rate (DeepSeek 2×). Streaming replies count the same as non-streaming.
Cache hit (read) 0.1× input
Reused context from the provider's prompt cache — charged at 10% of the input rate. Keep your system prompt stable to hit the cache; this is the single biggest cost saver.
Cache write 1.25× input
Storing context for reuse (Claude models). Charged once at 125% of the input rate, then reads are 10%.
Thinking tokens = output
Reasoning tokens from extended-thinking models are billed exactly like output tokens — no hidden multiplier.
Discounts standing
Discounts are policy, not promotions: Claude & GPT at 66% of official rates; DeepSeek · Kimi · GLM · Qwen at official rates. They do not expire and are not tied to a promotion window.

Every call is itemised in the console — input, output, cache and thinking tokens per request. Verify live rates with GET /v1/models.