Every major model on one key
Claude for reasoning, DeepSeek for code, Qwen for value, Kimi and GLM for multilingual workflows. Claude & GPT billed at 66% of official rates; other major models at competitive discounts. Live list per key via GET /v1/models.
Near-frontier intelligence at roughly half the price, 1M context.
Anthropic's frontier model, 1M context, for the hardest reasoning work.
The most agentic Sonnet, 1M context, balanced cost and capability.
Fastest and cheapest Claude, SWE-bench 73.3%.
Flagship reasoning model, Responses and Chat both supported.
Dedicated line for code generation, works with Codex CLI.
First 3T-class open model: 2.8T params, native vision, 1M context.
All gains from post-training, coding +50% over 5.2, strongest open.
Fast, low-cost variant of GLM-5.3 for high volume.
MIT-licensed flagship with true 1M context.
V4-Pro (1.6T/49B active) rivals top closed models, native 1M context.
V4-Flash (284B/13B) efficient and affordable, native 1M context.
2.4T MoE open flagship, can autonomously code 10+ day projects.
6B active params per token with 1M context.
Vision understanding with structured image and video analysis.
How usage is billed
Usage-based, metered per call. No subscription, credits never expire.
Every call is itemised in the console — input, output, cache and thinking tokens per request. Verify live rates with GET /v1/models.