Model Catalog

AIOrouter is an All-In-One AI Router — providing a single OpenAI-compatible API endpoint that routes to Western models (Google Gemini, Anthropic Claude), Chinese models (DeepSeek, Alibaba Qwen, Moonshot Kimi, Zhipu GLM), and more — all with built-in privacy protection (bidirectional PII pseudonymization, technical secret redaction, AI Firewall) and Canada-resident infrastructure.

LLM Public Token Pricing

Prices below are public provider rates in USD per 1 million tokens before CAD conversion, taxes, and any account-specific retail presentation. The Dashboard and API model response are the customer-facing source for your exact billable rate.

Pricing basis. Model pricing is based on the rates published on this page. AIOrouter reserves the right to adjust published rates from time to time; adjustments take effect once announced (see the Terms of Service).

Model Input (USD/1M tokens) Output (USD/1M tokens) Cache pricing
claude-fable-5 $10.00 $50.00 Cache write: $12.50; Cache read: $1.00
claude-haiku-4.5 $1.00 $5.00 Cache write: $1.25; Cache read: $0.100
claude-opus-5 $5.00 $25.00 Cache write: $6.25; Cache read: $0.500
claude-sonnet-5 $2.00 $10.00 Cache write: $2.50; Cache read: $0.200
deepseek-v4-flash (version 0731) $0.138 $0.275 Implicit cache: $0.028
deepseek-v4-pro → deepseek-v4-pro-0813 $0.636 $1.91 Cache read: $0.064; Implicit cache: $0.064; explicit creation $0.795
gemini-2.5-flash $0.300 $2.50 Implicit cache: $0.030
gemini-2.5-pro $1.25 up to 200K / $2.50 above $10.00 Implicit cache: $0.125
glm-5.1 $0.825 $3.30 Cache read: $0.083; Implicit cache: $0.165; explicit creation $1.03
glm-5.2 $1.40 $4.40 Cache read: $0.140; Implicit cache: $0.260; explicit creation $1.75
grok-4.6 $2.00 $6.00 Implicit cache: $0.500
kimi-k2.6 $0.950 $4.00 Implicit cache: $0.160
kimi-k2.7-code $0.950 $4.00 Implicit cache: $0.190
kimi-k3 $3.00 $15.00 Implicit cache: $0.300
qwen3.6-flash $0.165 $0.990 Cache read: $0.017; Implicit cache: $0.033; explicit creation $0.206
qwen3.6-plus → qwen3.7-max-2026-05-20 $0.276 $1.65 Cache read: $0.028; Implicit cache: $0.055; explicit creation $0.345
qwen3.7-max $0.825 $2.48 Cache read: $0.083; Implicit cache: $0.165; explicit creation $1.03
qwen3.7-plus $0.221 $0.881 Cache read: $0.022; Implicit cache: $0.045; explicit creation $0.275
qwen3.8-max $1.65 $4.95 Cache read: $0.137; Implicit cache: $0.206; explicit creation $2.06

Time-of-day pricing (peak/off-peak): requests during the off-peak hours (22:00–08:00 (Asia/Shanghai, UTC+8)) are billed at the prices in the table above; during peak hours the gateway automatically applies the 2× peak rate below (X-Billing-Model response header shows the resolved variant).

Model Off-peak (USD/1M in/out) Peak (USD/1M in/out)
deepseek-v4-flash (version 0731) $0.220 / $0.660 $0.440 / $1.32
deepseek-v4-pro → deepseek-v4-pro-0813 $0.636 / $1.91 $1.27 / $3.82

Current Exchange Rate

Date Source USD → CAD Rate FX Buffer Applied
2026-08-18 Bank of Canada — Bank of Canada VALET API 1.4142 Yes

Token costs in CAD are computed as: USD_price × CAD_USD_rate. The rate above is refreshed daily at 09:00 EST from the Bank of Canada VALET API with the FX buffer (5% execution fee covering currency conversion and payment processing) applied per pricing policy.

Current Exchange Rate

Date Source USD → CAD Rate 2% Buffer Applied
2026-08-20 Bank of Canada — Bank of Canada VALET API 1.4515 Yes

Token costs in CAD are computed as: USD_price × CAD_USD_rate. The rate above is refreshed daily at 09:00 EST from the Bank of Canada VALET API with a 5% buffer applied per pricing policy.

Official Pricing Sources

Data freshness: This BETA catalog was updated on 2026-08-20. Token rates are public provider rates.