375 SAR / mo
15% off- Discounted token allowance
- Per-second pod billing
- Email support
No procurement, no quota, no flat monthly GPU. Inference is billed per million tokens; pods are billed per GPU-second, cost-plus from the live market, and the unused time is refunded the instant you stop. Everything is priced in Saudi Riyal, shown before you commit.
Each chat-completion response also returns per-call usage pricing in USD and SAR.
GPU pod prices are cost-plus from the live market and refresh every few minutes; the rate at launch is the rate you pay for that pod.
Run it in your own VPC, with a DPA, MSA, and data-flow appendix. Dedicated capacity and a CSM.
Pay-as-you-go remains the default — subscriptions are optional and do not lock you in. Unused subscription tokens do not roll over.
GPU rental is billed prepaid per GPU-second in Saudi Riyal, cost-plus from the live market. Indicative on-demand hourly rates: NVIDIA RTX 3090 from 2.5 SAR/hr, RTX 4090 from 3.62 SAR/hr, RTX 5090 from 5.2 SAR/hr, L40S (48 GB) from 5.2 SAR/hr, A100 (80 GB) from 7.3 SAR/hr, H100 (80 GB) from 17.27 SAR/hr, and H200 (141 GB) from 23.05 SAR/hr. You are billed only for the seconds a verified GPU is actually serving you, and the unused time is refunded the instant you stop.
Inference is billed per million tokens in Saudi Riyal. Rates are by model class — from about 5 halala per 1M tokens for embedding models, 15 halala for tiny, 30 for small, 150 for medium, up to around 400 halala for large. Each chat-completion response carries per-call usage pricing in both USD and SAR. New renter accounts start with 100 SAR of credit and no card is required to begin.
Both. Pay-as-you-go is the default — per million tokens for inference and per GPU-second for pods, with a prorated refund when you stop early. Optional monthly subscriptions give a discounted token allowance for teams with steady usage: Starter at 375 SAR/mo, Growth at 1,500 SAR/mo, and Scale at 5,625 SAR/mo. Unused subscription tokens do not roll over.
A chargeable call returns HTTP 402 with a machine-readable body ({ code: "insufficient_balance", required_sar, balance_sar, topup_url, retryable: true }). No pod or charge is created, so you can top up and retry safely. Running pods are stopped (not killed silently) so you can resume after topping up.
GPU pod prices are cost-plus from the live market and refresh every few minutes, so the per-second rate you see at launch is the rate you pay for that pod. Per-token inference rates are stable per model class. Prices are always shown in SAR before you commit; nothing is billed opaquely.
No. Inference, pods, and storage run in-Kingdom on Saudi-owned hardware, so there is no egress fee. Cross-border frontier models are available only by explicit per-tenant opt-in and never incur hidden transfer charges — any cross-border cost is disclosed before you enable it.