200+ models through one OpenAI-compatible API
Last reviewed Aug 28, 2026
Novita AI is an inference platform serving 200+ models through a single OpenAI-compatible API, alongside serverless GPU jobs, on-demand GPU instances, and bare-metal clusters. Its catalogue is weighted toward open-weight families — Qwen, DeepSeek, GLM, Kimi, MiniMax, Llama and Gemma — with per-token billing and prompt caching on most chat models.
Hourly on-demand pricing. Click column headers to sort.
Prices last updated: September 5, 2026
Configurations, price rank, and alternatives for one GPU at a time.
Pay-per-token pricing. Prices shown per 1M tokens.
Prices last updated: September 5, 2026
| Model | Input/1M | Output/1M | |||
|---|---|---|---|---|---|
| $0.020 | $0.020 | ||||
| $0.020 | $0.050 | ||||
| $0.040 | $0.170 | ||||
| $0.040 | $0.150 | ||||
| $0.050 | $0.200 | ||||
| $0.050 | $0.250 | ||||
| $0.050 | $0.100 | ||||
| $0.060 | $0.090 | ||||
| $0.060 | $0.180 | ||||
| $0.070 | $0.270 | ||||
Input, output, and batch rates, plus alternatives, for one model at a time.
One endpoint for 200+ text, image, audio and video models, callable with existing OpenAI client libraries
Qwen, DeepSeek, GLM, Kimi, MiniMax, Llama, Gemma and Nemotron families served serverlessly with per-token billing
Cache-read pricing published separately from standard input pricing on most chat models
Per-second serverless GPU jobs, on-demand instances, and bare-metal clusters with H100 and H200 hardware
Dedicated deployments for workloads that need reserved capacity rather than shared serverless throughput
Isolated runtimes for agent workloads that call models and execute tools, billed per second
Serverless per-token inference across text, image, audio and video models through an OpenAI-compatible API
On-demand and serverless GPU compute, plus bare-metal clusters for larger deployments
| Option | Details |
|---|---|
Regional availability is not published on the marketing site.
Documentation site, dashboard guidance, and email support at support@novita.ai.
Sign up at novita.ai and generate an API key from the dashboard
Set the base URL to Novita's OpenAI-compatible endpoint and pass the API key as the bearer token
Browse the model library for the model id, then pass it as the model parameter
Move to a private endpoint or dedicated GPU instance when shared serverless throughput is not enough
Novita AI offers various GPU types including H100 SXM, H100 SXM, RTX 4090, RTX 4090, L40S, RTX 5090, RTX 5090. Check the pricing table above for current availability and pricing.
, , ,
Novita AI's main advantages include: Broad open-weight model coverage under one API, so several model families can be compared without changing providers, Catalogue endpoint publishes per-model pricing openly, with no account required, Prompt caching rates published alongside standard input rates, Both token-billed inference and GPU rental available from one account.
Novita AI's main limitations include: Model catalogue turns over quickly, and retired models stay listed with a status flag rather than being removed, A few models use tiered billing where the rate depends on prompt length, Headquarters and data-centre regions are not published on the marketing site.
Find the best prices for the same GPUs and models from other providers