Novita AI
200+ models through one OpenAI-compatible API
Last reviewed Aug 28, 2026
Novita AI is an inference platform serving 200+ models through a single OpenAI-compatible API, alongside serverless GPU jobs, on-demand GPU instances, and bare-metal clusters. Its catalogue is weighted toward open-weight families — Qwen, DeepSeek, GLM, Kimi, MiniMax, Llama and Gemma — with per-token billing and prompt caching on most chat models.
We're actively tracking prices for Novita AI. Check back soon, or browse other providers with current pricing.
Pros & Cons
Advantages
- Broad open-weight model coverage under one API, so several model families can be compared without changing providers
- Catalogue endpoint publishes per-model pricing openly, with no account required
- Prompt caching rates published alongside standard input rates
- Both token-billed inference and GPU rental available from one account
Limitations
- Model catalogue turns over quickly, and retired models stay listed with a status flag rather than being removed
- A few models use tiered billing where the rate depends on prompt length
- Headquarters and data-centre regions are not published on the marketing site
Key Features
OpenAI-Compatible Model API
One endpoint for 200+ text, image, audio and video models, callable with existing OpenAI client libraries
Open-Weight Catalogue
Qwen, DeepSeek, GLM, Kimi, MiniMax, Llama, Gemma and Nemotron families served serverlessly with per-token billing
Prompt Caching
Cache-read pricing published separately from standard input pricing on most chat models
Serverless and Dedicated GPUs
Per-second serverless GPU jobs, on-demand instances, and bare-metal clusters with H100 and H200 hardware
Private Endpoints
Dedicated deployments for workloads that need reserved capacity rather than shared serverless throughput
Agent Sandbox
Isolated runtimes for agent workloads that call models and execute tools, billed per second
Compute Services
Model APIs
Serverless per-token inference across text, image, audio and video models through an OpenAI-compatible API
GPU Instances
On-demand and serverless GPU compute, plus bare-metal clusters for larger deployments
Pricing Options
| Option | Details |
|---|---|
Availability & Support
Regions
Regional availability is not published on the marketing site.
Support
Documentation site, dashboard guidance, and email support at support@novita.ai.
Getting Started
- 1
Sign up at novita.ai and generate an API key from the dashboard
- 2
Set the base URL to Novita's OpenAI-compatible endpoint and pass the API key as the bearer token
- 3
Browse the model library for the model id, then pass it as the model parameter
- 4
Move to a private endpoint or dedicated GPU instance when shared serverless throughput is not enough
Frequently Asked Questions
What GPU types does Novita AI offer?
Check the pricing table above for Novita AI's current GPU availability and pricing.
How do I get started with Novita AI?
, , ,
What are Novita AI's main advantages?
Novita AI's main advantages include: Broad open-weight model coverage under one API, so several model families can be compared without changing providers, Catalogue endpoint publishes per-model pricing openly, with no account required, Prompt caching rates published alongside standard input rates, Both token-billed inference and GPU rental available from one account.
What are Novita AI's limitations?
Novita AI's main limitations include: Model catalogue turns over quickly, and retired models stay listed with a status flag rather than being removed, A few models use tiered billing where the rate depends on prompt length, Headquarters and data-centre regions are not published on the marketing site.