Skip to main content
Novita AI logo

Novita AI

200+ models through one OpenAI-compatible API

Rapidly-catching neocloudinferenceopen-sourcebudget

Last reviewed Aug 28, 2026

Novita AI is an inference platform serving 200+ models through a single OpenAI-compatible API, alongside serverless GPU jobs, on-demand GPU instances, and bare-metal clusters. Its catalogue is weighted toward open-weight families — Qwen, DeepSeek, GLM, Kimi, MiniMax, Llama and Gemma — with per-token billing and prompt caching on most chat models.

We're actively tracking prices for Novita AI. Check back soon, or browse other providers with current pricing.

Pros & Cons

Advantages

  • Broad open-weight model coverage under one API, so several model families can be compared without changing providers
  • Catalogue endpoint publishes per-model pricing openly, with no account required
  • Prompt caching rates published alongside standard input rates
  • Both token-billed inference and GPU rental available from one account

Limitations

  • Model catalogue turns over quickly, and retired models stay listed with a status flag rather than being removed
  • A few models use tiered billing where the rate depends on prompt length
  • Headquarters and data-centre regions are not published on the marketing site

Key Features

OpenAI-Compatible Model API

One endpoint for 200+ text, image, audio and video models, callable with existing OpenAI client libraries

Open-Weight Catalogue

Qwen, DeepSeek, GLM, Kimi, MiniMax, Llama, Gemma and Nemotron families served serverlessly with per-token billing

Prompt Caching

Cache-read pricing published separately from standard input pricing on most chat models

Serverless and Dedicated GPUs

Per-second serverless GPU jobs, on-demand instances, and bare-metal clusters with H100 and H200 hardware

Private Endpoints

Dedicated deployments for workloads that need reserved capacity rather than shared serverless throughput

Agent Sandbox

Isolated runtimes for agent workloads that call models and execute tools, billed per second

Compute Services

Model APIs

Serverless per-token inference across text, image, audio and video models through an OpenAI-compatible API

GPU Instances

On-demand and serverless GPU compute, plus bare-metal clusters for larger deployments

Pricing Options

OptionDetails

Availability & Support

Regions

Regional availability is not published on the marketing site.

Support

Documentation site, dashboard guidance, and email support at support@novita.ai.

Getting Started

  1. 1

    Sign up at novita.ai and generate an API key from the dashboard

  2. 2

    Set the base URL to Novita's OpenAI-compatible endpoint and pass the API key as the bearer token

  3. 3

    Browse the model library for the model id, then pass it as the model parameter

  4. 4

    Move to a private endpoint or dedicated GPU instance when shared serverless throughput is not enough

Frequently Asked Questions

What GPU types does Novita AI offer?

Check the pricing table above for Novita AI's current GPU availability and pricing.

How do I get started with Novita AI?

, , ,

What are Novita AI's main advantages?

Novita AI's main advantages include: Broad open-weight model coverage under one API, so several model families can be compared without changing providers, Catalogue endpoint publishes per-model pricing openly, with no account required, Prompt caching rates published alongside standard input rates, Both token-billed inference and GPU rental available from one account.

What are Novita AI's limitations?

Novita AI's main limitations include: Model catalogue turns over quickly, and retired models stay listed with a status flag rather than being removed, A few models use tiered billing where the rate depends on prompt length, Headquarters and data-centre regions are not published on the marketing site.