Prime Intellect
Compute aggregation, inference, and RL training
Last reviewed Sep 28, 2026
Prime Intellect is an AI research company and compute platform. It aggregates on-demand GPUs and reserved clusters from many providers behind one account, and runs Prime Intellect Inference, an OpenAI-compatible API that serves open-weight models directly and routes to proprietary models, plus hosted reinforcement-learning training through its Lab platform.
We're actively tracking prices for Prime Intellect. Check back soon, or browse other providers with current pricing.
Pros & Cons
Advantages
- One account for GPUs across many underlying providers
- Per-token inference rates are published through a public models endpoint
- Inference, training, and GPU rental on the same platform
- Publishes open-source models, datasets, and RL tooling
Limitations
- GPU rates are only available after sign-in through the availability API, so they are not tracked here
- As an aggregator, hardware, networking, and support depend on the underlying provider
- Gateway models are routed to other inference providers rather than served on Prime Intellect's own hardware
Key Features
Multi-provider GPU access
On-demand access to 1 to 256 GPUs sourced from many clouds, managed from one dashboard, CLI, and API
Reserved clusters
Large-cluster requests sent to 50+ data centers for parallel quotes, from H100 and H200 to B200, B300, and GB300 NVL72
Prime Intellect Inference
OpenAI-compatible API for hosted open-weight models and gateway access to OpenAI, Anthropic, Google, xAI, and other model families, billed per token
Hosted RL training
Lab runs reinforcement-learning post-training against environments from the Environments Hub, with LoRA adapters deployable for inference
Orchestration and networking
SLURM and Kubernetes orchestration, InfiniBand networking, and Grafana monitoring for clusters
Spot resale
Idle reserved nodes can be put on Prime Intellect's spot market
Pricing Options
| Option | Details |
|---|---|
| On-demand GPUs | Hourly rates set per underlying provider, with community and spot tiers where available |
| Reserved clusters | Quoted per request from participating data centers |
| Per token | Inference billed per 1M input and output tokens, with cached-input rates on hosted models |
Availability & Support
Support
Documentation, plus direct assistance from the research and infrastructure engineering team and a dedicated solutions engineer for reserved-cluster customers
Company
Published by Prime Intellect to identify the company behind this listing.
- Headquarters
- 🇺🇸 San Francisco, United States
Getting Started
- 1
Create an account
Sign up at app.primeintellect.ai
- 2
Rent GPUs
Pick a GPU type, count, and provider from the availability list, then provision from the dashboard or the prime CLI
- 3
Call the inference API
Create an API key and point an OpenAI-compatible client at api.pinference.ai
Frequently Asked Questions
What GPU types does Prime Intellect offer?
Check the pricing table above for Prime Intellect's current GPU availability and pricing.
How do I get started with Prime Intellect?
Create an account, Rent GPUs, Call the inference API
What are Prime Intellect's main advantages?
Prime Intellect's main advantages include: One account for GPUs across many underlying providers, Per-token inference rates are published through a public models endpoint, Inference, training, and GPU rental on the same platform, Publishes open-source models, datasets, and RL tooling.
What are Prime Intellect's limitations?
Prime Intellect's main limitations include: GPU rates are only available after sign-in through the availability API, so they are not tracked here, As an aggregator, hardware, networking, and support depend on the underlying provider, Gateway models are routed to other inference providers rather than served on Prime Intellect's own hardware.