Runcrate
AI cloud for inference and compute
Last reviewed Aug 18, 2026
Runcrate runs a token-billed inference API over a catalog of open-source text, image, video, audio and embedding models alongside on-demand GPU instances and reserved clusters, billed from a single credit balance. It is operated by Aeonmind, LLC.
Available GPUs
Hourly on-demand pricing. Click column headers to sort.
Prices last updated: October 3, 2026
Runcrate pricing by GPU
Configurations, price rank, and alternatives for one GPU at a time.
LLM API Pricing
Pay-per-token pricing. Prices shown per 1M tokens.
Prices last updated: October 3, 2026
| Model | Input/1M | Output/1M | |||
|---|---|---|---|---|---|
| $0.0030/min | - | ||||
| $0.020/min | - | ||||
| $0.030 | $0.140 | ||||
| $0.040 | $0.160 | ||||
| $0.040/img | - | ||||
| $0.050 | $0.200 | ||||
| $0.060 | $0.400 | ||||
| $0.070 | $0.140 | ||||
| $0.100 | $0.150 | ||||
| $0.100 | $0.950 | ||||
Runcrate pricing by model
Input, output, and batch rates, plus alternatives, for one model at a time.
Pros & Cons
Advantages
- Inference API and GPU compute share one account, one credit balance and one bill
- Public rate card for both tokens and GPU hours, with no negotiation required to start
- OpenAI-compatible API makes migration from other inference providers a base-URL change
- Full root access and custom images on GPU instances
- No egress or data transfer fees
Limitations
- Volume discounts, uptime SLAs and reserved capacity require contacting sales rather than self-service
- Region pinning and compliance commitments (HIPAA eligibility, SOC 2 Type II datacenter partners) are Enterprise-tier only
- GPU capacity is sourced multi-cloud rather than from owned datacenters, so placement control is limited on self-serve plans
- Blackwell parts appear in marketing copy and cluster examples but are not on the public self-serve GPU rate card
Key Features
Unified inference and compute billing
Token-billed API calls and GPU hours draw on one credit balance with a single invoice, auto-recharge, and credits that do not expire
OpenAI-compatible endpoint
Chat, image, video, text-to-speech, speech-to-text and embedding models served behind OpenAI-compatible REST endpoints, so existing clients switch by changing the base URL
Root access on GPU instances
Bare-metal instances with SSH, Docker and custom images, provisioned in about 60 seconds
Public GPU pricing API
An unauthenticated endpoint publishes the live GPU rate card, refreshed every few minutes
Dedicated clusters
Reserved bare-metal from 16 to 128+ nodes with NVLink fabric and managed Slurm job submission
Developer tooling
Python and TypeScript SDKs, a Vercel AI SDK provider, a CLI for instance and volume management, and an MCP server for controlling the platform from AI assistants
Persistent storage
Volumes that mount to instances, with a built-in file explorer
No data transfer fees
Egress is not charged, so embedding, vector storage and generation can run on one platform without cross-provider transfer costs
Compute Services
GPU Instances
Self-service bare-metal GPU instances with root access, custom images and persistent volumes, deployable in about 60 seconds and scalable from 1 to 128 nodes
Dedicated Clusters
Reserved bare-metal clusters from 16 to 128+ nodes with NVLink fabric and managed Slurm, aimed at multi-node pre-training and fine-tuning
Pricing Options
| Option | Details |
|---|---|
| Pay-per-token inference | Published per-million-token input and output rates on every model in the catalog, with no minimum and no commitment |
| Usage-based GPU compute | GPU instances metered while running with no minimum commitment — stopping the instance stops the meter |
| Dedicated capacity | Reserved GPU capacity sized to traffic in exchange for a monthly minimum, priced 40–60% below the public rate card and carrying a 99.9% uptime SLA |
| Enterprise contracts | Custom contract pricing with bring-your-own-cloud and self-hosted deployments, region pinning, and 99.95%/99.99% SLAs with service credits |
Availability & Support
Regions
Multi-cloud capacity advertised across North America (Los Angeles, Chicago), Europe (Amsterdam, Frankfurt) and Asia-Pacific (Singapore, Mumbai, Sydney, Tokyo); region pinning to US, EU or APAC is an Enterprise-tier option
Support
Documentation, quickstarts and SDK references; email and community Discord on the self-serve plan; Slack Connect with engineers on Dedicated; named customer success manager and on-call engineering on Enterprise
Getting Started
- 1
Create an account
Sign up through the console and generate an API key
- 2
Add credits
Fund one balance used by both API calls and GPU hours, optionally with auto-recharge
- 3
Call the API or deploy an instance
Point an OpenAI-compatible client at the inference endpoint, or launch a GPU instance from the console or CLI
Frequently Asked Questions
What GPU types does Runcrate offer?
Runcrate offers various GPU types including RTX A5000, RTX A6000, A100 SXM, H100 SXM, GH200, H200, A100 SXM 40GB, L40, A10, RTX 4090, RTX 4090, A40, RTX 6000 Ada, L40S, RTX 5090, Tesla V100, L4, RTX PRO 6000. Check the pricing table above for current availability and pricing.
How do I get started with Runcrate?
Create an account, Add credits, Call the API or deploy an instance
What are Runcrate's main advantages?
Runcrate's main advantages include: Inference API and GPU compute share one account, one credit balance and one bill, Public rate card for both tokens and GPU hours, with no negotiation required to start, OpenAI-compatible API makes migration from other inference providers a base-URL change, Full root access and custom images on GPU instances, No egress or data transfer fees.
What are Runcrate's limitations?
Runcrate's main limitations include: Volume discounts, uptime SLAs and reserved capacity require contacting sales rather than self-service, Region pinning and compliance commitments (HIPAA eligibility, SOC 2 Type II datacenter partners) are Enterprise-tier only, GPU capacity is sourced multi-cloud rather than from owned datacenters, so placement control is limited on self-serve plans, Blackwell parts appear in marketing copy and cluster examples but are not on the public self-serve GPU rate card.
Compare Providers
Find the best prices for the same GPUs and models from other providers