Skip to main content
Runcrate logo

Runcrate

AI cloud for inference and compute

Inference specialistinferencecomputeaiopen-sourcebare-metal

Last reviewed Aug 18, 2026

Runcrate runs a token-billed inference API over a catalog of open-source text, image, video, audio and embedding models alongside on-demand GPU instances and reserved clusters, billed from a single credit balance. It is operated by Aeonmind, LLC.

GPU Models
17
From / hour
$0.52
LLM Models
23
From / 1M input
$0.03

Available GPUs

Hourly on-demand pricing. Click column headers to sort.

Prices last updated: October 3, 2026

A10
$1.68/hr
1×
A100 SXM
$1.76/hr
1×
A100 SXM 40GB
$1.82/hr
1×
A40
$2.42/hr
1×
GH200
$2.98/hr
1×
H100 SXM
$2.59/hr
1×
H200
$4.24/hr
1×
L4
$1.23/hr
1×
L40
$1.14/hr
1×
L40S
$1.14/hr
1×

Runcrate pricing by GPU

Configurations, price rank, and alternatives for one GPU at a time.

LLM API Pricing

Pay-per-token pricing. Prices shown per 1M tokens.

Prices last updated: October 3, 2026

ModelInput/1MOutput/1M
$0.0030/min-
$0.020/min-
$0.030$0.140
$0.040$0.160
$0.040/img-
$0.050$0.200
$0.060$0.400
$0.070$0.140
$0.100$0.150
$0.100$0.950

Runcrate pricing by model

Input, output, and batch rates, plus alternatives, for one model at a time.

Pros & Cons

Advantages

  • Inference API and GPU compute share one account, one credit balance and one bill
  • Public rate card for both tokens and GPU hours, with no negotiation required to start
  • OpenAI-compatible API makes migration from other inference providers a base-URL change
  • Full root access and custom images on GPU instances
  • No egress or data transfer fees

Limitations

  • Volume discounts, uptime SLAs and reserved capacity require contacting sales rather than self-service
  • Region pinning and compliance commitments (HIPAA eligibility, SOC 2 Type II datacenter partners) are Enterprise-tier only
  • GPU capacity is sourced multi-cloud rather than from owned datacenters, so placement control is limited on self-serve plans
  • Blackwell parts appear in marketing copy and cluster examples but are not on the public self-serve GPU rate card

Key Features

Unified inference and compute billing

Token-billed API calls and GPU hours draw on one credit balance with a single invoice, auto-recharge, and credits that do not expire

OpenAI-compatible endpoint

Chat, image, video, text-to-speech, speech-to-text and embedding models served behind OpenAI-compatible REST endpoints, so existing clients switch by changing the base URL

Root access on GPU instances

Bare-metal instances with SSH, Docker and custom images, provisioned in about 60 seconds

Public GPU pricing API

An unauthenticated endpoint publishes the live GPU rate card, refreshed every few minutes

Dedicated clusters

Reserved bare-metal from 16 to 128+ nodes with NVLink fabric and managed Slurm job submission

Developer tooling

Python and TypeScript SDKs, a Vercel AI SDK provider, a CLI for instance and volume management, and an MCP server for controlling the platform from AI assistants

Persistent storage

Volumes that mount to instances, with a built-in file explorer

No data transfer fees

Egress is not charged, so embedding, vector storage and generation can run on one platform without cross-provider transfer costs

Compute Services

GPU Instances

Self-service bare-metal GPU instances with root access, custom images and persistent volumes, deployable in about 60 seconds and scalable from 1 to 128 nodes

Dedicated Clusters

Reserved bare-metal clusters from 16 to 128+ nodes with NVLink fabric and managed Slurm, aimed at multi-node pre-training and fine-tuning

Pricing Options

OptionDetails
Pay-per-token inferencePublished per-million-token input and output rates on every model in the catalog, with no minimum and no commitment
Usage-based GPU computeGPU instances metered while running with no minimum commitment — stopping the instance stops the meter
Dedicated capacityReserved GPU capacity sized to traffic in exchange for a monthly minimum, priced 40–60% below the public rate card and carrying a 99.9% uptime SLA
Enterprise contractsCustom contract pricing with bring-your-own-cloud and self-hosted deployments, region pinning, and 99.95%/99.99% SLAs with service credits

Availability & Support

Regions

Multi-cloud capacity advertised across North America (Los Angeles, Chicago), Europe (Amsterdam, Frankfurt) and Asia-Pacific (Singapore, Mumbai, Sydney, Tokyo); region pinning to US, EU or APAC is an Enterprise-tier option

Support

Documentation, quickstarts and SDK references; email and community Discord on the self-serve plan; Slack Connect with engineers on Dedicated; named customer success manager and on-call engineering on Enterprise

Getting Started

  1. 1

    Create an account

    Sign up through the console and generate an API key

  2. 2

    Add credits

    Fund one balance used by both API calls and GPU hours, optionally with auto-recharge

  3. 3

    Call the API or deploy an instance

    Point an OpenAI-compatible client at the inference endpoint, or launch a GPU instance from the console or CLI

Frequently Asked Questions

What GPU types does Runcrate offer?

Runcrate offers various GPU types including RTX A5000, RTX A6000, A100 SXM, H100 SXM, GH200, H200, A100 SXM 40GB, L40, A10, RTX 4090, RTX 4090, A40, RTX 6000 Ada, L40S, RTX 5090, Tesla V100, L4, RTX PRO 6000. Check the pricing table above for current availability and pricing.

How do I get started with Runcrate?

Create an account, Add credits, Call the API or deploy an instance

What are Runcrate's main advantages?

Runcrate's main advantages include: Inference API and GPU compute share one account, one credit balance and one bill, Public rate card for both tokens and GPU hours, with no negotiation required to start, OpenAI-compatible API makes migration from other inference providers a base-URL change, Full root access and custom images on GPU instances, No egress or data transfer fees.

What are Runcrate's limitations?

Runcrate's main limitations include: Volume discounts, uptime SLAs and reserved capacity require contacting sales rather than self-service, Region pinning and compliance commitments (HIPAA eligibility, SOC 2 Type II datacenter partners) are Enterprise-tier only, GPU capacity is sourced multi-cloud rather than from owned datacenters, so placement control is limited on self-serve plans, Blackwell parts appear in marketing copy and cluster examples but are not on the public self-serve GPU rate card.

Compare Providers

Find the best prices for the same GPUs and models from other providers

Omega Gradient logo

Omega Gradient

16 shared GPUs with Runcrate

Compare Prices
Sesterce logo

Sesterce

16 shared GPUs with Runcrate

Compare Prices
Shadeform logo

Shadeform

16 shared GPUs with Runcrate

Compare Prices

All Runcrate Comparisons

Runcrate vs Omega GradientRuncrate vs SesterceRuncrate vs ShadeformRuncrate vs RunpodRuncrate vs AquanodeRuncrate vs gpu.aiRuncrate vs Vast.aiRuncrate vs AtmosComputeRuncrate vs IO.NETRuncrate vs Hexgrid CloudRuncrate vs Lium.ioRuncrate vs Massed ComputeRuncrate vs SpheronRuncrate vs VerdaRuncrate vs VoltageGPURuncrate vs 1LegionRuncrate vs AceCloudRuncrate vs Amazon AWSRuncrate vs CoreWeaveRuncrate vs Google CloudRuncrate vs Hugging FaceRuncrate vs Lambda LabsRuncrate vs ModalRuncrate vs BasetenRuncrate vs HyperstackRuncrate vs OblivusRuncrate vs Oracle CloudRuncrate vs SeewebRuncrate vs Together AIRuncrate vs DaytonaRuncrate vs Denvr DataworksRuncrate vs Jarvis LabsRuncrate vs KoyebRuncrate vs OVHcloudRuncrate vs PaperspaceRuncrate vs CivoRuncrate vs CrusoeRuncrate vs DigitalOceanRuncrate vs LyceumRuncrate vs Microsoft AzureRuncrate vs NebiusRuncrate vs Novita AIRuncrate vs Thunder ComputeRuncrate vs UpCloudRuncrate vs BeamRuncrate vs GcoreRuncrate vs HinodeRuncrate vs Impossible CloudRuncrate vs Packet AIRuncrate vs Salad CloudRuncrate vs ScalewayRuncrate vs SwissGPURuncrate vs VultrRuncrate vs Atlas CloudRuncrate vs Compute CheapRuncrate vs Cudo ComputeRuncrate vs GMI CloudRuncrate vs GPU OutletRuncrate vs HyperbolicRuncrate vs Latitude.shRuncrate vs Theta EdgeCloudRuncrate vs Akamai CloudRuncrate vs EcoHashRuncrate vs fal.aiRuncrate vs Voltage ParkRuncrate vs ZettabyteRuncrate vs BentausRuncrate vs Beyond.plRuncrate vs Deep InfraRuncrate vs Hot AisleRuncrate vs QuickPodRuncrate vs TheAI CloudRuncrate vs OpenRouterRuncrate vs Prime IntellectRuncrate vs HeabsyRuncrate vs VelokeyRuncrate vs GeoddRuncrate vs GroqRuncrate vs ReplicateRuncrate vs AnthropicRuncrate vs CerebrasRuncrate vs CohereRuncrate vs DeepSeekRuncrate vs Fireworks AIRuncrate vs Mistral AIRuncrate vs OpenAIRuncrate vs Pareto InferenceRuncrate vs PerplexityRuncrate vs SambaNovaRuncrate vs xAI