GMI Cloud
Inference-first cloud with dedicated NVIDIA GPU clusters
Last reviewed May 8, 2026
GMI Cloud is a neocloud combining dedicated NVIDIA GPU compute (H100, H200, B200, GB200, with GB300 on pre-order) and a serverless, OpenAI-compatible inference API for text, image, video, and audio models.
Available GPUs
Hourly on-demand pricing. Click column headers to sort.
Prices last updated: September 29, 2026
GMI Cloud pricing by GPU
Configurations, price rank, and alternatives for one GPU at a time.
LLM API Pricing
Pay-per-token pricing. Prices shown per 1M tokens.
Prices last updated: September 29, 2026
| Model | Input/1M | Output/1M | |||
|---|---|---|---|---|---|
| $0.050 | $0.250 | ||||
| $0.087 | $0.350 | ||||
| $0.090 | $0.300 | ||||
| $0.091 | $0.182 | ||||
| $0.100 | $0.500 | ||||
| $0.119 | $0.238 | ||||
| $0.130 | $0.400 | ||||
| $0.140 | $0.280 | ||||
| $0.140 | $0.400 | ||||
| $0.140 | $0.580 | ||||
GMI Cloud pricing by model
Input, output, and batch rates, plus alternatives, for one model at a time.
Pros & Cons
Advantages
- Transparent published per-GPU hourly rates
- Access to current Blackwell-generation systems
- OpenAI-compatible inference API alongside dedicated GPU clusters
Limitations
- Smaller global footprint than hyperscalers
- Newer entrant relative to long-established providers
- Published GPU catalog covers only datacenter-class parts, with no lower-tier options such as L4 or L40S
Key Features
Serverless Inference
OpenAI-compatible endpoints for LLM and multimodal models with request batching and scaling to zero
Dedicated GPU Compute
Managed Kubernetes clusters, container instances, and bare-metal servers with RDMA-ready networking
Blackwell Capacity
GB200 available and GB300 on pre-order alongside H100, H200, and B200 systems
Studio and AgentBox
Visual workflow builder for multi-step model pipelines plus a marketplace for publishing and using AI agents
Reserved + On-Demand
Both hourly on-demand capacity and longer-term committed reservations are published
Pricing Options
| Option | Details |
|---|---|
| On-Demand Containers | Hourly billing for self-serve GPU containers |
| Reserved Private Cloud | Discounted longer-term reservations of dedicated GPU clusters |
| Serverless Inference | Per-token billing for LLM endpoints and per-request billing for image and video models |
Availability & Support
Regions
Data centers in North America and Asia, with region-aware pricing and unified billing
Support
Documentation, self-service console, Discord community, and enterprise support via sales
Getting Started
- 1
Create an account
Sign up for the GMI Cloud console
- 2
Pick a GPU and region
Select an on-demand container, bare-metal cluster, or inference endpoint
- 3
Deploy your workload
Launch via the console or programmatically through the GMI API
Frequently Asked Questions
What GPU types does GMI Cloud offer?
GMI Cloud offers various GPU types including H100 SXM, H200, GB200, B200. Check the pricing table above for current availability and pricing.
How do I get started with GMI Cloud?
Create an account, Pick a GPU and region, Deploy your workload
What are GMI Cloud's main advantages?
GMI Cloud's main advantages include: Transparent published per-GPU hourly rates, Access to current Blackwell-generation systems, OpenAI-compatible inference API alongside dedicated GPU clusters.
What are GMI Cloud's limitations?
GMI Cloud's main limitations include: Smaller global footprint than hyperscalers, Newer entrant relative to long-established providers, Published GPU catalog covers only datacenter-class parts, with no lower-tier options such as L4 or L40S.
Compare Providers
Find the best prices for the same GPUs and models from other providers