Optimized inference for open-source models
Last reviewed Mar 14, 2026
Deep Infra is an AI inference cloud that provides serverless APIs for 100+ models and dedicated GPU rentals. Features OpenAI-compatible endpoints, custom model deployments, and infrastructure optimized for scale.
Hourly on-demand pricing. Click column headers to sort.
Prices last updated: September 17, 2026
Configurations, price rank, and alternatives for one GPU at a time.
Pay-per-token pricing. Prices shown per 1M tokens.
Prices last updated: September 17, 2026
| Model | Input/1M | Output/1M | |||
|---|---|---|---|---|---|
| $0.019 | $0.030 | ||||
| $0.020 | $0.100 | ||||
| $0.030 | $0.120 | ||||
| $0.030 | $0.140 | ||||
| $0.050 | $0.100 | ||||
| $0.050 | $0.200 | ||||
| $0.050 | $0.150 | ||||
| $0.060 | $0.180 | ||||
| $0.060 | $0.180 | ||||
| $0.060 | $0.180 | ||||
Input, output, and batch rates, plus alternatives, for one model at a time.
OpenAI-compatible API for 100+ models including DeepSeek, Qwen, Llama 4, Claude, and Gemini families with autoscaling
B200 instances with SSH access spin up in about 10 seconds and bill hourly
Deploy your own Hugging Face models onto dedicated A100, H100, H200, B200, or B300 GPUs
Published per-GPU hourly rates for A100, H100, H200, B200, and B300 with competitive pricing
Runs on its own inference-optimized hardware in US-based data centers rather than rented capacity
Support for text generation, vision and OCR, embeddings and reranking, image, video and music generation, and speech recognition and synthesis
Asynchronous batch endpoints with a files API, plus prompt caching billed at a reduced cached-input rate
Documented integrations for LangChain, LlamaIndex, the Vercel AI SDK, AutoGen, and the Anthropic SDK including Claude Code
Zero retention policy with enterprise-grade security certifications for data privacy and protection
Hosted model APIs with autoscaling on Deep Infra's own inference infrastructure.
On-demand GPU nodes with SSH access for custom workloads.
Multi-node B200 and B300 clusters with SSH access for training and full-control workloads.
Run your own fine-tuned weights on dedicated hardware behind a private endpoint.
| Option | Details |
|---|---|
| Serverless pay-per-token | OpenAI-compatible inference APIs billed per input and output token, with no idle GPU time or minimums |
| Dedicated GPU hourly rates | Published transparent hourly pricing for A100, H100, H200, B200, and B300 GPUs with pay-as-you-go billing |
| Per-execution-time billing | Non-LLM models are billed for inference execution time rather than per token |
| Cached input pricing | Prompt-cached input tokens are billed at a lower rate than uncached input tokens |
| No long-term commitments | Flexible hourly billing for dedicated instances with no prepayments, contracts, or minimums required |
Runs on self-operated, inference-optimized infrastructure in US-based data centers; no international region list is published.
Documentation site, dashboard guidance, Discord community, feedback email (feedback@deepinfra.com), and contact-sales options.
Sign up (GitHub-supported) and open the Deep Infra dashboard
Add a payment method to unlock GPU rentals and API usage
Choose serverless APIs or dedicated A100, H100, H200, B200, or B300 instances
Start instances with SSH access or call the OpenAI-compatible API endpoints
Track spend and instance status from the dashboard and shut down when idle
Deep Infra offers various GPU types including A100 SXM, H100 SXM, H200, HGX B300, B200. Check the pricing table above for current availability and pricing.
Create an account, Enable billing, Pick a GPU option, Launch and connect, Monitor usage
Deep Infra's main advantages include: Simple OpenAI-compatible API alongside controllable GPU rentals, Competitive hourly rates for flagship NVIDIA GPUs including latest B200 and B300, Fast provisioning with SSH access for dedicated instances (ready in ~10 seconds), Supports custom deployments in addition to hosted public models.
Deep Infra's main limitations include: Infrastructure is US-based only, with no published international regions, Primarily focused on inference and GPU rentals rather than broader cloud services, Newer player compared to established cloud providers.
Find the best prices for the same GPUs and models from other providers