Google Cloud
Enterprise cloud with advanced AI/ML services
Last reviewed Mar 14, 2026
GCP provides powerful GPU instances with flexible pricing and integration with Google's AI and machine learning tools. It's a major cloud provider known for its innovation in Kubernetes, AI/ML, and data analytics.
Available GPUs
Hourly on-demand pricing. Click column headers to sort.
Prices last updated: August 18, 2026
GPU Model↑ | Price / hr↑ | |
|---|---|---|
| A100 SXM | $3.67/hr 1×2×4×8×16× | 1×2×4×8×16× |
| H100 SXM | $11.06/hr 8× | 8× |
| H200 | $10.60/hr 8× | 8× |
| L4 | $0.560/hr 1×2×4×8× | 1×2×4×8× |
| RTX PRO 6000 | $0.647/hr 1×2×4×8× | 1×2×4×8× |
| Tesla T4 | $0.350/hr 1× | 1× |
| Tesla V100 | $2.48/hr 1× | 1× |
Google Cloud pricing by GPU
Configurations, price rank, and alternatives for one GPU at a time.
LLM API Pricing
Pay-per-token pricing. Prices shown per 1M tokens.
Prices last updated: August 17, 2026
| Model | Input/1M | Output/1M |
|---|---|---|
| $1.50 | $7.50 | |
| $1.50 | $7.50 |
Google Cloud pricing by model
Input, output, and batch rates, plus alternatives, for one model at a time.
Pros & Cons
Advantages
- Flexible pricing options, including sustained use discounts
- Strong AI and machine learning tools (Vertex AI)
- Good integration with other Google services
- Mature managed Kubernetes offering (GKE)
- Strong global network infrastructure
- Broad AI/ML and data analytics service portfolio
Limitations
- Limited availability in some regions compared to AWS
- Complexity in managing resources
- Support can be costly
- Steeper learning curve for some services
Key Features
Compute Engine
Scalable virtual machines with a wide range of machine types, including GPUs.
Google Kubernetes Engine (GKE)
Managed Kubernetes service for deploying and managing containerized applications.
Cloud Functions
Event-driven serverless compute platform.
Cloud Run
Fully managed serverless platform for containerized applications.
Vertex AI
Unified ML platform for building, deploying, and managing ML models.
Spot VMs
Spare compute capacity at discounted Spot prices, suitable for fault-tolerant workloads that can be interrupted.
Flexible machine sizing
Balance vCPUs, memory, and disk with up to 8 GPUs per instance, billed per second after a one-minute minimum.
Cloud Storage
Scalable and durable object storage.
Persistent Disk
Block storage for Compute Engine instances.
Cloud Load Balancing
High-performance, scalable load balancing.
Virtual Private Cloud (VPC)
Software-defined networking for your cloud resources.
Compute Services
Compute Engine
Offers customizable virtual machines running in Google's data centers.
Google Kubernetes Engine (GKE)
Managed Kubernetes service for running containerized applications.
- Automated Kubernetes operations
- Integration with Google Cloud services
- Advanced cluster management features
Cloud Functions
Serverless compute platform for running code in response to events.
- Automatic scaling and high availability
- Pay only for the compute time consumed
- Supports multiple programming languages
Cloud Run
Fully managed serverless platform for deploying and scaling containerized applications.
- Runs stateless containers on a fully managed environment
- Automatic scaling and high availability
- Pay only for the resources used
Inference Services
Vertex AI
Access to Google's Gemini models and other foundation models through a fully managed platform with enterprise security, MLOps tools, and Google Cloud integration.
- Gemini Models: Access Google's latest Gemini Pro and Flash models with multimodal capabilities
- Model Garden: Curated collection of open-source and Google-developed models
- Grounding: Connect models to Google Search or your own data for accurate responses
- Context Caching: Cache large context windows for cost savings on repeated queries
Pricing Models
- Pay-per-token: Standard per-token pricing for input and output
- Context Caching: Reduced rates for cached context in long conversations
- Provisioned Throughput: Reserved capacity for predictable performance
Pricing Options
| Option | Details |
|---|---|
| On-Demand | Pay for vCPUs, GPUs, and memory with a one-minute minimum and per-second billing thereafter, with no long-term commitments. |
| Sustained Use Discounts | Automatic discounts for running instances for a significant portion of the month. |
| Resource-based Committed Use Discounts | Discounted rates in exchange for a 1-year or 3-year commitment to a minimum level of resource usage in a region, and combinable with reservations. |
| Compute Flexible Committed Use Discounts | Spend-based commitments that apply across eligible machine families and regions rather than to specific resources. |
| Spot VMs | Spare capacity at discounted Spot prices for fault-tolerant workloads that can be preempted. |
Availability & Support
Regions
40+ regions and 120+ zones worldwide.
Support
Role-based (free), Standard, Enhanced and Premium support plans. Comprehensive documentation, community forums, and training resources.
Getting Started
- 1
Create a Google Cloud project
Set up a project in the Google Cloud Console.
- 2
Enable billing
Set up a billing account to pay for resource usage.
- 3
Choose a compute service
Select Compute Engine, GKE, Cloud Functions, or Cloud Run based on your needs.
- 4
Create and configure an instance
Launch a VM instance, configure a Kubernetes cluster, or deploy a function/application.
- 5
Manage resources
Use the Cloud Console, command-line tools, or APIs to manage your resources.
Compare Providers
Find the best prices for the same GPUs and models from other providers