Skip to main content
Google Cloud logo

Google Cloud

Enterprise cloud with advanced AI/ML services

Classical hyperscaler🇺🇸 USinferenceenterprisemultimodal

Last reviewed Mar 14, 2026

GCP provides powerful GPU instances with flexible pricing and integration with Google's AI and machine learning tools. It's a major cloud provider known for its innovation in Kubernetes, AI/ML, and data analytics.

GPU Models
8
From / hour
$0.16
LLM Models
4
From / 1M input
$0.38

Available GPUs

Hourly on-demand pricing. Click column headers to sort.

Prices last updated: August 18, 2026

Pricing
GPU Model
Price / hr
A100 SXM
$3.67/hr
1×2×4×8×16×
H100 SXM
$11.06/hr
8×
H200
$10.60/hr
8×
L4
$0.560/hr
1×2×4×8×
RTX PRO 6000
$0.647/hr
1×2×4×8×
Tesla T4
$0.350/hr
1×
Tesla V100
$2.48/hr
1×

Google Cloud pricing by GPU

Configurations, price rank, and alternatives for one GPU at a time.

LLM API Pricing

Pay-per-token pricing. Prices shown per 1M tokens.

Prices last updated: August 17, 2026

Pricing
ModelInput/1MOutput/1M
$1.50$7.50
$1.50$7.50

Google Cloud pricing by model

Input, output, and batch rates, plus alternatives, for one model at a time.

Pros & Cons

Advantages

  • Flexible pricing options, including sustained use discounts
  • Strong AI and machine learning tools (Vertex AI)
  • Good integration with other Google services
  • Mature managed Kubernetes offering (GKE)
  • Strong global network infrastructure
  • Broad AI/ML and data analytics service portfolio

Limitations

  • Limited availability in some regions compared to AWS
  • Complexity in managing resources
  • Support can be costly
  • Steeper learning curve for some services

Key Features

Compute Engine

Scalable virtual machines with a wide range of machine types, including GPUs.

Google Kubernetes Engine (GKE)

Managed Kubernetes service for deploying and managing containerized applications.

Cloud Functions

Event-driven serverless compute platform.

Cloud Run

Fully managed serverless platform for containerized applications.

Vertex AI

Unified ML platform for building, deploying, and managing ML models.

Spot VMs

Spare compute capacity at discounted Spot prices, suitable for fault-tolerant workloads that can be interrupted.

Flexible machine sizing

Balance vCPUs, memory, and disk with up to 8 GPUs per instance, billed per second after a one-minute minimum.

Cloud Storage

Scalable and durable object storage.

Persistent Disk

Block storage for Compute Engine instances.

Cloud Load Balancing

High-performance, scalable load balancing.

Virtual Private Cloud (VPC)

Software-defined networking for your cloud resources.

Compute Services

Compute Engine

Offers customizable virtual machines running in Google's data centers.

Google Kubernetes Engine (GKE)

Managed Kubernetes service for running containerized applications.

  • Automated Kubernetes operations
  • Integration with Google Cloud services
  • Advanced cluster management features

Cloud Functions

Serverless compute platform for running code in response to events.

  • Automatic scaling and high availability
  • Pay only for the compute time consumed
  • Supports multiple programming languages

Cloud Run

Fully managed serverless platform for deploying and scaling containerized applications.

  • Runs stateless containers on a fully managed environment
  • Automatic scaling and high availability
  • Pay only for the resources used

Inference Services

Vertex AI

Access to Google's Gemini models and other foundation models through a fully managed platform with enterprise security, MLOps tools, and Google Cloud integration.

  • Gemini Models: Access Google's latest Gemini Pro and Flash models with multimodal capabilities
  • Model Garden: Curated collection of open-source and Google-developed models
  • Grounding: Connect models to Google Search or your own data for accurate responses
  • Context Caching: Cache large context windows for cost savings on repeated queries

Pricing Models

  • Pay-per-token: Standard per-token pricing for input and output
  • Context Caching: Reduced rates for cached context in long conversations
  • Provisioned Throughput: Reserved capacity for predictable performance

Pricing Options

OptionDetails
On-DemandPay for vCPUs, GPUs, and memory with a one-minute minimum and per-second billing thereafter, with no long-term commitments.
Sustained Use DiscountsAutomatic discounts for running instances for a significant portion of the month.
Resource-based Committed Use DiscountsDiscounted rates in exchange for a 1-year or 3-year commitment to a minimum level of resource usage in a region, and combinable with reservations.
Compute Flexible Committed Use DiscountsSpend-based commitments that apply across eligible machine families and regions rather than to specific resources.
Spot VMsSpare capacity at discounted Spot prices for fault-tolerant workloads that can be preempted.

Availability & Support

Regions

40+ regions and 120+ zones worldwide.

Support

Role-based (free), Standard, Enhanced and Premium support plans. Comprehensive documentation, community forums, and training resources.

Getting Started

  1. 1

    Create a Google Cloud project

    Set up a project in the Google Cloud Console.

  2. 2

    Enable billing

    Set up a billing account to pay for resource usage.

  3. 3

    Choose a compute service

    Select Compute Engine, GKE, Cloud Functions, or Cloud Run based on your needs.

  4. 4

    Create and configure an instance

    Launch a VM instance, configure a Kubernetes cluster, or deploy a function/application.

  5. 5

    Manage resources

    Use the Cloud Console, command-line tools, or APIs to manage your resources.

Compare Providers

Find the best prices for the same GPUs and models from other providers

GPU Outlet logo

GPU Outlet

8 shared GPUs with Google Cloud

Compare Prices
IO.NET logo

IO.NET

7 shared GPUs with Google Cloud

Compare Prices
Modal logo

Modal

7 shared GPUs with Google Cloud

Compare Prices

All Google Cloud Comparisons

Google Cloud vs GPU OutletGoogle Cloud vs IO.NETGoogle Cloud vs ModalGoogle Cloud vs RunpodGoogle Cloud vs AtmosComputeGoogle Cloud vs SesterceGoogle Cloud vs VerdaGoogle Cloud vs AceCloudGoogle Cloud vs Amazon AWSGoogle Cloud vs CoreWeaveGoogle Cloud vs HyperstackGoogle Cloud vs Jarvis LabsGoogle Cloud vs Massed ComputeGoogle Cloud vs Oracle CloudGoogle Cloud vs SeewebGoogle Cloud vs CivoGoogle Cloud vs fal.aiGoogle Cloud vs KoyebGoogle Cloud vs Lambda LabsGoogle Cloud vs LyceumGoogle Cloud vs Microsoft AzureGoogle Cloud vs Together AIGoogle Cloud vs BeamGoogle Cloud vs CrusoeGoogle Cloud vs Denvr DataworksGoogle Cloud vs Genesis CloudGoogle Cloud vs GMI CloudGoogle Cloud vs HyperbolicGoogle Cloud vs OblivusGoogle Cloud vs PaperspaceGoogle Cloud vs Theta EdgeCloudGoogle Cloud vs UpCloudGoogle Cloud vs Atlas CloudGoogle Cloud vs DigitalOceanGoogle Cloud vs HinodeGoogle Cloud vs Latitude.shGoogle Cloud vs NebiusGoogle Cloud vs ScalewayGoogle Cloud vs Thunder ComputeGoogle Cloud vs 1LegionGoogle Cloud vs Akamai CloudGoogle Cloud vs Beyond.plGoogle Cloud vs Cudo ComputeGoogle Cloud vs Deep InfraGoogle Cloud vs Packet AIGoogle Cloud vs Vast.aiGoogle Cloud vs Voltage ParkGoogle Cloud vs VultrGoogle Cloud vs ZettabyteGoogle Cloud vs Hot AisleGoogle Cloud vs Salad CloudGoogle Cloud vs SwissGPUGoogle Cloud vs OpenRouterGoogle Cloud vs AnthropicGoogle Cloud vs CohereGoogle Cloud vs Fireworks AIGoogle Cloud vs GroqGoogle Cloud vs Mistral AIGoogle Cloud vs OpenAIGoogle Cloud vs PerplexityGoogle Cloud vs ReplicateGoogle Cloud vs Wafer