Skip to main content
OpenRelay logo

Gemma 4 31B API pricing on OpenRelay

Every OpenRelay Gemma 4 31B rate we track, compared against 12 providers serving the same model.

Input from /1M
$0.990
Output /1M
$1.49
Price rank
#12 of 12
Last updated
October 6, 2026

The cheapest Gemma 4 31B input price we currently track is $0.090/1M on Deep Infra.

OpenRelay Gemma 4 31B rates

Prices per 1M tokens. Batch and cached-input rates are shown where the provider publishes them.

OpenRelay Gemma 4 31B pricing by mode
ModeInputOutput
Standard$0.990/1M$1.49/1M

About Gemma 4 31B

Creator
Google
Context
262K
Modalities
text, image, video
Open weights
Yes
Full Gemma 4 31B details and every provider →

Gemma 4 31B on other providers

About OpenRelay

OpenRelay is a GPU cloud and inference API. Rent NVIDIA and AMD GPUs by the hour, from the RTX 3090 and RTX 4090 to the A100, H100 SXM, B200, B300 and AMD Instinct MI355X, as VMs with the GPU passed through or as bring-your-own-container pods. There is no commitment, quota request or reservation, and no egress fees; CPU, RAM and disk are included in the hourly rate. The same account calls open-weight LLMs (GPT-OSS, DeepSeek, GLM, Qwen, Gemma) through an OpenAI-compatible inference API billed per token. Based in San Francisco, backed by Y Combinator (S26).

Frequently Asked Questions

How much does Gemma 4 31B cost on OpenRelay?

OpenRelay serves Gemma 4 31B from $0.990 per 1M input tokens, across 1 tracked pricing mode. Prices are collected daily — see the table above for current input, output, and cached-input rates.

Is OpenRelay the cheapest way to run Gemma 4 31B?

OpenRelay ranks #12 of 12 providers we track serving Gemma 4 31B. The cheapest input price right now is on Deep Infra. Price is only one factor — throughput, latency, and context limits differ between providers.

Does OpenRelay offer batch or cached pricing for Gemma 4 31B?

We track 1 Gemma 4 31B offering from OpenRelay. Batch and cached-input rates appear in the table above when the provider publishes them; blank cells mean we have no current figure for that mode.