Gemma 4 31B API pricing on OpenRelay
Every OpenRelay Gemma 4 31B rate we track, compared against 12 providers serving the same model.
The cheapest Gemma 4 31B input price we currently track is $0.090/1M on Deep Infra.
OpenRelay Gemma 4 31B rates
Prices per 1M tokens. Batch and cached-input rates are shown where the provider publishes them.
| Mode | Input | Output |
|---|---|---|
| Standard | $0.990/1M | $1.49/1M |
About Gemma 4 31B
- Creator
- Context
- 262K
- Modalities
- text, image, video
- Open weights
- Yes
Gemma 4 31B on other providers
About OpenRelay
OpenRelay is a GPU cloud and inference API. Rent NVIDIA and AMD GPUs by the hour, from the RTX 3090 and RTX 4090 to the A100, H100 SXM, B200, B300 and AMD Instinct MI355X, as VMs with the GPU passed through or as bring-your-own-container pods. There is no commitment, quota request or reservation, and no egress fees; CPU, RAM and disk are included in the hourly rate. The same account calls open-weight LLMs (GPT-OSS, DeepSeek, GLM, Qwen, Gemma) through an OpenAI-compatible inference API billed per token. Based in San Francisco, backed by Y Combinator (S26).
Frequently Asked Questions
How much does Gemma 4 31B cost on OpenRelay?
OpenRelay serves Gemma 4 31B from $0.990 per 1M input tokens, across 1 tracked pricing mode. Prices are collected daily — see the table above for current input, output, and cached-input rates.
Is OpenRelay the cheapest way to run Gemma 4 31B?
OpenRelay ranks #12 of 12 providers we track serving Gemma 4 31B. The cheapest input price right now is on Deep Infra. Price is only one factor — throughput, latency, and context limits differ between providers.
Does OpenRelay offer batch or cached pricing for Gemma 4 31B?
We track 1 Gemma 4 31B offering from OpenRelay. Batch and cached-input rates appear in the table above when the provider publishes them; blank cells mean we have no current figure for that mode.