Skip to main content
Open SourceGoogle

Gemma 4 E4B

Gemma 4 E4B is a compact chat model from Google in the Gemma family, benchmarked at roughly 59 output tokens per second with a time to first token near 258 ms.

License Open Source
Input from
$0.020 / 1M tokens
across 1 provider

API Pricing

ProviderInput / 1MOutput / 1M
$0.020$0.100

Prices updated daily. Last check: Oct 11, 2026

Compare API pricing for every Gemma model →

Gemma 4 E4B pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
7.5 / 100
Output Speed
32.9 t/s
Latency (TTFT)
411ms

Reasoning & Knowledge

  • GPQA Diamond54.9%
  • Humanity's Last Exam4.8%

Agentic & Tool Use

  • Terminal-Bench Hard7.6%
  • τ²-bench26.0%

Instruction & Long Context

  • IFBench40.5%
  • Long-Context Reasoning24.0%

Benchmarks measured Oct 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
Google
Modalities
Text

Capabilities

Open Source
Yes

Strengths & Limitations

Strengths

  • Compact Gemma-class model sized for latency-sensitive and high-volume workloads rather than maximum model size
  • Measured time to first token of roughly 258 ms, suitable for streaming chat interfaces where perceived responsiveness matters
  • Independent throughput benchmark of about 58.6 output tokens per second from Artificial Analysis gives a concrete performance reference
  • Part of Google's Gemma line, which has an established tooling and documentation ecosystem around it
  • Small-model footprint generally means lower per-token serving cost than large frontier models, letting you run many more requests per budget
  • Fits well as a first-pass or router model in tiered pipelines that escalate hard requests to a larger model

Limitations

  • Compact models in this size class generally trail large frontier models on multi-step reasoning, long-form coding, and specialist knowledge
  • We do not have a confirmed context window for this entry, so long-document workloads need verification against provider documentation
  • Modality support is not tracked in our data — do not assume image, audio, or video input without checking the serving provider
  • Tool calling and structured output support are unconfirmed in our metadata and may vary by provider
  • Measured throughput of roughly 59 tokens per second is mid-range; some small models on optimized inference stacks serve considerably faster

Key Features

•Chat/instruction-following text model in Google's Gemma 4 generation
•Compact "E4B" configuration targeting a small effective parameter count
•Measured output throughput around 58.6 tokens per second (Artificial Analysis)
•Measured time to first token around 258 ms (Artificial Analysis)
•Streaming response support typical of chat-completion endpoints
•Available through multiple inference providers — see the pricing table on this page
•Positioned as a lightweight tier within the Gemma family rather than a large-scale variant

About Gemma 4 E4B

Gemma 4 E4B is a text chat model from Google and part of the Gemma line, Google's family of smaller models that sits apart from the larger Gemini series. The "E4B" suffix follows the naming convention Google has used for compact Gemma variants, where the label points to a small effective parameter count rather than a full-size frontier configuration. Within the Gemma 4 generation it is positioned as one of the lightweight entries rather than a large-scale model. Third-party measurements from Artificial Analysis put Gemma 4 E4B at approximately 58.6 output tokens per second with a time to first token of about 258 ms. Those figures describe a model that returns the start of a response quickly, which matters for interactive chat and streaming interfaces. Throughput and latency in practice depend heavily on which provider serves the model, the hardware behind the endpoint, and how heavily that endpoint is loaded, so treat the benchmark as a reference point rather than a guarantee. We do not currently track a confirmed context window, modality list, or tool-calling support for this specific entry, so those details should be checked against the serving provider's documentation before you build around them. In practice, small Gemma-class models are most often used where cost per token and response latency matter more than maximum reasoning depth: assistants embedded in products, summarization and rewriting pipelines, classification and extraction jobs, and drafting steps that feed into a larger model. Compare the providers listed in the pricing table on this page to see who serves Gemma 4 E4B and on what terms.

Common Use Cases

Gemma 4 E4B is aimed at workloads where per-request cost and response latency dominate the decision: in-product chat assistants, customer-facing FAQ and support bots, text summarization and rewriting at volume, tagging and classification jobs, structured field extraction from short documents, and synthetic data or draft generation that a human or larger model reviews afterward. Its sub-300 ms time to first token makes it a reasonable fit for interfaces that stream tokens back to a user, where the first visible word matters more than raw completion speed. It is also a common choice for the cheap tier of a cascading pipeline — handle the easy majority of traffic here, and route only the hard or ambiguous requests to a larger model. For deep multi-step reasoning, competitive coding tasks, or long-context document analysis, a larger model in Google's Gemini line or a comparable frontier model from another creator is generally the better match.

Frequently Asked Questions

How much does Gemma 4 E4B cost to use?

Pricing depends on which provider serves the model and on the pricing type — input tokens, output tokens, batch or cached rates, and dedicated versus shared capacity are all priced differently, and rates change often. Check the live pricing table on this page for current per-provider figures.

What is Gemma 4 E4B best used for?

It suits high-volume, latency-sensitive text work: product chat assistants, summarization, rewriting, classification, short-document extraction, and first-pass drafting in pipelines that escalate difficult requests to a larger model. Its roughly 258 ms time to first token makes it usable in streaming chat UIs.

How fast is Gemma 4 E4B?

Artificial Analysis measured approximately 58.6 output tokens per second with a time to first token of about 258 ms. Actual numbers vary by provider, hardware, request size, and endpoint load, so use these as a reference point rather than a guarantee.

What does the "E4B" in the name mean?

Google uses the E-prefixed suffix in Gemma naming for compact variants, where the number refers to a small effective parameter count — here around 4B. It signals that this is a lightweight entry in the Gemma 4 generation rather than a large configuration.

Should I use Gemma 4 E4B or a larger Google model?

Choose Gemma 4 E4B when you are optimizing for cost per request and fast first-token latency across many simple calls. Choose a larger model from Google's Gemini line when the task requires deep multi-step reasoning, substantial coding work, or analysis of long documents, since compact models in this size class generally trail larger ones on those tasks.

Does Gemma 4 E4B accept image input?

Our metadata does not confirm modality support for this entry, so we do not list it as multimodal. Check the documentation of the specific provider you plan to use before designing around image, audio, or video input.