Gemma 4 E4B
Gemma 4 E4B is a compact chat model from Google in the Gemma family, benchmarked at roughly 59 output tokens per second with a time to first token near 258 ms.
API Pricing
| Provider | Input / 1M | Output / 1M |
|---|---|---|
| $0.020 | $0.100 |
Prices updated daily. Last check: Oct 11, 2026
Compare API pricing for every Gemma model →Gemma 4 E4B pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- GPQA Diamond54.9%
- Humanity's Last Exam4.8%
Agentic & Tool Use
- Terminal-Bench Hard7.6%
- τ²-bench26.0%
Instruction & Long Context
- IFBench40.5%
- Long-Context Reasoning24.0%
Benchmarks measured Oct 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- Modalities
- Text
Capabilities
- Open Source
- Yes
Strengths & Limitations
Strengths
- Compact Gemma-class model sized for latency-sensitive and high-volume workloads rather than maximum model size
- Measured time to first token of roughly 258 ms, suitable for streaming chat interfaces where perceived responsiveness matters
- Independent throughput benchmark of about 58.6 output tokens per second from Artificial Analysis gives a concrete performance reference
- Part of Google's Gemma line, which has an established tooling and documentation ecosystem around it
- Small-model footprint generally means lower per-token serving cost than large frontier models, letting you run many more requests per budget
- Fits well as a first-pass or router model in tiered pipelines that escalate hard requests to a larger model
Limitations
- Compact models in this size class generally trail large frontier models on multi-step reasoning, long-form coding, and specialist knowledge
- We do not have a confirmed context window for this entry, so long-document workloads need verification against provider documentation
- Modality support is not tracked in our data — do not assume image, audio, or video input without checking the serving provider
- Tool calling and structured output support are unconfirmed in our metadata and may vary by provider
- Measured throughput of roughly 59 tokens per second is mid-range; some small models on optimized inference stacks serve considerably faster
Key Features
About Gemma 4 E4B
Common Use Cases
Gemma 4 E4B is aimed at workloads where per-request cost and response latency dominate the decision: in-product chat assistants, customer-facing FAQ and support bots, text summarization and rewriting at volume, tagging and classification jobs, structured field extraction from short documents, and synthetic data or draft generation that a human or larger model reviews afterward. Its sub-300 ms time to first token makes it a reasonable fit for interfaces that stream tokens back to a user, where the first visible word matters more than raw completion speed. It is also a common choice for the cheap tier of a cascading pipeline — handle the easy majority of traffic here, and route only the hard or ambiguous requests to a larger model. For deep multi-step reasoning, competitive coding tasks, or long-context document analysis, a larger model in Google's Gemini line or a comparable frontier model from another creator is generally the better match.
Frequently Asked Questions
How much does Gemma 4 E4B cost to use?
Pricing depends on which provider serves the model and on the pricing type — input tokens, output tokens, batch or cached rates, and dedicated versus shared capacity are all priced differently, and rates change often. Check the live pricing table on this page for current per-provider figures.
What is Gemma 4 E4B best used for?
It suits high-volume, latency-sensitive text work: product chat assistants, summarization, rewriting, classification, short-document extraction, and first-pass drafting in pipelines that escalate difficult requests to a larger model. Its roughly 258 ms time to first token makes it usable in streaming chat UIs.
How fast is Gemma 4 E4B?
Artificial Analysis measured approximately 58.6 output tokens per second with a time to first token of about 258 ms. Actual numbers vary by provider, hardware, request size, and endpoint load, so use these as a reference point rather than a guarantee.
What does the "E4B" in the name mean?
Google uses the E-prefixed suffix in Gemma naming for compact variants, where the number refers to a small effective parameter count — here around 4B. It signals that this is a lightweight entry in the Gemma 4 generation rather than a large configuration.
Should I use Gemma 4 E4B or a larger Google model?
Choose Gemma 4 E4B when you are optimizing for cost per request and fast first-token latency across many simple calls. Choose a larger model from Google's Gemini line when the task requires deep multi-step reasoning, substantial coding work, or analysis of long documents, since compact models in this size class generally trail larger ones on those tasks.
Does Gemma 4 E4B accept image input?
Our metadata does not confirm modality support for this entry, so we do not list it as multimodal. Check the documentation of the specific provider you plan to use before designing around image, audio, or video input.