Granite 4.2 3B is a small chat model from IBM's Granite 4.2 family, positioned as a compact option for high-throughput and latency-sensitive text workloads.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.030 | $0.120 | $0.0075 |
Prices updated daily. Last check: Sep 5, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Granite 4.2 3B fits high-volume, latency-sensitive text work where per-request cost and response speed dominate the requirements: intent classification, ticket routing, entity and field extraction from forms and documents, short summarization, autocomplete and suggestion features, and first-draft generation that a human or a larger model reviews. Its measured sub-quarter-second time to first token makes it usable for streaming chat surfaces where perceived responsiveness matters. It is also a reasonable candidate as the cheap tier in a routing setup, handling the majority of simple requests while escalating harder reasoning, long-context analysis, or complex coding tasks to a larger Granite variant or another model. For workloads that require sustained multi-step planning or deep domain reasoning, evaluate a larger model before committing.
Pricing varies by provider and by pricing model — some bill per million input and output tokens, others bill for dedicated capacity or hourly GPU time if you self-host. Because rates change frequently and differ between hosts, check the pricing table on this page for current figures rather than relying on a fixed number.
It suits high-throughput, latency-sensitive text tasks: classification, extraction, routing, short summarization, and interactive chat surfaces. Its ~231 tokens/second output rate and ~206 ms time to first token make it well matched to workloads where many small requests need fast responses.
As the small end of the Granite 4.2 lineup, the 3B model prioritizes speed and low resource use. Larger siblings generally handle more complex reasoning and longer, more intricate instructions better. A common pattern is to run the 3B model as the default and escalate difficult requests to a larger variant.
We do not have a confirmed context window figure for this model in our database. Check IBM's official model card or your inference provider's documentation for the supported maximum, since some hosts serve smaller windows than the model supports.
Its measured time to first token of about 206 ms and output rate near 231 tokens per second are consistent with responsive streaming chat. Actual latency depends on the provider, prompt length, and load, so benchmark against your own traffic pattern before committing.