Granite 4.2 3B
Granite 4.2 3B is a small chat model from IBM's Granite 4.2 family, positioned as a compact option for high-throughput and latency-sensitive text workloads.
API Pricing
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.030 | $0.120 | $0.0075 |
Prices updated daily. Last check: Aug 29, 2026
Granite 4.2 3B pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- GPQA Diamond55.9%
- Humanity's Last Exam6.6%
Coding
- SciCode24.9%
Agentic & Tool Use
- Terminal-Bench v2.113.9%
- τ-bench Banking5.6%
Instruction & Long Context
- Long-Context Reasoning24.3%
Benchmarks measured Aug 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- IBM
- Modalities
- Text
Capabilities
- Tool Calling
- No
- Open Source
- No
Strengths & Limitations
Strengths
- Measured output throughput of approximately 231 tokens per second in Artificial Analysis testing
- Time to first token around 206 ms, suitable for interactive and streaming interfaces
- Roughly 3B parameter class, which keeps memory requirements low relative to mid-size and large chat models
- Part of IBM's Granite family, which is developed with enterprise document and assistant workloads in mind
- Small footprint makes it a candidate for self-hosting or private deployment alongside hosted API access
- Family versioning (4.2) means it can be swapped against sibling Granite sizes with similar prompting conventions
Limitations
- A ~3B parameter model has less capacity for complex multi-step reasoning than larger models in the same family
- We do not have a confirmed context window figure for this model in our database
- No published accuracy benchmarks (reasoning, coding, math) are tracked for this entry, so quality must be validated on your own task
- Provider availability for smaller Granite models is narrower than for widely hosted open models
- Modality support beyond text is not something we currently track for this model
Key Features
About Granite 4.2 3B
Common Use Cases
Granite 4.2 3B fits high-volume, latency-sensitive text work where per-request cost and response speed dominate the requirements: intent classification, ticket routing, entity and field extraction from forms and documents, short summarization, autocomplete and suggestion features, and first-draft generation that a human or a larger model reviews. Its measured sub-quarter-second time to first token makes it usable for streaming chat surfaces where perceived responsiveness matters. It is also a reasonable candidate as the cheap tier in a routing setup, handling the majority of simple requests while escalating harder reasoning, long-context analysis, or complex coding tasks to a larger Granite variant or another model. For workloads that require sustained multi-step planning or deep domain reasoning, evaluate a larger model before committing.
Frequently Asked Questions
How much does Granite 4.2 3B cost to run?
Pricing varies by provider and by pricing model — some bill per million input and output tokens, others bill for dedicated capacity or hourly GPU time if you self-host. Because rates change frequently and differ between hosts, check the pricing table on this page for current figures rather than relying on a fixed number.
What is Granite 4.2 3B best used for?
It suits high-throughput, latency-sensitive text tasks: classification, extraction, routing, short summarization, and interactive chat surfaces. Its ~231 tokens/second output rate and ~206 ms time to first token make it well matched to workloads where many small requests need fast responses.
How does it compare to larger models in the Granite family?
As the small end of the Granite 4.2 lineup, the 3B model prioritizes speed and low resource use. Larger siblings generally handle more complex reasoning and longer, more intricate instructions better. A common pattern is to run the 3B model as the default and escalate difficult requests to a larger variant.
What context window does Granite 4.2 3B support?
We do not have a confirmed context window figure for this model in our database. Check IBM's official model card or your inference provider's documentation for the supported maximum, since some hosts serve smaller windows than the model supports.
Is Granite 4.2 3B fast enough for real-time chat?
Its measured time to first token of about 206 ms and output rate near 231 tokens per second are consistent with responsive streaming chat. Actual latency depends on the provider, prompt length, and load, so benchmark against your own traffic pattern before committing.