Skip to main content
IBM

Granite 4.2 3B

Granite 4.2 3B is a small chat model from IBM's Granite 4.2 family, positioned as a compact option for high-throughput and latency-sensitive text workloads.

Input from
$0.030 / 1M tokens
across 1 provider

API Pricing

ProviderInput / 1MOutput / 1MCached / 1M
$0.030$0.120$0.0075

Prices updated daily. Last check: Aug 29, 2026

Granite 4.2 3B pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
14.3 / 100
Coding
17.5 / 100
Output Speed
231 t/s
Latency (TTFT)
206ms

Reasoning & Knowledge

  • GPQA Diamond55.9%
  • Humanity's Last Exam6.6%

Coding

  • SciCode24.9%

Agentic & Tool Use

  • Terminal-Bench v2.113.9%
  • τ-bench Banking5.6%

Instruction & Long Context

  • Long-Context Reasoning24.3%

Benchmarks measured Aug 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
IBM
Modalities
Text

Capabilities

Tool Calling
No
Open Source
No

Strengths & Limitations

Strengths

  • Measured output throughput of approximately 231 tokens per second in Artificial Analysis testing
  • Time to first token around 206 ms, suitable for interactive and streaming interfaces
  • Roughly 3B parameter class, which keeps memory requirements low relative to mid-size and large chat models
  • Part of IBM's Granite family, which is developed with enterprise document and assistant workloads in mind
  • Small footprint makes it a candidate for self-hosting or private deployment alongside hosted API access
  • Family versioning (4.2) means it can be swapped against sibling Granite sizes with similar prompting conventions

Limitations

  • A ~3B parameter model has less capacity for complex multi-step reasoning than larger models in the same family
  • We do not have a confirmed context window figure for this model in our database
  • No published accuracy benchmarks (reasoning, coding, math) are tracked for this entry, so quality must be validated on your own task
  • Provider availability for smaller Granite models is narrower than for widely hosted open models
  • Modality support beyond text is not something we currently track for this model

Key Features

Chat and instruction-following interface
Approximately 3B parameter compact model class
~231 output tokens/second measured throughput
~206 ms time to first token
Member of the IBM Granite 4.2 model family
Sized for low-memory serving and high request concurrency
Available through hosted inference providers listed in the pricing table

About Granite 4.2 3B

Granite 4.2 3B is a chat-oriented language model released by IBM as part of its Granite 4.2 model family. At roughly 3 billion parameters, it sits at the small end of that family, where IBM typically pairs a compact footprint with enterprise-oriented instruction tuning. The Granite line has historically been aimed at business document processing, summarization, and assistant workloads rather than open-ended consumer chat. In measured serving performance from Artificial Analysis, Granite 4.2 3B produces roughly 231 output tokens per second with a time to first token of about 206 milliseconds. Those figures put it in the responsive tier of hosted models, which is the main practical consequence of its size: short prompts return quickly and long generations complete in less wall-clock time than larger models handling the same request. We do not currently track a confirmed context window, modality list, or benchmark accuracy scores for this specific model in our database, so buyers evaluating it for a particular task should verify those details against IBM's own model card. In practice, small models in this size class are used where volume matters more than depth of reasoning: classification, extraction, routing, drafting, and on-device or private-deployment scenarios. Compared with larger Granite variants or with mid-size models from other vendors, Granite 4.2 3B trades headroom on complex multi-step reasoning for lower latency and lower serving cost per request. Check the pricing table on this page for how providers currently price it.

Common Use Cases

Granite 4.2 3B fits high-volume, latency-sensitive text work where per-request cost and response speed dominate the requirements: intent classification, ticket routing, entity and field extraction from forms and documents, short summarization, autocomplete and suggestion features, and first-draft generation that a human or a larger model reviews. Its measured sub-quarter-second time to first token makes it usable for streaming chat surfaces where perceived responsiveness matters. It is also a reasonable candidate as the cheap tier in a routing setup, handling the majority of simple requests while escalating harder reasoning, long-context analysis, or complex coding tasks to a larger Granite variant or another model. For workloads that require sustained multi-step planning or deep domain reasoning, evaluate a larger model before committing.

Frequently Asked Questions

How much does Granite 4.2 3B cost to run?

Pricing varies by provider and by pricing model — some bill per million input and output tokens, others bill for dedicated capacity or hourly GPU time if you self-host. Because rates change frequently and differ between hosts, check the pricing table on this page for current figures rather than relying on a fixed number.

What is Granite 4.2 3B best used for?

It suits high-throughput, latency-sensitive text tasks: classification, extraction, routing, short summarization, and interactive chat surfaces. Its ~231 tokens/second output rate and ~206 ms time to first token make it well matched to workloads where many small requests need fast responses.

How does it compare to larger models in the Granite family?

As the small end of the Granite 4.2 lineup, the 3B model prioritizes speed and low resource use. Larger siblings generally handle more complex reasoning and longer, more intricate instructions better. A common pattern is to run the 3B model as the default and escalate difficult requests to a larger variant.

What context window does Granite 4.2 3B support?

We do not have a confirmed context window figure for this model in our database. Check IBM's official model card or your inference provider's documentation for the supported maximum, since some hosts serve smaller windows than the model supports.

Is Granite 4.2 3B fast enough for real-time chat?

Its measured time to first token of about 206 ms and output rate near 231 tokens per second are consistent with responsive streaming chat. Actual latency depends on the provider, prompt length, and load, so benchmark against your own traffic pattern before committing.