Skip to main content
IBM

Granite 4.2 30B

Granite 4.2 30B is a chat-oriented large language model from IBM, part of the Granite family, positioned at roughly 30B parameter scale within the 4.2 generation.

Input from
$0.160 / 1M tokens
across 1 provider

API Pricing

ProviderInput / 1MOutput / 1MCached / 1M
$0.160$0.650$0.040

Prices updated daily. Last check: Aug 29, 2026

Granite 4.2 30B pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
23.7 / 100
Coding
29.9 / 100
Output Speed
77.5 t/s
Latency (TTFT)
249ms

Reasoning & Knowledge

  • GPQA Diamond64.4%
  • Humanity's Last Exam11.2%

Coding

  • SciCode36.6%

Agentic & Tool Use

  • Terminal-Bench v2.126.6%
  • τ-bench Banking14.4%

Instruction & Long Context

  • Long-Context Reasoning46.7%

Benchmarks measured Aug 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
IBM
Modalities
Text

Capabilities

Tool Calling
No
Open Source
No

Strengths & Limitations

Strengths

  • Measured time to first token of approximately 249 ms, supporting responsive streaming chat interfaces
  • Measured output throughput of about 77 tokens per second in third-party Artificial Analysis benchmarking
  • Mid-size ~30B scale offers a middle point between compact Granite variants and much larger models on cost and latency
  • Part of IBM's Granite series, a family with multiple size tiers that allows moving up or down in scale without changing vendor ecosystems
  • Chat-tuned rather than base-only, so it can be used for instruction-following and multi-turn dialogue without additional fine-tuning
  • Sized to be a practical target for self-hosting on multi-GPU or single high-memory-GPU configurations, giving deployment flexibility

Limitations

  • Context window is not confirmed in our data — verify the supported length with your serving provider before designing long-document workflows
  • Modality support (image or audio input) is not tracked for this entry; treat it as a text chat model unless the provider states otherwise
  • Tool calling and structured output support are not confirmed in our metadata and should be checked against provider documentation
  • No standardized quality benchmark scores (reasoning, coding, math) are recorded for this model in our database, making direct capability comparison with peers difficult
  • At roughly 30B scale it is unlikely to match much larger frontier models on the hardest reasoning and long-horizon agentic tasks
  • Provider availability for Granite models is generally narrower than for the most widely hosted open model families

Key Features

Chat and instruction-following text generation
Approximately 30B parameter scale within the Granite 4.2 generation
~249 ms measured time to first token (Artificial Analysis)
~77 output tokens per second measured throughput (Artificial Analysis)
Part of IBM's multi-tier Granite model family
Streaming token output suitable for interactive applications
Available through API providers tracked in the pricing table on this page

About Granite 4.2 30B

Granite 4.2 30B is a text chat model released by IBM as part of its Granite series of general-purpose language models. The name places it in the 4.2 generation of Granite at approximately a 30-billion-parameter scale, sitting in the mid-size range of the family — larger than the compact Granite variants intended for edge and CPU deployment, and smaller than the largest configurations IBM has shipped under the Granite name. IBM has historically developed Granite for enterprise and business-application workloads rather than as a consumer chat assistant. On measured serving performance, third-party benchmarking from Artificial Analysis records Granite 4.2 30B at roughly 77 output tokens per second with a time to first token of about 249 milliseconds. That combination — sub-quarter-second first-token latency and a throughput figure in the high-70s — puts it in a range suitable for interactive chat and streaming interfaces, though actual numbers will vary by hosting provider, hardware, batch size, and prompt length. We do not currently track a confirmed context window, modality set, or tool-calling specification for this entry, so buyers evaluating it for a specific integration should confirm those details against the serving provider's own documentation. In practice, a 30B-class chat model like Granite 4.2 30B is typically deployed where a mid-size model offers a better cost and latency profile than a much larger frontier model, while still handling multi-turn conversation, summarization, and instruction-following work. Compared with the very small Granite variants, it trades higher per-token serving cost for more headroom on reasoning-heavy prompts; compared with much larger models, it trades some capability ceiling for faster responses and cheaper inference. Check the pricing table on this page to see which providers currently serve it and how their rates compare.

Common Use Cases

Granite 4.2 30B fits workloads where a mid-size chat model is the right trade-off: customer-facing assistants and internal help desks that need fast first-token response, document and meeting summarization, drafting and rewriting business text, structured extraction from unstructured records, and retrieval-augmented question answering over an enterprise knowledge base. Its measured sub-250 ms time to first token makes it a reasonable candidate for latency-sensitive interactive UIs, while its throughput supports streaming long responses without noticeable stalling. Teams already standardizing on IBM's Granite family may use the 30B tier as the general-purpose workhorse, routing simple classification or routing calls to smaller Granite variants and escalating only the hardest reasoning tasks to a larger model. Before committing it to long-context document pipelines or tool-calling agent frameworks, confirm the supported context length and function-calling behavior with your chosen provider, since our metadata does not record those specifications.

Frequently Asked Questions

How much does Granite 4.2 30B cost to use?

Pricing depends on which provider serves the model and whether you are billed per input token, per output token, per hour of dedicated capacity, or through a self-hosted GPU deployment. Rates differ between providers and change over time, so consult the live pricing table on this page for current figures rather than relying on any fixed number.

What is Granite 4.2 30B best used for?

It suits general enterprise chat and text workloads — assistants, summarization, drafting, information extraction, and retrieval-augmented question answering — where a mid-size model gives acceptable quality at lower latency and cost than a much larger model. Its measured ~249 ms time to first token makes it a reasonable fit for interactive, streaming interfaces.

How fast is Granite 4.2 30B?

Third-party measurements from Artificial Analysis put it at roughly 77 output tokens per second with a time to first token of about 249 milliseconds. Real-world performance varies with the hosting provider, hardware, prompt length, and concurrent load, so treat these as reference points rather than guarantees.

What context window does Granite 4.2 30B support?

We do not have a confirmed context window recorded for this model in our database. Check the documentation of the specific provider you plan to use, since serving platforms sometimes cap the usable context below a model's maximum.

How does it compare to other models in the Granite family?

At roughly 30B parameters it sits above IBM's compact Granite variants, which are aimed at edge and low-cost high-volume use, and below the largest Granite configurations. The practical trade-off is more capability headroom than the small tiers at higher per-token serving cost, and faster, cheaper inference than the largest options at a lower capability ceiling.

Can Granite 4.2 30B accept images or other non-text input?

Our metadata records it as a chat model and does not confirm image, audio, or other non-text modality support. Unless your provider explicitly documents multimodal input, plan for text-only usage.