Granite 4.2 30B
Granite 4.2 30B is a chat-oriented large language model from IBM, part of the Granite family, positioned at roughly 30B parameter scale within the 4.2 generation.
API Pricing
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.160 | $0.650 | $0.040 |
Prices updated daily. Last check: Aug 29, 2026
Granite 4.2 30B pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- GPQA Diamond64.4%
- Humanity's Last Exam11.2%
Coding
- SciCode36.6%
Agentic & Tool Use
- Terminal-Bench v2.126.6%
- τ-bench Banking14.4%
Instruction & Long Context
- Long-Context Reasoning46.7%
Benchmarks measured Aug 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- IBM
- Modalities
- Text
Capabilities
- Tool Calling
- No
- Open Source
- No
Strengths & Limitations
Strengths
- Measured time to first token of approximately 249 ms, supporting responsive streaming chat interfaces
- Measured output throughput of about 77 tokens per second in third-party Artificial Analysis benchmarking
- Mid-size ~30B scale offers a middle point between compact Granite variants and much larger models on cost and latency
- Part of IBM's Granite series, a family with multiple size tiers that allows moving up or down in scale without changing vendor ecosystems
- Chat-tuned rather than base-only, so it can be used for instruction-following and multi-turn dialogue without additional fine-tuning
- Sized to be a practical target for self-hosting on multi-GPU or single high-memory-GPU configurations, giving deployment flexibility
Limitations
- Context window is not confirmed in our data — verify the supported length with your serving provider before designing long-document workflows
- Modality support (image or audio input) is not tracked for this entry; treat it as a text chat model unless the provider states otherwise
- Tool calling and structured output support are not confirmed in our metadata and should be checked against provider documentation
- No standardized quality benchmark scores (reasoning, coding, math) are recorded for this model in our database, making direct capability comparison with peers difficult
- At roughly 30B scale it is unlikely to match much larger frontier models on the hardest reasoning and long-horizon agentic tasks
- Provider availability for Granite models is generally narrower than for the most widely hosted open model families
Key Features
About Granite 4.2 30B
Common Use Cases
Granite 4.2 30B fits workloads where a mid-size chat model is the right trade-off: customer-facing assistants and internal help desks that need fast first-token response, document and meeting summarization, drafting and rewriting business text, structured extraction from unstructured records, and retrieval-augmented question answering over an enterprise knowledge base. Its measured sub-250 ms time to first token makes it a reasonable candidate for latency-sensitive interactive UIs, while its throughput supports streaming long responses without noticeable stalling. Teams already standardizing on IBM's Granite family may use the 30B tier as the general-purpose workhorse, routing simple classification or routing calls to smaller Granite variants and escalating only the hardest reasoning tasks to a larger model. Before committing it to long-context document pipelines or tool-calling agent frameworks, confirm the supported context length and function-calling behavior with your chosen provider, since our metadata does not record those specifications.
Frequently Asked Questions
How much does Granite 4.2 30B cost to use?
Pricing depends on which provider serves the model and whether you are billed per input token, per output token, per hour of dedicated capacity, or through a self-hosted GPU deployment. Rates differ between providers and change over time, so consult the live pricing table on this page for current figures rather than relying on any fixed number.
What is Granite 4.2 30B best used for?
It suits general enterprise chat and text workloads — assistants, summarization, drafting, information extraction, and retrieval-augmented question answering — where a mid-size model gives acceptable quality at lower latency and cost than a much larger model. Its measured ~249 ms time to first token makes it a reasonable fit for interactive, streaming interfaces.
How fast is Granite 4.2 30B?
Third-party measurements from Artificial Analysis put it at roughly 77 output tokens per second with a time to first token of about 249 milliseconds. Real-world performance varies with the hosting provider, hardware, prompt length, and concurrent load, so treat these as reference points rather than guarantees.
What context window does Granite 4.2 30B support?
We do not have a confirmed context window recorded for this model in our database. Check the documentation of the specific provider you plan to use, since serving platforms sometimes cap the usable context below a model's maximum.
How does it compare to other models in the Granite family?
At roughly 30B parameters it sits above IBM's compact Granite variants, which are aimed at edge and low-cost high-volume use, and below the largest Granite configurations. The practical trade-off is more capability headroom than the small tiers at higher per-token serving cost, and faster, cheaper inference than the largest options at a lower capability ceiling.
Can Granite 4.2 30B accept images or other non-text input?
Our metadata records it as a chat model and does not confirm image, audio, or other non-text modality support. Unless your provider explicitly documents multimodal input, plan for text-only usage.