Llama 3.2 Instruct 11B
Llama 3.2 Instruct 11B is an open-weight instruction-tuned chat model from Meta, positioned as a mid-small option in the Llama 3.2 generation.
API Pricing
| Provider | Input / 1M | Output / 1M |
|---|---|---|
| $0.160 | $0.160 |
Prices updated daily. Last check: Oct 10, 2026
Compare API pricing for every Llama model →Llama 3.2 Instruct 11B pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- MMLU-Pro46.4%
- GPQA Diamond22.1%
- Humanity's Last Exam5.5%
Coding
- LiveCodeBench11.0%
Math
- AIME 20251.7%
- AIME9.3%
- MATH-50051.6%
Agentic & Tool Use
- Terminal-Bench Hard0.8%
- τ²-bench14.6%
Instruction & Long Context
- IFBench30.4%
- Long-Context Reasoning12.7%
Benchmarks measured Oct 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- Meta
- Modalities
- Text
Capabilities
- Open Source
- Yes
Strengths & Limitations
Strengths
- Open weights under Meta's Llama community license, enabling self-hosting and on-premises or air-gapped deployment
- Served by multiple independent inference providers, so you can compare hosts and switch without rewriting prompts
- Measured time to first token of roughly 393 ms, suitable for streaming interactive chat
- Measured output throughput of roughly 55 tokens per second in our benchmark data
- Mid-small 11B parameter size fits on a single modern accelerator in common quantized configurations
- Part of the 11B/90B branch of Llama 3.2, the sizes Meta shipped with vision-enabled variants
- Instruction-tuned out of the box, so it follows chat-style system and user prompts without extra fine-tuning
Limitations
- An 11B-parameter model will trail much larger Llama models and frontier closed models on complex multi-step reasoning and long-form code generation
- Effective context length, image input, and tool-calling support vary by provider rather than being guaranteed by the model name
- Throughput near 55 tokens per second is modest compared with heavily optimized small models on the same hardware class
- Llama's community license carries acceptable-use and attribution terms that need review before commercial deployment
- We do not track benchmark accuracy scores for this model, so quality comparisons should be validated on your own evaluation set
Key Features
About Llama 3.2 Instruct 11B
Common Use Cases
Llama 3.2 Instruct 11B suits general-purpose assistant workloads where an open-weight, mid-small model gives an acceptable quality-to-cost ratio: customer-facing chat and support triage, document summarization, retrieval-augmented question answering over an internal knowledge base, structured data extraction, content drafting, and classification or tagging pipelines. Its sub-400 ms measured time to first token makes it a reasonable fit for interactive UIs where perceived responsiveness matters more than peak reasoning depth. Because the weights are openly distributed, it is also a common choice for regulated or privacy-sensitive deployments that must run inference inside their own infrastructure, and for teams that plan to fine-tune on domain data. For very high-volume, latency-critical classification, the smaller Llama 3.2 1B and 3B models are usually the better economics; for hard reasoning, long agentic chains, or advanced code generation, a larger Llama or a frontier-tier model is the more appropriate step up.
Frequently Asked Questions
How much does Llama 3.2 Instruct 11B cost to run?
Because the weights are openly available, the model is hosted by several inference providers, and pricing differs between them and by pricing type (per-token serverless, dedicated capacity, or your own GPU rental for self-hosting). Rather than quoting a figure that would go stale, see the pricing table on this page for current per-provider rates, and compare those against GPU hourly costs if you plan to self-host.
What is Llama 3.2 Instruct 11B best used for?
General-purpose chat and assistant tasks, summarization, retrieval-augmented question answering, extraction, and classification — workloads where an 11B open-weight model is sufficient and you want the option to self-host or fine-tune. Its measured ~393 ms time to first token makes it usable for streaming interactive interfaces.
Can Llama 3.2 Instruct 11B accept image input?
The 11B size sits in the branch of Llama 3.2 that Meta shipped with vision-enabled variants, but whether a given endpoint exposes image input depends on which checkpoint the provider deployed and how their API is configured. Check the specific provider's documentation before building an image-dependent workflow.
How does it compare with the smaller Llama 3.2 1B and 3B models?
The 1B and 3B models are text-oriented and cheaper and faster to serve, which makes them attractive for high-volume, simple tasks like routing and tagging. The 11B model has more parameter capacity for instruction following and multi-step prompts, at higher compute cost per request. If your task is narrow and high-volume, test the smaller sizes first; if quality is the constraint, 11B is the next step up.
Can I self-host Llama 3.2 Instruct 11B instead of using an API?
Yes. Meta distributes Llama weights under its community license, and an 11B model fits on a single modern accelerator in common quantized configurations. Self-hosting shifts the cost from per-token API billing to GPU hours plus operational overhead, and it requires reviewing the license's acceptable-use and attribution terms for your deployment.
How fast is it?
Our benchmark data records roughly 55 output tokens per second with a time to first token of about 393 ms. Actual figures vary by provider, hardware, quantization, batch size, and prompt length, so treat these as a reference point rather than a guarantee for any particular endpoint.