Skip to main content
Open SourceMeta

Llama 3.2 Instruct 11B

Llama 3.2 Instruct 11B is an open-weight instruction-tuned chat model from Meta, positioned as a mid-small option in the Llama 3.2 generation.

License Open Source
Input from
$0.160 / 1M tokens
across 1 provider

API Pricing

ProviderInput / 1MOutput / 1M
$0.160$0.160

Prices updated daily. Last check: Oct 10, 2026

Compare API pricing for every Llama model →

Llama 3.2 Instruct 11B pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
5.4 / 100
Math
1.7 / 100
Output Speed
22.5 t/s
Latency (TTFT)
816ms

Reasoning & Knowledge

  • MMLU-Pro46.4%
  • GPQA Diamond22.1%
  • Humanity's Last Exam5.5%

Coding

  • LiveCodeBench11.0%

Math

  • AIME 20251.7%
  • AIME9.3%
  • MATH-50051.6%

Agentic & Tool Use

  • Terminal-Bench Hard0.8%
  • τ²-bench14.6%

Instruction & Long Context

  • IFBench30.4%
  • Long-Context Reasoning12.7%

Benchmarks measured Oct 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
Meta
Modalities
Text

Capabilities

Open Source
Yes

Strengths & Limitations

Strengths

  • Open weights under Meta's Llama community license, enabling self-hosting and on-premises or air-gapped deployment
  • Served by multiple independent inference providers, so you can compare hosts and switch without rewriting prompts
  • Measured time to first token of roughly 393 ms, suitable for streaming interactive chat
  • Measured output throughput of roughly 55 tokens per second in our benchmark data
  • Mid-small 11B parameter size fits on a single modern accelerator in common quantized configurations
  • Part of the 11B/90B branch of Llama 3.2, the sizes Meta shipped with vision-enabled variants
  • Instruction-tuned out of the box, so it follows chat-style system and user prompts without extra fine-tuning

Limitations

  • An 11B-parameter model will trail much larger Llama models and frontier closed models on complex multi-step reasoning and long-form code generation
  • Effective context length, image input, and tool-calling support vary by provider rather than being guaranteed by the model name
  • Throughput near 55 tokens per second is modest compared with heavily optimized small models on the same hardware class
  • Llama's community license carries acceptable-use and attribution terms that need review before commercial deployment
  • We do not track benchmark accuracy scores for this model, so quality comparisons should be validated on your own evaluation set

Key Features

•Instruction-tuned chat model in Meta's Llama 3.2 generation
•Open weights available for self-hosting and fine-tuning
•Long context support as documented by Meta for the Llama 3.2 series (verify the limit your provider serves)
•Part of the 11B/90B vision-enabled branch of Llama 3.2
•Streaming generation with ~393 ms measured time to first token
•~55 output tokens per second measured throughput
•Available across multiple third-party inference APIs for price and latency comparison
•Deployable on a single accelerator in common quantized formats

About Llama 3.2 Instruct 11B

Llama 3.2 Instruct 11B is an instruction-tuned chat model released by Meta as part of the Llama 3.2 family. Within that family it sits between the compact 1B/3B text models and the larger 11B/90B-class options, making it a mid-small tier choice for teams that want an openly distributed model they can either self-host or rent through a hosted inference API. Because Meta publishes Llama weights under a community license, the same model is served by many independent providers, which is why the pricing table on this page typically shows multiple entries for it. In our measured data (sourced from Artificial Analysis), the model generates roughly 55 output tokens per second with a time to first token of about 393 ms. That combination puts it in the range where streaming responses feel responsive in interactive chat and assistant interfaces, though throughput on any given endpoint depends on the host's hardware, batching, and quantization choices. Meta's documentation for the Llama 3.2 series describes a long context window (commonly cited as 128K tokens), but individual providers frequently serve shorter effective limits, so the served context length should be confirmed against the specific provider's documentation. The 11B size in the Llama 3.2 generation is the smaller of Meta's two vision-enabled Llama 3.2 models; whether image input is exposed depends on the endpoint you use. In practice, Llama 3.2 Instruct 11B is used for general-purpose assistant work, summarization, drafting, retrieval-augmented question answering, and lightweight classification or extraction where a very large frontier model would be more capacity than the task requires. Compared with the 1B and 3B Llama 3.2 models it offers more headroom on reasoning-heavy prompts; compared with larger Llama models and closed frontier models it trades peak quality for lower resource requirements and the flexibility of open weights.

Common Use Cases

Llama 3.2 Instruct 11B suits general-purpose assistant workloads where an open-weight, mid-small model gives an acceptable quality-to-cost ratio: customer-facing chat and support triage, document summarization, retrieval-augmented question answering over an internal knowledge base, structured data extraction, content drafting, and classification or tagging pipelines. Its sub-400 ms measured time to first token makes it a reasonable fit for interactive UIs where perceived responsiveness matters more than peak reasoning depth. Because the weights are openly distributed, it is also a common choice for regulated or privacy-sensitive deployments that must run inference inside their own infrastructure, and for teams that plan to fine-tune on domain data. For very high-volume, latency-critical classification, the smaller Llama 3.2 1B and 3B models are usually the better economics; for hard reasoning, long agentic chains, or advanced code generation, a larger Llama or a frontier-tier model is the more appropriate step up.

Frequently Asked Questions

How much does Llama 3.2 Instruct 11B cost to run?

Because the weights are openly available, the model is hosted by several inference providers, and pricing differs between them and by pricing type (per-token serverless, dedicated capacity, or your own GPU rental for self-hosting). Rather than quoting a figure that would go stale, see the pricing table on this page for current per-provider rates, and compare those against GPU hourly costs if you plan to self-host.

What is Llama 3.2 Instruct 11B best used for?

General-purpose chat and assistant tasks, summarization, retrieval-augmented question answering, extraction, and classification — workloads where an 11B open-weight model is sufficient and you want the option to self-host or fine-tune. Its measured ~393 ms time to first token makes it usable for streaming interactive interfaces.

Can Llama 3.2 Instruct 11B accept image input?

The 11B size sits in the branch of Llama 3.2 that Meta shipped with vision-enabled variants, but whether a given endpoint exposes image input depends on which checkpoint the provider deployed and how their API is configured. Check the specific provider's documentation before building an image-dependent workflow.

How does it compare with the smaller Llama 3.2 1B and 3B models?

The 1B and 3B models are text-oriented and cheaper and faster to serve, which makes them attractive for high-volume, simple tasks like routing and tagging. The 11B model has more parameter capacity for instruction following and multi-step prompts, at higher compute cost per request. If your task is narrow and high-volume, test the smaller sizes first; if quality is the constraint, 11B is the next step up.

Can I self-host Llama 3.2 Instruct 11B instead of using an API?

Yes. Meta distributes Llama weights under its community license, and an 11B model fits on a single modern accelerator in common quantized configurations. Self-hosting shifts the cost from per-token API billing to GPU hours plus operational overhead, and it requires reviewing the license's acceptable-use and attribution terms for your deployment.

How fast is it?

Our benchmark data records roughly 55 output tokens per second with a time to first token of about 393 ms. Actual figures vary by provider, hardware, quantization, batch size, and prompt length, so treat these as a reference point rather than a guarantee for any particular endpoint.