Skip to main content
Nous Research

Hermes 3 - Llama-3.1 70B

Hermes 3 - Llama-3.1 70B is Nous Research's instruction-tuned chat model built on Meta's Llama 3.1 70B base, positioned as the mid-size member of the Hermes 3 series.

Input from
$0.700 / 1M tokens
across 2 providers

API Pricing

ProviderInput / 1MOutput / 1M
$0.700$0.700
$0.700$0.700

Prices updated daily. Last check: Sep 29, 2026

Hermes 3 - Llama-3.1 70B pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
6.0 / 100

Reasoning & Knowledge

  • MMLU-Pro57.1%
  • GPQA Diamond40.1%
  • Humanity's Last Exam4.0%

Coding

  • LiveCodeBench18.8%

Math

  • AIME2.3%
  • MATH-50053.8%

Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
Nous Research
Modalities
Text

Capabilities

Tool Calling
No
Open Source
No

Strengths & Limitations

Strengths

  • Open-weight model — weights are downloadable and can be self-hosted or fine-tuned further, subject to Meta's Llama 3.1 license terms
  • Built on the Llama 3.1 70B base, so it runs on the mature Llama inference tooling ecosystem (vLLM, TGI, llama.cpp, TensorRT-LLM) without custom code
  • Tuned specifically for system-prompt adherence, which helps applications that need a persistent persona or strict output format
  • Chat template includes conventions for tool/function calling and structured response blocks, suitable for agent scaffolds
  • Mid-size 70B footprint is servable on a single multi-GPU node, unlike the 405B member of the same family
  • Inherits the long context window of the Llama 3.1 base, with the exact served limit depending on the provider
  • Available from multiple inference providers as well as for local deployment, so pricing and latency can be compared across hosts

Limitations

  • We do not track verified throughput or time-to-first-token benchmarks for this model, so speed claims should be validated on your chosen endpoint
  • As a 70B dense model, it needs substantial GPU memory to self-host at full precision — quantization is usually required for single-GPU setups
  • Capability ceiling is bounded by the Llama 3.1 70B base it was fine-tuned from; it does not add new pretraining knowledge
  • Fewer hosted providers than mainstream Meta-released Llama instruct models, which can limit availability and price competition
  • Behavior tuned toward steerability means safety guardrails differ from vendor-aligned models — applications may need their own moderation layer

Key Features

•Instruction-tuned chat model in the Hermes 3 series from Nous Research
•Fine-tuned on Meta's Llama 3.1 70B base architecture
•Open weights available for self-hosting and further fine-tuning
•Long context inherited from the Llama 3.1 base (exact served limit varies by provider)
•System-prompt steerability for persona and format control
•Tool/function-call formatting conventions in the chat template
•Structured output support for JSON and tagged response blocks
•Mid-tier sizing between the Hermes 3 8B and 405B variants

About Hermes 3 - Llama-3.1 70B

Hermes 3 - Llama-3.1 70B is a chat model released by Nous Research as part of the Hermes 3 series, a line of generalist instruction-tuned assistants fine-tuned on top of Meta's Llama 3.1 base models. The 70B variant sits between the smaller 8B release and the larger 405B release in the same family, making it the mid-tier option for teams that want Hermes-style behavior without serving a very large model. Because it is a fine-tune rather than a new pretrain, it inherits the Llama 3.1 architecture, tokenizer, and licensing terms from Meta while replacing the post-training with Nous Research's own instruction and alignment data. The Hermes line is characterized by strong system-prompt steerability: the tuning emphasizes following detailed persona and format instructions rather than overriding them with a fixed assistant voice. Hermes 3 also targets structured outputs, multi-turn agentic patterns, and function/tool-call formatting via conventions defined in its chat template. Context length follows the underlying Llama 3.1 base model, so hosted deployments commonly expose a long context window, though the exact limit served varies by provider — check the provider listings on this page rather than assuming a single number. We do not track independent throughput or latency measurements for this model, so serving speed should be evaluated against the specific endpoint you plan to use. In practice, Hermes 3 70B is used where an open-weight Llama 3.1 derivative is preferred over a closed API model: self-hosted assistants, roleplay and persona-driven applications, synthetic data generation, and agent scaffolds that depend on predictable formatting. Compared with Meta's own Llama 3.1 70B Instruct, the differences are behavioral rather than architectural — the same parameter count and inference cost profile, with different refusal characteristics, instruction adherence, and output style. Compared with the 405B Hermes 3, it trades some capability for substantially lower serving requirements.

Common Use Cases

Hermes 3 - Llama-3.1 70B fits teams that want an open-weight assistant with predictable, controllable behavior: persona-driven chat products, interactive fiction and roleplay platforms, and internal assistants where a detailed system prompt must be respected turn after turn. Its tool-call formatting and structured output conventions make it usable as the reasoning core of agent scaffolds and retrieval pipelines, and the long inherited context supports document question answering and multi-turn sessions with large histories. Because the weights are downloadable, it is also chosen for on-premises or VPC deployments where data cannot leave the environment, and for synthetic data generation or distillation workflows where per-call API terms would be restrictive. For very high-volume, latency-sensitive classification or routing, the smaller 8B Hermes 3 or another lightweight model is usually the more economical choice; for the hardest reasoning tasks, the 405B variant or a larger frontier-class model may be more appropriate.

Frequently Asked Questions

How much does Hermes 3 - Llama-3.1 70B cost to use?

Pricing depends on which inference provider you use and whether you are billed per token, per GPU-hour for a dedicated deployment, or not at all because you are self-hosting the open weights. Rates for hosted endpoints also differ between input and output tokens. See the pricing table on this page for the current per-provider figures.

What is Hermes 3 - Llama-3.1 70B best used for?

It suits persona-driven assistants, roleplay and creative applications, agentic workflows that rely on consistent tool-call and structured-output formatting, and self-hosted deployments where open weights are a requirement. The 70B size is a middle ground between the cheaper 8B Hermes 3 and the much heavier 405B variant.

How is it different from Meta's Llama 3.1 70B Instruct?

Both share the same base model, parameter count, and architecture. The difference is post-training: Nous Research replaced Meta's instruction tuning with its own Hermes 3 data mixture, which emphasizes system-prompt steerability, persona consistency, and structured/tool-call output formats. Serving cost and hardware requirements are effectively the same.

Can I run Hermes 3 - Llama-3.1 70B on my own hardware?

Yes — the weights are published and usable with standard Llama-compatible runtimes such as vLLM, Text Generation Inference, and llama.cpp. A 70B dense model at full precision needs multiple high-memory GPUs; quantized builds reduce that requirement at some quality cost. Use is subject to Meta's Llama 3.1 license terms.

How fast is it?

We do not currently have verified output-tokens-per-second or time-to-first-token measurements for this model in our database. Throughput for a 70B dense model varies widely with the provider's hardware, batching, and quantization, so benchmark the specific endpoint listed in the pricing table before committing.