Hermes 3 - Llama-3.1 70B
Hermes 3 - Llama-3.1 70B is Nous Research's instruction-tuned chat model built on Meta's Llama 3.1 70B base, positioned as the mid-size member of the Hermes 3 series.
API Pricing
| Provider | Input / 1M | Output / 1M |
|---|---|---|
| $0.700 | $0.700 | |
| $0.700 | $0.700 |
Prices updated daily. Last check: Sep 29, 2026
Hermes 3 - Llama-3.1 70B pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- MMLU-Pro57.1%
- GPQA Diamond40.1%
- Humanity's Last Exam4.0%
Coding
- LiveCodeBench18.8%
Math
- AIME2.3%
- MATH-50053.8%
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- Nous Research
- Modalities
- Text
Capabilities
- Tool Calling
- No
- Open Source
- No
Strengths & Limitations
Strengths
- Open-weight model — weights are downloadable and can be self-hosted or fine-tuned further, subject to Meta's Llama 3.1 license terms
- Built on the Llama 3.1 70B base, so it runs on the mature Llama inference tooling ecosystem (vLLM, TGI, llama.cpp, TensorRT-LLM) without custom code
- Tuned specifically for system-prompt adherence, which helps applications that need a persistent persona or strict output format
- Chat template includes conventions for tool/function calling and structured response blocks, suitable for agent scaffolds
- Mid-size 70B footprint is servable on a single multi-GPU node, unlike the 405B member of the same family
- Inherits the long context window of the Llama 3.1 base, with the exact served limit depending on the provider
- Available from multiple inference providers as well as for local deployment, so pricing and latency can be compared across hosts
Limitations
- We do not track verified throughput or time-to-first-token benchmarks for this model, so speed claims should be validated on your chosen endpoint
- As a 70B dense model, it needs substantial GPU memory to self-host at full precision — quantization is usually required for single-GPU setups
- Capability ceiling is bounded by the Llama 3.1 70B base it was fine-tuned from; it does not add new pretraining knowledge
- Fewer hosted providers than mainstream Meta-released Llama instruct models, which can limit availability and price competition
- Behavior tuned toward steerability means safety guardrails differ from vendor-aligned models — applications may need their own moderation layer
Key Features
About Hermes 3 - Llama-3.1 70B
Common Use Cases
Hermes 3 - Llama-3.1 70B fits teams that want an open-weight assistant with predictable, controllable behavior: persona-driven chat products, interactive fiction and roleplay platforms, and internal assistants where a detailed system prompt must be respected turn after turn. Its tool-call formatting and structured output conventions make it usable as the reasoning core of agent scaffolds and retrieval pipelines, and the long inherited context supports document question answering and multi-turn sessions with large histories. Because the weights are downloadable, it is also chosen for on-premises or VPC deployments where data cannot leave the environment, and for synthetic data generation or distillation workflows where per-call API terms would be restrictive. For very high-volume, latency-sensitive classification or routing, the smaller 8B Hermes 3 or another lightweight model is usually the more economical choice; for the hardest reasoning tasks, the 405B variant or a larger frontier-class model may be more appropriate.
Frequently Asked Questions
How much does Hermes 3 - Llama-3.1 70B cost to use?
Pricing depends on which inference provider you use and whether you are billed per token, per GPU-hour for a dedicated deployment, or not at all because you are self-hosting the open weights. Rates for hosted endpoints also differ between input and output tokens. See the pricing table on this page for the current per-provider figures.
What is Hermes 3 - Llama-3.1 70B best used for?
It suits persona-driven assistants, roleplay and creative applications, agentic workflows that rely on consistent tool-call and structured-output formatting, and self-hosted deployments where open weights are a requirement. The 70B size is a middle ground between the cheaper 8B Hermes 3 and the much heavier 405B variant.
How is it different from Meta's Llama 3.1 70B Instruct?
Both share the same base model, parameter count, and architecture. The difference is post-training: Nous Research replaced Meta's instruction tuning with its own Hermes 3 data mixture, which emphasizes system-prompt steerability, persona consistency, and structured/tool-call output formats. Serving cost and hardware requirements are effectively the same.
Can I run Hermes 3 - Llama-3.1 70B on my own hardware?
Yes — the weights are published and usable with standard Llama-compatible runtimes such as vLLM, Text Generation Inference, and llama.cpp. A 70B dense model at full precision needs multiple high-memory GPUs; quantized builds reduce that requirement at some quality cost. Use is subject to Meta's Llama 3.1 license terms.
How fast is it?
We do not currently have verified output-tokens-per-second or time-to-first-token measurements for this model in our database. Throughput for a 70B dense model varies widely with the provider's hardware, batching, and quantization, so benchmark the specific endpoint listed in the pricing table before committing.