Llama 3.3 70B API pricing on Heabsy
Every Heabsy Llama 3.3 70B rate we track, compared against 10 providers serving the same model.
The cheapest Llama 3.3 70B input price we currently track is $0.100/1M on Deep Infra.
Heabsy Llama 3.3 70B rates
Prices per 1M tokens. Batch and cached-input rates are shown where the provider publishes them.
| Mode | Input | Output |
|---|---|---|
| Standard | $0.150/1M | $0.490/1M |
About Llama 3.3 70B
- Creator
- Meta
- Context
- 128K
- Modalities
- text
- Knowledge cutoff
- Mar 2024
- Tool calling
- Yes
- Open weights
- Yes
Llama 3.3 70B on other providers
About Heabsy
Heabsy is an OpenAI- and Anthropic-compatible inference API for open-weight models, operated by FEYA, s.r.o. in Bratislava, Slovakia, an EU company with no US parent. Models in the EEA tier, including Qwen3.8 27B, run on GPUs that Heabsy owns and operates inside the European Economic Area, with zero data retention: prompts and completions are processed in memory and never written to logs. A routed tier adds more than 25 open models (DeepSeek, Kimi, GLM, gpt-oss, Llama, Gemma, Qwen and others) bought from global providers under a European contract; every model in the catalog is labelled with where it is processed. One API key, one EU invoice, per-token pricing with no minimum spend or subscription.
Frequently Asked Questions
How much does Llama 3.3 70B cost on Heabsy?
Heabsy serves Llama 3.3 70B from $0.150 per 1M input tokens, across 1 tracked pricing mode. Prices are collected daily — see the table above for current input, output, and cached-input rates.
Is Heabsy the cheapest way to run Llama 3.3 70B?
Heabsy ranks #4 of 10 providers we track serving Llama 3.3 70B. The cheapest input price right now is on Deep Infra. Price is only one factor — throughput, latency, and context limits differ between providers.
Does Heabsy offer batch or cached pricing for Llama 3.3 70B?
We track 1 Llama 3.3 70B offering from Heabsy. Batch and cached-input rates appear in the table above when the provider publishes them; blank cells mean we have no current figure for that mode.