Llama 3.1 8B pricing
$ Affordable$0.05/M input · $0.08/M output. List prices from provider data; in idapt, chat is metered per token with transparent, usage-based pricing at the prices shown in the app.
Full rate table
| Rate | Price | Unit |
|---|---|---|
| Input | 0.049999999999999996 | /M tokens |
| Output | 0.08 | /M tokens |
| Cache read | 0.024999999999999998 | /M tokens |
What typical work costs
| Workload | Tokens | List price |
|---|---|---|
| Quick question | 1K tokens in, 500 out | $0.000090 |
| Long-document summary | 60K tokens in, 2K out | $0.0032 |
| Agent working session | 400K tokens in, 40K out across a day | $0.02 |
Price your own workload
- Per call
- $0.000090
- Per month
- $0.0090
Compare up to four models at once in the cost calculator.
Frequently asked
What does Llama 3.1 8B cost per 1M tokens?
Llama 3.1 8B lists at $0.05 per 1M input tokens and $0.08 per 1M output tokens.
Does Llama 3.1 8B discount cached input tokens?
Yes. Cached input tokens are billed at $0.02 per 1M, versus $0.05 for fresh input. Prompts that repeat a long prefix (a system prompt, attached files) benefit the most.
What does a typical chat message with Llama 3.1 8B cost?
A message with about 1,000 input tokens and a 500-token reply costs roughly $0.0001 at list price. Long documents and long replies scale that linearly with token counts.
How does idapt bill Llama 3.1 8B?
Llama 3.1 8B is in idapt's free lane: signed-in users run it against the free daily allowance, with transparent, usage-based pricing beyond it at the prices shown in the app.
Part of the Llama 3.1 8B model page · run it locally · cheapest models board · price changes