Why Obsidian takes you higher than Berget
Berget AI sells a per-token API; we rent you the GPU. Per token suits small or bursty volume; a GPU of your own suits steady volume, isolation, custom stacks and fine-tuning. The table, the break-even and the questions below say where the line goes.
Side by side
Six differences.
Berget AI sells a per-token API; we rent you the GPU.
| Per-token API vs your own GPU | Per-token API | Your own GPU |
|---|---|---|
| What you buy | Tokens | The GPU |
| Who else runs on the hardware | A shared pool | Nobody during your term |
| Models | Their shelf | Any |
| Access | API only | Root + API on your box |
| Pricing | Per token, prepaid plan required | Per minute, published |
| When it wins | Small or bursty volume | Steady volume, isolation, custom stacks, fine-tuning |
When each wins
Pick by volume and control.
A per-token API wins when
- Volume is small or bursty: you pay per token and run nothing.
- A model on its shelf is exactly what you need.
- You want no ops: nothing to keep alive, nothing to configure.
Your own GPU wins when
- Volume is steady: above the break-even, tokens cost more than the machine.
- You need any model: Hugging Face or our mirror, fine-tuned or not.
- Isolation matters: nobody else on the hardware during your term.
- You run a custom stack: your vLLM settings, embeddings and speech on one box.
- You want root access and a fixed monthly price to budget.
Break-even
Where the line goes.
One dedicated 5090 at 8 000 SEK/month costs about the same as 1.5 billion output tokens at a typical Swedish per-token API’s list price. Below that, tokens are cheaper. Above it, or when you need isolation, a custom stack or fine-tuning, the GPU wins.
List price used: €0.50 per million output tokens, the published price of a 31B open model at a Swedish per-token API on 23 Sep 2026; 8 000 SEK/month converted at about 11 SEK per euro. Ex VAT.
Questions
Questions and answers.
- What is the difference between a per-token API and renting a GPU?
- With a per-token API you buy output tokens from a shared pool of hardware running the provider’s shelf of models. With us you rent the GPU itself: nobody else runs on it during your term, you run any model, and you have root access and an API endpoint on your own box.
- When is a per-token API the better choice?
- Small or bursty volume, a model that is on its shelf, and no need for isolation: you pay only for the tokens you use and run nothing yourself.
- When does renting the GPU win?
- Steady volume, a model that is not on the API’s shelf, fine-tuning, a custom stack, or when isolation matters: one machine, one customer, during your term.
- Where is the break-even?
- One dedicated 5090 at 8 000 SEK/month costs about the same as 1.5 billion output tokens at a typical Swedish per-token API’s list price (€0.50 per million output tokens on 23 Sep 2026). Below that, tokens are cheaper; above it, the GPU is.
- Can I get an OpenAI-compatible API from a rented GPU?
- Yes. The dedicated inference box serves the model you pick on an OpenAI-compatible endpoint (vLLM) behind your own key, so anything that speaks the OpenAI API can point at your box.
- How do I pay, and can I stop?
- On demand: a prepaid balance billed by the minute, the price locked 7 days at a time; stop whenever you like, and unused balance is refunded on request. Dedicated: a monthly price under a signed agreement, monthly or on a 12-month term.
- Is my data in Sweden?
- Yes. The hardware is ours and it is in Sweden; the contract, the invoice and the support come from Obsidian Peaks Technologies AB under Swedish law; we do not move your data out of the EU/EEA.
