Self-hosted LLM for a small team: what it takes, and when it pays

By Arshad Ansari

A self-hosted LLM is an open-weight language model running on hardware you control, answering requests over your own network. Nothing you send it leaves the building. That is the main reason teams ask about it, and it is a good reason.

The second reason people give is cost, and that one is often wrong. I ran my own system's thinking on an 8 GB GPU for months, then moved it to a hosted API, and kept one job at home. Which job stayed, and why, is most of what a small team needs to know before buying a card.

Is a self-hosted LLM worth it?

It depends on the job, not on the team. Three questions settle most cases:

  1. Must the data stay inside? Contracts, health data, client documents under an NDA. If yes, self-hosting may be the only option, and cost is secondary.
  2. Is the work steady or bursty? A GPU costs the same idle as busy. Steady, high-volume work uses what you paid for. A few hundred requests a day, arriving in bursts, does not.
  3. Does the job need judgement? Embeddings, labelling, short summaries: small models do these well. Extracting the right number from a messy document, or deciding what matters: small models do these badly.

Here is what that looked like in my own system, AEGIS. On a replay of 36 real financial emails with known answers, a hosted model extracted 32 correctly. A local 4B model got 13. The local model also averaged about 23 seconds a call against 2 seconds hosted. And the hosted bill for the whole agent fleet came to about five cents a day.

So I moved the judgement work to a hosted API. I did not move the embeddings. The full numbers are in why AEGIS moved off my local GPU; I won't repeat them here.

Hardware: how much VRAM a self-hosted LLM needs

The rule is simple. The model, plus its context, has to fit in GPU memory. If it does not, the server runs part of it on the CPU, and it slows down sharply.

Model size depends on parameter count and quantisation (how many bits store each weight). From Ollama's Llama 3.1 tags page, checked on 1 October 2026:

ModelBuildDownload size
Llama 3.1 8B4-bit (q4_K_M), the default4.9 GB
Llama 3.1 8B8-bit (q8_0)8.5 GB
Llama 3.1 8B16-bit (fp16)16 GB
Llama 3.1 70B4-bit (q4_K_M)43 GB

The context needs memory on top of the weights. Ollama's FAQ says required memory scales with the number of parallel requests times the context length. So four users with long prompts need much more than one.

Ollama also picks a default context from your VRAM. Its context length docs say: under 24 GiB of VRAM, 4k tokens; 24–48 GiB, 32k; 48 GiB or more, 256k. That matters for RAG, below.

My reading of those numbers, as a guide and not a promise:

  • 8 GB card: embeddings, and small models (up to about 8B at 4-bit) with short context. This is what I have.
  • 24 GB card: 8B-class models with long context, or somewhat larger models at 4-bit.
  • 70B-class models: about 43 GB of weights at 4-bit before any context, so two large cards or a workstation-class one.

Check the real split after you load a model. ollama ps shows a PROCESSOR column; you want 100% GPU.

And plan for heat. My GPU runbook fires an alert above 95°C. The first step is to scale the Ollama service to zero and wait until the card is below 70°C before starting it again. A runaway batch of requests can hold a card at 100% long enough to get there.

How to deploy a self-hosted LLM

Three servers cover most small-team setups. All three speak an OpenAI-shaped API, so code written against a hosted API can point at them with a base URL change.

Ollama: the simplest. One container, models pulled by name. From the Docker docs, after installing the NVIDIA Container Toolkit:

docker run -d --gpus=all -v ollama:/root/.ollama \
  -p 11434:11434 --name ollama ollama/ollama
docker exec -it ollama ollama pull llama3.1:8b

It binds to 127.0.0.1:11434 by default. It keeps a model in memory for five minutes after the last request, then unloads it; OLLAMA_KEEP_ALIVE changes that. It has its own API (/api/generate, /api/chat, /api/embed) and an OpenAI-compatible one.

llama.cpp's llama-server: more control. It runs a GGUF model file directly. From its README:

llama-server -m model.gguf -c 8192 -ngl 99 -np 2 --api-key "$LLAMA_API_KEY"

-c sets the context, -ngl how many layers go to the GPU, -np the number of parallel slots. It listens on 127.0.0.1:8080 and serves /v1/chat/completions and /v1/embeddings.

vLLM: for throughput. Built for many concurrent requests on a proper GPU. The quickstart lists Linux and Python 3.10–3.13, and serves on port 8000:

vllm serve Qwen/Qwen2.5-1.5B-Instruct --api-key "$VLLM_API_KEY"

For a team of five sharing one card, Ollama is enough. vLLM earns its extra setup when many people or jobs hit the model at once.

Three things that bit me running it

These come from running Ollama behind my Dagster pipeline, which uses it for short instrument summaries. The client is about 450 lines of Python.

A login page is not JSON. The Ollama endpoint sits behind an access proxy. When the service token was missing, the proxy answered with its HTML login page and HTTP 200. The status check passed. The JSON parse failed with "Expecting value: line 1 column 1", which says nothing about authentication. The client now sends the service-token headers, and treats a non-JSON response as a retryable failure.

Put auth in front, every time. Ollama's authentication docs say the local API "does not require authentication". Mine is reached over HTTPS through a reverse proxy with Basic Auth, and the password comes from a Docker secret file, never from code. llama-server and vLLM both take an --api-key. Use it.

Bound everything. The client sets a 120-second request timeout, three retries with backoff, and a cap of 1,024 output tokens. Without a cap, one model that decides to keep talking holds the GPU for everyone else. On the AEGIS side, a reasoning model once locked the inference server for an hour and a half.

Can a self-hosted LLM do RAG over your documents?

Yes. RAG (retrieval-augmented generation) has two halves, and a small GPU is good at one of them.

Retrieval: easy on small hardware. You embed each chunk of each document once, store the vectors, and at question time embed the question and find the nearest chunks. Embedding models are small. nomic-embed-text is 274 MB in Ollama's library. AEGIS keeps more than 500,000 chunks embedded with it on my own hardware, stored in Postgres with pgvector. One Ollama call does the embedding:

curl http://localhost:11434/api/embed \
  -d '{"model": "nomic-embed-text", "input": ["first chunk", "second chunk"]}'

You don't need a separate vector database for a team-sized corpus. DuckDB or SQLite can hold the vectors in-process; search your documents without a vector database shows the build.

One rule to set early: the stored vectors and the question vector must come from the same embedding model. Changing the embedding model later means re-embedding everything.

Generation: where small models struggle. The model reads the retrieved passages and writes an answer. Two problems show up:

  • The context may not fit. Under 24 GiB of VRAM, Ollama defaults to a 4k-token context. Five retrieved passages of 800 tokens each, plus the question and instructions, can pass that. Set the context length yourself and check memory again.
  • The answer can still be wrong. Retrieval finding the right passage does not mean a 4B model reads it correctly. Mine got 13 of 36 extractions right with the right email in front of it.

So test with questions you already know the answer to before anyone relies on it. That is the cheapest evaluation there is.

What it costs

Briefly, because the full sums are in local model or paid API: the honest math. Owned hardware is a fixed cost; an API is a per-token cost. Low, bursty volume favours the API. Steady, high volume favours the hardware. A model run over millions of rows overnight is the clearest case for owning the GPU, and running a local model over millions of rows covers how to do that without it falling over.

The part people skip: measure your volume first. I protected a five-cent-a-day cost for months because I hadn't counted.

A short checklist for a small team

  1. Write down which jobs the model will do. Split embeddings, labelling and judgement.
  2. Count the requests per day for each, and when they arrive.
  3. Pick the model, then check it fits: weights plus context, 100% GPU in ollama ps.
  4. Put authentication in front of the server before anyone else can reach it.
  5. Set a timeout, a retry limit and an output-token cap in the client.
  6. Route calls through one seam (a proxy, or one client) so moving a job between local and hosted is a config change.
  7. Test against questions with known answers, and keep the score.

So: should a small team self-host an LLM?

Self-host the jobs that are small, steady or private: embeddings, retrieval, labelling, bulk runs. Think hard before self-hosting judgement on a small card. My system does both today: questions are embedded at home, and answered by a hosted model. That isn't a compromise. It is the answer per job.

If you are deciding this for a team, the AI workflow teardown linked below runs through the same questions against your own setup. If your data can leave the building and your volume is low, you may not need a GPU at all, and you don't need me to tell you that.

Common questions

Is self-hosting an LLM worth it?
For some jobs, yes. Self-hosting pays when the work is steady and high-volume, when data must not leave your network, or when the job is small enough for a small model, such as embeddings, classification or summaries. It tends not to pay for bursty, low-volume judgement work on a small GPU. In my own system an 8 GB GPU handled embeddings well but lost badly on extraction: on 36 real financial emails a hosted model got 32 right and a local 4B model got 13, at about 23 seconds a call against 2. I moved generation to a hosted API and kept embeddings local.
What hardware do I need to self-host an LLM?
Mainly enough GPU memory to hold the model plus its context. Ollama's library lists Llama 3.1 8B at 4.9 GB in its default 4-bit (q4_K_M) build, 8.5 GB at 8-bit and 16 GB at fp16, and Llama 3.1 70B at 43 GB in 4-bit. The context needs memory on top, and Ollama's docs say required memory grows with the number of parallel requests times the context length. A model that does not fit is partly run on the CPU and slows down sharply; `ollama ps` shows the GPU/CPU split. An 8 GB card runs small models; a 24 GB card is a more comfortable floor for 8B-class models with long context.
How do I deploy a self-hosted LLM?
Three common servers. Ollama is the simplest: one install or one Docker container (`ollama/ollama`), it pulls models by name and serves an API on port 11434, including an OpenAI-compatible one. llama.cpp's `llama-server` runs a GGUF model file and serves OpenAI-compatible endpoints on port 8080, with flags for context size, GPU layers, parallel slots and an API key. vLLM (`vllm serve <model>`) is built for throughput on Linux with a GPU and serves on port 8000. All three bind to localhost by default; put authentication in front before you expose one.
Can a self-hosted LLM do RAG over company documents?
Yes, and the retrieval half is the easy part. Embedding models are small: nomic-embed-text is 274 MB in Ollama's library, and my own system keeps more than 500,000 chunks embedded with it on home hardware. Store the vectors in Postgres with pgvector, DuckDB or SQLite. The hard half is the answer: a small local model given retrieved passages can still answer badly, and Ollama's default context on a GPU under 24 GiB is 4k tokens, so long retrieved context may not fit unless you raise it. Test answers against questions you already know the answer to.
How much does a self-hosted LLM cost compared with an API?
It depends on volume more than anything. The hardware and electricity are fixed whether you send one request a day or a million; an API charges per token. At low or bursty volume an API is usually cheaper: my whole agent system spent about five cents a day on a hosted model. At steady high volume, such as a model run over millions of rows overnight, owned hardware can win. Work out your own token volume before you buy a GPU.

Get new posts by email

Data engineering notes like this one — what breaks and what it costs, in production.

What breaks and what it costs — pipelines, warehouse bills, and the failures that only show up in production. A few a month, never padded to hit a schedule. No sequence, no pitch deck. Reply 'stop' once and you're off — it reaches me, not a queue.

Want the whole playbook?

If this was useful, the long version is my book. Local-First Analytics — 313 pages, runnable code for every chapter — is the full build: DuckDB, Parquet and Arrow, from install to production. On Amazon, and chapter 1 is free to read here.

Get the book

Not ready to buy? Read chapter 1 free — the whole chapter, no email required.

Rather talk it through? Book a free 30-minute call. No slot that suits your time zone? Email info@hikmahtech.in.