Joinery

Private AI

Running an open language model on your own server: what a small firm needs to know

A workstation in a cupboard can now do most of what your team pastes into ChatGPT. Here is what the machine looks like, which models run on it, what it costs to keep running, and where it still falls short.

The question arrives in almost every first call now, usually from the partner who signs the data-protection forms: can we have this without sending our files to an American server? The answer in 2026 is yes, for most of the everyday work, and the machine that does it is smaller than people expect.

"On-premise LLM" is the industry term. It means a large language model, the kind of program behind ChatGPT, running as software on hardware you control: a workstation in your office, a server in your rack, or a dedicated server rented under your company's name in a data centre in your country. The model is a file. You download it once, and from then on every question and every document stays on that machine.

What the machine looks like

Three sizes cover nearly every firm with 5 to 50 people.

SetupWhat it isWho it servesHardware budget, roughly
WorkstationA desktop with one strong graphics card (24 to 48 GB of graphics memory), under a desk or in a cupboardA team of 10 to 20 with everyday tasks: mail, drafts, summaries, questions over documents3,000 to 8,000 EUR
ServerA rack machine with two to four graphics cards, or one workstation-class card with 96 GBFirms with heavy volume, or several agents running at once15,000 to 40,000 EUR
Rented dedicated serverA machine with a graphics card in a data centre in your country, under your name, nobody else on itFirms without a server room, or with a remote teamA monthly fee, no purchase

Apple's desktop machines with large unified memory are a fourth option that suits small teams well. They are quiet and low on power, and they hold a large model in memory, at the cost of slower output than a dedicated graphics card.

The graphics memory is the number that matters. It decides which models fit and how many people can ask at the same time. Everything else is ordinary office hardware.

Which models run on it

Open-weight models are the ones you can download and run yourself. Several families are mature enough for office work: Meta's Llama, Alibaba's Qwen, Mistral's models, Google's Gemma, and a growing number of others. They come in sizes, measured in billions of parameters. A model in the 7 to 14 billion class runs fast on a single card and handles sorting and extraction well. Models in the 30 to 70 billion class write and reason noticeably better and want a stronger card or two. Above that you are in server territory.

Most firms run a mid-sized model in a compressed form ("quantized" is the technical word), which cuts the memory need roughly in half with little visible loss in quality. The setup should be built so the model can be swapped: the open models improve every few months, and the one you install in autumn will be replaced by spring.

What it does reliably today

  • Reading and sorting incoming mail, and drafting replies from your own documents and price list.
  • Summarising long files: contracts, tenders, meeting recordings that were transcribed.
  • Answering questions over an indexed set of company documents, with the source shown next to the answer.
  • Pulling fields out of invoices, forms and delivery notes.
  • Translating and rewriting text in your house style.

Where the largest cloud models are still ahead: long chains of reasoning, hard mathematics, unusual programming tasks, and very long documents that must be held in memory at once. For a law firm summarising a 300-page file or an engineering office checking a complex calculation, a cloud model may still win. A sensible setup uses the local model by default and asks a person before anything is sent to a cloud model.

What it costs to keep running

Electricity: a workstation under load draws about as much as a space heater on its low setting, and it idles most of the day. Budget tens of euros per month, more for a server. Maintenance: someone has to install updates, watch the disk fill up, and swap the model when a better one arrives. That is a few hours a month for an IT partner, or a retainer. There are no per-user fees for an open model, which is the line item that makes cloud subscriptions expensive once a team grows past ten.

The fair comparison is against the subscription you would otherwise buy for the whole team over three years. For firms with more than ten users the machine often wins on cost by the second year, and it wins on the data question from the first day.

The data protection angle

With the model on your own hardware, the AI vendor drops off your list of processors for this work. Client data stays in the country, and the request log sits on a disk your data protection officer can read. Your own duties remain: the processing still has to be documented, access still has to be controlled, and staff still need to know what may go into the system. A local model makes those answers shorter. It does not make them disappear.

When it is not worth it

A firm of five that uses AI a few times a week is better served by a careful cloud subscription with the right settings. A firm whose work is mostly hard reasoning will be frustrated by a mid-sized model. And a firm with nobody willing to own the machine, not even through an IT partner, should not buy one. In each of those cases we say so on the call and suggest the smaller step.

Joinery sets up private AI for firms with 5 to 50 people, indexes their documents on the machine, and builds the first agents that work from it. The hardware is yours, the models are open, the logs stay on your disk. How the private AI setup works.

Frederik Theissen
Frederik Theissen

Founder of Joinery. Former neuroscience researcher, then quantitative risk data scientist at a crypto hedge fund and Lead Data Scientist at Accointing by Glassnode. Building with large language models daily since December 2022. LinkedIn