Text Link
1/5

Desired Service

PROJECT DETAILS

2/5

BUDGET

3/5

TIMEFRAME

4/5

CONTACT DETAILS

We use your details solely to process your enquiry. Details in our privacy policy.

5/5

THANK YOU

Our AI agent was quick: the reply to your initial enquiry is already in your inbox.

ERROR

THANK YOU!

We will get back to you shortly.

Local LLM · build and run

Local LLM: your data stays in house

Many companies have stopped cloud AI by policy because nobody can say where the inputs end up. A language model on your own hardware answers that question with the architecture itself. The policy stays as it is.

Runs inside your networkOpen modelsConnected to existing systemsFirst call free of charge
0
Transfers outside
with a fully local setup
3
Ways to build it
local, dedicated, mixed
2
Weeks to a test environment
before real data goes in
1
Server to start with
grown as needed
Local LLM
A local LLM is a large language model that runs on your own hardware or in a dedicated environment, so that no input is sent to an external service. Business data never leaves your network, which sits comfortably with policies that forbid passing data on. The same setup is also called an on-premise LLM or a self-hosted model.
01 The problem

What is banned is rarely the AI. It is the data leaving.

Read your own IT policy again and you will usually find a sentence like: business data must not leave the company. Generative AI is rarely named.

What blocks the request is an open question: where does an input go, how long does it stay there, and does anyone train on it? As long as nobody can answer that, the approval sits in a drawer.

If the model runs inside your network, the architecture answers the question on its own. That is a shorter route than renegotiating the policy.

48.83 % cite data protection
Among enterprises that considered AI and decided against it, 48.83 % name concerns about data protection as a reason and 52.52 % a lack of clarity about legal consequences. A setup where nothing leaves the building answers the first point with the architecture.
Source: Eurostat, Use of artificial intelligence in enterprises, 2024 survey, published December 2025. ec.europa.eu/eurostat
02 Decision

Cloud or local is the wrong first question.

Four points decide whether a setup meets your policy. If all four hold, a cloud service sometimes passes too.

01

Where it is processed

Physically: which country, which server. With default settings, requests end up in a different region more often than people expect.

RegionData centre
02

What stays behind

Input history, logs, caches. Deleted often only means deleted in that one place.

LogsCacheRetention
03

Whether it trains on it

That is set in the contract and in the admin console. One of the two is not enough.

ContractSettings
04

Who can prove it

When the auditor asks: do you have to write to the provider, or do you show your own log and you are done?

AuditEvidence
03 Setup

Three ways to build a local LLM

None of them is always right. It comes down to how sensitive the data is and what quality you need.

01

Fully local

An open model on a GPU server inside your network, served with vLLM or Ollama. Business data never leaves the network. The choice for HR files, legal and engineering data.

No transferOpen modelHigher upfront cost
02

Dedicated environment

A dedicated environment in an EU region, contractually without training on your data. Current commercial models with a fixed place of processing.

Fixed regionNo trainingHigh quality
03

Mixed

Sensitive data locally, general text externally. A rule decides what goes where and logs the decision. In practice the most common setup.

RoutingLogPractical
04 Comparison

What changes compared with the usual setup

The same function, a different answer at the next audit.

PointUsual setupWithout data leaving
Document storageConsolidated in an external storeStays on your file server, only the index is built
Transcribing meetingsRecordings go to an external serviceRuns on your own GPU, the recording stays inside
ModelA provider's APIOpen model on your own hardware
Search indexExternal vector databaseInside your network, deletions handled by you
Auditor's questionAnswered after asking the providerShow your own log, immediately
Model changesWhen the provider decidesWhen you decide, with a bit more effort

Scroll sideways for all columns →

05 Plainly

When a local LLM is the wrong choice

We do not want to sell a setup, we want to sell a system that runs. If local does not fit, we say so in the first call.

One reason often remains anyway: the request gets approved, and at the next audit you answer yourself.

Upfront cost
A server with an 80 GB GPU costs five figures. The money is due at the start rather than monthly.
Updates
A new model does not swap itself in. Your team does that, or the maintenance contract covers it.
Quality
For long reasoning chains and subtle phrasing in several languages the large commercial models are better. For search, summaries and forms an open model is enough.
Proportion
For public material a local setup is overkill. A cloud service checked against the four points above will do.
06 In practice

Speech stays in house, in production for months

In the meeting project the speech recognition runs on the customer's own hardware. Meetings contain HR matters, prices and plans nobody is supposed to know yet. Processing them outside the network was never an option.

Name matching runs against the internal staff list. Who took on which task is nothing anyone outside needs to know.

Details are in the meeting dashboard case study. What a multilingual knowledge system looks like when index and search stay in house is shown in the knowledge agent case study.

Speech recognition
Open model on the customer's GPU, no external service.
Name matching
Against the internal list, in the same network.
Index
PostgreSQL with vector and full text search on the customer's hardware.
Handover
Manual and technical documentation for the EU AI Act.
07 Common questions

Local LLM: common questions

Our policy bans ChatGPT. Is a local LLM the replacement?+
In many cases, yes. Most policies forbid passing business data to third parties and do not mention generative AI at all. If the model runs inside your network, the policy stays exactly as written. We start by reading the wording together.
What hardware does a local LLM need?+
As a rule of thumb: a model with 7 to 8 billion parameters runs on a single GPU with 24 GB of memory, such as an RTX 4090 or an L4. Models around 70 billion parameters need 80 GB or more, which means an H100 or two smaller cards. Search and summarisation are fine on the small class. We start with one server and grow it when usage justifies that.
How does this sit with GDPR?+
With a fully local setup there is no transfer to a processor, so no Art. 28 GDPR contract is needed for the model. You still need the record of processing activities, a deletion concept and access rights. We prepare those with you.
Is it cheaper than an API?+
It depends on volume. Below a few thousand requests a day the API is usually cheaper, because the hardware costs five figures up front and power and maintenance come on top. Above that the maths flips. For sensitive data the price matters less than whether the request gets approved at all.
How good are open models compared with cloud models?+
For searching documents, summaries and standard texts in English and German, current open models are good enough. For long chains of reasoning, rare languages and subtle phrasing the large commercial models are ahead. So you can judge for yourself, you get a test environment with your own data first.
Can it connect to our existing systems?+
Yes. File servers, groupware and ERP are the usual candidates. The connection is the part whose effort is hardest to estimate, so we test it first and only then put the model in.
Who maintains the system?+
Model swaps, rebuilding the index and fixing outages are covered by the maintenance contract. The aim is that your team can operate it. At handover you get a manual, and we walk the responsible person through the routine steps.
08 Read on

Related topics

Next step

Tell us which sentence in your policy it is stuck on.

Thirty minutes are enough to say which setup gets through. If the cloud is enough, we say that too.

DEJPEN