Private LLM for your business — your data never leaves

Your team wants AI. Your legal and security people want a straight answer about where the data goes. A private LLM gives you both: capable AI that runs inside your own network or cloud account, on infrastructure you control, with nothing sent to a third-party model vendor.

Why private

The question isn't whether AI helps. It's where your data ends up.

Most AI tools are someone else's service. Every document, contract, and customer record your team pastes in leaves your company and lands in an account you don't control. For a lot of businesses in the EU that is not a preference problem — it's a compliance problem.

Your data leaves the company

Client files, internal documents, and personal data go to a vendor's servers — often outside the EU. You can't inspect what's stored, for how long, or who has access.

Your data can train someone else's model

Consumer and prosumer AI plans have shifted their training terms more than once. What was opt-out last year may not be this year, and the decision isn't yours to make.

You can't answer the compliance questions

GDPR asks where personal data is processed, on what legal basis, and how it gets deleted. "It's in a chatbot somewhere" fails an audit, a security review, and an enterprise procurement questionnaire.

A private LLM removes the question. The model runs in your environment, in an EU region if you need one, and the data never crosses your boundary. Your answer to the auditor is a diagram, not a hope.

What we deploy

Not just a model — the controls that make it safe to hand to your team.

Running a model is the easy part. What makes a private AI rollout survive the first quarter is knowing who used what, what it cost, and being able to stop it.

One controlled entry point

Every request goes through a gateway that sits in front of the model. One place to enforce access rules, route traffic, swap models, and keep an audit trail — instead of a dozen teams wiring up their own connections.

Budgets per team, with a hard stop

Each team gets its own budget. When it's spent, their access stops — automatically, not after someone notices the invoice. No surprise bills at the end of the month.

You can see what it costs and who used it

Cost, usage, and response times traced per request, per team, per project. When finance asks what AI cost last month, you have the number and the breakdown.

It fits the tools your team already uses

Developers keep their editors, your staff keep their internal apps. The gateway plugs in behind them, so adoption doesn't depend on anyone learning a new tool.

Under the hood: an LLM gateway (LiteLLM-class) in front of self-hosted open-weight models or a private cloud endpoint such as AWS Bedrock, with PostgreSQL and Langfuse for usage and cost tracing — deployed with containers into your account.

Want to know what a private LLM would look like on your infrastructure?

Start a conversation

From our work

We built this before we offered it.

R&D we ran · LLM cost governance

An LLM gateway with real budget control

Challenge
An engineering organisation wanted many teams using LLMs without unpredictable spend — and without blocking developers behind an approval queue.
Solution
We designed and built a gateway in front of the model provider, with per-key spending limits, per-team budgets, streaming-accurate usage counting, and full tracing of cost and latency.
Result
Budget overshoot verified at under 1% on real budgets; per-team hard cut-off confirmed (when a team's budget is gone, all of its keys stop at once); the gateway's added latency measured at roughly two seconds, with clear guidance on which traffic to route through it; validated against the real tools engineers use every day.

This was R&D we ran to production-readiness, not a client system we operate — we're naming it here because it's the capability we deliver, and the numbers are ours and measured.

Read the full case →

Deployment options

Two ways to keep the data inside.

Which one fits depends on your obligations and your existing infrastructure — we'll tell you honestly which one your case actually needs.

Your own cloud account

The model and the gateway run in your AWS, Azure, or Google Cloud account, in the region you choose — an EU region when data residency matters. Your account, your keys, your logs. The fastest route for most companies, and it scales without buying hardware.

On-premise, inside your network

For regulated work or a strict no-cloud policy: open-weight models on your own servers, with no outbound connection to any model vendor at all. Higher upfront effort and real hardware, in exchange for the strongest possible answer to "where is our data?"

How we work together

Pick the engagement that fits the job

Fixed scope

For a well-defined piece of work — a pilot deployment, a security review, one gateway. We agree on the outcome upfront, estimate each feature, and do our best to stay within 20% of the estimate.

Discovery first

For a bigger or fuzzier rollout: a paid Discovery (3–5 days, $1,500–3,000, fixed price agreed upfront) → artefacts, an architecture plan, and a priced proposal. You keep the artefacts either way.

Dedicated capacity

For ongoing work: a monthly retainer of reserved days for new use cases, model updates, and support — how we run our longest engagement today.

FAQ

Common questions.

What is a private LLM, in plain terms?

A large language model that runs on infrastructure you control — your cloud account or your own servers — instead of being a service you send your data to. Same kind of capability your team already knows from public AI tools, but the documents and prompts stay inside your boundary.

Is a self-hosted LLM for business as good as the big public models?

For a lot of business work — summarising, extracting data from documents, answering questions over internal content, drafting — open-weight models are now good enough that most users can't tell. For the hardest reasoning tasks the frontier models still lead. We benchmark on your actual task before recommending anything, and a gateway lets you mix: sensitive work stays private, non-sensitive work can go elsewhere.

Does private AI for small business make sense, or is this only for large companies?

It makes sense whenever the data is the constraint — a small accounting firm handling client financials has the same obligations as a large one. Small teams usually start in their own cloud account rather than on hardware, which keeps the entry cost close to ordinary infrastructure spend.

Can it answer questions from our internal documents?

Yes — that's one of the most common first use cases. We set up RAG for internal documents: your policies, contracts, and knowledge base get indexed so the model answers from your content with references, instead of guessing. The index lives in your environment along with everything else.

How does this help with GDPR and EU data residency?

You can say exactly where personal data is processed, keep it in an EU region or on your own hardware, delete it on request, and show the access controls and logs behind it. We're engineers, not your legal counsel — but we build the deployment so your legal and security teams have real answers to give, and documentation to point at.

How long does it take, and what does it cost?

A working pilot on your infrastructure typically lands in weeks, not quarters. Cost has two parts: our work to design and deploy it, and your own infrastructure bill — which the budgets and usage tracing we build are there to keep predictable. We estimate each feature upfront and do our best to stay within 20% of the estimate.

What if we already use a public AI vendor?

Most companies end up with both. The gateway becomes the single entry point either way: sensitive workloads route to the private model, everything else keeps going where it goes today — and you finally get one place where usage, cost, and access rules are visible.

Let's find out whether a private LLM fits your case.

Tell us what your team wants to do with AI and what your data rules are. We'll come back with the deployment that fits — and say so if you don't need one.

Start a conversation