Your data leaves the company
Client files, internal documents, and personal data go to a vendor's servers — often outside the EU. You can't inspect what's stored, for how long, or who has access.
Your team wants AI. Your legal and security people want a straight answer about where the data goes. A private LLM gives you both: capable AI that runs inside your own network or cloud account, on infrastructure you control, with nothing sent to a third-party model vendor.
Why private
Most AI tools are someone else's service. Every document, contract, and customer record your team pastes in leaves your company and lands in an account you don't control. For a lot of businesses in the EU that is not a preference problem — it's a compliance problem.
Client files, internal documents, and personal data go to a vendor's servers — often outside the EU. You can't inspect what's stored, for how long, or who has access.
Consumer and prosumer AI plans have shifted their training terms more than once. What was opt-out last year may not be this year, and the decision isn't yours to make.
GDPR asks where personal data is processed, on what legal basis, and how it gets deleted. "It's in a chatbot somewhere" fails an audit, a security review, and an enterprise procurement questionnaire.
A private LLM removes the question. The model runs in your environment, in an EU region if you need one, and the data never crosses your boundary. Your answer to the auditor is a diagram, not a hope.
What we deploy
Running a model is the easy part. What makes a private AI rollout survive the first quarter is knowing who used what, what it cost, and being able to stop it.
Every request goes through a gateway that sits in front of the model. One place to enforce access rules, route traffic, swap models, and keep an audit trail — instead of a dozen teams wiring up their own connections.
Each team gets its own budget. When it's spent, their access stops — automatically, not after someone notices the invoice. No surprise bills at the end of the month.
Cost, usage, and response times traced per request, per team, per project. When finance asks what AI cost last month, you have the number and the breakdown.
Developers keep their editors, your staff keep their internal apps. The gateway plugs in behind them, so adoption doesn't depend on anyone learning a new tool.
Under the hood: an LLM gateway (LiteLLM-class) in front of self-hosted open-weight models or a private cloud endpoint such as AWS Bedrock, with PostgreSQL and Langfuse for usage and cost tracing — deployed with containers into your account.
Want to know what a private LLM would look like on your infrastructure?
Start a conversationFrom our work
R&D we ran · LLM cost governance
This was R&D we ran to production-readiness, not a client system we operate — we're naming it here because it's the capability we deliver, and the numbers are ours and measured.
Read the full case →Deployment options
Which one fits depends on your obligations and your existing infrastructure — we'll tell you honestly which one your case actually needs.
The model and the gateway run in your AWS, Azure, or Google Cloud account, in the region you choose — an EU region when data residency matters. Your account, your keys, your logs. The fastest route for most companies, and it scales without buying hardware.
For regulated work or a strict no-cloud policy: open-weight models on your own servers, with no outbound connection to any model vendor at all. Higher upfront effort and real hardware, in exchange for the strongest possible answer to "where is our data?"
How we work together
For a well-defined piece of work — a pilot deployment, a security review, one gateway. We agree on the outcome upfront, estimate each feature, and do our best to stay within 20% of the estimate.
For a bigger or fuzzier rollout: a paid Discovery (3–5 days, $1,500–3,000, fixed price agreed upfront) → artefacts, an architecture plan, and a priced proposal. You keep the artefacts either way.
For ongoing work: a monthly retainer of reserved days for new use cases, model updates, and support — how we run our longest engagement today.
FAQ
A large language model that runs on infrastructure you control — your cloud account or your own servers — instead of being a service you send your data to. Same kind of capability your team already knows from public AI tools, but the documents and prompts stay inside your boundary.
For a lot of business work — summarising, extracting data from documents, answering questions over internal content, drafting — open-weight models are now good enough that most users can't tell. For the hardest reasoning tasks the frontier models still lead. We benchmark on your actual task before recommending anything, and a gateway lets you mix: sensitive work stays private, non-sensitive work can go elsewhere.
It makes sense whenever the data is the constraint — a small accounting firm handling client financials has the same obligations as a large one. Small teams usually start in their own cloud account rather than on hardware, which keeps the entry cost close to ordinary infrastructure spend.
Yes — that's one of the most common first use cases. We set up RAG for internal documents: your policies, contracts, and knowledge base get indexed so the model answers from your content with references, instead of guessing. The index lives in your environment along with everything else.
You can say exactly where personal data is processed, keep it in an EU region or on your own hardware, delete it on request, and show the access controls and logs behind it. We're engineers, not your legal counsel — but we build the deployment so your legal and security teams have real answers to give, and documentation to point at.
A working pilot on your infrastructure typically lands in weeks, not quarters. Cost has two parts: our work to design and deploy it, and your own infrastructure bill — which the budgets and usage tracing we build are there to keep predictable. We estimate each feature upfront and do our best to stay within 20% of the estimate.
Most companies end up with both. The gateway becomes the single entry point either way: sensitive workloads route to the private model, everything else keeps going where it goes today — and you finally get one place where usage, cost, and access rules are visible.
Tell us what your team wants to do with AI and what your data rules are. We'll come back with the deployment that fits — and say so if you don't need one.