The great privacy pivot: why developers are moving RAG offline
Retrieval-Augmented Generation lets a model answer from your own documents. It is also the point where your most sensitive data gets sent to someone else's servers. More developers are choosing to run the whole thing locally instead.
Retrieval-Augmented Generation, or RAG, is one of the most useful patterns in applied AI. Instead of relying on what a model learned during training, you retrieve relevant passages from your own documents and pass them to the model as context.1 It is what powers "chat with your codebase" and "ask the company wiki." The catch is that doing RAG against a cloud API means sending those private documents to the API. For a growing number of teams, that is reason enough to keep the whole pipeline offline.
01What actually leaves your network
With a hosted model, every chunk of context you attach is sent to the provider: source code, financial records, customer data, internal plans. Major providers say API inputs are not used to train their models by default,2 which does matter. But it does not change the basic fact that the data has left your control and now sits, at least briefly, on systems you do not own and cannot inspect. For regulated data, or for a company's core source code, that transfer is the risk on its own.
This has already gone wrong in public. Samsung restricted employee use of generative-AI tools after staff pasted internal code into ChatGPT,3 and the average cost of a data breach keeps rising year over year.4 Data that never leaves is data that cannot leak.
02A fully offline RAG pipeline
Usefully, every part of a RAG pipeline has a mature local option, so the whole thing can run without a network connection. The pieces are:
- A local embedding model turns each document chunk into a vector — a list of numbers that represents its meaning — on your own hardware.
- A local vector database stores those vectors and finds the closest matches to a query. Chroma is a good starting point for developers,5 and Milvus is a better fit when you need to handle large volumes.6
- A local model (a Qwen3 or Gemma 4 running through the tools in part three) takes the retrieved chunks and your question and writes the answer.78
None of those steps needs the internet. The documents are embedded, stored, retrieved, and answered locally, so the data never crosses your network boundary. That is exactly what regulated and security-conscious teams are after.
03When the cost math changes
Privacy is the main reason, but cost is often what settles the argument. Cloud APIs charge per token, which is cheap for occasional use and expensive at sustained scale.9 Local inference has the opposite shape: a real upfront hardware cost, then very low cost per token afterwards. For an individual or a low-volume app, the cloud is almost always cheaper. The point where local starts to win is heavy, steady usage — on the order of millions of tokens a day, every day — where the per-token fees add up to more than amortised hardware would cost. If your usage is that high and that consistent, the economics start to favour running locally.
The cloud is cheapest when you use it lightly. Local hardware is cheapest when you use it heavily and want the data to stay put.
04The limitations
Offline RAG has trade-offs. A local model is smaller than a frontier cloud model, so it will be weaker on the hardest reasoning (part four is about matching tasks to that reality). You take on the maintenance, the embeddings pipeline, and the hardware. And a basic local setup can be slower than a well-tuned cloud endpoint. Even so, for answering questions over a bounded, private set of documents — your code, your files, your records — a local RAG stack is often the right choice, because it keeps the data with you.
05Where OcxlyDev lands
We keep RAG local when the documents are sensitive, and we are comfortable using the cloud when they are not. The approach is simple: put private, high-value documents behind a fully local pipeline (Chroma or Milvus plus a local model), and use cloud APIs for public or low-stakes data where convenience matters more. You get answers grounded in your own material, a bill that does not grow with every query, and the assurance that the data stayed on your own machines.