Search

· AI Engineering Â· 11 min read

Local-First MCP: Air-Gapped Tool Calling with llama.cpp

Local-first MCP with llama.cpp makes offline tool-calling agents viable, and a 32k context budget forces tool discovery at runtime, not upfront.

Featured image for: Local-First MCP: Air-Gapped Tool Calling with llama.cpp

I’ve sat in more than one review where a team had a genuinely useful agent prototype, and a security lead asked one question that ended the meeting: “Where does the prompt go?” The honest answer was a third-party API. The data was clinical notes, or trading positions, or a network inventory. The prototype was shelved, and the team went back to writing scripts by hand.

That gap is what a local-first Model Context Protocol (MCP) setup is meant to close. llama.cpp, the C/C++ inference engine most people already use to run open-weights models on their own hardware, now ships an experimental MCP client in its server. The model, the tool runner and the data can all sit inside one network boundary with no outbound calls. For a certain class of workload, that turns “not allowed” into “allowed, with caveats.”

The caveats are the interesting part. A local model usually runs with a small context window, and 32k tokens is a common and realistic budget on modest hardware. Fit a dozen MCP servers’ worth of tool schemas into that and you’ve spent a third of your memory before the user says hello. So the real design question is less “can I call tools offline” and more “how do I let the model find tools without carrying all of them.”

Key Takeaways
  • llama.cpp's server can act as an MCP client against stdio and HTTP servers, which makes fully offline, tool-using agents practical for air-gapped or regulated data.
  • With roughly 32k tokens of context, loading every tool schema upfront crowds out the task itself. Discover tools at runtime and load only what the step needs.
  • Local-first buys data residency and control. It does not buy the capability ceiling or vendor support of hosted frontier models, so pick by workload, not ideology.
  • For leaders: the business case is unlocking use cases that compliance currently blocks, not saving money on tokens.

What llama.cpp actually gives you

Be careful here, because this area moves quickly and the feature is flagged experimental. From the official llama-server documentation in the llama.cpp repository, the pieces that matter are:

  • --mcp-servers-config PATH loads MCP server definitions from a file in the Cursor-compatible format.
  • --mcp-servers-json JSON takes the same definitions inline.
  • --jinja enables the Jinja chat-template engine, which is what lets a model’s own tool-calling template drive function calls. The docs list it as enabled by default.
  • --webui-mcp-proxy turns on an experimental CORS proxy for the web UI’s MCP support. It is off by default, and the documentation warns against enabling it in untrusted environments.

I haven’t pinned a release number here on purpose. Check llama-server --help on the build you actually deploy, because flag names and defaults in an experimental feature can change between builds. Treat anything in this article as a pattern to verify against your binary, not a spec.

What the integration means in practice: the same process that serves your model can speak MCP to local tool servers, run the agent loop (model asks for a tool, the tool runs, the result goes back in), and present it through the built-in web UI or the CLI. There’s no cloud middleware in the path. Because MCP supports local stdio servers, a tool can be a plain subprocess on the same machine, with no listening port at all.

The air gap, drawn

An air-gapped deployment is easy to describe and easy to get wrong. The model weights have to arrive by some controlled route (a vetted transfer, a signed artifact), and after that, nothing leaves.

Supplemental Explainer

Every arrow in that diagram stays inside the box. That is the whole value proposition, and it’s also the whole threat model: you now own patching, tool permissions and audit logging for every component inside it. A hosted API would have handed you some of that for free.

Why 32k forces a different design

Here is the arithmetic that changes how you build. An MCP tool definition is a name, a description and a JSON Schema for the arguments. A simple tool might run 100 to 200 tokens. A real one, with enums, nested objects and the long descriptions people write so the model behaves, can run 500 or more. These are my rough estimates from tools I’ve worked with, not a benchmark, and your tokenizer will shift them.

Take a plausible internal setup: five MCP servers exposing about 12 tools each. That’s 60 tools. At 300 tokens apiece you’re at 18,000 tokens of schema. In a 32k window, that leaves about 14,000 for the system prompt, the conversation, retrieved documents and every tool result that comes back. A single query returning a few hundred rows of SQL output can blow through that on its own.

It gets worse in a way that’s less obvious. Small and mid-sized local models get noticeably less reliable at picking the right function as the menu grows. You pay in context and in accuracy at the same time. Hosted frontier models with very large windows can often tolerate a crowded tool list; a quantized model on a single workstation GPU usually can’t.

So the naive pattern, where the client loads every schema at startup and hands them all to the model, is the wrong default for local-first. The better one is discovery at runtime.

The discovery pattern

Instead of exposing 60 tools, expose a very small fixed surface and let the model look the rest up:

  1. A catalog tool. list_capabilities returns one line per tool group: a name and a short purpose. For 60 tools in 5 groups, that’s five lines, maybe 150 tokens.
  2. A describe tool. describe_tools(group) returns the full schemas for one group only, on request.
  3. Execution through the normal path. The model calls the real tool once it has seen the real schema.
  4. Eviction. When the model moves on, the schemas for the finished group drop out of the working context instead of piling up.

The cost drops from “all schemas, always” to “a short index plus the one group in play.” In the example above, that’s roughly 150 tokens of index plus about 3,600 for a 12-tool group at the same assumed 300 per tool, so around 4,000 instead of 18,000. Again, that’s my estimate under stated assumptions. Measure your own.

A thin gateway makes this practical. It is itself a small MCP server that fronts the others, holds the real schemas in memory, and answers the catalog and describe calls. The model sees one server. You keep the plumbing flexible behind it.

{
  "mcpServers": {
    "gateway": {
      "command": "./tool-gateway",
      "args": ["--registry", "/etc/agent/tools.d"]
    }
  }
}

That follows the Cursor-style mcpServers shape the llama.cpp flags describe, and you’d point --mcp-servers-config at the file. The tool-gateway binary is your code, not something that ships with llama.cpp. The point is where the logic lives: in a component you can test, log and lock down, not in the prompt.

A worked example: incident triage in a closed network

Picture a utilities operator with a control-room network that has no internet path. Engineers want to ask, in plain language, “which substations had breaker alarms in the last six hours, and what did the maintenance log say about the last time?”

The pieces: a quantized open-weights model served by llama-server on a workstation; a read-only SQL tool over the alarm historian; a document search tool over scanned maintenance logs; and the gateway above.

A turn goes like this. The model sees only the index of capability groups. It asks for the alarms group, gets the schema for two tools, and calls the query tool. The result comes back trimmed to the first 50 rows with a count, because the gateway enforces a row cap. The model then asks for the documents group, searches, and answers with citations to file paths. Peak schema load in context is one group at a time. Total prompt stays well inside the budget.

Two details made it work. First, tool results are capped and summarized by the tool, not left for the model to cope with. Second, the SQL tool is read-only at the database permission level, not just by instruction. A local model can be talked into things by text inside a maintenance log, and prompt injection doesn’t care whether you’re air-gapped.

What can go wrong

I’d rather you hear the failure modes from me than from your pager.

The feature is experimental. The flags say so. Don’t build a compliance story on behaviour you haven’t pinned to a specific build and tested.

Tool poisoning still applies. An offline network removes the internet from the picture, not malicious content. A hostile string in a document or a compromised internal tool description can steer the model. Treat tool descriptions as untrusted input, review them like code, and keep write-capable tools behind a human confirmation.

Template mismatch. Tool calling through --jinja depends on the model’s chat template actually supporting tools. Some models ship templates that don’t, or that format calls in ways the parser doesn’t expect. Test the exact model file you plan to ship, with the exact quantization.

Small models stumble on multi-step plans. Choosing among five tools is fine. Chaining eight calls with dependencies is where a smaller local model drifts. Keep workflows short, or encode the sequence in a tool.

Operational burden. Model updates, GPU drivers, log shipping and access control become your job. That’s a real staffing cost, and it belongs in the business case.

Where hosted frontier models are still the right call

This isn’t an argument that local beats hosted. It’s an argument that local is now a credible option for one particular constraint.

If your data can legally and contractually go to a hosted provider, the strongest hosted models will generally give you a higher capability ceiling on hard reasoning, longer and more reliable multi-step tool use, much larger context windows that make the upfront-schema problem less pressing, and vendor support with service commitments. Those are real advantages, and for many teams they outweigh data residency, which they weren’t worried about in the first place.

The honest decision rule, as I’d put it to a steering group:

SituationBetter fit
Data cannot leave the network (air-gapped, classified, strict residency)Local-first with llama.cpp and MCP
Tasks are narrow, tools are few, workflows are shortLocal-first is often enough
Tasks need deep reasoning or long autonomous chainsHosted frontier model
You need support guarantees and no appetite for running GPUsHosted frontier model
Mixed: sensitive retrieval, hard reasoningHybrid, with local tools redacting or summarizing before anything leaves

The hybrid row deserves attention. Some teams will run the local model as the gatekeeper that touches raw data, and send only reviewed, minimized output to a hosted model when policy allows. That’s a legitimate architecture, though it reopens the compliance question, so get sign-off first.

The executive view

If you’re funding this, the pitch isn’t cheaper tokens. At low volume, a hosted API is often cheaper than buying and operating GPU hardware, and I wouldn’t claim otherwise without your numbers. The pitch is access: use cases that sit in a compliance backlog because nobody could answer the data-residency question.

Ask three things before approving a pilot. Which specific workflow is blocked today, and what’s its value if unblocked? Who owns patching and monitoring of the inference host and tool servers? And what’s the exit if a later hosted offering meets your residency terms? A pilot with a clear answer to each is a small, reversible bet. One without them is a science project.

For the engineers, the checklist is shorter. Pin the llama.cpp build and the model file. Keep the exposed tool surface tiny and discover the rest at runtime. Cap and summarize tool output in the tool. Enforce permissions in the systems the tools touch, not in prompts. And test the failure cases on purpose, including a poisoned document, before anyone calls it done.

The protocol itself is documented at modelcontextprotocol.io, and the engine lives in the llama.cpp repository. Read the server README for your build before you trust any flag I’ve quoted. In a closed network, the constraint you design around isn’t bandwidth or even model quality. It’s the small, finite window the model has to think in, and the discipline is deciding what earns a place in it.

Enjoying this insight?

Join the distribution list to get deep dives on AI transitions and agency economics directly in your inbox. No spam, ever.

Back to Blog

Related Posts

View All Posts »
A2UI: The Interface is Now a Variable

A2UI: The Interface is Now a Variable

We’re moving past static dashboards and iframes. We explore the A2UI protocol, how models choose their own blueprints, and the future of morphing interfaces.

Structured Agent Memory vs Vector Search

Structured Agent Memory vs Vector Search

Pinecone's Nexus Engine compiles business context into structured knowledge graphs for agents. Nemotron 3 Embed tops RTEB. Vector search alone is now insufficient. Here is the memory architecture...