Local LLM

Local LLM Deployment Team Handoff Runbook 2026

A practical handoff runbook for local LLM deployments: model weights, Ollama configs, endpoint handoff, SSH patterns, rollback, and team continuity.

FreeLast tested: 2026-06-19Audience: Engineering leads, platform engineers, AI-infra teams

Why handoff breaks local LLM deployments

Local LLM deployments do not fail because the model is bad. They fail because the next engineer inherits a Mac mini with Ollama installed, a model pulled six months ago, an SSH tunnel that no longer works, and a README that says "it just runs." In practice, that is not runnable. It is a guessing game.

The difference between a working local deployment and a shelved project is usually documentation discipline: exact model tags, quant formats, GPU offload settings, endpoint URLs, fallback behavior, and who to call when the model server is down. This runbook is meant to shrink that gap.

Pre-flight handoff checklist

Before the outgoing engineer logs off, confirm the receiving engineer can answer four questions without asking for help: which model is running, how to reach it, how to restart it, and how to verify it is healthy.

ArtifactMust includeWhy it matters
Model manifestExact tags, quant types, size on disk, export pathPrevents "I pulled the wrong GGUF"
Service configOllama env, Modelfile, GPU layers, context windowReproducible startup on a different machine
Endpoint mapHost, port, OpenAI-compatible path, auth, Tailscale/public IPConsumers need one stable URL
RunbookStart, stop, logs, health check, rollback, on-callReduces mean-time-to-recovery

Model weights and provenance

Treat model weights like dependency locks. Record the exact source, tag, quantization, and license. For Ollama, that means the FROM line in the Modelfile plus the registry tag. For llama.cpp, it means the GGUF filename, quantization method, and whether it was converted from a safetensors release.

Also capture hardware assumptions. A 4-bit quant that fits on an M2 Ultra does not necessarily fit on an M2 Air. If the deployment depends on partial GPU offload, record the NUM_GPU layer count and expected RAM usage. Those numbers change the next person's hardware choice.

Environment and config handoff

Separate the model from its environment. Many teams hand off a model directory without handing off the runtime config, system prompt wrapper, and adapter files. The result is a model that runs but behaves differently. Record Ollama environment variables, prompt templates, and any LoRA or adapter mounts.

SettingExampleNotes
OLLAMA_HOST0.0.0.0Bind address for team access
OLLAMA_MAX_LOADED_MODELS1Prevent memory contention on small hosts
Context window8192Should match model training context
Prompt templateModelfile TEMPLATE blockVersion-control this separately

Endpoint handoff patterns

Local LLM deployments usually end up serving consumers through three transport shapes: localhost-only for the developer, Tailscale for the team, or a reverse proxy for external tooling. Pick one and document it.

For Tailscale deployments, the handoff pattern is straightforward: the new engineer joins the Tailscale network, resolves the hostname, and hits http://host:11434. For reverse proxies, hand off the TLS cert path, nginx config, and expected upstream health endpoint. Without those, the next person will proxy to a dead backend and spend hours debugging DNS instead of the model.

See local-llm-deployment-remote-ssh-patterns-2026.html for SSH tunneling and multi-host access patterns that often get missed during handoff.

Smoke test before sign-off

The outgoing engineer should leave a smoke test that the incoming engineer can run in under two minutes. It does not need to be comprehensive. It needs to prove the model answers, the endpoint responds, and the hardware offload is active.

curl -s http://localhost:11434/api/generate \ -d '{"model":"MODEL_TAG","prompt":"Say OK","stream":false}' \ | jq -r '.response'

If that command returns text, the deployment is alive. If it hangs or errors, the handoff is incomplete. Pair it with a health endpoint check and a quick GPU offload log line, and the next person has a minimum viable proof of life.

Rollback and continuity

Local models get updated. Quant formats change. Regressions happen. A handoff runbook should include the previous known-good model tag, how to switch back, and how to notify downstream consumers. Without rollback, an update becomes an outage with no recovery path.

For teams running multiple local models side by side, version the endpoint paths by model tag rather than by alias. That makes A/B testing and rollback a config change instead of a migration.

Related reading

If you are setting up local models for the first time, start with local-llm-deployment-guide-practical-smoke-test.html. For team rollout and adoption patterns, see ai-workflow-handoff-audit-engineering-teams.html. For broader AI infrastructure decisions, read local-llm-vs-cloud-api-cost-comparison.html.

Limits and notes

This runbook covers handoff for self-hosted local LLMs. It does not cover managed API keys, cloud model providers, or inference clusters. For those environments, the handoff checklist is similar in spirit but different in detail: rotate keys, document rate limits, and verify remaining quota before sign-off.