Running a Homelab with Claude Code: What an AI Operator Is Actually Good For

Share
Cover illustration

Introduction

For the past few months I have been managing my homelab and a small fleet of public VPS servers with Claude Code, Anthropic's command line coding agent, running on a Windows management box. It is not a chatbot I paste logs into. It has a shell, SSH access to my servers, and the ability to read and write files, so it does the actual work: hardening hosts, building monitoring, configuring alerts, automating patching, debugging outages and writing the documentation afterwards.

This article is an honest account of how that works in practice, what it did well, and the habits that made it safe enough to trust with production infrastructure. Every host name, address and credential in this guide is a placeholder.

By the end you will know:

  • How to set Claude Code up as an operator for a homelab, rather than as a code assistant
  • The working pattern that makes changes safe: backup, change, verify, document
  • What it built for me in a few weeks of evenings, and what that would normally cost in time
  • The guardrails and habits that keep an AI operator safe on production systems

What I actually asked it to do

I did not start with a grand plan. I started with one small question: here is a fresh VPS, what is wrong with it? That turned into this list over the following weeks.

  • Audit and hardening. A full inspection of each server, followed by a prioritised fix list. The headline finding on the first VPS was that Docker published ports bypass the host firewall entirely, so admin interfaces I thought were private were reachable by the whole internet. Internet scanners were already showing up in the logs.
  • A real firewall model. Host firewall zones restricted to my home address, plus an allow list in the DOCKER-USER chain for container ports, with a systemd unit to reapply it at boot.
  • Standardised container deployments. A written house style for compose files, then every existing stack rebuilt to it: pinned image digests, no-new-privileges, dropped capabilities, memory and process limits, log rotation, healthchecks and secrets kept out of the compose file.
  • CIS benchmark hardening. The CIS Level 1 Server profile applied with the vendor build kit, with the results triaged by hand. Most failures were false negatives, and one fix broke every container on the box (more on that below).
  • Intrusion detection. CrowdSec for SSH and web attacks, and AIDE for file integrity.
  • Monitoring and alerting. A health check on a five minute systemd timer that checks containers, the firewall rules, certificates, disk, pending reboots and whether patching has lapsed. It sends mail through my existing mail relay with a consistent subject format, so severity is visible in the inbox before I open anything.
  • Patching. Unattended security updates widened to the regular updates pocket, a weekly patch run, and mail on failure or on change.
  • Backups before every change. A configuration backup pulled back to the management box before anything risky.
  • Everything else. Reverse proxy and tunnel setup, DNS, a home automation dashboard, a GPU server, scripted firewall rule changes through an API, even this wiki.

None of this is exotic. It is the checklist any sysadmin knows. The difference is that one person with an evening free can now get through it in a week instead of a quarter, and every step is written down.

The setup

The management box is a Windows server I already had. Claude Code runs in a project folder that holds three kinds of file, and that folder is the whole secret.

A project instructions file. Claude Code reads a CLAUDE.md at the start of every session. Mine says what the folder is for and how I like to work.

A persistent memory directory. Facts that must survive between sessions are saved as small files with an index: how each server is reached, what was built, what was learned, and what is still pending. The agent has no memory of yesterday unless it is written down, so this is what turns a stateless tool into a colleague who knows the estate.

Running logs per server. One markdown file per host, with a numbered change log: what was changed, why, how it was verified. Plus a standards document for the compose style. These double as my own documentation, and as the source for articles like this one.

For access, I use plain SSH with host keys pinned on first connect, read from a fingerprint I verified myself before any credential was used. Credentials live in a text file that the scripts read, never on a command line. Where a command needs elevated rights, sudo is used with the password supplied on standard input, not as an argument, because command arguments end up in logs.

The pattern that makes it safe

Every change follows the same four steps, and I insist on them.

  1. Back up first. A copy of the file or a configuration archive pulled off the server. There is no snapshot facility on my VPS provider, so this is the only undo button.
  2. Make one change. Not a batch. If something breaks, the cause is obvious.
  3. Verify with evidence. Not "it should work". Fetch the page, check the port from outside, read the counters, run a throwaway container. The firewall rules, for example, were proven by deliberately pointing the allow list at a fake address, confirming the admin page became unreachable, then restoring it.
  4. Write it down. The numbered change log entry is part of the change, not an afterthought.

The agent is genuinely good at steps 3 and 4. It does not get bored of verifying, and it writes the log entry while the detail is still in context. That discipline is worth more than the speed.

What it was brilliant at

Systematic debugging with a bias for evidence. When every container on a freshly hardened host refused to start with fork/exec /proc/self/fd/6: permission denied, it spent a while on plausible but wrong theories. It then traced the container runtime's system calls, read the actual failing call, and found that the hardening kit had switched the runtime's AppArmor profile into complain mode, which this version of runc cannot cope with. The fix took one line once the cause was known.

Comparing against a known good. A tunnel client connected perfectly but never received its configuration. After DNS, firewall, routing and keys had all been ruled out, I asked for a line by line comparison against my other server, where the same setup worked. One line differed: the tunnel endpoint was set to a hostname proxied through Cloudflare, which does not carry the raw UDP that WireGuard needs. A diff beat hours of theory.

Tireless reading of boring material. Vendor documentation, entrypoint scripts, benchmark reports with hundreds of lines. Before deploying an unfamiliar container image it reads the image's entrypoint and layer history, checks that the vendor is real, and only then runs it.

Guardrails that actually mattered

Claude Code asks permission before risky actions, and in the mode I use a separate classifier can block them outright. It will not use stored credentials or loosen a security profile unless I have explicitly authorised that specific task. When it is blocked, the right behaviour is to hand the command to me, not to find a way around.

Honest cost and benefit

Benefit. Work that would have taken me weeks of evenings, and that I would have done less thoroughly, now gets done properly. The environment is documented better than it has ever been, because documenting is part of the loop. The alerting has already caught real problems before I noticed them.

Cost. You must understand what it is doing. If you cannot judge whether a firewall rule is right, you cannot supervise the agent writing it. It makes confident mistakes, and the confident tone does not correlate with correctness. I read every destructive command before approving it.

Not a replacement for judgement. It is excellent labour with good instincts, working under the supervision of someone who owns the outcome.

A starting checklist

  • Give it a dedicated project folder, and let it write the instructions file and notes as you go.
  • Pin SSH host keys yourself, before it connects the first time.
  • Keep secrets in files read by scripts, never in commands or chat.
  • Insist on backup, one change, verification, and a log entry. Every time.
  • Keep out of band access (provider console) open for anything touching SSH, firewall or networking.
  • Keep secrets out of commands and chat, and rotate anything that is ever exposed.
  • Read the destructive commands.
  • After any change that matters, test the thing that matters: run a container, load the page, send the alert.

Closing thoughts

I expected a faster way to type commands. What I got was something closer to a diligent junior colleague who never tires of verification, writes everything down, and occasionally needs stopping before they do something clever. My infrastructure is more secure, better monitored and better documented than it has ever been, and I understand it better because I had to review every step. If you run a homelab with more than a couple of servers, and you are willing to supervise properly, it is well worth trying.

Read more