Rook, My Agent Workstation

Most days, working on my homelab starts on my Mac with ssh rook and promptly stops being work on my Mac.

Rook is the machine where Claude Code, Codex, and opencode run. The infrastructure repositories are checked out there alongside the wiki and application code, so an agent can read the documentation, inspect the current configuration, query a service API, and find the repository that owns whatever it has been asked to change.

That proximity is useful. It is also why I don’t treat rook as a disposable AI sandbox. It is an administration workstation with enough access to be dangerous, so the route from an agent’s change to a live change matters more than the list of models installed on it.

The normal result is a branch and a pull request on my Forgejo instance. CI runs the checks and renders the OpenTofu plan. I read the diff and the plan before merging. The agent prepares the change; it does not merge it.

From wikibot to workstation

Rook started as wikibot, an 8 GB LXC whose main job was to hold my homelab wiki and give Claude Code somewhere close to the systems it documented. The name stopped fitting once the blog, Surgite, and other development repositories moved onto it. By then it was doing about a third wiki work and two-thirds everything else, so I renamed it after another corvid and admitted it was a workstation.

Docker forced the hardware question. Surgite ships a Compose stack with PostgreSQL 16, and its Alembic migrations need to be tested against PostgreSQL rather than a convenient SQLite substitute. Installing Docker on an 8 GB root disk that was already 59% full, with 3.1 GB free, was asking for a fairly boring outage.

I changed the declared disk to 40 GB. Rook now looks like this:

GuestProxmox LXC 206 on corvus
OSDebian 13, unprivileged
CPU4 virtual cores
Memory8 GiB RAM, 4 GiB swap
Disk40 GB
LXC flagsnesting=1, keyctl disabled
Docker storageoverlayfs through containerd

The container is unprivileged, so root inside it is mapped away from root on the Proxmox host. That is a useful boundary, but it does not make the contents of the guest unimportant. Rook holds source checkouts, agent sessions, service credentials, and SSH keys. Root inside rook can still read or alter all of those.

What lives on it

The agent CLIs are only the visible part of the machine. The rest of the toolchain is what lets them finish a job instead of writing a plausible suggestion and stopping there.

Rook has Node.js 22, Python through uv, Docker CE with the Compose plugin, git, go-task, ripgrep, fd, jq, GitHub CLI, and OpenTofu. Go, Rust, and Swift are absent because nothing on the box currently needs them. I would rather add a tool when a project asks for it than maintain a workstation built for every hypothetical project.

Docker inside an unprivileged LXC is one of those things that either works quietly or sends you into several hours of storage-driver archaeology. On rook it runs with nesting=1 and no keyctl flag. Docker 29 uses the containerd image store and reports overlayfs rather than the older overlay2 name. The Ansible check accepts either and fails if Docker falls back to vfs, because vfs copies complete layers and would chew through the disk I had just enlarged.

The repository layout matters just as much. The OpenTofu configuration, Ansible playbooks, homelab wiki, blog, and application repositories are all local. When an alert says a service is unhealthy, the agent can move from the runbook to the live API response and then to the configuration that owns the service without me pasting five files into a chat window.

There are two ways to inspect the live environment. Proxmox, Forgejo, and Komodo have API credentials templated onto rook by Ansible. SSH to the cluster nodes also permits a set of status and configuration commands. That means the agent can check whether the documentation still matches reality before changing a file. I do not want an agent fixing a stale wiki page while the service is failing for a completely different reason.

The same access is why repository ownership stays explicit. An API response is evidence about what exists now. It is not automatically the source of truth. If OpenTofu owns a guest’s memory, the fix goes into hosts.yaml; if Ansible owns an SSH setting, the fix goes into the playbook. A manual command can be useful for diagnosis, but it is not the finished change.

The flow from there is deliberately split:

%%{init: {"flowchart": {"useMaxWidth": false, "nodeSpacing": 16, "rankSpacing": 32, "diagramPadding": 4}}}%%
flowchart TB
    subgraph prepare["1. Prepare"]
        direction LR
        mac["Mac\nssh rook"] --> rook["rook\nagents + repositories"] --> pr["related Forgejo\npull requests"]
    end
    subgraph review["2. Review"]
        direction LR
        ci["Actions\nplan + tests"] --> human["human review\nand merge"]
    end
    subgraph configure["3. Apply and configure"]
        direction LR
        apply["OpenTofu apply\non main"] --> ansible["manual dispatch\nAnsible configures guest"]
    end
    prepare --> review --> configure

The repositories are separate, so a change that spans infrastructure, configuration, and documentation normally means related pull requests rather than one atomic pull request. The agent can prepare them as one piece of work and link them, but Git cannot make three repositories transactional.

OpenTofu owns the outside

OpenTofu owns whether a guest exists and what Proxmox should give it: node placement, VMID, IP address, cores, memory, swap, disk, network settings, and container flags. The source of truth is hosts.yaml in the infrastructure repository. Rook’s entry resolves to LXC 206 at 192.168.0.206, on corvus, with the resources in the table above.

A pull request runs tofu fmt, tofu validate, and tofu plan on a dedicated Forgejo Actions runner. The plan is read-only. Merging to main starts the separate tofu apply workflow, using state stored in MinIO with native locking.

The pull-request plan is speculative. My apply workflow generates a fresh plan from main rather than applying a saved plan artifact from the pull request. Usually those are the same, but they can differ if the live infrastructure or another branch changes in between. The reviewed plan is therefore a warning and review tool, not a cryptographic promise about the later apply. The apply job still logs the actions it takes, and only a merge starts it.

This is an important distinction in my setup: the agent on rook has OpenTofu installed and can inspect or validate a change, but the routine apply happens on the CI runner after merge. The credentials used by that workflow live in Forgejo Actions secrets, not in the repository.

OpenTofu does not quite own all of rook’s outside. Tailscale needs a TUN device, so the Proxmox container config has two raw lines:

lxc.cgroup2.devices.allow: c 10:200 rwm
lxc.mount.entry: /dev/net/tun dev/net/tun none bind,create=file

Tailscale uses /dev/net/tun as its Linux tunnel device. Those two lines are not represented in my OpenTofu resource. Recreating the container from the declaration alone would therefore bring rook back without Tailscale, which would also break its .lan name resolution through MagicDNS.

That is configuration drift, even though it is documented drift. I kept the rename as a state moved operation rather than destroying and recreating the LXC partly because I did not want that omission to become a surprise rebuild problem. The proper fix is to bring the device mapping under declared management or build an explicit reprovisioning step around it. For now the wiki calls it out in large letters.

Ansible owns the inside

Ansible takes over once the guest is reachable. Its rook playbook installs the packages and agent CLIs, creates the account, hardens SSH, configures Tailscale, clones the repositories, installs the Forgejo deploy key, and writes the service credentials the agents use at runtime.

The credentials are templated from Ansible Vault into the agent configuration. The values never belong in git, shell history, or this post. The useful point is that they are reproducible. Rebuilding rook no longer leaves me with a working machine that cannot query Proxmox, Forgejo, or Komodo until I remember which tokens I placed by hand six months earlier.

Provisioning a new guest is intentionally a two-stage handoff. OpenTofu creates it and seeds the SSH keys needed for first contact. After that apply has completed, I dispatch the run-playbook workflow in Forgejo and Ansible configures the operating system. I could chain the playbook directly onto the apply, but two separate runs are easier to inspect and recover. The extra click has not become annoying enough to earn more automation.

An agent adding a service therefore has a fairly concrete job. It adds the guest to hosts.yaml, writes or updates the matching playbook, adjusts the wiki, runs the repository checks, and opens the related pull requests. I review the proposed Proxmox shape before merge. OpenTofu creates the empty guest, then I trigger the Ansible run that turns it into the intended service.

Keeping 40 GB from filling again

Moving from 8 GB to 40 GB removed the immediate pressure. It did not make Docker stop accumulating old layers and build cache.

Rook has a daily workstation-cleanup.timer which calls a small shell script. The part that does most of the work is:

docker image prune -af --filter until=168h
docker builder prune -af --filter until=168h

apt-get clean
uv cache prune
npm cache clean --force

The seven-day filter means an image used for current work is not thrown away immediately. More importantly, the script never calls docker volume prune and never passes a volume-pruning flag. Surgite’s PostgreSQL data lives in a development volume. Deleting an old image costs a pull; deleting that volume costs data and a much worse afternoon.

The distinction is small enough to disappear inside a generic docker system prune habit, which is why rook has its own cleanup script rather than borrowing the more aggressive one from my disposable CI runner. At the moment the 40 GB filesystem is 33% used, with about 26 GB free. The timer is there to keep that boring.

The access boundary is not finished

SSH into rook is key-only. Its separate cluster key is accepted on the Proxmox nodes only when the connection comes from rook’s address:

from="192.168.0.206" ssh-ed25519 AAAA... rook@rook

That stops a copied key being used directly from another machine. It does not reduce what the key can do after a process on rook uses it.

There are two privilege edges here. Inside rook, nick currently has passwordless sudo for every command. Because the LXC is unprivileged, that is not the same as root on corvus, but it is complete control of the workstation and everything stored on it.

The cluster-side edge is worse than an old “read-only” description in my wiki implied. On corvus, the key can use passwordless pct set * and pct exec *. pct exec is root command execution inside any LXC on that node, and pct set can change their configuration. The source-address restriction narrows where the key works; it does not turn those commands into read-only access.

I am working on replacing that with narrower, task-specific access. The likely split is read-only inspection for agents, with mutation credentials kept on the Forgejo runner and exposed only through reviewed workflows. I have not finished that design, so I am not going to describe the current setup as least privilege.

Forgejo review still matters, but it solves a different problem. It gives me a legible diff, runs the checks, records the plan, and requires my merge before the normal apply pipeline starts. It does not sandbox an agent that already holds an operational credential. Until the cluster permissions are narrowed, the boundary is partly technical and partly a rule about how I use the workstation.

That is why rook is treated as an administration machine rather than a throwaway prompt box. The agents do the repository archaeology and assemble the OpenTofu, Ansible, and documentation changes. I get the related pull requests, the CI output, and the final decision about what reaches main. Then the automation takes over.

Discussion