2.1 KiB
2.1 KiB
Rebuild plan & decisions log
Where we started (2026-09-19)
- Old cluster: Docker Swarm (manager planck001) with Portainer, Traefik, and
what appears to be CapRover remnants. Broken: workers stuck in "pending"
because the manager's IP changed (workers dial stale
192.168.1.173:2377). - Decision: wipe everything and rebuild clean rather than repair.
Hardware
- 20× Raspberry Pi 4 Model B Rev 1.4, 8GB RAM, 4 cores
- Boot: 238.5GB USB SSD on every node (no SD cards installed)
- Bootloader EEPROM: 2023/01/11 on 19 nodes; planck019 has Feb 2021 (update it)
- All nodes healthy: temps 43–52°C, disks ~4% used
Network
- Ubiquiti UDM Pro, 2 switches, all nodes on the same router
- IPs currently drift across 192.168.0/1/2.x subnets via DHCP — planck001 sits on 192.168.30.x (different subnet, cause unknown, doesn't matter: we're rebuilding)
- Decision: dedicated static addressing for the cluster via UniFi DHCP
reservations (MAC table in
docs/MACS.txt)
Storage
- Synology NAS on same LAN — shared storage target (NFS) for cluster workloads
Phases
- Phase 0 — Access + inventory (SSH keys on all 20, docs/)
- Phase 1 — Network design: static IPs / reservations in UniFi
- Phase 2 — Reimage all 20 nodes with fresh Raspberry Pi OS
- Phase 3 — Cluster runtime install (candidates: k3s vs Docker Swarm)
- Phase 4 — Storage integration with Synology (NFS CSI)
- Phase 5 — Apps: monitoring, Portainer/UI, whatever the lab is for
Open decisions
- Reimage method — options: a. PXE/network boot: EEPROM boot-order set to netboot; DHCP next-server points at a TFTP/HTTP installer. Zero-touch per node, reusable forever. Installer server can run on the Synology (Docker) or a laptop. b. SD-card bootstrap installer: flash one SD image, insert per node, it wipes + installs the USB SSD, removes itself. Simple, physical walk. c. Pull SSDs, image on a PC with Raspberry Pi Imager ×20. Most manual.
- Cluster runtime — k3s (lighter, k8s-compatible, we manage it with automation) vs Docker Swarm + Portainer (simplest).