Files
pi-swarm/docs/PLAN.md
2026-09-19 10:37:53 -07:00

6.6 KiB
Raw Blame History

Rebuild plan & decisions log

Where we started (2026-09-19)

  • Old cluster: Docker Swarm (manager planck001) with Portainer, Traefik, and what appears to be CapRover remnants. Broken: workers stuck in "pending" because the manager's IP changed (workers dial stale 192.168.1.173:2377).
  • Decision: wipe everything and rebuild clean rather than repair.

Hardware

  • 20× Raspberry Pi 4 Model B Rev 1.4, 8GB RAM, 4 cores
  • Boot: 238.5GB USB SSD on every node (no SD cards installed)
  • Bootloader EEPROM: 2023/01/11 on 19 nodes; planck019 has Feb 2021 (update it)
  • All nodes healthy: temps 4352°C, disks ~4% used

Network

  • Ubiquiti UDM Pro, 2 switches, all nodes on the same router
  • IPs currently drift across 192.168.0/1/2.x subnets via DHCP — planck001 sits on 192.168.30.x (different subnet, cause unknown, doesn't matter: we're rebuilding)
  • Decision: dedicated static addressing for the cluster via UniFi DHCP reservations (MAC table in docs/MACS.txt)

Storage

  • Synology NAS on same LAN — shared storage target (NFS) for cluster workloads

Rebuild architecture (decided 2026-09-19)

Fully unattended network reimage, no per-node physical access:

  • Synology (192.168.1.157) hosts three services (all auto-start, see configs/):
    • planck-tftp.service — custom python TFTP server on :69 serving /volume1/plancknetboot/tftp/ (firmware, kernel, installer initramfs)
    • planck-netboot.service — python HTTP helper on :8000 (GET rootfs.tar template, GET /events/<name> progress log, POST = upload)
    • NFS read-only export of the share for future use
  • Installer initramfs (configs/installer-init.sh): busybox + sfdisk + e2fsprogs. Identifies the node by CPU serial (map in configs/hostmap), wipes /dev/sda, downloads the Debian 13 Raspberry Pi OS Lite template over HTTP, installs it, sets hostname/ssh keys/user, POSTs pieeprom-revert.upd/.sig to tftp/<serial>/, reboots.
  • Bootloader flow: nodes' EEPROM is flashed (via staged pieeprom.upd from the old OS) to install-mode: BOOT_ORDER=0xf142 (network first), TFTP_IP=192.168.1.157, TFTP_PREFIX=1 (per-serial dirs), ENABLE_SELF_UPDATE=1. After install, the TFTP-served bootloader_update=1 config + <serial>/pieeprom.upd reverts the node to BOOT_ORDER=0xf14 (USB first). Fresh OS boots from SSD; firstboot service tidies up.
  • Reinstall a node anytime: stage install-mode EEPROM again + reboot (delete tftp/<serial>/pieeprom.* on the Synology first if it exists).
  • Node password: see configs/node-password.txt (gitignored). SSH is key-based; password exists for console/sudo recovery.
  • UniFi MFA blocks API access; DHCP reservations deferred — mDNS (planck0NN.local) works across subnets today.

Current debug state (paused here, swap to stronger model)

First live test on planck020 (EEPROM install-mode staged, works — confirmed via vcgencmd bootloader_config showing BOOT_ORDER=0xf142, TFTP_IP, TFTP_PREFIX=1):

  • TFTP chain works: bootloader + firmware fetch firmware, config.txt, cmdline.txt, kernel8.img, and the initramfs (initramfs8 via auto_initramfs=1 — the explicit initramfs <file> followkernel syntax did NOT work on this bootloader; use auto_initramfs + file named initramfs8). All transfers show DONE on the Synology TFTP server (journalctl -u planck-tftp). "STALLED" log lines are harmless — they're duplicate/speculative client sessions that never get ACKed.
  • Custom TFTP server gotcha: firmware requires tsize in the OACK — without it transfers stall. Fixed in configs/planck-tftp-server.py (deployed version on Synology is current).
  • Kernel boots (node pings ~25s after reboot with a DHCP-assigned IP — init reached DHCP), but no HTTP requests ever arrive at the planck-netboot helper (no /events/ GETs, no template GET, journal always empty). Then the node reboots into the OLD OS (~2-3 min cycle). So initramfs /init is either crashing before/inside wget, or HTTP egress from the initramfs fails silently, and the test-mode poweroff -f acts like a reboot (node returns on old OS — actually likely USB fallback after a failed netboot pass).
  • planck020 EEPROM is still in install-mode; every reboot retries netboot then falls back to old OS. To restore it: stage revert via rpi-eeprom-config --config revert.boot.conf flow from the old OS (see EEPROM configs in git history / regenerate from /lib/firmware/raspberrypi/bootloader/stable/pieeprom-2023-01-11.bin).

Next steps for whoever continues

  1. Get console visibility on the netbooted node: add per-log-line HTTP reporting in configs/installer-init.sh (each log() also does wget -q -O /dev/null "$HTTP/events/<serial>-<msg>" once DHCP is up), plus flush a tmpfs log file after DHCP. Alternatively boot with console=serial0,115200 and attach a USB-serial cable to planck020.
  2. Suspects for the silent init death: busybox wget HTTP behavior in initramfs (try wget -O- verbose, test with numeric IP + port 8000 — verified reachable from other nodes), or udhcpc script path issues. Consider testing initramfs content standalone (gunzip + cpio listing).
  3. Template tarball + all services are already on the Synology (/volume1/plancknetboot/): rootfs.tar (sha256 e37837bee03ae2b3...), tftp/ (firmware, initramfs8, auto_initramfs config), HTTP helper :8000, TFTP :69. Credentials for Synology SSH are session-only (not stored).
  4. Once one node completes end-to-end, the rest is a loop: stage install-mode EEPROM per node + reboot (script this in scripts/).

Phases

  • Phase 0 — Access + inventory (SSH keys on all 20, docs/)
  • Phase 1 — Network design: static IPs / reservations in UniFi
  • Phase 2 — Reimage all 20 nodes with fresh Raspberry Pi OS
  • Phase 3 — Cluster runtime install (candidates: k3s vs Docker Swarm)
  • Phase 4 — Storage integration with Synology (NFS CSI)
  • Phase 5 — Apps: monitoring, Portainer/UI, whatever the lab is for

Open decisions

  1. Reimage method — options: a. PXE/network boot: EEPROM boot-order set to netboot; DHCP next-server points at a TFTP/HTTP installer. Zero-touch per node, reusable forever. Installer server can run on the Synology (Docker) or a laptop. b. SD-card bootstrap installer: flash one SD image, insert per node, it wipes + installs the USB SSD, removes itself. Simple, physical walk. c. Pull SSDs, image on a PC with Raspberry Pi Imager ×20. Most manual.
  2. Cluster runtime — k3s (lighter, k8s-compatible, we manage it with automation) vs Docker Swarm + Portainer (simplest).