Documentation audit: refresh README, resolve stale PLAN.md items, list open items
This commit is contained in:
35
docs/PLAN.md
35
docs/PLAN.md
@@ -70,7 +70,11 @@ Fully unattended network reimage, no per-node physical access:
|
||||
- NOTE: Pi firmware injects `cgroup_disable=memory` — the node cmdline.txt
|
||||
overrides with `cgroup_enable=cpuset cgroup_memory=1 cgroup_enable=memory`
|
||||
(already applied fleet-wide; required for k3s).
|
||||
- Domain for websites: carr.pub (DNS not yet pointed; pending)
|
||||
- Domain: carr.pub is LIVE — DNSimple apex A record (id 84723206) →
|
||||
50.46.44.67, UDM port-forwards 80/443 → planck002, Traefik ingress,
|
||||
cert-manager Let's Encrypt issuer `letsencrypt-prod`, DDNS CronJob
|
||||
`ddns-carr-pub` keeps DNS fresh. DNSimple token: k8s secret
|
||||
`dnsimple-token` in namespace `monitoring`
|
||||
- Next up: monitoring (light Prometheus+Grafana), Synology NFS storageclass,
|
||||
carr.pub DNS + cert-manager, load-test rig, agent harness namespaces
|
||||
|
||||
@@ -165,18 +169,21 @@ TFTP_PREFIX=1):
|
||||
- [x] Phase 0 — Access + inventory (SSH keys on all 20, docs/)
|
||||
- [ ] Phase 1 — Network design: static IPs / reservations in UniFi
|
||||
- [ ] Phase 2 — Reimage all 20 nodes with fresh Raspberry Pi OS
|
||||
- [ ] Phase 3 — Cluster runtime install (candidates: k3s vs Docker Swarm)
|
||||
- [ ] Phase 4 — Storage integration with Synology (NFS CSI)
|
||||
- [ ] Phase 5 — Apps: monitoring, Portainer/UI, whatever the lab is for
|
||||
- [x] Phase 3 — Cluster runtime install (k3s — decided and live)
|
||||
- [x] Phase 4 — Fixed IPs pinned on nodes (Synology NFS storage class still open)
|
||||
- [x] Phase 5 — Monitoring (Prometheus/Grafana), load-test rig, https://carr.pub
|
||||
|
||||
## Open decisions
|
||||
## Resolved decisions (were "open" on day one)
|
||||
|
||||
1. **Reimage method** — options:
|
||||
a. PXE/network boot: EEPROM boot-order set to netboot; DHCP next-server
|
||||
points at a TFTP/HTTP installer. Zero-touch per node, reusable forever.
|
||||
Installer server can run on the Synology (Docker) or a laptop.
|
||||
b. SD-card bootstrap installer: flash one SD image, insert per node, it
|
||||
wipes + installs the USB SSD, removes itself. Simple, physical walk.
|
||||
c. Pull SSDs, image on a PC with Raspberry Pi Imager ×20. Most manual.
|
||||
2. **Cluster runtime** — k3s (lighter, k8s-compatible, we manage it with
|
||||
automation) vs Docker Swarm + Portainer (simplest).
|
||||
1. **Reimage method** — pull-and-image on the laptop won (`flash-one.sh`).
|
||||
The network-boot experiment is retired but documented below as a
|
||||
postmortem; its Synology services linger (tftp disabled by design).
|
||||
2. **Cluster runtime** — k3s, 3 servers (002-004) + 17 workers.
|
||||
|
||||
## Open items (next sessions)
|
||||
|
||||
- Synology NFS StorageClass (survives node swaps; local-path is current default)
|
||||
- Grafana dashboards worth looking at (stack up, dashboards not built)
|
||||
- Rotate the Grafana admin password (`planck-lab-admin` is a committed
|
||||
placeholder in manifests/monitoring/grafana.yaml)
|
||||
- UniFi DHCP reservations for the 20 node IPs (see docs/IPS.md)
|
||||
|
||||
Reference in New Issue
Block a user