Compare commits

..

17 Commits

Author SHA1 Message Date
Adam Carr
258284a0c7 Backups complete: Glacier offsite tier covers plancknetboot 2026-09-20 17:03:57 -07:00
Adam Carr
e25d9add77 Backups: etcd snapshots (6h/14d), NAS btrfs daily snapshots, secrets off-laptop 2026-09-20 16:58:45 -07:00
Adam Carr
3df157a774 Portal: link to DSM directly 2026-09-20 16:46:22 -07:00
Adam Carr
a7a841988f Tailscale on Synology: NAS + File Station reachable remotely via tailnet 2026-09-20 16:11:36 -07:00
Adam Carr
e6d7b265f8 Tailscale on all 20 nodes: remote kubectl + SSH from anywhere (docs/IPS.md) 2026-09-20 14:42:46 -07:00
Adam Carr
b2878f43c6 act-runner: move registration token to k8s secret 2026-09-20 14:33:50 -07:00
Adam Carr
de0bb0ad7d Gitea + CI: git.carr.pub, NFS storage class, 3 runners 2026-09-20 14:33:39 -07:00
Adam Carr
c284205db4 k3s cluster dashboard: kube-state-metrics + kubelet/cAdvisor scraping, fix empty-label relabel bug 2026-09-20 13:30:57 -07:00
Adam Carr
e3f381f807 Node health dashboard + fix node-exporter scrape config (role:pod, instance=node name) 2026-09-20 13:12:29 -07:00
Adam Carr
4dab6cdb97 tokamak load test + dashboard unit fix (k6 RW emits seconds) 2026-09-20 11:29:05 -07:00
Adam Carr
97d5cda6f9 k6 step-load rig: staircase scenarios, per-step/per-endpoint tagging, live Prometheus remote-write, Grafana dashboard 2026-09-20 11:12:03 -07:00
Adam Carr
2bc85826f1 Admin stack: Authelia SSO, portal at admin.carr.pub, gated grafana.carr.pub 2026-09-19 21:56:58 -07:00
Adam Carr
282fd3db8c Documentation audit: refresh README, resolve stale PLAN.md items, list open items 2026-09-19 20:59:19 -07:00
Adam Carr
a1ee870f67 carr.pub live: cert-manager, hello site, DDNS; monitoring + loadtest committed 2026-09-19 20:57:04 -07:00
Adam Carr
8451333c3a Load-test rig: k6 swarm script + in-cluster target drone 2026-09-19 18:16:57 -07:00
Adam Carr
e46cf90e69 k3s cluster live: 3 servers + 17 workers, all 20 Ready 2026-09-19 17:50:08 -07:00
Adam Carr
08c40856fd k3s-join.sh: join servers 003/004 + 17 agents 2026-09-19 17:44:50 -07:00
27 changed files with 1917 additions and 46 deletions

4
.gitignore vendored
View File

@@ -1,2 +1,6 @@
configs/node-password.txt
.synology-password
configs/admin-password.txt
configs/authelia-session-secret.txt
configs/authelia-jwt-secret.txt
configs/gitea-password.txt

View File

@@ -1,20 +1,43 @@
# planck — 20-node Raspberry Pi 4 home lab
20× Raspberry Pi 4 Model B (8GB), each booting a 240GB USB SSD. Named
`planck001``planck020`. Network: Ubiquiti UDM Pro, 2 switches, one router.
Synology NAS on the same LAN for shared storage.
20× Raspberry Pi 4 (8GB, ARM64), each booting a 240GB NVMe-in-USB stick.
Fresh Raspberry Pi OS Lite (Debian 13) everywhere, running a k3s Kubernetes
cluster: servers planck002/003/004 (embedded etcd), workers 001 + 005020.
## Docs
**Home page: https://carr.pub** (served by this cluster).
- `docs/INVENTORY.md` — hardware inventory (regenerate: `./scripts/inventory.sh > docs/INVENTORY.md`)
- `docs/MACS.txt` — eth0 MAC addresses (for DHCP reservations)
- `docs/PLAN.md` — the rebuild plan and decisions log
## The map
## Scripts
| Path | What it is |
|------|-----------|
| `docs/PLAN.md` | Full history: what was built, every bug and its fix, postmortems. READ FIRST |
| `docs/IPS.md` | Node ↔ IP ↔ MAC table (static IPs, pinned on nodes) |
| `docs/INVENTORY.md` | Hardware inventory per node |
| `scripts/flash-one.sh` | Reimage a node's NVMe stick (run on laptop via pkexec) |
| `scripts/k3s-join.sh` | Join servers/agents to the cluster (needs K3S_TOKEN) |
| `scripts/loadtest.sh` | Spawn a k6 load swarm across N nodes |
| `scripts/inventory.sh` | Regenerate hardware inventory |
| `manifests/monitoring/` | Prometheus, node-exporter, Grafana, DDNS CronJob |
| `manifests/loadtest/` | k6 rig + practice target |
| `manifests/web/` | The carr.pub site + Ingress |
| `manifests/cluster-issuer.yaml` | Let's Encrypt issuer |
| `configs/` | Node password (gitignored), installer internals, retired netboot servers |
- `scripts/inventory.sh` — collects hardware/network inventory from all nodes
## Daily operations
## Node access
- kubectl on the laptop: `~/.local/bin/kubectl` (kubeconfig `~/.kube/config`)
- SSH any node: `ssh adamcarr@planck0NN.local` (passwordless sudo)
- Grafana: `kubectl -n monitoring port-forward svc/grafana 3000:3000`
- Load test: `scripts/loadtest.sh <url> [pods] [vus-per-pod] [duration]`
- New public site: Ingress in `web` namespace on `<name>.carr.pub` with the
`cert-manager.io/cluster-issuer: letsencrypt-prod` annotation + DNSimple
record
SSH: `ssh adamcarr@planck0NN.local` (key auth set up 2026-09-19). mDNS
resolves `planck0NN.local` across subnets.
Agent documentation lives in `~/.agents/skills/planck-cluster/SKILL.md`
(loaded automatically by opencode sessions).
## Safety rules
- Never reboot more than ONE of planck002/003/004 at once (etcd quorum)
- Images must be linux/arm64 multi-arch
- The cluster room switch shares a PSU with nodes 011015

View File

@@ -3,29 +3,34 @@
All nodes: static IP configured on the node (NetworkManager, "Wired
connection 1", manual, /22, gateway+DNS 192.168.1.1 = UDM Pro).
Every node is also on the owner's Tailscale tailnet (always-on) — the second
IP is reachable from anywhere via the tailnet, including the laptop
kubeconfig (API: https://100.124.112.13:6443 = planck002's Tailscale IP,
added to k3s tls-san on all three servers).
RECOMMENDED follow-up: add these as DHCP reservations in UniFi too, so the
router never leases these addresses to other devices. (UniFi UI: each client
device → Fixed IP. API is blocked by MFA.)
| Node | IP | MAC |
| Node | LAN IP | Tailscale IP | MAC |
|------|----|----|
| planck001 | 192.168.2.245 | dc:a6:32:e9:7f:75 |
| planck002 | 192.168.0.110 | dc:a6:32:e9:7f:de |
| planck003 | 192.168.1.109 | dc:a6:32:e9:7d:fa |
| planck004 | 192.168.2.250 | dc:a6:32:e9:7f:7b |
| planck005 | 192.168.2.178 | dc:a6:32:e9:7f:33 |
| planck006 | 192.168.1.2 | dc:a6:32:e9:7d:8b |
| planck007 | 192.168.1.48 | dc:a6:32:e9:7d:bd |
| planck008 | 192.168.2.43 | dc:a6:32:e9:7e:b2 |
| planck009 | 192.168.1.82 | dc:a6:32:e9:7d:df |
| planck010 | 192.168.2.91 | dc:a6:32:e9:7e:e2 |
| planck011 | 192.168.0.185 | dc:a6:32:e9:7d:46 |
| planck012 | 192.168.0.235 | dc:a6:32:e9:7d:78 |
| planck013 | 192.168.2.230 | dc:a6:32:e9:7f:67 |
| planck014 | 192.168.2.4 | dc:a6:32:e9:78:d1 |
| planck015 | 192.168.2.31 | dc:a6:32:e9:78:ec |
| planck016 | 192.168.2.94 | dc:a6:32:e9:7e:e5 |
| planck017 | 192.168.1.94 | dc:a6:32:e9:7d:eb |
| planck018 | 192.168.2.102 | dc:a6:32:e9:7e:ed |
| planck019 | 192.168.2.32 | dc:a6:32:e9:7e:a6 |
| planck020 | 192.168.1.23 | dc:a6:32:e9:7d:a4 |
| planck001 | 192.168.2.245 | 100.95.205.28 | dc:a6:32:e9:7f:75 |
| planck002 | 192.168.0.110 | 100.124.112.13 | dc:a6:32:e9:7f:de |
| planck003 | 192.168.1.109 | 100.98.49.13 | dc:a6:32:e9:7d:fa |
| planck004 | 192.168.2.250 | 100.93.201.48 | dc:a6:32:e9:7f:7b |
| planck005 | 192.168.2.178 | 100.98.22.50 | dc:a6:32:e9:7f:33 |
| planck006 | 192.168.1.2 | 100.117.72.42 | dc:a6:32:e9:7d:8b |
| planck007 | 192.168.1.48 | 100.116.208.66 | dc:a6:32:e9:7d:bd |
| planck008 | 192.168.2.43 | 100.125.106.86 | dc:a6:32:e9:7e:b2 |
| planck009 | 192.168.1.82 | 100.91.196.17 | dc:a6:32:e9:7d:df |
| planck010 | 192.168.2.91 | 100.65.93.90 | dc:a6:32:e9:7e:e2 |
| planck011 | 192.168.0.185 | 100.68.40.111 | dc:a6:32:e9:7d:46 |
| planck012 | 192.168.0.235 | 100.98.229.124 | dc:a6:32:e9:7d:78 |
| planck013 | 192.168.2.230 | 100.121.93.41 | dc:a6:32:e9:7f:67 |
| planck014 | 192.168.2.4 | 100.68.35.118 | dc:a6:32:e9:78:d1 |
| planck015 | 192.168.2.31 | 100.108.69.67 | dc:a6:32:e9:78:ec |
| planck016 | 192.168.2.94 | 100.88.118.95 | dc:a6:32:e9:7e:e5 |
| planck017 | 192.168.1.94 | 100.72.37.121 | dc:a6:32:e9:7d:eb |
| planck018 | 192.168.2.102 | 100.114.149.34 | dc:a6:32:e9:7e:ed |
| planck019 | 192.168.2.32 | 100.113.193.126 | dc:a6:32:e9:7e:a6 |
| planck020 | 192.168.1.23 | 100.104.220.31 | dc:a6:32:e9:7d:a4 |

View File

@@ -58,6 +58,26 @@ Fully unattended network reimage, no per-node physical access:
- UniFi MFA blocks API access; DHCP reservations deferred — mDNS
(`planck0NN.local`) works across subnets today.
## K3S CLUSTER LIVE (2026-09-19)
- **3 servers** (planck002/003/004, embedded etcd HA) + **17 workers**
- Kubernetes v1.36.4+k3s1, all 20 nodes Ready, static IPs
- kubectl on the laptop: `~/.local/bin/kubectl`, kubeconfig at `~/.kube/config`
(server URL https://192.168.0.110:6443). k3s node-token lives on planck002
at /var/lib/rancher/k3s/server/node-token.
- Join flow: `scripts/k3s-join.sh` (needs K3S_TOKEN env)
- Traefik ingress + local-path storage ship with k3s; default
- NOTE: Pi firmware injects `cgroup_disable=memory` — the node cmdline.txt
overrides with `cgroup_enable=cpuset cgroup_memory=1 cgroup_enable=memory`
(already applied fleet-wide; required for k3s).
- Domain: carr.pub is LIVE — DNSimple apex A record (id 84723206) →
50.46.44.67, UDM port-forwards 80/443 → planck002, Traefik ingress,
cert-manager Let's Encrypt issuer `letsencrypt-prod`, DDNS CronJob
`ddns-carr-pub` keeps DNS fresh. DNSimple token: k8s secret
`dnsimple-token` in namespace `monitoring`
- Next up: monitoring (light Prometheus+Grafana), Synology NFS storageclass,
carr.pub DNS + cert-manager, load-test rig, agent harness namespaces
## REBUILD COMPLETE (2026-09-19)
All 20 nodes re-imaged with Raspberry Pi OS Lite (Debian 13 trixie), via
@@ -149,18 +169,50 @@ TFTP_PREFIX=1):
- [x] Phase 0 — Access + inventory (SSH keys on all 20, docs/)
- [ ] Phase 1 — Network design: static IPs / reservations in UniFi
- [ ] Phase 2 — Reimage all 20 nodes with fresh Raspberry Pi OS
- [ ] Phase 3 — Cluster runtime install (candidates: k3s vs Docker Swarm)
- [ ] Phase 4 — Storage integration with Synology (NFS CSI)
- [ ] Phase 5 — Apps: monitoring, Portainer/UI, whatever the lab is for
- [x] Phase 3 — Cluster runtime install (k3s — decided and live)
- [x] Phase 4 — Fixed IPs pinned on nodes (Synology NFS storage class still open)
- [x] Phase 5 — Monitoring (Prometheus/Grafana), load-test rig, https://carr.pub
## Open decisions
## Resolved decisions (were "open" on day one)
1. **Reimage method**options:
a. PXE/network boot: EEPROM boot-order set to netboot; DHCP next-server
points at a TFTP/HTTP installer. Zero-touch per node, reusable forever.
Installer server can run on the Synology (Docker) or a laptop.
b. SD-card bootstrap installer: flash one SD image, insert per node, it
wipes + installs the USB SSD, removes itself. Simple, physical walk.
c. Pull SSDs, image on a PC with Raspberry Pi Imager ×20. Most manual.
2. **Cluster runtime** — k3s (lighter, k8s-compatible, we manage it with
automation) vs Docker Swarm + Portainer (simplest).
1. **Reimage method**pull-and-image on the laptop won (`flash-one.sh`).
The network-boot experiment is retired but documented below as a
postmortem; its Synology services linger (tftp disabled by design).
2. **Cluster runtime** — k3s, 3 servers (002-004) + 17 workers.
## Backups (established 2026-09-20)
Three layers:
1. **etcd snapshots** (cluster state — all k3s objects, secrets, config):
every 6h on each server node, 14-day retention
(`/etc/rancher/k3s/config.yaml`: `etcd-snapshot-schedule-cron: "0 */6 * * *"`).
Files at `/var/lib/rancher/k3s/server/db/snapshots/`. Manual test snapshot:
`sudo k3s etcd-snapshot save --name test`. Restore: stop k3s on all
servers, `sudo k3s server --cluster-init --cluster-reset
--cluster-reset-restore-path=<snapshot>` on one, then start all.
2. **NAS btrfs snapshots** (Gitea repos, k8s-volumes, netboot artifacts):
`/usr/local/bin/planck-backup.sh` on the Synology, daily at 3am, 14-day
retention, read-only btrfs snapshots of the `plancknetboot` share
(`plancknetboot@daily-YYYY-MM-DD`). Manual run: the script itself.
Restore: `btrfs subvolume snapshot /volume1/plancknetboot@daily-<date> /volume1/restore`.
3. **Laptop-only secret files** (node/gitea/admin passwords, authelia
session/jwt secrets — the gitignored ones): tarball at
`/volume1/plancknetboot/cluster-secrets/secrets.tar.gz` (root-only).
Refresh whenever a password is created/changed.
4. **Offsite**: Glacier Backup to AWS now includes the `plancknetboot` share
(owner verified the task 2026-09-20). Slow-restore tier for real
disasters; layers 1-2 are the fast-recovery tiers.
Grafana dashboards and all manifests live in git (Gitea + laptop), so they
are covered by layers 1-3 transitively. Prometheus history (15d) is
deliberately not backed up — monitoring data is expendable.
## Open items (next sessions)
- Synology NFS StorageClass (survives node swaps; local-path is current default)
- Grafana dashboards worth looking at (stack up, dashboards not built)
- Rotate the Grafana admin password (`planck-lab-admin` is a committed
placeholder in manifests/monitoring/grafana.yaml)
- UniFi DHCP reservations for the 20 node IPs (see docs/IPS.md)

View File

@@ -0,0 +1,133 @@
apiVersion: v1
kind: Namespace
metadata:
name: admin
---
apiVersion: v1
kind: ConfigMap
metadata:
name: authelia-config
namespace: admin
data:
configuration.yml: |
server:
address: tcp://:9091
theme: dark
log:
level: info
authentication_backend:
file:
path: /config/users_database.yml
password_policy:
standard:
enabled: true
min_length: 12
session:
cookies:
- domain: carr.pub
authelia_url: https://auth.carr.pub
default_redirection_url: https://admin.carr.pub
same_site: lax
storage:
local:
path: /config/data/db.sqlite3
notifier:
filesystem:
filename: /config/data/notifications.txt
totp:
issuer: carr.pub
webauthn:
disable: true
access_control:
default_policy: one_factor
rules: []
---
apiVersion: v1
kind: Secret
metadata:
name: authelia-users
namespace: admin
stringData:
users_database.yml: |
users:
adamcarr:
displayname: "Adam Carr"
password: "<argon2-hash from configs/admin-password.txt via deploy script>"
groups: ["admins"]
---
apiVersion: v1
kind: Secret
metadata:
name: authelia-secrets
namespace: admin
stringData:
session-secret: <from configs/authelia-session-secret.txt>
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: authelia-data
namespace: admin
spec:
accessModes: [ReadWriteOnce]
storageClassName: local-path
resources:
requests: {storage: 2Gi}
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: authelia
namespace: admin
spec:
replicas: 1
selector:
matchLabels: {app: authelia}
template:
metadata:
labels: {app: authelia}
spec:
containers:
- name: authelia
image: authelia/authelia:v4.39.28
args: ["--config", "/config/configuration.yml"]
env:
- name: AUTHELIA_SESSION_SECRET
valueFrom: {secretKeyRef: {name: authelia-secrets, key: session-secret}}
ports: [{containerPort: 9091}]
resources:
requests: {cpu: 50m, memory: 64Mi}
limits: {memory: 256Mi}
volumeMounts:
- {name: config, mountPath: /config/configuration.yml, subPath: configuration.yml}
- {name: users, mountPath: /config/users_database.yml, subPath: users_database.yml}
- {name: data, mountPath: /config/data}
volumes:
- name: config
configMap: {name: authelia-config}
- name: users
secret: {secretName: authelia-users}
- name: data
persistentVolumeClaim: {claimName: authelia-data}
---
apiVersion: v1
kind: Service
metadata:
name: authelia
namespace: admin
spec:
selector: {app: authelia}
ports:
- port: 9091
targetPort: 9091
---
apiVersion: traefik.io/v1alpha1
kind: Middleware
metadata:
name: authelia-auth
namespace: admin
spec:
forwardAuth:
address: "http://authelia.admin.svc:9091/api/verify?rd=https://auth.carr.pub"
trustForwardHeader: true
authResponseHeaders: [Remote-User, Remote-Groups, Remote-Email, Remote-Name]

View File

@@ -0,0 +1,61 @@
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: auth
namespace: admin
annotations:
cert-manager.io/cluster-issuer: letsencrypt-prod
spec:
tls:
- hosts: [auth.carr.pub]
secretName: auth-carr-pub-tls
rules:
- host: auth.carr.pub
http:
paths:
- path: /
pathType: Prefix
backend:
service: {name: authelia, port: {number: 9091}}
---
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: admin-portal
namespace: admin
annotations:
cert-manager.io/cluster-issuer: letsencrypt-prod
traefik.ingress.kubernetes.io/router.middlewares: admin-authelia-auth@kubernetescrd
spec:
tls:
- hosts: [admin.carr.pub]
secretName: admin-carr-pub-tls
rules:
- host: admin.carr.pub
http:
paths:
- path: /
pathType: Prefix
backend:
service: {name: portal, port: {number: 80}}
---
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: grafana
namespace: monitoring
annotations:
cert-manager.io/cluster-issuer: letsencrypt-prod
traefik.ingress.kubernetes.io/router.middlewares: admin-authelia-auth@kubernetescrd
spec:
tls:
- hosts: [grafana.carr.pub]
secretName: grafana-carr-pub-tls
rules:
- host: grafana.carr.pub
http:
paths:
- path: /
pathType: Prefix
backend:
service: {name: grafana, port: {number: 3000}}

View File

@@ -0,0 +1,68 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: portal-index
namespace: admin
data:
index.html: |
<!doctype html>
<html>
<head><meta charset="utf-8"><title>planck admin</title>
<style>
body { font-family: system-ui, sans-serif; background: #111; color: #eee;
display: grid; place-items: center; min-height: 100vh; margin: 0; }
.card { text-align: center; }
h1 { font-size: 2.5rem; margin: 0 0 1.5rem; }
a { display: block; margin: 0.6rem auto; padding: 0.8rem 1.5rem;
background: #1e1e2e; border: 1px solid #333; border-radius: 10px;
color: #8be9fd; text-decoration: none; font-size: 1.1rem; width: 260px; }
a:hover { background: #2a2a3e; }
p { color: #888; }
</style>
</head>
<body><div class="card">
<h1>planck</h1>
<a href="https://grafana.carr.pub">Grafana</a>
<a href="https://git.carr.pub">Gitea (git + CI)</a>
<a href="https://100.123.214.83:5001">Synology DSM + files</a>
<a href="https://auth.carr.pub">Account &amp; 2FA settings</a>
<p>20 nodes · k3s · carr.pub</p>
</div></body>
</html>
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: portal
namespace: admin
spec:
replicas: 2
selector:
matchLabels: {app: portal}
template:
metadata:
labels: {app: portal}
spec:
containers:
- name: web
image: nginx:1.29-alpine
ports: [{containerPort: 80}]
resources:
requests: {cpu: 10m, memory: 16Mi}
limits: {memory: 64Mi}
volumeMounts:
- {name: html, mountPath: /usr/share/nginx/html}
volumes:
- name: html
configMap: {name: portal-index}
---
apiVersion: v1
kind: Service
metadata:
name: portal
namespace: admin
spec:
selector: {app: portal}
ports:
- port: 80
targetPort: 80

View File

@@ -0,0 +1,14 @@
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: letsencrypt-prod
spec:
acme:
server: https://acme-v02.api.letsencrypt.org/directory
email: adam.carr@hey.com
privateKeySecretRef:
name: letsencrypt-prod-account-key
solvers:
- http01:
ingress:
class: traefik

View File

@@ -0,0 +1,75 @@
# act_runner x3 with docker-in-docker sidecars, so Gitea Actions jobs
# run as containers on the worker nodes (ARM64).
apiVersion: apps/v1
kind: Deployment
metadata:
name: act-runner
namespace: git
spec:
replicas: 3
selector:
matchLabels: {app: act-runner}
template:
metadata:
labels: {app: act-runner}
spec:
containers:
- name: runner
image: gitea/act_runner:latest
env:
- {name: GITEA_INSTANCE_URL, value: "http://gitea.git.svc:3000"}
- {name: GITEA_RUNNER_REGISTRATION_TOKEN, valueFrom: {secretKeyRef: {name: act-runner-token, key: token}}}
- {name: DOCKER_HOST, value: "tcp://127.0.0.1:2375"}
- {name: CONFIG_FILE, value: /config/config.yaml}
resources:
requests: {cpu: 50m, memory: 64Mi}
limits: {memory: 256Mi}
volumeMounts:
- {name: work, mountPath: /data}
- {name: config, mountPath: /config}
- name: dind
image: docker:27-dind
securityContext:
privileged: true
env:
- {name: DOCKER_TLS_CERTDIR, value: ""}
resources:
requests: {cpu: 100m, memory: 256Mi}
limits: {memory: 1Gi}
volumeMounts:
- {name: work, mountPath: /data}
volumes:
- name: work
emptyDir: {}
- name: config
configMap: {name: act-runner-config}
affinity:
podAntiAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
podAffinityTerm:
labelSelector:
matchLabels: {app: act-runner}
topologyKey: kubernetes.io/hostname
---
apiVersion: v1
kind: ConfigMap
metadata:
name: act-runner-config
namespace: git
data:
config.yaml: |
runner:
capacity: 4
timeout: 30m
insecure: false
fetch_timeout: 5s
labels:
- "arm64:host"
- "ubuntu-latest:docker://node:22-bookworm"
cache:
enabled: false
container:
privileged: false
options: ""
valid_volumes: []

86
manifests/git/gitea.yaml Normal file
View File

@@ -0,0 +1,86 @@
# Gitea: git hosting + CI, with repo data on Synology NFS.
apiVersion: v1
kind: Namespace
metadata:
name: git
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: gitea-data
namespace: git
spec:
accessModes: [ReadWriteOnce]
storageClassName: nfs-synology
resources:
requests: {storage: 20Gi}
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: gitea
namespace: git
spec:
replicas: 1
strategy:
type: Recreate
selector:
matchLabels: {app: gitea}
template:
metadata:
labels: {app: gitea}
spec:
containers:
- name: gitea
image: gitea/gitea:1.24
env:
- {name: USER_UID, value: "1000"}
- {name: USER_GID, value: "1000"}
- {name: GITEA__database__DB_TYPE, value: sqlite3}
- {name: GITEA__server__DOMAIN, value: git.carr.pub}
- {name: GITEA__server__ROOT_URL, value: "https://git.carr.pub/"}
- {name: GITEA__server__HTTP_PORT, value: "3000"}
- {name: GITEA__actions__ENABLED, value: "true"}
- {name: GITEA__service__DISABLE_REGISTRATION, value: "true"}
ports: [{containerPort: 3000}]
resources:
requests: {cpu: 100m, memory: 256Mi}
limits: {memory: 1Gi}
volumeMounts:
- {name: data, mountPath: /data}
volumes:
- name: data
persistentVolumeClaim: {claimName: gitea-data}
---
apiVersion: v1
kind: Service
metadata:
name: gitea
namespace: git
spec:
selector: {app: gitea}
ports:
- port: 3000
targetPort: 3000
---
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: gitea
namespace: git
annotations:
cert-manager.io/cluster-issuer: letsencrypt-prod
# NOTE: intentionally NOT behind the SSO gate - git clients and CI
# need direct token auth; Gitea has its own accounts
spec:
tls:
- hosts: [git.carr.pub]
secretName: git-carr-pub-tls
rules:
- host: git.carr.pub
http:
paths:
- path: /
pathType: Prefix
backend:
service: {name: gitea, port: {number: 3000}}

View File

@@ -0,0 +1,38 @@
apiVersion: v1
kind: Namespace
metadata:
name: loadtest
---
apiVersion: v1
kind: ConfigMap
metadata:
name: k6-script
namespace: loadtest
data:
test.js: |
// Generic load generator. Configured by env:
// TARGET_URL - full URL to hit
// VUS - virtual users per pod (default 20)
// DURATION - k6 duration string (default "60s")
import http from 'k6/http';
import { check, sleep } from 'k6';
const target = __ENV.TARGET_URL || 'http://example.com';
const vus = parseInt(__ENV.VUS || '20');
const duration = __ENV.DURATION || '60s';
export const options = {
stages: [
{ duration: '10s', target: vus },
{ duration: duration, target: vus },
{ duration: '5s', target: 0 },
],
thresholds: { http_req_failed: ['rate<0.01'] },
noConnectionReuse: false,
};
export default function () {
const res = http.get(target, { timeout: '10s' });
check(res, { 'status 2xx/3xx': (r) => r.status < 300 });
sleep(1 / vus * 10);
}

View File

@@ -0,0 +1,94 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: k6-step-script
namespace: loadtest
data:
step.js: |
// Step-load (staircase) test. Hits a collection of endpoints, stepping
// the request rate up over time, tagging every request with its step and
// endpoint so latency can be sliced either way in Grafana.
//
// Env:
// BASE_URL - target base URL
// ENDPOINTS_JSON - [{"name":"list","method":"GET","path":"/x","weight":2}, ...]
// START_RPS - rate of step 1
// STEP_RPS - rate added each step
// STEP_DURATION - hold time per step, e.g. "2m"
// STEPS - number of steps
import http from 'k6/http';
import { check } from 'k6';
import exec from 'k6/execution';
const BASE = __ENV.BASE_URL;
let endpoints;
try {
endpoints = JSON.parse(__ENV.ENDPOINTS_JSON);
} catch (e) {
throw new Error('ENDPOINTS_JSON is not valid JSON: ' + e.message);
}
// expand weights
const pool = [];
for (const ep of endpoints) {
for (let i = 0; i < (ep.weight || 1); i++) pool.push(ep);
}
const startRps = parseInt(__ENV.START_RPS || '50');
const stepRps = parseInt(__ENV.STEP_RPS || '50');
const stepDurMs = ms(__ENV.STEP_DURATION || '2m');
const steps = parseInt(__ENV.STEPS || '5');
const rampMs = 5000;
// Build stage list and per-step [startMs, endMs, label] windows.
// Each step: 5s quick ramp to the rate, then a hold.
const stages = [];
const windows = [];
let clock = 0;
for (let i = 0; i < steps; i++) {
const rate = startRps + i * stepRps;
stages.push({ duration: ms2str(rampMs), target: rate });
clock += rampMs;
const wStart = clock;
stages.push({ duration: ms2str(stepDurMs), target: rate });
clock += stepDurMs;
windows.push([wStart, clock, `step${i + 1} @ ${rate} rps`]);
}
const totalMs = clock;
const maxRate = startRps + (steps - 1) * stepRps;
const vus = Math.max(10, Math.ceil(maxRate * 2));
function ms(s) {
const m = /^(\d+)(ms|s|m)$/.exec(String(s));
if (!m) throw new Error('bad duration: ' + s);
return +m[1] * ({ ms: 1, s: 1000, m: 60000 })[m[2]];
}
function ms2str(n) { return n + 'ms'; }
export const options = {
discardResponseBodies: true,
scenarios: {
staircase: {
executor: 'ramping-arrival-rate',
startRate: 0,
preAllocatedVUs: vus,
maxVUs: vus,
stages: stages,
},
},
thresholds: { http_req_failed: ['rate<0.05'] },
};
export default function () {
const elapsed = exec.scenario.progress * totalMs;
let step = 'ramp-up';
for (const [s, e, label] of windows) {
if (elapsed >= s && elapsed < e) { step = label; break; }
}
const ep = pool[exec.scenario.iterationInTest % pool.length];
const url = BASE + ep.path;
const params = { tags: { name: ep.name, step: step }, timeout: '15s' };
const res = (ep.method || 'GET') === 'POST'
? http.post(url, ep.body || '', params)
: http.get(url, params);
check(res, { ok: (r) => r.status < 300 });
}

View File

@@ -0,0 +1,134 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: k6-tokamak-script
namespace: loadtest
data:
tokamak.js: |
// Staircase load test for the tokamak public API.
// Setup harvests real IDs from the list endpoints, then iterations fan out
// across the whole endpoint collection, weighted so light calls dominate
// and the two heavy data endpoints are sampled less often.
//
// Env: BASE_URL, TOKAMAK_API_KEY, START_RPS, STEP_RPS, STEP_DURATION, STEPS
import http from 'k6/http';
import { check, fail } from 'k6';
import exec from 'k6/execution';
const BASE = __ENV.BASE_URL;
const auth = { headers: { Authorization: 'Bearer ' + __ENV.TOKAMAK_API_KEY } };
const startRps = parseInt(__ENV.START_RPS || '5');
const stepRps = parseInt(__ENV.STEP_RPS || '5');
const stepDurMs = ms(__ENV.STEP_DURATION || '1m');
const steps = parseInt(__ENV.STEPS || '4');
const rampMs = 5000;
function ms(s) {
const m = /^(\d+)(ms|s|m)$/.exec(String(s));
if (!m) throw new Error('bad duration: ' + s);
return +m[1] * ({ ms: 1, s: 1000, m: 60000 })[m[2]];
}
function ms2str(n) { return n + 'ms'; }
// staircase stages + per-step windows (5s ramp + hold each)
const stages = [];
const windows = [];
let clock = 0;
for (let i = 0; i < steps; i++) {
const rate = startRps + i * stepRps;
stages.push({ duration: ms2str(rampMs), target: rate });
clock += rampMs;
const wStart = clock;
stages.push({ duration: ms2str(stepDurMs), target: rate });
clock += stepDurMs;
windows.push([wStart, clock, 'step' + (i + 1) + ' @ ' + rate + ' rps']);
}
const totalMs = clock;
const maxRate = startRps + (steps - 1) * stepRps;
const vus = Math.max(10, Math.ceil(maxRate * 1.5));
export const options = {
discardResponseBodies: false,
scenarios: {
staircase: {
executor: 'ramping-arrival-rate',
startRate: 0,
preAllocatedVUs: vus,
maxVUs: vus,
stages: stages,
},
},
thresholds: { http_req_failed: ['rate<0.05'] },
};
// defensive id harvesting - handle data/items/results/root-array shapes
function idsOf(json) {
if (!json) return [];
const arr = json.data || json.items || json.results
|| json.brands || json.cohorts || json.metrics || json.frameworks || json;
if (!Array.isArray(arr)) return [];
return arr.map(function (x) {
return x && (x.id || x.brand_id || x.cohort_id || x.metric_id || x.framework_id);
}).filter(Boolean);
}
export function setup() {
const me = http.get(BASE + '/me', auth);
if (me.status !== 200) {
fail('auth check failed: /me returned ' + me.status + ' - check API key and environment');
}
const brands = idsOf(http.get(BASE + '/brands?per_page=100', auth).json());
const cohorts = idsOf(http.get(BASE + '/cohorts?per_page=100', auth).json());
const metrics = idsOf(http.get(BASE + '/metrics', auth).json());
const frameworks = idsOf(http.get(BASE + '/frameworks', auth).json());
console.log('harvested: brands=' + brands.length + ' cohorts=' + cohorts.length
+ ' metrics=' + metrics.length + ' frameworks=' + frameworks.length);
if (!brands.length) fail('no brand ids harvested - cannot build data calls');
return { brands: brands, cohorts: cohorts, metrics: metrics, frameworks: frameworks };
}
// weighted call pool: light lists and detail reads, heavy data sampled less
function pickCall(d, i) {
const b = d.brands[i % d.brands.length];
const c = d.cohorts.length ? d.cohorts[i % d.cohorts.length] : null;
const m = d.metrics.length ? d.metrics[i % d.metrics.length] : null;
const f = d.frameworks.length ? d.frameworks[i % d.frameworks.length] : null;
const r = Math.random();
if (r < 0.10) return { name: 'getStatus', path: '/' };
if (r < 0.20) return { name: 'getMe', path: '/me' };
if (r < 0.32) return { name: 'listBrands', path: '/brands?per_page=25' };
if (r < 0.42) return { name: 'listMetrics', path: '/metrics' };
if (r < 0.50) return { name: 'listFrameworks', path: '/frameworks' };
if (r < 0.62) return { name: 'getBrand', path: '/brands/' + b };
if (r < 0.70) return c
? { name: 'getCohort', path: '/cohorts/' + c }
: { name: 'listBrands', path: '/brands?per_page=25' };
if (r < 0.76) return c
? { name: 'listCohortBrands', path: '/cohorts/' + c + '/brands?per_page=25' }
: { name: 'listCohorts', path: '/cohorts?per_page=25' };
if (r < 0.78) return f
? { name: 'getFramework', path: '/frameworks/' + f }
: { name: 'listFrameworks', path: '/frameworks' };
if (r < 0.90) return m
? { name: 'getMetricData', path: '/metric/data/' + m + '/' + b }
: { name: 'listMetrics', path: '/metrics' };
return f
? { name: 'getFrameworkData', path: '/framework/data/' + f + '/' + b }
: { name: 'listFrameworks', path: '/frameworks' };
}
export default function (d) {
const elapsed = exec.scenario.progress * totalMs;
let step = 'ramp-up';
for (const w of windows) {
if (elapsed >= w[0] && elapsed < w[1]) { step = w[2]; break; }
}
const call = pickCall(d, exec.scenario.iterationInTest);
const res = http.get(BASE + call.path, {
headers: auth.headers,
tags: { name: call.name, step: step },
timeout: '30s',
});
check(res, { ok: function (r) { return r.status < 300; } });
}

View File

@@ -0,0 +1,31 @@
apiVersion: apps/v1
kind: Deployment
metadata:
name: target-drone
namespace: loadtest
spec:
replicas: 3
selector:
matchLabels: {app: target-drone}
template:
metadata:
labels: {app: target-drone}
spec:
containers:
- name: nginx
image: nginx:1.29-alpine
ports: [{containerPort: 80}]
resources:
requests: {cpu: 10m, memory: 16Mi}
limits: {memory: 64Mi}
---
apiVersion: v1
kind: Service
metadata:
name: target-drone
namespace: loadtest
spec:
selector: {app: target-drone}
ports:
- port: 80
targetPort: 80

View File

@@ -0,0 +1,69 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: ddns-script
namespace: monitoring
data:
ddns.sh: |
#!/bin/sh
# Keeps the carr.pub apex A record pointing at this network's public IP.
RECORD_ID=84723206
ZONE="carr.pub"
ACCOUNT=6300
API="https://api.dnsimple.com/v2"
AUTH="Authorization: Bearer ${DNSIMPLE_TOKEN}"
IP=$(curl -s -m 15 https://ifconfig.me)
case "$IP" in
*.*.*.*) : ;;
*) echo "no valid public ip: '$IP'"; exit 0 ;;
esac
CUR=$(curl -s -m 15 -H "$AUTH" "$API/$ACCOUNT/zones/$ZONE/records/$RECORD_ID" \
| sed -n 's/.*"content":"\([^"]*\)".*/\1/p')
if [ "$IP" = "$CUR" ]; then
echo "$(date -Is) up to date ($IP)"
exit 0
fi
RES=$(curl -s -m 15 -X PATCH -H "$AUTH" -H "Content-Type: application/json" \
"$API/$ACCOUNT/zones/$ZONE/records/$RECORD_ID" \
-d "{\"content\":\"$IP\"}")
echo "$(date -Is) updated carr.pub -> $IP"
case "$RES" in
*'"id"'*) exit 0 ;;
*) echo "API error: $RES"; exit 1 ;;
esac
---
apiVersion: batch/v1
kind: CronJob
metadata:
name: ddns-carr-pub
namespace: monitoring
spec:
schedule: "*/5 * * * *"
concurrencyPolicy: Forbid
successfulJobsHistoryLimit: 1
failedJobsHistoryLimit: 3
jobTemplate:
spec:
backoffLimit: 1
template:
spec:
restartPolicy: Never
containers:
- name: ddns
image: curlimages/curl:8.11.1
command: ["/bin/sh", "/scripts/ddns.sh"]
env:
- name: DNSIMPLE_TOKEN
valueFrom:
secretKeyRef: {name: dnsimple-token, key: token}
resources:
requests: {cpu: 10m, memory: 16Mi}
limits: {memory: 32Mi}
volumeMounts:
- {name: scripts, mountPath: /scripts}
volumes:
- name: scripts
configMap: {name: ddns-script, defaultMode: 0744}

View File

@@ -0,0 +1,212 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: grafana-dashboard-provider
namespace: monitoring
data:
dashboards.yaml: |
apiVersion: 1
providers:
- name: default
orgId: 1
folder: ""
type: file
disableDeletion: false
updateIntervalSeconds: 30
options:
path: /var/lib/grafana/dashboards
---
apiVersion: v1
kind: ConfigMap
metadata:
name: grafana-dashboards
namespace: monitoring
data:
k3s-cluster.json: |
{
"title": "planck k3s cluster",
"uid": "planckk3s",
"timezone": "browser",
"refresh": "15s",
"time": { "from": "now-3h", "to": "now" },
"panels": [
{
"id": 1, "type": "timeseries", "title": "Pods running by node",
"gridPos": {"x": 0, "y": 0, "w": 12, "h": 9},
"targets": [{"expr": "count by (node) (kube_pod_status_phase{phase=\"Running\"} == 1)", "legendFormat": "{{node}}"}],
"fieldConfig": {"defaults": {"unit": "short", "min": 0}, "overrides": []}
},
{
"id": 2, "type": "timeseries", "title": "Deployment replicas: ready vs desired",
"gridPos": {"x": 12, "y": 0, "w": 12, "h": 9},
"targets": [
{"expr": "sum by (deployment, namespace) (kube_deployment_status_replicas_ready)", "legendFormat": "ready {{namespace}}/{{deployment}}"},
{"expr": "sum by (deployment, namespace) (kube_deployment_spec_replicas)", "legendFormat": "want {{namespace}}/{{deployment}}"}
],
"fieldConfig": {"defaults": {"unit": "short", "min": 0}, "overrides": []}
},
{
"id": 3, "type": "timeseries", "title": "Top pods by CPU (cores)",
"gridPos": {"x": 0, "y": 9, "w": 12, "h": 9},
"targets": [{"expr": "topk(12, sum by (pod, namespace) (rate(container_cpu_usage_seconds_total{container!=\"\", container!=\"POD\"}[2m])))", "legendFormat": "{{namespace}}/{{pod}}"}],
"fieldConfig": {"defaults": {"unit": "none"}, "overrides": []}
},
{
"id": 4, "type": "timeseries", "title": "Top pods by memory (working set)",
"gridPos": {"x": 12, "y": 9, "w": 12, "h": 9},
"targets": [{"expr": "topk(12, sum by (pod, namespace) (container_memory_working_set_bytes{container!=\"\", container!=\"POD\"}))", "legendFormat": "{{namespace}}/{{pod}}"}],
"fieldConfig": {"defaults": {"unit": "bytes"}, "overrides": []}
},
{
"id": 5, "type": "timeseries", "title": "Container restarts (cumulative)",
"gridPos": {"x": 0, "y": 18, "w": 12, "h": 9},
"targets": [{"expr": "sum by (pod, namespace) (kube_pod_container_status_restarts_total)", "legendFormat": "{{namespace}}/{{pod}}"}],
"fieldConfig": {"defaults": {"unit": "short", "min": 0}, "overrides": []}
},
{
"id": 6, "type": "timeseries", "title": "CPU: requested vs allocatable (per node, top 10 nodes)",
"gridPos": {"x": 12, "y": 18, "w": 12, "h": 9},
"targets": [
{"expr": "topk(10, sum by (node) (kube_pod_container_resource_requests{resource=\"cpu\", unit=\"core\"}))", "legendFormat": "requested {{node}}"},
{"expr": "kube_node_status_allocatable{resource=\"cpu\", unit=\"core\"}", "legendFormat": "allocatable {{node}}"}
],
"fieldConfig": {"defaults": {"unit": "none"}, "overrides": []}
},
{
"id": 7, "type": "timeseries", "title": "Nodes Ready",
"gridPos": {"x": 0, "y": 27, "w": 12, "h": 7},
"targets": [{"expr": "count(kube_node_status_condition{condition=\"Ready\", status=\"true\"} == 1)", "legendFormat": "nodes ready"}],
"fieldConfig": {"defaults": {"unit": "short", "min": 0, "max": 20}, "overrides": []}
},
{
"id": 8, "type": "timeseries", "title": "Pods not Running (pending/failed/succeeded)",
"gridPos": {"x": 12, "y": 27, "w": 12, "h": 7},
"targets": [{"expr": "count by (phase) (kube_pod_status_phase == 1) unless on(pod, namespace, phase) label_replace(kube_pod_status_phase{phase=\"Running\"}, \"phase\", \"running\", \"\", \"\") * 0", "legendFormat": "{{phase}}"}],
"fieldConfig": {"defaults": {"unit": "short", "min": 0}, "overrides": []}
}
],
"schemaVersion": 39,
"version": 1
}
node-health.json: |
{
"title": "planck node health",
"uid": "plancknodes",
"timezone": "browser",
"refresh": "10s",
"time": { "from": "now-1h", "to": "now" },
"panels": [
{
"id": 1, "type": "timeseries", "title": "CPU temperature (all nodes)",
"gridPos": {"x": 0, "y": 0, "w": 12, "h": 8},
"targets": [{"expr": "node_thermal_zone_temp", "legendFormat": "{{instance}}"}],
"fieldConfig": {"defaults": {"unit": "celsius"}, "overrides": []}
},
{
"id": 2, "type": "timeseries", "title": "CPU usage % (all nodes)",
"gridPos": {"x": 12, "y": 0, "w": 12, "h": 8},
"targets": [{"expr": "100 * (1 - avg by (instance) (rate(node_cpu_seconds_total{mode=\"idle\"}[2m])))", "legendFormat": "{{instance}}"}],
"fieldConfig": {"defaults": {"unit": "percent", "min": 0, "max": 100}, "overrides": []}
},
{
"id": 3, "type": "timeseries", "title": "Memory used % (all nodes)",
"gridPos": {"x": 0, "y": 8, "w": 12, "h": 8},
"targets": [{"expr": "100 * (1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes)", "legendFormat": "{{instance}}"}],
"fieldConfig": {"defaults": {"unit": "percent", "min": 0, "max": 100}, "overrides": []}
},
{
"id": 4, "type": "timeseries", "title": "Disk used % (root, all nodes)",
"gridPos": {"x": 12, "y": 8, "w": 12, "h": 8},
"targets": [{"expr": "100 * (1 - node_filesystem_avail_bytes{mountpoint=\"/\"} / node_filesystem_size_bytes{mountpoint=\"/\"})", "legendFormat": "{{instance}}"}],
"fieldConfig": {"defaults": {"unit": "percent", "min": 0, "max": 100}, "overrides": []}
},
{
"id": 5, "type": "timeseries", "title": "Network eth0: RX / TX",
"gridPos": {"x": 0, "y": 16, "w": 12, "h": 8},
"targets": [
{"expr": "sum by (instance) (rate(node_network_receive_bytes_total{device=\"eth0\"}[1m]))", "legendFormat": "rx {{instance}}"},
{"expr": "sum by (instance) (rate(node_network_transmit_bytes_total{device=\"eth0\"}[1m]))", "legendFormat": "tx {{instance}}"}
],
"fieldConfig": {"defaults": {"unit": "Bps"}, "overrides": []}
},
{
"id": 6, "type": "timeseries", "title": "SSD I/O: read / write",
"gridPos": {"x": 12, "y": 16, "w": 12, "h": 8},
"targets": [
{"expr": "sum by (instance) (rate(node_disk_read_bytes_total{device=~\"sda.*\"}[1m]))", "legendFormat": "read {{instance}}"},
{"expr": "sum by (instance) (rate(node_disk_written_bytes_total{device=~\"sda.*\"}[1m]))", "legendFormat": "write {{instance}}"}
],
"fieldConfig": {"defaults": {"unit": "Bps"}, "overrides": []}
},
{
"id": 7, "type": "timeseries", "title": "Load (1 min, all nodes)",
"gridPos": {"x": 0, "y": 24, "w": 24, "h": 7},
"targets": [{"expr": "node_load1", "legendFormat": "{{instance}}"}],
"fieldConfig": {"defaults": {"unit": "short"}, "overrides": []}
}
],
"schemaVersion": 39,
"version": 1
}
k6-step-load.json: |
{
"title": "k6 Step Load",
"uid": "k6step",
"timezone": "browser",
"refresh": "5s",
"time": { "from": "now-30m", "to": "now" },
"panels": [
{
"id": 1, "type": "timeseries", "title": "Request rate (the staircase)",
"gridPos": {"x": 0, "y": 0, "w": 24, "h": 8},
"targets": [
{"expr": "sum(rate(k6_http_reqs_total[30s]))", "legendFormat": "total rps"},
{"expr": "sum by (name) (rate(k6_http_reqs_total[30s]))", "legendFormat": "{{name}}"}
],
"fieldConfig": {"defaults": {"unit": "reqps"}, "overrides": []}
},
{
"id": 2, "type": "timeseries", "title": "Latency avg / p50 / p95 / p99",
"gridPos": {"x": 0, "y": 8, "w": 12, "h": 8},
"targets": [
{"expr": "k6_http_req_duration_avg", "legendFormat": "avg"},
{"expr": "k6_http_req_duration_p50", "legendFormat": "p50"},
{"expr": "k6_http_req_duration_p95", "legendFormat": "p95"},
{"expr": "k6_http_req_duration_p99", "legendFormat": "p99"}
],
"fieldConfig": {"defaults": {"unit": "s"}, "overrides": []}
},
{
"id": 3, "type": "timeseries", "title": "p95 latency per step",
"gridPos": {"x": 12, "y": 8, "w": 12, "h": 8},
"targets": [
{"expr": "k6_http_req_duration_p95", "legendFormat": "{{step}}"}
],
"fieldConfig": {"defaults": {"unit": "s"}, "overrides": []}
},
{
"id": 4, "type": "timeseries", "title": "p95 latency per endpoint",
"gridPos": {"x": 0, "y": 16, "w": 12, "h": 8},
"targets": [
{"expr": "k6_http_req_duration_p95", "legendFormat": "{{name}}"}
],
"fieldConfig": {"defaults": {"unit": "s"}, "overrides": []}
},
{
"id": 5, "type": "timeseries", "title": "Error rate",
"gridPos": {"x": 12, "y": 16, "w": 12, "h": 8},
"targets": [
{"expr": "sum(rate(k6_http_req_failed_total[30s])) / sum(rate(k6_http_reqs_total[30s]))", "legendFormat": "error fraction"}
],
"fieldConfig": {"defaults": {"unit": "percentunit", "min": 0, "max": 1}, "overrides": []}
},
{
"id": 6, "type": "timeseries", "title": "Active VUs",
"gridPos": {"x": 0, "y": 24, "w": 24, "h": 6},
"targets": [{"expr": "k6_vus", "legendFormat": "vus"}],
"fieldConfig": {"defaults": {"unit": "short"}, "overrides": []}
}
],
"schemaVersion": 39,
"version": 1
}

View File

@@ -0,0 +1,71 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: grafana-datasources
namespace: monitoring
data:
datasources.yaml: |
apiVersion: 1
datasources:
- name: Prometheus
type: prometheus
access: proxy
url: http://prometheus.monitoring.svc:9090
isDefault: true
---
apiVersion: v1
kind: Secret
metadata:
name: grafana-admin
namespace: monitoring
stringData:
admin-user: adamcarr
admin-password: planck-lab-admin
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: grafana
namespace: monitoring
spec:
replicas: 1
selector:
matchLabels: {app: grafana}
template:
metadata:
labels: {app: grafana}
spec:
containers:
- name: grafana
image: grafana/grafana:12.0.2
env:
- {name: GF_SECURITY_ADMIN_USER, valueFrom: {secretKeyRef: {name: grafana-admin, key: admin-user}}}
- {name: GF_SECURITY_ADMIN_PASSWORD, valueFrom: {secretKeyRef: {name: grafana-admin, key: admin-password}}}
- {name: GF_AUTH_ANONYMOUS_ENABLED, value: "true"}
ports:
- containerPort: 3000
resources:
requests: {cpu: 50m, memory: 128Mi}
limits: {memory: 512Mi}
volumeMounts:
- {name: datasources, mountPath: /etc/grafana/provisioning/datasources}
- {name: dashboard-provider, mountPath: /etc/grafana/provisioning/dashboards}
- {name: dashboards, mountPath: /var/lib/grafana/dashboards}
volumes:
- name: datasources
configMap: {name: grafana-datasources}
- name: dashboard-provider
configMap: {name: grafana-dashboard-provider}
- name: dashboards
configMap: {name: grafana-dashboards}
---
apiVersion: v1
kind: Service
metadata:
name: grafana
namespace: monitoring
spec:
selector: {app: grafana}
ports:
- port: 3000
targetPort: 3000

View File

@@ -0,0 +1,71 @@
apiVersion: v1
kind: ServiceAccount
metadata:
name: kube-state-metrics
namespace: monitoring
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: kube-state-metrics
rules:
- apiGroups: [""]
resources: [pods, nodes, namespaces, services, endpoints, configmaps, secrets, persistentvolumeclaims, persistentvolumes, resourcequotas, limitranges]
verbs: [list, watch]
- apiGroups: [apps]
resources: [deployments, daemonsets, replicasets, statefulsets]
verbs: [list, watch]
- apiGroups: [batch]
resources: [jobs, cronjobs]
verbs: [list, watch]
- apiGroups: [autoscaling]
resources: [horizontalpodautoscalers]
verbs: [list, watch]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: kube-state-metrics
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: kube-state-metrics
subjects:
- kind: ServiceAccount
name: kube-state-metrics
namespace: monitoring
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: kube-state-metrics
namespace: monitoring
spec:
replicas: 1
selector:
matchLabels: {app: kube-state-metrics}
template:
metadata:
labels: {app: kube-state-metrics}
spec:
serviceAccountName: kube-state-metrics
containers:
- name: kube-state-metrics
image: registry.k8s.io/kube-state-metrics/kube-state-metrics:v2.17.0
ports: [{containerPort: 8080}]
resources:
requests: {cpu: 30m, memory: 64Mi}
limits: {memory: 512Mi}
---
apiVersion: v1
kind: Service
metadata:
name: kube-state-metrics
namespace: monitoring
labels: {app: kube-state-metrics}
spec:
selector: {app: kube-state-metrics}
ports:
- name: http-metrics
port: 8080
targetPort: 8080

View File

@@ -0,0 +1,114 @@
# NFS dynamic provisioning from the Synology (192.168.1.157).
# Base dir /volume1/plancknetboot/k8s-volumes; each PVC gets a subdir.
apiVersion: v1
kind: ServiceAccount
metadata:
name: nfs-provisioner
namespace: monitoring
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: nfs-provisioner
rules:
- apiGroups: [""]
resources: [persistentvolumes]
verbs: [get, list, watch, create, delete]
- apiGroups: [""]
resources: [persistentvolumeclaims]
verbs: [get, list, watch, update]
- apiGroups: ["storage.k8s.io"]
resources: [storageclasses, volumeattributesclasses]
verbs: [get, list, watch]
- apiGroups: [""]
resources: [events]
verbs: [create, update, patch]
- apiGroups: ["storage.k8s.io"]
resources: [csidrivers]
verbs: [get, list]
- apiGroups: [""]
resources: [nodes]
verbs: [get, list, watch]
- apiGroups: ["storage.k8s.io"]
resources: [volumeattachments]
verbs: [get, list, watch, delete]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: nfs-provisioner
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: nfs-provisioner
subjects:
- kind: ServiceAccount
name: nfs-provisioner
namespace: monitoring
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: nfs-provisioner
namespace: monitoring
rules:
- apiGroups: [""]
resources: [endpoints]
verbs: [get, list, watch, create, update, patch]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: nfs-provisioner
namespace: monitoring
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: Role
name: nfs-provisioner
subjects:
- kind: ServiceAccount
name: nfs-provisioner
namespace: monitoring
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: nfs-provisioner
namespace: monitoring
spec:
replicas: 1
selector:
matchLabels: {app: nfs-provisioner}
template:
metadata:
labels: {app: nfs-provisioner}
spec:
serviceAccountName: nfs-provisioner
containers:
- name: nfs-provisioner
image: registry.k8s.io/sig-storage/nfs-subdir-external-provisioner:v4.0.2
volumeMounts:
- {name: nfs, mountPath: /persistentvolumes}
env:
- {name: PROVISIONER_NAME, value: nfs-synology}
- {name: NFS_SERVER, value: "192.168.1.157"}
- {name: NFS_PATH, value: /volume1/plancknetboot/k8s-volumes}
- {name: ENABLE_LEADER_ELECTION, value: "false"}
resources:
requests: {cpu: 20m, memory: 32Mi}
limits: {memory: 128Mi}
volumes:
- name: nfs
nfs:
server: "192.168.1.157"
path: /volume1/plancknetboot/k8s-volumes
---
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: nfs-synology
provisioner: nfs-synology
reclaimPolicy: Delete
volumeBindingMode: Immediate
parameters:
archiveOnDelete: "true"

View File

@@ -0,0 +1,41 @@
apiVersion: apps/v1
kind: DaemonSet
metadata:
name: node-exporter
namespace: monitoring
labels: {app: node-exporter}
spec:
selector:
matchLabels: {app: node-exporter}
template:
metadata:
labels: {app: node-exporter}
spec:
hostNetwork: true
hostPID: true
tolerations:
- operator: Exists
containers:
- name: node-exporter
image: quay.io/prometheus/node-exporter:v1.9.1
args:
- --path.procfs=/host/proc
- --path.sysfs=/host/sys
- --collector.filesystem.mount-points-exclude=^/(dev|proc|sys|var/lib/docker/.+|var/lib/kubelet/.+)($|/)
ports:
- containerPort: 9100
hostPort: 9100
resources:
requests: {cpu: 20m, memory: 32Mi}
limits: {memory: 64Mi}
volumeMounts:
- {name: proc, mountPath: /host/proc, readOnly: true}
- {name: sys, mountPath: /host/sys, readOnly: true}
- {name: root, mountPath: /rootfs, readOnly: true}
volumes:
- name: proc
hostPath: {path: /proc}
- name: sys
hostPath: {path: /sys}
- name: root
hostPath: {path: /}

View File

@@ -0,0 +1,160 @@
apiVersion: v1
kind: Namespace
metadata:
name: monitoring
---
apiVersion: v1
kind: ServiceAccount
metadata:
name: prometheus
namespace: monitoring
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: prometheus
rules:
- apiGroups: [""]
resources: [nodes, nodes/proxy, services, endpoints, pods]
verbs: [get, list, watch]
- apiGroups: ["extensions", "networking.k8s.io"]
resources: [ingresses]
verbs: [get, list, watch]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: prometheus
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: prometheus
subjects:
- kind: ServiceAccount
name: prometheus
namespace: monitoring
---
apiVersion: v1
kind: ConfigMap
metadata:
name: prometheus-config
namespace: monitoring
data:
prometheus.yml: |
global:
scrape_interval: 15s
external_labels:
cluster: planck
scrape_configs:
- job_name: prometheus
static_configs:
- targets: ["localhost:9090"]
- job_name: node-exporter
kubernetes_sd_configs:
- role: pod
selectors:
- role: pod
label: "app=node-exporter"
relabel_configs:
- source_labels: [__meta_kubernetes_pod_node_name]
target_label: instance
- source_labels: [__meta_kubernetes_pod_container_port_number]
action: keep
regex: "9100"
- job_name: kube-state-metrics
kubernetes_sd_configs:
- role: pod
selectors:
- role: pod
label: "app=kube-state-metrics"
relabel_configs:
- source_labels: [__meta_kubernetes_pod_container_port_number]
action: keep
regex: "8080"
- job_name: kubelet
scheme: https
tls_config:
insecure_skip_verify: true
authorization:
credentials_file: /var/run/secrets/kubernetes.io/serviceaccount/token
kubernetes_sd_configs:
- role: node
relabel_configs:
- source_labels: [__meta_kubernetes_node_name]
target_label: instance
- target_label: __address__
replacement: kubernetes.default.svc:443
- source_labels: [__meta_kubernetes_node_name]
target_label: __metrics_path__
regex: (.+)
replacement: /api/v1/nodes/${1}/proxy/metrics
- job_name: cadvisor
scheme: https
tls_config:
insecure_skip_verify: true
authorization:
credentials_file: /var/run/secrets/kubernetes.io/serviceaccount/token
kubernetes_sd_configs:
- role: node
relabel_configs:
- source_labels: [__meta_kubernetes_node_name]
target_label: instance
- target_label: __address__
replacement: kubernetes.default.svc:443
- source_labels: [__meta_kubernetes_node_name]
target_label: __metrics_path__
regex: (.+)
replacement: /api/v1/nodes/${1}/proxy/metrics/cadvisor
---
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: prometheus
namespace: monitoring
spec:
serviceName: prometheus
replicas: 1
selector:
matchLabels: {app: prometheus}
template:
metadata:
labels: {app: prometheus}
spec:
serviceAccountName: prometheus
containers:
- name: prometheus
image: prom/prometheus:v3.6.0
args:
- --config.file=/etc/prometheus/prometheus.yml
- --storage.tsdb.retention.time=15d
- --storage.tsdb.path=/data
- --web.enable-remote-write-receiver
ports:
- containerPort: 9090
resources:
requests: {cpu: 150m, memory: 512Mi}
limits: {memory: 1Gi}
volumeMounts:
- {name: config, mountPath: /etc/prometheus}
- {name: data, mountPath: /data}
volumes:
- name: config
configMap: {name: prometheus-config}
volumeClaimTemplates:
- metadata: {name: data}
spec:
accessModes: [ReadWriteOnce]
storageClassName: local-path
resources:
requests: {storage: 20Gi}
---
apiVersion: v1
kind: Service
metadata:
name: prometheus
namespace: monitoring
spec:
selector: {app: prometheus}
ports:
- port: 9090
targetPort: 9090

View File

@@ -0,0 +1,88 @@
apiVersion: v1
kind: Namespace
metadata:
name: web
---
apiVersion: v1
kind: ConfigMap
metadata:
name: hello-index
namespace: web
data:
index.html: |
<!doctype html>
<html>
<head><meta charset="utf-8"><title>carr.pub</title>
<style>
body { font-family: system-ui, sans-serif; background: #111; color: #eee;
display: grid; place-items: center; height: 100vh; margin: 0; }
h1 { font-size: 3rem; margin: 0 0 .5rem; }
p { color: #888; font-size: 1.2rem; }
</style>
</head>
<body>
<div style="text-align:center">
<h1>carr.pub</h1>
<p>Served by the planck cluster — 20 Raspberry Pis that finally earn their rack space.</p>
</div>
</body>
</html>
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: hello-web
namespace: web
spec:
replicas: 6
selector:
matchLabels: {app: hello-web}
template:
metadata:
labels: {app: hello-web}
spec:
containers:
- name: web
image: nginx:1.29-alpine
ports: [{containerPort: 80}]
resources:
requests: {cpu: 10m, memory: 16Mi}
limits: {memory: 64Mi}
volumeMounts:
- {name: html, mountPath: /usr/share/nginx/html}
volumes:
- name: html
configMap: {name: hello-index}
---
apiVersion: v1
kind: Service
metadata:
name: hello-web
namespace: web
spec:
selector: {app: hello-web}
ports:
- port: 80
targetPort: 80
---
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: carr-pub
namespace: web
annotations:
cert-manager.io/cluster-issuer: letsencrypt-prod
spec:
tls:
- hosts: [carr.pub]
secretName: carr-pub-tls
rules:
- host: carr.pub
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: hello-web
port: {number: 80}

24
scripts/deploy-admin.sh Normal file
View File

@@ -0,0 +1,24 @@
#!/bin/bash
# Redeploy the admin stack (Authelia SSO + portal). Secrets are regenerated
# from the gitignored files in configs/ so nothing sensitive is committed.
# First-time setup: run scripts/gen-admin-secrets.sh before this.
set -euo pipefail
cd "$(dirname "$0")/.."
HASH=$(/tmp/opencode/authelia/authelia crypto hash generate argon2 \
--password "$(cat configs/admin-password.txt)" 2>/dev/null | grep -o '\$argon2id\$.*')
SESSION=$(cat configs/authelia-session-secret.txt)
JWT=$(cat configs/authelia-jwt-secret.txt)
ENCKEY=$(openssl rand -base64 24 | tr -d "=+/" | cut -c1-32)
export HASH SESSION JWT ENCKEY
for f in manifests/admin/*.yaml; do
envsubst '\$HASH \$SESSION \$JWT \$ENCKEY' < "$f" | kubectl apply -f -
done
kubectl -n admin set env deployment/authelia \
AUTHELIA_SESSION_SECRET="$SESSION" \
AUTHELIA_STORAGE_ENCRYPTION_KEY="$ENCKEY" \
AUTHELIA_IDENTITY_VALIDATION_RESET_PASSWORD_JWT_SECRET="$JWT" \
>/dev/null
echo "admin stack deployed"

32
scripts/k3s-join.sh Executable file
View File

@@ -0,0 +1,32 @@
#!/bin/bash
# Joins the remaining 19 nodes to the k3s cluster (planck002 = first server).
# Usage: K3S_TOKEN=<node-token> ./scripts/k3s-join.sh
set -euo pipefail
SERVER_URL="https://192.168.0.110:6443" # planck002
SERVERS=(planck003 planck004)
AGENTS=(planck001 planck005 planck006 planck007 planck008 planck009 planck010
planck011 planck012 planck013 planck014 planck015 planck016
planck017 planck018 planck019 planck020)
[ -n "${K3S_TOKEN:-}" ] || { echo "K3S_TOKEN env required"; exit 1; }
export K3S_TOKEN
join_server() {
ssh -o BatchMode=yes adamcarr@$1.local \
"curl -sfL https://get.k3s.io | sudo K3S_TOKEN='$K3S_TOKEN' sh -s - server --server '$SERVER_URL' --write-kubeconfig-mode 644" \
2>&1 | grep -viE "rptl.io|newuser" | tail -1
echo "joined server: $1"
}
join_agent() {
ssh -o BatchMode=yes adamcarr@$1.local \
"curl -sfL https://get.k3s.io | sudo K3S_URL='$SERVER_URL' K3S_TOKEN='$K3S_TOKEN' sh -s -" \
2>&1 | grep -viE "rptl.io|newuser" | tail -1
echo "joined agent: $1"
}
for s in "${SERVERS[@]}"; do join_server "$s"; done
for a in "${AGENTS[@]}"; do join_agent "$a" & done
wait
echo "all join commands issued"

47
scripts/loadtest.sh Executable file
View File

@@ -0,0 +1,47 @@
#!/bin/bash
# Spawn a k6 load swarm against a target.
# Usage: ./scripts/loadtest.sh <target-url> [pods] [vus-per-pod] [duration]
# Example: ./scripts/loadtest.sh https://staging.example.com 20 50 5m
# -> 20 pods x 50 virtual users = 1000 concurrent requests, 5 minutes
set -euo pipefail
TARGET=${1:?usage: loadtest.sh <target-url> [pods] [vus-per-pod] [duration]}
PODS=${2:-20}
VUS=${3:-20}
DURATION=${4:-60s}
NAME="loadtest-$(date +%s)"
kubectl -n loadtest apply -f - <<EOF
apiVersion: batch/v1
kind: Job
metadata:
name: $NAME
spec:
parallelism: $PODS
completions: $PODS
ttlSecondsAfterFinished: 300
template:
spec:
restartPolicy: Never
containers:
- name: k6
image: grafana/k6:1.4.0
command: ["k6", "run", "/scripts/test.js"]
env:
- {name: TARGET_URL, value: "$TARGET"}
- {name: VUS, value: "$VUS"}
- {name: DURATION, value: "$DURATION"}
resources:
requests: {cpu: 100m, memory: 128Mi}
limits: {memory: 256Mi}
volumeMounts:
- {name: scripts, mountPath: /scripts}
volumes:
- name: scripts
configMap: {name: k6-script}
EOF
echo "Swarm '$NAME' launched: $PODS pods x $VUS vus for $DURATION against $TARGET"
echo "Watch: kubectl -n loadtest get pods -w"
echo "Logs: kubectl -n loadtest logs -f job/$NAME --all-pods=true --tail=5"
echo "Stop: kubectl -n loadtest delete job $NAME"

61
scripts/stepload.sh Executable file
View File

@@ -0,0 +1,61 @@
#!/bin/bash
# Run a step-load (staircase) test across a collection of endpoints.
# Usage:
# ./scripts/stepload.sh <base-url> <endpoints-json> [start-rps] [step-rps] [step-duration] [steps]
# Example:
# ./scripts/stepload.sh https://api.example.com \
# '[{"name":"login","method":"POST","path":"/login"},{"name":"list","path":"/items","weight":3}]' \
# 25 25 2m 6
# -> starts at 25 rps, +25 rps every 2 minutes, 6 steps (ending at 150 rps)
set -euo pipefail
BASE=${1:?usage: stepload.sh <base-url> <endpoints-json> [start-rps] [step-rps] [step-duration] [steps]}
EPS=${2:?endpoints-json required - see script header for format}
START=${3:-50}
STEP=${4:-50}
DUR=${5:-2m}
STEPS=${6:-5}
NAME="stepload-$(date +%s)"
PROM_RW="http://prometheus.monitoring.svc:9090/api/v1/write"
kubectl -n loadtest apply -f - <<EOF
apiVersion: batch/v1
kind: Job
metadata:
name: $NAME
spec:
parallelism: 1
ttlSecondsAfterFinished: 1800
template:
spec:
restartPolicy: Never
containers:
- name: k6
image: grafana/k6:1.4.0
command: ["k6", "run", "/scripts/step.js", "-o", "experimental-prometheus-rw"]
env:
- {name: BASE_URL, value: "$BASE"}
- {name: ENDPOINTS_JSON, value: '$EPS'}
- {name: START_RPS, value: "$START"}
- {name: STEP_RPS, value: "$STEP"}
- {name: STEP_DURATION, value: "$DUR"}
- {name: STEPS, value: "$STEPS"}
- {name: K6_PROMETHEUS_RW_SERVER_URL, value: "$PROM_RW"}
- {name: K6_PROMETHEUS_RW_TREND_STATS, value: "avg,p(50),p(95),p(99),min,max"}
resources:
requests: {cpu: 200m, memory: 256Mi}
limits: {memory: 512Mi}
volumeMounts:
- {name: scripts, mountPath: /scripts}
volumes:
- name: scripts
configMap: {name: k6-step-script}
EOF
MAX=$(( START + (STEPS - 1) * STEP ))
echo "Staircase '$NAME': $STEPS steps, $START -> $MAX rps (+$STEP each), $DUR holds"
echo "Total duration: $STEPS x $DUR (plus 5s ramps per step)"
echo ""
echo "Watch live: Grafana -> the 'k6 Step Load' dashboard (or kubectl -n loadtest logs -f job/$NAME)"
echo "Logs: kubectl -n loadtest logs -f job/$NAME"
echo "Stop early: kubectl -n loadtest delete job $NAME"

63
scripts/tokamak-load.sh Executable file
View File

@@ -0,0 +1,63 @@
#!/bin/bash
# Staircase load test against the tokamak public API (prod or dev).
# Usage: ./scripts/tokamak-load.sh <start-rps> <step-rps> <step-duration> <steps> [env]
# Example (gentle first run):
# ./scripts/tokamak-load.sh 2 2 1m 4
set -euo pipefail
START=${1:?start rps required}
STEP=${2:?step rps required}
DUR=${3:?step duration required, e.g. 2m}
STEPS=${4:?number of steps required}
ENVIRONMENT=${5:-prod}
NAME="tokamak-$(date +%s)"
PROM_RW="http://prometheus.monitoring.svc:9090/api/v1/write"
case "$ENVIRONMENT" in
prod) BASE="https://api.blueocean.ai/public/v1" ;;
dev) BASE="https://api.dev.blueocean.ai/public/v1" ;;
*) echo "env must be prod or dev"; exit 1 ;;
esac
kubectl -n loadtest apply -f - <<EOF
apiVersion: batch/v1
kind: Job
metadata:
name: $NAME
spec:
parallelism: 1
ttlSecondsAfterFinished: 1800
template:
spec:
restartPolicy: Never
containers:
- name: k6
image: grafana/k6:1.4.0
command: ["k6", "run", "/scripts/tokamak.js", "-o", "experimental-prometheus-rw", "--tag", "testid=$NAME"]
env:
- {name: BASE_URL, value: "$BASE"}
- {name: START_RPS, value: "$START"}
- {name: STEP_RPS, value: "$STEP"}
- {name: STEP_DURATION, value: "$DUR"}
- {name: STEPS, value: "$STEPS"}
- {name: K6_PROMETHEUS_RW_SERVER_URL, value: "$PROM_RW"}
- {name: K6_PROMETHEUS_RW_TREND_STATS, value: "avg,p(50),p(95),p(99),min,max"}
- name: TOKAMAK_API_KEY
valueFrom: {secretKeyRef: {name: tokamak-api-key, key: api-key}}
resources:
requests: {cpu: 200m, memory: 256Mi}
limits: {memory: 1Gi}
volumeMounts:
- {name: scripts, mountPath: /scripts}
volumes:
- name: scripts
configMap: {name: k6-tokamak-script}
EOF
MAX=$(( START + (STEPS - 1) * STEP ))
echo "tokamak staircase '$NAME' against $ENVIRONMENT ($BASE)"
echo " $STEPS steps: $START -> $MAX rps (+$STEP each), $DUR holds"
echo " auth: k8s secret tokamak-api-key"
echo "Watch: Grafana -> 'k6 Step Load' dashboard (filter testid=$NAME)"
echo "Logs: kubectl -n loadtest logs -f job/$NAME"
echo "Stop: kubectl -n loadtest delete job $NAME"