0028. Codify PVE Reboot Recovery Procedures¶
- Status: accepted
- Date: 2026-06-25
- Deciders: lusoris
- Related: PR #1902
Context¶
The PVE rolling-reboot incident left the home API server timing out against its local etcd member, office Longhorn disks unable to salvage at the configured free-space threshold, and Seerr's 4K servers without the recovered default flags. PR #1902 codified the corresponding desired-state recovery settings after office quorum was restored.
Decision¶
Add a bounded per-host etcd endpoint override for k8s-cp-home-1, lower Longhorn storageMinimalAvailablePercentage from 25 to 18, and reconcile both Radarr 4K and Sonarr 4K as is4k: true and isDefault: true in the Seerr bootstrap.
Alternatives considered¶
Leave Longhorn's minimum-free threshold at 25%; rejected because it made the affected 256 GiB office disk unschedulable during recovery. No other alternatives were recorded for the API-server override or Seerr flags.
Consequences¶
argocd-apps/infra/longhorn.yaml now sets the threshold to 18, and apps/media/seerr-config-bootstrap/values.yaml writes and verifies both 4K services as defaults. ansible/inventory/group_vars/all.yml now carries the bounded k8s-cp-home-1 override; subsequent history expanded it from the original office pair to all three current etcd voters.
References¶
- PR #1902 — fix: codify PVE reboot recovery