Skip to content

0028. Codify PVE Reboot Recovery Procedures

  • Status: accepted
  • Date: 2026-06-25
  • Deciders: lusoris
  • Related: PR #1902

Context

The PVE rolling-reboot incident left the home API server timing out against its local etcd member, office Longhorn disks unable to salvage at the configured free-space threshold, and Seerr's 4K servers without the recovered default flags. PR #1902 codified the corresponding desired-state recovery settings after office quorum was restored.

Decision

Add a bounded per-host etcd endpoint override for k8s-cp-home-1, lower Longhorn storageMinimalAvailablePercentage from 25 to 18, and reconcile both Radarr 4K and Sonarr 4K as is4k: true and isDefault: true in the Seerr bootstrap.

Alternatives considered

Leave Longhorn's minimum-free threshold at 25%; rejected because it made the affected 256 GiB office disk unschedulable during recovery. No other alternatives were recorded for the API-server override or Seerr flags.

Consequences

argocd-apps/infra/longhorn.yaml now sets the threshold to 18, and apps/media/seerr-config-bootstrap/values.yaml writes and verifies both 4K services as defaults. ansible/inventory/group_vars/all.yml now carries the bounded k8s-cp-home-1 override; subsequent history expanded it from the original office pair to all three current etcd voters.

References

  • PR #1902 — fix: codify PVE reboot recovery