Running services, network usage, memory usage, bandwidth, disk I/O, successful logins, whether the thing is even alive, etc…
So far my only method has been “hope and pray”.
Running services, network usage, memory usage, bandwidth, disk I/O, successful logins, whether the thing is even alive, etc…
So far my only method has been “hope and pray”.
When logged in locally, I use
btopto see an overview of what’s happening.Other than that, I have relevant Prometheus exporters in every machine (node exporter in all machines, specific exporters by the workload), hooked up over Wireguard to my monitoring solution offsite.
The phone I actually carry around has a
ntfyclient talking tontfyserver on the aforementioned monitoring solution, so I get buzzes when something goes down.Btw, does anybody happen to know where I could get a pre-cooked comprehensive alert system for my nodes? Surely many people have already written all these rules: