Running services, network usage, memory usage, bandwidth, disk I/O, successful logins, whether the thing is even alive, etc…
So far my only method has been “hope and pray”.

  • dihutenosa@piefed.social
    link
    fedilink
    English
    arrow-up
    1
    ·
    21 hours ago

    When logged in locally, I use btop to see an overview of what’s happening.

    Other than that, I have relevant Prometheus exporters in every machine (node exporter in all machines, specific exporters by the workload), hooked up over Wireguard to my monitoring solution offsite.

    The phone I actually carry around has a ntfy client talking to ntfy server on the aforementioned monitoring solution, so I get buzzes when something goes down.

    Btw, does anybody happen to know where I could get a pre-cooked comprehensive alert system for my nodes? Surely many people have already written all these rules:

    • if disk space > 80% consumed, send a low-priority alert
    • if disk space > 95% consumed, send an urgent alert
    • … everything else, there’s so much to check…