Running services, network usage, memory usage, bandwidth, disk I/O, successful logins, whether the thing is even alive, etc…
So far my only method has been “hope and pray”.

  • truxnell@quokk.au
    link
    fedilink
    English
    arrow-up
    1
    ·
    16 hours ago

    Cockpit on each server (microos) and bezel for a lite centralized monitoring/alert stack. I have a bespoke gitops on each server running ansible hourly, and ntfy pings for failure on these. Also healthchecks.Io for heartbeat/backups, other critical infra checks