Running services, network usage, memory usage, bandwidth, disk I/O, successful logins, whether the thing is even alive, etc…
So far my only method has been “hope and pray”.
Running services, network usage, memory usage, bandwidth, disk I/O, successful logins, whether the thing is even alive, etc…
So far my only method has been “hope and pray”.
Cockpit on each server (microos) and bezel for a lite centralized monitoring/alert stack. I have a bespoke gitops on each server running ansible hourly, and ntfy pings for failure on these. Also healthchecks.Io for heartbeat/backups, other critical infra checks