Running services, network usage, memory usage, bandwidth, disk I/O, successful logins, whether the thing is even alive, etc…
So far my only method has been “hope and pray”.
You guys are monitoring your servers?
I plan to perhaps, just maybe, I am not sure, to try self-hosting an email server. But I’ll need a good uptime.
Although I am really scared about security.Well, it’s a good learning experience. But…… don’t. Good luck getting around the major email providers’ smart screen shit. It’s probably not gonna happen. I tried a few years ago, and even microslop couldn’t figure out why their shit screen crap wouldn’t allow my emails. It’s not in there best interest to fix it. So they don’t. But it’s worth doing it in the short term to learn how to do it. Teaches you about email, spf, dkim etc. which actually does translate to real world skills. I used what I learnt setting it up in my real life job.
“can’t access this thing, server must be down”
I notice something stopped working, or someone in the household notices.
Then I go fix it.
Acronyms, initialisms, abbreviations, contractions, and other phrases which expand to something larger, that I’ve seen in this thread:
Fewer Letters More Letters AP WiFi Access Point DNS Domain Name Service/System LXC Linux Containers
3 acronyms in this thread; the most compressed thread commented on today has 6 acronyms.
[Thread #122 for this comm, first seen 9th Oct 2026, 13:00] [FAQ] [Full list] [Contact] [Source code]
Uptime Kuma for outages, Prometheus for metrics, rsyslog to Loki backend for logs. Grafana for ingesting Prometheus and Loki data.
My primary monitoring system is Beszel, but since I just migrated from VMs to LXCs it can’t properly see the memory/CPU usage, so it just checks for disk usage and liveliness. Home assistant with Proxmox & Dockhand integrations gives me the rest of my monitoring needs.
I am also using Beszel after using checkmk. For me checkmk was way too complicated (but is also way more powerfull) to setup and maintain propperly. So I switcheed to beszel because it is much easier and good enough for my use case which is:
- Is the thing running at all?
- Disk full?
- CPU, network, etc. on full throttle for some strange reason?
- send me a mail if somethings wrong
Besides that I don’t need anything. The only thing missing is SNMP support to check switches, routers and AP’s.
It’s running the private DNS for my phone. If it goes down, I realise quite quickly.
That’s the neat part, I don’t.
Zabbix
Prometheus + grafana. It’s overkill tbh, most of the services restart automatically and I mostly ignore it. (it’s an artifact from life where I cared about it)
Exactly this scenario at my place as well. 😄
I use nagios
Cockpit on each server (microos) and bezel for a lite centralized monitoring/alert stack. I have a bespoke gitops on each server running ansible hourly, and ntfy pings for failure on these. Also healthchecks.Io for heartbeat/backups, other critical infra checks
When logged in locally, I use
btopto see an overview of what’s happening.Other than that, I have relevant Prometheus exporters in every machine (node exporter in all machines, specific exporters by the workload), hooked up over Wireguard to my monitoring solution offsite.
The phone I actually carry around has a
ntfyclient talking tontfyserver on the aforementioned monitoring solution, so I get buzzes when something goes down.Btw, does anybody happen to know where I could get a pre-cooked comprehensive alert system for my nodes? Surely many people have already written all these rules:
- if disk space > 80% consumed, send a low-priority alert
- if disk space > 95% consumed, send an urgent alert
- … everything else, there’s so much to check…
I have a robust monitoring system for my Jellyfin server, been running for the last few years. Checks in periodically, at least once per 24 hours and notifies me if it’s down. Doesn’t use any electricity, but does consume a good amount of Cheerios and mac & cheese.
Does this monitoring system also happen to be a dependent on your taxes?
Lies! It uses electricity!
Indirectly. But as another commenter pointed out, it balances out with the tax rebates.
A script checks container status and accessibility, and pushes a ntfy alert if it goes down. Healthchecks.io reports if the whole Pi goes down.
I usually just walk down the hall and move the mouse. I don’t really need remote monitoring. 😅








