Running services, network usage, memory usage, bandwidth, disk I/O, successful logins, whether the thing is even alive, etc…
So far my only method has been “hope and pray”.

      • bertlewirth@lemmy.world
        link
        fedilink
        English
        arrow-up
        1
        ·
        5 minutes ago

        Well, it’s a good learning experience. But…… don’t. Good luck getting around the major email providers’ smart screen shit. It’s probably not gonna happen. I tried a few years ago, and even microslop couldn’t figure out why their shit screen crap wouldn’t allow my emails. It’s not in there best interest to fix it. So they don’t. But it’s worth doing it in the short term to learn how to do it. Teaches you about email, spf, dkim etc. which actually does translate to real world skills. I used what I learnt setting it up in my real life job.

  • Decronym@lemmy.decronym.xyzB
    link
    fedilink
    English
    arrow-up
    1
    ·
    edit-2
    3 minutes ago

    Acronyms, initialisms, abbreviations, contractions, and other phrases which expand to something larger, that I’ve seen in this thread:

    Fewer Letters More Letters
    AP WiFi Access Point
    DNS Domain Name Service/System
    LXC Linux Containers

    3 acronyms in this thread; the most compressed thread commented on today has 6 acronyms.

    [Thread #122 for this comm, first seen 9th Oct 2026, 13:00] [FAQ] [Full list] [Contact] [Source code]

  • Dultas@lemmy.world
    link
    fedilink
    English
    arrow-up
    1
    ·
    1 hour ago

    Uptime Kuma for outages, Prometheus for metrics, rsyslog to Loki backend for logs. Grafana for ingesting Prometheus and Loki data.

  • dimjim@sh.itjust.works
    link
    fedilink
    English
    arrow-up
    3
    ·
    3 hours ago

    My primary monitoring system is Beszel, but since I just migrated from VMs to LXCs it can’t properly see the memory/CPU usage, so it just checks for disk usage and liveliness. Home assistant with Proxmox & Dockhand integrations gives me the rest of my monitoring needs.

    • sakphul@discuss.tchncs.de
      link
      fedilink
      English
      arrow-up
      1
      ·
      46 minutes ago

      I am also using Beszel after using checkmk. For me checkmk was way too complicated (but is also way more powerfull) to setup and maintain propperly. So I switcheed to beszel because it is much easier and good enough for my use case which is:

      • Is the thing running at all?
      • Disk full?
      • CPU, network, etc. on full throttle for some strange reason?
      • send me a mail if somethings wrong

      Besides that I don’t need anything. The only thing missing is SNMP support to check switches, routers and AP’s.

  • pHr34kY@lemmy.world
    link
    fedilink
    English
    arrow-up
    4
    ·
    4 hours ago

    It’s running the private DNS for my phone. If it goes down, I realise quite quickly.

  • esc@piefed.social
    link
    fedilink
    English
    arrow-up
    4
    ·
    6 hours ago

    Prometheus + grafana. It’s overkill tbh, most of the services restart automatically and I mostly ignore it. (it’s an artifact from life where I cared about it)

  • truxnell@quokk.au
    link
    fedilink
    English
    arrow-up
    1
    ·
    5 hours ago

    Cockpit on each server (microos) and bezel for a lite centralized monitoring/alert stack. I have a bespoke gitops on each server running ansible hourly, and ntfy pings for failure on these. Also healthchecks.Io for heartbeat/backups, other critical infra checks

  • dihutenosa@piefed.social
    link
    fedilink
    English
    arrow-up
    1
    ·
    5 hours ago

    When logged in locally, I use btop to see an overview of what’s happening.

    Other than that, I have relevant Prometheus exporters in every machine (node exporter in all machines, specific exporters by the workload), hooked up over Wireguard to my monitoring solution offsite.

    The phone I actually carry around has a ntfy client talking to ntfy server on the aforementioned monitoring solution, so I get buzzes when something goes down.

    Btw, does anybody happen to know where I could get a pre-cooked comprehensive alert system for my nodes? Surely many people have already written all these rules:

    • if disk space > 80% consumed, send a low-priority alert
    • if disk space > 95% consumed, send an urgent alert
    • … everything else, there’s so much to check…
  • owenfromcanada@lemmy.ca
    link
    fedilink
    English
    arrow-up
    38
    ·
    13 hours ago

    I have a robust monitoring system for my Jellyfin server, been running for the last few years. Checks in periodically, at least once per 24 hours and notifies me if it’s down. Doesn’t use any electricity, but does consume a good amount of Cheerios and mac & cheese.