I made a Grafana Master-dashboard to monitor my Kubernetes cluster.

Features:

  1. Health bar at the top, that will turn amber or red when things go wrong, and send push notifications to my phone
  2. CPU, Memory and Network utilization
  3. Database and Storage backup statuses

I also added CPU/Memory usage request monitoring so that I can tune how much CPU/Memory a pod requests in the cluster.
image

And active alerts
image
I’ve muted these two, I need to add a 3rd node with more CPU and ram, currently if one node goes down there isn’t enough redundancy for the cluster to just keep working.

Lots more detail here: https://erasmus.works/ It’s all Open-Source, leave a Star on my repo if you like what you see.

  • Legion739@piefed.zipOP
    link
    fedilink
    English
    arrow-up
    2
    ·
    10 hours ago

    See the “monitoring” namespace

    image

    kube-prometheus-stack plus VictoriaLogs for logs, about 100 pods on 2 nodes.

    The whole stack uses about 2.2 GB RAM. Prometheus takes about 1.1 GB of that and Grafana about 700 MB. CPU is basically idle, around 0.2 cores.

    Prometheus has about 150k series and uses about 7 GB for 15 days of metrics, so roughly 14 GB a month. Logs are about 3 GB a month. Grafana itself is under 1 GB.

    It’s mostly the default exporters, so a smaller lab would need less. A 20 GB volume would be plenty.

    Prometheus is set to retention: 15d and retentionSize: 18GB. Anything older than 15 days is dropped, and so is the oldest data if the total goes over 18 GB, whichever happens first. It’s sitting at a steady ~7 GB now that it has a full 15 days stored.