

Thanks!
I took a lot of inspiration from here https://github.com/onedr0p/home-ops, there is also a discord server on there where you can ask people for help, they are quite helpful.
I mostly shopped around and found parts and implementations I liked and coppied that and made it my own



See the “monitoring” namespace
kube-prometheus-stack plus VictoriaLogs for logs, about 100 pods on 2 nodes.
The whole stack uses about 2.2 GB RAM. Prometheus takes about 1.1 GB of that and Grafana about 700 MB. CPU is basically idle, around 0.2 cores.
Prometheus has about 150k series and uses about 7 GB for 15 days of metrics, so roughly 14 GB a month. Logs are about 3 GB a month. Grafana itself is under 1 GB.
It’s mostly the default exporters, so a smaller lab would need less. A 20 GB volume would be plenty.
Prometheus is set to retention: 15d and retentionSize: 18GB. Anything older than 15 days is dropped, and so is the oldest data if the total goes over 18 GB, whichever happens first. It’s sitting at a steady ~7 GB now that it has a full 15 days stored.