Prometheus scrapes metrics from everything that exposes them and stores them as a time series. Grafana turns that into dashboards and alerts. Together they're the observability stack that lets you answer "is it slow because of CPU, memory, or Postgres" with a graph instead of a hunch.
What it is
Prometheus is a pull-based scraper: you give it a list of URLs and an interval, and it stores what it finds. Grafana is the query and visualisation layer, with a large community dashboard library so you rarely build from scratch. The node exporter feeds it host metrics, which is where most of the immediate value is.
Before you start
- Retention is disk.
--storage.tsdb.retention.time=30don a 2 GB host will fill a volume. Start at 7 days and add anode_exporterbefore you add anything heavy. - Pin major versions. Prometheus 3.x will not read a 2.x data directory, and it will refuse to start rather than migrate it for you.
- Set the admin password through an environment variable, not the first-run wizard, or you'll be prompted for it every time the config is reset.
1 — Create the Prometheus config
mkdir -p ~/services/monitoring/{prometheus,grafana}
cd ~/services/monitoring
cat > prometheus/prometheus.yml <<'EOF'
global:
scrape_interval: 30s
evaluation_interval: 30s
external_labels:
monitor: trinity
scrape_configs:
- job_name: prometheus
static_configs:
- targets: ["localhost:9090"]
- job_name: node
static_configs:
- targets: ["node-exporter:9100"]
- job_name: cadvisor
static_configs:
- targets: ["cadvisor:8080"]
- job_name: postgres
static_configs:
- targets: ["postgres-exporter:9187"]
- job_name: containers
static_configs:
- targets: ["docker-exporter:9323"]
EOF
2 — Write the compose file
services:
prometheus:
image: prom/prometheus:v3.1.0
container_name: prometheus
restart: unless-stopped
user: "65534:65534"
ports:
- "127.0.0.1:9090:9090"
volumes:
- ./prometheus/prometheus.yml:/etc/prometheus/prometheus.yml:ro
- ./prometheus/data:/prometheus
command:
- --config.file=/etc/prometheus/prometheus.yml
- --storage.tsdb.path=/prometheus
- --storage.tsdb.retention.time=15d
- --storage.tsdb.retention.size=4GB
- --web.enable-lifecycle
healthcheck:
test: ["CMD", "wget", "--spider", "-q", "http://localhost:9090/-/healthy"]
interval: 30s
timeout: 5s
retries: 3
node-exporter:
image: prom/node-exporter:latest
container_name: node-exporter
restart: unless-stopped
command:
- --path.rootfs=/host
- --collector.filesystem.mount-points-exclude=^/(sys|proc|dev|host|etc)($$|/)
volumes:
- /:/host:ro,rslave
pid: host
healthcheck:
test: ["CMD", "wget", "--spider", "-q", "http://localhost:9100/metrics"]
interval: 30s
timeout: 5s
retries: 3
cadvisor:
image: gcr.io/cadvisor/cadvisor:latest
container_name: cadvisor
restart: unless-stopped
volumes:
- /:/rootfs:ro
- /var/run:/var/run:ro
- /sys:/sys:ro
- /var/lib/docker/:/var/lib/docker:ro
privileged: true
device_cgroup_rules:
- "c 1:3 rwm"
grafana:
image: grafana/grafana:latest
container_name: grafana
restart: unless-stopped
depends_on:
- prometheus
ports:
- "127.0.0.1:3000:3000"
environment:
GF_SECURITY_ADMIN_USER: admin
GF_SECURITY_ADMIN_PASSWORD: CHANGE_ME
GF_USERS_ALLOW_SIGN_UP: "false"
GF_INSTALL_PLUGINS: ""
GF_ANALYTICS_REPORTING_ENABLED: "false"
GF_LOG_LEVEL: warn
volumes:
- ./grafana/data:/var/lib/grafana
- ./grafana/dashboards:/var/lib/grafana/dashboards
healthcheck:
test: ["CMD-SHELL", "wget -qO- http://localhost:3000/api/health | grep -q ok"]
interval: 30s
timeout: 5s
retries: 3
3 — Start it
docker compose up -d
docker compose exec prometheus promtool check config /etc/prometheus/prometheus.yml
curl -s http://127.0.0.1:9090/api/v1/targets | jq '.data.activeTargets[] | {job: .labels.job, health: .health}'
4 — First-run setup
- Log into Grafana at
http://yourhost:3000asadmin, change the password, and turn off signup in Configuration → Users. - Add a Prometheus data source pointing at
http://prometheus:9090. Click Save & Test — it must say "Successfully queried the Prometheus API". - Import the Node Exporter Full dashboard (ID
1860) and the Docker / cAdvisor one (ID893). You now have host and per-container CPU, memory, disk and network without writing a single query. - Build one alert that matters: container filesystem over 85% on the volumes that actually hold data. Route it to the same channel you set up in Uptime Kuma.
- For a quick sanity check, run
node_memory_MemAvailable_byteson the node dashboard and compare it against the container panels — that's how you find a container quietly eating the host.
5 — Keep it from filling the disk
# what is actually taking space?
docker exec prometheus du -sh /prometheus
# reload config without restarting
curl -X POST http://127.0.0.1:9090/-/reload
# trim hard
docker exec prometheus promtool tsdb clean --time=24h \
--dry-run
--storage.tsdb.retention.time and --storage.tsdb.retention.size hits first. On a busy host 15 days can be far more than 4 GB, and on a quiet one it can be far less. Set the size cap and let the time cap be a ceiling, not a promise.