Grafana Forge A Self-Configuring Grafana Stack
Modern observability stacks are powerful, but getting one running locally can still involve a surprising amount of work.
You install Grafana.
You install Prometheus.
You configure exporters.
You configure dashboards.
You configure datasources.
You configure Loki if you want logs.
You resolve port conflicts.
You discover that one container is unhealthy.
Then you open Grafana and start configuring everything manually.
I wanted something different.
One command. A complete observability environment. Configuration as code. Automatic diagnostics. Automatic repair.
That is the idea behind Grafana Forge and its open-source.
The idea
Grafana Forge is a standalone shell utility that creates and manages a complete Grafana-based observability stack using Docker Compose.
The goal is simple:
./grafana-forge
And instead of asking the developer:
"What should I configure?"
Grafana Forge should answer:
"I detected your environment, configured the stack, started it, and here are the URLs."
The project currently supports:
- Grafana
- Prometheus
- Node Exporter
- cAdvisor
- Loki
- Grafana Alloy
- automatic dashboards
- datasource provisioning
- alert rules
- automatic port selection
- configuration validation
- health diagnostics
- repair and rebuild operations
- backup of generated configuration
The important part is not simply running containers.
The important part is automating the operational knowledge around them.
Why another Grafana setup tool?
There are already excellent tools for running Grafana.
Grafana itself provides provisioning, dashboards, datasources, integrations and a large ecosystem of exporters.
So this project is not trying to replace Grafana.
Instead, Grafana Forge sits one layer above the individual components.
Think of it as a small observability control plane for developers.
Instead of remembering:
How do I configure Prometheus?
Where is my Grafana port?
Did I provision the datasource?
Why is cAdvisor down?
Did my dashboard load?
Is Loki running?
What container is using this port?
The developer interacts with one command:
./grafana-forge doctor
or:
./grafana-forge fix
The architecture
The default architecture looks roughly like this:
┌─────────────────────┐
│ Grafana │
│ Dashboards / UI │
└──────────┬──────────┘
│
┌───────────┴───────────┐
│ │
▼ ▼
┌─────────────┐ ┌─────────────┐
│ Prometheus │ │ Loki │
│ Metrics │ │ Logs │
└──────┬──────┘ └──────┬──────┘
│ ▲
┌─────────┴─────────┐ │
│ │ │
▼ ▼ │
┌─────────────┐ ┌─────────────┐ │
│Node Exporter│ │ cAdvisor │ │
│ Host Metrics│ │ Containers │ │
└─────────────┘ └─────────────┘ │
│
┌─────┴─────┐
│ Alloy │
│ Log Agent │
└───────────┘
Each component has a specific responsibility.
Grafana
The visualization and exploration layer.
Prometheus
The metrics database and query engine.
Node Exporter
Provides Linux host metrics such as:
- CPU
- memory
- filesystem
- disk I/O
- network
- load
cAdvisor
Provides container-level metrics such as:
- CPU usage
- memory usage
- filesystem usage
- network traffic
- container lifecycle information
Loki
Stores logs.
Grafana Alloy
Collects and forwards telemetry, including container logs.
The result is a small but useful observability environment without requiring a manually assembled stack.
One command, multiple profiles
Not every developer needs everything.
Grafana Forge therefore uses profiles.
./grafana-forge up host
Starts host monitoring.
./grafana-forge up host docker
Adds Docker/container monitoring.
./grafana-forge up host docker logs
Adds centralized logs.
The default behavior can also detect the local environment and select appropriate components.
This is important because a monitoring tool should not blindly deploy services that the current machine cannot support.
Automatic port selection
One of the most annoying problems with local observability stacks is port collisions.
You already have something running on:
3000
Grafana wants:
3000
Prometheus wants:
9090
Another project already uses it.
Instead of forcing the developer to edit Compose files, Grafana Forge searches for available ports.
For example:
Grafana: http://127.0.0.1:3000
Prometheus: http://127.0.0.1:9090
Loki: http://127.0.0.1:3100
If those ports are unavailable, the generated environment selects alternatives.
The selected values are persisted in .env.
That means subsequent commands continue using the same ports.
Configuration as code
Another important design decision is that Grafana Forge does not depend on clicking around the Grafana UI.
The stack generates:
compose.yml
prometheus.yml
loki-config.yml
provisioning/
├── datasources/
└── dashboards/
dashboards/
├── forge-home.json
├── forge-server.json
├── forge-docker.json
└── forge-logs.json
alerts/
└── forge.yml
This makes the environment reproducible.
Delete the generated environment.
Run the command again.
The configuration can be reconstructed.
That is much closer to infrastructure-as-code than manually configuring a local Grafana instance.
Automatic dashboards
A monitoring stack is not very useful if the developer still has to spend 20 minutes building dashboards.
Grafana Forge provisions dashboards automatically.
The main dashboards are designed around different operational questions.
Home
A high-level overview.
System Health
CPU
Memory
Disk
Network
Containers
Alerts
Server
Focuses on the underlying machine.
Examples include:
- CPU utilization
- memory utilization
- swap
- filesystem usage
- disk I/O
- network traffic
- system load
Docker
Focuses on containers.
For example:
Container CPU
Container Memory
Network Traffic
Filesystem Usage
Restarts
Container State
Logs
When the logs profile is enabled, Loki and Alloy provide centralized container logs.
The idea is to move from:
docker logs container-name
to:
Grafana → Explore → Loki
Automatic datasource provisioning
Grafana Forge also provisions datasources.
For example, Grafana can communicate with Prometheus internally using:
http://prometheus:9090
and Loki using:
http://loki:3100
Notice that these are Docker service names.
They are not host addresses.
That distinction matters.
From the host:
http://127.0.0.1:9090
Inside the Docker network:
http://prometheus:9090
Grafana Forge keeps those two concepts separate.
That means Grafana does not need to know which host port was automatically selected.
The developer's real problem: something breaks
Starting the stack is easy.
The harder problem is what happens when something doesn't work.
For example:
cadvisor: down
node: up
prometheus: up
A normal Compose workflow might be:
docker compose ps
docker compose logs cadvisor
docker compose inspect cadvisor
docker compose config
Then you start investigating manually.
Grafana Forge introduces:
./grafana-forge doctor
The goal of doctor is to turn an opaque failure into an actionable diagnosis.
Doctor: detect before you repair
The diagnostic workflow checks multiple layers.
Docker
Is Docker available?
Compose
Is Docker Compose available?
Generated configuration
Does the generated Compose configuration validate?
docker compose config
Prometheus
Is Prometheus responding?
/-/ready
Grafana
Is Grafana responding?
/api/health
Prometheus targets
Are exporters actually reachable?
For example:
prometheus: up
node: up
cadvisor: down
Provisioning
Does the datasource configuration exist?
Do dashboard provisioning files exist?
Are the expected dashboards available?
Alert rules
Are the generated alert rules present?
The result becomes a simple health report:
[PASS] Docker
[PASS] Docker Compose
[PASS] Compose configuration
[PASS] Prometheus API
[PASS] Grafana API
[PASS] Datasource provisioning
[PASS] Dashboard provisioning
[FAIL] Prometheus targets
[FAIL] Alert rules
RESULT: UNHEALTHY
That is much more useful than simply seeing:
container exited with code 1
AutoFix
This is where Grafana Forge becomes more than a Compose wrapper.
If the environment is unhealthy, run:
./grafana-forge fix
The repair process can:
- regenerate configuration
- validate Compose
- pull required images
- recreate services
- wait for services to become available
- run diagnostics again
- report the remaining problem
The important design principle is:
Repair should be bounded and observable.
The tool should not continuously restart containers forever.
A repair attempt should have a clear beginning and end.
Normal repair vs hard repair
There is also a difference between fixing a stack and destroying it.
The normal repair command should preserve data:
./grafana-forge fix
This is the safe option.
For example, Prometheus data and Grafana state should not disappear simply because cAdvisor stopped working.
For a complete rebuild:
./grafana-forge fix --hard
The hard mode can:
stop containers
remove generated runtime state
remove volumes
regenerate configuration
pull images
recreate everything
run diagnostics
This gives the developer two clear recovery levels.
fix
↓
safe recovery
fix --hard
↓
clean rebuild
The CLI becomes a developer tool
The final interface is intentionally simple.
./grafana-forge
Start the default environment.
./grafana-forge up host docker logs
Start specific profiles.
./grafana-forge status
Show service status.
./grafana-forge urls
Show access URLs.
./grafana-forge doctor
Diagnose the environment.
./grafana-forge debug
Show detailed diagnostic information.
./grafana-forge config
Show generated configuration.
./grafana-forge fix
Automatically repair the environment.
./grafana-forge fix --hard
Perform a clean rebuild.
./grafana-forge logs grafana
Inspect a service.
./grafana-forge backup
Backup generated configuration.
./grafana-forge down
Stop the stack.
./grafana-forge reset
Remove the generated environment.
The goal is for a developer to remember one command, not a collection of Docker commands.
Grafana plugins
Grafana Forge also supports optional Grafana plugins through configuration.
For example:
FORGE_GRAFANA_PLUGINS=some-plugin-id
The plugin configuration becomes part of the generated Grafana environment.
This is useful because the infrastructure remains reproducible.
Instead of:
"I installed this plugin manually on my laptop."
The environment becomes:
"This environment declares the plugin it needs."
Alerts are generated too
The stack also includes Prometheus alert rules.
Examples include:
Target down
Detect when an exporter or monitoring target stops responding.
High CPU
Detect sustained high CPU usage.
High memory
Detect sustained memory pressure.
Low disk space
Detect filesystem capacity approaching a dangerous threshold.
Container restarting
Detect containers repeatedly restarting.
The purpose is not to build a complete enterprise alerting platform.
The purpose is to provide sensible defaults immediately after installation.
Why the shell script?
A reasonable question is:
Why not build this in Python, Go or Rust?
Because the orchestration layer is fundamentally operating system and Docker oriented.
The shell script can directly interact with:
docker
docker compose
curl
filesystem
environment variables
Unix utilities
It also makes the project extremely portable.
There is no runtime dependency such as:
Python virtual environment
Node.js
npm
compiled binary
The architecture is intentionally simple:
grafana-forge
│
├── Docker
├── Docker Compose
├── generated YAML
├── generated dashboards
└── diagnostics
The script is the control layer.
Docker is the execution layer.
Grafana, Prometheus, Loki and the exporters do the actual observability work.
The interesting engineering problem
The interesting part of Grafana Forge is not:
"I created a Docker Compose file."
That is easy.
The interesting problem is:
How do you make infrastructure behave like a developer tool?
A developer tool should:
- detect its environment
- choose sensible defaults
- avoid unnecessary configuration
- explain failures
- repair common problems
- preserve state when possible
- provide destructive operations explicitly
- be reproducible
- expose useful information immediately
That is the philosophy behind Grafana Forge.
Detect → Explain → Repair → Verify
The overall workflow can be summarized as:
┌─────────────┐
│ Start │
└──────┬──────┘
│
▼
┌─────────────┐
│ Detect │
│ environment │
└──────┬──────┘
│
▼
┌─────────────┐
│ Generate │
│ config │
└──────┬──────┘
│
▼
┌─────────────┐
│ Start │
│ stack │
└──────┬──────┘
│
▼
┌─────────────┐
│ Doctor │
└──────┬──────┘
│
┌──────┴──────┐
│ │
Healthy Unhealthy
│ │
▼ ▼
Done Explain
│
▼
AutoFix
│
▼
Verify
This is the part I think is worth exploring further.
Observability tools should not only show problems.
They should increasingly help developers understand and recover from them.
What I want Grafana Forge to become
The current implementation is intentionally small.
But the direction is larger.
Potential future capabilities include:
Smarter diagnostics
Instead of:
cadvisor: down
provide:
cAdvisor is unhealthy.
Likely cause:
The container failed to start because /dev/kmsg
is unavailable on this host.
Suggested action:
Remove the /dev/kmsg dependency or run cAdvisor
with the required host permissions.
Dependency-aware repair
Instead of blindly restarting everything:
Grafana
↓
Prometheus
↓
cAdvisor
the tool could understand service dependencies and repair only the affected path.
Configuration validation
Before deployment:
✓ Compose syntax
✓ Required environment variables
✓ Port availability
✓ Profiles
✓ Datasources
✓ Dashboard files
✓ Alert rules
More integrations
The same architecture could eventually support:
- PostgreSQL
- Redis
- Nginx
- Kubernetes
- systemd
- cloud infrastructure
- application-specific exporters
without turning Grafana Forge into another huge platform.
The bigger idea
Grafana Forge started from a simple frustration:
Setting up observability should not require becoming an expert in the observability stack before you can observe your application.
Grafana, Prometheus, Loki and Alloy are powerful individually.
The opportunity is to combine them into a developer experience that feels much simpler.
./grafana-forge
should be enough to go from:
Nothing
to:
Metrics
Logs
Dashboards
Alerts
Diagnostics
And when something goes wrong:
./grafana-forge doctor
should tell you what happened.
Then:
./grafana-forge fix
should attempt to recover it.
That is the idea behind Grafana Forge:
One command to build it.
One command to understand it.
One command to repair it.