Grafana Forge A Self-Configuring Grafana Stack

Modern observability stacks are powerful, but getting one running locally can still involve a surprising amount of work.

You install Grafana.
You install Prometheus.
You configure exporters.
You configure dashboards.
You configure datasources.
You configure Loki if you want logs.
You resolve port conflicts.
You discover that one container is unhealthy.
Then you open Grafana and start configuring everything manually.

I wanted something different.

One command. A complete observability environment. Configuration as code. Automatic diagnostics. Automatic repair.

That is the idea behind Grafana Forge and its open-source.


The idea

Grafana Forge is a standalone shell utility that creates and manages a complete Grafana-based observability stack using Docker Compose.

The goal is simple:

./grafana-forge

And instead of asking the developer:

"What should I configure?"

Grafana Forge should answer:

"I detected your environment, configured the stack, started it, and here are the URLs."

The project currently supports:

  • Grafana
  • Prometheus
  • Node Exporter
  • cAdvisor
  • Loki
  • Grafana Alloy
  • automatic dashboards
  • datasource provisioning
  • alert rules
  • automatic port selection
  • configuration validation
  • health diagnostics
  • repair and rebuild operations
  • backup of generated configuration

The important part is not simply running containers.

The important part is automating the operational knowledge around them.


Why another Grafana setup tool?

There are already excellent tools for running Grafana.

Grafana itself provides provisioning, dashboards, datasources, integrations and a large ecosystem of exporters.

So this project is not trying to replace Grafana.

Instead, Grafana Forge sits one layer above the individual components.

Think of it as a small observability control plane for developers.

Instead of remembering:

How do I configure Prometheus?
Where is my Grafana port?
Did I provision the datasource?
Why is cAdvisor down?
Did my dashboard load?
Is Loki running?
What container is using this port?

The developer interacts with one command:

./grafana-forge doctor

or:

./grafana-forge fix

The architecture

The default architecture looks roughly like this:

                     ┌─────────────────────┐
                     │      Grafana        │
                     │   Dashboards / UI   │
                     └──────────┬──────────┘
                                │
                    ┌───────────┴───────────┐
                    │                       │
                    ▼                       ▼
             ┌─────────────┐         ┌─────────────┐
             │ Prometheus  │         │    Loki     │
             │   Metrics   │         │    Logs     │
             └──────┬──────┘         └──────┬──────┘
                    │                       ▲
          ┌─────────┴─────────┐             │
          │                   │             │
          ▼                   ▼             │
   ┌─────────────┐     ┌─────────────┐      │
   │Node Exporter│     │  cAdvisor   │      │
   │ Host Metrics│     │ Containers  │      │
   └─────────────┘     └─────────────┘      │
                                            │
                                      ┌─────┴─────┐
                                      │   Alloy   │
                                      │ Log Agent │
                                      └───────────┘

Each component has a specific responsibility.

Grafana

The visualization and exploration layer.

Prometheus

The metrics database and query engine.

Node Exporter

Provides Linux host metrics such as:

  • CPU
  • memory
  • filesystem
  • disk I/O
  • network
  • load

cAdvisor

Provides container-level metrics such as:

  • CPU usage
  • memory usage
  • filesystem usage
  • network traffic
  • container lifecycle information

Loki

Stores logs.

Grafana Alloy

Collects and forwards telemetry, including container logs.

The result is a small but useful observability environment without requiring a manually assembled stack.


One command, multiple profiles

Not every developer needs everything.

Grafana Forge therefore uses profiles.

./grafana-forge up host

Starts host monitoring.

./grafana-forge up host docker

Adds Docker/container monitoring.

./grafana-forge up host docker logs

Adds centralized logs.

The default behavior can also detect the local environment and select appropriate components.

This is important because a monitoring tool should not blindly deploy services that the current machine cannot support.


Automatic port selection

One of the most annoying problems with local observability stacks is port collisions.

You already have something running on:

3000

Grafana wants:

3000

Prometheus wants:

9090

Another project already uses it.

Instead of forcing the developer to edit Compose files, Grafana Forge searches for available ports.

For example:

Grafana:    http://127.0.0.1:3000
Prometheus: http://127.0.0.1:9090
Loki:       http://127.0.0.1:3100

If those ports are unavailable, the generated environment selects alternatives.

The selected values are persisted in .env.

That means subsequent commands continue using the same ports.


Configuration as code

Another important design decision is that Grafana Forge does not depend on clicking around the Grafana UI.

The stack generates:

compose.yml
prometheus.yml
loki-config.yml

provisioning/
├── datasources/
└── dashboards/

dashboards/
├── forge-home.json
├── forge-server.json
├── forge-docker.json
└── forge-logs.json

alerts/
└── forge.yml

This makes the environment reproducible.

Delete the generated environment.

Run the command again.

The configuration can be reconstructed.

That is much closer to infrastructure-as-code than manually configuring a local Grafana instance.


Automatic dashboards

A monitoring stack is not very useful if the developer still has to spend 20 minutes building dashboards.

Grafana Forge provisions dashboards automatically.

The main dashboards are designed around different operational questions.

Home

A high-level overview.

System Health
CPU
Memory
Disk
Network
Containers
Alerts

Server

Focuses on the underlying machine.

Examples include:

  • CPU utilization
  • memory utilization
  • swap
  • filesystem usage
  • disk I/O
  • network traffic
  • system load

Docker

Focuses on containers.

For example:

Container CPU
Container Memory
Network Traffic
Filesystem Usage
Restarts
Container State

Logs

When the logs profile is enabled, Loki and Alloy provide centralized container logs.

The idea is to move from:

docker logs container-name

to:

Grafana → Explore → Loki

Automatic datasource provisioning

Grafana Forge also provisions datasources.

For example, Grafana can communicate with Prometheus internally using:

http://prometheus:9090

and Loki using:

http://loki:3100

Notice that these are Docker service names.

They are not host addresses.

That distinction matters.

From the host:

http://127.0.0.1:9090

Inside the Docker network:

http://prometheus:9090

Grafana Forge keeps those two concepts separate.

That means Grafana does not need to know which host port was automatically selected.


The developer's real problem: something breaks

Starting the stack is easy.

The harder problem is what happens when something doesn't work.

For example:

cadvisor: down
node: up
prometheus: up

A normal Compose workflow might be:

docker compose ps
docker compose logs cadvisor
docker compose inspect cadvisor
docker compose config

Then you start investigating manually.

Grafana Forge introduces:

./grafana-forge doctor

The goal of doctor is to turn an opaque failure into an actionable diagnosis.


Doctor: detect before you repair

The diagnostic workflow checks multiple layers.

Docker

Is Docker available?

Compose

Is Docker Compose available?

Generated configuration

Does the generated Compose configuration validate?

docker compose config

Prometheus

Is Prometheus responding?

/-/ready

Grafana

Is Grafana responding?

/api/health

Prometheus targets

Are exporters actually reachable?

For example:

prometheus: up
node: up
cadvisor: down

Provisioning

Does the datasource configuration exist?

Do dashboard provisioning files exist?

Are the expected dashboards available?

Alert rules

Are the generated alert rules present?

The result becomes a simple health report:

[PASS] Docker
[PASS] Docker Compose
[PASS] Compose configuration
[PASS] Prometheus API
[PASS] Grafana API
[PASS] Datasource provisioning
[PASS] Dashboard provisioning
[FAIL] Prometheus targets
[FAIL] Alert rules

RESULT: UNHEALTHY

That is much more useful than simply seeing:

container exited with code 1

AutoFix

This is where Grafana Forge becomes more than a Compose wrapper.

If the environment is unhealthy, run:

./grafana-forge fix

The repair process can:

  1. regenerate configuration
  2. validate Compose
  3. pull required images
  4. recreate services
  5. wait for services to become available
  6. run diagnostics again
  7. report the remaining problem

The important design principle is:

Repair should be bounded and observable.

The tool should not continuously restart containers forever.

A repair attempt should have a clear beginning and end.


Normal repair vs hard repair

There is also a difference between fixing a stack and destroying it.

The normal repair command should preserve data:

./grafana-forge fix

This is the safe option.

For example, Prometheus data and Grafana state should not disappear simply because cAdvisor stopped working.

For a complete rebuild:

./grafana-forge fix --hard

The hard mode can:

stop containers
remove generated runtime state
remove volumes
regenerate configuration
pull images
recreate everything
run diagnostics

This gives the developer two clear recovery levels.

fix
    ↓
safe recovery

fix --hard
    ↓
clean rebuild

The CLI becomes a developer tool

The final interface is intentionally simple.

./grafana-forge

Start the default environment.

./grafana-forge up host docker logs

Start specific profiles.

./grafana-forge status

Show service status.

./grafana-forge urls

Show access URLs.

./grafana-forge doctor

Diagnose the environment.

./grafana-forge debug

Show detailed diagnostic information.

./grafana-forge config

Show generated configuration.

./grafana-forge fix

Automatically repair the environment.

./grafana-forge fix --hard

Perform a clean rebuild.

./grafana-forge logs grafana

Inspect a service.

./grafana-forge backup

Backup generated configuration.

./grafana-forge down

Stop the stack.

./grafana-forge reset

Remove the generated environment.

The goal is for a developer to remember one command, not a collection of Docker commands.


Grafana plugins

Grafana Forge also supports optional Grafana plugins through configuration.

For example:

FORGE_GRAFANA_PLUGINS=some-plugin-id

The plugin configuration becomes part of the generated Grafana environment.

This is useful because the infrastructure remains reproducible.

Instead of:

"I installed this plugin manually on my laptop."

The environment becomes:

"This environment declares the plugin it needs."

Alerts are generated too

The stack also includes Prometheus alert rules.

Examples include:

Target down

Detect when an exporter or monitoring target stops responding.

High CPU

Detect sustained high CPU usage.

High memory

Detect sustained memory pressure.

Low disk space

Detect filesystem capacity approaching a dangerous threshold.

Container restarting

Detect containers repeatedly restarting.

The purpose is not to build a complete enterprise alerting platform.

The purpose is to provide sensible defaults immediately after installation.


Why the shell script?

A reasonable question is:

Why not build this in Python, Go or Rust?

Because the orchestration layer is fundamentally operating system and Docker oriented.

The shell script can directly interact with:

docker
docker compose
curl
filesystem
environment variables
Unix utilities

It also makes the project extremely portable.

There is no runtime dependency such as:

Python virtual environment
Node.js
npm
compiled binary

The architecture is intentionally simple:

grafana-forge
       │
       ├── Docker
       ├── Docker Compose
       ├── generated YAML
       ├── generated dashboards
       └── diagnostics

The script is the control layer.

Docker is the execution layer.

Grafana, Prometheus, Loki and the exporters do the actual observability work.


The interesting engineering problem

The interesting part of Grafana Forge is not:

"I created a Docker Compose file."

That is easy.

The interesting problem is:

How do you make infrastructure behave like a developer tool?

A developer tool should:

  • detect its environment
  • choose sensible defaults
  • avoid unnecessary configuration
  • explain failures
  • repair common problems
  • preserve state when possible
  • provide destructive operations explicitly
  • be reproducible
  • expose useful information immediately

That is the philosophy behind Grafana Forge.


Detect → Explain → Repair → Verify

The overall workflow can be summarized as:

                 ┌─────────────┐
                 │    Start    │
                 └──────┬──────┘
                        │
                        ▼
                 ┌─────────────┐
                 │   Detect    │
                 │ environment │
                 └──────┬──────┘
                        │
                        ▼
                 ┌─────────────┐
                 │  Generate   │
                 │    config   │
                 └──────┬──────┘
                        │
                        ▼
                 ┌─────────────┐
                 │    Start    │
                 │    stack    │
                 └──────┬──────┘
                        │
                        ▼
                 ┌─────────────┐
                 │    Doctor   │
                 └──────┬──────┘
                        │
                 ┌──────┴──────┐
                 │             │
              Healthy       Unhealthy
                 │             │
                 ▼             ▼
              Done          Explain
                               │
                               ▼
                            AutoFix
                               │
                               ▼
                            Verify

This is the part I think is worth exploring further.

Observability tools should not only show problems.

They should increasingly help developers understand and recover from them.


What I want Grafana Forge to become

The current implementation is intentionally small.

But the direction is larger.

Potential future capabilities include:

Smarter diagnostics

Instead of:

cadvisor: down

provide:

cAdvisor is unhealthy.

Likely cause:
The container failed to start because /dev/kmsg
is unavailable on this host.

Suggested action:
Remove the /dev/kmsg dependency or run cAdvisor
with the required host permissions.

Dependency-aware repair

Instead of blindly restarting everything:

Grafana
    ↓
Prometheus
    ↓
cAdvisor

the tool could understand service dependencies and repair only the affected path.

Configuration validation

Before deployment:

✓ Compose syntax
✓ Required environment variables
✓ Port availability
✓ Profiles
✓ Datasources
✓ Dashboard files
✓ Alert rules

More integrations

The same architecture could eventually support:

  • PostgreSQL
  • Redis
  • Nginx
  • Kubernetes
  • systemd
  • cloud infrastructure
  • application-specific exporters

without turning Grafana Forge into another huge platform.


The bigger idea

Grafana Forge started from a simple frustration:

Setting up observability should not require becoming an expert in the observability stack before you can observe your application.

Grafana, Prometheus, Loki and Alloy are powerful individually.

The opportunity is to combine them into a developer experience that feels much simpler.

./grafana-forge

should be enough to go from:

Nothing

to:

Metrics
Logs
Dashboards
Alerts
Diagnostics

And when something goes wrong:

./grafana-forge doctor

should tell you what happened.

Then:

./grafana-forge fix

should attempt to recover it.

That is the idea behind Grafana Forge:

One command to build it.
One command to understand it.
One command to repair it.