Back to blog

Monitoring

Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide

Installs Node Exporter and Prometheus, configures scrape intervals, builds Grafana dashboards, and sets up alerting rules.

  • Linux
  • Monitoring
  • Prometheus
  • Grafana
  • Node Exporter
  • Alerting

SEO Metadata

SEO Title Options

  1. Linux Monitoring With Prometheus & Grafana: Full Stack
  2. Linux Monitoring With Prometheus: Practical 2026 Guide
  3. Monitoring Playbook: Linux Monitoring With Prometheus

Meta Description Options

  1. Learn Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide with a practical Monitoring framework, expert mistakes, implementation steps.
  2. Installs Node Exporter and Prometheus, configures scrape intervals, builds Grafana dashboards, and sets up alerting rules.

URL Slug

linux-monitoring-prometheus-grafana-full-stack-setup-guide

Focus Keyword

Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide

Additional LSI Keywords

  • Monitoring
  • Linux
  • Prometheus
  • Grafana
  • Node Exporter
  • Alerting
  • Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide
  • production checklist
  • implementation guide
  • best practices
  • architecture decisions
  • testing strategy

Table of Contents

Article overview

Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide is the kind of topic that looks simple until it reaches production. Teams usually discover the real cost late: unclear boundaries, weak defaults, hidden maintenance work, and decisions that seemed harmless when the codebase was small.

The problem gets worse when the article, tutorial, or implementation guide only explains the happy path. This guide closes that gap with a practical framework, a comparison table, common mistakes, and a deep technical section you can use while planning real work.

Keep reading for the non-obvious part: the safest implementation is rarely the most impressive-looking one. It is the one your team can debug, test, document, and evolve without turning every future change into archaeology.

Key Takeaways

  • Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide should be evaluated as a production decision, not only as a syntax or tooling choice.
  • The best implementation keeps responsibilities visible, with clear ownership, tests, documentation, and rollback paths.
  • Search visibility improves when practical depth, structured answers, and expert examples live on the same page.

[IMAGE: A mobile-first technical article layout showing the main concept, decision table, implementation checklist, and FAQ blocks. Alt: Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide expert guide for Monitoring]

What Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide means

Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide means applying monitoring knowledge to a concrete engineering decision, then turning that decision into reliable code, documentation, and operational behavior. In practice, it combines the topic's core concepts with trade-off analysis, implementation boundaries, testing strategy, and maintenance discipline.

This is the definition worth optimizing for featured snippets because it avoids hype. It tells the reader what the topic does and what a professional implementation must include.

Why it matters now

The technical web is more crowded than it was a few years ago. Thin tutorials can still get indexed, but they rarely earn trust from senior developers, buyers, AI answer systems, or teams that need production guidance.

For monitoring topics, the strongest content now has three layers:

  • a clear answer for fast scanning
  • a practical framework for implementation
  • expert context that explains what breaks later

That same structure helps search engines understand the page. It also helps readers decide whether the advice fits their project.

Implementation framework

Use this framework before adopting the approach described in this article.

  1. Define the user problem and the production risk.
  2. Identify the smallest reliable implementation boundary.
  3. Keep configuration, secrets, and environment-specific behavior outside the article's core logic.
  4. Add tests for the behavior that would hurt if it regressed.
  5. Document the trade-off, not only the final code.
  6. Measure the result with logs, metrics, or user-facing outcomes.
  7. Revisit the decision after real usage exposes edge cases.

The sequence is deliberately conservative. It keeps the work grounded in outcomes instead of novelty.

[IMAGE: A seven-step implementation framework with discovery, boundary design, configuration, tests, documentation, measurement, and iteration. Alt: Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide implementation framework]

Practical comparison

Decision areaStrong approachWeak approachWhy it matters
ScopeSolve one clear problemMix unrelated concernsFocus improves testing and search intent
ArchitecturePut logic in explicit classes or documented boundariesHide behavior in templates or incidental callbacksFuture changes stay easier to review
Data flowPass prepared data into the view or endpointQuery or compute in presentation codeReduces regressions and performance surprises
TestingCover the risky behavior directlyTest only the happy pathCatches production failures earlier
DocumentationExplain trade-offs and limitsRepeat generic definitionsBuilds E-E-A-T and reader trust
OperationsTrack logs, metrics, and rollback stepsShip without measurementMakes the decision reversible

This table is intentionally practical. It gives a reviewer something to check before the implementation becomes expensive to change.

Expert workflow

Expert tip: "Treat Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide as a system boundary. If the next developer cannot find where the decision lives, how it is tested, and when it should be avoided, the implementation is not finished."

A useful workflow is simple:

  • Start with the smallest working example.
  • Add the constraints that exist in your real project.
  • Remove anything that only demonstrates cleverness.
  • Write down the failure modes.
  • Add links to related decisions so future readers can navigate the topic cluster.

That last point matters for both humans and search systems. A single article can answer a question; a cluster proves authority.

Common mistakes

Mistake 1: Copying a pattern without its context

A pattern that works in a small demo can fail in a real application. The missing context is usually data volume, team experience, deployment process, security requirements, or observability.

Before copying the pattern, ask what assumption made it safe in the original example.

Mistake 2: Putting business logic in the wrong layer

This is the fastest way to make future debugging expensive. In Laravel, PHP, and server-rendered websites, presentation should receive prepared data, not discover rules on its own.

Keep decision logic in models, actions, services, policies, requests, jobs, or documented helpers where it can be tested directly.

Mistake 3: Optimizing for novelty instead of maintainability

Newer tools and language features can be valuable. They can also hide simple behavior behind unfamiliar syntax.

Use the option that makes the next production incident easier to understand.

Mistake 4: Publishing without a measurement plan

If the article describes a performance, SEO, security, or architecture improvement, define how success will be checked. Logs, tests, crawl diagnostics, analytics, and user behavior are all stronger than assumptions.

[IMAGE: A common-mistakes board with context loss, wrong layer, novelty bias, and missing measurement highlighted. Alt: Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide common mistakes]

Image placeholders

  • [IMAGE: A concept diagram for Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide with input, decision boundary, implementation, tests, and production feedback. Alt: Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide concept diagram]
  • [IMAGE: A mobile screenshot-style checklist for Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide. Alt: Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide mobile checklist]
  • [IMAGE: A comparison table visualization for strong versus weak implementation choices. Alt: Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide comparison table]

Video placeholder

[VIDEO: Insert a 5-8 minute YouTube walkthrough that demonstrates the main decision, the implementation boundary, the test strategy, and the production caveats for Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide.]

Internal linking opportunities

Original Technical Deep Dive

The short version

A practical Linux monitoring stack has four parts:

LayerToolJob
Host metricsNode ExporterExpose CPU, memory, disk, filesystem, network, and kernel metrics on each Linux host
Time series databasePrometheusScrape exporters, store samples, evaluate PromQL and alerting rules
Notification routingAlertmanagerGroup, deduplicate, silence, and route firing alerts
DashboardsGrafanaQuery Prometheus and visualize host health

The baseline topology:

Linux hosts: node_exporter on :9100
Monitoring host: prometheus on :9090
Monitoring host: alertmanager on :9093
Monitoring host: grafana-server on :3000

Use this guide for a small to medium Linux fleet where Prometheus scrapes hosts over a private network. Do not expose exporter, Prometheus, Alertmanager, or Grafana ports directly to the public internet.

This guide was reviewed on May 7, 2026 against the Prometheus and Grafana documentation. The examples use the current Prometheus download versions at that date:

Prometheus: 3.11.3
Node Exporter: 1.11.1
Alertmanager: 0.32.1

Check the official download page before installing on a real host and adjust the version variables.

Decide what to monitor first

Start with host signals that predict outages:

SignalUseful questions
upIs Prometheus able to scrape the target?
CPU saturationIs the host out of CPU or stuck in iowait?
Memory pressureIs available memory low?
Disk spaceWill a filesystem fill soon?
Disk I/OAre disks overloaded or slow?
Network trafficIs a host receiving or sending unusual traffic?
RebootsDid the host restart unexpectedly?

Avoid trying to alert on every metric. Alerts should represent conditions that need action.

Good first alerts:

target cannot be scraped for 5 minutes
root filesystem has less than 10% free space for 15 minutes
memory available is below 10% for 10 minutes
CPU iowait is high for 15 minutes
host rebooted unexpectedly

Dashboards are for investigation. Alerts are for interruption.

Prepare the hosts

These commands assume Ubuntu or Debian. The systemd units work the same on most modern Linux distributions.

Install basic tools:

sudo apt update
sudo apt install -y curl tar wget ca-certificates gnupg

Create dedicated users:

sudo useradd --system --no-create-home --shell /usr/sbin/nologin prometheus
sudo useradd --system --no-create-home --shell /usr/sbin/nologin node_exporter
sudo useradd --system --no-create-home --shell /usr/sbin/nologin alertmanager

If your distribution does not have /usr/sbin/nologin, use /bin/false.

Create directories on the monitoring host:

sudo mkdir -p /etc/prometheus /var/lib/prometheus
sudo mkdir -p /etc/alertmanager /var/lib/alertmanager
sudo chown prometheus:prometheus /etc/prometheus /var/lib/prometheus
sudo chown alertmanager:alertmanager /etc/alertmanager /var/lib/alertmanager
sudo chmod 0750 /etc/prometheus /var/lib/prometheus /etc/alertmanager /var/lib/alertmanager

Install Node Exporter on each Linux host

Node Exporter should run on every Linux server you want to monitor.

Set the version:

NODE_EXPORTER_VERSION=1.11.1

Download and install:

cd /tmp
wget "https://github.com/prometheus/node_exporter/releases/download/v${NODE_EXPORTER_VERSION}/node_exporter-${NODE_EXPORTER_VERSION}.linux-amd64.tar.gz"
tar -xzf "node_exporter-${NODE_EXPORTER_VERSION}.linux-amd64.tar.gz"
sudo install -m 0755 "node_exporter-${NODE_EXPORTER_VERSION}.linux-amd64/node_exporter" /usr/local/bin/node_exporter

Create the systemd service:

sudoedit /etc/systemd/system/node_exporter.service

Use:

[Unit]
Description=Prometheus Node Exporter
Documentation=https://prometheus.io/docs/guides/node-exporter/
After=network-online.target
Wants=network-online.target

[Service]
User=node_exporter
Group=node_exporter
Type=simple
ExecStart=/usr/local/bin/node_exporter \
  --web.listen-address=:9100
Restart=on-failure
RestartSec=5s
NoNewPrivileges=true
ProtectSystem=full
ProtectHome=true
PrivateTmp=true

[Install]
WantedBy=multi-user.target

Start it:

sudo systemctl daemon-reload
sudo systemctl enable --now node_exporter
sudo systemctl status node_exporter --no-pager

Verify metrics locally:

curl -s http://127.0.0.1:9100/metrics | head
curl -s http://127.0.0.1:9100/metrics | grep '^node_' | head

Open port 9100 only from the Prometheus server:

sudo ufw allow from 10.0.0.10 to any port 9100 proto tcp

Replace 10.0.0.10 with the Prometheus server IP.

Add optional Node Exporter collectors

The default collectors are enough for most hosts. Add collectors only when dashboards or alerts need them.

Common additions:

ExecStart=/usr/local/bin/node_exporter \
  --web.listen-address=:9100 \
  --collector.systemd \
  --collector.processes \
  --collector.textfile.directory=/var/lib/node_exporter/textfile_collector

Create the textfile collector directory if you enable it:

sudo mkdir -p /var/lib/node_exporter/textfile_collector
sudo chown node_exporter:node_exporter /var/lib/node_exporter/textfile_collector
sudo chmod 0755 /var/lib/node_exporter/textfile_collector

[IMAGE: Supporting visual 1 for Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide, showing Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide decisions, examples, and Linux, Monitoring, Prometheus. Alt: Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide linux-monitoring-prometheus-grafana-full-stack-setup-guide visual 1]

[IMAGE: Supporting visual 1 for Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide, showing Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide decisions, examples, and Linux, Monitoring, Prometheus. Alt: Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide linux-monitoring-prometheus-grafana-full-stack-setup-guide visual 1]

Example custom metric:

tmpfile=$(mktemp)
printf 'backup_last_success_timestamp_seconds %s\n' "$(date +%s)" > "$tmpfile"
sudo mv "$tmpfile" /var/lib/node_exporter/textfile_collector/backup.prom
sudo chown node_exporter:node_exporter /var/lib/node_exporter/textfile_collector/backup.prom

Then query:

time() - backup_last_success_timestamp_seconds

This is useful for cron jobs, backups, certificate renewal, and other host-level checks that do not have their own exporter.

Install Prometheus on the monitoring host

Set the version:

PROMETHEUS_VERSION=3.11.3

Download and install:

cd /tmp
wget "https://github.com/prometheus/prometheus/releases/download/v${PROMETHEUS_VERSION}/prometheus-${PROMETHEUS_VERSION}.linux-amd64.tar.gz"
tar -xzf "prometheus-${PROMETHEUS_VERSION}.linux-amd64.tar.gz"

sudo install -m 0755 "prometheus-${PROMETHEUS_VERSION}.linux-amd64/prometheus" /usr/local/bin/prometheus
sudo install -m 0755 "prometheus-${PROMETHEUS_VERSION}.linux-amd64/promtool" /usr/local/bin/promtool
sudo cp -r "prometheus-${PROMETHEUS_VERSION}.linux-amd64/consoles" /etc/prometheus/
sudo cp -r "prometheus-${PROMETHEUS_VERSION}.linux-amd64/console_libraries" /etc/prometheus/
sudo chown -R prometheus:prometheus /etc/prometheus /var/lib/prometheus

Create the config:

sudoedit /etc/prometheus/prometheus.yml

Use:

global:
  scrape_interval: 15s
  scrape_timeout: 10s
  evaluation_interval: 15s
  external_labels:
    environment: production
    monitor: primary

rule_files:
  - /etc/prometheus/rules/*.yml

alerting:
  alertmanagers:
    - static_configs:
        - targets:
            - 127.0.0.1:9093

scrape_configs:
  - job_name: prometheus
    static_configs:
      - targets:
          - 127.0.0.1:9090

  - job_name: node
    scrape_interval: 15s
    static_configs:
      - targets:
          - web-01.internal:9100
          - web-02.internal:9100
          - db-01.internal:9100
        labels:
          role: linux
          site: primary

Create the rules directory:

sudo mkdir -p /etc/prometheus/rules
sudo chown prometheus:prometheus /etc/prometheus/rules
sudo chmod 0750 /etc/prometheus/rules

Validate:

sudo -u prometheus promtool check config /etc/prometheus/prometheus.yml

Run Prometheus with systemd

Create:

sudoedit /etc/systemd/system/prometheus.service

Use:

[Unit]
Description=Prometheus
Documentation=https://prometheus.io/docs/prometheus/latest/getting_started/
After=network-online.target
Wants=network-online.target

[Service]
User=prometheus
Group=prometheus
Type=simple
ExecStart=/usr/local/bin/prometheus \
  --config.file=/etc/prometheus/prometheus.yml \
  --storage.tsdb.path=/var/lib/prometheus \
  --storage.tsdb.retention.time=30d \
  --storage.tsdb.retention.size=40GB \
  --web.console.templates=/etc/prometheus/consoles \
  --web.console.libraries=/etc/prometheus/console_libraries \
  --web.listen-address=127.0.0.1:9090 \
  --web.enable-lifecycle
ExecReload=/bin/kill -HUP $MAINPID
Restart=on-failure
RestartSec=5s
NoNewPrivileges=true
ProtectHome=true
ProtectSystem=full
ReadWritePaths=/var/lib/prometheus
PrivateTmp=true

[Install]
WantedBy=multi-user.target

Start:

sudo systemctl daemon-reload
sudo systemctl enable --now prometheus
sudo systemctl status prometheus --no-pager

Check:

curl -s http://127.0.0.1:9090/-/ready
curl -s http://127.0.0.1:9090/api/v1/targets | head

Prometheus is bound to 127.0.0.1 in this example. Put Nginx, Caddy, SSH tunneling, or VPN access in front of it if administrators need browser access.

Reload Prometheus safely

Validate first:

sudo -u prometheus promtool check config /etc/prometheus/prometheus.yml
sudo -u prometheus promtool check rules /etc/prometheus/rules/*.yml

Reload with systemd:

sudo systemctl reload prometheus

Or use the lifecycle endpoint:

curl -X POST http://127.0.0.1:9090/-/reload

If validation fails, fix the YAML before reloading. Prometheus keeps the old config if a runtime reload receives a malformed config, but your deploy process should still fail before that point.

Understand scrape intervals

Use 15s for hosts unless you have a reason to do otherwise.

IntervalUse caseTradeoff
5sShort-lived spikes, test environmentsMore samples, more storage, more query load
15sGeneral Linux host monitoringGood default
30sLarger fleets or slower linksLess precision
60sLow-value, low-change targetsSlower alert detection

Prometheus storage roughly depends on:

retention_seconds * ingested_samples_per_second * bytes_per_sample

Reducing high-cardinality series usually saves more than increasing the scrape interval. A label like user_id, request_id, session, or full path can explode series count.

Write useful PromQL queries

Host availability:

up{job="node"}

CPU busy percent:

100 - (avg by (instance) (rate(node_cpu_seconds_total{job="node",mode="idle"}[5m])) * 100)

Memory available percent:

node_memory_MemAvailable_bytes{job="node"} / node_memory_MemTotal_bytes{job="node"} * 100

Root filesystem free percent:

node_filesystem_avail_bytes{job="node",mountpoint="/",fstype!~"tmpfs|overlay"} /
node_filesystem_size_bytes{job="node",mountpoint="/",fstype!~"tmpfs|overlay"} * 100

Disk read/write rate:

rate(node_disk_read_bytes_total{job="node"}[5m])
rate(node_disk_written_bytes_total{job="node"}[5m])

Network receive/transmit rate:

rate(node_network_receive_bytes_total{job="node",device!~"lo|veth.*|docker.*"}[5m])
rate(node_network_transmit_bytes_total{job="node",device!~"lo|veth.*|docker.*"}[5m])

Unexpected reboot:

time() - node_boot_time_seconds{job="node"}

Start in Prometheus table view before graphing broad selectors. A bare metric can return thousands of series.

Add recording rules for dashboard queries

Recording rules precompute expensive expressions. Create:

sudoedit /etc/prometheus/rules/node-recording.yml

Use:

groups:
  - name: node-recording
    interval: 30s
    rules:
      - record: instance:node_cpu_busy:ratio
        expr: 1 - avg by (instance) (rate(node_cpu_seconds_total{job="node",mode="idle"}[5m]))

      - record: instance:node_memory_available:ratio
        expr: node_memory_MemAvailable_bytes{job="node"} / node_memory_MemTotal_bytes{job="node"}

      - record: instance:node_filesystem_root_available:ratio
        expr: |
          node_filesystem_avail_bytes{job="node",mountpoint="/",fstype!~"tmpfs|overlay"} /
          node_filesystem_size_bytes{job="node",mountpoint="/",fstype!~"tmpfs|overlay"}

Validate and reload:

sudo -u prometheus promtool check rules /etc/prometheus/rules/node-recording.yml
sudo systemctl reload prometheus

Then dashboards can query:

instance:node_cpu_busy:ratio * 100
instance:node_memory_available:ratio * 100
instance:node_filesystem_root_available:ratio * 100

Install Alertmanager

Set the version:

ALERTMANAGER_VERSION=0.32.1

Download and install:

cd /tmp
wget "https://github.com/prometheus/alertmanager/releases/download/v${ALERTMANAGER_VERSION}/alertmanager-${ALERTMANAGER_VERSION}.linux-amd64.tar.gz"
tar -xzf "alertmanager-${ALERTMANAGER_VERSION}.linux-amd64.tar.gz"

sudo install -m 0755 "alertmanager-${ALERTMANAGER_VERSION}.linux-amd64/alertmanager" /usr/local/bin/alertmanager
sudo install -m 0755 "alertmanager-${ALERTMANAGER_VERSION}.linux-amd64/amtool" /usr/local/bin/amtool

Create a minimal config:

sudoedit /etc/alertmanager/alertmanager.yml

Use a webhook first so you can test the flow without accidentally paging anyone:

route:
  receiver: local-webhook
  group_by:
    - alertname
    - cluster
    - instance
  group_wait: 30s
  group_interval: 5m
  repeat_interval: 4h

receivers:
  - name: local-webhook
    webhook_configs:
      - url: http://127.0.0.1:5001/alert
        send_resolved: true

Validate:

sudo -u alertmanager amtool check-config /etc/alertmanager/alertmanager.yml

Create the service:

sudoedit /etc/systemd/system/alertmanager.service

Use:

[Unit]
Description=Prometheus Alertmanager
Documentation=https://prometheus.io/docs/alerting/latest/alertmanager/
After=network-online.target
Wants=network-online.target

[Service]
User=alertmanager
Group=alertmanager
Type=simple
ExecStart=/usr/local/bin/alertmanager \
  --config.file=/etc/alertmanager/alertmanager.yml \
  --storage.path=/var/lib/alertmanager \
  --web.listen-address=127.0.0.1:9093
Restart=on-failure
RestartSec=5s
NoNewPrivileges=true
ProtectHome=true
ProtectSystem=full
ReadWritePaths=/var/lib/alertmanager
PrivateTmp=true

[Install]
WantedBy=multi-user.target

Start:

sudo systemctl daemon-reload
sudo systemctl enable --now alertmanager
sudo systemctl status alertmanager --no-pager
curl -s http://127.0.0.1:9093/-/ready

Add Prometheus alerting rules

Create:

sudoedit /etc/prometheus/rules/node-alerts.yml

Use:

groups:
  - name: node-alerts
    rules:
      - alert: NodeDown
        expr: up{job="node"} == 0
        for: 5m
        labels:
          severity: page
        annotations:
          summary: "Node is down"
          description: "Prometheus cannot scrape {{ $labels.instance }} for 5 minutes."

      - alert: NodeLowRootFilesystem
        expr: |
          (
            node_filesystem_avail_bytes{job="node",mountpoint="/",fstype!~"tmpfs|overlay"} /
            node_filesystem_size_bytes{job="node",mountpoint="/",fstype!~"tmpfs|overlay"}
          ) < 0.10
        for: 15m
        labels:
          severity: page
        annotations:
          summary: "Root filesystem is low"
          description: "{{ $labels.instance }} has less than 10% free space on /."

      - alert: NodeLowMemory
        expr: |
          (
            node_memory_MemAvailable_bytes{job="node"} /
            node_memory_MemTotal_bytes{job="node"}
          ) < 0.10
        for: 10m
        labels:
          severity: warning
        annotations:
          summary: "Host memory is low"
          description: "{{ $labels.instance }} has less than 10% available memory."

      - alert: NodeHighCpuIowait
        expr: |
          avg by (instance) (rate(node_cpu_seconds_total{job="node",mode="iowait"}[5m])) > 0.20
        for: 15m
        labels:
          severity: warning
        annotations:
          summary: "High CPU iowait"
          description: "{{ $labels.instance }} has high CPU iowait for 15 minutes."

      - alert: NodeRecentlyRebooted
        expr: time() - node_boot_time_seconds{job="node"} < 600
        for: 0m
        labels:
          severity: info
        annotations:
          summary: "Node recently rebooted"
          description: "{{ $labels.instance }} rebooted in the last 10 minutes."

Validate and reload:

sudo -u prometheus promtool check rules /etc/prometheus/rules/node-alerts.yml
sudo -u prometheus promtool check config /etc/prometheus/prometheus.yml
sudo systemctl reload prometheus

[IMAGE: Supporting visual 2 for Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide, showing Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide decisions, examples, and Linux, Monitoring, Prometheus. Alt: Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide linux-monitoring-prometheus-grafana-full-stack-setup-guide visual 2]

Check alerts in Prometheus:

curl -s http://127.0.0.1:9090/api/v1/rules | head
curl -s http://127.0.0.1:9090/api/v1/alerts | head

Test alert delivery without noise

Run a local webhook receiver:

python3 -m http.server 5001

That simple server will return 501 for POST requests, so Alertmanager delivery will fail. For a real local test, use nc:

nc -lk 127.0.0.1 5001

Then temporarily add a test alert:

groups:
  - name: test-alerts
    rules:
      - alert: AlwaysFiringTest
        expr: vector(1)
        for: 1m
        labels:
          severity: test
        annotations:
          summary: "Test alert"
          description: "Remove this rule after validating Alertmanager delivery."

Reload, confirm that the alert reaches Alertmanager, then delete the test rule.

[IMAGE: Supporting visual 2 for Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide, showing Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide decisions, examples, and Linux, Monitoring, Prometheus. Alt: Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide linux-monitoring-prometheus-grafana-full-stack-setup-guide visual 2]

Do not leave permanent alerts with expr: vector(1).

Install Grafana

On Debian or Ubuntu, install from the Grafana APT repository:

sudo apt-get install -y apt-transport-https wget gnupg
sudo mkdir -p /etc/apt/keyrings
sudo wget -O /etc/apt/keyrings/grafana.asc https://apt.grafana.com/gpg-full.key
sudo chmod 644 /etc/apt/keyrings/grafana.asc
echo "deb [signed-by=/etc/apt/keyrings/grafana.asc] https://apt.grafana.com stable main" | sudo tee /etc/apt/sources.list.d/grafana.list
sudo apt-get update
sudo apt-get install -y grafana

Start:

sudo systemctl enable --now grafana-server
sudo systemctl status grafana-server --no-pager

Grafana listens on http://localhost:3000 by default. The initial login is usually admin / admin; change it immediately.

For production, put Grafana behind TLS and an access control layer:

admin browser -> HTTPS reverse proxy or VPN -> Grafana :3000
Grafana -> Prometheus http://127.0.0.1:9090

Provision the Prometheus data source

You can add the data source through the UI, but provisioning keeps it repeatable.

Create:

sudo mkdir -p /etc/grafana/provisioning/datasources
sudoedit /etc/grafana/provisioning/datasources/prometheus.yml

Use:

apiVersion: 1

datasources:
  - name: Prometheus
    type: prometheus
    access: proxy
    url: http://127.0.0.1:9090
    isDefault: true
    editable: false
    jsonData:
      timeInterval: 15s

Restart:

sudo systemctl restart grafana-server

Verify in Grafana:

Connections -> Data sources -> Prometheus -> Save & test

Grafana has native Prometheus support, so no plugin is required.

Build the first dashboard

Create one dashboard for Linux hosts before importing large community dashboards.

Use these panels:

PanelQuery
Scrape statusup{job="node"}
CPU busyinstance:node_cpu_busy:ratio * 100
Memory availableinstance:node_memory_available:ratio * 100
Root filesystem freeinstance:node_filesystem_root_available:ratio * 100
Network receiverate(node_network_receive_bytes_total{job="node",device!~"lo|veth.*|docker.*"}[5m])
Network transmitrate(node_network_transmit_bytes_total{job="node",device!~"lo|veth.*|docker.*"}[5m])
Disk writesrate(node_disk_written_bytes_total{job="node"}[5m])
Load averagenode_load1{job="node"}

Recommended panel settings:

unit for CPU and memory ratios: percent
unit for bytes per second: bytes/sec
legend: {{instance}}
min: 0
dashboard variable: instance

Add a dashboard variable:

Name: instance
Type: Query
Data source: Prometheus
Query: label_values(up{job="node"}, instance)
Refresh: On dashboard load
Multi-value: enabled
Include All: enabled

Then filter queries:

up{job="node",instance=~"$instance"}

Import a Node Exporter dashboard carefully

Grafana.com has shared dashboards for Prometheus and Node Exporter. They are useful, but review them before treating them as production truth.

Common problems:

dashboard expects collectors you did not enable
dashboard assumes different job labels
dashboard uses old metric names
dashboard uses broad queries that are expensive in large fleets
dashboard mixes Linux and container filesystem devices incorrectly

After importing, check:

data source is Prometheus
job label matches your config
filesystem filters exclude tmpfs, overlay, squashfs, and container mounts where appropriate
network filters exclude lo and virtual interfaces where appropriate
query time range is reasonable

The dashboard is not the monitoring system. Your scrape config, labels, rules, and alert routing are the monitoring system.

Choose Prometheus alerts vs Grafana alerts

For infrastructure alerts, prefer Prometheus alerting rules plus Alertmanager:

Prometheus evaluates near the data
rules are files and can be reviewed
Alertmanager handles grouping, silence, inhibition, and routing
alerts survive Grafana downtime

[IMAGE: Supporting visual 3 for Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide, showing Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide decisions, examples, and Linux, Monitoring, Prometheus. Alt: Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide linux-monitoring-prometheus-grafana-full-stack-setup-guide visual 3]

Grafana alerting is useful when:

the alert depends on a Grafana-only data source
the team manages all alerts through Grafana provisioning
the rule combines dashboard context with notification policy

Do not define the same alert in both places unless you intentionally want duplicate notifications.

Secure the stack

Minimum production controls:

Node Exporter :9100 reachable only from Prometheus
Prometheus :9090 bound to localhost or private network
Alertmanager :9093 bound to localhost or private network
Grafana :3000 behind HTTPS, VPN, or SSO
no anonymous Grafana admin access
secrets not stored in world-readable provisioning files
regular backup of Grafana database and Prometheus rules
Prometheus data on local disk, not NFS

Example UFW rules on a monitored host:

sudo ufw allow from 10.0.0.10 to any port 9100 proto tcp
sudo ufw deny 9100/tcp

Example UFW rules on the monitoring host:

sudo ufw allow from 10.0.0.0/24 to any port 3000 proto tcp
sudo ufw deny 9090/tcp
sudo ufw deny 9093/tcp

If multiple networks need access, use explicit rules rather than opening ports globally.

Back up what matters

Back up:

/etc/prometheus/prometheus.yml
/etc/prometheus/rules
/etc/alertmanager/alertmanager.yml
/etc/grafana/provisioning
Grafana database, usually /var/lib/grafana/grafana.db for SQLite installs

Prometheus metrics are operational data. For many teams, rebuilding from exporters is acceptable. If you need durable historical metrics, use snapshots or remote storage instead of copying the live TSDB directory blindly.

Check Prometheus storage usage:

sudo du -sh /var/lib/prometheus
sudo find /var/lib/prometheus -maxdepth 1 -type d -ls | head

Check Grafana data:

sudo ls -lah /var/lib/grafana

Troubleshooting

[IMAGE: Supporting visual 3 for Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide, showing Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide decisions, examples, and Linux, Monitoring, Prometheus. Alt: Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide linux-monitoring-prometheus-grafana-full-stack-setup-guide visual 3]

Node Exporter is not reachable:

systemctl status node_exporter --no-pager
journalctl -u node_exporter -b --no-pager
ss -ltnp | grep 9100
curl -v http://127.0.0.1:9100/metrics

Prometheus target is down:

curl -v http://web-01.internal:9100/metrics
curl -s http://127.0.0.1:9090/api/v1/targets | head
sudo -u prometheus promtool check config /etc/prometheus/prometheus.yml

Prometheus will not start:

systemctl status prometheus --no-pager
journalctl -u prometheus -b --no-pager
sudo -u prometheus promtool check config /etc/prometheus/prometheus.yml
sudo ls -lah /var/lib/prometheus
sudo namei -l /var/lib/prometheus

Rules do not load:

sudo -u prometheus promtool check rules /etc/prometheus/rules/*.yml
curl -s http://127.0.0.1:9090/api/v1/rules | head

Alertmanager receives no alerts:

systemctl status alertmanager --no-pager
journalctl -u alertmanager -b --no-pager
curl -s http://127.0.0.1:9093/api/v2/status | head
curl -s http://127.0.0.1:9090/api/v1/alerts | head

Grafana has no data:

systemctl status grafana-server --no-pager
journalctl -u grafana-server -b --no-pager
curl -s http://127.0.0.1:9090/-/ready

Then check:

Grafana data source URL is reachable from the Grafana server
Prometheus target labels match dashboard queries
dashboard time range includes recent samples
queries do not filter on a missing job or instance label

Production checklist

Before relying on this stack:

all exporters are reachable only from Prometheus
Prometheus config is under version control
rules are validated with promtool in CI
Grafana provisioning is under version control
Grafana admin password is rotated
alerts are tested end to end
runbooks exist for page-level alerts
Prometheus retention fits disk capacity
Grafana database is backed up
monitoring host itself has disk and service alerts

The first useful monitoring stack is not the one with the most panels. It is the one that tells you when a host is down, when disks will fill, when memory is unsafe, and where to look next.

FAQ

What is Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide?

Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide is a practical monitoring topic that should be evaluated through implementation scope, production risk, testing, documentation, and long-term maintainability.

When should a team use Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide?

Use Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide when it solves a real project constraint, improves clarity, or reduces operational risk. Avoid it when it only adds novelty or hides behavior from future maintainers.

What is the biggest risk with Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide?

The biggest risk is copying a pattern without its context. Production systems need clear boundaries, rollback options, tests, and observability before a technique becomes dependable.

How do you test Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide?

Test the smallest unit that owns the behavior, then add integration coverage for the path users or systems actually rely on. Include failure cases, configuration differences, and regression checks.

How does Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide affect SEO and AI search visibility?

It improves visibility when the article gives a direct answer, expert context, structured headings, internal links, trustworthy references, and FAQ content that matches the visible page.

Conclusion

Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide is worth doing when the implementation improves clarity, reliability, or delivery speed. It is not worth doing when it hides ownership, increases operational risk, or makes the system harder to explain.

Use the framework above as a review checklist. Then connect this topic to the rest of the project documentation so readers can move from concept to implementation without losing context.

Top