SEO Metadata
SEO Title Options
- Linux Monitoring With Prometheus & Grafana: Full Stack
- Linux Monitoring With Prometheus: Practical 2026 Guide
- Monitoring Playbook: Linux Monitoring With Prometheus
Meta Description Options
- Learn Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide with a practical Monitoring framework, expert mistakes, implementation steps.
- Installs Node Exporter and Prometheus, configures scrape intervals, builds Grafana dashboards, and sets up alerting rules.
URL Slug
linux-monitoring-prometheus-grafana-full-stack-setup-guide
Focus Keyword
Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide
Additional LSI Keywords
- Monitoring
- Linux
- Prometheus
- Grafana
- Node Exporter
- Alerting
- Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide
- production checklist
- implementation guide
- best practices
- architecture decisions
- testing strategy
Table of Contents
- Article overview
- What Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide means
- Why it matters now
- Implementation framework
- Practical comparison
- Expert workflow
- Common mistakes
- Media and link plan
- Original technical deep dive
- FAQ
- Structured data
- Conclusion
Article overview
Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide is the kind of topic that looks simple until it reaches production. Teams usually discover the real cost late: unclear boundaries, weak defaults, hidden maintenance work, and decisions that seemed harmless when the codebase was small.
The problem gets worse when the article, tutorial, or implementation guide only explains the happy path. This guide closes that gap with a practical framework, a comparison table, common mistakes, and a deep technical section you can use while planning real work.
Keep reading for the non-obvious part: the safest implementation is rarely the most impressive-looking one. It is the one your team can debug, test, document, and evolve without turning every future change into archaeology.
Key Takeaways
- Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide should be evaluated as a production decision, not only as a syntax or tooling choice.
- The best implementation keeps responsibilities visible, with clear ownership, tests, documentation, and rollback paths.
- Search visibility improves when practical depth, structured answers, and expert examples live on the same page.
[IMAGE: A mobile-first technical article layout showing the main concept, decision table, implementation checklist, and FAQ blocks. Alt: Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide expert guide for Monitoring]
What Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide means
Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide means applying monitoring knowledge to a concrete engineering decision, then turning that decision into reliable code, documentation, and operational behavior. In practice, it combines the topic's core concepts with trade-off analysis, implementation boundaries, testing strategy, and maintenance discipline.
This is the definition worth optimizing for featured snippets because it avoids hype. It tells the reader what the topic does and what a professional implementation must include.
Why it matters now
The technical web is more crowded than it was a few years ago. Thin tutorials can still get indexed, but they rarely earn trust from senior developers, buyers, AI answer systems, or teams that need production guidance.
For monitoring topics, the strongest content now has three layers:
- a clear answer for fast scanning
- a practical framework for implementation
- expert context that explains what breaks later
That same structure helps search engines understand the page. It also helps readers decide whether the advice fits their project.
Implementation framework
Use this framework before adopting the approach described in this article.
- Define the user problem and the production risk.
- Identify the smallest reliable implementation boundary.
- Keep configuration, secrets, and environment-specific behavior outside the article's core logic.
- Add tests for the behavior that would hurt if it regressed.
- Document the trade-off, not only the final code.
- Measure the result with logs, metrics, or user-facing outcomes.
- Revisit the decision after real usage exposes edge cases.
The sequence is deliberately conservative. It keeps the work grounded in outcomes instead of novelty.
[IMAGE: A seven-step implementation framework with discovery, boundary design, configuration, tests, documentation, measurement, and iteration. Alt: Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide implementation framework]
Practical comparison
| Decision area | Strong approach | Weak approach | Why it matters |
|---|---|---|---|
| Scope | Solve one clear problem | Mix unrelated concerns | Focus improves testing and search intent |
| Architecture | Put logic in explicit classes or documented boundaries | Hide behavior in templates or incidental callbacks | Future changes stay easier to review |
| Data flow | Pass prepared data into the view or endpoint | Query or compute in presentation code | Reduces regressions and performance surprises |
| Testing | Cover the risky behavior directly | Test only the happy path | Catches production failures earlier |
| Documentation | Explain trade-offs and limits | Repeat generic definitions | Builds E-E-A-T and reader trust |
| Operations | Track logs, metrics, and rollback steps | Ship without measurement | Makes the decision reversible |
This table is intentionally practical. It gives a reviewer something to check before the implementation becomes expensive to change.
Expert workflow
Expert tip: "Treat Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide as a system boundary. If the next developer cannot find where the decision lives, how it is tested, and when it should be avoided, the implementation is not finished."
A useful workflow is simple:
- Start with the smallest working example.
- Add the constraints that exist in your real project.
- Remove anything that only demonstrates cleverness.
- Write down the failure modes.
- Add links to related decisions so future readers can navigate the topic cluster.
That last point matters for both humans and search systems. A single article can answer a question; a cluster proves authority.
Common mistakes
Mistake 1: Copying a pattern without its context
A pattern that works in a small demo can fail in a real application. The missing context is usually data volume, team experience, deployment process, security requirements, or observability.
Before copying the pattern, ask what assumption made it safe in the original example.
Mistake 2: Putting business logic in the wrong layer
This is the fastest way to make future debugging expensive. In Laravel, PHP, and server-rendered websites, presentation should receive prepared data, not discover rules on its own.
Keep decision logic in models, actions, services, policies, requests, jobs, or documented helpers where it can be tested directly.
Mistake 3: Optimizing for novelty instead of maintainability
Newer tools and language features can be valuable. They can also hide simple behavior behind unfamiliar syntax.
Use the option that makes the next production incident easier to understand.
Mistake 4: Publishing without a measurement plan
If the article describes a performance, SEO, security, or architecture improvement, define how success will be checked. Logs, tests, crawl diagnostics, analytics, and user behavior are all stronger than assumptions.
[IMAGE: A common-mistakes board with context loss, wrong layer, novelty bias, and missing measurement highlighted. Alt: Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide common mistakes]
Media and link plan
Image placeholders
- [IMAGE: A concept diagram for Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide with input, decision boundary, implementation, tests, and production feedback. Alt: Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide concept diagram]
- [IMAGE: A mobile screenshot-style checklist for Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide. Alt: Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide mobile checklist]
- [IMAGE: A comparison table visualization for strong versus weak implementation choices. Alt: Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide comparison table]
Video placeholder
[VIDEO: Insert a 5-8 minute YouTube walkthrough that demonstrates the main decision, the implementation boundary, the test strategy, and the production caveats for Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide.]
Trustworthy outbound links
- Linux manual pages - use this as the trust reference for operating-system reference.
- Google Search quality guidance - use this as the trust reference for people-first content and E-E-A-T alignment.
Internal linking opportunities
- Internal guide: Linux Log Management: rsyslog, logrotate - use this when readers need a related Monitoring follow-up.
- Internal guide: Configuring PostgreSQL on Linux - use this when readers need a related Database follow-up.
Original Technical Deep Dive
The short version
A practical Linux monitoring stack has four parts:
| Layer | Tool | Job |
|---|---|---|
| Host metrics | Node Exporter | Expose CPU, memory, disk, filesystem, network, and kernel metrics on each Linux host |
| Time series database | Prometheus | Scrape exporters, store samples, evaluate PromQL and alerting rules |
| Notification routing | Alertmanager | Group, deduplicate, silence, and route firing alerts |
| Dashboards | Grafana | Query Prometheus and visualize host health |
The baseline topology:
Linux hosts: node_exporter on :9100
Monitoring host: prometheus on :9090
Monitoring host: alertmanager on :9093
Monitoring host: grafana-server on :3000
Use this guide for a small to medium Linux fleet where Prometheus scrapes hosts over a private network. Do not expose exporter, Prometheus, Alertmanager, or Grafana ports directly to the public internet.
This guide was reviewed on May 7, 2026 against the Prometheus and Grafana documentation. The examples use the current Prometheus download versions at that date:
Prometheus: 3.11.3
Node Exporter: 1.11.1
Alertmanager: 0.32.1
Check the official download page before installing on a real host and adjust the version variables.
Decide what to monitor first
Start with host signals that predict outages:
| Signal | Useful questions |
|---|---|
up | Is Prometheus able to scrape the target? |
| CPU saturation | Is the host out of CPU or stuck in iowait? |
| Memory pressure | Is available memory low? |
| Disk space | Will a filesystem fill soon? |
| Disk I/O | Are disks overloaded or slow? |
| Network traffic | Is a host receiving or sending unusual traffic? |
| Reboots | Did the host restart unexpectedly? |
Avoid trying to alert on every metric. Alerts should represent conditions that need action.
Good first alerts:
target cannot be scraped for 5 minutes
root filesystem has less than 10% free space for 15 minutes
memory available is below 10% for 10 minutes
CPU iowait is high for 15 minutes
host rebooted unexpectedly
Dashboards are for investigation. Alerts are for interruption.
Prepare the hosts
These commands assume Ubuntu or Debian. The systemd units work the same on most modern Linux distributions.
Install basic tools:
sudo apt update
sudo apt install -y curl tar wget ca-certificates gnupg
Create dedicated users:
sudo useradd --system --no-create-home --shell /usr/sbin/nologin prometheus
sudo useradd --system --no-create-home --shell /usr/sbin/nologin node_exporter
sudo useradd --system --no-create-home --shell /usr/sbin/nologin alertmanager
If your distribution does not have /usr/sbin/nologin, use /bin/false.
Create directories on the monitoring host:
sudo mkdir -p /etc/prometheus /var/lib/prometheus
sudo mkdir -p /etc/alertmanager /var/lib/alertmanager
sudo chown prometheus:prometheus /etc/prometheus /var/lib/prometheus
sudo chown alertmanager:alertmanager /etc/alertmanager /var/lib/alertmanager
sudo chmod 0750 /etc/prometheus /var/lib/prometheus /etc/alertmanager /var/lib/alertmanager
Install Node Exporter on each Linux host
Node Exporter should run on every Linux server you want to monitor.
Set the version:
NODE_EXPORTER_VERSION=1.11.1
Download and install:
cd /tmp
wget "https://github.com/prometheus/node_exporter/releases/download/v${NODE_EXPORTER_VERSION}/node_exporter-${NODE_EXPORTER_VERSION}.linux-amd64.tar.gz"
tar -xzf "node_exporter-${NODE_EXPORTER_VERSION}.linux-amd64.tar.gz"
sudo install -m 0755 "node_exporter-${NODE_EXPORTER_VERSION}.linux-amd64/node_exporter" /usr/local/bin/node_exporter
Create the systemd service:
sudoedit /etc/systemd/system/node_exporter.service
Use:
[Unit]
Description=Prometheus Node Exporter
Documentation=https://prometheus.io/docs/guides/node-exporter/
After=network-online.target
Wants=network-online.target
[Service]
User=node_exporter
Group=node_exporter
Type=simple
ExecStart=/usr/local/bin/node_exporter \
--web.listen-address=:9100
Restart=on-failure
RestartSec=5s
NoNewPrivileges=true
ProtectSystem=full
ProtectHome=true
PrivateTmp=true
[Install]
WantedBy=multi-user.target
Start it:
sudo systemctl daemon-reload
sudo systemctl enable --now node_exporter
sudo systemctl status node_exporter --no-pager
Verify metrics locally:
curl -s http://127.0.0.1:9100/metrics | head
curl -s http://127.0.0.1:9100/metrics | grep '^node_' | head
Open port 9100 only from the Prometheus server:
sudo ufw allow from 10.0.0.10 to any port 9100 proto tcp
Replace 10.0.0.10 with the Prometheus server IP.
Add optional Node Exporter collectors
The default collectors are enough for most hosts. Add collectors only when dashboards or alerts need them.
Common additions:
ExecStart=/usr/local/bin/node_exporter \
--web.listen-address=:9100 \
--collector.systemd \
--collector.processes \
--collector.textfile.directory=/var/lib/node_exporter/textfile_collector
Create the textfile collector directory if you enable it:
sudo mkdir -p /var/lib/node_exporter/textfile_collector
sudo chown node_exporter:node_exporter /var/lib/node_exporter/textfile_collector
sudo chmod 0755 /var/lib/node_exporter/textfile_collector
[IMAGE: Supporting visual 1 for Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide, showing Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide decisions, examples, and Linux, Monitoring, Prometheus. Alt: Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide linux-monitoring-prometheus-grafana-full-stack-setup-guide visual 1]
[IMAGE: Supporting visual 1 for Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide, showing Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide decisions, examples, and Linux, Monitoring, Prometheus. Alt: Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide linux-monitoring-prometheus-grafana-full-stack-setup-guide visual 1]
Example custom metric:
tmpfile=$(mktemp)
printf 'backup_last_success_timestamp_seconds %s\n' "$(date +%s)" > "$tmpfile"
sudo mv "$tmpfile" /var/lib/node_exporter/textfile_collector/backup.prom
sudo chown node_exporter:node_exporter /var/lib/node_exporter/textfile_collector/backup.prom
Then query:
time() - backup_last_success_timestamp_seconds
This is useful for cron jobs, backups, certificate renewal, and other host-level checks that do not have their own exporter.
Install Prometheus on the monitoring host
Set the version:
PROMETHEUS_VERSION=3.11.3
Download and install:
cd /tmp
wget "https://github.com/prometheus/prometheus/releases/download/v${PROMETHEUS_VERSION}/prometheus-${PROMETHEUS_VERSION}.linux-amd64.tar.gz"
tar -xzf "prometheus-${PROMETHEUS_VERSION}.linux-amd64.tar.gz"
sudo install -m 0755 "prometheus-${PROMETHEUS_VERSION}.linux-amd64/prometheus" /usr/local/bin/prometheus
sudo install -m 0755 "prometheus-${PROMETHEUS_VERSION}.linux-amd64/promtool" /usr/local/bin/promtool
sudo cp -r "prometheus-${PROMETHEUS_VERSION}.linux-amd64/consoles" /etc/prometheus/
sudo cp -r "prometheus-${PROMETHEUS_VERSION}.linux-amd64/console_libraries" /etc/prometheus/
sudo chown -R prometheus:prometheus /etc/prometheus /var/lib/prometheus
Create the config:
sudoedit /etc/prometheus/prometheus.yml
Use:
global:
scrape_interval: 15s
scrape_timeout: 10s
evaluation_interval: 15s
external_labels:
environment: production
monitor: primary
rule_files:
- /etc/prometheus/rules/*.yml
alerting:
alertmanagers:
- static_configs:
- targets:
- 127.0.0.1:9093
scrape_configs:
- job_name: prometheus
static_configs:
- targets:
- 127.0.0.1:9090
- job_name: node
scrape_interval: 15s
static_configs:
- targets:
- web-01.internal:9100
- web-02.internal:9100
- db-01.internal:9100
labels:
role: linux
site: primary
Create the rules directory:
sudo mkdir -p /etc/prometheus/rules
sudo chown prometheus:prometheus /etc/prometheus/rules
sudo chmod 0750 /etc/prometheus/rules
Validate:
sudo -u prometheus promtool check config /etc/prometheus/prometheus.yml
Run Prometheus with systemd
Create:
sudoedit /etc/systemd/system/prometheus.service
Use:
[Unit]
Description=Prometheus
Documentation=https://prometheus.io/docs/prometheus/latest/getting_started/
After=network-online.target
Wants=network-online.target
[Service]
User=prometheus
Group=prometheus
Type=simple
ExecStart=/usr/local/bin/prometheus \
--config.file=/etc/prometheus/prometheus.yml \
--storage.tsdb.path=/var/lib/prometheus \
--storage.tsdb.retention.time=30d \
--storage.tsdb.retention.size=40GB \
--web.console.templates=/etc/prometheus/consoles \
--web.console.libraries=/etc/prometheus/console_libraries \
--web.listen-address=127.0.0.1:9090 \
--web.enable-lifecycle
ExecReload=/bin/kill -HUP $MAINPID
Restart=on-failure
RestartSec=5s
NoNewPrivileges=true
ProtectHome=true
ProtectSystem=full
ReadWritePaths=/var/lib/prometheus
PrivateTmp=true
[Install]
WantedBy=multi-user.target
Start:
sudo systemctl daemon-reload
sudo systemctl enable --now prometheus
sudo systemctl status prometheus --no-pager
Check:
curl -s http://127.0.0.1:9090/-/ready
curl -s http://127.0.0.1:9090/api/v1/targets | head
Prometheus is bound to 127.0.0.1 in this example. Put Nginx, Caddy, SSH tunneling, or VPN access in front of it if administrators need browser access.
Reload Prometheus safely
Validate first:
sudo -u prometheus promtool check config /etc/prometheus/prometheus.yml
sudo -u prometheus promtool check rules /etc/prometheus/rules/*.yml
Reload with systemd:
sudo systemctl reload prometheus
Or use the lifecycle endpoint:
curl -X POST http://127.0.0.1:9090/-/reload
If validation fails, fix the YAML before reloading. Prometheus keeps the old config if a runtime reload receives a malformed config, but your deploy process should still fail before that point.
Understand scrape intervals
Use 15s for hosts unless you have a reason to do otherwise.
| Interval | Use case | Tradeoff |
|---|---|---|
5s | Short-lived spikes, test environments | More samples, more storage, more query load |
15s | General Linux host monitoring | Good default |
30s | Larger fleets or slower links | Less precision |
60s | Low-value, low-change targets | Slower alert detection |
Prometheus storage roughly depends on:
retention_seconds * ingested_samples_per_second * bytes_per_sample
Reducing high-cardinality series usually saves more than increasing the scrape interval. A label like user_id, request_id, session, or full path can explode series count.
Write useful PromQL queries
Host availability:
up{job="node"}
CPU busy percent:
100 - (avg by (instance) (rate(node_cpu_seconds_total{job="node",mode="idle"}[5m])) * 100)
Memory available percent:
node_memory_MemAvailable_bytes{job="node"} / node_memory_MemTotal_bytes{job="node"} * 100
Root filesystem free percent:
node_filesystem_avail_bytes{job="node",mountpoint="/",fstype!~"tmpfs|overlay"} /
node_filesystem_size_bytes{job="node",mountpoint="/",fstype!~"tmpfs|overlay"} * 100
Disk read/write rate:
rate(node_disk_read_bytes_total{job="node"}[5m])
rate(node_disk_written_bytes_total{job="node"}[5m])
Network receive/transmit rate:
rate(node_network_receive_bytes_total{job="node",device!~"lo|veth.*|docker.*"}[5m])
rate(node_network_transmit_bytes_total{job="node",device!~"lo|veth.*|docker.*"}[5m])
Unexpected reboot:
time() - node_boot_time_seconds{job="node"}
Start in Prometheus table view before graphing broad selectors. A bare metric can return thousands of series.
Add recording rules for dashboard queries
Recording rules precompute expensive expressions. Create:
sudoedit /etc/prometheus/rules/node-recording.yml
Use:
groups:
- name: node-recording
interval: 30s
rules:
- record: instance:node_cpu_busy:ratio
expr: 1 - avg by (instance) (rate(node_cpu_seconds_total{job="node",mode="idle"}[5m]))
- record: instance:node_memory_available:ratio
expr: node_memory_MemAvailable_bytes{job="node"} / node_memory_MemTotal_bytes{job="node"}
- record: instance:node_filesystem_root_available:ratio
expr: |
node_filesystem_avail_bytes{job="node",mountpoint="/",fstype!~"tmpfs|overlay"} /
node_filesystem_size_bytes{job="node",mountpoint="/",fstype!~"tmpfs|overlay"}
Validate and reload:
sudo -u prometheus promtool check rules /etc/prometheus/rules/node-recording.yml
sudo systemctl reload prometheus
Then dashboards can query:
instance:node_cpu_busy:ratio * 100
instance:node_memory_available:ratio * 100
instance:node_filesystem_root_available:ratio * 100
Install Alertmanager
Set the version:
ALERTMANAGER_VERSION=0.32.1
Download and install:
cd /tmp
wget "https://github.com/prometheus/alertmanager/releases/download/v${ALERTMANAGER_VERSION}/alertmanager-${ALERTMANAGER_VERSION}.linux-amd64.tar.gz"
tar -xzf "alertmanager-${ALERTMANAGER_VERSION}.linux-amd64.tar.gz"
sudo install -m 0755 "alertmanager-${ALERTMANAGER_VERSION}.linux-amd64/alertmanager" /usr/local/bin/alertmanager
sudo install -m 0755 "alertmanager-${ALERTMANAGER_VERSION}.linux-amd64/amtool" /usr/local/bin/amtool
Create a minimal config:
sudoedit /etc/alertmanager/alertmanager.yml
Use a webhook first so you can test the flow without accidentally paging anyone:
route:
receiver: local-webhook
group_by:
- alertname
- cluster
- instance
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
receivers:
- name: local-webhook
webhook_configs:
- url: http://127.0.0.1:5001/alert
send_resolved: true
Validate:
sudo -u alertmanager amtool check-config /etc/alertmanager/alertmanager.yml
Create the service:
sudoedit /etc/systemd/system/alertmanager.service
Use:
[Unit]
Description=Prometheus Alertmanager
Documentation=https://prometheus.io/docs/alerting/latest/alertmanager/
After=network-online.target
Wants=network-online.target
[Service]
User=alertmanager
Group=alertmanager
Type=simple
ExecStart=/usr/local/bin/alertmanager \
--config.file=/etc/alertmanager/alertmanager.yml \
--storage.path=/var/lib/alertmanager \
--web.listen-address=127.0.0.1:9093
Restart=on-failure
RestartSec=5s
NoNewPrivileges=true
ProtectHome=true
ProtectSystem=full
ReadWritePaths=/var/lib/alertmanager
PrivateTmp=true
[Install]
WantedBy=multi-user.target
Start:
sudo systemctl daemon-reload
sudo systemctl enable --now alertmanager
sudo systemctl status alertmanager --no-pager
curl -s http://127.0.0.1:9093/-/ready
Add Prometheus alerting rules
Create:
sudoedit /etc/prometheus/rules/node-alerts.yml
Use:
groups:
- name: node-alerts
rules:
- alert: NodeDown
expr: up{job="node"} == 0
for: 5m
labels:
severity: page
annotations:
summary: "Node is down"
description: "Prometheus cannot scrape {{ $labels.instance }} for 5 minutes."
- alert: NodeLowRootFilesystem
expr: |
(
node_filesystem_avail_bytes{job="node",mountpoint="/",fstype!~"tmpfs|overlay"} /
node_filesystem_size_bytes{job="node",mountpoint="/",fstype!~"tmpfs|overlay"}
) < 0.10
for: 15m
labels:
severity: page
annotations:
summary: "Root filesystem is low"
description: "{{ $labels.instance }} has less than 10% free space on /."
- alert: NodeLowMemory
expr: |
(
node_memory_MemAvailable_bytes{job="node"} /
node_memory_MemTotal_bytes{job="node"}
) < 0.10
for: 10m
labels:
severity: warning
annotations:
summary: "Host memory is low"
description: "{{ $labels.instance }} has less than 10% available memory."
- alert: NodeHighCpuIowait
expr: |
avg by (instance) (rate(node_cpu_seconds_total{job="node",mode="iowait"}[5m])) > 0.20
for: 15m
labels:
severity: warning
annotations:
summary: "High CPU iowait"
description: "{{ $labels.instance }} has high CPU iowait for 15 minutes."
- alert: NodeRecentlyRebooted
expr: time() - node_boot_time_seconds{job="node"} < 600
for: 0m
labels:
severity: info
annotations:
summary: "Node recently rebooted"
description: "{{ $labels.instance }} rebooted in the last 10 minutes."
Validate and reload:
sudo -u prometheus promtool check rules /etc/prometheus/rules/node-alerts.yml
sudo -u prometheus promtool check config /etc/prometheus/prometheus.yml
sudo systemctl reload prometheus
[IMAGE: Supporting visual 2 for Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide, showing Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide decisions, examples, and Linux, Monitoring, Prometheus. Alt: Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide linux-monitoring-prometheus-grafana-full-stack-setup-guide visual 2]
Check alerts in Prometheus:
curl -s http://127.0.0.1:9090/api/v1/rules | head
curl -s http://127.0.0.1:9090/api/v1/alerts | head
Test alert delivery without noise
Run a local webhook receiver:
python3 -m http.server 5001
That simple server will return 501 for POST requests, so Alertmanager delivery will fail. For a real local test, use nc:
nc -lk 127.0.0.1 5001
Then temporarily add a test alert:
groups:
- name: test-alerts
rules:
- alert: AlwaysFiringTest
expr: vector(1)
for: 1m
labels:
severity: test
annotations:
summary: "Test alert"
description: "Remove this rule after validating Alertmanager delivery."
Reload, confirm that the alert reaches Alertmanager, then delete the test rule.
[IMAGE: Supporting visual 2 for Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide, showing Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide decisions, examples, and Linux, Monitoring, Prometheus. Alt: Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide linux-monitoring-prometheus-grafana-full-stack-setup-guide visual 2]
Do not leave permanent alerts with expr: vector(1).
Install Grafana
On Debian or Ubuntu, install from the Grafana APT repository:
sudo apt-get install -y apt-transport-https wget gnupg
sudo mkdir -p /etc/apt/keyrings
sudo wget -O /etc/apt/keyrings/grafana.asc https://apt.grafana.com/gpg-full.key
sudo chmod 644 /etc/apt/keyrings/grafana.asc
echo "deb [signed-by=/etc/apt/keyrings/grafana.asc] https://apt.grafana.com stable main" | sudo tee /etc/apt/sources.list.d/grafana.list
sudo apt-get update
sudo apt-get install -y grafana
Start:
sudo systemctl enable --now grafana-server
sudo systemctl status grafana-server --no-pager
Grafana listens on http://localhost:3000 by default. The initial login is usually admin / admin; change it immediately.
For production, put Grafana behind TLS and an access control layer:
admin browser -> HTTPS reverse proxy or VPN -> Grafana :3000
Grafana -> Prometheus http://127.0.0.1:9090
Provision the Prometheus data source
You can add the data source through the UI, but provisioning keeps it repeatable.
Create:
sudo mkdir -p /etc/grafana/provisioning/datasources
sudoedit /etc/grafana/provisioning/datasources/prometheus.yml
Use:
apiVersion: 1
datasources:
- name: Prometheus
type: prometheus
access: proxy
url: http://127.0.0.1:9090
isDefault: true
editable: false
jsonData:
timeInterval: 15s
Restart:
sudo systemctl restart grafana-server
Verify in Grafana:
Connections -> Data sources -> Prometheus -> Save & test
Grafana has native Prometheus support, so no plugin is required.
Build the first dashboard
Create one dashboard for Linux hosts before importing large community dashboards.
Use these panels:
| Panel | Query |
|---|---|
| Scrape status | up{job="node"} |
| CPU busy | instance:node_cpu_busy:ratio * 100 |
| Memory available | instance:node_memory_available:ratio * 100 |
| Root filesystem free | instance:node_filesystem_root_available:ratio * 100 |
| Network receive | rate(node_network_receive_bytes_total{job="node",device!~"lo|veth.*|docker.*"}[5m]) |
| Network transmit | rate(node_network_transmit_bytes_total{job="node",device!~"lo|veth.*|docker.*"}[5m]) |
| Disk writes | rate(node_disk_written_bytes_total{job="node"}[5m]) |
| Load average | node_load1{job="node"} |
Recommended panel settings:
unit for CPU and memory ratios: percent
unit for bytes per second: bytes/sec
legend: {{instance}}
min: 0
dashboard variable: instance
Add a dashboard variable:
Name: instance
Type: Query
Data source: Prometheus
Query: label_values(up{job="node"}, instance)
Refresh: On dashboard load
Multi-value: enabled
Include All: enabled
Then filter queries:
up{job="node",instance=~"$instance"}
Import a Node Exporter dashboard carefully
Grafana.com has shared dashboards for Prometheus and Node Exporter. They are useful, but review them before treating them as production truth.
Common problems:
dashboard expects collectors you did not enable
dashboard assumes different job labels
dashboard uses old metric names
dashboard uses broad queries that are expensive in large fleets
dashboard mixes Linux and container filesystem devices incorrectly
After importing, check:
data source is Prometheus
job label matches your config
filesystem filters exclude tmpfs, overlay, squashfs, and container mounts where appropriate
network filters exclude lo and virtual interfaces where appropriate
query time range is reasonable
The dashboard is not the monitoring system. Your scrape config, labels, rules, and alert routing are the monitoring system.
Choose Prometheus alerts vs Grafana alerts
For infrastructure alerts, prefer Prometheus alerting rules plus Alertmanager:
Prometheus evaluates near the data
rules are files and can be reviewed
Alertmanager handles grouping, silence, inhibition, and routing
alerts survive Grafana downtime
[IMAGE: Supporting visual 3 for Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide, showing Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide decisions, examples, and Linux, Monitoring, Prometheus. Alt: Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide linux-monitoring-prometheus-grafana-full-stack-setup-guide visual 3]
Grafana alerting is useful when:
the alert depends on a Grafana-only data source
the team manages all alerts through Grafana provisioning
the rule combines dashboard context with notification policy
Do not define the same alert in both places unless you intentionally want duplicate notifications.
Secure the stack
Minimum production controls:
Node Exporter :9100 reachable only from Prometheus
Prometheus :9090 bound to localhost or private network
Alertmanager :9093 bound to localhost or private network
Grafana :3000 behind HTTPS, VPN, or SSO
no anonymous Grafana admin access
secrets not stored in world-readable provisioning files
regular backup of Grafana database and Prometheus rules
Prometheus data on local disk, not NFS
Example UFW rules on a monitored host:
sudo ufw allow from 10.0.0.10 to any port 9100 proto tcp
sudo ufw deny 9100/tcp
Example UFW rules on the monitoring host:
sudo ufw allow from 10.0.0.0/24 to any port 3000 proto tcp
sudo ufw deny 9090/tcp
sudo ufw deny 9093/tcp
If multiple networks need access, use explicit rules rather than opening ports globally.
Back up what matters
Back up:
/etc/prometheus/prometheus.yml
/etc/prometheus/rules
/etc/alertmanager/alertmanager.yml
/etc/grafana/provisioning
Grafana database, usually /var/lib/grafana/grafana.db for SQLite installs
Prometheus metrics are operational data. For many teams, rebuilding from exporters is acceptable. If you need durable historical metrics, use snapshots or remote storage instead of copying the live TSDB directory blindly.
Check Prometheus storage usage:
sudo du -sh /var/lib/prometheus
sudo find /var/lib/prometheus -maxdepth 1 -type d -ls | head
Check Grafana data:
sudo ls -lah /var/lib/grafana
Troubleshooting
[IMAGE: Supporting visual 3 for Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide, showing Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide decisions, examples, and Linux, Monitoring, Prometheus. Alt: Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide linux-monitoring-prometheus-grafana-full-stack-setup-guide visual 3]
Node Exporter is not reachable:
systemctl status node_exporter --no-pager
journalctl -u node_exporter -b --no-pager
ss -ltnp | grep 9100
curl -v http://127.0.0.1:9100/metrics
Prometheus target is down:
curl -v http://web-01.internal:9100/metrics
curl -s http://127.0.0.1:9090/api/v1/targets | head
sudo -u prometheus promtool check config /etc/prometheus/prometheus.yml
Prometheus will not start:
systemctl status prometheus --no-pager
journalctl -u prometheus -b --no-pager
sudo -u prometheus promtool check config /etc/prometheus/prometheus.yml
sudo ls -lah /var/lib/prometheus
sudo namei -l /var/lib/prometheus
Rules do not load:
sudo -u prometheus promtool check rules /etc/prometheus/rules/*.yml
curl -s http://127.0.0.1:9090/api/v1/rules | head
Alertmanager receives no alerts:
systemctl status alertmanager --no-pager
journalctl -u alertmanager -b --no-pager
curl -s http://127.0.0.1:9093/api/v2/status | head
curl -s http://127.0.0.1:9090/api/v1/alerts | head
Grafana has no data:
systemctl status grafana-server --no-pager
journalctl -u grafana-server -b --no-pager
curl -s http://127.0.0.1:9090/-/ready
Then check:
Grafana data source URL is reachable from the Grafana server
Prometheus target labels match dashboard queries
dashboard time range includes recent samples
queries do not filter on a missing job or instance label
Production checklist
Before relying on this stack:
all exporters are reachable only from Prometheus
Prometheus config is under version control
rules are validated with promtool in CI
Grafana provisioning is under version control
Grafana admin password is rotated
alerts are tested end to end
runbooks exist for page-level alerts
Prometheus retention fits disk capacity
Grafana database is backed up
monitoring host itself has disk and service alerts
The first useful monitoring stack is not the one with the most panels. It is the one that tells you when a host is down, when disks will fill, when memory is unsafe, and where to look next.
FAQ
What is Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide?
Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide is a practical monitoring topic that should be evaluated through implementation scope, production risk, testing, documentation, and long-term maintainability.
When should a team use Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide?
Use Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide when it solves a real project constraint, improves clarity, or reduces operational risk. Avoid it when it only adds novelty or hides behavior from future maintainers.
What is the biggest risk with Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide?
The biggest risk is copying a pattern without its context. Production systems need clear boundaries, rollback options, tests, and observability before a technique becomes dependable.
How do you test Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide?
Test the smallest unit that owns the behavior, then add integration coverage for the path users or systems actually rely on. Include failure cases, configuration differences, and regression checks.
How does Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide affect SEO and AI search visibility?
It improves visibility when the article gives a direct answer, expert context, structured headings, internal links, trustworthy references, and FAQ content that matches the visible page.
Conclusion
Linux Monitoring With Prometheus & Grafana: Full Stack Setup Guide is worth doing when the implementation improves clarity, reliability, or delivery speed. It is not worth doing when it hides ownership, increases operational risk, or makes the system harder to explain.
Use the framework above as a review checklist. Then connect this topic to the rest of the project documentation so readers can move from concept to implementation without losing context.