Back to blog

Networking

Linux High Availability With Keepalived: VRRP & Failover Configuration

Sets up Keepalived VRRP for virtual IP failover, configures health check scripts, tracks interface state, and tests manual failover.

  • Linux
  • Keepalived
  • VRRP
  • High Availability
  • Networking
  • Failover

SEO Metadata

SEO Title Options

  1. Linux High Availability With Keepalived: VRRP & Failover
  2. Linux Networking: Practical 2026 Guide
  3. Networking Playbook: Linux Networking

Meta Description Options

  1. Learn Linux Networking with a practical Networking framework, expert mistakes, implementation steps, examples, FAQ, and schema-ready guidance.
  2. Sets up Keepalived VRRP for virtual IP failover, configures health check scripts, tracks interface state, and tests manual failover.

URL Slug

linux-high-availability-keepalived-vrrp-failover-configuration

Focus Keyword

Linux Networking

Additional LSI Keywords

  • Networking
  • Linux
  • Keepalived
  • VRRP
  • High Availability
  • Failover
  • Linux High Availability With Keepalived: VRRP & Failover Configuration
  • production checklist
  • implementation guide
  • best practices
  • architecture decisions
  • testing strategy

Table of Contents

Article overview

Linux Networking is the kind of topic that looks simple until it reaches production. Teams usually discover the real cost late: unclear boundaries, weak defaults, hidden maintenance work, and decisions that seemed harmless when the codebase was small.

The problem gets worse when the article, tutorial, or implementation guide only explains the happy path. This guide closes that gap with a practical framework, a comparison table, common mistakes, and a deep technical section you can use while planning real work.

Keep reading for the non-obvious part: the safest implementation is rarely the most impressive-looking one. It is the one your team can debug, test, document, and evolve without turning every future change into archaeology.

Key Takeaways

  • Linux Networking should be evaluated as a production decision, not only as a syntax or tooling choice.
  • The best implementation keeps responsibilities visible, with clear ownership, tests, documentation, and rollback paths.
  • Search visibility improves when practical depth, structured answers, and expert examples live on the same page.

[IMAGE: A mobile-first technical article layout showing the main concept, decision table, implementation checklist, and FAQ blocks. Alt: Linux Networking expert guide for Networking]

What Linux Networking means

Linux Networking means applying networking knowledge to a concrete engineering decision, then turning that decision into reliable code, documentation, and operational behavior. In practice, it combines the topic's core concepts with trade-off analysis, implementation boundaries, testing strategy, and maintenance discipline.

This is the definition worth optimizing for featured snippets because it avoids hype. It tells the reader what the topic does and what a professional implementation must include.

Why it matters now

The technical web is more crowded than it was a few years ago. Thin tutorials can still get indexed, but they rarely earn trust from senior developers, buyers, AI answer systems, or teams that need production guidance.

For networking topics, the strongest content now has three layers:

  • a clear answer for fast scanning
  • a practical framework for implementation
  • expert context that explains what breaks later

That same structure helps search engines understand the page. It also helps readers decide whether the advice fits their project.

Implementation framework

Use this framework before adopting the approach described in this article.

  1. Define the user problem and the production risk.
  2. Identify the smallest reliable implementation boundary.
  3. Keep configuration, secrets, and environment-specific behavior outside the article's core logic.
  4. Add tests for the behavior that would hurt if it regressed.
  5. Document the trade-off, not only the final code.
  6. Measure the result with logs, metrics, or user-facing outcomes.
  7. Revisit the decision after real usage exposes edge cases.

The sequence is deliberately conservative. It keeps the work grounded in outcomes instead of novelty.

[IMAGE: A seven-step implementation framework with discovery, boundary design, configuration, tests, documentation, measurement, and iteration. Alt: Linux Networking implementation framework]

Practical comparison

Decision areaStrong approachWeak approachWhy it matters
ScopeSolve one clear problemMix unrelated concernsFocus improves testing and search intent
ArchitecturePut logic in explicit classes or documented boundariesHide behavior in templates or incidental callbacksFuture changes stay easier to review
Data flowPass prepared data into the view or endpointQuery or compute in presentation codeReduces regressions and performance surprises
TestingCover the risky behavior directlyTest only the happy pathCatches production failures earlier
DocumentationExplain trade-offs and limitsRepeat generic definitionsBuilds E-E-A-T and reader trust
OperationsTrack logs, metrics, and rollback stepsShip without measurementMakes the decision reversible

This table is intentionally practical. It gives a reviewer something to check before the implementation becomes expensive to change.

Expert workflow

Expert tip: "Treat Linux Networking as a system boundary. If the next developer cannot find where the decision lives, how it is tested, and when it should be avoided, the implementation is not finished."

A useful workflow is simple:

  • Start with the smallest working example.
  • Add the constraints that exist in your real project.
  • Remove anything that only demonstrates cleverness.
  • Write down the failure modes.
  • Add links to related decisions so future readers can navigate the topic cluster.

That last point matters for both humans and search systems. A single article can answer a question; a cluster proves authority.

Common mistakes

Mistake 1: Copying a pattern without its context

A pattern that works in a small demo can fail in a real application. The missing context is usually data volume, team experience, deployment process, security requirements, or observability.

Before copying the pattern, ask what assumption made it safe in the original example.

Mistake 2: Putting business logic in the wrong layer

This is the fastest way to make future debugging expensive. In Laravel, PHP, and server-rendered websites, presentation should receive prepared data, not discover rules on its own.

Keep decision logic in models, actions, services, policies, requests, jobs, or documented helpers where it can be tested directly.

Mistake 3: Optimizing for novelty instead of maintainability

Newer tools and language features can be valuable. They can also hide simple behavior behind unfamiliar syntax.

Use the option that makes the next production incident easier to understand.

Mistake 4: Publishing without a measurement plan

If the article describes a performance, SEO, security, or architecture improvement, define how success will be checked. Logs, tests, crawl diagnostics, analytics, and user behavior are all stronger than assumptions.

[IMAGE: A common-mistakes board with context loss, wrong layer, novelty bias, and missing measurement highlighted. Alt: Linux Networking common mistakes]

Image placeholders

  • [IMAGE: A concept diagram for Linux Networking with input, decision boundary, implementation, tests, and production feedback. Alt: Linux Networking concept diagram]
  • [IMAGE: A mobile screenshot-style checklist for Linux High Availability With Keepalived: VRRP & Failover Configuration. Alt: Linux Networking mobile checklist]
  • [IMAGE: A comparison table visualization for strong versus weak implementation choices. Alt: Linux Networking comparison table]

Video placeholder

[VIDEO: Insert a 5-8 minute YouTube walkthrough that demonstrates the main decision, the implementation boundary, the test strategy, and the production caveats for Linux Networking.]

Internal linking opportunities

Original Technical Deep Dive

The short version

Keepalived uses VRRP to move a virtual IP address between Linux hosts. It is a good fit for active-passive failover of a load balancer, reverse proxy, DNS resolver, NFS endpoint, or small internal service.

The basic pattern:

client -> virtual IP
             |
             +-- node A owns VIP while healthy
             +-- node B takes VIP if node A fails

Use this baseline:

same L2 network for both nodes
unique virtual_router_id per VIP group
higher priority on the preferred node
health check script that lowers priority or enters FAULT
track_interface for link failure
firewall permits VRRP protocol 112 between peers
manual failover test before production
monitor state transitions

Keepalived does not replicate application state. It moves an IP. The service behind that IP must already be safe to run on either node.

Know when Keepalived fits

Good fits:

  • HAProxy or Nginx active-passive frontend;
  • internal DNS resolver VIP;
  • small NFS or SMB endpoint where storage is already shared or replicated;
  • a router/firewall pair on the same LAN;
  • lab and edge systems where a full load balancer platform is unnecessary.

Bad fits:

  • cloud networks that block multicast, gratuitous ARP, or IP ownership changes;
  • services that keep local unsynchronized state;
  • database primary election by moving an IP without a real database failover plan;
  • multi-region failover;
  • public Internet anycast or BGP failover.

Many cloud providers do not support classic VRRP failover on ordinary VM NICs. Use the provider's floating IP, private IP reassignment API, load balancer, or route-table update mechanism there.

Example topology

Two load balancer nodes:

RoleHostInterfaceReal IPVIP
Preferred masterlb1ens16010.20.10.11/2410.20.10.50/24
Backuplb2ens16010.20.10.12/2410.20.10.50/24
Gatewayrouter-10.20.10.1-

Both nodes must be on the same network segment for the VIP:

ip address show dev ens160
ip route get 10.20.10.12
ping -c 3 10.20.10.12

Check from lb2 too:

ip route get 10.20.10.11
ping -c 3 10.20.10.11

Do not start with Keepalived until the real addresses can reach each other.

Install Keepalived

Ubuntu or Debian:

sudo apt update
sudo apt install -y keepalived iproute2 tcpdump curl

RHEL, Rocky Linux, AlmaLinux, or Fedora:

sudo dnf install -y keepalived iproute tcpdump curl

Enable the service, but do not start it until the config is ready:

sudo systemctl enable keepalived

Check version and options:

keepalived --version
keepalived --help

The main config file is:

/etc/keepalived/keepalived.conf

Prepare a script user

Keepalived can run notification and health scripts. Do not run scripts as root unless they need root.

Create a script user:

sudo useradd --system --no-create-home --shell /usr/sbin/nologin keepalived_script

Create a script directory:

sudo install -d -o root -g root -m 0755 /usr/local/libexec/keepalived

Keepalived supports enable_script_security, which refuses unsafe script paths. Keep scripts owned by root and avoid writable directories in the path.

Create a health check script

This example checks an HAProxy readiness endpoint. Replace the URL with the local service that must be healthy before this node should own the VIP.

[IMAGE: Supporting visual 1 for Linux High Availability With Keepalived: VRRP & Failover Configuration, showing Linux Networking decisions, examples, and Linux, Keepalived, VRRP. Alt: Linux Networking linux-high-availability-keepalived-vrrp-failover-configuration visual 1]

[IMAGE: Supporting visual 1 for Linux High Availability With Keepalived: VRRP & Failover Configuration, showing Linux Networking decisions, examples, and Linux, Keepalived, VRRP. Alt: Linux Networking linux-high-availability-keepalived-vrrp-failover-configuration visual 1]

Create:

sudoedit /usr/local/libexec/keepalived/check_haproxy.sh

Content:

#!/usr/bin/env bash
set -euo pipefail

curl -fsS --max-time 1 http://127.0.0.1:8404/healthz >/dev/null

Install permissions:

sudo chown root:root /usr/local/libexec/keepalived/check_haproxy.sh
sudo chmod 0755 /usr/local/libexec/keepalived/check_haproxy.sh

Test it:

sudo -u keepalived_script /usr/local/libexec/keepalived/check_haproxy.sh
echo $?

If the script fails from the script user but succeeds as root, fix permissions or the endpoint. Do not hide the problem by running checks as root.

Configure lb1

Back up the config:

sudo cp /etc/keepalived/keepalived.conf /etc/keepalived/keepalived.conf.$(date +%F).bak 2>/dev/null || true

Open the file on lb1:

sudoedit /etc/keepalived/keepalived.conf

Use:

global_defs {
    router_id lb1
    enable_script_security
    script_user keepalived_script
}

vrrp_script chk_haproxy {
    script "/usr/local/libexec/keepalived/check_haproxy.sh"
    interval 2
    timeout 1
    fall 2
    rise 2
    weight -80
    user keepalived_script
}

vrrp_instance VI_10 {
    state BACKUP
    interface ens160
    virtual_router_id 10
    priority 150
    advert_int 1

    unicast_src_ip 10.20.10.11
    unicast_peer {
        10.20.10.12
    }

    virtual_ipaddress {
        10.20.10.50/24 dev ens160
    }

    track_interface {
        ens160
    }

    track_script {
        chk_haproxy
    }

    notify "/usr/local/libexec/keepalived/notify.sh"
}

This uses unicast VRRP packets between the two real node IPs. If your network supports VRRP multicast reliably, unicast is not required, but unicast is easier to reason about on many server networks.

The important settings:

SettingMeaning
state BACKUPInitial state; priority elects the master
interface ens160Interface that owns the VIP
virtual_router_id 10VRRP ID shared by both peers for this VIP
priority 150Higher priority wins while healthy
advert_int 1Send VRRP advertisements every second
unicast_src_ipReal local IP used for VRRP packets
unicast_peerReal peer IPs that receive VRRP packets
virtual_ipaddressVIP added on MASTER and removed on BACKUP
track_interfaceLink failure moves instance to FAULT by default
track_scriptHealth script can reduce effective priority

The script weight is -80. If lb1 fails its check, its effective priority becomes 70, which is lower than lb2 priority 100. That makes failover possible without stopping Keepalived.

Configure lb2

Use the same script and directory setup on lb2.

Open:

sudoedit /etc/keepalived/keepalived.conf

Use:

global_defs {
    router_id lb2
    enable_script_security
    script_user keepalived_script
}

vrrp_script chk_haproxy {
    script "/usr/local/libexec/keepalived/check_haproxy.sh"
    interval 2
    timeout 1
    fall 2
    rise 2
    weight -80
    user keepalived_script
}

vrrp_instance VI_10 {
    state BACKUP
    interface ens160
    virtual_router_id 10
    priority 100
    advert_int 1

    unicast_src_ip 10.20.10.12
    unicast_peer {
        10.20.10.11
    }

    virtual_ipaddress {
        10.20.10.50/24 dev ens160
    }

    track_interface {
        ens160
    }

    track_script {
        chk_haproxy
    }

    notify "/usr/local/libexec/keepalived/notify.sh"
}

The two nodes must share:

  • virtual_router_id;
  • VIP;
  • interface network;
  • health check semantics.

They must differ in:

  • router_id;
  • priority;
  • unicast_src_ip;
  • unicast_peer.

Add state notifications

Create a simple notify script on both nodes:

sudoedit /usr/local/libexec/keepalived/notify.sh

Content:

#!/usr/bin/env bash
set -euo pipefail

TYPE="${1:-unknown}"
NAME="${2:-unknown}"
STATE="${3:-unknown}"

logger -t keepalived-notify "type=${TYPE} name=${NAME} state=${STATE}"

Install permissions:

sudo chown root:root /usr/local/libexec/keepalived/notify.sh
sudo chmod 0755 /usr/local/libexec/keepalived/notify.sh

Test:

sudo -u keepalived_script /usr/local/libexec/keepalived/notify.sh INSTANCE VI_10 MASTER
journalctl -t keepalived-notify --since '5 minutes ago' --no-pager

Keepalived passes transition details to notify scripts. Use this for logs, alerts, or controlled service hooks. Do not put slow deployment logic in state transition scripts.

Validate and start

Check config syntax on each node:

sudo keepalived --config-test=/etc/keepalived/keepalived.conf

Some older packages use:

sudo keepalived -t -f /etc/keepalived/keepalived.conf

Start lb2 first:

sudo systemctl start keepalived
sudo systemctl status keepalived --no-pager

Then start lb1:

sudo systemctl start keepalived
sudo systemctl status keepalived --no-pager

Check logs:

journalctl -u keepalived -b --no-pager
journalctl -t keepalived-notify -b --no-pager

On lb1, the VIP should appear:

ip -brief address show dev ens160
ip address show dev ens160 | grep 10.20.10.50

On lb2, the VIP should not appear while lb1 is healthy:

ip address show dev ens160 | grep 10.20.10.50 || true

Open the firewall

VRRP uses IP protocol 112. It is not TCP and not UDP.

[IMAGE: Supporting visual 2 for Linux High Availability With Keepalived: VRRP & Failover Configuration, showing Linux Networking decisions, examples, and Linux, Keepalived, VRRP. Alt: Linux Networking linux-high-availability-keepalived-vrrp-failover-configuration visual 2]

For unicast VRRP between these two nodes, nftables example:

table inet filter {
  chain input {
    type filter hook input priority 0;
    ip protocol vrrp ip saddr { 10.20.10.11, 10.20.10.12 } accept
  }
}

iptables example:

sudo iptables -A INPUT -p vrrp -s 10.20.10.11 -j ACCEPT
sudo iptables -A INPUT -p vrrp -s 10.20.10.12 -j ACCEPT

If using multicast VRRP, packets use the VRRP multicast destination for IPv4:

224.0.0.18

Capture packets:

sudo tcpdump -ni ens160 proto 112

[IMAGE: Supporting visual 2 for Linux High Availability With Keepalived: VRRP & Failover Configuration, showing Linux Networking decisions, examples, and Linux, Keepalived, VRRP. Alt: Linux Networking linux-high-availability-keepalived-vrrp-failover-configuration visual 2]

If no VRRP packets arrive on the backup node, fix firewall, routing, multicast support, or unicast peer addresses before debugging Keepalived priority logic.

Test manual failover

Watch both nodes:

watch -n1 'hostname; ip -brief address show dev ens160; systemctl is-active keepalived'

From a third host on the same network:

ping -c 3 10.20.10.50
arp -an | grep 10.20.10.50 || true
curl -I http://10.20.10.50/

Stop Keepalived on the current master:

sudo systemctl stop keepalived

On the backup node, confirm the VIP moved:

ip address show dev ens160 | grep 10.20.10.50
journalctl -u keepalived -b --no-pager | tail -80

Start Keepalived again on the original master:

sudo systemctl start keepalived

Because lb1 has higher priority, it may preempt and take the VIP back after it is healthy. If you do not want automatic failback, configure nopreempt and design a manual failback process.

Test health-check failover

On lb1, stop the checked service:

sudo systemctl stop haproxy

Watch Keepalived logs:

journalctl -u keepalived -f

Expected behavior:

chk_haproxy fails
lb1 effective priority drops
lb2 becomes MASTER
VIP appears on lb2

Confirm:

ip address show dev ens160 | grep 10.20.10.50 || true

Restart the service:

sudo systemctl start haproxy

If the VIP does not move, check the priority math:

lb1 priority 150 - weight 80 = 70
lb2 priority 100

The failed preferred node must end up below the backup node, or the backup will not take over.

Track interface state

This block:

track_interface {
    ens160
}

puts the instance into FAULT when the tracked interface goes down.

To test in a lab:

sudo ip link set ens160 down

Bring it back:

sudo ip link set ens160 up

Do not run this over SSH unless you have console access. Bringing down the management interface can lock you out.

For multi-interface systems, track the interface that determines whether this node can serve traffic:

track_interface {
    ens160
    bond0
}

Unweighted interface tracking moves to FAULT on failure. Weighted tracking can adjust priority instead, but use that only when partial service is meaningful.

Decide preempt behavior

Default behavior usually allows a higher-priority node to take the VIP back when it recovers.

That is fine when:

  • the preferred node should normally own the VIP;
  • failback is safe;
  • connection disruption is acceptable;
  • the service is stateless or clients retry cleanly.

Use nopreempt when:

  • you want the current master to stay master until a manual maintenance window;
  • failback disrupts long-lived sessions;
  • you prefer fewer transitions over preferred-node ownership.

Example:

vrrp_instance VI_10 {
    state BACKUP
    interface ens160
    virtual_router_id 10
    priority 150
    advert_int 1
    nopreempt

    virtual_ipaddress {
        10.20.10.50/24 dev ens160
    }
}

If you use nopreempt, document the manual failback command sequence.

Avoid split brain

Split brain means both nodes believe they are MASTER and both configure the VIP.

[IMAGE: Supporting visual 3 for Linux High Availability With Keepalived: VRRP & Failover Configuration, showing Linux Networking decisions, examples, and Linux, Keepalived, VRRP. Alt: Linux Networking linux-high-availability-keepalived-vrrp-failover-configuration visual 3]

Common causes:

  • VRRP packets blocked by firewall;
  • multicast blocked by switches or clouds;
  • wrong unicast_peer;
  • different virtual_router_id values;
  • different VIP definitions;
  • CPU starvation delaying VRRP processing;
  • duplicate Keepalived instances;
  • network partition.

Detect it:

ip address show dev ens160 | grep 10.20.10.50
sudo tcpdump -ni ens160 proto 112
journalctl -u keepalived -b --no-pager | grep -E 'MASTER|BACKUP|FAULT|priority'

From a third host:

arping -I eth0 10.20.10.50

If the VIP answers from two MAC addresses, stop Keepalived on one node and fix packet exchange before re-enabling automatic failover.

[IMAGE: Supporting visual 3 for Linux High Availability With Keepalived: VRRP & Failover Configuration, showing Linux Networking decisions, examples, and Linux, Keepalived, VRRP. Alt: Linux Networking linux-high-availability-keepalived-vrrp-failover-configuration visual 3]

Keep services bound correctly

A service can bind to:

0.0.0.0:80

or a real local IP:

10.20.10.11:80
10.20.10.12:80

If a service tries to bind only the VIP while the node is BACKUP, it may fail because the VIP is not present.

For HAProxy or Nginx, it is usually simpler to bind all local addresses:

bind :80
bind :443

If you must bind a service to an address before the VIP exists, you may need:

sudo sysctl -w net.ipv4.ip_nonlocal_bind=1

Persistent setting:

sudoedit /etc/sysctl.d/60-nonlocal-bind.conf

Content:

net.ipv4.ip_nonlocal_bind = 1

Apply:

sudo sysctl --system

Use this deliberately. Binding to non-local addresses can hide configuration mistakes.

Troubleshooting order

Use this sequence:

  1. Confirm real node IP connectivity.
  2. Confirm service health check succeeds locally.
  3. Validate Keepalived config.
  4. Confirm VRRP packets are exchanged.
  5. Confirm only one node owns the VIP.
  6. Confirm the service listens on the node with the VIP.
  7. Confirm clients can reach the VIP.
  8. Test stop, start, health failure, and recovery.

Commands:

sudo keepalived --config-test=/etc/keepalived/keepalived.conf
systemctl status keepalived --no-pager
journalctl -u keepalived -b --no-pager
journalctl -t keepalived-notify -b --no-pager
ip -brief address show dev ens160
ip route get 10.20.10.50
sudo tcpdump -ni ens160 proto 112
sudo -u keepalived_script /usr/local/libexec/keepalived/check_haproxy.sh
curl -I http://10.20.10.50/

Common failures:

SymptomLikely cause
Both nodes become MASTERVRRP packets blocked, wrong unicast peers, multicast unsupported
No VIP anywhereboth health checks fail, config error, service stopped, interface fault
VIP moves but service failsservice not listening, local firewall, wrong bind address
VIP does not move on health failurescript weight does not overcome priority gap
VIP moves repeatedlyflaky health check, low fall/rise values, overloaded node
Backup never takes overpriority too low, VRRP packets not received, backup service unhealthy
Clients keep old MACARP cache delay, switch behavior, gratuitous ARP issue

Production checklist

[ ] Both real node IPs are reachable.
[ ] VIP is unused before Keepalived starts.
[ ] virtual_router_id is unique on this LAN.
[ ] Both nodes define the same VIP and VRID.
[ ] Priority gap and script weights are intentional.
[ ] VRRP protocol 112 is allowed between peers.
[ ] Health checks run as a non-root script user where possible.
[ ] Health checks test real readiness, not only process existence.
[ ] Interface tracking is configured.
[ ] Manual failover has been tested.
[ ] Health-check failover has been tested.
[ ] Split-brain detection is documented.
[ ] Application state is safe on either node.
[ ] Monitoring alerts on MASTER, BACKUP, and FAULT transitions.

FAQ

What is Linux Networking?

Linux Networking is a practical networking topic that should be evaluated through implementation scope, production risk, testing, documentation, and long-term maintainability.

When should a team use Linux Networking?

Use Linux Networking when it solves a real project constraint, improves clarity, or reduces operational risk. Avoid it when it only adds novelty or hides behavior from future maintainers.

What is the biggest risk with Linux Networking?

The biggest risk is copying a pattern without its context. Production systems need clear boundaries, rollback options, tests, and observability before a technique becomes dependable.

How do you test Linux Networking?

Test the smallest unit that owns the behavior, then add integration coverage for the path users or systems actually rely on. Include failure cases, configuration differences, and regression checks.

How does Linux Networking affect SEO and AI search visibility?

It improves visibility when the article gives a direct answer, expert context, structured headings, internal links, trustworthy references, and FAQ content that matches the visible page.

Conclusion

Linux Networking is worth doing when the implementation improves clarity, reliability, or delivery speed. It is not worth doing when it hides ownership, increases operational risk, or makes the system harder to explain.

Use the framework above as a review checklist. Then connect this topic to the rest of the project documentation so readers can move from concept to implementation without losing context.

Top