Skip to content

使用 Keepalived 实现高可用 ​

对于 IoT 和工业物联网(IIoT)部署,确保边缘消息服务持续可用非常关键。本指南介绍一种基于 Keepalived 和 VRRP(Virtual Router Redundancy Protocol)的 EMQX Edge 主备高可用(HA)方案。

通过该模式,你可以提供一个浮动虚拟 IP(VIP)地址。MQTT 客户端和边缘应用始终连接到该 VIP。Keepalived 会自动将流量路由到当前活跃的(Primary)EMQX Edge 节点。如果主节点发生硬件或软件故障,Keepalived 会检测到该问题,并在约 5 秒内自动将 VIP 迁移到 Standby 节点,从而实现平滑故障切换。

重要说明

这种基于 VRRP 和 VIP 的架构适用于本地裸金属服务器、本地虚拟机和私有边缘网络。标准 VRRP 和浮动 VIP 通常无法在大多数公有云 VM 环境(例如 AWS、Azure、Google Cloud)中原生工作。公有云通常会阻止 VRRP 依赖的组播/广播流量,并在虚拟交换层限制任意 MAC/IP 地址重新分配。

如果你在公有云中部署 EMQX Edge,请使用云厂商原生负载均衡器(例如 AWS NLB、Azure Load Balancer)将流量路由到节点,而不是使用 Keepalived。

架构 ​

MQTT Clients / Edge Devices
         │ (Connect strictly to VIP)
         ▼
   VIP: 192.168.1.100:1883
         │
    ┌────┴─────────────────────────┐
    │                              │
[Primary Node]              [Standby Node]
EMQX Edge (MASTER)          EMQX Edge (BACKUP)
IP: 192.168.1.10            IP: 192.168.1.11
Keepalived                  Keepalived
    │                              │
    └─────────── VRRP ─────────────┘
          (Health Heartbeat)

Keepalived 会持续监控 EMQX Edge 服务的健康状态。如果主节点未通过健康检查,备用节点会被提升并接管 VIP。

前置条件 ​

在两个节点上下载并解压 EMQX Edge。本指南以 1.3.0 版本为例:

bash
wget https://www.emqx.com/en/downloads/emqx-edge/1.3.0/emqx-edge-1.3.0-linux-amd64.zip
unzip emqx-edge-1.3.0-linux-amd64.zip

核心 HA 脚本 ​

无论部署在裸金属机器上还是 Docker 中,Keepalived 都需要两个核心脚本。

健康检查脚本 ​

该脚本通过检查 EMQX Edge HTTP API 验证服务是否正常运行,并在必要时回退到 TCP 端口检查。

创建 check_emqx_edge.sh:

bash
#!/bin/bash
# Check EMQX Edge health via its HTTP API
# Returns 0 = healthy, non-zero = unhealthy
EMQX_EDGE_HOST="127.0.0.1"
HTTP_PORT="8081"
TIMEOUT=3

# Try the EMQX Edge HTTP health endpoint
HTTP_STATUS=$(curl -s -o /dev/null -w "%{http_code}" \
  --max-time $TIMEOUT \
  "http://${EMQX_EDGE_HOST}:${HTTP_PORT}/api/v4/")

# 401 means auth is required, but the broker is successfully responding to HTTP requests
if [ "$HTTP_STATUS" = "104" ] || [ "$HTTP_STATUS" = "200" ] || [ "$HTTP_STATUS" = "401" ]; then
  exit 0
fi

# Fallback: check if the MQTT port (1883) is accepting connections
if nc -z -w $TIMEOUT $EMQX_EDGE_HOST 1883 2>/dev/null; then
  exit 0
fi

# EMQX Edge is down
echo "EMQX Edge health check FAILED (HTTP: $HTTP_STATUS)"
exit 1

状态通知脚本 ​

该脚本记录状态转换,便于审计故障切换何时发生。

创建 notify.sh:

bash
#!/bin/bash
STATE=$1
HOSTNAME=$(hostname)
echo "[$(date '+%H:%M:%S')] $HOSTNAME -> $STATE"
case "$STATE" in
  MASTER)
    echo ">>> This node is now ACTIVE - VIP acquired"
    # Optional: Add webhook triggers or email alerts here
    ;;
  BACKUP)
    echo ">>> This node is now STANDBY"
    ;;
  FAULT)
    echo ">>> This node is FAULTED"
    ;;
esac

为两个脚本添加执行权限:

bash
chmod +x check_emqx_edge.sh notify.sh

在裸金属或虚拟机上部署 ​

本节介绍直接在物理机或虚拟机上部署的方式,以获得最高性能和直接网络访问能力。

本指南假设:

  • 网络接口:eth0
  • Primary IP:192.168.1.10
  • Standby IP:192.168.1.11
  • Virtual IP (VIP):192.168.1.100

安装 Keepalived ​

在两台 Ubuntu/Debian 机器上安装 Keepalived 和所需工具:

bash
sudo apt-get update
sudo apt-get install -y keepalived curl netcat-openbsd iproute2

部署脚本 ​

将核心 HA 脚本中的脚本复制到两台机器的 Keepalived 目录:

bash
sudo mkdir -p /etc/keepalived
sudo cp check_emqx_edge.sh /etc/keepalived/
sudo cp notify.sh /etc/keepalived/
sudo chmod +x /etc/keepalived/*.sh

配置 Primary 节点 ​

在 Primary 节点(192.168.1.10)上创建 /etc/keepalived/keepalived.conf:

global_defs {
  router_id EMQX_EDGE_PRIMARY
  script_user root
  enable_script_security
}

vrrp_script check_emqx_edge {
  script       "/etc/keepalived/check_emqx_edge.sh"
  interval     2     # Check every 2 seconds
  timeout      5     # Timeout after 5 seconds
  fall         2     # Require 2 failures to mark as down
  rise         1     # Require 1 success to mark as up
  weight       -30   # Reduce priority by 30 if check fails
}

vrrp_instance EMQX_EDGE_HA {
  state             MASTER
  interface         eth0
  virtual_router_id 51
  priority          100  # Higher priority than Standby
  advert_int        1
  preempt

  # Unicast is recommended to avoid multicast blocking
  unicast_src_ip 192.168.1.10
  unicast_peer {
    192.168.1.11
  }

  authentication {
    auth_type PASS
    auth_pass REPLACE_WITH_YOUR_OWN_SHARED_VRRP_PASSWORD
  }

  virtual_ipaddress {
    192.168.1.100/24 dev eth0
  }

  track_script {
    check_emqx_edge
  }

  notify_master "/etc/keepalived/notify.sh MASTER"
  notify_backup "/etc/keepalived/notify.sh BACKUP"
  notify_fault  "/etc/keepalived/notify.sh FAULT"
}

配置 Standby 节点 ​

在 Standby 节点(192.168.1.11)上创建 /etc/keepalived/keepalived.conf:

global_defs {
  router_id EMQX_EDGE_STANDBY
  script_user root
  enable_script_security
}

vrrp_script check_emqx_edge {
  script       "/etc/keepalived/check_emqx_edge.sh"
  interval     2
  timeout      5
  fall         2
  rise         1
  weight       -30
}

vrrp_instance EMQX_EDGE_HA {
  state             BACKUP
  interface         eth0
  virtual_router_id 51
  priority          90   # Lower priority than Primary
  advert_int        1

  unicast_src_ip 192.168.1.11
  unicast_peer {
    192.168.1.10
  }

  authentication {
    auth_type PASS
    auth_pass REPLACE_WITH_YOUR_OWN_SHARED_VRRP_PASSWORD
  }

  virtual_ipaddress {
    192.168.1.100/24 dev eth0
  }

  track_script {
    check_emqx_edge
  }

  notify_master "/etc/keepalived/notify.sh MASTER"
  notify_backup "/etc/keepalived/notify.sh BACKUP"
  notify_fault  "/etc/keepalived/notify.sh FAULT"
}

启动服务 ​

在两台机器上启动 EMQX Edge 和 Keepalived:

bash
# Start EMQX Edge
cd /path/to/emqx-edge-1.3.0-linux-amd64
./nanomq start -d

# Start Keepalived
sudo systemctl enable keepalived
sudo systemctl start keepalived

使用 Docker 部署 ​

Docker 可将 EMQX Edge 和 Keepalived 一起打包,便于部署。Docker 环境需要特定网络权限(NET_ADMIN)用于 VIP 路由。

项目结构 ​

按如下结构创建目录:

text
emqx-edge-ha/
├── docker-compose.yml
├── Dockerfile
├── entrypoint.sh
├── emqx-edge-1.3.0-linux-amd64/
│   └── (extracted EMQX Edge files)
├── HA/
│   ├── keepalived-primary.conf
│   ├── keepalived-standby.conf
│   ├── check_emqx_edge.sh
│   └── notify.sh

将核心 HA 脚本中的脚本放入 HA/ 目录。

Docker 的 Keepalived 配置 ​

Docker 部署使用子网 172.22.0.0/24。请在 HA/ 中创建配置文件,结构与裸金属示例相同,但 IP 值按下表更新。

HA/keepalived-primary.conf

ParameterValue
unicast_src_ip172.22.0.10
unicast_peer172.22.0.11
virtual_ipaddress172.22.0.100/24 dev eth0

HA/keepalived-standby.conf

ParameterValue
unicast_src_ip172.22.0.11
unicast_peer172.22.0.10
virtual_ipaddress172.22.0.100/24 dev eth0

Docker Entrypoint ​

创建 entrypoint.sh,用于在容器内启动两个服务:

bash
#!/bin/bash
set -e

# Clean up stale PID files from unclean shutdowns
if [ -f "/tmp/nanomq/nanomq.pid" ]; then
  rm -f /tmp/nanomq/nanomq.pid
fi

echo "Starting EMQX Edge..."
./nanomq start &

echo "Starting Keepalived..."
keepalived --dont-fork --log-console --log-detail &

wait

Dockerfile ​

创建 Dockerfile 构建统一镜像:

dockerfile
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y \
    unzip curl netcat-openbsd iproute2 keepalived \
    && rm -rf /var/lib/apt/lists/*

RUN mkdir -p /opt/emqx-edge
COPY emqx-edge-1.3.0-linux-amd64/. /opt/emqx-edge/
RUN chmod +x /opt/emqx-edge/nanomq

COPY HA/check_emqx_edge.sh /etc/keepalived/check_emqx_edge.sh
COPY HA/notify.sh /etc/keepalived/notify.sh
RUN chmod +x /etc/keepalived/check_emqx_edge.sh /etc/keepalived/notify.sh

COPY entrypoint.sh /entrypoint.sh
RUN chmod +x /entrypoint.sh

WORKDIR /opt/emqx-edge
ENTRYPOINT ["/entrypoint.sh"]

Docker Compose ​

创建 docker-compose.yml 编排双节点集群:

yaml
networks:
  emqx-edge-ha-net:
    driver: bridge
    ipam:
      config:
        - subnet: 172.22.0.0/24

services:
  emqx-edge-primary:
    build: .
    container_name: emqx-edge-primary
    hostname: emqx-edge-primary
    cap_add:
      - NET_ADMIN       # Required: lets Keepalived add/remove the VIP
      - NET_BROADCAST   # Required: VRRP advertisements
    networks:
      emqx-edge-ha-net:
        ipv4_address: 172.22.0.10
    volumes:
      - ./HA/keepalived-primary.conf:/etc/keepalived/keepalived.conf
    ports:
      - "1883:1883"
    restart: unless-stopped

  emqx-edge-standby:
    build: .
    container_name: emqx-edge-standby
    hostname: emqx-edge-standby
    cap_add:
      - NET_ADMIN
      - NET_BROADCAST
    networks:
      emqx-edge-ha-net:
        ipv4_address: 172.22.0.11
    volumes:
      - ./HA/keepalived-standby.conf:/etc/keepalived/keepalived.conf
    ports:
      - "1884:1883"     # Exposed on a different host port for debugging
    restart: unless-stopped

构建并启动容器:

bash
cd emqx-edge-ha
docker compose build
docker compose up -d

测试 HA 故障切换 ​

按照以下步骤验证故障切换是否正常工作。

验证初始 VIP 分配 ​

检查当前哪个节点持有 VIP。

裸金属:

bash
# Run on the primary node
ip addr show eth0 | grep 192.168.1.100

Docker:

bash
docker exec emqx-edge-primary ip addr show eth0 | grep 172.22.0.100
docker exec emqx-edge-standby ip addr show eth0 | grep 172.22.0.100

模拟故障 ​

停止 Primary 节点上的 EMQX Edge 进程以触发故障切换。

裸金属:

bash
killall nanomq

Docker:

bash
docker exec emqx-edge-primary pkill nanomq

观察故障切换 ​

大约 3 到 4 秒内,Standby 节点上的 Keepalived 会通过 check_emqx_edge.sh 检测到故障并接管 VIP。

裸金属:

bash
# Run on the standby node
ip addr show eth0 | grep 192.168.1.100

Docker:

bash
docker exec emqx-edge-standby ip addr show eth0 | grep 172.22.0.100

任何连接到 VIP 的活跃 MQTT 客户端都会短暂断开并重连,随后流量会路由到 Standby 节点。

恢复 Primary 节点 ​

重启 Primary 节点上的 EMQX Edge 服务。由于启用了 preempt,Primary 节点通过健康检查后会重新接管 VIP。

裸金属:

bash
./nanomq start -d

Docker:

bash
docker exec emqx-edge-primary ./nanomq start