使用 Keepalived 实现高可用
对于 IoT 和工业物联网(IIoT)部署,确保边缘消息服务持续可用非常关键。本指南介绍一种基于 Keepalived 和 VRRP(Virtual Router Redundancy Protocol)的 EMQX Edge 主备高可用(HA)方案。
通过该模式,你可以提供一个浮动虚拟 IP(VIP)地址。MQTT 客户端和边缘应用始终连接到该 VIP。Keepalived 会自动将流量路由到当前活跃的(Primary)EMQX Edge 节点。如果主节点发生硬件或软件故障,Keepalived 会检测到该问题,并在约 5 秒内自动将 VIP 迁移到 Standby 节点,从而实现平滑故障切换。
重要说明
这种基于 VRRP 和 VIP 的架构适用于本地裸金属服务器、本地虚拟机和私有边缘网络。标准 VRRP 和浮动 VIP 通常无法在大多数公有云 VM 环境(例如 AWS、Azure、Google Cloud)中原生工作。公有云通常会阻止 VRRP 依赖的组播/广播流量,并在虚拟交换层限制任意 MAC/IP 地址重新分配。
如果你在公有云中部署 EMQX Edge,请使用云厂商原生负载均衡器(例如 AWS NLB、Azure Load Balancer)将流量路由到节点,而不是使用 Keepalived。
架构
MQTT Clients / Edge Devices
│ (Connect strictly to VIP)
▼
VIP: 192.168.1.100:1883
│
┌────┴─────────────────────────┐
│ │
[Primary Node] [Standby Node]
EMQX Edge (MASTER) EMQX Edge (BACKUP)
IP: 192.168.1.10 IP: 192.168.1.11
Keepalived Keepalived
│ │
└─────────── VRRP ─────────────┘
(Health Heartbeat)Keepalived 会持续监控 EMQX Edge 服务的健康状态。如果主节点未通过健康检查,备用节点会被提升并接管 VIP。
前置条件
在两个节点上下载并解压 EMQX Edge。本指南以 1.3.0 版本为例:
wget https://www.emqx.com/en/downloads/emqx-edge/1.3.0/emqx-edge-1.3.0-linux-amd64.zip
unzip emqx-edge-1.3.0-linux-amd64.zip核心 HA 脚本
无论部署在裸金属机器上还是 Docker 中,Keepalived 都需要两个核心脚本。
健康检查脚本
该脚本通过检查 EMQX Edge HTTP API 验证服务是否正常运行,并在必要时回退到 TCP 端口检查。
创建 check_emqx_edge.sh:
#!/bin/bash
# Check EMQX Edge health via its HTTP API
# Returns 0 = healthy, non-zero = unhealthy
EMQX_EDGE_HOST="127.0.0.1"
HTTP_PORT="8081"
TIMEOUT=3
# Try the EMQX Edge HTTP health endpoint
HTTP_STATUS=$(curl -s -o /dev/null -w "%{http_code}" \
--max-time $TIMEOUT \
"http://${EMQX_EDGE_HOST}:${HTTP_PORT}/api/v4/")
# 401 means auth is required, but the broker is successfully responding to HTTP requests
if [ "$HTTP_STATUS" = "104" ] || [ "$HTTP_STATUS" = "200" ] || [ "$HTTP_STATUS" = "401" ]; then
exit 0
fi
# Fallback: check if the MQTT port (1883) is accepting connections
if nc -z -w $TIMEOUT $EMQX_EDGE_HOST 1883 2>/dev/null; then
exit 0
fi
# EMQX Edge is down
echo "EMQX Edge health check FAILED (HTTP: $HTTP_STATUS)"
exit 1状态通知脚本
该脚本记录状态转换,便于审计故障切换何时发生。
创建 notify.sh:
#!/bin/bash
STATE=$1
HOSTNAME=$(hostname)
echo "[$(date '+%H:%M:%S')] $HOSTNAME -> $STATE"
case "$STATE" in
MASTER)
echo ">>> This node is now ACTIVE - VIP acquired"
# Optional: Add webhook triggers or email alerts here
;;
BACKUP)
echo ">>> This node is now STANDBY"
;;
FAULT)
echo ">>> This node is FAULTED"
;;
esac为两个脚本添加执行权限:
chmod +x check_emqx_edge.sh notify.sh在裸金属或虚拟机上部署
本节介绍直接在物理机或虚拟机上部署的方式,以获得最高性能和直接网络访问能力。
本指南假设:
- 网络接口:
eth0 - Primary IP:
192.168.1.10 - Standby IP:
192.168.1.11 - Virtual IP (VIP):
192.168.1.100
安装 Keepalived
在两台 Ubuntu/Debian 机器上安装 Keepalived 和所需工具:
sudo apt-get update
sudo apt-get install -y keepalived curl netcat-openbsd iproute2部署脚本
将核心 HA 脚本中的脚本复制到两台机器的 Keepalived 目录:
sudo mkdir -p /etc/keepalived
sudo cp check_emqx_edge.sh /etc/keepalived/
sudo cp notify.sh /etc/keepalived/
sudo chmod +x /etc/keepalived/*.sh配置 Primary 节点
在 Primary 节点(192.168.1.10)上创建 /etc/keepalived/keepalived.conf:
global_defs {
router_id EMQX_EDGE_PRIMARY
script_user root
enable_script_security
}
vrrp_script check_emqx_edge {
script "/etc/keepalived/check_emqx_edge.sh"
interval 2 # Check every 2 seconds
timeout 5 # Timeout after 5 seconds
fall 2 # Require 2 failures to mark as down
rise 1 # Require 1 success to mark as up
weight -30 # Reduce priority by 30 if check fails
}
vrrp_instance EMQX_EDGE_HA {
state MASTER
interface eth0
virtual_router_id 51
priority 100 # Higher priority than Standby
advert_int 1
preempt
# Unicast is recommended to avoid multicast blocking
unicast_src_ip 192.168.1.10
unicast_peer {
192.168.1.11
}
authentication {
auth_type PASS
auth_pass REPLACE_WITH_YOUR_OWN_SHARED_VRRP_PASSWORD
}
virtual_ipaddress {
192.168.1.100/24 dev eth0
}
track_script {
check_emqx_edge
}
notify_master "/etc/keepalived/notify.sh MASTER"
notify_backup "/etc/keepalived/notify.sh BACKUP"
notify_fault "/etc/keepalived/notify.sh FAULT"
}配置 Standby 节点
在 Standby 节点(192.168.1.11)上创建 /etc/keepalived/keepalived.conf:
global_defs {
router_id EMQX_EDGE_STANDBY
script_user root
enable_script_security
}
vrrp_script check_emqx_edge {
script "/etc/keepalived/check_emqx_edge.sh"
interval 2
timeout 5
fall 2
rise 1
weight -30
}
vrrp_instance EMQX_EDGE_HA {
state BACKUP
interface eth0
virtual_router_id 51
priority 90 # Lower priority than Primary
advert_int 1
unicast_src_ip 192.168.1.11
unicast_peer {
192.168.1.10
}
authentication {
auth_type PASS
auth_pass REPLACE_WITH_YOUR_OWN_SHARED_VRRP_PASSWORD
}
virtual_ipaddress {
192.168.1.100/24 dev eth0
}
track_script {
check_emqx_edge
}
notify_master "/etc/keepalived/notify.sh MASTER"
notify_backup "/etc/keepalived/notify.sh BACKUP"
notify_fault "/etc/keepalived/notify.sh FAULT"
}启动服务
在两台机器上启动 EMQX Edge 和 Keepalived:
# Start EMQX Edge
cd /path/to/emqx-edge-1.3.0-linux-amd64
./nanomq start -d
# Start Keepalived
sudo systemctl enable keepalived
sudo systemctl start keepalived使用 Docker 部署
Docker 可将 EMQX Edge 和 Keepalived 一起打包,便于部署。Docker 环境需要特定网络权限(NET_ADMIN)用于 VIP 路由。
项目结构
按如下结构创建目录:
emqx-edge-ha/
├── docker-compose.yml
├── Dockerfile
├── entrypoint.sh
├── emqx-edge-1.3.0-linux-amd64/
│ └── (extracted EMQX Edge files)
├── HA/
│ ├── keepalived-primary.conf
│ ├── keepalived-standby.conf
│ ├── check_emqx_edge.sh
│ └── notify.sh将核心 HA 脚本中的脚本放入 HA/ 目录。
Docker 的 Keepalived 配置
Docker 部署使用子网 172.22.0.0/24。请在 HA/ 中创建配置文件,结构与裸金属示例相同,但 IP 值按下表更新。
HA/keepalived-primary.conf
| Parameter | Value |
|---|---|
unicast_src_ip | 172.22.0.10 |
unicast_peer | 172.22.0.11 |
virtual_ipaddress | 172.22.0.100/24 dev eth0 |
HA/keepalived-standby.conf
| Parameter | Value |
|---|---|
unicast_src_ip | 172.22.0.11 |
unicast_peer | 172.22.0.10 |
virtual_ipaddress | 172.22.0.100/24 dev eth0 |
Docker Entrypoint
创建 entrypoint.sh,用于在容器内启动两个服务:
#!/bin/bash
set -e
# Clean up stale PID files from unclean shutdowns
if [ -f "/tmp/nanomq/nanomq.pid" ]; then
rm -f /tmp/nanomq/nanomq.pid
fi
echo "Starting EMQX Edge..."
./nanomq start &
echo "Starting Keepalived..."
keepalived --dont-fork --log-console --log-detail &
waitDockerfile
创建 Dockerfile 构建统一镜像:
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y \
unzip curl netcat-openbsd iproute2 keepalived \
&& rm -rf /var/lib/apt/lists/*
RUN mkdir -p /opt/emqx-edge
COPY emqx-edge-1.3.0-linux-amd64/. /opt/emqx-edge/
RUN chmod +x /opt/emqx-edge/nanomq
COPY HA/check_emqx_edge.sh /etc/keepalived/check_emqx_edge.sh
COPY HA/notify.sh /etc/keepalived/notify.sh
RUN chmod +x /etc/keepalived/check_emqx_edge.sh /etc/keepalived/notify.sh
COPY entrypoint.sh /entrypoint.sh
RUN chmod +x /entrypoint.sh
WORKDIR /opt/emqx-edge
ENTRYPOINT ["/entrypoint.sh"]Docker Compose
创建 docker-compose.yml 编排双节点集群:
networks:
emqx-edge-ha-net:
driver: bridge
ipam:
config:
- subnet: 172.22.0.0/24
services:
emqx-edge-primary:
build: .
container_name: emqx-edge-primary
hostname: emqx-edge-primary
cap_add:
- NET_ADMIN # Required: lets Keepalived add/remove the VIP
- NET_BROADCAST # Required: VRRP advertisements
networks:
emqx-edge-ha-net:
ipv4_address: 172.22.0.10
volumes:
- ./HA/keepalived-primary.conf:/etc/keepalived/keepalived.conf
ports:
- "1883:1883"
restart: unless-stopped
emqx-edge-standby:
build: .
container_name: emqx-edge-standby
hostname: emqx-edge-standby
cap_add:
- NET_ADMIN
- NET_BROADCAST
networks:
emqx-edge-ha-net:
ipv4_address: 172.22.0.11
volumes:
- ./HA/keepalived-standby.conf:/etc/keepalived/keepalived.conf
ports:
- "1884:1883" # Exposed on a different host port for debugging
restart: unless-stopped构建并启动容器:
cd emqx-edge-ha
docker compose build
docker compose up -d测试 HA 故障切换
按照以下步骤验证故障切换是否正常工作。
验证初始 VIP 分配
检查当前哪个节点持有 VIP。
裸金属:
# Run on the primary node
ip addr show eth0 | grep 192.168.1.100Docker:
docker exec emqx-edge-primary ip addr show eth0 | grep 172.22.0.100
docker exec emqx-edge-standby ip addr show eth0 | grep 172.22.0.100模拟故障
停止 Primary 节点上的 EMQX Edge 进程以触发故障切换。
裸金属:
killall nanomqDocker:
docker exec emqx-edge-primary pkill nanomq观察故障切换
大约 3 到 4 秒内,Standby 节点上的 Keepalived 会通过 check_emqx_edge.sh 检测到故障并接管 VIP。
裸金属:
# Run on the standby node
ip addr show eth0 | grep 192.168.1.100Docker:
docker exec emqx-edge-standby ip addr show eth0 | grep 172.22.0.100任何连接到 VIP 的活跃 MQTT 客户端都会短暂断开并重连,随后流量会路由到 Standby 节点。
恢复 Primary 节点
重启 Primary 节点上的 EMQX Edge 服务。由于启用了 preempt,Primary 节点通过健康检查后会重新接管 VIP。
裸金属:
./nanomq start -dDocker:
docker exec emqx-edge-primary ./nanomq start