This commit is contained in:
2026-02-04 17:20:10 +08:00
23 changed files with 3095 additions and 48 deletions
+26 -3
View File
@@ -11,7 +11,9 @@ description: g3fo 系统相关文档。管理业务流程说明、服务职责
- **查业务流程**:查阅 `references/business_flows/` (如订单生命周期、用户入金等跨服务流程)。 - **查业务流程**:查阅 `references/business_flows/` (如订单生命周期、用户入金等跨服务流程)。
- **查服务职责**:查阅 `references/services/` (如 `g3fo-trade-service.md` 定义的具体功能)。 - **查服务职责**:查阅 `references/services/` (如 `g3fo-trade-service.md` 定义的具体功能)。
- **查架构概览**:查阅 `references/architecture/domain_overview.md`。 - **查架构概览**:查阅 `references/architecture/domain_overview.md`。
- **分析测试流程**:参考 `references/business_flows/_TEST_FLOW_TEMPLATE.md`。
- **查标准运维流程**:去 `references/middleware/` 或 `references/system-init/`。 - **查标准运维流程**:去 `references/middleware/` 或 `references/system-init/`。
- **查中间件总览/高可用架构**:查阅 `references/middleware/middleware-comprehensive-guide.md`。
- **查具体环境/客户资产**: - **查具体环境/客户资产**:
- 内部环境(Dev/UAT):查阅 `references/inventory/internal.md`。 - 内部环境(Dev/UAT):查阅 `references/inventory/internal.md`。
- 外部客户(客户A、B等):查阅 `references/inventory/clients/[客户名].md`。 - 外部客户(客户A、B等):查阅 `references/inventory/clients/[客户名].md`。
@@ -19,7 +21,7 @@ description: g3fo 系统相关文档。管理业务流程说明、服务职责
## 2. 问答逻辑 (AI 引导) ## 2. 问答逻辑 (AI 引导)
### A. 业务逻辑问答 ### A. 业务逻辑问答
1. **识别范围**:判断用户问题涉及哪些业务流程或服务。 1. **识别范围**:判断用户问题涉及哪些业务流程 or 服务。
2. **加载文档**: 2. **加载文档**:
- 优先查找 `business_flows/` 下的相关流程文档。 - 优先查找 `business_flows/` 下的相关流程文档。
- 结合 `services/` 下涉及的服务文档,深入了解具体职责。 - 结合 `services/` 下涉及的服务文档,深入了解具体职责。
@@ -45,23 +47,44 @@ description: g3fo 系统相关文档。管理业务流程说明、服务职责
- **格式 A (Docker 模式)**:完整的 `docker exec -it mysql mysql -uroot -pafe123456 -h127.0.0.1 -e "SQL_CONTENT"`。 - **格式 A (Docker 模式)**:完整的 `docker exec -it mysql mysql -uroot -pafe123456 -h127.0.0.1 -e "SQL_CONTENT"`。
- **格式 B (纯 SQL 模式)**:仅包含 SQL 语句本身,方便在 GUI 工具中执行。 - **格式 B (纯 SQL 模式)**:仅包含 SQL 语句本身,方便在 GUI 工具中执行。
### C. 测试流程分析 (New)
1. **需求解析**:识别用户提出的业务关键词(如“下单”、“撤单”、“充值”)。
2. **文档检索**:
- 查找 `business_flows/` 下的业务流程文档。
- 查找 `services/` 下涉及的服务定义。
3. **代码溯源**:
- 根据文档中提到的服务,在项目中搜索对应的 Controller、Service 或核心逻辑代码。
- **重点搜索路径**:
- REST 接口:`com.afe.g3fo.*.controller` 或 `com.afe.g3fo.*.api`。
- Dubbo 接口:`com.afe.g3fo.*.facade`(定义)及 `com.afe.g3fo.*.facade.impl`(实现)。
- **重点搜索注解**:`@RequestMapping`, `@PostMapping`, `@GetMapping`, `@Service`, `@DubboService` 等。
4. **表格生成**:汇总信息,输出包含“步骤”、“涉及服务”、“关键接口/代码逻辑”、“测试内容/预期结果”的表格。
- **格式参考**:`references/business_flows/_TEST_FLOW_TEMPLATE.md`。
## 3. 示例 Prompt (用户可参考) ## 3. 示例 Prompt (用户可参考)
- **业务咨询**: - **业务咨询**:
- “我想测试下单业务流程,请分析代码并列出详细的测试步骤表格。”
- “g3fo 系统中,一个订单从下单到成交会经过哪些服务?请结合 `references/business_flows/` 下的相关文档回答。” - “g3fo 系统中,一个订单从下单到成交会经过哪些服务?请结合 `references/business_flows/` 下的相关文档回答。”
- “`g3fo-margin-service` 负责哪些核心逻辑?它的上下游依赖是谁?” - “`g3fo-margin-service` 负责哪些核心逻辑?它的上下游依赖是谁?”
- **部署配置**: - **部署配置**:
- “请结合 `references/inventory/internal.md` 中的 `UAT` 环境变量,参考 `references/middleware/keepalived/deploy.md` 部署文档,为我生成 Node A 的 `keepalived.conf` 配置文件。” - “请结合 `references/inventory/internal.md` 中的 `UAT` 环境变量,参考 `references/middleware/keepalived/deploy.md` 部署文档,为我生成 Node A 的 `keepalived.conf` 配置文件。”
- “请根据 `references/middleware/middleware-comprehensive-guide.md` 和 `references/middleware/nacos/deploy.md`,结合客户资产清单生成 Nacos 部署配置。”
- **故障排查**: - **故障排查**:
- “我的 MySQL 出现了复制冲突,报错 `Duplicate entry`,请根据 `references/middleware/mysql/fault-analysis.md` 提供排查脚本和修复建议。” - “我的 MySQL 出现了复制冲突,报错 `Duplicate entry`,请根据 `references/middleware/mysql/fault-analysis.md` 提供排查脚本 and 修复建议。”
## 4. 资源地图 ## 4. 资源地图
- **业务与服务**: - **业务与服务**:
- **服务职责**: `services/` (g3fo-trade-service, g3fo-margin-service 等) - **服务职责**: `services/` (g3fo-trade-service, g3fo-margin-service 等)
- **跨服务流程**: `business_flows/` (订单生命周期等) - **跨服务流程**: `business_flows/` (订单生命周期等)
- **架构概览**: `architecture/domain_overview.md` - **架构概览**: `architecture/domain_overview.md`
- **中间件标准**: - **中间件标准**(部署/故障分析等详见各子目录):
- **总览**:`middleware/middleware-comprehensive-guide.md`
- **Keepalived**: `middleware/keepalived/` - **Keepalived**: `middleware/keepalived/`
- **MySQL**: `middleware/mysql/` - **MySQL**: `middleware/mysql/`
- **Nacos**: `middleware/nacos/`
- **PowerJob**: `middleware/powerjob/`
- **Redis**: `middleware/redis/` - **Redis**: `middleware/redis/`
- **RocketMQ**: `middleware/rocketmq/`
- **Stunnel4**: `middleware/stunnel4/`
- **系统规范**: 内核优化, 安全加固... - **系统规范**: 内核优化, 安全加固...
- **资产清单**: 内部测试机, 客户生产环境... - **资产清单**: 内部测试机, 客户生产环境...
@@ -0,0 +1,15 @@
# Test Flow Template
## Business Flow: [Flow Name]
| 步骤 | 涉及服务 | 关键接口/代码逻辑 | 测试内容/预期结果 |
| :--- | :--- | :--- | :--- |
| 1. [步骤名称] | `[服务名]` | `[类名.方法名()]` | [测试要点及预期行为描述] |
| 2. [步骤名称] | `[服务名]` | `[类名.方法名()]` | [测试要点及预期行为描述] |
### Testing Notes
- **Prerequisites**: [e.g., User must be logged in, Balance > 0]
- **Key Data Points**: [e.g., OrderID, TransactionID]
- **Edge Cases to Cover**:
- [Edge Case 1]
- [Edge Case 2]
@@ -0,0 +1,44 @@
# 待添加的中间件文档
## 待添加中间件列表
以下中间件的文档尚未添加,需要后续补充:
### 1. RocketMQ
- [x] 部署文档 (deploy.md)
- [ ] 故障分析 (fault-analysis.md)
- [ ] 使用指南 (usage.md)
### 2. Nacos
- [x] 部署文档 (deploy.md)
- [ ] 故障分析 (fault-analysis.md)
- [ ] 使用指南 (usage.md)
### 3. Dufs
- [ ] 部署文档 (deploy.md)
- [ ] 故障分析 (fault-analysis.md)
- [ ] 使用指南 (usage.md)
### 4. PowerJob
- [x] 部署文档 (deploy.md)
- [ ] 故障分析 (fault-analysis.md)
- [ ] 使用指南 (usage.md)
## 已完成的中间件
- [x] Keepalived
- [x] deploy.md
- [x] fault-analysis.md
- [x] MySQL
- [x] deploy.md
- [x] fault-analysis.md
- [x] log-management.md
- [x] Redis
- [x] deploy.md
- [x] usage.md
@@ -138,6 +138,15 @@ vrrp_script check_mysql {
rise 2 # 连续 2 次成功判定为恢复 rise 2 # 连续 2 次成功判定为恢复
} }
# 主库抢占前置检查(仅 Node A 需要)
vrrp_script check_preempt {
script "/etc/keepalived/check-preempt.sh"
interval 3 # 检测间隔:3 秒
weight -20 # 检测失败时,优先级扣 20(确保 A 低于 B)
fall 3
rise 2
}
vrrp_instance VI_1 { vrrp_instance VI_1 {
state BACKUP # 所有节点统一设为 BACKUP,靠优先级决定 Master state BACKUP # 所有节点统一设为 BACKUP,靠优先级决定 Master
interface ${INTERFACE} # 替换为宿主机实际网卡名(如 ens33/br0 等) interface ${INTERFACE} # 替换为宿主机实际网卡名(如 ens33/br0 等)
@@ -167,6 +176,7 @@ vrrp_instance VI_1 {
# 绑定检测脚本 # 绑定检测脚本
track_script { track_script {
check_mysql check_mysql
check_preempt
} }
} }
``` ```
@@ -275,6 +285,41 @@ fi
chmod +x /data/keepalived/check-mysql.sh chmod +x /data/keepalived/check-mysql.sh
``` ```
### 3.5 主库抢占前置检查脚本(仅 Node A 需要)
当主库恢复后,希望仅在“副库端口正常 + 本地复制线程正常”时才允许抢占。若副库端口不通,则仍允许 30 秒后抢占。
路径:`/data/keepalived/check-preempt.sh`
如 Keepalived 以 Docker 方式运行,请在 Node A 的容器挂载中增加:
`/data/keepalived/check-preempt.sh:/etc/keepalived/check-preempt.sh:ro`
```bash
#!/bin/sh
# 逻辑说明:
# 1. 副库端口可达时,检查本地复制线程(IO/SQL)是否正常
# 2. 副库端口不可达时,直接放行抢占(避免无可用 DB)
PEER_IP=${NODE_B_IP}
if timeout 2 nc -z "${PEER_IP}" 3306 > /dev/null 2>&1; then
STATUS=$(docker exec -i mysql mysql -uroot -pafe123456 -h127.0.0.1 -e "SHOW REPLICA STATUS\G" 2>/dev/null)
if echo "$STATUS" | grep -q "Replica_IO_Running: Yes" \
&& echo "$STATUS" | grep -q "Replica_SQL_Running: Yes"; then
exit 0
else
exit 1
fi
else
exit 0
fi
```
添加执行权限:
```bash
chmod +x /data/keepalived/check-preempt.sh
```
## 4. 启动与验证 ## 4. 启动与验证
### 4.1 启动容器 ### 4.1 启动容器
三台机器执行相同命令: 三台机器执行相同命令:
@@ -322,4 +367,5 @@ ip addr show ${INTERFACE} # 替换为实际网卡名
+ 所有节点 `virtual_router_id` 和认证信息必须一致; + 所有节点 `virtual_router_id` 和认证信息必须一致;
+ 网卡名需替换为宿主机实际名称; + 网卡名需替换为宿主机实际名称;
+ 仲裁节点 C 不配置 `virtual_ipaddress`,不挂载 MySQL 检测脚本。 + 仲裁节点 C 不配置 `virtual_ipaddress`,不挂载 MySQL 检测脚本。
+ 若使用 `check_preempt`,`weight` 需确保 A 检测失败时优先级低于 B(例如 A=110,B=100,weight=-20 → A=90)。
@@ -23,3 +23,4 @@
| 15 | C Keepalived 停机 | 正常/正常 | 正常/正常 | 故障(离线) | A=110、B=100、C=离线 | Node A | C 仅为仲裁,离线不影响 A/B 选举 | | 15 | C Keepalived 停机 | 正常/正常 | 正常/正常 | 故障(离线) | A=110、B=100、C=离线 | Node A | C 仅为仲裁,离线不影响 A/B 选举 |
| 16 | A、B MySQL 均停机 + C 停机 | 故障/正常 | 故障/正常 | 故障(离线) | A=60、B=50、C=离线 | Node A | C 离线不影响 A/B 选举 | | 16 | A、B MySQL 均停机 + C 停机 | 故障/正常 | 故障/正常 | 故障(离线) | A=60、B=50、C=离线 | Node A | C 离线不影响 A/B 选举 |
| 17 | 所有节点 Keepalived 均停机 | 正常/故障(离线) | 正常/故障(离线) | 故障(离线) | 全离线 | 无节点绑定 VIP | 无 Keepalived 参与选举,VIP 失联 | | 17 | 所有节点 Keepalived 均停机 | 正常/故障(离线) | 正常/故障(离线) | 故障(离线) | 全离线 | 无节点绑定 VIP | 无 Keepalived 参与选举,VIP 失联 |
| 18 | A 恢复但复制异常(check_preempt 失败) | 恢复/正常 | 正常/正常 | 正常 | A=90、B=100、C=40 | Node B | 副库端口正常但 A 本地复制异常,A 降权不抢占 |
@@ -0,0 +1,373 @@
# 中间件高可用部署综合指南
## 1. 文档概述
本指南基于两台高性能服务器加一台低配置服务器的组合方案,详细介绍各中间件的核心功能、高可用部署方案以及灾备转换的基本原理。所有中间件均采用容器化部署,确保部署一致性和可维护性。
## 2. 部署架构总览
### 2.1 节点配置
| 节点类型 | 数量 | 配置 | 角色分配 |
| ------------ | ---- | --------------------- | ------------------------------------ |
| 高性能服务器 | 2 | CPU/内存/存储配置较高 | 核心业务处理节点,部署所有中间件服务 |
| 低配置服务器 | 1 | CPU/内存/存储配置较低 | 仅参与选举和监控,不部署核心业务服务 |
### 2.2 整体架构图
```mermaid
graph TB
%% 机器 A
subgraph Srv_A [高性能服务器 A: .230]
direction TB
KA[Keepalived M]
AppA[业务应用集群]
subgraph MW_A [中间件/数据库]
MA[(MySQL M1)]
RA[Redis Master]
RMA[RocketMQ Broker M]
end
subgraph Cluster_Comp_A [集群仲裁/管理 A]
RS1[Redis Sentinel 1]
RNC1[RMQ NS/Controller 1]
NA[Nacos/PowerJob]
end
end
%% 机器 B
subgraph Srv_B [高性能服务器 B: .200]
direction TB
KB[Keepalived B]
AppB[业务应用集群]
subgraph MW_B [中间件/数据库]
MB[(MySQL M2)]
RB[Redis Slave]
RMB[RocketMQ Broker S]
end
subgraph Cluster_Comp_B [集群仲裁/管理 B]
RS2[Redis Sentinel 2]
RNC2[RMQ NS/Controller 2]
NB[Nacos/PowerJob]
end
end
%% 机器 C
subgraph Srv_C [仲裁服务器 C: .110]
KC[Keepalived Arbiter]
RS3[Redis Sentinel 3]
RNC3[RMQ NS/Controller 3]
end
%% VIP 逻辑
VIP((Virtual IP: .233))
KA -.->|管理| VIP
KB -.->|管理| VIP
%% 核心访问流
AppA & AppB ==> VIP
VIP ==> MA & MB
%% 跨机数据同步
MA <==>|双主复制| MB
RA --->|主从同步| RB
RMA <==>|同步| RMB
NA <==>|同步| NB
%% 集群内部协调 (逻辑表示)
RS1 --- RS2 --- RS3
RNC1 --- RNC2 --- RNC3
%% 监控选主关系 (虚线)
RS1 & RS2 & RS3 -.->|监控/选主| RA & RB
RNC1 & RNC2 & RNC3 -.->|路由/选主| RMA & RMB
%% 样式
classDef server fill:#f0f5ff,stroke:#2f54eb;
classDef arbiter fill:#fff1f0,stroke:#ff4d4f;
classDef vip fill:#f6ffed,stroke:#52c41a,stroke-width:2px;
class Srv_A,Srv_B server;
class Srv_C arbiter;
class VIP vip;
```
## 3. 中间件详细说明
### 3.1 Keepalived
#### 3.1.1 核心功能
- 实现VIP(虚拟IP)的高可用管理
- 自动故障检测和VIP切换
- 防止脑裂问题
#### 3.1.2 高可用部署方案
- **部署模式**:三节点集群(2台高性能服务器+1台低配置服务器)
- **角色分配**:
- 高性能服务器A:Master节点(基础优先级110)
- 高性能服务器B:Backup节点(基础优先级100)
- 低配置服务器C:Arbiter节点(基础优先级40,仅参与选举)
- **核心配置**:
- VIP:192.168.3.233
- 采用单播通信方式
- 集成MySQL健康检查脚本
#### 3.1.3 灾备转换原理
- **优先级规则**:VIP归属由有效优先级决定,基础优先级A > B > C
- **故障检测**:通过脚本检测MySQL端口,故障时优先级扣减50
- **脑裂防护**:Arbiter节点确保至少2/3节点正常才能进行VIP切换
- **切换流程**:
1. 检测到Master节点故障
2. Backup节点发起选举
3. 获得Arbiter节点支持后成为新Master
4. 绑定VIP并提供服务
### 3.2 MySQL
#### 3.2.1 核心功能
- 关系型数据库服务
- 数据持久化存储
- 双主双向复制
#### 3.2.2 高可用部署方案
- **部署模式**:双主(Source-Source)复制集群
- **节点分配**:仅部署在2台高性能服务器上
- **核心配置**:
- GTID + AUTO_POSITION:自动定位同步位点
- ROW模式binlog:保证复制一致性
- 自增键隔离:A节点生成奇数主键,B节点生成偶数主键
- 持久化配置:innodb_flush_log_at_trx_commit=1,sync_binlog=1
#### 3.2.3 灾备转换原理
- **双主同步**:两台节点互为主从,实时同步数据
- **故障检测**:通过Keepalived的健康检查脚本检测MySQL端口
- **切换流程**:
1. Keepalived检测到MySQL故障
2. 降低对应节点优先级
3. VIP自动切换到健康节点
4. 应用通过VIP继续访问MySQL服务
### 3.3 Redis
#### 3.3.1 核心功能
- 分布式缓存服务
- 数据持久化存储
- 高可用故障转移
#### 3.3.2 高可用部署方案
- **部署模式**:Redis Sentinel集群
- **节点分配**:
- 2台高性能服务器:部署Redis主从节点
- 3台服务器:均部署Sentinel节点
- **核心配置**:
- 主从复制:1主1从
- Sentinel集群:3节点,quorum=2
- 持久化:AOF + RDB混合模式
- 密码认证:统一密码管理
#### 3.3.3 灾备转换原理
- **健康检测**:Sentinel节点定期检查Redis主从状态
- **故障判定**:当quorum个Sentinel节点判定主节点故障时触发故障转移
- **切换流程**:
1. Sentinel集群选举Leader
2. Leader选择最优Slave节点
3. 将Slave提升为新Master
4. 配置其他Slave指向新Master
5. 通知客户端更新主节点信息
### 3.4 Nacos
#### 3.4.1 核心功能
- 服务注册与发现
- 配置中心
- Dubbo注册中心
#### 3.4.2 高可用部署方案
- **部署模式**:集群部署
- **节点分配**:部署在2台高性能服务器上,配置1个虚拟节点
- **核心配置**:
- 数据库存储:MySQL持久化元数据
- 集群规模:3节点(2实际+1虚拟)
- 数据同步:基于数据库的共享存储
#### 3.4.3 灾备转换原理
- **无状态设计**:所有节点共享同一MySQL数据库
- **故障恢复**:节点故障后,其他节点自动接管服务
- **客户端容错**:客户端配置多个Nacos地址,自动切换
### 3.5 PowerJob
#### 3.5.1 核心功能
- 分布式任务调度
- 多样化任务类型支持
- 任务管理与监控
#### 3.5.2 高可用部署方案
- **部署模式**:集群部署
- **节点分配**:部署在2台高性能服务器上
- **核心配置**:
- 数据库存储:MySQL持久化任务信息
- 集群通信:基于Akka实现节点间通信
- 负载均衡:任务自动分配到可用节点
#### 3.5.3 灾备转换原理
- **无状态设计**:所有节点共享同一MySQL数据库
- **任务容错**:节点故障时,任务自动转移到其他节点执行
- **客户端配置**:Worker节点配置多个Server地址,自动切换
### 3.6 RocketMQ
#### 3.6.1 核心功能
- 分布式消息中间件
- 高吞吐、高可用
- 支持事务消息、延时消息
- 自动主从切换
#### 3.6.2 高可用部署方案
- **部署模式**:Controller模式集群
- **节点分配**:
- 3台服务器:均部署NameServer和Controller
- 2台高性能服务器:部署Broker(1主1从)
- **核心配置**:
- Controller集群:基于jRaft实现,3节点
- Broker配置:ASYNC_MASTER + SLAVE
- 持久化:异步刷盘
#### 3.6.3 灾备转换原理
- **Controller集群**:基于Raft协议选举Leader,管理Broker元数据
- **自动主从切换**:
1. Controller监控Master Broker状态
2. 检测到故障后,选举最优Slave
3. 将Slave提升为新Master
4. 更新NameServer中的路由信息
5. 客户端自动获取新的Master地址
## 4. 灾备转换流程
### 4.1 单节点故障场景
```mermaid
graph TB
%% 故障离线节点 (服务器 A)
subgraph Srv_A [高性能服务器 A: .230 <br/> 🔴 全机离线]
direction TB
KA[Keepalived - 离线]
AppA[业务应用 - 离线]
subgraph MW_A [中间件/数据库]
MA[(MySQL - 离线)]
RA[Redis - 离线]
RMA[RMQ Broker - 离线]
end
end
%% 故障转移后的活动节点 (服务器 B)
subgraph Srv_B [高性能服务器 B: .200 <br/> 🟢 承载全量业务]
direction TB
KB[Keepalived - Master]
AppB[业务应用 - 正常]
subgraph MW_B [中间件/数据库]
MB[(MySQL - Master)]
RB[Redis - NEW Master]
RMB[RMQ Broker - NEW Master]
end
subgraph Cluster_B [管理节点]
NB[Nacos/PowerJob]
RNCB[RMQ NS/Controller]
end
end
%% 仲裁节点 (服务器 C)
subgraph Srv_C [仲裁服务器 C: .110]
KC[Keepalived Arbiter]
RS[Redis Sentinel 集群]
RNCC[RMQ NS/Controller]
end
%% VIP 与流量重定向
VIP((Virtual IP: .233))
KB ==>|接管| VIP
AppB ==>|内部访问| VIP
VIP ==>|读写| MB
%% 故障转移逻辑 (核心连线)
RS -.->|1.检测到故障| RA
RS ==>|2.提升为主| RB
RNCC -.->|1.检测到故障| RMA
RNCC ==>|2.提升为主| RMB
%% 样式定义
classDef offline fill:#f5f5f5,stroke:#d9d9d9,stroke-dasharray: 5 5,color:#bfbfbf;
classDef active fill:#e6f7ff,stroke:#1890ff,stroke-width:2px;
classDef arbiter fill:#fff1f0,stroke:#ff4d4f;
classDef highlight fill:#f6ffed,stroke:#52c41a,stroke-width:2px;
class Srv_A,KA,AppA,MA,RA,RMA offline;
class Srv_B,KB,AppB,MB,RB,RMB,NB,RNCB active;
class Srv_C,KC,RS,RNCC arbiter;
class VIP highlight;
```
### 4.2 灾备转换步骤
1. **故障检测**:各中间件自身的健康检查机制或外部监控系统检测到节点故障
2. **优先级调整**:Keepalived根据健康状态调整节点优先级
3. **VIP切换**:Keepalived将VIP绑定到优先级最高的健康节点
4. **服务接管**:
- Redis:Sentinel选举新Master
- RocketMQ:Controller自动将Slave提升为Master
- MySQL:应用通过VIP访问健康节点
- Nacos/PowerJob:健康节点自动接管服务
5. **客户端切换**:客户端通过配置的多节点地址自动连接到健康节点
## 5. 监控与维护
### 5.1 监控指标
| 中间件 | 核心监控指标 |
| ---------- | ------------------------------------------ |
| Keepalived | VIP状态、节点优先级、健康检查结果 |
| MySQL | 主从复制状态、连接数、查询响应时间、慢查询 |
| Redis | 主从状态、内存使用率、命中率、连接数 |
| Nacos | 服务注册数量、配置更新次数、节点状态 |
| PowerJob | 任务执行成功率、任务堆积量、节点状态 |
| RocketMQ | 消息堆积量、发送/消费TPS、Broker状态 |
### 5.2 维护建议
1. **定期备份**:定期备份数据库、配置文件和重要数据
2. **日志管理**:配置日志轮转,定期清理过期日志
3. **安全加固**:
- 开启访问认证
- 配置防火墙规则
- 定期更新密码
4. **性能优化**:
- 根据负载调整资源配置
- 优化中间件参数
- 定期进行性能测试
5. **灾备演练**:定期进行故障模拟演练,验证灾备转换流程
## 6. 总结
本指南基于两台高性能服务器加一台低配置服务器的组合方案,详细介绍了各中间件的高可用部署方案和灾备转换原理。通过合理的角色分配和架构设计,实现了资源的优化利用和系统的高可用性。
各中间件均采用容器化部署,通过Keepalived实现VIP的统一管理,确保了系统在节点故障时能够自动进行灾备转换,保障业务的连续性和可用性。
建议在实际部署过程中,根据业务需求和资源情况,对各中间件的配置进行适当调整和优化,以达到最佳的性能和可用性。
@@ -0,0 +1,132 @@
#!/bin/bash
TAG=""
while getopts "t:" opt; do
case $opt in
t) TAG="$OPTARG" ;;
*) echo "Invalid option: -$OPTARG" >&2; exit 1 ;;
esac
done
shift $((OPTIND-1))
cd /data/backup/mysql || exit 1
if [ -n "$TAG" ]; then
BACKUP_DIR="./$(date +%Y%m%d%H%M%S)_${TAG}"
else
BACKUP_DIR="./$(date +%Y%m%d%H%M%S)"
fi
mkdir -p "$BACKUP_DIR"
MYSQL_USER="root"
MYSQL_PASSWORD="ENC(VN13kACOdhyl7xYqbgDMTMP0fK4dwoCkFe0AQW0VEt7/tpwx96GzRVY6eIpwlASM)"
MYSQL_HOST="127.0.0.1"
if [[ "$MYSQL_PASSWORD" =~ ^ENC\(.*\)$ ]]; then
encrypted_pwd="${MYSQL_PASSWORD:4:-1}"
if ! MYSQL_PASSWORD=$(java -cp /data/secure/jasypt-agent.jar com.jasypt.agent.DecryptionCLI $encrypted_pwd \
2>&1); then
echo -e "\033[31m[Fatal Error] Password decryption failed. Reason:"
echo "$MYSQL_PASSWORD" | grep -iE "error|exception"
exit 1
fi
fi
# Function to execute MySQL command
mysql_exec() {
docker exec -i mysql mysql -u "$MYSQL_USER" -p"$MYSQL_PASSWORD" -h "$MYSQL_HOST" -N -e "$1" 2>/dev/null
}
# Function to restore MySQL state (called on exit)
restore_mysql_state() {
if [ "$READ_ONLY_SET" = "true" ]; then
echo "Restoring MySQL state..."
mysql_exec "SET GLOBAL read_only = OFF;" || echo "[Warning] Failed to set read_only = OFF"
echo "MySQL state restored"
fi
}
# Set trap to ensure MySQL state is restored even on error
trap restore_mysql_state EXIT ERR
# Pre-backup: Set read-only and stop slave
echo "Setting MySQL to read-only mode..."
if ! mysql_exec "SET GLOBAL read_only = ON;"; then
echo "[Critical Error] Failed to set read_only = ON"
exit 1
fi
READ_ONLY_SET="true"
echo "Stopping slave replication..."
if ! mysql_exec "STOP REPLICA;"; then
echo "[Warning] Failed to stop replica (may not be a slave server or already stopped)"
fi
# Get GTID position using SELECT @@global.gtid_executed;
echo "Getting GTID position..."
GTID_EXECUTED=$(mysql_exec "SELECT @@global.gtid_executed;")
if [ -n "$GTID_EXECUTED" ]; then
# Remove newlines and literal \n strings
GTID_EXECUTED=$(echo "$GTID_EXECUTED" | tr -d '\n' | sed 's/\\n//g' | tr -s ' ')
echo "GTID Executed: $GTID_EXECUTED"
# Write GTID to file for restore use
echo "$GTID_EXECUTED" > "$BACKUP_DIR/gtid_executed.txt"
else
echo "[Warning] Failed to get GTID executed information"
fi
# Get database list for backup
if [ $# -eq 0 ]; then
# When no arguments, get all non-system databases
databases=$(mysql_exec "SHOW DATABASES;" | grep -Ev "(Database|information_schema|performance_schema|mysql|skywalking|sys)")
# Convert to array and check empty values
readarray -t db_list <<< "$databases" || {
echo "ERROR: No valid databases found or failed to retrieve database list"
restore_mysql_state
exit 1
}
else
# Use arguments as database list when provided
db_list=("$@")
fi
# Check if database list is empty
if [ ${#db_list[@]} -eq 0 ]; then
echo "ERROR: No databases to backup"
restore_mysql_state
exit 1
fi
backup_list_str="$BACKUP_DIR.7z"
# Perform backup
for db in "${db_list[@]}"; do
echo "Backing up database: $db"
backup_list_str="$backup_list_str $db"
if ! docker exec -i mysql mysqldump -u "$MYSQL_USER" -p"$MYSQL_PASSWORD" -h "$MYSQL_HOST" --set-gtid-purged=OFF --single-transaction --databases "$db" > "$BACKUP_DIR/$db.sql"; then
echo "[Critical Error] Failed to backup database $db, aborting"
restore_mysql_state
exit 1
fi
done
# Post-backup: Restore MySQL state (before compression)
echo "Restoring MySQL to read-write mode..."
mysql_exec "SET GLOBAL read_only = OFF;" || echo "[Warning] Failed to set read_only = OFF"
# Clear the trap since we've successfully restored
READ_ONLY_SET="false"
trap - EXIT ERR
# Compress backup files
if ! 7zz a "$BACKUP_DIR.7z" "$BACKUP_DIR" >/dev/null; then
echo "Compression failed. Please check disk space or 7z installation"
exit 1
fi
# Cleanup temporary files
rm -rf "$BACKUP_DIR"
echo "Backup completed: $BACKUP_DIR.7z"
echo $backup_list_str >> backup_list.txt
echo "Backup process completed successfully"
@@ -4,24 +4,49 @@
> 1. **强制参数**: 所有 `docker exec` 命令必须包含 `-h127.0.0.1` 参数。 > 1. **强制参数**: 所有 `docker exec` 命令必须包含 `-h127.0.0.1` 参数。
> 2. **双重输出**: 涉及 SQL 操作时,必须同时输出 `Docker 执行命令` 和 `纯 SQL 脚本`。 > 2. **双重输出**: 涉及 SQL 操作时,必须同时输出 `Docker 执行命令` 和 `纯 SQL 脚本`。
> 3. **默认连接**: 默认使用 `-uroot -pafe123456 -h127.0.0.1`。 > 3. **默认连接**: 默认使用 `-uroot -pafe123456 -h127.0.0.1`。
> 4. **排障连贯性**: 诊断复制冲突时,在提供 **A. 查看复制状态概要** 后,**必须紧跟** 提供 **B. 自动提取错误详情** 脚本,以便用户直接获取修复建议。
## 1. 复制冲突解决(主键/唯一键重复) ## 1. 复制冲突解决(主键/唯一键重复或记录缺失)
### 1.1 快速定位冲突信息 ### 1.1 快速定位冲突信息
提取冲突 GTID、表、值(直接执行,自动解析错误日志):
当复制中断时,**必须按顺序执行以下两步**:首先查看概要,然后提取详细错误以获取修复脚本。
**A. 查看复制状态概要:**
```bash
# 方式 1: 传统状态查看 (重点关注 Last_SQL_Error)
docker exec -it mysql mysql -uroot -pafe123456 -h127.0.0.1 -e "SHOW REPLICA STATUS\G" | grep -E "Replica_.*_Running|Last_SQL_Error|Retrieved_Gtid_Set|Executed_Gtid_Set"
# 方式 2: 多线程环境下查看详情 (若启用并行复制,SHOW REPLICA STATUS 只能看到协调线程错误)
docker exec -it mysql mysql -uroot -pafe123456 -h127.0.0.1 -e "SELECT * FROM performance_schema.replication_applier_status_by_worker WHERE LAST_ERROR_NUMBER > 0\G"
```
**B. 自动提取错误详情(支持 1062 重复键 / 1032 记录未找到):**
直接在宿主机执行以下脚本,它会自动解析错误日志并生成对应的修复方案:
```bash ```bash
grep "Duplicate entry" /data/mysql/logs/mysql_error.log | perl -nle ' # 提取冲突 GTID、表、冲突值及修复建议
if (/^(\S+).*?transaction\s+'\''([^'\'']+)'\''.*?table\s+([^\s;]+);.*?Duplicate entry\s+'\''([^'\'']+)'\''\s+for\s+key\s+'\''([^'\'']+)'\''/) { grep -aE "Duplicate entry|Error_code: 1032" /data/mysql/logs/mysql_error.log | perl -nle '
print "\n" . "="x50; if (/^(\S+).*?transaction\s+[\x27"]([^\x27"]+)[\x27"].*?table\s+([^\s;]+);/) {
print "时间: $1"; my ($time, $gtid, $table) = ($1, $2, $3);
print "GTID: $2"; print "\n" . "="x60;
print "表名: $3"; print "时间: $time";
print "冲突值: $4 (索引: $5)"; print "GTID: $gtid";
print "\n方案 A (删除冲突行):"; print "表名: $table";
print "DELETE FROM $3 WHERE [主键列] = \x27$4\x27;";
print "\n方案 B (跳过此事务):"; if (/Duplicate entry\s+[\x27"]([^\x27"]+)[\x27"]\s+for\s+key\s+[\x27"]([^\x27"]+)[\x27"]/) {
print "STOP REPLICA; SET GTID_NEXT=\x27$2\x27; BEGIN; COMMIT; SET GTID_NEXT=\x27AUTOMATIC\x27; START REPLICA;"; print "类型: 1062 (主键/唯一键重复)";
print "冲突值: $1 (索引: $2)";
print "\n方案 A (手动修复 - 删除从库冲突行):";
print "DELETE FROM $table WHERE [主键列] = \x27$1\x27;";
print "\n方案 B (跳过事务 - 保持从库现状):";
print "STOP REPLICA; SET GTID_NEXT=\x27$gtid\x27; BEGIN; COMMIT; SET GTID_NEXT=\x27AUTOMATIC\x27; START REPLICA;";
} elsif (/Error_code: 1032|HA_ERR_KEY_NOT_FOUND/) {
print "类型: 1032 (记录未找到 - 通常是 Delete/Update 目标不存在)";
print "说明: 目标行在从库已不存在,Delete 操作已实质生效。";
print "\n方案 A (直接跳过 - 安全):";
print "STOP REPLICA; SET GTID_NEXT=\x27$gtid\x27; BEGIN; COMMIT; SET GTID_NEXT=\x27AUTOMATIC\x27; START REPLICA;";
}
} }
' '
``` ```
@@ -74,3 +99,122 @@ BEGIN; COMMIT;
SET GTID_NEXT = 'AUTOMATIC'; SET GTID_NEXT = 'AUTOMATIC';
START REPLICA; START REPLICA;
``` ```
## 2. 全量备份还原解决冲突(适用于严重冲突场景)
当复制冲突过多、数据不一致严重,或手动修复成本过高时,可以使用全量备份还原的方式彻底解决冲突问题。此方法通过从主库(或备份库)全量备份数据,然后在从库上完全还原,实现数据一致性。
### 2.1 相关脚本
- **备份脚本**: [`backup-sync.sh`](./backup-sync.sh) - 在主库或备份库上执行全量备份
- **还原脚本**: [`restore-sync.sh`](./restore-sync.sh) - 在从库上执行全量还原
### 2.2 操作流程
#### 步骤 1: 在备份机上执行备份
在需要备份的 MySQL 服务器(通常是主库或数据一致的节点)上执行备份脚本:
```bash
# 备份所有数据库(默认)
cd /data/backup/mysql
./backup-sync.sh
# 或指定标签备份
./backup-sync.sh -t conflict_resolve
# 或备份指定数据库
./backup-sync.sh database1 database2
```
备份脚本会自动:
- 设置 MySQL 为只读模式
- 停止复制(如果是从库)
- 记录当前 GTID 位置
- 执行全量备份并压缩为 `.7z` 文件
- 恢复 MySQL 为读写模式
备份完成后,会在 `/data/backup/mysql/` 目录下生成 `YYYYMMDDHHMMSS.7z` 或 `YYYYMMDDHHMMSS_TAG.7z` 格式的备份文件。
#### 步骤 2: 传输备份文件到还原机
使用 `scp` 将备份文件推送到需要还原的 DB 机的 `/data/backup/mysql/` 目录:
```bash
# 从备份机执行(将备份文件推送到还原机)
scp /data/backup/mysql/YYYYMMDDHHMMSS.7z user@restore-server:/data/backup/mysql/
# 示例
scp /data/backup/mysql/20250127120000.7z root@192.168.1.100:/data/backup/mysql/
```
**注意**: 确保还原机的 `/data/backup/mysql/` 目录存在且有写入权限。
#### 步骤 3: 在还原机上执行还原
在需要还原的 MySQL 服务器(从库)上执行还原脚本:
```bash
cd /data/backup/mysql
# 还原所有数据库
./restore-sync.sh YYYYMMDDHHMMSS.7z
# 或仅还原指定数据库
./restore-sync.sh YYYYMMDDHHMMSS.7z database1.sql
```
还原脚本会自动:
- 设置 MySQL 为只读模式
- 停止复制
- 重置二进制日志和 GTID
- 导入备份数据(导入时禁用 binlog,避免循环复制)
- 设置 GTID purged(从备份时的 GTID 位置)
- 恢复 MySQL 为读写模式
- **自动启动复制并检查状态**
#### 步骤 4: 在备份机上启动同步(重要)
还原完成后,还原脚本会自动启动复制并显示同步状态。确认看到同步成功后,**需要在备份机(主库)上执行以下操作**:
**A. Docker 执行模式 (包含 -h127.0.0.1):**
```bash
docker exec -it mysql mysql -uroot -pafe123456 -h127.0.0.1 -e "
START REPLICA;
"
```
**B. 纯 SQL 模式:**
```sql
START REPLICA;
```
等待 3 秒后,检查同步状态:
**A. Docker 执行模式 (包含 -h127.0.0.1):**
```bash
sleep 3
docker exec -it mysql mysql -uroot -pafe123456 -h127.0.0.1 -e "SHOW REPLICA STATUS\G"
```
**B. 纯 SQL 模式:**
```sql
-- 等待 3 秒后执行
SHOW REPLICA STATUS\G;
```
检查关键指标:
- `Replica_IO_Running`: 应为 `Yes`
- `Replica_SQL_Running`: 应为 `Yes`
- `Last_IO_Error`: 应为空
- `Last_SQL_Error`: 应为空
- `Seconds_Behind_Source`: 延迟秒数(应逐渐减小)
### 2.3 注意事项
1. **备份时机**: 建议在业务低峰期执行备份,减少对业务的影响
2. **网络传输**: 确保备份机和还原机之间网络畅通,备份文件较大时传输可能需要较长时间
3. **磁盘空间**: 确保还原机有足够的磁盘空间存放备份文件和临时解压文件
4. **权限要求**: 脚本需要 MySQL root 权限,以及 `/data/backup/mysql/` 目录的读写权限
5. **GTID 一致性**: 还原脚本会自动处理 GTID,确保复制能够正常继续
6. **备份机同步**: 还原完成后务必在备份机(主库)上执行 `START REPLICA` 并检查状态,确保双向复制正常
@@ -0,0 +1,227 @@
#!/bin/bash
# Define variables
MYSQL_USER="root"
MYSQL_PASSWORD="ENC(VN13kACOdhyl7xYqbgDMTMP0fK4dwoCkFe0AQW0VEt7/tpwx96GzRVY6eIpwlASM)"
MYSQL_HOST="127.0.0.1"
BACKUP_DIR="/data/backup/mysql"
TEMP_DIR="/tmp/mysql_restore_temp"
if [[ "$MYSQL_PASSWORD" =~ ^ENC\(.*\)$ ]]; then
encrypted_pwd="${MYSQL_PASSWORD:4:-1}"
# Decrypt password
if ! MYSQL_PASSWORD=$(java -cp /data/secure/jasypt-agent.jar com.jasypt.agent.DecryptionCLI $encrypted_pwd \
2>&1); then
echo -e "\033[31m[Fatal Error] Password decryption failed. Reason:"
echo "$MYSQL_PASSWORD" | grep -iE "error|exception"
exit 1
fi
fi
# Function to execute MySQL command
mysql_exec() {
docker exec -i mysql mysql -u "$MYSQL_USER" -p"$MYSQL_PASSWORD" -h "$MYSQL_HOST" -N -e "$1" 2>/dev/null
}
# Function to restore MySQL state (called on exit)
restore_mysql_state() {
if [ "$READ_ONLY_SET" = "true" ]; then
echo "Restoring MySQL state..."
mysql_exec "SET GLOBAL read_only = OFF;" || echo "[Warning] Failed to set read_only = OFF"
echo "MySQL state restored"
fi
}
# Check parameters
if [ $# -lt 1 ]; then
echo "Usage: $0 <backup_file.7z> [sql_file]"
echo "Example: $0 20250414105525.7z nacos.sql"
exit 1
fi
BACKUP_FILE="$1"
SQL_FILE="$2"
# Verify backup file existence
if [ ! -f "$BACKUP_DIR/$BACKUP_FILE" ]; then
echo "Error: Backup file $BACKUP_DIR/$BACKUP_FILE not found"
exit 1
fi
# Extract base backup name
BACKUP_NAME=$(basename "$BACKUP_FILE" .7z)
# Create temporary directory
mkdir -p "$TEMP_DIR"
echo "Creating temporary directory: $TEMP_DIR"
# Extract backup file
echo "Extracting $BACKUP_FILE to $TEMP_DIR..."
7zz x "$BACKUP_DIR/$BACKUP_FILE" -o"$TEMP_DIR" -y
# Check extraction status
if [ $? -ne 0 ]; then
echo "Error: Decompression failed"
rm -rf "$TEMP_DIR"
exit 1
fi
echo "Extraction completed"
# Handle nested directory
NESTED_DIR="$TEMP_DIR/$BACKUP_NAME"
if [ -d "$NESTED_DIR" ]; then
echo "Detected nested directory: $NESTED_DIR"
SQL_DIR="$NESTED_DIR"
else
echo "No nested directory found, using root extraction path"
SQL_DIR="$TEMP_DIR"
fi
# Verify SQL files existence
SQL_COUNT=$(find "$SQL_DIR" -name "*.sql" | wc -l)
if [ "$SQL_COUNT" -eq 0 ]; then
echo "Error: No SQL files found in extracted directory"
rm -rf "$TEMP_DIR"
exit 1
fi
# Pre-restore: Set read-only, stop replica, and reset GTID
echo "Setting MySQL to read-only mode..."
if ! mysql_exec "SET GLOBAL read_only = ON;"; then
echo "[Critical Error] Failed to set read_only = ON"
rm -rf "$TEMP_DIR"
exit 1
fi
READ_ONLY_SET="true"
# Set trap to ensure MySQL state is restored even on error
trap restore_mysql_state EXIT ERR
echo "Stopping replica replication..."
if ! mysql_exec "STOP REPLICA;"; then
echo "[Warning] Failed to stop replica (may not be a replica server or already stopped)"
fi
echo "Resetting binary logs and GTIDs..."
if ! mysql_exec "RESET BINARY LOGS AND GTIDS;"; then
echo "[Critical Error] Failed to reset binary logs and GTIDs"
restore_mysql_state
rm -rf "$TEMP_DIR"
exit 1
fi
# Database import logic
if [ -n "$SQL_FILE" ]; then
# Import specific SQL file
if [ -f "$SQL_DIR/$SQL_FILE" ]; then
echo "Importing $SQL_FILE..."
db_name=$(basename "$SQL_FILE" .sql)
echo "Using database: $db_name"
# Create database if not exists
docker exec -i mysql mysql --init-command="SET SQL_LOG_BIN=0;" -u"$MYSQL_USER" -p"$MYSQL_PASSWORD" -h"$MYSQL_HOST" --connect-expired-password -e "CREATE DATABASE IF NOT EXISTS \`$db_name\`;"
# Import SQL file
docker exec -i mysql mysql --init-command="SET SQL_LOG_BIN=0;" -u"$MYSQL_USER" -p"$MYSQL_PASSWORD" -h"$MYSQL_HOST" --connect-expired-password "$db_name" < "$SQL_DIR/$SQL_FILE"
if [ $? -eq 0 ]; then
echo "Successfully imported $SQL_FILE"
else
echo "Error: Failed to import $SQL_FILE"
restore_mysql_state
rm -rf "$TEMP_DIR"
exit 1
fi
else
echo "Error: SQL file $SQL_DIR/$SQL_FILE not found"
restore_mysql_state
rm -rf "$TEMP_DIR"
exit 1
fi
else
# Import all SQL files
echo "Importing all SQL files..."
for sql_file in $(find "$SQL_DIR" -name "*.sql" | sort); do
db_name=$(basename "$sql_file" .sql)
echo "Importing $sql_file to database $db_name..."
# Create database if not exists
docker exec -i mysql mysql --init-command="SET SQL_LOG_BIN=0;" -u"$MYSQL_USER" -p"$MYSQL_PASSWORD" -h"$MYSQL_HOST" --connect-expired-password -e "CREATE DATABASE IF NOT EXISTS \`$db_name\`;"
# Import SQL file
docker exec -i mysql mysql --init-command="SET SQL_LOG_BIN=0;" -u"$MYSQL_USER" -p"$MYSQL_PASSWORD" -h"$MYSQL_HOST" --connect-expired-password "$db_name" < "$sql_file"
if [ $? -eq 0 ]; then
echo "Successfully imported $sql_file to $db_name"
else
echo "Error: Failed to import $sql_file to $db_name"
restore_mysql_state
rm -rf "$TEMP_DIR"
exit 1
fi
done
fi
# Post-restore: Set GTID purged from backup file
GTID_FILE="$SQL_DIR/gtid_executed.txt"
if [ -f "$GTID_FILE" ]; then
echo "Reading GTID from $GTID_FILE..."
GTID_EXECUTED=$(cat "$GTID_FILE" | tr -d '\n' | sed 's/\\n//g' | tr -s ' ' | xargs)
if [ -n "$GTID_EXECUTED" ]; then
echo "Setting GTID purged to: $GTID_EXECUTED"
if ! mysql_exec "SET GLOBAL gtid_purged = '$GTID_EXECUTED';"; then
echo "[Critical Error] Failed to set gtid_purged"
restore_mysql_state
rm -rf "$TEMP_DIR"
exit 1
fi
echo "GTID purged set successfully"
else
echo "[Warning] GTID file is empty"
fi
else
echo "[Warning] GTID file $GTID_FILE not found, skipping GTID purged setting"
fi
# Restore MySQL to read-write mode
echo "Restoring MySQL to read-write mode..."
mysql_exec "SET GLOBAL read_only = OFF;" || echo "[Warning] Failed to set read_only = OFF"
READ_ONLY_SET="false"
# Start replica
echo "Starting replica replication..."
if ! mysql_exec "START REPLICA;"; then
echo "[Warning] Failed to start replica (may not be a replica server)"
else
echo "Waiting 3 seconds for replica to initialize..."
sleep 3
echo "Checking replica status..."
REPLICA_STATUS=$(mysql_exec "SHOW REPLICA STATUS\G")
if [ -n "$REPLICA_STATUS" ]; then
echo "Replica Status:"
echo "$REPLICA_STATUS"
# Check for errors in replica status
ERROR_COUNT=$(echo "$REPLICA_STATUS" | grep -i "Last_Error\|Last_IO_Error\|Last_SQL_Error" | grep -v "No error" | wc -l)
if [ "$ERROR_COUNT" -gt 0 ]; then
echo "[Warning] Replica status shows errors, please check the output above"
else
echo "Replica status appears normal"
fi
else
echo "[Warning] Failed to get replica status"
fi
fi
# Clear the trap since we've successfully restored
trap - EXIT ERR
# Cleanup
echo "Cleaning temporary directory..."
rm -rf "$TEMP_DIR"
echo "Database restore operation completed"
exit 0
@@ -0,0 +1,201 @@
# Nacos 部署配置
## Nacos 用途说明
Nacos 是一个更易于构建云原生应用的动态服务发现、配置管理和服务管理平台。在项目中,Nacos 主要用于以下场景:
### 1. 注册中心
- **服务注册**:所有微服务实例启动时会自动向 Nacos 注册自己的信息,包括服务名称、IP 地址、端口号等
- **服务健康检查**:Nacos 会定期检查注册的服务实例是否健康,自动剔除不健康的实例
- **高可用设计**:通过集群部署确保注册中心的可用性,避免单点故障
### 2. 服务发现
- **服务消费者查找服务**:服务消费者通过 Nacos 获取所需服务的可用实例列表
- **负载均衡**:提供服务实例的负载均衡策略,支持权重配置
- **实时更新**:服务实例变化时,Nacos 会实时通知消费者,确保服务列表的准确性
### 3. Dubbo 注册中心
- **Dubbo 服务注册**:Dubbo 服务提供者将服务注册到 Nacos
- **Dubbo 服务发现**:Dubbo 服务消费者从 Nacos 获取服务列表
- **元数据管理**:Dubbo 服务的元数据信息(如接口定义、方法签名等)存储在 Nacos 中
### 4. 配置中心
- **集中式配置管理**:所有服务的配置集中存储在 Nacos 中,避免配置分散在各个服务中
- **动态配置更新**:支持配置的动态修改,无需重启服务即可生效
- **配置版本管理**:记录配置的历史版本,支持回滚操作
- **配置分组与命名空间**:支持按环境、应用等维度管理配置
## Docker Compose 配置
```yaml
nacos:
image: nacos-registry.cn-hangzhou.cr.aliyuncs.com/nacos/nacos-server:v3.1.1
container_name: nacos
restart: always
depends_on:
- mysql
deploy:
resources:
limits:
memory: 2G
replicas: 1
placement:
constraints:
- node.role == manager
ports:
- 8848:8848
- 8849:8080
- 9848:9848
- 9849:9849
- 7848:7848
- 6848:9080
volumes:
- /etc/localtime:/etc/localtime
- /etc/timezone:/etc/timezone
- /data/nacos/conf/application.properties:/home/nacos/conf/application.properties
- /data/nacos/conf/cluster.conf:/home/nacos/conf/cluster.conf
- /data/nacos/data:/home/nacos/data
- /data/nacos/logs:/home/nacos/logs
- /data/nacos/lib:/home/nacos/lib
environment:
- JAVA_OPT=-javaagent:/home/nacos/lib/jasypt-agent.jar
- TZ=Asia/Hong_Kong
- SPRING_DATASOURCE_PLATFORM=mysql
- MODE=cluster
- NACOS_AUTH_TOKEN=LldB9CGYrmmbhIQOIFlg3L3avFB1wbUw8s1E0lL6WZc=
- NACOS_AUTH_IDENTITY_KEY=G3SF
- NACOS_AUTH_IDENTITY_VALUE=afe123456
- NACOS_SERVERS=192.168.3.230:8848 192.168.3.200:8848 192.168.4.250:8848
```
## 集群配置文件 (`/data/nacos/conf/cluster.conf`)
```
#2025-12-29T09:21:31.265596226
192.168.3.200:8848
192.168.3.230:8848
192.168.4.250:8848
```
## 应用配置文件 (`/data/nacos/conf/application.properties`)
```properties
nacos.server.main.port=8848
nacos.inetutils.prefer-hostname-over-ip=false
### Specify local server's IP:
nacos.inetutils.ip-address=192.168.3.230
spring.sql.init.platform=mysql
### Count of DB:
db.num=1
db.url.0=jdbc:mysql://192.168.3.233:3306/nacos?characterEncoding=utf8&connectTimeout=1000&socketTimeout=3000&autoReconnect=true&useUnicode=true&useSSL=false&serverTimezone=UTC
db.user=root
db.password=afe123456
management.metrics.export.elastic.enabled=false
management.metrics.export.influx.enabled=false
nacos.config.push.maxRetryTime=50
nacos.naming.empty-service.auto-clean=true
nacos.naming.empty-service.clean.initial-delay-ms=50000
nacos.naming.empty-service.clean.period-time-ms=30000
nacos.ai.mcp.registry.port=9080
nacos.server.contextPath=/nacos
server.tomcat.accesslog.enabled=true
### accesslog automatic cleaning time
server.tomcat.accesslog.max-days=30
### The access log pattern:
server.tomcat.accesslog.pattern=%h %l %u %t "%r" %s %b %D %{User-Agent}i %{Request-Source}i
server.tomcat.basedir=file:.
#*************** API Related Configurations ***************#
### Include message field
server.error.include-message=ALWAYS
#*************** Nacos Console Related Configurations ***************#
### Nacos Console Main port
nacos.console.port=8080
### Nacos Server Web context path:
nacos.console.contextPath=
### Nacos Server context path, which link to nacos server `nacos.server.contextPath`, works when deployment type is `console`
nacos.console.remote.server.context-path=/nacos
nacos.security.ignore.urls=/,/error,/**/*.css,/**/*.js,/**/*.html,/**/*.map,/**/*.svg,/**/*.png,/**/*.ico,/console-ui/public/**,/v1/auth/**,/v1/console/health/**,/actuator/**,/v1/console/server/**
### The auth system to use, default 'nacos' and 'ldap' is supported, other type should be implemented by yourself:
nacos.core.auth.system.type=nacos
### If turn on auth system:
# Whether open nacos server API auth system
nacos.core.auth.enabled=false
# Whether open nacos admin API auth system
nacos.core.auth.admin.enabled=false
# Whether open nacos console API auth system
nacos.core.auth.console.enabled=false
### Turn on/off caching of auth information. By turning on this switch, the update of auth information would have a 15 seconds delay.
nacos.core.auth.caching.enabled=true
nacos.core.auth.server.identity.key=
nacos.core.auth.server.identity.value=
nacos.core.auth.plugin.nacos.token.cache.enable=false
nacos.core.auth.plugin.nacos.token.expire.seconds=18000
### The default token (Base64 string):
#nacos.core.auth.plugin.nacos.token.secret.key=VGhpc0lzTXlDdXN0b21TZWNyZXRLZXkwMTIzNDU2Nzg=
nacos.core.auth.plugin.nacos.token.secret.key=
nacos.istio.mcp.server.enabled=false
nacos.k8s.sync.enabled=false
nacos.deployment.type=merged
```
## Spring Boot 配置连接集群
### Spring Cloud 配置
```yaml
spring:
cloud:
nacos:
# 注册中心
discovery:
enabled: true
server-addr: 192.168.3.200:8848,192.168.3.230:8848
service: ${app.service-name}-app
namespace: dev
group: app-server
```
### Dubbo 配置
```yaml
# Dubbo配置
dubbo:
registry:
address: nacos://192.168.3.230:8848,192.168.3.200:8848
metadata-report:
address: nacos://192.168.3.230:8848,192.168.3.200:8848
```
## 部署说明
1. **当前部署情况**:
- 实际部署了两台主机:`192.168.3.200` 和 `192.168.3.230`
- `192.168.4.250` 是虚拟节点,用于解决集群节点显示问题
2. **集群显示问题**:
- 配置3个节点(实际2个+1个虚拟)可以解决节点状态显示异常问题
- 当只配置2个实际节点时,会出现节点互相看不到的情况
- 添加虚拟节点后,实际节点可以正常显示在线状态
3. **数据同步**:
- 配置3个节点后,数据可以在实际节点间正常同步
- 在任一实际节点修改配置,其他实际节点可以获取到最新数据
4. **注意事项**:
- 确保MySQL数据库已正确配置并可访问
- 调整 `nacos.inetutils.ip-address` 为当前服务器的实际IP
- 根据实际部署环境调整 volumes 映射路径
- 监控Nacos集群状态,确保至少两个实际节点正常运行
@@ -0,0 +1,173 @@
# PowerJob 部署配置
## PowerJob 用途说明
PowerJob 是一个分布式任务调度与计算框架,专注于解决大规模任务调度问题,提供丰富的任务类型和强大的调度能力。在项目中,PowerJob 主要用于以下场景:
### 1. 分布式任务调度
- **定时任务调度**:支持 cron 表达式、固定延迟、固定频率等多种调度方式
- **任务分片**:将大任务拆分为多个小任务,并行执行,提高处理效率
- **任务依赖**:支持任务间的依赖关系配置,实现复杂的工作流
- **高可用设计**:通过集群部署确保任务调度的可靠性,避免单点故障
### 2. 多样化任务类型
- **HTTP 任务**:直接调用 HTTP 接口执行任务
- **Java 任务**:执行 Java 代码,支持动态加载和执行
- **Shell 任务**:执行 Shell 脚本
- **Python 任务**:执行 Python 脚本
- **MapReduce 任务**:支持大数据处理场景
### 3. 任务管理与监控
- **可视化管理界面**:提供 Web 控制台,方便查看和管理任务
- **实时监控**:实时监控任务执行状态、日志和结果
- **告警机制**:支持任务失败、超时等告警通知
- **历史记录**:保存任务执行历史,便于追溯和分析
## Docker Compose 配置
```yaml
powerjob-server:
image: powerjob/powerjob-server:v5.1.2
container_name: powerjob-server
restart: always
depends_on:
- mysql
deploy:
resources:
limits:
memory: 1536M
replicas: 1
placement:
constraints:
- node.role == manager
networks:
- my-net
ports:
- 7700:7700
- 10086:10086
- 10010:10010
- 10077:10077
volumes:
- /etc/localtime:/etc/localtime
- /etc/timezone:/etc/timezone
- /data/powerjob/data:/root/powerjob/server
- /data/powerjob/application.properties:/application.properties
- /data/powerjob/lib:/app/lib
environment:
- JVMOPTIONS=-Dpowerjob.network.external.address=192.168.3.230 -Dpowerjob.network.local.address=0.0.0.0 -javaagent:/app/lib/jasypt-agent.jar
```
## 应用配置文件 (`/data/powerjob/application.properties`)
```properties
spring.datasource.core.jdbc-url=jdbc:mysql://192.168.3.233:3306/powerjob?useUnicode=true&characterEncoding=UTF-8
spring.datasource.core.username=root
spring.datasource.core.password=ENC(RjBbbZtMkQ1XoESFu6Wo7XeC29GcDkeIiIq92ONWx2odVz9NAmClnxjRXJwq1dNI)
#spring.datasource.core.password=afe123456
jasypt.encryptor.password=1234
#jasypt.encryptor.algorithm=PBEWITHHMACSHA512ANDAES_256
#jasypt.encryptor.iv-generator-classname=org.jasypt.iv.RandomIvGenerator
oms.mongodb.enable=false
spring.profiles.active=product
```
## Spring Boot 配置连接集群
```yaml
powerjob:
worker:
app-name: ${app.service-name}
protocol: http
server-address: 192.168.3.230:7700,192.168.3.200:7700
```
## 部署说明
### 1. 端口说明
| 端口 | 作用 | 说明 |
|------|------|------|
| 7700 | PowerJob 服务器的 Web 服务端口 | 用于访问 Web 控制台,接收 Worker 注册和任务请求,必须打开 |
| 10086 | Akka 端口 | PowerJob 配置的 Akka 端口,用于内部通信 |
| 10010 | 多语言客户端 HTTP 端口 | PowerJob 配置的多语言客户端 HTTP 端口 |
| 10077 | MU 协议端口 | PowerJob 配置的 MU 协议端口,可选打开 |
### 端口配置原则
1. **最省事的方法**:所有端口(7700 + 10086 + 10010 + 10077)全打开
2. **精细化控制原则**:
- 对于任何用户,7700 端口必须打开,这是调度服务器的 Web 服务端口
- `oms.协议.port` 的端口按需打开,考虑 server-server 和 server-worker 通讯场景:
- server-server 默认通过 HTTP 协议交互(由 `oms.transporter.main.protocol` 控制),必须打开 HTTP 10010 端口
- server-worker 部分通过 HTTP,部分通过 AKKA,需要打开 AKKA 的 10086 端口
### 主要配置项说明
| 配置项 | 含义 | 是否必填 |
|--------|------|----------|
| server.port | SpringBoot 配置,HTTP 端口号,默认 7700 | 否,且不建议更改 |
| oms.transporter.active.protocols | server 需要激活的通讯协议,建议激活全部支持的协议 | 否,且不建议更改 |
| oms.transporter.main.protocol | 主要通讯协议,用于 server 与 server 之间的通讯 | 否 |
| oms.akka.port | PowerJob 配置,Akka 端口号,默认 10086 | 否,且不建议更改 |
| oms.http.port | PowerJob 配置,多语言客户端 HTTP 端口号,默认 10010 | 否,且不建议更改 |
| oms.mu.port | PowerJob 配置,MU 协议端口号,默认 10077 | 是 |
| oms.table-prefix | 自定义数据库表名前缀 | 是 |
| spring.datasource.core.xxx | 关系型数据库连接配置 | 否 |
| spring.mail.xxx | 邮件配置 | 是,未配置情况下将无法使用邮件报警功能 |
| oms.container.retention.local | 本地容器保留天数,负数代表永久保留 | 是 |
| oms.container.retention.remote | 远程容器保留天数,负数代表永久保留 | 是 |
| oms.instanceinfo.retention | 任务实例和工作流实例信息的保留天数 | 是,推荐使用默认配置,生产环境保留 7 天 |
| oms.auth.initiliaze.admin.password | 系统初始化时默认创建的超级管理员密码 | 是,默认值 powerjob_admin,无此配置时,随机生成密码 |
| oms.auth.dingtalk.* | 钉钉用户账号体系相关配置内容 | 否,仅启用钉钉账号登录体系时需要配置 |
### 2. 目录准备
在部署 PowerJob 之前,需要创建以下目录:
```bash
mkdir -p /data/powerjob/{data,lib}
```
### 3. 环境变量说明
- **JVMOPTIONS**:
- `-Dpowerjob.network.external.address`:PowerJob 服务器的外部访问地址,用于 Worker 连接
- `-Dpowerjob.network.local.address`:PowerJob 服务器的本地绑定地址
- `-javaagent:/app/lib/jasypt-agent.jar`:Jasypt 加密代理,用于解密配置文件中的敏感信息
### 4. 配置文件说明
- **spring.datasource.core.jdbc-url**:MySQL 数据库连接 URL
- **spring.datasource.core.username**:MySQL 数据库用户名
- **spring.datasource.core.password**:MySQL 数据库密码(支持 Jasypt 加密)
- **jasypt.encryptor.password**:Jasypt 加密密钥
- **oms.mongodb.enable**:是否启用 MongoDB(用于存储任务日志,默认 false)
- **spring.profiles.active**:Spring Boot 激活的配置文件
### 5. 注意事项
1. **数据库配置**:
- 确保 MySQL 数据库已创建,并且用户具有足够的权限
- 首次启动时,PowerJob 会自动创建所需的表结构
2. **集群部署**:
- 可以部署多个 PowerJob 服务器实例,通过负载均衡提高可用性
- 所有实例共享同一个 MySQL 数据库
- Worker 配置中指定多个服务器地址,实现高可用连接
3. **安全配置**:
- 生产环境中建议使用 Jasypt 加密数据库密码等敏感信息
- 调整 Jasypt 加密密钥,避免使用默认密钥
4. **资源限制**:
- 根据实际任务量调整内存限制(当前配置为 1536M)
- 监控服务器资源使用情况,及时调整配置
## 相关文档
- [PowerJob 官方文档](https://www.yuque.com/powerjob/guidence/deploy_server)
- [PowerJob GitHub 仓库](https://github.com/PowerJob/PowerJob)
@@ -0,0 +1,961 @@
# RocketMQ 5.3.2 集群部署文档
## 一、概述
### 1.1 核心功能
RocketMQ 是一款分布式消息中间件,具有高吞吐、高可用、支持事务消息、延时消息等特性。RocketMQ 5.x 引入了 Controller 模式,实现了自动主从切换,提升了集群的高可用能力。
### 1.2 Controller 模式说明
RocketMQ 5.x 引入了基于 jRaft 的 Controller 模式,实现了以下核心功能:
- **自动主从切换**:当 Master 节点故障时,Controller 自动选举新的 Master,无需人工干预
- **元数据管理**:统一管理 Topic、订阅组等元数据,避免元数据不一致
- **负载均衡**:自动进行 Broker 负载均衡,优化集群资源利用率
### 1.3 适用场景
- 生产环境下的高可用消息队列需求
- 需要事务消息、延时消息的分布式系统
- 大规模消息吞吐场景(百万级 TPS)
### 1.4 前置条件
- **运行环境**:Docker & Docker Compose
- **镜像版本**:apache/rocketmq:5.3.2
- **网络规划**:所有节点需网络互通,且需明确各宿主机的外部 IP
- **集群规模**:建议至少 3 个节点组成集群(NameServer、Broker、Controller 各 3 个实例)
---
## 二、环境准备
### 2.1 节点信息规划
| 节点 | 主机IP | brokerId | jRaftServerId | 机器配置 | 角色 |
| :---- | :------------ | :------- | :----------------- | :------- | :-------------------------------- |
| 节点1 | 192.168.3.230 | 0 | 192.168.3.230:9880 | 高性能 | Master + NameServer + Controller |
| 节点2 | 192.168.3.200 | 1 | 192.168.3.200:9880 | 高性能 | Slave + NameServer + Controller |
| 节点3 | 192.168.3.110 | - | 192.168.3.110:9880 | 低配 | NameServer + Controller(仅选举) |
**架构说明**:
- **3台机均部署 NameServer**:NameServer 作为服务注册发现中心,所有节点都需要部署以避免单点
- **2台高性能机部署 Broker**:Broker 负责消息存储和转发,对性能要求高,仅在两台高性能机上部署
- **Controller 内嵌于 NameServer**:Controller 基于 jRaft 实现,随 NameServer 一起启动,用于 Broker 的自动主从切换
- **低配机仅用于选举**:192.168.3.110 节点仅部署 NameServer(包含 Controller),参与 Controller 集群选举,不部署 Broker
### 2.2 目录结构配置
#### 2.2.1 高性能机节点(192.168.3.230、192.168.3.200)执行
```bash
# 创建 Broker 相关目录
sudo mkdir -p /data/rocketmq/broker/{logs,store}
# 创建 NameServer 相关目录
sudo mkdir -p /data/rocketmq/nameserver/{logs,data}
# 设置目录权限(RocketMQ 容器默认用户 UID/GID 为 3000)
sudo chown 3000:3000 /data/rocketmq -R
```
#### 2.2.2 低配机节点(192.168.3.110)执行
```bash
# 仅创建 NameServer 相关目录(无需 Broker 目录)
sudo mkdir -p /data/rocketmq/nameserver/{logs,data}
# 设置目录权限
sudo chown 3000:3000 /data/rocketmq -R
```
---
## 三、部署架构说明
### 3.1 单节点部署(开发/测试环境)
适用于开发测试环境,部署单个 NameServer 和 Broker 实例:
- 1 个 NameServer 实例
- 1 个 Broker 实例(Master)
- 1 个 Dashboard 实例(可选)
**优点**:部署简单,资源占用少
**缺点**:无高可用,单点故障风险
### 3.2 集群部署(生产环境推荐)
适用于生产环境,部署多个实例组成高可用集群:
- 3 个 NameServer 实例(避免单点,3台机均部署)
- 2 个 Broker 实例(1 Master + 1 Slave,仅在2台高性能机部署)
- 3 个 Controller 实例(基于 jRaft,内嵌于 NameServer)
- 1 个 Dashboard 实例(可选,建议部署在高性能机)
**优点**:高可用、自动故障转移、负载均衡、资源优化
**缺点**:部署复杂,资源占用较多
**架构优势**:
- 低配机仅承担 NameServer 和 Controller 选举职责,资源占用低
- 高性能机专注处理 Broker 消息存储和转发,性能最大化
- Controller 集群跨3台机部署,确保选举的高可用性
---
## 四、Docker Compose 配置
### 4.1 NameServer 配置(所有3台节点统一执行)
创建 `/data/rocketmq/docker-compose.yml`,添加 NameServer 服务:
```yaml
version: "3.8"
networks:
my-net:
driver: bridge
services:
rocketmq-nameserver:
image: apache/rocketmq:5.3.2
container_name: rocketmq-nameserver
restart: always
deploy:
resources:
limits:
memory: 1G
replicas: 1
placement:
constraints:
- node.role == manager
networks:
- my-net
ports:
- "9876:9876" # NameServer 默认端口
- "9880:9880" # jRaft 内部通信端口
- "9770:9770" # Controller 外部通信端口
volumes:
- /etc/localtime:/etc/localtime
- /etc/timezone:/etc/timezone
- /data/rocketmq/nameserver/logs:/home/rocketmq/logs/rocketmqlogs
- /data/rocketmq/nameserver/data:/home/rocketmq/data
- /data/rocketmq/nameserver/namesrv.conf:/home/rocketmq/conf/namesrv.conf
environment:
- TZ=Asia/Hong_Kong
command: sh mqnamesrv -c /home/rocketmq/conf/namesrv.conf
```
**注意**:NameServer 配置在所有3台节点(192.168.3.230、192.168.3.200、192.168.3.110)都需要执行。
### 4.2 Broker 配置(仅2台高性能机执行)
在 `/data/rocketmq/docker-compose.yml` 中添加 Broker 服务:
```yaml
rocketmq-broker:
image: apache/rocketmq:5.3.2
container_name: rocketmq-broker
restart: always
deploy:
resources:
limits:
memory: 2G
networks:
- my-net
environment:
- TZ=Asia/Hong_Kong
ports:
- "10911:10911" # Broker 默认服务端口
- "10909:10909" # HA 端口(Master-Slave 同步)
- "18081:18081" # gRPC 代理端口
- "18080:18080" # HTTP 代理端口
volumes:
- /etc/localtime:/etc/localtime
- /etc/timezone:/etc/timezone
- /data/rocketmq/broker/logs:/home/rocketmq/logs/rocketmqlogs
- /data/rocketmq/broker/store:/home/rocketmq/store
- /data/rocketmq/broker/broker.conf:/home/rocketmq/broker.conf
- /data/rocketmq/broker/rmq-proxy.json:/home/rocketmq/rocketmq-5.3.2/conf/rmq-proxy.json
command: sh mqbroker -c /home/rocketmq/broker.conf --enable-proxy
```
**注意**:Broker 配置仅在2台高性能机(192.168.3.230、192.168.3.200)执行,低配机(192.168.3.110)不需要部署 Broker。
### 4.3 Dashboard 配置(可选,建议在高性能机部署)
在 `/data/rocketmq/docker-compose.yml` 中添加 Dashboard 服务:
```yaml
rocketmq-dashboard:
image: apacherocketmq/rocketmq-dashboard
container_name: rocketmq-dashboard
restart: always
deploy:
resources:
limits:
memory: 512M
replicas: 1
placement:
constraints:
- node.role == manager
networks:
- my-net
ports:
- "8086:8080"
volumes:
- /etc/localtime:/etc/localtime
- /etc/timezone:/etc/timezone
environment:
- JAVA_OPTS=-Xmx512M -Xms256M -Xmn128M -Drocketmq.namesrv.addr=192.168.3.230:9876;192.168.3.200:9876;192.168.3.110:9876 -Dcom.rocketmq.sendMessageWithVIPChannel=false
```
**注意**:Dashboard 建议部署在高性能机(如 192.168.3.230),仅需部署1个实例即可。
---
## 五、核心配置文件
### 5.1 NameServer 配置(所有3台节点统一执行)
创建 `/data/rocketmq/nameserver/namesrv.conf`:
```properties
# NameServer 监听端口
listenPort = 9876
# 启用内嵌 Controller(5.x 新特性)
enableControllerInNamesrv = true
# jRaft 内核配置(三台节点统一)
controllerType = jRaft
# 集群唯一标识
jRaftGroupId = controller-jraft-group
# Raft 内部通信地址(对应容器 9880 端口)
jRaftInitConf = 192.168.3.230:9880,192.168.3.200:9880,192.168.3.110:9880
# Controller 外部通信地址(对应容器 9770 端口)
jRaftControllerRPCAddr = 192.168.3.230:9770,192.168.3.200:9770,192.168.3.110:9770
# 标志自己节点的 ServerId,必须出现在 jRaftInitConf 中
# 每个节点需要修改为对应的主机 IP
jRaftServerId = 192.168.3.200:9880
# Controller 数据存储路径
controllerStorePath = /home/rocketmq/data
```
**注意**:每个节点的 `jRaftServerId` 需要修改为对应的主机 IP,例如:
- 节点1(192.168.3.230):`jRaftServerId = 192.168.3.230:9880`
- 节点2(192.168.3.200):`jRaftServerId = 192.168.3.200:9880`
- 节点3(192.168.3.110):`jRaftServerId = 192.168.3.110:9880`
### 5.2 Broker 配置(仅2台高性能机执行)
创建 `/data/rocketmq/broker/broker.conf`:
```properties
# ========== 节点标识配置(每个节点不同) ==========
# 节点 ID,0 表示 Master,其他正整数表示 Slave
brokerId = 1
# Broker 节点名称,集群部署时同一主从对的名称相同
brokerName = broker
# ========== Controller 模式配置 ==========
# Broker Controller 模式的总开关,只有该值为 true,自动主从切换模式才会打开
enableControllerMode = true
# Controller 地址列表(多个用 ; 隔开)
controllerAddr = 192.168.3.230:9770;192.168.3.200:9770;192.168.3.110:9770
# ========== NameServer 配置 ==========
# NameServer 地址列表(多个用 ; 隔开)
namesrvAddr = 192.168.3.230:9876;192.168.3.200:9876;192.168.3.110:9876
# ========== 集群配置 ==========
# 集群名称,同一集群中必须一致
brokerClusterName = DefaultCluster
# ========== 网络配置 ==========
# Broker 对外服务的监听端口(默认 10911)
# 注意:Broker 启动后会占用 3 个端口(listenPort-2、listenPort、listenPort+1)
listenPort = 10911
# Broker 服务地址(内部使用填内网 IP,外部使用填公网 IP)
brokerIP1 = 192.168.3.200
# BrokerHAIP 地址,供 Slave 同步消息的地址
# brokerIP2 = 127.0.0.1
# ========== 主从复制配置 ==========
# Broker 角色
# ASYNC_MASTER:异步复制 Master,主写成功即响应,可能丢失少量数据
# SYNC_MASTER:同步双写 Master,主从都写成功才响应,不会丢失数据
# SLAVE:从节点
brokerRole = ASYNC_MASTER
# 刷盘方式
# SYNC_FLUSH:同步刷新,性能较差但可靠性高
# ASYNC_FLUSH:异步刷新,性能好但可能丢失少量数据
flushDiskType = ASYNC_FLUSH
# ========== 消息存储配置 ==========
# 每天什么时间删除超过保留时间的 commit log(默认 04 点)
deleteWhen = 04
# 文件保留时间(小时,默认 72 小时)
fileReservedTime = 24
# 消息最大大小(字节,默认 4MB)
maxMessageSize = 4194304
# ========== Topic 配置 ==========
# 自动创建 Topic 时的默认队列数
defaultTopicQueueNums = 4
# 是否允许 Broker 自动创建 Topic(建议线下开启,线上关闭)
autoCreateTopicEnable = true
# 是否允许 Broker 自动创建订阅组(建议线下开启,线上关闭)
autoCreateSubscriptionGroup = true
# ========== 延时消息配置 ==========
# 延时等级(从 1 开始,可自定义添加如 1d)
messageDelayLevel = 1s 5s 10s 30s 1m 2m 3m 4m 5m 6m 7m 8m 9m 10m 20m 30m 1h 2h
# ========== 事务消息配置 ==========
# TM 在 20 秒内应将最终确认状态发送给 TC,否则引发消息回查(默认 60 秒)
transactionTimeout = 20
# 最多回查 5 次,超过后丢弃消息并记录错误日志(默认 15 次)
transactionCheckMax = 5
# 消息回查的时间间隔为 10 秒(默认 60 秒)
transactionCheckInterval = 10
# ========== 其他配置 ==========
# 延时消息最大时长(秒,默认 3 天,这里设置为 7 天)
timerMaxDelaySec = 604800
# 开启消息追踪
traceTopicEnable = true
```
**注意**:每个节点需要修改以下参数:
- `brokerId`:高性能机1(192.168.3.230)为 0(Master),高性能机2(192.168.3.200)为 1(Slave)
- `brokerIP1`:当前节点的主机 IP
**配置示例**:
- 高性能机1(192.168.3.230):`brokerId = 0`,`brokerIP1 = 192.168.3.230`
- 高性能机2(192.168.3.200):`brokerId = 1`,`brokerIP1 = 192.168.3.200`
- 低配机(192.168.3.110):不需要部署 Broker
### 5.3 Proxy 配置(仅2台高性能机执行)
创建 `/data/rocketmq/broker/rmq-proxy.json`:
```json
{
"rocketMQClusterName": "DefaultCluster",
"remotingListenPort": 18080,
"grpcServerPort": 18081,
"namesrvAddr": "192.168.3.200:9876;192.168.3.230:9876;192.168.3.110:9876"
}
```
**注意**:`namesrvAddr` 可以配置为当前节点优先的 NameServer 地址列表。
---
## 六、端口说明
### 6.1 NameServer 端口
| 端口 | 用途 | 说明 |
| :--- | :---------------------- | :----------------------------------------- |
| 9876 | NameServer 服务端口 | 客户端和 Broker 连接 NameServer 的默认端口 |
| 9880 | jRaft 内部通信端口 | Controller 集群内部 Raft 协议通信 |
| 9770 | Controller 外部通信端口 | Broker 连接 Controller 的 RPC 端口 |
### 6.2 Broker 端口
| 端口 | 用途 | 说明 |
| :---- | :------------------ | :------------------------------------------ |
| 10911 | Broker 默认服务端口 | 客户端发送和接收消息的默认端口 |
| 10909 | HA 端口 | Master-Slave 主从同步端口(listenPort - 2) |
| 10912 | Fast Fail 端口 | 快速失败端口(listenPort + 1,自动占用) |
| 18080 | HTTP 代理端口 | Proxy HTTP 协议接入端口 |
| 18081 | gRPC 代理端口 | Proxy gRPC 协议接入端口 |
### 6.3 Dashboard 端口
| 端口 | 用途 | 说明 |
| :--- | :----------------- | :-------------------------- |
| 8086 | Dashboard Web 界面 | RocketMQ 管理控制台访问端口 |
**注意**:部署集群时需要确保所有节点的端口不冲突,建议提前规划端口分配。
---
## 七、Controller 模式详解
### 7.1 Controller 模式架构
Controller 模式基于 jRaft 实现,采用 Raft 一致性算法,确保元数据的一致性和高可用:
```
+-----------------+
| Client App |
+--------+--------+
|
v
+--------------------+--------------------+
| | |
+-------v-------+ +-------v-------+ +-------v-------+
| NameServer | | NameServer | | NameServer |
| (高性能机1) | | (高性能机2) | | (低配机) |
| (Node1) | | (Node2) | | (Node3) |
+-------+-------+ +-------+-------+ +-------+-------+
| | |
+--------------------+--------------------+
|
+--------v--------+
| Controller | (jRaft 集群)
| (Leader) |
+--------+--------+
|
+--------------------+
|
+-------v-------+
| Broker | (高性能机1)
| (Master) |
+---------------+
|
v
+---------------+
| Broker | (高性能机2)
| (Slave) |
+---------------+
```
**架构说明**:
- **3台 NameServer**:所有节点都部署 NameServer,确保服务注册发现的高可用
- **Controller 内嵌于 NameServer**:Controller 基于 jRaft 实现,随 NameServer 一起启动
- **2台 Broker**:仅在2台高性能机部署 Broker,低配机仅参与选举
- **自动主从切换**:当 Master 故障时,Controller 自动选举 Slave 为新 Master
### 7.2 Controller 核心功能
1. **自动主从切换**
- 监控 Broker 健康状态
- Master 故障时自动选举新 Master
- 更新 NameServer 路由信息
2. **元数据管理**
- 统一管理 Topic、订阅组等元数据
- 避免元数据不一致问题
- 支持动态扩缩容
3. **负载均衡**
- 自动进行 Broker 负载均衡
- 优化集群资源利用率
- 支持流量调度
### 7.3 jRaft 配置说明
| 配置项 | 说明 | 示例值 |
| :----------------------- | :-------------------------- | :--------------------------------------------------------- |
| `controllerType` | Controller 实现类型 | `jRaft` |
| `jRaftGroupId` | Raft 集群唯一标识 | `controller-jraft-group` |
| `jRaftInitConf` | Raft 内部通信地址列表 | `192.168.3.230:9880,192.168.3.200:9880,192.168.3.110:9880` |
| `jRaftControllerRPCAddr` | Controller 外部通信地址列表 | `192.168.3.230:9770,192.168.3.200:9770,192.168.3.110:9770` |
| `jRaftServerId` | 当前节点的 ServerId | `192.168.3.200:9880` |
**注意**:
- `jRaftServerId` 必须出现在 `jRaftInitConf` 中
- 建议至少 3 个节点组成 Raft 集群,确保高可用
- Raft 集群会自动选举 Leader,无需手动指定
---
## 八、安全配置(可选)
### 8.1 ACL 认证配置
生产环境建议开启 ACL 认证,防止未授权访问。
#### 8.1.1 启用 ACL 认证
在 `broker.conf` 中添加:
```properties
# 开启 ACL 认证
aclEnable = true
# 指定 ACL 配置文件路径
globalWhiteRemoteAddresses = 127.0.0.1
```
#### 8.1.2 创建 ACL 配置文件
创建 `/data/rocketmq/broker/plain_acl.yml`:
```yaml
# 全局白名单(允许访问的 IP 地址)
globalWhiteRemoteAddresses:
- 10.*.*.*
- 192.168.*.*
# 账户配置
accounts:
# 管理员账户
- accessKey: admin
secretKey: admin123
whiteRemoteAddress:
admin: true
# 普通用户账户
- accessKey: appuser
secretKey: appuser123
whiteRemoteAddress:
admin: false
defaultTopicPerm: PUB|SUB
defaultGroupPerm: PUB|SUB
topicPerms:
- topicA=PUB
- topicB=SUB
groupPerms:
- groupA=PUB|SUB
```
#### 8.1.3 挂载 ACL 配置文件
在 `docker-compose.yml` 中添加 ACL 配置文件挂载:
```yaml
volumes:
- /data/rocketmq/broker/plain_acl.yml:/home/rocketmq/conf/plain_acl.yml
```
### 8.2 TLS/SSL 加密(可选)
生产环境建议开启 TLS/SSL 加密,保障数据传输安全。
#### 8.2.1 生成证书
```bash
# 生成 CA 证书
openssl genrsa -out ca.key 2048
openssl req -new -x509 -days 3650 -key ca.key -out ca.crt -subj "/CN=RocketMQ CA"
# 生成服务器证书
openssl genrsa -out server.key 2048
openssl req -new -key server.key -out server.csr -subj "/CN=192.168.3.200"
openssl x509 -req -days 3650 -in server.csr -CA ca.crt -CAkey ca.key -CAcreateserial -out server.crt
# 生成客户端证书
openssl genrsa -out client.key 2048
openssl req -new -key client.key -out client.csr -subj "/CN=Client"
openssl x509 -req -days 3650 -in client.csr -CA ca.crt -CAkey ca.key -CAcreateserial -out client.crt
```
#### 8.2.2 配置 TLS
在 `broker.conf` 中添加:
```properties
# 开启 TLS
tlsTestModeEnable = false
tlsServerCert = /home/rocketmq/conf/server.crt
tlsServerKey = /home/rocketmq/conf/server.key
tlsServerAuthClient = true
tlsClientCertPath = /home/rocketmq/conf/ca.crt
```
挂载证书文件:
```yaml
volumes:
- /data/rocketmq/broker/server.crt:/home/rocketmq/conf/server.crt
- /data/rocketmq/broker/server.key:/home/rocketmq/conf/server.key
- /data/rocketmq/broker/ca.crt:/home/rocketmq/conf/ca.crt
```
---
## 九、启动和验证
### 9.1 启动服务
#### 9.1.1 启动 NameServer(所有3台节点执行)
```bash
# 进入目录
cd /data/rocketmq
# 启动 NameServer
docker compose up -d rocketmq-nameserver
# 查看启动状态
docker compose ps
```
#### 9.1.2 启动 Broker(仅2台高性能机执行)
```bash
# 进入目录
cd /data/rocketmq
# 启动 Broker
docker compose up -d rocketmq-broker
# 查看启动状态
docker compose ps
```
#### 9.1.3 启动 Dashboard(可选,建议在高性能机执行)
```bash
# 进入目录
cd /data/rocketmq
# 启动 Dashboard
docker compose up -d rocketmq-dashboard
# 查看启动状态
docker compose ps
```
**启动顺序建议**:
1. 先在所有3台节点启动 NameServer
2. 等待 NameServer 启动完成(约 30 秒)
3. 再在2台高性能机启动 Broker
4. 最后启动 Dashboard(可选)
### 9.2 验证 NameServer
```bash
# 查看日志(所有3台节点执行)
docker logs rocketmq-nameserver
# 测试连接
telnet 192.168.3.230 9876
telnet 192.168.3.200 9876
telnet 192.168.3.110 9876
```
### 9.3 验证 Broker
```bash
# 查看日志(仅2台高性能机执行)
docker logs rocketmq-broker
# 检查 Broker 是否注册到 NameServer
docker exec rocketmq-nameserver sh mqadmin clusterList -n 192.168.3.230:9876
# 查看集群状态(应该看到 2 个 Broker)
docker exec rocketmq-nameserver sh mqadmin brokerStatus -n 192.168.3.230:9876
```
### 9.4 验证 Controller
```bash
# 查看 Controller 状态(所有3台节点执行)
docker exec rocketmq-nameserver sh mqadmin getControllerMode -n 192.168.3.230:9876
# 查看 Controller 集群状态
docker exec rocketmq-nameserver sh mqadmin getControllerInfo -n 192.168.3.230:9876
```
**预期结果**:
- Controller 模式应显示为 `ENABLED`
- Controller 集群应选举出 Leader(3台节点中选1个)
### 9.5 访问 Dashboard
浏览器访问:`http://192.168.3.230:8086`
默认账号密码:`admin / admin`
**验证内容**:
- 集群概览中应显示 2 个 Broker
- NameServer 列表中应显示 3 个节点
- Controller 状态应为正常
---
## 十、常见问题
### 10.1 Broker 无法连接 NameServer
**现象**:Broker 日志显示连接 NameServer 失败
**排查步骤**:
1. 检查 `namesrvAddr` 配置是否正确
2. 检查防火墙是否开放 9876 端口
3. 检查网络连通性:`telnet <nameserver-ip> 9876`
### 10.2 Controller 集群无法选举 Leader
**现象**:Controller 集群一直处于选举状态
**排查步骤**:
1. 检查 `jRaftInitConf` 和 `jRaftServerId` 配置是否正确
2. 检查 9880 和 9770 端口是否开放
3. 检查节点时间是否同步(NTP)
4. 查看 NameServer 日志:`docker logs rocketmq-nameserver`
### 10.3 主从切换失败
**现象**:Master 故障后无法自动切换
**排查步骤**:
1. 检查 `enableControllerMode` 是否为 `true`
2. 检查 Controller 集群状态是否正常
3. 检查 Slave 节点是否正常运行
4. 查看 Broker 日志:`docker logs rocketmq-broker`
### 10.4 消息发送失败
**现象**:客户端发送消息超时或失败
**排查步骤**:
1. 检查 NameServer 地址配置是否正确
2. 检查 Topic 是否存在(`autoCreateTopicEnable` 是否开启)
3. 检查 Broker 是否正常注册到 NameServer
4. 检查网络连通性和防火墙规则
### 10.5 磁盘空间不足
**现象**:Broker 日志显示磁盘空间不足
**解决方案**:
1. 调整 `fileReservedTime` 参数,缩短消息保留时间
2. 调整 `deleteWhen` 参数,增加清理频率
3. 扩容磁盘空间
4. 手动清理过期消息:`docker exec rocketmq-broker sh mqadmin cleanExpiredCQFile -n <nameserver-addr>`
---
## 十一、性能优化建议
### 11.1 JVM 参数优化
在 `docker-compose.yml` 中调整 JVM 参数:
```yaml
environment:
- JAVA_OPT_EXT=-Xms2g -Xmx2g -Xmn1g -XX:MetaspaceSize=128m -XX:MaxMetaspaceSize=320m
```
### 11.2 操作系统优化
```bash
# 增加文件描述符限制
ulimit -n 65535
# 优化 TCP 参数
echo 'net.core.somaxconn = 1024' >> /etc/sysctl.conf
echo 'net.ipv4.tcp_max_syn_backlog = 2048' >> /etc/sysctl.conf
sysctl -p
```
### 11.3 网络优化
- 使用万兆网络(10Gbps)
- 部署在同一机房,减少网络延迟
- 使用专用网络,避免公网访问
### 11.4 存储优化
- 使用 SSD 存储,提升 IO 性能
- 将日志和数据分离到不同磁盘
- 定期清理过期消息,避免磁盘满
---
## 十二、监控和告警
### 12.1 关键监控指标
| 指标 | 说明 | 告警阈值 |
| :---------------- | :------------------- | :------- |
| 消息堆积量 | 消息未消费数量 | > 10000 |
| 消息发送 TPS | 每秒发送消息数 | 异常波动 |
| 消息消费 TPS | 每秒消费消息数 | 异常波动 |
| Broker CPU 使用率 | Broker 进程 CPU 占用 | > 80% |
| Broker 内存使用率 | Broker 进程内存占用 | > 80% |
| 磁盘使用率 | 数据目录磁盘占用 | > 80% |
| 网络流量 | 网络入出流量 | 异常波动 |
### 12.2 日志监控
- Broker 日志:`/data/rocketmq/broker/logs/`
- NameServer 日志:`/data/rocketmq/nameserver/logs/`
建议使用 ELK 或 Loki 等日志收集系统进行集中管理。
---
## 十三、备份和恢复
### 13.1 数据备份
```bash
# 备份 Broker 数据
tar -czf rocketmq-broker-backup-$(date +%Y%m%d).tar.gz /data/rocketmq/broker/store
# 备份配置文件
tar -czf rocketmq-config-backup-$(date +%Y%m%d).tar.gz /data/rocketmq/broker/*.conf /data/rocketmq/broker/*.json
```
### 13.2 数据恢复
```bash
# 停止 Broker
docker compose stop rocketmq-broker
# 恢复数据
tar -xzf rocketmq-broker-backup-20240121.tar.gz -C /
# 启动 Broker
docker compose start rocketmq-broker
```
---
## 十四、升级和扩容
### 14.1 版本升级
```bash
# 1. 停止服务
docker compose down
# 2. 备份数据
tar -czf rocketmq-backup-$(date +%Y%m%d).tar.gz /data/rocketmq
# 3. 更新镜像版本
sed -i 's/apache\/rocketmq:5.3.2/apache\/rocketmq:5.3.3/g' docker-compose.yml
# 4. 启动服务
docker compose up -d
# 5. 验证服务
docker compose ps
```
### 14.2 集群扩容
```bash
# 1. 在新节点创建目录
sudo mkdir -p /data/rocketmq/broker/{logs,store}
sudo mkdir -p /data/rocketmq/nameserver/{logs,data}
sudo chown 3000:3000 /data/rocketmq -R
# 2. 复制配置文件到新节点
scp /data/rocketmq/broker/broker.conf root@<new-node>:/data/rocketmq/broker/
scp /data/rocketmq/nameserver/namesrv.conf root@<new-node>:/data/rocketmq/nameserver/
# 3. 修改新节点配置(brokerId、brokerIP1、jRaftServerId 等)
# 4. 在新节点启动服务
cd /data/rocketmq && docker compose up -d
# 5. 验证新节点注册
docker exec rocketmq-nameserver sh mqadmin clusterList -n 192.168.3.230:9876
```
---
## 十五、总结
### 15.1 部署检查清单
#### 15.1.1 所有3台节点(192.168.3.230、192.168.3.200、192.168.3.110)
- [ ] NameServer 目录创建完成(`/data/rocketmq/nameserver/{logs,data}`)
- [ ] NameServer 目录权限设置正确(`chown 3000:3000 /data/rocketmq -R`)
- [ ] NameServer 配置文件正确(`/data/rocketmq/nameserver/namesrv.conf`)
- [ ] jRaftServerId 配置唯一(每个节点对应自己的 IP)
- [ ] Docker Compose 配置文件正确(包含 NameServer 服务)
- [ ] 防火墙规则配置正确(9876、9880、9770 端口开放)
- [ ] NameServer 启动成功,日志无错误
- [ ] NameServer 之间网络互通
#### 15.1.2 高性能机节点(192.168.3.230、192.168.3.200)
- [ ] Broker 目录创建完成(`/data/rocketmq/broker/{logs,store}`)
- [ ] Broker 目录权限设置正确(`chown 3000:3000 /data/rocketmq -R`)
- [ ] Broker 配置文件正确(`/data/rocketmq/broker/broker.conf`)
- [ ] brokerId 配置正确(192.168.3.230 为 0,192.168.3.200 为 1)
- [ ] brokerIP1 配置正确(对应当前节点 IP)
- [ ] Proxy 配置文件正确(`/data/rocketmq/broker/rmq-proxy.json`)
- [ ] Docker Compose 配置文件正确(包含 Broker 服务)
- [ ] 防火墙规则配置正确(10911、10909、18080、18081 端口开放)
- [ ] Broker 启动成功,日志无错误
- [ ] Broker 成功注册到所有 NameServer
#### 15.1.3 Dashboard 节点(建议在 192.168.3.230)
- [ ] Docker Compose 配置文件正确(包含 Dashboard 服务)
- [ ] 防火墙规则配置正确(8086 端口开放)
- [ ] Dashboard 启动成功
- [ ] Dashboard 可以正常访问(`http://192.168.3.230:8086`)
#### 15.1.4 集群整体验证
- [ ] 3 个 NameServer 都正常运行
- [ ] 2 个 Broker 都正常运行
- [ ] Controller 集群选举成功(有 1 个 Leader)
- [ ] Controller 模式已启用(`ENABLED`)
- [ ] Broker 主从关系正常(1 Master + 1 Slave)
- [ ] Dashboard 显示集群状态正常
### 15.2 最佳实践
1. **架构设计**:
- 3台 NameServer 确保服务注册发现的高可用
- 2台 Broker 部署在高性能机,低配机仅参与选举
- Controller 跨3台机部署,确保选举的高可用性
2. **数据可靠性**:根据业务需求选择 `brokerRole` 和 `flushDiskType`
3. **监控告警**:建立完善的监控和告警机制,及时发现异常
4. **定期备份**:定期备份配置文件和数据,防止数据丢失
5. **安全加固**:生产环境建议开启 ACL 认证和 TLS 加密
6. **性能优化**:根据实际负载调整 JVM 参数和系统参数
7. **日志管理**:定期清理日志,避免磁盘满
8. **资源规划**:低配机仅部署 NameServer,高性能机部署 Broker,资源利用率最大化
---
**参考资料**:
- [RocketMQ 官方文档](https://rocketmq.apache.org/zh/docs/)
- [RocketMQ 5.x Controller 模式介绍](https://rocketmq.apache.org/zh/docs/featureBehavior/05controller)
- [jRaft 官方文档](https://github.com/sofastack/sofa-jraft)
@@ -0,0 +1,171 @@
# stunnel4 中间件部署与配置指南
## 1. 概述
stunnel4 是一个用于加密 TCP 连接的开源工具,在本项目中仅用于连接 OSL 交易所。它通过在客户端和服务器之间建立 SSL/TLS 隧道,确保数据传输的安全性。
## 2. 安装
### 2.1 下载地址
- 通用下载地址:[https://pkgs.org/search/?q=stunnel4](https://pkgs.org/search/?q=stunnel4)
- Ubuntu 24.04 LTS 版本下载地址:[http://archive.ubuntu.com/ubuntu/pool/universe/s/stunnel4/stunnel4_5.72-1build2_amd64.deb](http://archive.ubuntu.com/ubuntu/pool/universe/s/stunnel4/stunnel4_5.72-1build2_amd64.deb)
### 2.2 安装流程
1. 下载 deb 包到服务器
```bash
wget http://archive.ubuntu.com/ubuntu/pool/universe/s/stunnel4/stunnel4_5.72-1build2_amd64.deb
```
2. 安装 deb 包
```bash
sudo dpkg -i stunnel4_5.72-1build2_amd64.deb
```
3. 安装依赖(如果需要)
```bash
sudo apt-get install -f
```
## 3. 配置
### 3.1 UAT 环境配置
配置文件路径:`/etc/stunnel/stunnel.conf`
```ini
socket = l:TCP_NODELAY=1
socket = r:TCP_NODELAY=1
sslVersionMin = TLSv1.2
sslVersionMax = all
TIMEOUTconnect = 30
delay = yes
debug = 7
cert = /etc/stunnel/stunnel_client_dlsec.crt
key = /etc/stunnel/stunnel_client.key
output = /var/log/stunnel4/stunnel.log
[om]
sni = uat-ndsdlsecom-osl
client = yes
accept = 0.0.0.0:440
connect = fix-test.oslsandbox.com:443
[dc]
sni = uat-ndsdlsecdc-osl
client = yes
accept = 0.0.0.0:441
connect = fix-test.oslsandbox.com:443
```
### 3.2 生产环境配置
配置文件路径:`/etc/stunnel/stunnel.conf.pro`
```ini
socket = l:TCP_NODELAY=1
socket = r:TCP_NODELAY=1
sslVersionMin = TLSv1.2
sslVersionMax = all
TIMEOUTconnect = 30
delay = yes
debug = 7
cert = /etc/stunnel/stunnel_client.crt
key = /etc/stunnel/stunnel_client.key
output = /var/log/stunnel4/stunnel.log
[om]
sni = prod-ndsdlsecom-osl
client = yes
accept = 0.0.0.0:440
connect = fixtrade.osl.com:443
[dc]
sni = prod-ndsdlsecdc-osl
client = yes
accept = 0.0.0.0:441
connect = fixtrade.osl.com:443
```
## 4. 环境切换
### 4.1 从 UAT 切换到生产环境
1. 停止 stunnel4 服务
```bash
sudo systemctl stop stunnel4
```
2. 查找并终止占用端口 440 的进程
```bash
sudo netstat -aonp | grep 440
# 找到对应的 PID,使用以下命令终止进程
sudo kill -9 {PID}
```
3. 重命名配置文件
```bash
sudo mv /etc/stunnel/stunnel.conf /etc/stunnel/stunnel.conf.uat
sudo mv /etc/stunnel/stunnel.conf.pro /etc/stunnel/stunnel.conf
```
4. 启动 stunnel4 服务
```bash
sudo systemctl start stunnel4
```
### 4.2 其他切换注意事项
1. **OMC 配置修改**:
- 进入 OMC 系统
- 导航至 `Base -> Market API Session`
- 修改对应的 session 账号密码
2. **Nacos 配置确认**:
- 打开 `quick-fix-client.yml` 配置
- 确认 `enabledSession` 连接的 session
- 确认配置 session 的 `SocketConnectHost` 是否为 web server 的 IP
## 5. 部署架构
- stunnel4 安装在 web server 上
- 开启端口 440 与 441
- AP server 通过 web server 的内网 IP 连接
## 6. 常见问题与排查
### 6.1 连接问题
- 确认 stunnel4 服务是否正常运行:`sudo systemctl status stunnel4`
- 检查端口是否正常监听:`sudo netstat -aonp | grep 440`
- 查看日志文件:`tail -f /var/log/stunnel4/stunnel.log`
### 6.2 证书问题
- 确认证书文件路径是否正确
- 检查证书文件权限是否正确
- 验证证书是否过期
## 7. 相关命令
- 启动服务:`sudo systemctl start stunnel4`
- 停止服务:`sudo systemctl stop stunnel4`
- 重启服务:`sudo systemctl restart stunnel4`
- 查看服务状态:`sudo systemctl status stunnel4`
- 查看端口占用:`sudo netstat -aonp | grep 440`
- 查看日志:`tail -f /var/log/stunnel4/stunnel.log`
## 8. 注意事项
- stunnel4 仅用于连接 OSL 交易所
- 切换环境时务必按照上述步骤操作,确保服务正常运行
- 定期检查 stunnel4 服务状态和日志,确保连接稳定
- 确保证书文件的安全性,避免泄露
@@ -0,0 +1,120 @@
# FIX Engine 高可用 (HA) 架构方案说明文档
## 1. 架构演进概述
### 1.1 原单实例架构
- **消息存储**: 使用本地文件系统(FileStore),消息持久化在本地磁盘。
- **灾备能力**: 无自动切换机制。若实例宕机,需手动启动备机,且由于文件存储不共享,消息序列号(Sequence Number)难以同步。
- **会话管理**: 启动即连接,缺乏灵活性。
### 1.2 现高可用 (HA) 架构
- **消息存储**: 迁移至 **MySQL 数据库**(JdbcStore)。多实例共享同一数据库,确保消息和序列号在节点切换后保持一致。
- **领导权选举**: 引入 **Redis 分布式锁**(Redisson)。以 Session 为粒度进行选举,确保每个 FIX 会话在全网只有一个活动实例。
- **自动灾备**: 备机定时检查锁状态,主节点宕机后锁释放,备机自动接管并恢复连接。
- **动态会话**: 结合 DynamicSession 机制,仅在获得领导权后才创建并启动 FIX 连接。
---
## 2. 核心组件分析
### 2.1 MySQL 消息持久化 (JdbcStore)
通过 `QuickFixJConfig` 配置,系统支持将消息存储从文件切换到 MySQL。
- **核心类**: `quickfix.JdbcStoreFactory`
- **配置触发**: 当 `StorageType=mysql` 时启用。
- **数据表要求**: 需要在数据库中创建 QuickFIX/J 标准表(如 `messages`, `sessions` 等)。
- **优势**:
- **共享状态**: 所有实例访问同一份数据,切换节点无需手动同步序列号。
- **可靠性**: 数据库事务保障消息持久化。
### 2.2 Redis 分布式锁选举 (LeaderElectionService)
使用 Redisson 实现 Session 级别的分布式锁管理。
- **锁 Key 格式**: `g3fo:fix-engine:session-lock:{SenderCompID}@{TargetCompID}`
- **看门狗机制**: 租约时间设为 -1,Redisson 自动续期。只要进程存活,锁就不会过期。
- **检查机制**: 定时线程(默认 10 秒)遍历所有配置的 Session。
- **宕机延迟 (Failover Delay)**: 检测到锁释放后,等待 5 秒再尝试抢占,防止因网络抖动引起的频繁切换。
### 2.3 动态会话控制 (Dynamic Session)
解决 QuickFIX/J 在 `start()` 时会自动连接所有会话的问题。
- **机制**: 在配置文件中将 Session 标记为 Dynamic(不预先加载)。
- **流程**:
1. `LeaderElectionService` 获得 Redis 锁。
2. 调用 `quickFixJConfig.startSession(sessionCode)`。
3. 内部通过 `socketInitiator.createDynamicSession(sessionId)` 动态创建会话对象。
4. 调用 `session.logon()` 触发连接。
---
## 3. 完整业务流程
### 3.1 服务启动流程
```mermaid
flowchart TD
Start([服务启动]) --> InitInitiator[初始化 QuickFIX Initiator]
InitInitiator --> StartElection[启动 LeaderElectionService]
StartElection --> LoadSessions[加载 enabledSession 配置]
LoadSessions --> LoopCheck{遍历每个 Session}
LoopCheck --> TryLock[尝试获取 Redis 分布式锁]
TryLock -- 成功获得锁 --> StartComp[启动会话组件]
StartComp --> CreateDynamic[创建 DynamicSession 对象]
CreateDynamic --> FixLogon[发送 FIX Logon]
FixLogon --> StartMQ[启动对应的 RocketMQ Listener]
TryLock -- 失败 --> Standby[进入待命状态]
Standby --> Wait[等待下一个检查周期]
Wait --> LoopCheck
```
### 3.2 宕机自动切换流程 (Failover)
1. **主节点 (Node A)** 持有 `A@OCG` 的 Redis 锁,正常运行。
2. **主节点 (Node A)** 宕机或网络断开。
3. **Redis 锁失效**: 经过看门狗租约超时,Redis 中的锁 Key 消失。
4. **备节点 (Node B)** 定时任务检测到锁已释放。
5. **进入延迟等待**: 备节点等待 `failoverDelay` (如 5秒)。
6. **抢占锁**: 备节点尝试 `tryLock`,成功获得领导权。
7. **恢复连接**: 备节点动态创建 Session 并在 MySQL 中读取最新的序列号,发送 Logon。
8. **流量接管**: 启动 MQ Listener,开始处理业务消息。
---
## 4. 关键配置指南 (Nacos)
在 `quick-fix-client.yml` 中进行如下配置:
```yaml
quickfixj:
client:
# 启用会话列表
enabledSession: A@OCG,A@OSL
# 选举参数
sessionElection:
checkInterval: 10 # 检查周期(秒)
failoverDelay: 5 # 接管延迟(秒)
# 存储配置(需配合数据源)
# 确保 SessionSettings 中的 StorageType=mysql
```
---
## 5. 日志监控与诊断
通过监控日志可以实时掌握 HA 状态:
| 日志内容 | 含义 | 级别 |
| :--- | :--- | :--- |
| `>>> 获得Session A@OCG 领导权 <<<` | 当前实例成功竞争到主节点 | INFO |
| `!!! Session A@OCG 领导权锁意外丢失 !!!` | 锁异常丢失,可能是 Redis 连接断开 | ERROR |
| `正在启动Session A@OCG 组件...` | 获得锁后开始加载 FIX 和 MQ | INFO |
| `✅ Session A@OCG 组件启动完成` | 成功恢复服务 | INFO |
| `未获得领导权,作为备用实例运行中...` | 当前为备份节点,正常待命 | DEBUG |
---
## 6. 总结
本方案通过 **MySQL 共享存储 + Redis 分布式选举 + 动态会话加载** 的组合,解决了 FIX 引擎单点故障问题。它不仅保证了消息的连续性(序列号一致),还实现了会话级别的细粒度高可用,能够灵活应对多种部署场景。
@@ -1,25 +1,31 @@
# Service Name: [Service ID] # Service Name: g3fo-exchange-fix-engine-service
## 1. Overview ## 1. Overview
[A brief description of what this service does and its primary responsibility.] g3fo-exchange-fix-engine-service 是 g3fo 系统的 FIX 协议接入引擎。它负责与外部交易所或交易对手建立 FIX 连接,进行消息的编码、解码、序列号维护以及消息的持久化存储。
## 2. Key Responsibilities ## 2. Key Responsibilities
- [Responsibility 1] - **会话管理**: 维护与外部实体的 FIX 会话,处理 Logon, Logout, Heartbeat, ResendRequest 等协议级消息。
- [Responsibility 2] - **消息路由**: 将接收到的 FIX 消息转换为内部业务格式并转发给下游服务,同时将下游服务的指令转换为 FIX 消息发送给外部。
- [Responsibility 3] - **消息持久化**: 记录所有发送和接收的 FIX 消息,确保在异常重启后能够恢复会话状态。
- **高可用 (HA)**: 支持基于 Redis 选举和 MySQL 共享存储的多实例部署。
## 3. Key Data Entities ## 3. Key Data Entities
[List major database tables or domain objects managed by this service.] - **FIX 消息存储**: 存储在 MySQL 中的 `messages` 表(基于 JdbcStore)。
- **Entity A**: [Description] - **FIX 会话状态**: 存储在 MySQL 中的 `sessions` 表。
- **Entity B**: [Description]
## 4. Dependencies ## 4. Dependencies
- **Upstream**: [Services that call this service] - **Upstream**: `g3fo-trade-service` (发送交易指令)
- **Downstream**: [Services called by this service] - **Downstream**: 外部交易所 FIX Gateway
- **Middleware**: [e.g., MySQL, Redis, Kafka] - **Middleware**:
- **MySQL**: 用于消息持久化。
- **Redis**: 用于分布式锁选举。
- **RocketMQ**: 用于与其他微服务通信。
## 5. Critical Configurations ## 5. Critical Configurations
[Key environment variables or config items.] - **QuickFIX/J 配置**: `quick-fix-client.yml` 中定义的会话参数、存储类型、选举参数等。
- **Nacos**: 动态配置中心,存储服务运行参数。
## 6. Common Operations / Troubleshooting ## 6. Common Operations / Troubleshooting
[Service-specific health check URLs, log locations, etc.] - **高可用架构说明**: 详细的 HA 方案请参考 [FIX Engine 高可用 (HA) 架构方案说明文档](./fix-engine/ha-architecture.md)。
- **日志监控**: 重点关注 `LeaderElectionService` 的选举日志和 `quickfix.JdbcStore` 的数据库操作日志。
- **会话重置**: 手动重置序列号通常需要清理数据库中的 `sessions` 表对应记录并重启服务。
+202
View File
@@ -0,0 +1,202 @@
Apache License
Version 2.0, January 2004
http://www.apache.org/licenses/
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
1. Definitions.
"License" shall mean the terms and conditions for use, reproduction,
and distribution as defined by Sections 1 through 9 of this document.
"Licensor" shall mean the copyright owner or entity authorized by
the copyright owner that is granting the License.
"Legal Entity" shall mean the union of the acting entity and all
other entities that control, are controlled by, or are under common
control with that entity. For the purposes of this definition,
"control" means (i) the power, direct or indirect, to cause the
direction or management of such entity, whether by contract or
otherwise, or (ii) ownership of fifty percent (50%) or more of the
outstanding shares, or (iii) beneficial ownership of such entity.
"You" (or "Your") shall mean an individual or Legal Entity
exercising permissions granted by this License.
"Source" form shall mean the preferred form for making modifications,
including but not limited to software source code, documentation
source, and configuration files.
"Object" form shall mean any form resulting from mechanical
transformation or translation of a Source form, including but
not limited to compiled object code, generated documentation,
and conversions to other media types.
"Work" shall mean the work of authorship, whether in Source or
Object form, made available under the License, as indicated by a
copyright notice that is included in or attached to the work
(an example is provided in the Appendix below).
"Derivative Works" shall mean any work, whether in Source or Object
form, that is based on (or derived from) the Work and for which the
editorial revisions, annotations, elaborations, or other modifications
represent, as a whole, an original work of authorship. For the purposes
of this License, Derivative Works shall not include works that remain
separable from, or merely link (or bind by name) to the interfaces of,
the Work and Derivative Works thereof.
"Contribution" shall mean any work of authorship, including
the original version of the Work and any modifications or additions
to that Work or Derivative Works thereof, that is intentionally
submitted to Licensor for inclusion in the Work by the copyright owner
or by an individual or Legal Entity authorized to submit on behalf of
the copyright owner. For the purposes of this definition, "submitted"
means any form of electronic, verbal, or written communication sent
to the Licensor or its representatives, including but not limited to
communication on electronic mailing lists, source code control systems,
and issue tracking systems that are managed by, or on behalf of, the
Licensor for the purpose of discussing and improving the Work, but
excluding communication that is conspicuously marked or otherwise
designated in writing by the copyright owner as "Not a Contribution."
"Contributor" shall mean Licensor and any individual or Legal Entity
on behalf of whom a Contribution has been received by Licensor and
subsequently incorporated within the Work.
2. Grant of Copyright License. Subject to the terms and conditions of
this License, each Contributor hereby grants to You a perpetual,
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
copyright license to reproduce, prepare Derivative Works of,
publicly display, publicly perform, sublicense, and distribute the
Work and such Derivative Works in Source or Object form.
3. Grant of Patent License. Subject to the terms and conditions of
this License, each Contributor hereby grants to You a perpetual,
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
(except as stated in this section) patent license to make, have made,
use, offer to sell, sell, import, and otherwise transfer the Work,
where such license applies only to those patent claims licensable
by such Contributor that are necessarily infringed by their
Contribution(s) alone or by combination of their Contribution(s)
with the Work to which such Contribution(s) was submitted. If You
institute patent litigation against any entity (including a
cross-claim or counterclaim in a lawsuit) alleging that the Work
or a Contribution incorporated within the Work constitutes direct
or contributory patent infringement, then any patent licenses
granted to You under this License for that Work shall terminate
as of the date such litigation is filed.
4. Redistribution. You may reproduce and distribute copies of the
Work or Derivative Works thereof in any medium, with or without
modifications, and in Source or Object form, provided that You
meet the following conditions:
(a) You must give any other recipients of the Work or
Derivative Works a copy of this License; and
(b) You must cause any modified files to carry prominent notices
stating that You changed the files; and
(c) You must retain, in the Source form of any Derivative Works
that You distribute, all copyright, patent, trademark, and
attribution notices from the Source form of the Work,
excluding those notices that do not pertain to any part of
the Derivative Works; and
(d) If the Work includes a "NOTICE" text file as part of its
distribution, then any Derivative Works that You distribute must
include a readable copy of the attribution notices contained
within such NOTICE file, excluding those notices that do not
pertain to any part of the Derivative Works, in at least one
of the following places: within a NOTICE text file distributed
as part of the Derivative Works; within the Source form or
documentation, if provided along with the Derivative Works; or,
within a display generated by the Derivative Works, if and
wherever such third-party notices normally appear. The contents
of the NOTICE file are for informational purposes only and
do not modify the License. You may add Your own attribution
notices within Derivative Works that You distribute, alongside
or as an addendum to the NOTICE text from the Work, provided
that such additional attribution notices cannot be construed
as modifying the License.
You may add Your own copyright statement to Your modifications and
may provide additional or different license terms and conditions
for use, reproduction, or distribution of Your modifications, or
for any such Derivative Works as a whole, provided Your use,
reproduction, and distribution of the Work otherwise complies with
the conditions stated in this License.
5. Submission of Contributions. Unless You explicitly state otherwise,
any Contribution intentionally submitted for inclusion in the Work
by You to the Licensor shall be under the terms and conditions of
this License, without any additional terms or conditions.
Notwithstanding the above, nothing herein shall supersede or modify
the terms of any separate license agreement you may have executed
with Licensor regarding such Contributions.
6. Trademarks. This License does not grant permission to use the trade
names, trademarks, service marks, or product names of the Licensor,
except as required for reasonable and customary use in describing the
origin of the Work and reproducing the content of the NOTICE file.
7. Disclaimer of Warranty. Unless required by applicable law or
agreed to in writing, Licensor provides the Work (and each
Contributor provides its Contributions) on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
implied, including, without limitation, any warranties or conditions
of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
PARTICULAR PURPOSE. You are solely responsible for determining the
appropriateness of using or redistributing the Work and assume any
risks associated with Your exercise of permissions under this License.
8. Limitation of Liability. In no event and under no legal theory,
whether in tort (including negligence), contract, or otherwise,
unless required by applicable law (such as deliberate and grossly
negligent acts) or agreed to in writing, shall any Contributor be
liable to You for damages, including any direct, indirect, special,
incidental, or consequential damages of any character arising as a
result of this License or out of the use or inability to use the
Work (including but not limited to damages for loss of goodwill,
work stoppage, computer failure or malfunction, or any and all
other commercial damages or losses), even if such Contributor
has been advised of the possibility of such damages.
9. Accepting Warranty or Additional Liability. While redistributing
the Work or Derivative Works thereof, You may choose to offer,
and charge a fee for, acceptance of support, warranty, indemnity,
or other liability obligations and/or rights consistent with this
License. However, in accepting such obligations, You may act only
on Your own behalf and on Your sole responsibility, not on behalf
of any other Contributor, and only if You agree to indemnify,
defend, and hold each Contributor harmless for any liability
incurred by, or claims asserted against, such Contributor by reason
of your accepting any such warranty or additional liability.
END OF TERMS AND CONDITIONS
APPENDIX: How to apply the Apache License to your work.
To apply the Apache License to your work, attach the following
boilerplate notice, with the fields enclosed by brackets "[]"
replaced with your own identifying information. (Don't include
the brackets!) The text should be enclosed in the appropriate
comment syntax for the file format. We also recommend that a
file or class name and description of purpose be included on the
same "printed page" as the copyright notice for easier
identification within third-party archives.
Copyright [yyyy] [name of copyright owner]
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
+32
View File
@@ -0,0 +1,32 @@
---
name: internal-comms
description: A set of resources to help me write all kinds of internal communications, using the formats that my company likes to use. Claude should use this skill whenever asked to write some sort of internal communications (status reports, leadership updates, 3P updates, company newsletters, FAQs, incident reports, project updates, etc.).
license: Complete terms in LICENSE.txt
---
## When to use this skill
To write internal communications, use this skill for:
- 3P updates (Progress, Plans, Problems)
- Company newsletters
- FAQ responses
- Status reports
- Leadership updates
- Project updates
- Incident reports
## How to use this skill
To write any internal communication:
1. **Identify the communication type** from the request
2. **Load the appropriate guideline file** from the `examples/` directory:
- `examples/3p-updates.md` - For Progress/Plans/Problems team updates
- `examples/company-newsletter.md` - For company-wide newsletters
- `examples/faq-answers.md` - For answering frequently asked questions
- `examples/general-comms.md` - For anything else that doesn't explicitly match one of the above
3. **Follow the specific instructions** in that file for formatting, tone, and content gathering
If the communication type doesn't match any existing guideline, ask for clarification or more context about the desired format.
## Keywords
3P updates, company newsletter, company comms, weekly update, faqs, common questions, updates, internal comms
@@ -0,0 +1,47 @@
## Instructions
You are being asked to write a 3P update. 3P updates stand for "Progress, Plans, Problems." The main audience is for executives, leadership, other teammates, etc. They're meant to be very succinct and to-the-point: think something you can read in 30-60sec or less. They're also for people with some, but not a lot of context on what the team does.
3Ps can cover a team of any size, ranging all the way up to the entire company. The bigger the team, the less granular the tasks should be. For example, "mobile team" might have "shipped feature" or "fixed bugs," whereas the company might have really meaty 3Ps, like "hired 20 new people" or "closed 10 new deals."
They represent the work of the team across a time period, almost always one week. They include three sections:
1) Progress: what the team has accomplished over the next time period. Focus mainly on things shipped, milestones achieved, tasks created, etc.
2) Plans: what the team plans to do over the next time period. Focus on what things are top-of-mind, really high priority, etc. for the team.
3) Problems: anything that is slowing the team down. This could be things like too few people, bugs or blockers that are preventing the team from moving forward, some deal that fell through, etc.
Before writing them, make sure that you know the team name. If it's not specified, you can ask explicitly what the team name you're writing for is.
## Tools Available
Whenever possible, try to pull from available sources to get the information you need:
- Slack: posts from team members with their updates - ideally look for posts in large channels with lots of reactions
- Google Drive: docs written from critical team members with lots of views
- Email: emails with lots of responses of lots of content that seems relevant
- Calendar: non-recurring meetings that have a lot of importance, like product reviews, etc.
Try to gather as much context as you can, focusing on the things that covered the time period you're writing for:
- Progress: anything between a week ago and today
- Plans: anything from today to the next week
- Problems: anything between a week ago and today
If you don't have access, you can ask the user for things they want to cover. They might also include these things to you directly, in which case you're mostly just formatting for this particular format.
## Workflow
1. **Clarify scope**: Confirm the team name and time period (usually past week for Progress/Problems, next
week for Plans)
2. **Gather information**: Use available tools or ask the user directly
3. **Draft the update**: Follow the strict formatting guidelines
4. **Review**: Ensure it's concise (30-60 seconds to read) and data-driven
## Formatting
The format is always the same, very strict formatting. Never use any formatting other than this. Pick an emoji that is fun and captures the vibe of the team and update.
[pick an emoji] [Team Name] (Dates Covered, usually a week)
Progress: [1-3 sentences of content]
Plans: [1-3 sentences of content]
Problems: [1-3 sentences of content]
Each section should be no more than 1-3 sentences: clear, to the point. It should be data-driven, and generally include metrics where possible. The tone should be very matter-of-fact, not super prose-heavy.
@@ -0,0 +1,65 @@
## Instructions
You are being asked to write a company-wide newsletter update. You are meant to summarize the past week/month of a company in the form of a newsletter that the entire company will read. It should be maybe ~20-25 bullet points long. It will be sent via Slack and email, so make it consumable for that.
Ideally it includes the following attributes:
- Lots of links: pulling documents from Google Drive that are very relevant, linking to prominent Slack messages in announce channels and from executives, perhgaps referencing emails that went company-wide, highlighting significant things that have happened in the company.
- Short and to-the-point: each bullet should probably be no longer than ~1-2 sentences
- Use the "we" tense, as you are part of the company. Many of the bullets should say "we did this" or "we did that"
## Tools to use
If you have access to the following tools, please try to use them. If not, you can also let the user know directly that their responses would be better if they gave them access.
- Slack: look for messages in channels with lots of people, with lots of reactions or lots of responses within the thread
- Email: look for things from executives that discuss company-wide announcements
- Calendar: if there were meetings with large attendee lists, particularly things like All-Hands meetings, big company announcements, etc. If there were documents attached to those meetings, those are great links to include.
- Documents: if there were new docs published in the last week or two that got a lot of attention, you can link them. These should be things like company-wide vision docs, plans for the upcoming quarter or half, things authored by critical executives, etc.
- External press: if you see references to articles or press we've received over the past week, that could be really cool too.
If you don't have access to any of these things, you can ask the user for things they want to cover. In this case, you'll mostly just be polishing up and fitting to this format more directly.
## Sections
The company is pretty big: 1000+ people. There are a variety of different teams and initiatives going on across the company. To make sure the update works well, try breaking it into sections of similar things. You might break into clusters like {product development, go to market, finance} or {recruiting, execution, vision}, or {external news, internal news} etc. Try to make sure the different areas of the company are highlighted well.
## Prioritization
Focus on:
- Company-wide impact (not team-specific details)
- Announcements from leadership
- Major milestones and achievements
- Information that affects most employees
- External recognition or press
Avoid:
- Overly granular team updates (save those for 3Ps)
- Information only relevant to small groups
- Duplicate information already communicated
## Example Formats
:megaphone: Company Announcements
- Announcement 1
- Announcement 2
- Announcement 3
:dart: Progress on Priorities
- Area 1
- Sub-area 1
- Sub-area 2
- Sub-area 3
- Area 2
- Sub-area 1
- Sub-area 2
- Sub-area 3
- Area 3
- Sub-area 1
- Sub-area 2
- Sub-area 3
:pillar: Leadership Updates
- Post 1
- Post 2
- Post 3
:thread: Social Updates
- Update 1
- Update 2
- Update 3
@@ -0,0 +1,30 @@
## Instructions
You are an assistant for answering questions that are being asked across the company. Every week, there are lots of questions that get asked across the company, and your goal is to try to summarize what those questions are. We want our company to be well-informed and on the same page, so your job is to produce a set of frequently asked questions that our employees are asking and attempt to answer them. Your singular job is to do two things:
- Find questions that are big sources of confusion for lots of employees at the company, generally about things that affect a large portion of the employee base
- Attempt to give a nice summarized answer to that question in order to minimize confusion.
Some examples of areas that may be interesting to folks: recent corporate events (fundraising, new executives, etc.), upcoming launches, hiring progress, changes to vision or focus, etc.
## Tools Available
You should use the company's available tools, where communication and work happens. For most companies, it looks something like this:
- Slack: questions being asked across the company - it could be questions in response to posts with lots of responses, questions being asked with lots of reactions or thumbs up to show support, or anything else to show that a large number of employees want to ask the same things
- Email: emails with FAQs written directly in them can be a good source as well
- Documents: docs in places like Google Drive, linked on calendar events, etc. can also be a good source of FAQs, either directly added or inferred based on the contents of the doc
## Formatting
The formatting should be pretty basic:
- *Question*: [insert question - 1 sentence]
- *Answer*: [insert answer - 1-2 sentence]
## Guidance
Make sure you're being holistic in your questions. Don't focus too much on just the user in question or the team they are a part of, but try to capture the entire company. Try to be as holistic as you can in reading all the tools available, producing responses that are relevant to all at the company.
## Answer Guidelines
- Base answers on official company communications when possible
- If information is uncertain, indicate that clearly
- Link to authoritative sources (docs, announcements, emails)
- Keep tone professional but approachable
- Flag if a question requires executive input or official response
@@ -0,0 +1,16 @@
## Instructions
You are being asked to write internal company communication that doesn't fit into the standard formats (3P
updates, newsletters, or FAQs).
Before proceeding:
1. Ask the user about their target audience
2. Understand the communication's purpose
3. Clarify the desired tone (formal, casual, urgent, informational)
4. Confirm any specific formatting requirements
Use these general principles:
- Be clear and concise
- Use active voice
- Put the most important information first
- Include relevant links and references
- Match the company's communication style
+10 -5
View File
@@ -1,21 +1,26 @@
--- ---
name: skill-sync name: skill-sync
description: Synchronize skills from the current project to the global Claude skills directory (~/.claude/skills). Use this when you have modified skills in the project and want to update the global installation. Supports Windows 11. description: Synchronize skills from the current project to multiple global skills directories (~/.claude/skills, ~/.trae/skills, ~/.trae-cn/skills). Use this when you have modified skills in the project and want to update the global installations. Supports Windows 11.
--- ---
# Skill Sync # Skill Sync
Synchronize skills from the project's `skills/` directory to the global Claude environment. Synchronize skills from the project's `skills/` directory to multiple global environments (Claude, Trae, and Trae-CN).
## Workflow ## Workflow
### 1. Synchronize All Skills ### 1. Synchronize All Skills
When changes are made to skills within this project, run the synchronization script to update the global installation. When changes are made to skills within this project, run the synchronization script to update all global installations.
- **Operation**: Run `powershell -ExecutionPolicy Bypass -File skills/skill-sync/scripts/sync.ps1` - **Operation**: Run `powershell -ExecutionPolicy Bypass -File skills/skill-sync/scripts/sync.ps1`
- **Effect**: All skill directories in `skills/` will be copied to `~\.claude\skills`. Existing skills in the target directory will be updated, while additional skills in the target that are not in the project remain untouched. - **Effect**: All skill directories in `skills/` will be copied to the following locations:
- `~\.claude\skills` (Claude)
- `~\.trae\skills` (Trae)
- `~\.trae-cn\skills` (Trae-CN)
Existing skills in the target directories will be updated, while additional skills in the targets that are not in the project remain untouched.
## Resources ## Resources
### scripts/ ### scripts/
- `sync.ps1`: PowerShell script for synchronizing skill directories to the global Claude skills folder. - `sync.ps1`: PowerShell script for synchronizing skill directories to multiple global skills folders (Claude, Trae, and Trae-CN).
+22 -9
View File
@@ -2,27 +2,40 @@
$ScriptDir = Split-Path -Parent $MyInvocation.MyCommand.Path $ScriptDir = Split-Path -Parent $MyInvocation.MyCommand.Path
$SkillDir = Split-Path -Parent $ScriptDir $SkillDir = Split-Path -Parent $ScriptDir
$ProjectSkillsDir = Split-Path -Parent $SkillDir $ProjectSkillsDir = Split-Path -Parent $SkillDir
$GlobalSkillsDir = Join-Path $HOME ".claude\skills"
Write-Host "Syncing skills from $ProjectSkillsDir to $GlobalSkillsDir..." -ForegroundColor Cyan # Define global skills directories for different platforms
$GlobalSkillsDirs = @(
@{ Name = "Claude"; Path = Join-Path $HOME ".claude\skills" }
@{ Name = "Trae"; Path = Join-Path $HOME ".trae\skills" }
@{ Name = "Trae-CN"; Path = Join-Path $HOME ".trae-cn\skills" }
)
if (-not (Test-Path $GlobalSkillsDir)) { Write-Host "Syncing skills from $ProjectSkillsDir..." -ForegroundColor Cyan
Write-Host "Global skills directory does not exist. Creating it..." -ForegroundColor Yellow
New-Item -ItemType Directory -Path $GlobalSkillsDir -Force
}
# Get all skill directories in the project # Get all skill directories in the project
$Skills = Get-ChildItem -Path $ProjectSkillsDir -Directory $Skills = Get-ChildItem -Path $ProjectSkillsDir -Directory
foreach ($DirConfig in $GlobalSkillsDirs) {
$PlatformName = $DirConfig.Name
$TargetBaseDir = $DirConfig.Path
Write-Host "`nSyncing to $PlatformName skills directory: $TargetBaseDir..." -ForegroundColor Magenta
if (-not (Test-Path $TargetBaseDir)) {
Write-Host "$PlatformName skills directory does not exist. Creating it..." -ForegroundColor Yellow
New-Item -ItemType Directory -Path $TargetBaseDir -Force
}
foreach ($Skill in $Skills) { foreach ($Skill in $Skills) {
$SkillName = $Skill.Name $SkillName = $Skill.Name
$TargetDir = Join-Path $GlobalSkillsDir $SkillName $TargetDir = Join-Path $TargetBaseDir $SkillName
Write-Host "Syncing skill: $SkillName..." -ForegroundColor Green Write-Host "Syncing skill: $SkillName..." -ForegroundColor Green
# Sync the skill directory # Sync the skill directory
# Note: Copy-Item -Recurse -Force will overwrite existing files # Note: Copy-Item -Recurse -Force will overwrite existing files
Copy-Item -Path $Skill.FullName -Destination $GlobalSkillsDir -Recurse -Force Copy-Item -Path $Skill.FullName -Destination $TargetBaseDir -Recurse -Force
}
} }
Write-Host "Sync completed successfully." -ForegroundColor Green Write-Host "`nSync completed successfully to all platforms!" -ForegroundColor Green