Initial commit with project setup and basic structure established.
This commit is contained in:
@@ -0,0 +1,67 @@
|
||||
---
|
||||
name: g3fo-docs
|
||||
description: g3fo 系统相关文档。管理业务流程说明、服务职责定义、中间件标准部署、系统初始化规范,以及跨环境(Dev/UAT)和多客户的资产清单。支持差异化配置查询与业务逻辑问答。
|
||||
---
|
||||
|
||||
# g3fo 系统相关文档 (g3fo-docs)
|
||||
|
||||
本 Skill 采用“标准模板 + 客户清单”的架构,并整合了**业务流程**与**服务职责**,旨在帮助 AI 快速理解并回答复杂的系统业务问题。
|
||||
|
||||
## 1. 查找指南
|
||||
- **查业务流程**:查阅 `references/business_flows/` (如订单生命周期、用户入金等跨服务流程)。
|
||||
- **查服务职责**:查阅 `references/services/` (如 `g3fo-trade-service.md` 定义的具体功能)。
|
||||
- **查架构概览**:查阅 `references/architecture/domain_overview.md`。
|
||||
- **查标准运维流程**:去 `references/middleware/` 或 `references/system-init/`。
|
||||
- **查具体环境/客户资产**:
|
||||
- 内部环境(Dev/UAT):查阅 `references/inventory/internal.md`。
|
||||
- 外部客户(客户A、B等):查阅 `references/inventory/clients/[客户名].md`。
|
||||
|
||||
## 2. 问答逻辑 (AI 引导)
|
||||
|
||||
### A. 业务逻辑问答
|
||||
1. **识别范围**:判断用户问题涉及哪些业务流程或服务。
|
||||
2. **加载文档**:
|
||||
- 优先查找 `business_flows/` 下的相关流程文档。
|
||||
- 结合 `services/` 下涉及的服务文档,深入了解具体职责。
|
||||
3. **关联分析**:
|
||||
- 使用流程文档中的 Mermaid 时序图理解服务间的交互。
|
||||
- 如果用户问及“某个功能由谁负责”,查阅 `architecture/domain_overview.md` 确认所属域。
|
||||
4. **回答准则**:
|
||||
- **基于事实**:仅根据已有的 Markdown 文档回答。
|
||||
- **明确边界**:如果文档中没有相关说明,必须回答:“根据现有文档,我无法确定 [具体业务点] 的实现细节,建议查阅代码或询问开发人员。”
|
||||
|
||||
### B. 运维配置生成
|
||||
1. **识别主体**:确定用户问的是哪个环境或哪个客户。
|
||||
2. **组合信息**:
|
||||
- **读取资产变量**:优先读取客户专属文件(如 `clients/client_a.md`)或内部清单(`internal.md`)中的“变量定义”表格。
|
||||
- **读取部署模版**:读取 `middleware/` 下的标准部署文档(如 `keepalived/deploy.md`)。
|
||||
- **变量替换**:将部署文档中的 `${VARIABLE_NAME}` 占位符替换为资产清单中对应的值。
|
||||
3. **输出要求**:
|
||||
- **环境标注**:在提供客户配置时,必须明确标注该配置适用于哪个特定环境,防止误操作。
|
||||
- **完整配置**:如果用户要求生成特定节点的配置文件,请直接输出替换后的完整配置内容。
|
||||
- **MySQL 命令规范**:
|
||||
- **强制包含 Host**:所有 `docker exec` 执行的 `mysql` 命令必须显式包含 `-h127.0.0.1` 参数,以确保在容器内连接成功。
|
||||
- **双格式输出**:当提供 SQL 修复或查询脚本时,必须同时提供以下两种格式:
|
||||
- **格式 A (Docker 模式)**:完整的 `docker exec -it mysql mysql -uroot -pafe123456 -h127.0.0.1 -e "SQL_CONTENT"`。
|
||||
- **格式 B (纯 SQL 模式)**:仅包含 SQL 语句本身,方便在 GUI 工具中执行。
|
||||
|
||||
## 3. 示例 Prompt (用户可参考)
|
||||
- **业务咨询**:
|
||||
- “g3fo 系统中,一个订单从下单到成交会经过哪些服务?请结合 `references/business_flows/` 下的相关文档回答。”
|
||||
- “`g3fo-margin-service` 负责哪些核心逻辑?它的上下游依赖是谁?”
|
||||
- **部署配置**:
|
||||
- “请结合 `references/inventory/internal.md` 中的 `UAT` 环境变量,参考 `references/middleware/keepalived/deploy.md` 部署文档,为我生成 Node A 的 `keepalived.conf` 配置文件。”
|
||||
- **故障排查**:
|
||||
- “我的 MySQL 出现了复制冲突,报错 `Duplicate entry`,请根据 `references/middleware/mysql/fault-analysis.md` 提供排查脚本和修复建议。”
|
||||
|
||||
## 4. 资源地图
|
||||
- **业务与服务**:
|
||||
- **服务职责**: `services/` (g3fo-trade-service, g3fo-margin-service 等)
|
||||
- **跨服务流程**: `business_flows/` (订单生命周期等)
|
||||
- **架构概览**: `architecture/domain_overview.md`
|
||||
- **中间件标准**:
|
||||
- **Keepalived**: `middleware/keepalived/`
|
||||
- **MySQL**: `middleware/mysql/`
|
||||
- **Redis**: `middleware/redis/`
|
||||
- **系统规范**: 内核优化, 安全加固...
|
||||
- **资产清单**: 内部测试机, 客户生产环境...
|
||||
@@ -0,0 +1,31 @@
|
||||
# G3FO System Architecture Overview
|
||||
|
||||
## 1. Domain Grouping
|
||||
The system is divided into the following functional domains:
|
||||
|
||||
### Trading Domain
|
||||
- `g3fo-trade-service`: Core matching and order management.
|
||||
- `g3fo-exchange-fix-engine-service`: Exchange connectivity.
|
||||
- `g3fo-gateway-service`: API entry point for clients.
|
||||
- `g3fo-dx-service`: [Description]
|
||||
|
||||
### Risk & Capital Domain
|
||||
- `g3fo-margin-service`: Margin calculations and liquidation.
|
||||
- `g3fo-base-service`: Basic master data.
|
||||
|
||||
### User & Admin Domain
|
||||
- `g3fo-user-service`: User authentication and authorization.
|
||||
- `g3fo-admin-service`: Back-office management.
|
||||
- `g3fo-product-service`: Trading instrument configuration.
|
||||
|
||||
### Utility & Infrastructure
|
||||
- `g3fo-utility-service`: Shared utilities.
|
||||
- `g3fo-push-service`: Real-time data broadcasting.
|
||||
- `g3fo-notification-service`: Email/SMS alerts.
|
||||
- `g3fo-monitor-service`: System health and metrics.
|
||||
|
||||
## 2. Global Tech Stack
|
||||
- **Language**: Java / Spring Boot
|
||||
- **Database**: MySQL
|
||||
- **Cache**: Redis
|
||||
- **Communication**: [e.g., REST, gRPC, Kafka]
|
||||
@@ -0,0 +1,30 @@
|
||||
# Business Flow: [Flow Name]
|
||||
|
||||
## 1. Overview
|
||||
[Description of the business objective this flow achieves.]
|
||||
|
||||
## 2. Involved Services
|
||||
- **Service A**: [Role in this flow]
|
||||
- **Service B**: [Role in this flow]
|
||||
|
||||
## 3. Sequence Diagram (Mermaid)
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant A as Service A
|
||||
participant B as Service B
|
||||
|
||||
A->>B: Request
|
||||
B-->>A: Response
|
||||
```
|
||||
|
||||
## 4. Detailed Step-by-Step
|
||||
1. **[Step 1]**: [Detailed description]
|
||||
2. **[Step 2]**: [Detailed description]
|
||||
|
||||
## 5. Error Handling & Edge Cases
|
||||
- **Failure at Step X**: [What happens?]
|
||||
- **Timeout**: [How is it handled?]
|
||||
|
||||
## 6. Related Documentation
|
||||
- [Link to Service A doc]
|
||||
- [Link to Service B doc]
|
||||
@@ -0,0 +1,43 @@
|
||||
# 示例客户 A 专属配置 (Client Overlay)
|
||||
|
||||
本文件包含客户 A 的 300+ 行专有配置及资产信息。
|
||||
|
||||
## 1. 资产信息与变量 (Keepalived)
|
||||
| 变量名 | 值 | 说明 |
|
||||
| --- | --- | --- |
|
||||
| `VIP_IP` | 172.20.10.100 | 虚拟 IP |
|
||||
| `NODE_A_IP` | 172.20.10.10 | 主节点 IP |
|
||||
| `NODE_B_IP` | 172.20.10.11 | 备节点 IP |
|
||||
| `NODE_C_IP` | 172.20.10.12 | 仲裁节点 IP |
|
||||
| `INTERFACE` | ens192 | 网卡名称 |
|
||||
| `ROUTER_ID_A` | LVS_CLIENTA_A | 节点 A 标识 |
|
||||
| `ROUTER_ID_B` | LVS_CLIENTA_B | 节点 B 标识 |
|
||||
| `ROUTER_ID_C` | LVS_CLIENTA_C | 节点 C 标识 |
|
||||
| `VIRTUAL_ROUTER_ID` | 101 | 虚拟路由 ID |
|
||||
| `AUTH_PASS` | aB3#dE5! | 认证密码 |
|
||||
|
||||
## 2. 专属初始化配置 (300行示例)
|
||||
该客户要求执行严格的安全加固脚本和内核调优:
|
||||
|
||||
```bash
|
||||
# --- 客户 A 专有安全脚本 START ---
|
||||
# 禁止 root 远程登录
|
||||
sed -i 's/PermitRootLogin yes/PermitRootLogin no/g' /etc/ssh/sshd_config
|
||||
# 专属防火墙规则 (100+行)
|
||||
iptables -A INPUT -p tcp --dport 22 -s 172.20.1.0/24 -j ACCEPT
|
||||
# ... 省略 150 行 ...
|
||||
# --- 客户 A 专有安全脚本 END ---
|
||||
```
|
||||
|
||||
## 3. 专属 Keepalived 配置
|
||||
由于环境限制,心跳间隔调整为 2s,且不开启抢占:
|
||||
|
||||
```conf
|
||||
vrrp_instance VI_1 {
|
||||
state BACKUP
|
||||
interface ens192
|
||||
nopreempt # 客户要求不抢占
|
||||
advert_int 2 # 间隔改为 2s
|
||||
# ... 其他 50 行专属配置 ...
|
||||
}
|
||||
```
|
||||
@@ -0,0 +1,32 @@
|
||||
# 内部环境资产清单 (Internal)
|
||||
|
||||
## 1. 开发环境 (Dev)
|
||||
| 变量名 | 值 | 说明 |
|
||||
| --- | --- | --- |
|
||||
| `VIP_IP` | 192.168.3.233 | 虚拟 IP |
|
||||
| `NODE_A_IP` | 192.168.3.200 | 主节点 IP |
|
||||
| `NODE_B_IP` | 192.168.3.230 | 备节点 IP |
|
||||
| `NODE_C_IP` | 192.168.3.110 | 仲裁节点 IP |
|
||||
| `INTERFACE` | eth0 | 网卡名称 |
|
||||
| `ROUTER_ID_A` | LVS_DEVEL_A | 节点 A 标识 |
|
||||
| `ROUTER_ID_B` | LVS_DEVEL_B | 节点 B 标识 |
|
||||
| `ROUTER_ID_C` | LVS_DEVEL_C | 节点 C 标识 |
|
||||
| `VIRTUAL_ROUTER_ID` | 51 | 虚拟路由 ID |
|
||||
| `AUTH_PASS` | 1111 | 认证密码 |
|
||||
| `REDIS_MASTER_IP` | 192.168.3.230 | Redis 主节点 IP |
|
||||
| `REDIS_SLAVE_IP` | 192.168.3.200 | Redis 从节点 IP |
|
||||
| `REDIS_SENTINEL_3_IP` | 192.168.3.110 | Redis 哨兵节点 3 IP |
|
||||
|
||||
## 2. 测试环境 (UAT)
|
||||
| 变量名 | 值 | 说明 |
|
||||
| --- | --- | --- |
|
||||
| `VIP_IP` | 10.50.1.100 | 虚拟 IP |
|
||||
| `NODE_A_IP` | 10.50.1.10 | 主节点 IP |
|
||||
| `NODE_B_IP` | 10.50.1.11 | 备节点 IP |
|
||||
| `NODE_C_IP` | 10.50.1.12 | 仲裁节点 IP |
|
||||
| `INTERFACE` | eth1 | 网卡名称 |
|
||||
| `ROUTER_ID_A` | LVS_UAT_A | 节点 A 标识 |
|
||||
| `ROUTER_ID_B` | LVS_UAT_B | 节点 B 标识 |
|
||||
| `ROUTER_ID_C` | LVS_UAT_C | 节点 C 标识 |
|
||||
| `VIRTUAL_ROUTER_ID` | 60 | 虚拟路由 ID |
|
||||
| `AUTH_PASS` | 2222 | 认证密码 |
|
||||
@@ -0,0 +1,325 @@
|
||||
# Docker 部署 Keepalived 2.3.4 实现三节点高可用 VIP
|
||||
本文档详细说明如何在三台机器上通过 Docker 部署 Keepalived 2.3.4,实现 VIP(${VIP_IP})的高可用,并通过仲裁节点防止脑裂问题。
|
||||
|
||||
## 1. 节点信息
|
||||
| 节点名称 | IP 地址 | 角色 | 基础优先级 | 备注 |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| Node A | ${NODE_A_IP} | 主节点 (Master) | 110 | MySQL 故障优先级扣 50 → 60 |
|
||||
| Node B | ${NODE_B_IP} | 备节点 (Backup) | 100 | MySQL 故障优先级扣 50 → 50 |
|
||||
| Node C | ${NODE_C_IP} | 仲裁节点 (Arbiter) | 40 | 参与选举但不持有 VIP |
|
||||
|
||||
|
||||
+ **VIP 地址**:${VIP_IP}
|
||||
+ **仲裁节点作用**:Node C 仅参与选举流程,防止 Node A 和 Node B 之间发生脑裂,正常场景下不持有 VIP。
|
||||
|
||||
## 2. 构建 Keepalived 镜像(解决 libip4tc 报错)
|
||||
采用多阶段构建方式,确保镜像最小化且包含所有运行时依赖。
|
||||
|
||||
### 2.1 Dockerfile 内容
|
||||
```dockerfile
|
||||
# Stage 1: 构建 Keepalived 二进制文件
|
||||
FROM alpine:3.23.2 AS builder
|
||||
|
||||
# 安装构建依赖
|
||||
RUN apk add --no-cache \
|
||||
gcc \
|
||||
make \
|
||||
libc-dev \
|
||||
openssl-dev \
|
||||
libnl3-dev \
|
||||
ipset-dev \
|
||||
iptables-dev \
|
||||
libnfnetlink-dev \
|
||||
libnftnl-dev \
|
||||
libmnl-dev \
|
||||
net-snmp-dev \
|
||||
libssh2-dev \
|
||||
pcre2-dev \
|
||||
autoconf \
|
||||
automake \
|
||||
linux-headers \
|
||||
curl \
|
||||
tar
|
||||
|
||||
# 下载并编译 Keepalived 2.3.4
|
||||
ENV KEEPALIVED_VERSION=2.3.4
|
||||
RUN curl -sSL https://www.keepalived.org/software/keepalived-${KEEPALIVED_VERSION}.tar.gz -o keepalived.tar.gz && \
|
||||
tar -zxf keepalived.tar.gz && \
|
||||
cd keepalived-${KEEPALIVED_VERSION} && \
|
||||
./configure --prefix=/usr --sysconfdir=/etc && \
|
||||
make && \
|
||||
make install DESTDIR=/install
|
||||
|
||||
# Stage 2: 构建最终运行镜像
|
||||
FROM alpine:3.23.2
|
||||
|
||||
# 安装运行时依赖
|
||||
RUN apk add --no-cache \
|
||||
libnl3 \
|
||||
ipset \
|
||||
iptables \
|
||||
libnftnl \
|
||||
libmnl \
|
||||
libnfnetlink \
|
||||
net-snmp-libs \
|
||||
libssh2 \
|
||||
pcre2 \
|
||||
openssl \
|
||||
ca-certificates \
|
||||
tzdata \
|
||||
libip4tc \
|
||||
libip6tc
|
||||
|
||||
# 从构建阶段复制编译好的二进制文件
|
||||
COPY --from=builder /install /
|
||||
|
||||
# 创建默认配置目录
|
||||
RUN mkdir -p /etc/keepalived
|
||||
|
||||
# 暴露 VRRP 协议端口 (112)
|
||||
EXPOSE 112
|
||||
|
||||
# 前台运行 Keepalived(适配 Docker 容器运行模式)
|
||||
# --dont-fork: 不后台运行
|
||||
# --log-console: 日志输出到控制台(便于 Docker 日志查看)
|
||||
ENTRYPOINT ["keepalived", "--dont-fork", "--log-console"]
|
||||
```
|
||||
|
||||
### 2.2 构建镜像命令
|
||||
```bash
|
||||
docker build -t keepalived:2.3.4 .
|
||||
```
|
||||
|
||||
## 3. 配置文件部署
|
||||
### 3.1 前置准备
|
||||
在三台机器上统一创建配置目录:
|
||||
|
||||
```bash
|
||||
mkdir -p /data/keepalived
|
||||
```
|
||||
|
||||
### 3.2 Docker Compose 配置(三台机器通用)
|
||||
创建 `/data/keepalived/docker-compose.yml`:
|
||||
|
||||
```yaml
|
||||
services:
|
||||
keepalived:
|
||||
image: keepalived:2.3.4
|
||||
container_name: keepalived
|
||||
restart: always
|
||||
network_mode: host # 主机网络模式,确保 VRRP 协议正常通信
|
||||
cap_add: # 添加网络管理权限
|
||||
- NET_ADMIN
|
||||
- NET_BROADCAST
|
||||
- NET_RAW
|
||||
volumes:
|
||||
- /data/keepalived/keepalived.conf:/etc/keepalived/keepalived.conf:ro # 只读挂载配置文件
|
||||
- /data/keepalived/check-mysql.sh:/etc/keepalived/check-mysql.sh:ro # 挂载 MySQL 检测脚本(仅 Node A/B 需要)
|
||||
```
|
||||
|
||||
### 3.3 各节点 Keepalived 配置
|
||||
#### 3.3.1 Node A(${NODE_A_IP})配置
|
||||
路径:`/data/keepalived/keepalived.conf`
|
||||
|
||||
```nginx
|
||||
global_defs {
|
||||
router_id ${ROUTER_ID_A}
|
||||
vrrp_skip_check_adv_addr
|
||||
script_user root # 指定脚本执行用户
|
||||
enable_script_security # 启用脚本安全校验,防止执行未授权脚本
|
||||
}
|
||||
|
||||
# MySQL 存活检测脚本定义
|
||||
vrrp_script check_mysql {
|
||||
script "/etc/keepalived/check-mysql.sh"
|
||||
interval 3 # 检测间隔:3 秒
|
||||
weight -50 # 检测失败时,优先级扣 50
|
||||
fall 3 # 连续 3 次失败判定为故障
|
||||
rise 2 # 连续 2 次成功判定为恢复
|
||||
}
|
||||
|
||||
vrrp_instance VI_1 {
|
||||
state BACKUP # 所有节点统一设为 BACKUP,靠优先级决定 Master
|
||||
interface ${INTERFACE} # 替换为宿主机实际网卡名(如 ens33/br0 等)
|
||||
virtual_router_id ${VIRTUAL_ROUTER_ID} # 虚拟路由 ID,集群内必须一致(1-255)
|
||||
priority 110 # 基础优先级
|
||||
preempt_delay 30 # 优先级恢复后,延迟 30 秒抢占 VIP(减少抖动)
|
||||
advert_int 1 # VRRP 通告间隔:1 秒
|
||||
|
||||
# 认证配置(集群内必须一致)
|
||||
authentication {
|
||||
auth_type PASS
|
||||
auth_pass ${AUTH_PASS}
|
||||
}
|
||||
|
||||
# 单播配置(替代组播,适配部分网络环境)
|
||||
unicast_src_ip ${NODE_A_IP} # 本机 IP
|
||||
unicast_peer { # 集群其他节点 IP
|
||||
${NODE_B_IP}
|
||||
${NODE_C_IP}
|
||||
}
|
||||
|
||||
# 虚拟 IP 配置
|
||||
virtual_ipaddress {
|
||||
${VIP_IP}/32 # VIP 地址(32 位掩码)
|
||||
}
|
||||
|
||||
# 绑定检测脚本
|
||||
track_script {
|
||||
check_mysql
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### 3.3.2 Node B(${NODE_B_IP})配置
|
||||
路径:`/data/keepalived/keepalived.conf`
|
||||
|
||||
```nginx
|
||||
global_defs {
|
||||
router_id ${ROUTER_ID_B}
|
||||
vrrp_skip_check_adv_addr
|
||||
script_user root
|
||||
enable_script_security
|
||||
}
|
||||
|
||||
# MySQL 存活检测脚本定义
|
||||
vrrp_script check_mysql {
|
||||
script "/etc/keepalived/check-mysql.sh"
|
||||
interval 3
|
||||
weight -50
|
||||
fall 3
|
||||
rise 2
|
||||
}
|
||||
|
||||
vrrp_instance VI_1 {
|
||||
state BACKUP
|
||||
interface ${INTERFACE} # 替换为宿主机实际网卡名
|
||||
virtual_router_id ${VIRTUAL_ROUTER_ID}
|
||||
priority 100 # 基础优先级(低于 Node A)
|
||||
advert_int 1
|
||||
|
||||
authentication {
|
||||
auth_type PASS
|
||||
auth_pass ${AUTH_PASS}
|
||||
}
|
||||
|
||||
unicast_src_ip ${NODE_B_IP}
|
||||
unicast_peer {
|
||||
${NODE_A_IP}
|
||||
${NODE_C_IP}
|
||||
}
|
||||
|
||||
virtual_ipaddress {
|
||||
${VIP_IP}/32
|
||||
}
|
||||
|
||||
track_script {
|
||||
check_mysql
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### 3.3.3 Node C(${NODE_C_IP},仲裁节点)配置
|
||||
路径:`/data/keepalived/keepalived.conf`
|
||||
|
||||
```nginx
|
||||
global_defs {
|
||||
router_id ${ROUTER_ID_C}
|
||||
vrrp_skip_check_adv_addr
|
||||
}
|
||||
|
||||
vrrp_instance VI_1 {
|
||||
state BACKUP
|
||||
interface ${INTERFACE} # 替换为宿主机实际网卡名
|
||||
virtual_router_id ${VIRTUAL_ROUTER_ID}
|
||||
priority 40 # 最低优先级(仅参与选举,不持有 VIP)
|
||||
advert_int 1
|
||||
|
||||
authentication {
|
||||
auth_type PASS
|
||||
auth_pass ${AUTH_PASS}
|
||||
}
|
||||
|
||||
unicast_src_ip ${NODE_C_IP}
|
||||
unicast_peer {
|
||||
${NODE_A_IP}
|
||||
${NODE_B_IP}
|
||||
}
|
||||
|
||||
# 仲裁节点不配置 virtual_ipaddress,不持有 VIP
|
||||
}
|
||||
```
|
||||
|
||||
### 3.4 MySQL 检测脚本(仅 Node A/B 需要)
|
||||
路径:`/data/keepalived/check-mysql.sh`
|
||||
|
||||
```bash
|
||||
#!/bin/sh
|
||||
# 检测逻辑:
|
||||
# 1. timeout 2:限制命令 2 秒内完成,避免脚本阻塞
|
||||
# 2. nc -z:仅检测端口连通性(无数据传输)
|
||||
# 3. 屏蔽所有输出,仅返回退出码
|
||||
|
||||
if timeout 2 nc -z 127.0.0.1 3306 > /dev/null 2>&1; then
|
||||
# 端口连通,返回 0(Keepalived 判定正常)
|
||||
exit 0
|
||||
else
|
||||
# 端口不通/超时,返回 1(Keepalived 触发优先级扣减)
|
||||
exit 1
|
||||
fi
|
||||
```
|
||||
|
||||
添加执行权限:
|
||||
|
||||
```bash
|
||||
chmod +x /data/keepalived/check-mysql.sh
|
||||
```
|
||||
|
||||
## 4. 启动与验证
|
||||
### 4.1 启动容器
|
||||
三台机器执行相同命令:
|
||||
|
||||
```bash
|
||||
cd /data/keepalived
|
||||
docker-compose up -d
|
||||
```
|
||||
|
||||
### 4.2 日志查看
|
||||
```bash
|
||||
docker logs -f keepalived
|
||||
```
|
||||
|
||||
> 正常日志会显示 VRRP 选举过程、VIP 绑定/释放、脚本检测结果等信息。
|
||||
>
|
||||
|
||||
### 4.3 VIP 验证
|
||||
在任意节点执行,查看 VIP 是否绑定:
|
||||
|
||||
```bash
|
||||
ip addr show ${INTERFACE} # 替换为实际网卡名
|
||||
```
|
||||
|
||||
+ 正常情况下,VIP(${VIP_IP})会绑定在优先级最高的节点(初始为 Node A)。
|
||||
+ 输出示例:`inet ${VIP_IP}/32 scope global ${INTERFACE}`
|
||||
|
||||
## 5. 故障场景分析
|
||||
详细的故障场景模拟与 VIP 切换逻辑分析,请参阅:[Keepalived 故障场景分析](./fault-analysis.md)
|
||||
|
||||
## 6. 核心总结
|
||||
### 6.1 优先级规则
|
||||
+ VIP 归属由 Keepalived **有效优先级** 决定:基础优先级 A(110) > B(100) > C(40);
|
||||
+ MySQL 故障会让 A/B 优先级扣 50,Keepalived 离线则失去选举资格。
|
||||
|
||||
### 6.2 仲裁节点作用
|
||||
+ Node C 仅参与选举流程,不持有 VIP;
|
||||
+ 核心作用是防止 A、B 脑裂,正常/单节点故障场景下不影响 A/B 选举逻辑。
|
||||
|
||||
### 6.3 风险与优化
|
||||
+ **风险**:A/B MySQL 均故障但 Keepalived 正常时,VIP 仍绑定到 A,此时无可用 MySQL;需在业务层增加数据库存活检测,避免连接失效的 VIP。
|
||||
+ **抢占优化**:A 恢复后会在 `preempt_delay`(30 秒)后抢占 VIP,若需避免业务抖动,可在 Node A 配置中添加 `nopreempt` 关闭抢占。
|
||||
|
||||
### 6.4 关键配置注意事项
|
||||
+ 所有节点 `virtual_router_id` 和认证信息必须一致;
|
||||
+ 网卡名需替换为宿主机实际名称;
|
||||
+ 仲裁节点 C 不配置 `virtual_ipaddress`,不挂载 MySQL 检测脚本。
|
||||
|
||||
@@ -0,0 +1,25 @@
|
||||
# Keepalived 故障场景分析
|
||||
|
||||
本文档详细分析了在三节点高可用架构下,不同故障场景对 VIP 归属和集群状态的影响。
|
||||
|
||||
## 故障场景表
|
||||
|
||||
| 场景编号 | 故障场景描述 | 节点 A 状态(MySQL/Keepalived) | 节点 B 状态(MySQL/Keepalived) | 节点 C 状态(Keepalived) | 各节点有效优先级 | VIP 最终归属 | 关键说明 |
|
||||
| --- | --- | --- | --- | --- | --- | --- | --- |
|
||||
| 1 | 初始正常状态 | 正常/正常 | 正常/正常 | 正常 | A=110、B=100、C=40 | Node A | A 优先级最高,成为 Master |
|
||||
| 2 | A 的 MySQL 停机,Keepalived 正常 | 故障/正常 | 正常/正常 | 正常 | A=60、B=100、C=40 | Node B | A 扣权重后,B 优先级更高接管 VIP |
|
||||
| 3 | 场景 2 后,A 的 MySQL 恢复 | 恢复/正常 | 正常/正常 | 正常 | A=110、B=100、C=40 | Node A | A 优先级恢复,30 秒后抢占 VIP |
|
||||
| 4 | A 的 Keepalived 停机 | 正常/故障(离线) | 正常/正常 | 正常 | A=离线、B=100、C=40 | Node B | A 无选举资格,B 接管 |
|
||||
| 5 | B 的 MySQL 停机 | 正常/正常 | 故障/正常 | 正常 | A=110、B=50、C=40 | Node A | B 扣权重后,A 仍为最高 |
|
||||
| 6 | B 的 Keepalived 停机 | 正常/正常 | 正常/故障(离线) | 正常 | A=110、B=离线、C=40 | Node A | B 无选举资格,A 保持 Master |
|
||||
| 7 | A、B MySQL 均停机 | 故障/正常 | 故障/正常 | 正常 | A=60、B=50、C=40 | Node A | 无可用 MySQL,但 Keepalived 仍按优先级选举 |
|
||||
| 8 | A MySQL 停机 + B Keepalived 停机 | 故障/正常 | 无意义/故障(离线) | 正常 | A=60、B=离线、C=40 | Node A | B 离线,A 优先级高于 C(但 MySQL 不可用) |
|
||||
| 9 | A Keepalived 停机 + B MySQL 停机 | 正常/故障(离线) | 故障/正常 | 正常 | A=离线、B=50、C=40 | Node B | A 离线,B 优先级高于 C(但 MySQL 不可用) |
|
||||
| 10 | A、B Keepalived 均停机 | 正常/故障(离线) | 正常/故障(离线) | 正常 | A=离线、B=离线、C=40 | 无节点绑定 VIP | C 无 VIP 配置,仅参与选举不持有 VIP |
|
||||
| 11 | A MySQL+Keepalived 均停机 | 故障/故障(离线) | 正常/正常 | 正常 | A=离线、B=100、C=40 | Node B | A 完全离线,B 正常接管 |
|
||||
| 12 | B MySQL+Keepalived 均停机 | 正常/正常 | 故障/故障(离线) | 正常 | A=110、B=离线、C=40 | Node A | B 完全离线,A 保持 Master |
|
||||
| 13 | 场景 7 后,A MySQL 恢复 | 恢复/正常 | 故障/正常 | 正常 | A=110、B=50、C=40 | Node A | A 优先级恢复最高,接管 VIP |
|
||||
| 14 | 场景 7 后,B MySQL 恢复 | 故障/正常 | 恢复/正常 | 正常 | A=60、B=100、C=40 | Node B | B 优先级高于 A,接管 VIP |
|
||||
| 15 | C Keepalived 停机 | 正常/正常 | 正常/正常 | 故障(离线) | A=110、B=100、C=离线 | Node A | C 仅为仲裁,离线不影响 A/B 选举 |
|
||||
| 16 | A、B MySQL 均停机 + C 停机 | 故障/正常 | 故障/正常 | 故障(离线) | A=60、B=50、C=离线 | Node A | C 离线不影响 A/B 选举 |
|
||||
| 17 | 所有节点 Keepalived 均停机 | 正常/故障(离线) | 正常/故障(离线) | 故障(离线) | 全离线 | 无节点绑定 VIP | 无 Keepalived 参与选举,VIP 失联 |
|
||||
@@ -0,0 +1,337 @@
|
||||
# MySQL 8.4 双主(Source-Source)同步部署文档(跨主机统一目录/容器名版)
|
||||
|
||||
> **AI 响应规范 (MySQL 执行策略)**:
|
||||
> 1. **强制参数**: 所有 `docker exec` 命令必须包含 `-h127.0.0.1` 参数。
|
||||
> 2. **双重输出**: 涉及 SQL 操作时,必须同时输出 `Docker 执行命令` 和 `纯 SQL 脚本`。
|
||||
|
||||
## 文档概述
|
||||
你需要部署跨两台独立主机的 MySQL 8.4 双主双向复制集群,核心要求是:两台主机的容器名、目录路径统一使用 `mysql`(仅通过 IP 区分节点),保证两节点数据一致性,支持初始化、校验、宕机重启恢复及节点重建全流程。核心设计遵循:
|
||||
|
||||
+ **GTID + AUTO_POSITION**:重启后自动定位同步位点,减少人工干预;
|
||||
+ **ROW 模式 binlog**:避免非确定性函数导致的数据不一致;
|
||||
+ **持久化数据卷**:独立 Volume 保障容器重启数据不丢失;
|
||||
+ **自增键隔离**:通过步长/偏移配置避免双写主键冲突。
|
||||
|
||||
## 1. 环境准备
|
||||
### 1.1 节点信息(核心区分点)
|
||||
| 节点 | 主机IP | server_id | auto_increment_offset | 核心标识 |
|
||||
| :--- | :--- | :--- | :--- | :--- |
|
||||
| 主节点A | ${NODE_A_IP} | 1 | 1 | 生成主键1、3、5… |
|
||||
| 主节点B | ${NODE_B_IP} | 2 | 2 | 生成主键2、4、6… |
|
||||
|
||||
|
||||
### 1.2 系统依赖(两台主机统一执行)
|
||||
```plain
|
||||
# 安装 Docker & Docker Compose、MySQL 客户端
|
||||
sudo apt update && sudo apt install -y docker.io docker-compose-plugin mysql-client
|
||||
# 验证依赖(确保版本正常)
|
||||
docker --version && docker compose version && mysql --version
|
||||
```
|
||||
|
||||
### 1.3 目录规划(两台主机完全一致)
|
||||
```plain
|
||||
# 两台主机均执行以下命令,目录统一为 /data/mysql
|
||||
sudo mkdir -p /data/mysql/{data,backup,logs}
|
||||
sudo chown 999:999 /data/mysql -R # MySQL 容器默认UID/GID为999,避免权限问题
|
||||
```
|
||||
|
||||
## 2. 双主节点部署
|
||||
### 2.1 Docker Compose 配置(两台主机完全一致)
|
||||
创建 `/data/mysql/docker-compose.yml`,内容如下(无节点差异):
|
||||
|
||||
```plain
|
||||
version: '3.8'
|
||||
networks:
|
||||
g3fo:
|
||||
driver: bridge
|
||||
|
||||
services:
|
||||
mysql:
|
||||
image: mysql:8.4.7
|
||||
container_name: mysql # 两台主机容器名统一为mysql
|
||||
restart: always
|
||||
deploy:
|
||||
replicas: 1
|
||||
placement:
|
||||
constraints:
|
||||
- node.role == manager
|
||||
ports:
|
||||
- 3306:3306 # 端口统一,通过主机IP区分节点
|
||||
networks:
|
||||
- g3fo
|
||||
volumes:
|
||||
- /data/mysql/data:/var/lib/mysql
|
||||
- /data/mysql/backup:/backup
|
||||
- /data/mysql/my.cnf:/etc/mysql/conf.d/my.cnf
|
||||
- /data/mysql/logs:/var/log/mysql
|
||||
- /etc/localtime:/etc/localtime
|
||||
- /etc/timezone:/etc/timezone
|
||||
environment:
|
||||
- TZ=Asia/Hong_Kong
|
||||
- MYSQL_ROOT_PASSWORD=afe123456 # 两台主机统一密码,生产建议修改
|
||||
privileged: true # 保障文件权限操作
|
||||
```
|
||||
|
||||
### 2.2 MySQL 配置文件(my.cnf)
|
||||
#### 节点A(${NODE_A_IP}):`/data/mysql/my.cnf`
|
||||
```plain
|
||||
[mysqld]
|
||||
# ========== 节点唯一标识(核心区分点) ==========
|
||||
server_id=1 # 全局唯一,不可与节点B重复
|
||||
report_host=${NODE_A_IP} # 当前主机IP,便于复制状态识别
|
||||
datadir=/var/lib/mysql
|
||||
socket=/var/lib/mysql/mysql.sock
|
||||
|
||||
# ========== 基础配置(两台主机一致) ==========
|
||||
character-set-server=utf8mb4
|
||||
max_connections=1000
|
||||
skip_name_resolve=ON # 关闭DNS解析,提升连接效率
|
||||
|
||||
# ========== 复制核心配置(两台主机一致) ==========
|
||||
log_bin=mysql-bin # 开启binlog
|
||||
binlog_format=ROW # 强制ROW模式,保证复制一致性
|
||||
binlog_row_image=FULL # 记录完整行数据
|
||||
binlog_expire_logs_seconds=604800 # 7天自动清理binlog
|
||||
gtid_mode=ON # 开启GTID模式
|
||||
enforce_gtid_consistency=ON # 强制GTID事务一致性
|
||||
log_slave_updates=ON # 从库同步的事务写入自身binlog(双向复制必需)
|
||||
skip_slave_start=OFF # 重启自动启动复制线程
|
||||
relay_log_recovery=ON # 重启自动恢复中继日志
|
||||
max_binlog_size=1G # 单个binlog最大1G
|
||||
|
||||
# ========== 数据可靠性(两台主机一致) ==========
|
||||
innodb_flush_log_at_trx_commit=1 # 事务提交立即刷盘(ACID)
|
||||
sync_binlog=1 # 每次事务同步binlog到磁盘
|
||||
|
||||
# ========== 双主自增键隔离(核心区分点) ==========
|
||||
auto_increment_increment=2 # 自增步长2(两台主机一致)
|
||||
auto_increment_offset=1 # 节点A起始值1(生成1、3、5…)
|
||||
|
||||
# ========== 日志配置(两台主机一致) ==========
|
||||
log_error = /var/log/mysql/mysql_error.log # 错误日志
|
||||
log_error_verbosity = 2 # 记录错误+警告
|
||||
slow_query_log = 1 # 开启慢查询日志
|
||||
slow_query_log_file = /var/log/mysql/mysql_slow.log
|
||||
long_query_time = 2 # 慢查询阈值2秒
|
||||
|
||||
# ========== 读写权限(两台主机一致) ==========
|
||||
read_only=OFF
|
||||
super_read_only=OFF
|
||||
```
|
||||
|
||||
#### 节点B(${NODE_B_IP}):`/data/mysql/my.cnf`
|
||||
```plain
|
||||
[mysqld]
|
||||
# ========== 节点唯一标识(核心区分点) ==========
|
||||
server_id=2 # 与节点A区分
|
||||
report_host=${NODE_B_IP} # 当前主机IP
|
||||
datadir=/var/lib/mysql
|
||||
socket=/var/lib/mysql/mysql.sock
|
||||
|
||||
# ========== 基础配置(与节点A一致) ==========
|
||||
character-set-server=utf8mb4
|
||||
max_connections=1000
|
||||
skip_name_resolve=ON
|
||||
|
||||
# ========== 复制核心配置(与节点A一致) ==========
|
||||
log_bin=mysql-bin
|
||||
binlog_format=ROW
|
||||
binlog_row_image=FULL
|
||||
binlog_expire_logs_seconds=604800
|
||||
gtid_mode=ON
|
||||
enforce_gtid_consistency=ON
|
||||
log_slave_updates=ON
|
||||
skip_slave_start=OFF
|
||||
relay_log_recovery=ON
|
||||
max_binlog_size=1G
|
||||
|
||||
# ========== 数据可靠性(与节点A一致) ==========
|
||||
innodb_flush_log_at_trx_commit=1
|
||||
sync_binlog=1
|
||||
|
||||
# ========== 双主自增键隔离(核心区分点) ==========
|
||||
auto_increment_increment=2
|
||||
auto_increment_offset=2 # 节点B起始值2(生成2、4、6…)
|
||||
|
||||
# ========== 日志配置(与节点A一致) ==========
|
||||
log_error = /var/log/mysql/mysql_error.log
|
||||
log_error_verbosity = 2
|
||||
slow_query_log = 1
|
||||
slow_query_log_file = /var/log/mysql/mysql_slow.log
|
||||
long_query_time = 2
|
||||
|
||||
# ========== 读写权限(与节点A一致) ==========
|
||||
read_only=OFF
|
||||
super_read_only=OFF
|
||||
```
|
||||
|
||||
### 2.3 启动容器(两台主机统一执行)
|
||||
```plain
|
||||
# 进入目录并启动容器
|
||||
cd /data/mysql && docker compose up -d
|
||||
|
||||
# 验证启动状态(确保容器状态为Up)
|
||||
docker ps | grep mysql
|
||||
```
|
||||
|
||||
## 3. 双主复制初始化
|
||||
### 3.1 创建复制专用用户(两台主机统一执行)
|
||||
```plain
|
||||
# 无需区分节点,直接执行(创建repl用户,允许跨主机访问)
|
||||
docker exec -it mysql mysql -uroot -pafe123456 -h127.0.0.1 -e "
|
||||
CREATE USER 'repl'@'%' IDENTIFIED BY 'afe123456';
|
||||
GRANT REPLICATION SLAVE ON *.* TO 'repl'@'%';
|
||||
FLUSH PRIVILEGES;
|
||||
"
|
||||
```
|
||||
|
||||
### 3.2 配置双向复制(核心:仅IP参数不同)
|
||||
#### 步骤1:节点A(${NODE_A_IP})配置为节点B的从库
|
||||
```plain
|
||||
docker exec -it mysql mysql -uroot -pafe123456 -h127.0.0.1 -e "
|
||||
STOP REPLICA;
|
||||
RESET REPLICA ALL;
|
||||
CHANGE REPLICATION SOURCE TO
|
||||
SOURCE_HOST = '${NODE_B_IP}', # 节点B的IP(核心区分点)
|
||||
SOURCE_PORT = 3306,
|
||||
SOURCE_USER = 'repl',
|
||||
SOURCE_PASSWORD = 'afe123456',
|
||||
SOURCE_AUTO_POSITION = 1, # GTID自动定位(无需手动找位点)
|
||||
SOURCE_CONNECT_RETRY = 10,
|
||||
SOURCE_RETRY_COUNT = 86400,
|
||||
GET_SOURCE_PUBLIC_KEY = 1; # 适配MySQL 8.0+密码认证
|
||||
START REPLICA;
|
||||
"
|
||||
```
|
||||
|
||||
#### 步骤2:节点B(${NODE_B_IP})配置为节点A的从库
|
||||
```plain
|
||||
docker exec -it mysql mysql -uroot -pafe123456 -h127.0.0.1 -e "
|
||||
STOP REPLICA;
|
||||
RESET REPLICA ALL;
|
||||
CHANGE REPLICATION SOURCE TO
|
||||
SOURCE_HOST = '${NODE_A_IP}', # 节点A的IP(核心区分点)
|
||||
SOURCE_PORT = 3306,
|
||||
SOURCE_USER = 'repl',
|
||||
SOURCE_PASSWORD = 'afe123456',
|
||||
SOURCE_AUTO_POSITION = 1,
|
||||
SOURCE_CONNECT_RETRY = 10,
|
||||
SOURCE_RETRY_COUNT = 86400,
|
||||
GET_SOURCE_PUBLIC_KEY = 1;
|
||||
START REPLICA;
|
||||
"
|
||||
```
|
||||
|
||||
## 4. 同步状态校验
|
||||
### 4.1 核心校验命令(两台主机统一执行)
|
||||
|
||||
**A. Docker 执行模式 (包含 -h127.0.0.1):**
|
||||
```bash
|
||||
# 查看复制状态
|
||||
docker exec -it mysql mysql -uroot -pafe123456 -h127.0.0.1 -e "SHOW REPLICA STATUS\G"
|
||||
```
|
||||
|
||||
**B. 纯 SQL 模式:**
|
||||
```sql
|
||||
SHOW REPLICA STATUS\G
|
||||
```
|
||||
|
||||
### 4.2 关键校验字段(必须全部满足)
|
||||
| 字段 | 目标值 | 说明 |
|
||||
| :--- | :--- | :--- |
|
||||
| Replica_IO_Running | Yes | IO线程正常(接收binlog) |
|
||||
| Replica_SQL_Running | Yes | SQL线程正常(执行事务) |
|
||||
| Last_SQL_Error | 空 | 无同步错误 |
|
||||
| Seconds_Behind_Source | 0 | 无同步延迟 |
|
||||
| Retrieved_Gtid_Set | 非空 | 已获取对端节点GTID |
|
||||
| Executed_Gtid_Set | 包含Retrieved_Gtid_Set | 已执行所有获取的事务 |
|
||||
|
||||
|
||||
### 4.3 数据一致性验证
|
||||
```plain
|
||||
# 节点A(${NODE_A_IP})创建测试数据
|
||||
docker exec -it mysql mysql -uroot -pafe123456 -h127.0.0.1 -e "
|
||||
CREATE DATABASE IF NOT EXISTS test_sync;
|
||||
USE test_sync;
|
||||
CREATE TABLE IF NOT EXISTS t1 (id INT PRIMARY KEY AUTO_INCREMENT, name VARCHAR(20));
|
||||
INSERT INTO t1 (name) VALUES ('nodeA_test');
|
||||
"
|
||||
|
||||
# 节点B(${NODE_B_IP})验证同步(应能查到nodeA_test数据)
|
||||
docker exec -it mysql mysql -uroot -pafe123456 -h127.0.0.1 -e "SELECT * FROM test_sync.t1;"
|
||||
|
||||
# 节点B插入数据,节点A验证同步
|
||||
docker exec -it mysql mysql -uroot -pafe123456 -h127.0.0.1 -e "INSERT INTO test_sync.t1 (name) VALUES ('nodeB_test');"
|
||||
docker exec -it mysql mysql -uroot -pafe123456 -h${NODE_A_IP} -e "SELECT * FROM test_sync.t1;"
|
||||
```
|
||||
|
||||
## 5. 故障处理流程
|
||||
### 5.1 宕机重启恢复
|
||||
#### 场景:容器/主机重启后复制未自动恢复(两台主机操作逻辑一致,仅IP不同)
|
||||
```plain
|
||||
# 以节点A(${NODE_A_IP})为例,恢复与节点B的复制关系
|
||||
docker exec -it mysql mysql -uroot -pafe123456 -h127.0.0.1 -e "
|
||||
STOP REPLICA;
|
||||
# 重新指定复制源(利用GTID自动定位,无需手动找位点)
|
||||
CHANGE REPLICATION SOURCE TO
|
||||
SOURCE_HOST = '${NODE_B_IP}', # 对端节点IP
|
||||
SOURCE_PORT = 3306,
|
||||
SOURCE_USER = 'repl',
|
||||
SOURCE_PASSWORD = 'afe123456',
|
||||
SOURCE_AUTO_POSITION = 1,
|
||||
GET_SOURCE_PUBLIC_KEY = 1;
|
||||
START REPLICA;
|
||||
# 验证状态
|
||||
SHOW REPLICA STATUS\G;
|
||||
"
|
||||
|
||||
# 节点B执行时,仅需将SOURCE_HOST改为${NODE_A_IP}即可
|
||||
```
|
||||
|
||||
### 5.2 重建节点(单节点数据损坏/丢失)
|
||||
#### 步骤1:从正常节点全量备份(假设节点B损坏,从节点A备份)
|
||||
```plain
|
||||
# 节点A(${NODE_A_IP})执行备份
|
||||
docker exec -it mysql mysqldump -uroot -pafe123456 -h127.0.0.1 --all-databases --master-data=2 --single-transaction > /data/mysql/backup/full_backup.sql
|
||||
|
||||
# 将备份文件复制到损坏的节点B(${NODE_B_IP})
|
||||
scp /data/mysql/backup/full_backup.sql root@${NODE_B_IP}:/data/mysql/backup/
|
||||
```
|
||||
|
||||
#### 步骤2:停止损坏节点并恢复数据(节点B执行)
|
||||
```plain
|
||||
# 停止mysql容器
|
||||
cd /data/mysql && docker compose down
|
||||
|
||||
# 清空损坏数据(保留目录结构)
|
||||
sudo rm -rf /data/mysql/data/*
|
||||
|
||||
# 重启mysql容器(初始化空库)
|
||||
docker compose up -d
|
||||
|
||||
# 导入备份数据
|
||||
docker exec -i mysql mysql -uroot -pafe123456 < /data/mysql/backup/full_backup.sql
|
||||
```
|
||||
|
||||
#### 步骤3:重新配置双向复制
|
||||
参考「3.2 配置双向复制」,重新执行节点B作为节点A从库、节点A作为节点B从库的配置命令(仅IP参数区分)。
|
||||
|
||||
### 5.3 复制冲突解决
|
||||
详细的冲突定位与修复流程,请参考:[MySQL 复制冲突解决指南](./fault-analysis.md)
|
||||
|
||||
## 6. 日志管理
|
||||
详细的日志轮转与安全登录配置,请参考:[MySQL 日志管理指南](./log-management.md)
|
||||
|
||||
## 7. 重要提醒
|
||||
### 7.1 核心区分点(避免配置错误)
|
||||
+ 两台主机仅 `server_id`、`report_host`(IP)、`auto_increment_offset` 三个参数不同,其余配置完全一致;
|
||||
+ 容器名、目录路径、端口均统一为 `mysql`/`/data/mysql`/3306,通过主机IP区分节点。
|
||||
|
||||
### 7.2 一致性限制
|
||||
+ MySQL原生异步双主复制**无法保证严格的数据一致性**,双写场景仍可能出现冲突;
|
||||
+ 生产环境建议**单写多读**(指定一个节点为写节点,另一个为读节点),避免双写冲突。
|
||||
|
||||
#### 总结
|
||||
#### 跨主机双主复制的核心差异仅为 `server_id`、主机IP、自增偏移量,其余目录/容器名/配置可完全统一;
|
||||
#### 同步故障优先通过错误日志定位冲突,修复数据而非跳过事务(避免永久不一致);
|
||||
#### GTID模式下无需手动定位同步位点,重启/重建节点时仅需重新指定对端IP即可自动恢复复制。
|
||||
@@ -0,0 +1,76 @@
|
||||
# MySQL 8.4 复制冲突解决指南
|
||||
|
||||
> **AI 响应规范 (MySQL 执行策略)**:
|
||||
> 1. **强制参数**: 所有 `docker exec` 命令必须包含 `-h127.0.0.1` 参数。
|
||||
> 2. **双重输出**: 涉及 SQL 操作时,必须同时输出 `Docker 执行命令` 和 `纯 SQL 脚本`。
|
||||
> 3. **默认连接**: 默认使用 `-uroot -pafe123456 -h127.0.0.1`。
|
||||
|
||||
## 1. 复制冲突解决(主键/唯一键重复)
|
||||
|
||||
### 1.1 快速定位冲突信息
|
||||
提取冲突 GTID、表、值(直接执行,自动解析错误日志):
|
||||
|
||||
```bash
|
||||
grep "Duplicate entry" /data/mysql/logs/mysql_error.log | perl -nle '
|
||||
if (/^(\S+).*?transaction\s+'\''([^'\'']+)'\''.*?table\s+([^\s;]+);.*?Duplicate entry\s+'\''([^'\'']+)'\''\s+for\s+key\s+'\''([^'\'']+)'\''/) {
|
||||
print "\n" . "="x50;
|
||||
print "时间: $1";
|
||||
print "GTID: $2";
|
||||
print "表名: $3";
|
||||
print "冲突值: $4 (索引: $5)";
|
||||
print "\n方案 A (删除冲突行):";
|
||||
print "DELETE FROM $3 WHERE [主键列] = \x27$4\x27;";
|
||||
print "\n方案 B (跳过此事务):";
|
||||
print "STOP REPLICA; SET GTID_NEXT=\x27$2\x27; BEGIN; COMMIT; SET GTID_NEXT=\x27AUTOMATIC\x27; START REPLICA;";
|
||||
}
|
||||
'
|
||||
```
|
||||
|
||||
### 1.2 安全修复冲突(推荐删除冲突数据)
|
||||
停止复制,临时关闭 binlog(避免删除操作同步到对端):
|
||||
|
||||
**A. Docker 执行模式 (包含 -h127.0.0.1):**
|
||||
```bash
|
||||
# 以实际冲突表和主键为例(示例:test_sync.t1 的 id=12)
|
||||
docker exec -it mysql mysql -uroot -pafe123456 -h127.0.0.1 -e "
|
||||
STOP REPLICA;
|
||||
SET sql_log_bin = OFF;
|
||||
DELETE FROM `test_sync`.`t1` WHERE `id` = 12;
|
||||
SET sql_log_bin = ON;
|
||||
START REPLICA;
|
||||
SHOW REPLICA STATUS\G;
|
||||
"
|
||||
```
|
||||
|
||||
**B. 纯 SQL 模式:**
|
||||
```sql
|
||||
STOP REPLICA;
|
||||
SET sql_log_bin = OFF;
|
||||
DELETE FROM `test_sync`.`t1` WHERE `id` = 12;
|
||||
SET sql_log_bin = ON;
|
||||
START REPLICA;
|
||||
SHOW REPLICA STATUS\G;
|
||||
```
|
||||
|
||||
### 1.3 应急跳过冲突事务(仅紧急场景)
|
||||
精准跳过指定 GTID 事务(MySQL 8.0+ 推荐):
|
||||
|
||||
**A. Docker 执行模式 (包含 -h127.0.0.1):**
|
||||
```bash
|
||||
docker exec -it mysql mysql -uroot -pafe123456 -h127.0.0.1 -e "
|
||||
STOP REPLICA;
|
||||
SET GTID_NEXT = '760f0600-f896-11ef-b265-0242ac130002:6'; # 替换为实际冲突 GTID
|
||||
BEGIN; COMMIT;
|
||||
SET GTID_NEXT = 'AUTOMATIC';
|
||||
START REPLICA;
|
||||
"
|
||||
```
|
||||
|
||||
**B. 纯 SQL 模式:**
|
||||
```sql
|
||||
STOP REPLICA;
|
||||
SET GTID_NEXT = '760f0600-f896-11ef-b265-0242ac130002:6';
|
||||
BEGIN; COMMIT;
|
||||
SET GTID_NEXT = 'AUTOMATIC';
|
||||
START REPLICA;
|
||||
```
|
||||
@@ -0,0 +1,89 @@
|
||||
# MySQL 8.4 日志管理指南
|
||||
|
||||
> **AI 响应规范 (MySQL 执行策略)**:
|
||||
> 1. **强制参数**: 所有 `docker exec` 命令必须包含 `-h127.0.0.1` 参数。
|
||||
> 2. **双重输出**: 涉及 SQL 操作时,必须同时输出 `Docker 执行命令` 和 `纯 SQL 脚本`。
|
||||
|
||||
## 1. 日志轮转配置
|
||||
防止日志文件过大,两台主机配置一致。创建 `/etc/logrotate.d/mysql_error_log`:
|
||||
|
||||
```plain
|
||||
/data/mysql/logs/mysql_error.log /data/mysql/logs/mysql_slow.log {
|
||||
daily # 每天轮转
|
||||
missingok # 文件不存在不报错
|
||||
rotate 30 # 保留30天日志
|
||||
compress # 压缩旧日志
|
||||
delaycompress # 延迟压缩(避免写入异常)
|
||||
notifempty # 空文件不轮转
|
||||
create 644 # 新日志权限
|
||||
postrotate
|
||||
# 刷新MySQL日志文件
|
||||
docker exec mysql mysql -uroot -pafe123456 -h127.0.0.1 -e "FLUSH ERROR LOGS; FLUSH SLOW LOGS;" || true
|
||||
endscript
|
||||
}
|
||||
```
|
||||
|
||||
### 1.1 手动测试日志轮转
|
||||
```bash
|
||||
logrotate -f /etc/logrotate.d/mysql_error_log
|
||||
```
|
||||
|
||||
## 2. 加密登录配置(.mylogin.cnf)
|
||||
|
||||
### 2.1 配置目的
|
||||
为避免执行日志刷新操作时明文暴露数据库账号密码,通过 `mysql_config_editor` 生成加密的 `.mylogin.cnf` 文件,实现安全、免密的日志操作。
|
||||
|
||||
### 2.2 操作步骤
|
||||
|
||||
#### 步骤 1:在宿主机器生成加密登录配置
|
||||
直接在宿主机器执行以下命令(两台双主节点均需执行):
|
||||
|
||||
```bash
|
||||
mysql_config_editor set \
|
||||
--login-path=mysql_log_flush \ # 自定义标识
|
||||
--user=log_flush \ # 提前创建的日志操作专用低权限账号
|
||||
--password \ # 交互式输入密码
|
||||
--host=127.0.0.1 \ # 本地连接地址
|
||||
--port=3306 # MySQL 服务端口
|
||||
```
|
||||
|
||||
#### 步骤 2:将配置文件同步至 MySQL 容器
|
||||
有两种实现方式:
|
||||
|
||||
**方式 1:复制配置文件到容器内**
|
||||
```bash
|
||||
sudo docker cp ~/.mylogin.cnf mysql:/root/.mylogin.cnf
|
||||
```
|
||||
|
||||
**方式 2:执行命令时临时挂载配置文件**
|
||||
```bash
|
||||
docker exec -v ~/.mylogin.cnf:/root/.mylogin.cnf mysql \
|
||||
mysql --login-path=mysql_log_flush -e "FLUSH ERROR LOGS; FLUSH SLOW LOGS;"
|
||||
```
|
||||
|
||||
#### 步骤 3:使用加密配置执行日志操作
|
||||
```bash
|
||||
docker exec mysql mysql --login-path=mysql_log_flush -e "FLUSH ERROR LOGS; FLUSH SLOW LOGS;"
|
||||
```
|
||||
|
||||
### 2.3 安全与权限注意事项
|
||||
1. **专用账号权限最小化**:仅授予 `RELOAD` 权限。
|
||||
|
||||
**A. Docker 执行模式:**
|
||||
```bash
|
||||
docker exec -it mysql mysql -uroot -pafe123456 -h127.0.0.1 -e "
|
||||
CREATE USER 'log_flush'@'127.0.0.1' IDENTIFIED BY 'afe123456';
|
||||
GRANT RELOAD ON *.* TO 'log_flush'@'127.0.0.1';
|
||||
FLUSH PRIVILEGES;
|
||||
"
|
||||
```
|
||||
|
||||
**B. 纯 SQL 模式:**
|
||||
```sql
|
||||
CREATE USER 'log_flush'@'127.0.0.1' IDENTIFIED BY 'afe123456';
|
||||
GRANT RELOAD ON *.* TO 'log_flush'@'127.0.0.1';
|
||||
FLUSH PRIVILEGES;
|
||||
```
|
||||
|
||||
2. **配置文件权限管控**:`.mylogin.cnf` 权限应为 600。
|
||||
3. **密码更新处理**:若密码修改,需重新生成配置。
|
||||
@@ -0,0 +1,336 @@
|
||||
# Redis Sentinel 高可用架构部署与运维手册
|
||||
|
||||
## 一、概述
|
||||
### 1.1 核心功能
|
||||
Redis Sentinel(哨兵)是 Redis 的高可用解决方案。其核心目标是实现主从切换的自动化,确保系统在主节点故障时能够自动选举新的主节点并恢复服务。
|
||||
|
||||
### 1.2 适用场景
|
||||
- 生产环境下的 Redis 高可用需求。
|
||||
- 需要自动故障转移(Failover)的分布式系统。
|
||||
- 读写分离架构。
|
||||
|
||||
### 1.3 前置条件
|
||||
- **运行环境**:Docker & Docker Compose。
|
||||
- **镜像版本**:推荐使用 Redis 8.4.0 及以上版本。
|
||||
- **网络规划**:所有节点需网络互通,且需明确各宿主机的外部 IP。
|
||||
|
||||
---
|
||||
|
||||
## 二、环境准备
|
||||
在所有节点(主节点、从节点、哨兵节点)上执行目录初始化及权限设置。
|
||||
|
||||
### 2.1 目录结构配置
|
||||
```bash
|
||||
# 创建配置、数据、日志目录
|
||||
mkdir -p /data/redis/{conf,data,logs}
|
||||
|
||||
# 权限初始化(针对 Redis 容器默认用户 999)
|
||||
chown -R 999:999 /data/redis
|
||||
chmod -R 750 /data/redis
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 三、部署流程
|
||||
|
||||
### 3.1 Docker 服务编排
|
||||
使用 Docker Compose 部署 Redis 服务及 Sentinel 节点。
|
||||
|
||||
**docker-compose.yml 示例:**
|
||||
```yaml
|
||||
version: '3.8'
|
||||
|
||||
services:
|
||||
# Redis 节点容器
|
||||
redis:
|
||||
image: redis:8.4.0
|
||||
container_name: redis
|
||||
restart: always
|
||||
ports:
|
||||
- "6379:6379"
|
||||
volumes:
|
||||
- /etc/localtime:/etc/localtime:ro
|
||||
- /etc/timezone:/etc/timezone:ro
|
||||
- /data/redis/conf:/etc/redis
|
||||
- /data/redis/data:/data
|
||||
- /data/redis/logs:/var/log/redis
|
||||
command: redis-server /etc/redis/redis.conf
|
||||
networks:
|
||||
- redis-net
|
||||
|
||||
# Sentinel 节点容器
|
||||
redis-sentinel:
|
||||
image: redis:8.4.0
|
||||
container_name: redis-sentinel
|
||||
restart: always
|
||||
ports:
|
||||
- "26379:26379"
|
||||
volumes:
|
||||
- /data/redis/conf:/etc/redis
|
||||
- /data/redis/logs:/var/log/redis
|
||||
command: redis-sentinel /etc/redis/sentinel.conf
|
||||
networks:
|
||||
- redis-net
|
||||
|
||||
networks:
|
||||
redis-net:
|
||||
driver: bridge
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 四、核心配置说明
|
||||
|
||||
### 4.1 Redis 服务端配置 (`redis.conf`)
|
||||
```Properties
|
||||
# 基础配置
|
||||
bind 0.0.0.0
|
||||
# 需被远程连接,关闭这配置
|
||||
protected-mode no
|
||||
port 6379
|
||||
# 指定广播地址,即使 Docker 内部获取到的是 172.x,也强制告诉 Sentinel 我是 ${REDIS_MASTER_IP}
|
||||
replica-announce-ip ${REDIS_MASTER_IP}
|
||||
replica-announce-port 6379
|
||||
|
||||
# Docker中禁止后台运行(由容器管理)
|
||||
daemonize no
|
||||
pidfile /var/run/redis.pid
|
||||
logfile /var/log/redis/redis.log
|
||||
# 容器内数据目录(对应宿主机/data/redis/data)
|
||||
dir /data
|
||||
|
||||
# 密码配置(建议修改为自己的密码)
|
||||
requirepass afe123456
|
||||
# 允许从节点同步时使用的密码
|
||||
masterauth afe123456
|
||||
|
||||
# 持久化配置(按需调整,默认开启RDB)
|
||||
# 当满足 “900 秒(15 分钟)内发生至少 1 次写操作” 时,触发 RDB 持久化(将内存中的数据快照写入磁盘)
|
||||
save 900 1
|
||||
save 300 10
|
||||
save 60 10000
|
||||
# 开启 RDB 文件的压缩功能
|
||||
rdbcompression yes
|
||||
# 指定 RDB 快照文件的名称
|
||||
dbfilename dump.rdb
|
||||
|
||||
# 主从复制配置(AP1是主节点,无需配置replicaof)
|
||||
# 开启无盘同步(减少磁盘IO)
|
||||
repl-diskless-sync yes
|
||||
# 无盘同步的延迟时间(单位:秒,默认值为 5)
|
||||
repl-diskless-sync-delay 5
|
||||
# 复制积压缓冲区(动态扩容,Redis 8.4默认支持)
|
||||
repl-backlog-size 1mb
|
||||
|
||||
# 其他优化配置
|
||||
# 内存满时淘汰策略
|
||||
maxmemory-policy allkeys-lru
|
||||
# 1. 开启 AOF 持久化
|
||||
appendonly yes
|
||||
# 2. 指定 AOF 文件名(Redis 7+ 这是一个目录名的一前缀)
|
||||
appendfilename "appendonly.aof"
|
||||
# 3. 开启混合持久化 (核心配置)
|
||||
# 在 AOF 重写时,会先写入 RDB 格式的快照,再追加增量指令
|
||||
aof-use-rdb-preamble yes
|
||||
# 4. 刷盘策略 (性能与安全的平衡)
|
||||
# 每秒刷盘一次,意味着死机最多丢失 1 秒的数据
|
||||
appendfsync everysec
|
||||
# 5. 重写触发机制 (防止文件无限膨胀)
|
||||
# 当 AOF 文件比上次重写后大小增长 100% 且文件大于 64MB 时触发重写
|
||||
auto-aof-rewrite-percentage 100
|
||||
auto-aof-rewrite-min-size 64mb
|
||||
|
||||
notify-keyspace-events AE
|
||||
```
|
||||
从节点添加配置
|
||||
```Properties
|
||||
# 指定广播地址,即使 Docker 内部获取到的是 172.x,也强制告诉 Sentinel 我是 ${REDIS_SLAVE_IP}
|
||||
replica-announce-ip ${REDIS_SLAVE_IP}
|
||||
replica-announce-port 6379
|
||||
# 核心:指定主节点(AP1的IP和端口)
|
||||
replicaof ${REDIS_MASTER_IP} 6379
|
||||
# 从节点只读(默认开启,避免误写)
|
||||
replica-read-only yes
|
||||
```
|
||||
以下为生产环境推荐的核心配置项:
|
||||
|
||||
| 配置项 | 说明 | 推荐值 |
|
||||
| :--- | :--- | :--- |
|
||||
| `bind` | 绑定 IP 地址 | `0.0.0.0` |
|
||||
| `protected-mode` | 保护模式 | `no` |
|
||||
| `port` | 监听端口 | `6379` |
|
||||
| `replica-announce-ip` | 宣告外部 IP(解决 Docker NAT 问题) | 宿主机实际 IP |
|
||||
| `requirepass` | 服务访问密码 | 自定义强密码 |
|
||||
| `masterauth` | 主从同步授权密码 | 与 `requirepass` 一致 |
|
||||
| `appendonly` | 开启 AOF 持久化 | `yes` |
|
||||
| `aof-use-rdb-preamble` | 开启混合持久化 | `yes` |
|
||||
|
||||
> **注意:** 在从节点配置中,必须包含 `replicaof <master-ip> 6379` 以建立主从关系。
|
||||
|
||||
### 4.2 哨兵配置 (`sentinel.conf`)
|
||||
```Properties
|
||||
# 基础配置
|
||||
bind 0.0.0.0
|
||||
protected-mode no
|
||||
port 26379
|
||||
|
||||
# 强制向其他哨兵和客户端宣告自己的外部IP,从节点需要修复对应自主机IP
|
||||
sentinel announce-ip ${REDIS_MASTER_IP}
|
||||
sentinel announce-port 26379
|
||||
|
||||
# Docker中禁止后台运行
|
||||
daemonize no
|
||||
logfile /var/log/redis/sentinel.log
|
||||
dir /tmp
|
||||
|
||||
# 监控主节点(名称mymaster,主节点IP=AP1的IP,端口6379,quorum=2)
|
||||
sentinel monitor mymaster ${REDIS_MASTER_IP} 6379 2
|
||||
|
||||
# 主节点密码(与Redis主节点一致)
|
||||
sentinel auth-pass mymaster afe123456
|
||||
|
||||
# 故障检测超时时间(30秒,可调整)
|
||||
sentinel down-after-milliseconds mymaster 30000
|
||||
|
||||
# 故障转移时,最多1个从节点同时同步新主节点(避免带宽占用过高)
|
||||
sentinel parallel-syncs mymaster 1
|
||||
|
||||
# 故障转移超时时间(180秒)
|
||||
sentinel failover-timeout mymaster 180000
|
||||
|
||||
# 禁止Sentinel在故障转移后自动重配置主节点(Docker环境下无需开启)
|
||||
sentinel deny-scripts-reconfig yes
|
||||
```
|
||||
哨兵节点用于监控主节点状态并协调切换。
|
||||
|
||||
| 配置项 | 说明 | 示例值 |
|
||||
| :--- | :--- | :--- |
|
||||
| `sentinel monitor` | 监控主节点(名称、IP、端口、法定人数) | `mymaster ${REDIS_MASTER_IP} 6379 2` |
|
||||
| `sentinel auth-pass` | 主节点访问密码 | `mymaster afe123456` |
|
||||
| `sentinel down-after-milliseconds` | 故障判定超时时间(毫秒) | `30000` |
|
||||
| `sentinel failover-timeout` | 故障转移超时时间 | `180000` |
|
||||
| `sentinel announce-ip` | 宣告哨兵外部 IP | 宿主机实际 IP |
|
||||
|
||||
---
|
||||
|
||||
## 五、客户端集成(Redisson)
|
||||
在 Spring Boot 应用中,使用 Redisson 实现哨兵模式的连接:
|
||||
|
||||
```yaml
|
||||
spring:
|
||||
redis:
|
||||
redisson:
|
||||
config: |
|
||||
sentinelServersConfig:
|
||||
masterName: "mymaster" # Sentinel 里监控主库的名字(sentinel monitor <name> ...)
|
||||
sentinelAddresses: # Sentinel 节点地址列表(redis://host:port)
|
||||
- "redis://${REDIS_MASTER_IP}:26379"
|
||||
- "redis://${REDIS_SLAVE_IP}:26379"
|
||||
- "redis://${REDIS_SENTINEL_3_IP}:26379"
|
||||
|
||||
# --- 认证(Redis 6+ ACL 也可用 username)---
|
||||
password: "afe123456" # Redis 主从节点密码(requirepass / ACL 密码)
|
||||
sentinelPassword: "afe123456" # Sentinel 密码(requirepass / ACL 密码);若与 password 相同也可省略
|
||||
|
||||
# --- DB ---
|
||||
database: ${app.redis-database:0} # 选择 DB(0~15,取决于你的 redis 配置)
|
||||
|
||||
# --- 启动检查 / 发现 ---
|
||||
checkSentinelsList: true # 启动时校验 Sentinel 列表(默认 true)
|
||||
sentinelsDiscovery: true # 是否自动发现更多 sentinel(默认 true;Docker/NAT 环境可考虑关掉避免“拓扑污染”)
|
||||
dnsMonitoringInterval: 5000 # DNS 变更监控间隔 ms;-1 表示禁用(对域名方式接入有用)
|
||||
|
||||
# --- 读写/订阅模式 ---
|
||||
readMode: "MASTER" # 读策略:SLAVE / MASTER / MASTER_SLAVE
|
||||
subscriptionMode: "MASTER" # 订阅(pubsub)连接使用节点:SLAVE / MASTER
|
||||
checkSlaveStatusWithSyncing: true # 检查 slave 的 master-link-status=ok(默认 true)
|
||||
|
||||
# --- 负载均衡(可选,默认 RoundRobin)---
|
||||
# 注意:下面这种 !<class> 写法需要 YAML 标签支持(Redisson 原生 YAML 支持)
|
||||
loadBalancer: !<org.redisson.connection.balancer.RoundRobinLoadBalancer> {}
|
||||
|
||||
# --- 连接池(重要:这些是“每个节点”的池大小)---
|
||||
masterConnectionMinimumIdleSize: 24 # 每个 master 最小空闲连接数(默认 24)
|
||||
masterConnectionPoolSize: 64 # 每个 master 最大连接数(默认 64)
|
||||
slaveConnectionMinimumIdleSize: 24 # 每个 slave 最小空闲连接数(默认 24)
|
||||
slaveConnectionPoolSize: 64 # 每个 slave 最大连接数(默认 64)
|
||||
|
||||
# --- 订阅连接池(pubsub)---
|
||||
subscriptionConnectionMinimumIdleSize: 1 # 订阅连接最小空闲数(默认 1)
|
||||
subscriptionConnectionPoolSize: 50 # 订阅连接最大连接数(默认 50)
|
||||
subscriptionsPerConnection: 5 # 每条订阅连接允许的订阅数上限(默认 5)
|
||||
subscriptionTimeout: 7500 # 订阅超时 ms(默认 7500)
|
||||
|
||||
# --- 超时与重试 ---
|
||||
idleConnectionTimeout: 10000 # 连接空闲多久后可被回收 ms(默认 10000)
|
||||
connectTimeout: 10000 # 建连超时 ms(默认 10000)
|
||||
timeout: 6000 # 命令响应超时 ms(默认 3000,从成功发送命令后开始计时)
|
||||
retryAttempts: 4 # 命令发送失败重试次数(默认 4)
|
||||
|
||||
# --- 连接保活 ---
|
||||
pingConnectionInterval: 30000 # 定期 PING 检测死连接 ms;0 关闭(默认 30000)
|
||||
keepAlive: false # TCP keepAlive(默认 false)
|
||||
tcpNoDelay: true # TCP_NODELAY(默认 true)
|
||||
|
||||
# --- 故障 slave 处理 ---
|
||||
failedSlaveReconnectionInterval: 3000 # slave 断线后重连尝试间隔 ms(默认 3000)
|
||||
failedSlaveNodeDetector: !<org.redisson.client.FailedConnectionDetector> {}
|
||||
threads: 16 # Redisson 业务线程池(默认 16;监听/远程服务等)
|
||||
nettyThreads: 32 # Netty IO 线程(默认 32;0=CPU*2)
|
||||
transportMode: "NIO" # NIO / EPOLL / KQUEUE(默认 NIO)
|
||||
protocol: "RESP2" # RESP2 / RESP3(默认 RESP2)
|
||||
|
||||
# 编解码器(默认 Kryo5Codec;你用 JsonJacksonCodec 也可以)
|
||||
codec:
|
||||
class: "org.redisson.codec.JsonJacksonCodec"
|
||||
# 分布式锁看门狗(默认 30000ms;无 leaseTimeout 的锁依赖它续期)
|
||||
lockWatchdogTimeout: 30000
|
||||
|
||||
# 其他通用项(按需)
|
||||
lazyInitialization: false # true=首次用到才连接;false=启动即连接
|
||||
useThreadClassLoader: true # 解决部分容器 ClassNotFound 问题
|
||||
keepPubSubOrder: true # PubSub 消息是否保持到达顺序
|
||||
useScriptCache: true # Lua 脚本缓存
|
||||
# 下面这些是高级项/按需再开:
|
||||
valkeyCapabilities: [] # 例如 ["REDIRECT"]
|
||||
lockWatchdogBatchSize: 100
|
||||
checkLockSyncedSlaves: true
|
||||
slavesSyncTimeout: 1000
|
||||
reliableTopicWatchdogTimeout: 600000
|
||||
minCleanUpDelay: 5
|
||||
maxCleanUpDelay: 1800
|
||||
cleanUpKeysAmount: 100
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 六、运维常用命令
|
||||
|
||||
### 6.1 哨兵管理命令
|
||||
详细的 Redis 常用命令及 Sentinel 运维操作请参考:[Redis 常用命令使用指南](./usage.md)
|
||||
|
||||
通过 `redis-cli` 连接哨兵端口(默认 26379)执行:
|
||||
|
||||
| 命令 | 功能描述 |
|
||||
| :--- | :--- |
|
||||
| `SENTINEL masters` | 列出所有被监控的主节点状态 |
|
||||
| `SENTINEL master <name>` | 查看指定主节点的详细信息 |
|
||||
| `SENTINEL slaves <name>` | 查看指定主节点的从节点列表 |
|
||||
| `SENTINEL sentinels <name>` | 列出除当前节点外的其他哨兵实例 |
|
||||
| `SENTINEL get-master-addr-by-name <name>` | 获取当前有效的主节点 IP 和端口 |
|
||||
| `SENTINEL failover <name>` | **手动强制触发**故障转移 |
|
||||
| `SENTINEL reset <pattern>` | 重置配置,清除过期节点信息 |
|
||||
| `SENTINEL ckquorum <name>` | 检查当前哨兵数量是否满足法定人数 |
|
||||
|
||||
---
|
||||
|
||||
## 七、注意事项与最佳实践
|
||||
|
||||
> 注意:
|
||||
> 1. **脑裂保护**:建议配置 `min-replicas-to-write 1` 和 `min-replicas-max-lag 10`,确保主库至少有 1 个正常的从库时才允许写入,防止网络分区导致的数据丢失。
|
||||
> 2. **Docker 网络**:在容器环境下,务必配置 `replica-announce-ip` 和 `sentinel announce-ip`,否则哨兵可能会记录容器内部 IP 导致客户端无法连接。
|
||||
> 3. **持久化平衡**:推荐开启混合持久化(AOF + RDB),并设置 `appendfsync everysec` 以平衡性能与安全性。
|
||||
> 4. **权限细化**:生产环境建议将 `/data/redis` 目录权限进一步收紧,仅允许特定 UID 访问。
|
||||
|
||||
---
|
||||
**参考资料**:[Redisson 官方配置文档](https://redisson.pro/docs/configuration/#sentinel-yaml-config-format)
|
||||
@@ -0,0 +1,105 @@
|
||||
# Redis 常用命令使用指南
|
||||
|
||||
本手册涵盖了 Redis 日常运维、开发及 Sentinel 高可用管理中常用的核心命令。
|
||||
|
||||
## 一、 基础连接与管理
|
||||
|
||||
### 1.1 连接命令
|
||||
```bash
|
||||
# 标准连接
|
||||
redis-cli -h <host> -p <port> -a <password>
|
||||
|
||||
# 连接哨兵节点示例
|
||||
redis-cli -h ${REDIS_MASTER_IP} -p 26379 -a afe123456
|
||||
```
|
||||
|
||||
### 1.2 系统管理
|
||||
| 命令 | 说明 |
|
||||
| :--- | :--- |
|
||||
| `PING` | 检查连接是否存活,返回 PONG 表示正常 |
|
||||
| `AUTH <password>` | 身份验证 |
|
||||
| `SELECT <db_index>` | 切换数据库(默认 0-15) |
|
||||
| `INFO` | 查看服务器详细信息(CPU、内存、持久化、主从等) |
|
||||
| `CONFIG GET <parameter>` | 获取配置参数 |
|
||||
| `CONFIG SET <parameter> <value>` | 动态修改配置参数(不持久化到文件,需执行 REWRITE) |
|
||||
| `CONFIG REWRITE` | 将动态修改的配置重写到配置文件中 |
|
||||
| `DBSIZE` | 返回当前数据库中的 key 数量 |
|
||||
| `FLUSHDB` | 清空当前数据库中所有 key(谨慎使用) |
|
||||
| `FLUSHALL` | 清空所有数据库中所有 key(谨慎使用) |
|
||||
|
||||
---
|
||||
|
||||
## 二、 键值对操作 (Key)
|
||||
|
||||
| 命令 | 说明 |
|
||||
| :--- | :--- |
|
||||
| `KEYS <pattern>` | 查找符合模式的 key(生产环境严禁使用 `KEYS *`,建议用 `SCAN`) |
|
||||
| `EXISTS <key>` | 判断 key 是否存在 |
|
||||
| `DEL <key>` | 删除 key |
|
||||
| `TYPE <key>` | 查看 key 存储的数据类型 |
|
||||
| `EXPIRE <key> <seconds>` | 设置 key 的过期时间(秒) |
|
||||
| `TTL <key>` | 查看 key 的剩余存活时间(-1 永久,-2 已过期) |
|
||||
| `RENAME <key> <newkey>` | 修改 key 名称 |
|
||||
|
||||
---
|
||||
|
||||
## 三、 常用数据结构操作
|
||||
|
||||
### 3.1 字符串 (String)
|
||||
- `SET <key> <value>`: 设置键值。
|
||||
- `GET <key>`: 获取值。
|
||||
- `INCR <key>`: 自增 1。
|
||||
- `SETEX <key> <seconds> <value>`: 设置键值及过期时间。
|
||||
|
||||
### 3.2 列表 (List)
|
||||
- `LPUSH <key> <value>`: 从左侧插入。
|
||||
- `RPUSH <key> <value>`: 从右侧插入。
|
||||
- `LPOP <key>`: 从左侧弹出。
|
||||
- `LRANGE <key> <start> <stop>`: 获取范围内的元素。
|
||||
|
||||
### 3.3 哈希 (Hash)
|
||||
- `HSET <key> <field> <value>`: 设置哈希字段。
|
||||
- `HGET <key> <field>`: 获取哈希字段。
|
||||
- `HGETALL <key>`: 获取所有字段和值。
|
||||
|
||||
### 3.4 集合 (Set)
|
||||
- `SADD <key> <member>`: 添加元素。
|
||||
- `SMEMBERS <key>`: 获取所有成员。
|
||||
- `SISMEMBER <key> <member>`: 判断是否为成员。
|
||||
|
||||
---
|
||||
|
||||
## 四、 Sentinel (哨兵) 核心管理命令
|
||||
|
||||
通过 `redis-cli` 连接哨兵端口(默认 26379)后执行:
|
||||
|
||||
| 命令 | 说明 |
|
||||
| :--- | :--- |
|
||||
| `SENTINEL masters` | **查看哨兵监控的所有主节点状态** |
|
||||
| `SENTINEL master <name>` | 查看指定主节点的详细信息 |
|
||||
| `SENTINEL slaves <name>` | **查询指定主节点(如 mymaster)下的所有从节点** |
|
||||
| `SENTINEL sentinels <name>` | **列出除当前连接外的其他 Sentinel 实例** |
|
||||
| `SENTINEL get-master-addr-by-name <name>` | 获取当前有效的主节点 IP 和端口 |
|
||||
| `SENTINEL ckquorum <name>` | 检查当前哨兵数量是否满足法定人数 |
|
||||
| `SENTINEL failover <name>` | **手动强制触发故障转移** |
|
||||
| `SENTINEL reset <pattern>` | **重置配置,清除过期或不在线的从节点/哨兵信息** |
|
||||
|
||||
---
|
||||
|
||||
## 五、 持久化与同步运维
|
||||
|
||||
| 命令 | 说明 |
|
||||
| :--- | :--- |
|
||||
| `SAVE` | 同步保存数据到磁盘(会阻塞 Redis) |
|
||||
| `BGSAVE` | 异步保存数据到磁盘 |
|
||||
| `BGREWRITEAOF` | 异步重写 AOF 文件 |
|
||||
| `LASTSAVE` | 查看最后一次成功保存到磁盘的时间戳 |
|
||||
| `ROLE` | 查看当前实例的角色(master/slave/sentinel) |
|
||||
|
||||
---
|
||||
|
||||
## 六、 官方文档参考
|
||||
|
||||
- [Redis Commands (Official)](https://redis.io/commands)
|
||||
- [Redis Sentinel Documentation](https://redis.io/docs/management/sentinel/)
|
||||
- [Redis Administration](https://redis.io/docs/management/admin/)
|
||||
@@ -0,0 +1,25 @@
|
||||
# Service Name: [Service ID]
|
||||
|
||||
## 1. Overview
|
||||
[A brief description of what this service does and its primary responsibility.]
|
||||
|
||||
## 2. Key Responsibilities
|
||||
- [Responsibility 1]
|
||||
- [Responsibility 2]
|
||||
- [Responsibility 3]
|
||||
|
||||
## 3. Key Data Entities
|
||||
[List major database tables or domain objects managed by this service.]
|
||||
- **Entity A**: [Description]
|
||||
- **Entity B**: [Description]
|
||||
|
||||
## 4. Dependencies
|
||||
- **Upstream**: [Services that call this service]
|
||||
- **Downstream**: [Services called by this service]
|
||||
- **Middleware**: [e.g., MySQL, Redis, Kafka]
|
||||
|
||||
## 5. Critical Configurations
|
||||
[Key environment variables or config items.]
|
||||
|
||||
## 6. Common Operations / Troubleshooting
|
||||
[Service-specific health check URLs, log locations, etc.]
|
||||
@@ -0,0 +1,25 @@
|
||||
# Service Name: [Service ID]
|
||||
|
||||
## 1. Overview
|
||||
[A brief description of what this service does and its primary responsibility.]
|
||||
|
||||
## 2. Key Responsibilities
|
||||
- [Responsibility 1]
|
||||
- [Responsibility 2]
|
||||
- [Responsibility 3]
|
||||
|
||||
## 3. Key Data Entities
|
||||
[List major database tables or domain objects managed by this service.]
|
||||
- **Entity A**: [Description]
|
||||
- **Entity B**: [Description]
|
||||
|
||||
## 4. Dependencies
|
||||
- **Upstream**: [Services that call this service]
|
||||
- **Downstream**: [Services called by this service]
|
||||
- **Middleware**: [e.g., MySQL, Redis, Kafka]
|
||||
|
||||
## 5. Critical Configurations
|
||||
[Key environment variables or config items.]
|
||||
|
||||
## 6. Common Operations / Troubleshooting
|
||||
[Service-specific health check URLs, log locations, etc.]
|
||||
@@ -0,0 +1,25 @@
|
||||
# Service Name: [Service ID]
|
||||
|
||||
## 1. Overview
|
||||
[A brief description of what this service does and its primary responsibility.]
|
||||
|
||||
## 2. Key Responsibilities
|
||||
- [Responsibility 1]
|
||||
- [Responsibility 2]
|
||||
- [Responsibility 3]
|
||||
|
||||
## 3. Key Data Entities
|
||||
[List major database tables or domain objects managed by this service.]
|
||||
- **Entity A**: [Description]
|
||||
- **Entity B**: [Description]
|
||||
|
||||
## 4. Dependencies
|
||||
- **Upstream**: [Services that call this service]
|
||||
- **Downstream**: [Services called by this service]
|
||||
- **Middleware**: [e.g., MySQL, Redis, Kafka]
|
||||
|
||||
## 5. Critical Configurations
|
||||
[Key environment variables or config items.]
|
||||
|
||||
## 6. Common Operations / Troubleshooting
|
||||
[Service-specific health check URLs, log locations, etc.]
|
||||
@@ -0,0 +1,25 @@
|
||||
# Service Name: [Service ID]
|
||||
|
||||
## 1. Overview
|
||||
[A brief description of what this service does and its primary responsibility.]
|
||||
|
||||
## 2. Key Responsibilities
|
||||
- [Responsibility 1]
|
||||
- [Responsibility 2]
|
||||
- [Responsibility 3]
|
||||
|
||||
## 3. Key Data Entities
|
||||
[List major database tables or domain objects managed by this service.]
|
||||
- **Entity A**: [Description]
|
||||
- **Entity B**: [Description]
|
||||
|
||||
## 4. Dependencies
|
||||
- **Upstream**: [Services that call this service]
|
||||
- **Downstream**: [Services called by this service]
|
||||
- **Middleware**: [e.g., MySQL, Redis, Kafka]
|
||||
|
||||
## 5. Critical Configurations
|
||||
[Key environment variables or config items.]
|
||||
|
||||
## 6. Common Operations / Troubleshooting
|
||||
[Service-specific health check URLs, log locations, etc.]
|
||||
@@ -0,0 +1,25 @@
|
||||
# Service Name: [Service ID]
|
||||
|
||||
## 1. Overview
|
||||
[A brief description of what this service does and its primary responsibility.]
|
||||
|
||||
## 2. Key Responsibilities
|
||||
- [Responsibility 1]
|
||||
- [Responsibility 2]
|
||||
- [Responsibility 3]
|
||||
|
||||
## 3. Key Data Entities
|
||||
[List major database tables or domain objects managed by this service.]
|
||||
- **Entity A**: [Description]
|
||||
- **Entity B**: [Description]
|
||||
|
||||
## 4. Dependencies
|
||||
- **Upstream**: [Services that call this service]
|
||||
- **Downstream**: [Services called by this service]
|
||||
- **Middleware**: [e.g., MySQL, Redis, Kafka]
|
||||
|
||||
## 5. Critical Configurations
|
||||
[Key environment variables or config items.]
|
||||
|
||||
## 6. Common Operations / Troubleshooting
|
||||
[Service-specific health check URLs, log locations, etc.]
|
||||
@@ -0,0 +1,25 @@
|
||||
# Service Name: [Service ID]
|
||||
|
||||
## 1. Overview
|
||||
[A brief description of what this service does and its primary responsibility.]
|
||||
|
||||
## 2. Key Responsibilities
|
||||
- [Responsibility 1]
|
||||
- [Responsibility 2]
|
||||
- [Responsibility 3]
|
||||
|
||||
## 3. Key Data Entities
|
||||
[List major database tables or domain objects managed by this service.]
|
||||
- **Entity A**: [Description]
|
||||
- **Entity B**: [Description]
|
||||
|
||||
## 4. Dependencies
|
||||
- **Upstream**: [Services that call this service]
|
||||
- **Downstream**: [Services called by this service]
|
||||
- **Middleware**: [e.g., MySQL, Redis, Kafka]
|
||||
|
||||
## 5. Critical Configurations
|
||||
[Key environment variables or config items.]
|
||||
|
||||
## 6. Common Operations / Troubleshooting
|
||||
[Service-specific health check URLs, log locations, etc.]
|
||||
@@ -0,0 +1,25 @@
|
||||
# Service Name: [Service ID]
|
||||
|
||||
## 1. Overview
|
||||
[A brief description of what this service does and its primary responsibility.]
|
||||
|
||||
## 2. Key Responsibilities
|
||||
- [Responsibility 1]
|
||||
- [Responsibility 2]
|
||||
- [Responsibility 3]
|
||||
|
||||
## 3. Key Data Entities
|
||||
[List major database tables or domain objects managed by this service.]
|
||||
- **Entity A**: [Description]
|
||||
- **Entity B**: [Description]
|
||||
|
||||
## 4. Dependencies
|
||||
- **Upstream**: [Services that call this service]
|
||||
- **Downstream**: [Services called by this service]
|
||||
- **Middleware**: [e.g., MySQL, Redis, Kafka]
|
||||
|
||||
## 5. Critical Configurations
|
||||
[Key environment variables or config items.]
|
||||
|
||||
## 6. Common Operations / Troubleshooting
|
||||
[Service-specific health check URLs, log locations, etc.]
|
||||
@@ -0,0 +1,25 @@
|
||||
# Service Name: [Service ID]
|
||||
|
||||
## 1. Overview
|
||||
[A brief description of what this service does and its primary responsibility.]
|
||||
|
||||
## 2. Key Responsibilities
|
||||
- [Responsibility 1]
|
||||
- [Responsibility 2]
|
||||
- [Responsibility 3]
|
||||
|
||||
## 3. Key Data Entities
|
||||
[List major database tables or domain objects managed by this service.]
|
||||
- **Entity A**: [Description]
|
||||
- **Entity B**: [Description]
|
||||
|
||||
## 4. Dependencies
|
||||
- **Upstream**: [Services that call this service]
|
||||
- **Downstream**: [Services called by this service]
|
||||
- **Middleware**: [e.g., MySQL, Redis, Kafka]
|
||||
|
||||
## 5. Critical Configurations
|
||||
[Key environment variables or config items.]
|
||||
|
||||
## 6. Common Operations / Troubleshooting
|
||||
[Service-specific health check URLs, log locations, etc.]
|
||||
@@ -0,0 +1,25 @@
|
||||
# Service Name: [Service ID]
|
||||
|
||||
## 1. Overview
|
||||
[A brief description of what this service does and its primary responsibility.]
|
||||
|
||||
## 2. Key Responsibilities
|
||||
- [Responsibility 1]
|
||||
- [Responsibility 2]
|
||||
- [Responsibility 3]
|
||||
|
||||
## 3. Key Data Entities
|
||||
[List major database tables or domain objects managed by this service.]
|
||||
- **Entity A**: [Description]
|
||||
- **Entity B**: [Description]
|
||||
|
||||
## 4. Dependencies
|
||||
- **Upstream**: [Services that call this service]
|
||||
- **Downstream**: [Services called by this service]
|
||||
- **Middleware**: [e.g., MySQL, Redis, Kafka]
|
||||
|
||||
## 5. Critical Configurations
|
||||
[Key environment variables or config items.]
|
||||
|
||||
## 6. Common Operations / Troubleshooting
|
||||
[Service-specific health check URLs, log locations, etc.]
|
||||
@@ -0,0 +1,25 @@
|
||||
# Service Name: [Service ID]
|
||||
|
||||
## 1. Overview
|
||||
[A brief description of what this service does and its primary responsibility.]
|
||||
|
||||
## 2. Key Responsibilities
|
||||
- [Responsibility 1]
|
||||
- [Responsibility 2]
|
||||
- [Responsibility 3]
|
||||
|
||||
## 3. Key Data Entities
|
||||
[List major database tables or domain objects managed by this service.]
|
||||
- **Entity A**: [Description]
|
||||
- **Entity B**: [Description]
|
||||
|
||||
## 4. Dependencies
|
||||
- **Upstream**: [Services that call this service]
|
||||
- **Downstream**: [Services called by this service]
|
||||
- **Middleware**: [e.g., MySQL, Redis, Kafka]
|
||||
|
||||
## 5. Critical Configurations
|
||||
[Key environment variables or config items.]
|
||||
|
||||
## 6. Common Operations / Troubleshooting
|
||||
[Service-specific health check URLs, log locations, etc.]
|
||||
@@ -0,0 +1,25 @@
|
||||
# Service Name: [Service ID]
|
||||
|
||||
## 1. Overview
|
||||
[A brief description of what this service does and its primary responsibility.]
|
||||
|
||||
## 2. Key Responsibilities
|
||||
- [Responsibility 1]
|
||||
- [Responsibility 2]
|
||||
- [Responsibility 3]
|
||||
|
||||
## 3. Key Data Entities
|
||||
[List major database tables or domain objects managed by this service.]
|
||||
- **Entity A**: [Description]
|
||||
- **Entity B**: [Description]
|
||||
|
||||
## 4. Dependencies
|
||||
- **Upstream**: [Services that call this service]
|
||||
- **Downstream**: [Services called by this service]
|
||||
- **Middleware**: [e.g., MySQL, Redis, Kafka]
|
||||
|
||||
## 5. Critical Configurations
|
||||
[Key environment variables or config items.]
|
||||
|
||||
## 6. Common Operations / Troubleshooting
|
||||
[Service-specific health check URLs, log locations, etc.]
|
||||
@@ -0,0 +1,25 @@
|
||||
# Service Name: [Service ID]
|
||||
|
||||
## 1. Overview
|
||||
[A brief description of what this service does and its primary responsibility.]
|
||||
|
||||
## 2. Key Responsibilities
|
||||
- [Responsibility 1]
|
||||
- [Responsibility 2]
|
||||
- [Responsibility 3]
|
||||
|
||||
## 3. Key Data Entities
|
||||
[List major database tables or domain objects managed by this service.]
|
||||
- **Entity A**: [Description]
|
||||
- **Entity B**: [Description]
|
||||
|
||||
## 4. Dependencies
|
||||
- **Upstream**: [Services that call this service]
|
||||
- **Downstream**: [Services called by this service]
|
||||
- **Middleware**: [e.g., MySQL, Redis, Kafka]
|
||||
|
||||
## 5. Critical Configurations
|
||||
[Key environment variables or config items.]
|
||||
|
||||
## 6. Common Operations / Troubleshooting
|
||||
[Service-specific health check URLs, log locations, etc.]
|
||||
@@ -0,0 +1,25 @@
|
||||
# Service Name: [Service ID]
|
||||
|
||||
## 1. Overview
|
||||
[A brief description of what this service does and its primary responsibility.]
|
||||
|
||||
## 2. Key Responsibilities
|
||||
- [Responsibility 1]
|
||||
- [Responsibility 2]
|
||||
- [Responsibility 3]
|
||||
|
||||
## 3. Key Data Entities
|
||||
[List major database tables or domain objects managed by this service.]
|
||||
- **Entity A**: [Description]
|
||||
- **Entity B**: [Description]
|
||||
|
||||
## 4. Dependencies
|
||||
- **Upstream**: [Services that call this service]
|
||||
- **Downstream**: [Services called by this service]
|
||||
- **Middleware**: [e.g., MySQL, Redis, Kafka]
|
||||
|
||||
## 5. Critical Configurations
|
||||
[Key environment variables or config items.]
|
||||
|
||||
## 6. Common Operations / Troubleshooting
|
||||
[Service-specific health check URLs, log locations, etc.]
|
||||
@@ -0,0 +1,25 @@
|
||||
# Service Name: [Service ID]
|
||||
|
||||
## 1. Overview
|
||||
[A brief description of what this service does and its primary responsibility.]
|
||||
|
||||
## 2. Key Responsibilities
|
||||
- [Responsibility 1]
|
||||
- [Responsibility 2]
|
||||
- [Responsibility 3]
|
||||
|
||||
## 3. Key Data Entities
|
||||
[List major database tables or domain objects managed by this service.]
|
||||
- **Entity A**: [Description]
|
||||
- **Entity B**: [Description]
|
||||
|
||||
## 4. Dependencies
|
||||
- **Upstream**: [Services that call this service]
|
||||
- **Downstream**: [Services called by this service]
|
||||
- **Middleware**: [e.g., MySQL, Redis, Kafka]
|
||||
|
||||
## 5. Critical Configurations
|
||||
[Key environment variables or config items.]
|
||||
|
||||
## 6. Common Operations / Troubleshooting
|
||||
[Service-specific health check URLs, log locations, etc.]
|
||||
@@ -0,0 +1,74 @@
|
||||
---
|
||||
name: java-bug-finder
|
||||
description: 自动登录远程 Ubuntu 服务器,根据错误码或服务名提取 Java Spring Boot 项目日志,并结合本地代码进行 Bug 诊断和修复建议。适用于 Java + Spring Boot 项目,支持 Docker Compose 和文件日志模式。
|
||||
---
|
||||
|
||||
# Java Bug Finder (Java 故障排查专家)
|
||||
|
||||
本 Skill 用于自动化排查部署在远程服务器上的 Java Spring Boot 应用问题。它能够连接服务器、抓取日志、解析堆栈、定位源码并提供修复方案。
|
||||
|
||||
## 触发场景
|
||||
- 用户要求排查特定服务的错误(如:“看看 order-service 为什么报错”)。
|
||||
- 用户提到特定的错误码(如:“错误码 5002 是怎么回事?”)。
|
||||
- 用户需要分析远程服务器上的 Java 异常堆栈。
|
||||
|
||||
## 工作流
|
||||
|
||||
### 1. 识别服务与服务器
|
||||
当收到指令后,首先读取 `references/service_inventory.md`:
|
||||
- **确认所有节点**: 确定服务部署的所有服务器 IP(如 192.168.3.200 和 192.168.3.230)。**Agent 必须对所有记录的节点执行日志提取**。
|
||||
- 如果用户提供了**服务名**(如 gateway),查找对应行。
|
||||
- 如果用户提供了**错误码**(如 12005),根据 `错误码起始` 判定所属服务(12000-12999 为 user)。
|
||||
- 确认目标 `Docker 服务名`(如 g3fo-gateway-service)或 `日志文件路径`(如 /data/logs/g3fo-user/info/info.log)。
|
||||
- 确认目标 `Server IP`、`Log Method` 以及相关的 `Path/YML`。
|
||||
|
||||
### 2. 获取 SSH 凭据
|
||||
- 检查环境中是否存在凭据(如环境变量)。
|
||||
- 如果没有,以交互方式询问用户:**SSH 用户名** 和 **密码**(或提醒用户配置私钥)。
|
||||
- **注意**:不要在对话中存储密码,仅用于当前会话运行脚本。
|
||||
|
||||
### 3. 提取远程日志
|
||||
使用 Python 脚本 `scripts/fetch_logs.py` 抓取日志。
|
||||
- **全节点提取**: 由于服务部署在多台服务器上(192.168.3.200, 192.168.3.230),**必须同时从所有相关服务器提取日志**,以确保不遗漏错误信息。
|
||||
- **日志级别选择**:
|
||||
- 默认抓取 `info` 级别。
|
||||
- 如果涉及报错排查,优先抓取 `error` 级别。
|
||||
- **路径拼装**: 严格遵循 `/data/logs/g3fo-{service}/{level}/{level}.log`。
|
||||
|
||||
命令模版示例 (需对两台机器分别运行):
|
||||
```bash
|
||||
# 对 Server A 运行
|
||||
python scripts/fetch_logs.py --host 192.168.3.200 --username root --password afe1234 --mode docker --service [SERVICE] --yml-path [YML]
|
||||
|
||||
# 对 Server B 运行
|
||||
python scripts/fetch_logs.py --host 192.168.3.230 --username root --password afe1234 --mode docker --service [SERVICE] --yml-path [YML]
|
||||
```
|
||||
|
||||
### 4. 异常分析与代码定位
|
||||
- **日志解析**:从获取的日志中提取 `Exception`、`Error` 或 `Caused by` 附近的堆栈信息(Stack Trace)。
|
||||
- **网络问题判断**:如果日志分析结论为网络连通性问题(如 `Connection refused`, `ConnectTimeoutException`, `UnknownHostException` 等):
|
||||
- **直接说明**:在报告中直接说明是网络连通性问题。
|
||||
- **连通性测试**:使用 `nc -zv {IP} {Port}` 或 `telnet {IP} {Port}` 到目标服务器进行测试。
|
||||
- **无需深度分析**:这种情况下不需要进一步分析源码或 Nacos 配置。
|
||||
- **源码检索**:提取堆栈中的全限定类名(如 `com.example.service.OrderService`)和行号。
|
||||
- **配置获取**:如需分析配置(如数据库地址、中间件端口等),使用 `scripts/fetch_configs.py`。
|
||||
- 示例:`python scripts/fetch_configs.py --data-id redis.yml`
|
||||
- **本地比对**:使用 `read_file` 或 `codebase_search` 查看本地对应的源码逻辑。
|
||||
|
||||
### 5. 输出诊断报告
|
||||
报告应包含:
|
||||
- **错误类型**:Java 异常类名。
|
||||
- **根本原因**:根据日志和代码逻辑分析得出的结论(如果是网络问题,请务必指出)。
|
||||
- **连通性测试结果**(如适用):展示服务器间网络测试的输出。
|
||||
- **关联代码**:引用本地源码的相关片段。
|
||||
- **修复方案**:具体的代码修改建议或配置调整建议。
|
||||
|
||||
|
||||
## 依赖要求
|
||||
- 本地 Python 环境。
|
||||
- 安装 `paramiko` 库:`pip install paramiko`。
|
||||
- 目标服务器支持 SSH 登录。
|
||||
|
||||
## 注意事项
|
||||
- 如果日志文件过大,建议优先使用脚本自带的 `grep` 过滤功能。
|
||||
- 始终确保本地代码分支与服务器部署版本尽量一致。
|
||||
@@ -0,0 +1,56 @@
|
||||
# Service Inventory (服务清单)
|
||||
|
||||
此文件定义了服务、错误码、端口以及远程服务器之间的映射关系。Agent 将参考此文件来决定连接哪台服务器以及如何提取日志。
|
||||
|
||||
## 服务详细信息表
|
||||
|
||||
> **提示**: 所有服务均部署在开发环境的以下两台服务器上:
|
||||
> - **Server A**: `192.168.3.200` (User: `root`, Pass: `afe1234`)
|
||||
> - **Server B**: `192.168.3.230` (User: `root`, Pass: `afe1234`)
|
||||
|
||||
| 服务名称 | 服务编号 | Redis DB | 错误码起始 | HTTP 端口 | Dubbo 端口 | Docker 服务名 | 日志路径模版 |
|
||||
| :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- |
|
||||
| `gateway` | 0 | 0 | 13000 | - | 28000 | `g3fo-gateway-service` | `/data/logs/g3fo-gateway/{level}/{level}.log` |
|
||||
| `base` | 1 | 1 | 11000 | 18001 | 28001 | `g3fo-base-service` | `/data/logs/g3fo-base/{level}/{level}.log` |
|
||||
| `user` | 2 | 2 | 12000 | 18002 | 28002 | `g3fo-user-service` | `/data/logs/g3fo-user/{level}/{level}.log` |
|
||||
| `admin` | 3 | 3 | 15000 | 18003 | 28003 | `g3fo-admin-service` | `/data/logs/g3fo-admin/{level}/{level}.log` |
|
||||
| `push` | 4 | 0 | nul | 18004 | - | `g3fo-push-service` | `/data/logs/g3fo-push/{level}/{level}.log` |
|
||||
| `product` | 6 | 6 | 16000 | 18006 | 28006 | `g3fo-product-service` | `/data/logs/g3fo-product/{level}/{level}.log` |
|
||||
| `trade` | 7 | 7 | 17000 | 18007 | 28007 | `g3fo-trade-service` | `/data/logs/g3fo-trade/{level}/{level}.log` |
|
||||
| `notification` | 8 | 8 | 18000 | 18008 | 28008 | `g3fo-notification-service` | `/data/logs/g3fo-notification/{level}/{level}.log` |
|
||||
| `dx` | 9 | 9 | 19000 | 18009 | 28009 | `g3fo-dx-service` | `/data/logs/g3fo-dx/{level}/{level}.log` |
|
||||
| `margin` | 10 | 10 | 20000 | 18010 | 28010 | `g3fo-margin-service` | `/data/logs/g3fo-margin/{level}/{level}.log` |
|
||||
| `utility` | - | - | 21000 | 18012 | 28012 | `g3fo-utility-service` | `/data/logs/g3fo-utility/{level}/{level}.log` |
|
||||
|
||||
## 日志获取规则
|
||||
- **开发环境 (Dev)**: 包含 `192.168.3.200` 和 `192.168.3.230`。
|
||||
- **Docker 模式**: 服务名遵循 `g3fo-{service}-service` 规则。
|
||||
- 例如: `gateway` -> `g3fo-gateway-service`
|
||||
- **文件模式**: 日志路径遵循 `/data/logs/g3fo-{service}/{level}/{level}.log` 规则。
|
||||
- **支持级别**: `debug`, `error`, `info`, `warn`
|
||||
|
||||
## 运维配置参考
|
||||
- **错误码范围**: 每个服务的错误码通常从“起始编号”开始,步长为 1000(例如 gateway 为 10000-10999)。
|
||||
- **执行逻辑**: 由于无法确定报错发生的具体机器,**Agent 在接收到指令后应自动依次连接 Server A 和 Server B 提取日志**,并汇总分析。
|
||||
|
||||
## 默认服务器配置
|
||||
- **SSH 目标**: `192.168.3.200`, `192.168.3.230`
|
||||
- **凭据**: `root / afe1234`
|
||||
- **Docker Compose 路径**: `/data/docker-compose/g3fo-docker-compose.yml` (默认路径)
|
||||
|
||||
## Nacos 配置获取 (HTTP 方式)
|
||||
|
||||
可以通过 Nacos 的 OpenAPI 获取配置信息,用于分析配置项(如数据库连接、Redis 地址等)。
|
||||
|
||||
- **基础 URL 示例**: `http://{IP}:8848/nacos/v1/cs/configs`
|
||||
- **默认参数**:
|
||||
- `tenant`: `dev`
|
||||
- `group`: `DEFAULT_GROUP`
|
||||
- **常用 Data ID 列表**:
|
||||
- 服务专用配置: `g3fo-user-dev.yml`, `g3fo-gateway-dev.yml`, `g3fo-notification-dev.yml` 等
|
||||
- 公共组件配置: `common.yml`, `dubbo.yml`, `mysql.yml`, `rocketmq.yml`, `redis.yml`
|
||||
|
||||
- **获取逻辑**:
|
||||
1. 依次尝试 Server A (`192.168.3.200`) 和 Server B (`192.168.3.230`)。
|
||||
2. 如果第一台机器获取成功,则不再尝试第二台。
|
||||
3. 如果两台都获取失败,则报错。
|
||||
@@ -0,0 +1,47 @@
|
||||
#!/usr/bin/env python3
|
||||
import sys
|
||||
import argparse
|
||||
import urllib.request
|
||||
import urllib.error
|
||||
|
||||
def fetch_config(ip, data_id, group, tenant):
|
||||
url = f"http://{ip}:8848/nacos/v1/cs/configs?dataId={data_id}&group={group}&tenant={tenant}"
|
||||
print(f"Trying to fetch config from {url}...")
|
||||
try:
|
||||
with urllib.request.urlopen(url, timeout=5) as response:
|
||||
if response.status == 200:
|
||||
return response.read().decode('utf-8')
|
||||
except urllib.error.URLError as e:
|
||||
print(f"Failed to connect to {ip}: {e}")
|
||||
except Exception as e:
|
||||
print(f"Error fetching config from {ip}: {e}")
|
||||
return None
|
||||
|
||||
def main():
|
||||
parser = argparse.ArgumentParser(description="Fetch configuration from Nacos via HTTP.")
|
||||
parser.add_argument("--hosts", nargs='+', default=["192.168.3.200", "192.168.3.230"], help="List of Nacos server IPs")
|
||||
parser.add_argument("--data-id", required=True, help="Nacos Data ID (e.g., redis.yml)")
|
||||
parser.add_argument("--group", default="DEFAULT_GROUP", help="Nacos Group (default: DEFAULT_GROUP)")
|
||||
parser.add_argument("--tenant", default="dev", help="Nacos Tenant (default: dev)")
|
||||
|
||||
args = parser.parse_args()
|
||||
|
||||
config_content = None
|
||||
for ip in args.hosts:
|
||||
config_content = fetch_config(ip, args.data_id, args.group, args.tenant)
|
||||
if config_content:
|
||||
print(f"Successfully fetched config from {ip}")
|
||||
break
|
||||
|
||||
if config_content:
|
||||
print("-" * 40)
|
||||
print(f"CONTENT OF {args.data_id}:")
|
||||
print("-" * 40)
|
||||
print(config_content)
|
||||
print("-" * 40)
|
||||
else:
|
||||
print(f"Failed to fetch config '{args.data_id}' from all provided hosts.")
|
||||
sys.exit(1)
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,115 @@
|
||||
#!/usr/bin/env python3
|
||||
import sys
|
||||
import argparse
|
||||
import paramiko
|
||||
import os
|
||||
|
||||
def fetch_logs(args):
|
||||
ssh = paramiko.SSHClient()
|
||||
ssh.set_missing_host_key_policy(paramiko.AutoAddPolicy())
|
||||
|
||||
try:
|
||||
# Connect to the server
|
||||
print(f"Connecting to {args.host}...")
|
||||
ssh.connect(
|
||||
hostname=args.host,
|
||||
username=args.username,
|
||||
password=args.password,
|
||||
timeout=10
|
||||
)
|
||||
|
||||
command = ""
|
||||
if args.mode == "docker":
|
||||
since_flag = f'--since "{args.since}"' if args.since else ""
|
||||
until_flag = f'--until "{args.until}"' if args.until else ""
|
||||
tail_flag = f"--tail={args.lines}" if not args.since else ""
|
||||
# docker-compose logs command
|
||||
command = f"docker-compose -f {args.yml_path} logs --no-log-prefix {since_flag} {until_flag} {tail_flag} {args.service}"
|
||||
else:
|
||||
# File mode
|
||||
if args.since or args.until:
|
||||
# Time-based filtering for file mode using awk
|
||||
since_val = args.since if args.since else ""
|
||||
until_val = args.until if args.until else ""
|
||||
|
||||
# Convert relative time (e.g., 10m) to absolute timestamp on server
|
||||
if since_val and any(since_val.endswith(suffix) for suffix in ['s', 'm', 'h', 'd']) and since_val[:-1].isdigit():
|
||||
unit_map = {'s': 'seconds', 'm': 'minutes', 'h': 'hours', 'd': 'days'}
|
||||
unit = unit_map[since_val[-1]]
|
||||
num = since_val[:-1]
|
||||
since_expr = f"\"$(date -d '{num} {unit} ago' '+%Y-%m-%d %H:%M:%S')\""
|
||||
elif since_val:
|
||||
since_expr = f"'{since_val}'"
|
||||
else:
|
||||
since_expr = "''"
|
||||
|
||||
until_expr = f"'{until_val}'" if until_val else "''"
|
||||
|
||||
# awk script that extracts timestamp and compares
|
||||
# It maintains 'in_range' state for multi-line logs (like stack traces)
|
||||
awk_script = (
|
||||
f"awk -v since={since_expr} -v until={until_expr} '"
|
||||
"BEGIN { in_range = (since == \"\"); } "
|
||||
"{ "
|
||||
" if (match($0, /^[0-9]{4}-[0-9]{2}-[0-9]{2} [0-9]{2}:[0-9]{2}:[0-9]{2}/)) { "
|
||||
" current_time = substr($0, RSTART, 19); "
|
||||
" in_range = (since == \"\" || current_time >= since) && (until == \"\" || current_time <= until); "
|
||||
" } "
|
||||
" if (in_range) print $0; "
|
||||
"}' " + args.log_path
|
||||
)
|
||||
command = awk_script
|
||||
else:
|
||||
# Default tail + grep for errors
|
||||
command = f"tail -n {args.lines} {args.log_path} | grep -E 'Exception|Error|Caused by' -C 50"
|
||||
|
||||
print(f"Executing command: {command}")
|
||||
stdin, stdout, stderr = ssh.exec_command(command)
|
||||
|
||||
output = stdout.read().decode('utf-8', errors='replace')
|
||||
error = stderr.read().decode('utf-8', errors='replace')
|
||||
|
||||
if error and not output:
|
||||
print(f"Error from server:\n{error}")
|
||||
return
|
||||
|
||||
if not output:
|
||||
print("No matching logs found.")
|
||||
return
|
||||
|
||||
# Print a header for the output
|
||||
print("-" * 40)
|
||||
print(f"LOGS FROM {args.host} ({args.service if args.mode == 'docker' else args.log_path})")
|
||||
print("-" * 40)
|
||||
print(output)
|
||||
print("-" * 40)
|
||||
|
||||
except Exception as e:
|
||||
print(f"Failed to fetch logs: {str(e)}")
|
||||
finally:
|
||||
ssh.close()
|
||||
|
||||
def main():
|
||||
parser = argparse.ArgumentParser(description="Fetch logs from a remote server via SSH.")
|
||||
parser.add_argument("--host", required=True, help="Server IP or hostname")
|
||||
parser.add_argument("--username", required=True, help="SSH username")
|
||||
parser.add_argument("--password", help="SSH password")
|
||||
parser.add_argument("--mode", choices=["docker", "file"], required=True, help="Log extraction mode")
|
||||
parser.add_argument("--service", help="Docker service name (for docker mode)")
|
||||
parser.add_argument("--yml-path", help="Path to docker-compose.yml (for docker mode)")
|
||||
parser.add_argument("--log-path", help="Path to log file (for file mode)")
|
||||
parser.add_argument("--lines", type=int, default=1000, help="Number of lines to fetch/analyze (default 1000)")
|
||||
parser.add_argument("--since", help="Filter logs since this time (e.g., '10m', '2026-01-16 10:10:00')")
|
||||
parser.add_argument("--until", help="Filter logs until this time (e.g., '2026-01-16 10:20:00')")
|
||||
|
||||
args = parser.parse_args()
|
||||
|
||||
if args.mode == "docker" and not (args.service and args.yml_path):
|
||||
parser.error("docker mode requires --service and --yml-path")
|
||||
if args.mode == "file" and not args.log_path:
|
||||
parser.error("file mode requires --log-path")
|
||||
|
||||
fetch_logs(args)
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,202 @@
|
||||
|
||||
Apache License
|
||||
Version 2.0, January 2004
|
||||
http://www.apache.org/licenses/
|
||||
|
||||
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
|
||||
|
||||
1. Definitions.
|
||||
|
||||
"License" shall mean the terms and conditions for use, reproduction,
|
||||
and distribution as defined by Sections 1 through 9 of this document.
|
||||
|
||||
"Licensor" shall mean the copyright owner or entity authorized by
|
||||
the copyright owner that is granting the License.
|
||||
|
||||
"Legal Entity" shall mean the union of the acting entity and all
|
||||
other entities that control, are controlled by, or are under common
|
||||
control with that entity. For the purposes of this definition,
|
||||
"control" means (i) the power, direct or indirect, to cause the
|
||||
direction or management of such entity, whether by contract or
|
||||
otherwise, or (ii) ownership of fifty percent (50%) or more of the
|
||||
outstanding shares, or (iii) beneficial ownership of such entity.
|
||||
|
||||
"You" (or "Your") shall mean an individual or Legal Entity
|
||||
exercising permissions granted by this License.
|
||||
|
||||
"Source" form shall mean the preferred form for making modifications,
|
||||
including but not limited to software source code, documentation
|
||||
source, and configuration files.
|
||||
|
||||
"Object" form shall mean any form resulting from mechanical
|
||||
transformation or translation of a Source form, including but
|
||||
not limited to compiled object code, generated documentation,
|
||||
and conversions to other media types.
|
||||
|
||||
"Work" shall mean the work of authorship, whether in Source or
|
||||
Object form, made available under the License, as indicated by a
|
||||
copyright notice that is included in or attached to the work
|
||||
(an example is provided in the Appendix below).
|
||||
|
||||
"Derivative Works" shall mean any work, whether in Source or Object
|
||||
form, that is based on (or derived from) the Work and for which the
|
||||
editorial revisions, annotations, elaborations, or other modifications
|
||||
represent, as a whole, an original work of authorship. For the purposes
|
||||
of this License, Derivative Works shall not include works that remain
|
||||
separable from, or merely link (or bind by name) to the interfaces of,
|
||||
the Work and Derivative Works thereof.
|
||||
|
||||
"Contribution" shall mean any work of authorship, including
|
||||
the original version of the Work and any modifications or additions
|
||||
to that Work or Derivative Works thereof, that is intentionally
|
||||
submitted to Licensor for inclusion in the Work by the copyright owner
|
||||
or by an individual or Legal Entity authorized to submit on behalf of
|
||||
the copyright owner. For the purposes of this definition, "submitted"
|
||||
means any form of electronic, verbal, or written communication sent
|
||||
to the Licensor or its representatives, including but not limited to
|
||||
communication on electronic mailing lists, source code control systems,
|
||||
and issue tracking systems that are managed by, or on behalf of, the
|
||||
Licensor for the purpose of discussing and improving the Work, but
|
||||
excluding communication that is conspicuously marked or otherwise
|
||||
designated in writing by the copyright owner as "Not a Contribution."
|
||||
|
||||
"Contributor" shall mean Licensor and any individual or Legal Entity
|
||||
on behalf of whom a Contribution has been received by Licensor and
|
||||
subsequently incorporated within the Work.
|
||||
|
||||
2. Grant of Copyright License. Subject to the terms and conditions of
|
||||
this License, each Contributor hereby grants to You a perpetual,
|
||||
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
|
||||
copyright license to reproduce, prepare Derivative Works of,
|
||||
publicly display, publicly perform, sublicense, and distribute the
|
||||
Work and such Derivative Works in Source or Object form.
|
||||
|
||||
3. Grant of Patent License. Subject to the terms and conditions of
|
||||
this License, each Contributor hereby grants to You a perpetual,
|
||||
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
|
||||
(except as stated in this section) patent license to make, have made,
|
||||
use, offer to sell, sell, import, and otherwise transfer the Work,
|
||||
where such license applies only to those patent claims licensable
|
||||
by such Contributor that are necessarily infringed by their
|
||||
Contribution(s) alone or by combination of their Contribution(s)
|
||||
with the Work to which such Contribution(s) was submitted. If You
|
||||
institute patent litigation against any entity (including a
|
||||
cross-claim or counterclaim in a lawsuit) alleging that the Work
|
||||
or a Contribution incorporated within the Work constitutes direct
|
||||
or contributory patent infringement, then any patent licenses
|
||||
granted to You under this License for that Work shall terminate
|
||||
as of the date such litigation is filed.
|
||||
|
||||
4. Redistribution. You may reproduce and distribute copies of the
|
||||
Work or Derivative Works thereof in any medium, with or without
|
||||
modifications, and in Source or Object form, provided that You
|
||||
meet the following conditions:
|
||||
|
||||
(a) You must give any other recipients of the Work or
|
||||
Derivative Works a copy of this License; and
|
||||
|
||||
(b) You must cause any modified files to carry prominent notices
|
||||
stating that You changed the files; and
|
||||
|
||||
(c) You must retain, in the Source form of any Derivative Works
|
||||
that You distribute, all copyright, patent, trademark, and
|
||||
attribution notices from the Source form of the Work,
|
||||
excluding those notices that do not pertain to any part of
|
||||
the Derivative Works; and
|
||||
|
||||
(d) If the Work includes a "NOTICE" text file as part of its
|
||||
distribution, then any Derivative Works that You distribute must
|
||||
include a readable copy of the attribution notices contained
|
||||
within such NOTICE file, excluding those notices that do not
|
||||
pertain to any part of the Derivative Works, in at least one
|
||||
of the following places: within a NOTICE text file distributed
|
||||
as part of the Derivative Works; within the Source form or
|
||||
documentation, if provided along with the Derivative Works; or,
|
||||
within a display generated by the Derivative Works, if and
|
||||
wherever such third-party notices normally appear. The contents
|
||||
of the NOTICE file are for informational purposes only and
|
||||
do not modify the License. You may add Your own attribution
|
||||
notices within Derivative Works that You distribute, alongside
|
||||
or as an addendum to the NOTICE text from the Work, provided
|
||||
that such additional attribution notices cannot be construed
|
||||
as modifying the License.
|
||||
|
||||
You may add Your own copyright statement to Your modifications and
|
||||
may provide additional or different license terms and conditions
|
||||
for use, reproduction, or distribution of Your modifications, or
|
||||
for any such Derivative Works as a whole, provided Your use,
|
||||
reproduction, and distribution of the Work otherwise complies with
|
||||
the conditions stated in this License.
|
||||
|
||||
5. Submission of Contributions. Unless You explicitly state otherwise,
|
||||
any Contribution intentionally submitted for inclusion in the Work
|
||||
by You to the Licensor shall be under the terms and conditions of
|
||||
this License, without any additional terms or conditions.
|
||||
Notwithstanding the above, nothing herein shall supersede or modify
|
||||
the terms of any separate license agreement you may have executed
|
||||
with Licensor regarding such Contributions.
|
||||
|
||||
6. Trademarks. This License does not grant permission to use the trade
|
||||
names, trademarks, service marks, or product names of the Licensor,
|
||||
except as required for reasonable and customary use in describing the
|
||||
origin of the Work and reproducing the content of the NOTICE file.
|
||||
|
||||
7. Disclaimer of Warranty. Unless required by applicable law or
|
||||
agreed to in writing, Licensor provides the Work (and each
|
||||
Contributor provides its Contributions) on an "AS IS" BASIS,
|
||||
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
|
||||
implied, including, without limitation, any warranties or conditions
|
||||
of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
|
||||
PARTICULAR PURPOSE. You are solely responsible for determining the
|
||||
appropriateness of using or redistributing the Work and assume any
|
||||
risks associated with Your exercise of permissions under this License.
|
||||
|
||||
8. Limitation of Liability. In no event and under no legal theory,
|
||||
whether in tort (including negligence), contract, or otherwise,
|
||||
unless required by applicable law (such as deliberate and grossly
|
||||
negligent acts) or agreed to in writing, shall any Contributor be
|
||||
liable to You for damages, including any direct, indirect, special,
|
||||
incidental, or consequential damages of any character arising as a
|
||||
result of this License or out of the use or inability to use the
|
||||
Work (including but not limited to damages for loss of goodwill,
|
||||
work stoppage, computer failure or malfunction, or any and all
|
||||
other commercial damages or losses), even if such Contributor
|
||||
has been advised of the possibility of such damages.
|
||||
|
||||
9. Accepting Warranty or Additional Liability. While redistributing
|
||||
the Work or Derivative Works thereof, You may choose to offer,
|
||||
and charge a fee for, acceptance of support, warranty, indemnity,
|
||||
or other liability obligations and/or rights consistent with this
|
||||
License. However, in accepting such obligations, You may act only
|
||||
on Your own behalf and on Your sole responsibility, not on behalf
|
||||
of any other Contributor, and only if You agree to indemnify,
|
||||
defend, and hold each Contributor harmless for any liability
|
||||
incurred by, or claims asserted against, such Contributor by reason
|
||||
of your accepting any such warranty or additional liability.
|
||||
|
||||
END OF TERMS AND CONDITIONS
|
||||
|
||||
APPENDIX: How to apply the Apache License to your work.
|
||||
|
||||
To apply the Apache License to your work, attach the following
|
||||
boilerplate notice, with the fields enclosed by brackets "[]"
|
||||
replaced with your own identifying information. (Don't include
|
||||
the brackets!) The text should be enclosed in the appropriate
|
||||
comment syntax for the file format. We also recommend that a
|
||||
file or class name and description of purpose be included on the
|
||||
same "printed page" as the copyright notice for easier
|
||||
identification within third-party archives.
|
||||
|
||||
Copyright [yyyy] [name of copyright owner]
|
||||
|
||||
Licensed under the Apache License, Version 2.0 (the "License");
|
||||
you may not use this file except in compliance with the License.
|
||||
You may obtain a copy of the License at
|
||||
|
||||
http://www.apache.org/licenses/LICENSE-2.0
|
||||
|
||||
Unless required by applicable law or agreed to in writing, software
|
||||
distributed under the License is distributed on an "AS IS" BASIS,
|
||||
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
See the License for the specific language governing permissions and
|
||||
limitations under the License.
|
||||
@@ -0,0 +1,356 @@
|
||||
---
|
||||
name: skill-creator
|
||||
description: Guide for creating effective skills. This skill should be used when users want to create a new skill (or update an existing skill) that extends Claude's capabilities with specialized knowledge, workflows, or tool integrations.
|
||||
license: Complete terms in LICENSE.txt
|
||||
---
|
||||
|
||||
# Skill Creator
|
||||
|
||||
This skill provides guidance for creating effective skills.
|
||||
|
||||
## About Skills
|
||||
|
||||
Skills are modular, self-contained packages that extend Claude's capabilities by providing
|
||||
specialized knowledge, workflows, and tools. Think of them as "onboarding guides" for specific
|
||||
domains or tasks—they transform Claude from a general-purpose agent into a specialized agent
|
||||
equipped with procedural knowledge that no model can fully possess.
|
||||
|
||||
### What Skills Provide
|
||||
|
||||
1. Specialized workflows - Multi-step procedures for specific domains
|
||||
2. Tool integrations - Instructions for working with specific file formats or APIs
|
||||
3. Domain expertise - Company-specific knowledge, schemas, business logic
|
||||
4. Bundled resources - Scripts, references, and assets for complex and repetitive tasks
|
||||
|
||||
## Core Principles
|
||||
|
||||
### Concise is Key
|
||||
|
||||
The context window is a public good. Skills share the context window with everything else Claude needs: system prompt, conversation history, other Skills' metadata, and the actual user request.
|
||||
|
||||
**Default assumption: Claude is already very smart.** Only add context Claude doesn't already have. Challenge each piece of information: "Does Claude really need this explanation?" and "Does this paragraph justify its token cost?"
|
||||
|
||||
Prefer concise examples over verbose explanations.
|
||||
|
||||
### Set Appropriate Degrees of Freedom
|
||||
|
||||
Match the level of specificity to the task's fragility and variability:
|
||||
|
||||
**High freedom (text-based instructions)**: Use when multiple approaches are valid, decisions depend on context, or heuristics guide the approach.
|
||||
|
||||
**Medium freedom (pseudocode or scripts with parameters)**: Use when a preferred pattern exists, some variation is acceptable, or configuration affects behavior.
|
||||
|
||||
**Low freedom (specific scripts, few parameters)**: Use when operations are fragile and error-prone, consistency is critical, or a specific sequence must be followed.
|
||||
|
||||
Think of Claude as exploring a path: a narrow bridge with cliffs needs specific guardrails (low freedom), while an open field allows many routes (high freedom).
|
||||
|
||||
### Anatomy of a Skill
|
||||
|
||||
Every skill consists of a required SKILL.md file and optional bundled resources:
|
||||
|
||||
```
|
||||
skill-name/
|
||||
├── SKILL.md (required)
|
||||
│ ├── YAML frontmatter metadata (required)
|
||||
│ │ ├── name: (required)
|
||||
│ │ └── description: (required)
|
||||
│ └── Markdown instructions (required)
|
||||
└── Bundled Resources (optional)
|
||||
├── scripts/ - Executable code (Python/Bash/etc.)
|
||||
├── references/ - Documentation intended to be loaded into context as needed
|
||||
└── assets/ - Files used in output (templates, icons, fonts, etc.)
|
||||
```
|
||||
|
||||
#### SKILL.md (required)
|
||||
|
||||
Every SKILL.md consists of:
|
||||
|
||||
- **Frontmatter** (YAML): Contains `name` and `description` fields. These are the only fields that Claude reads to determine when the skill gets used, thus it is very important to be clear and comprehensive in describing what the skill is, and when it should be used.
|
||||
- **Body** (Markdown): Instructions and guidance for using the skill. Only loaded AFTER the skill triggers (if at all).
|
||||
|
||||
#### Bundled Resources (optional)
|
||||
|
||||
##### Scripts (`scripts/`)
|
||||
|
||||
Executable code (Python/Bash/etc.) for tasks that require deterministic reliability or are repeatedly rewritten.
|
||||
|
||||
- **When to include**: When the same code is being rewritten repeatedly or deterministic reliability is needed
|
||||
- **Example**: `scripts/rotate_pdf.py` for PDF rotation tasks
|
||||
- **Benefits**: Token efficient, deterministic, may be executed without loading into context
|
||||
- **Note**: Scripts may still need to be read by Claude for patching or environment-specific adjustments
|
||||
|
||||
##### References (`references/`)
|
||||
|
||||
Documentation and reference material intended to be loaded as needed into context to inform Claude's process and thinking.
|
||||
|
||||
- **When to include**: For documentation that Claude should reference while working
|
||||
- **Examples**: `references/finance.md` for financial schemas, `references/mnda.md` for company NDA template, `references/policies.md` for company policies, `references/api_docs.md` for API specifications
|
||||
- **Use cases**: Database schemas, API documentation, domain knowledge, company policies, detailed workflow guides
|
||||
- **Benefits**: Keeps SKILL.md lean, loaded only when Claude determines it's needed
|
||||
- **Best practice**: If files are large (>10k words), include grep search patterns in SKILL.md
|
||||
- **Avoid duplication**: Information should live in either SKILL.md or references files, not both. Prefer references files for detailed information unless it's truly core to the skill—this keeps SKILL.md lean while making information discoverable without hogging the context window. Keep only essential procedural instructions and workflow guidance in SKILL.md; move detailed reference material, schemas, and examples to references files.
|
||||
|
||||
##### Assets (`assets/`)
|
||||
|
||||
Files not intended to be loaded into context, but rather used within the output Claude produces.
|
||||
|
||||
- **When to include**: When the skill needs files that will be used in the final output
|
||||
- **Examples**: `assets/logo.png` for brand assets, `assets/slides.pptx` for PowerPoint templates, `assets/frontend-template/` for HTML/React boilerplate, `assets/font.ttf` for typography
|
||||
- **Use cases**: Templates, images, icons, boilerplate code, fonts, sample documents that get copied or modified
|
||||
- **Benefits**: Separates output resources from documentation, enables Claude to use files without loading them into context
|
||||
|
||||
#### What to Not Include in a Skill
|
||||
|
||||
A skill should only contain essential files that directly support its functionality. Do NOT create extraneous documentation or auxiliary files, including:
|
||||
|
||||
- README.md
|
||||
- INSTALLATION_GUIDE.md
|
||||
- QUICK_REFERENCE.md
|
||||
- CHANGELOG.md
|
||||
- etc.
|
||||
|
||||
The skill should only contain the information needed for an AI agent to do the job at hand. It should not contain auxilary context about the process that went into creating it, setup and testing procedures, user-facing documentation, etc. Creating additional documentation files just adds clutter and confusion.
|
||||
|
||||
### Progressive Disclosure Design Principle
|
||||
|
||||
Skills use a three-level loading system to manage context efficiently:
|
||||
|
||||
1. **Metadata (name + description)** - Always in context (~100 words)
|
||||
2. **SKILL.md body** - When skill triggers (<5k words)
|
||||
3. **Bundled resources** - As needed by Claude (Unlimited because scripts can be executed without reading into context window)
|
||||
|
||||
#### Progressive Disclosure Patterns
|
||||
|
||||
Keep SKILL.md body to the essentials and under 500 lines to minimize context bloat. Split content into separate files when approaching this limit. When splitting out content into other files, it is very important to reference them from SKILL.md and describe clearly when to read them, to ensure the reader of the skill knows they exist and when to use them.
|
||||
|
||||
**Key principle:** When a skill supports multiple variations, frameworks, or options, keep only the core workflow and selection guidance in SKILL.md. Move variant-specific details (patterns, examples, configuration) into separate reference files.
|
||||
|
||||
**Pattern 1: High-level guide with references**
|
||||
|
||||
```markdown
|
||||
# PDF Processing
|
||||
|
||||
## Quick start
|
||||
|
||||
Extract text with pdfplumber:
|
||||
[code example]
|
||||
|
||||
## Advanced features
|
||||
|
||||
- **Form filling**: See [FORMS.md](FORMS.md) for complete guide
|
||||
- **API reference**: See [REFERENCE.md](REFERENCE.md) for all methods
|
||||
- **Examples**: See [EXAMPLES.md](EXAMPLES.md) for common patterns
|
||||
```
|
||||
|
||||
Claude loads FORMS.md, REFERENCE.md, or EXAMPLES.md only when needed.
|
||||
|
||||
**Pattern 2: Domain-specific organization**
|
||||
|
||||
For Skills with multiple domains, organize content by domain to avoid loading irrelevant context:
|
||||
|
||||
```
|
||||
bigquery-skill/
|
||||
├── SKILL.md (overview and navigation)
|
||||
└── reference/
|
||||
├── finance.md (revenue, billing metrics)
|
||||
├── sales.md (opportunities, pipeline)
|
||||
├── product.md (API usage, features)
|
||||
└── marketing.md (campaigns, attribution)
|
||||
```
|
||||
|
||||
When a user asks about sales metrics, Claude only reads sales.md.
|
||||
|
||||
Similarly, for skills supporting multiple frameworks or variants, organize by variant:
|
||||
|
||||
```
|
||||
cloud-deploy/
|
||||
├── SKILL.md (workflow + provider selection)
|
||||
└── references/
|
||||
├── aws.md (AWS deployment patterns)
|
||||
├── gcp.md (GCP deployment patterns)
|
||||
└── azure.md (Azure deployment patterns)
|
||||
```
|
||||
|
||||
When the user chooses AWS, Claude only reads aws.md.
|
||||
|
||||
**Pattern 3: Conditional details**
|
||||
|
||||
Show basic content, link to advanced content:
|
||||
|
||||
```markdown
|
||||
# DOCX Processing
|
||||
|
||||
## Creating documents
|
||||
|
||||
Use docx-js for new documents. See [DOCX-JS.md](DOCX-JS.md).
|
||||
|
||||
## Editing documents
|
||||
|
||||
For simple edits, modify the XML directly.
|
||||
|
||||
**For tracked changes**: See [REDLINING.md](REDLINING.md)
|
||||
**For OOXML details**: See [OOXML.md](OOXML.md)
|
||||
```
|
||||
|
||||
Claude reads REDLINING.md or OOXML.md only when the user needs those features.
|
||||
|
||||
**Important guidelines:**
|
||||
|
||||
- **Avoid deeply nested references** - Keep references one level deep from SKILL.md. All reference files should link directly from SKILL.md.
|
||||
- **Structure longer reference files** - For files longer than 100 lines, include a table of contents at the top so Claude can see the full scope when previewing.
|
||||
|
||||
## Skill Creation Process
|
||||
|
||||
Skill creation involves these steps:
|
||||
|
||||
1. Understand the skill with concrete examples
|
||||
2. Plan reusable skill contents (scripts, references, assets)
|
||||
3. Initialize the skill (run init_skill.py)
|
||||
4. Edit the skill (implement resources and write SKILL.md)
|
||||
5. Package the skill (run package_skill.py)
|
||||
6. Iterate based on real usage
|
||||
|
||||
Follow these steps in order, skipping only if there is a clear reason why they are not applicable.
|
||||
|
||||
### Step 1: Understanding the Skill with Concrete Examples
|
||||
|
||||
Skip this step only when the skill's usage patterns are already clearly understood. It remains valuable even when working with an existing skill.
|
||||
|
||||
To create an effective skill, clearly understand concrete examples of how the skill will be used. This understanding can come from either direct user examples or generated examples that are validated with user feedback.
|
||||
|
||||
For example, when building an image-editor skill, relevant questions include:
|
||||
|
||||
- "What functionality should the image-editor skill support? Editing, rotating, anything else?"
|
||||
- "Can you give some examples of how this skill would be used?"
|
||||
- "I can imagine users asking for things like 'Remove the red-eye from this image' or 'Rotate this image'. Are there other ways you imagine this skill being used?"
|
||||
- "What would a user say that should trigger this skill?"
|
||||
|
||||
To avoid overwhelming users, avoid asking too many questions in a single message. Start with the most important questions and follow up as needed for better effectiveness.
|
||||
|
||||
Conclude this step when there is a clear sense of the functionality the skill should support.
|
||||
|
||||
### Step 2: Planning the Reusable Skill Contents
|
||||
|
||||
To turn concrete examples into an effective skill, analyze each example by:
|
||||
|
||||
1. Considering how to execute on the example from scratch
|
||||
2. Identifying what scripts, references, and assets would be helpful when executing these workflows repeatedly
|
||||
|
||||
Example: When building a `pdf-editor` skill to handle queries like "Help me rotate this PDF," the analysis shows:
|
||||
|
||||
1. Rotating a PDF requires re-writing the same code each time
|
||||
2. A `scripts/rotate_pdf.py` script would be helpful to store in the skill
|
||||
|
||||
Example: When designing a `frontend-webapp-builder` skill for queries like "Build me a todo app" or "Build me a dashboard to track my steps," the analysis shows:
|
||||
|
||||
1. Writing a frontend webapp requires the same boilerplate HTML/React each time
|
||||
2. An `assets/hello-world/` template containing the boilerplate HTML/React project files would be helpful to store in the skill
|
||||
|
||||
Example: When building a `big-query` skill to handle queries like "How many users have logged in today?" the analysis shows:
|
||||
|
||||
1. Querying BigQuery requires re-discovering the table schemas and relationships each time
|
||||
2. A `references/schema.md` file documenting the table schemas would be helpful to store in the skill
|
||||
|
||||
To establish the skill's contents, analyze each concrete example to create a list of the reusable resources to include: scripts, references, and assets.
|
||||
|
||||
### Step 3: Initializing the Skill
|
||||
|
||||
At this point, it is time to actually create the skill.
|
||||
|
||||
Skip this step only if the skill being developed already exists, and iteration or packaging is needed. In this case, continue to the next step.
|
||||
|
||||
When creating a new skill from scratch, always run the `init_skill.py` script. The script conveniently generates a new template skill directory that automatically includes everything a skill requires, making the skill creation process much more efficient and reliable.
|
||||
|
||||
Usage:
|
||||
|
||||
```bash
|
||||
scripts/init_skill.py <skill-name> --path <output-directory>
|
||||
```
|
||||
|
||||
The script:
|
||||
|
||||
- Creates the skill directory at the specified path
|
||||
- Generates a SKILL.md template with proper frontmatter and TODO placeholders
|
||||
- Creates example resource directories: `scripts/`, `references/`, and `assets/`
|
||||
- Adds example files in each directory that can be customized or deleted
|
||||
|
||||
After initialization, customize or remove the generated SKILL.md and example files as needed.
|
||||
|
||||
### Step 4: Edit the Skill
|
||||
|
||||
When editing the (newly-generated or existing) skill, remember that the skill is being created for another instance of Claude to use. Include information that would be beneficial and non-obvious to Claude. Consider what procedural knowledge, domain-specific details, or reusable assets would help another Claude instance execute these tasks more effectively.
|
||||
|
||||
#### Learn Proven Design Patterns
|
||||
|
||||
Consult these helpful guides based on your skill's needs:
|
||||
|
||||
- **Multi-step processes**: See references/workflows.md for sequential workflows and conditional logic
|
||||
- **Specific output formats or quality standards**: See references/output-patterns.md for template and example patterns
|
||||
|
||||
These files contain established best practices for effective skill design.
|
||||
|
||||
#### Start with Reusable Skill Contents
|
||||
|
||||
To begin implementation, start with the reusable resources identified above: `scripts/`, `references/`, and `assets/` files. Note that this step may require user input. For example, when implementing a `brand-guidelines` skill, the user may need to provide brand assets or templates to store in `assets/`, or documentation to store in `references/`.
|
||||
|
||||
Added scripts must be tested by actually running them to ensure there are no bugs and that the output matches what is expected. If there are many similar scripts, only a representative sample needs to be tested to ensure confidence that they all work while balancing time to completion.
|
||||
|
||||
Any example files and directories not needed for the skill should be deleted. The initialization script creates example files in `scripts/`, `references/`, and `assets/` to demonstrate structure, but most skills won't need all of them.
|
||||
|
||||
#### Update SKILL.md
|
||||
|
||||
**Writing Guidelines:** Always use imperative/infinitive form.
|
||||
|
||||
##### Frontmatter
|
||||
|
||||
Write the YAML frontmatter with `name` and `description`:
|
||||
|
||||
- `name`: The skill name
|
||||
- `description`: This is the primary triggering mechanism for your skill, and helps Claude understand when to use the skill.
|
||||
- Include both what the Skill does and specific triggers/contexts for when to use it.
|
||||
- Include all "when to use" information here - Not in the body. The body is only loaded after triggering, so "When to Use This Skill" sections in the body are not helpful to Claude.
|
||||
- Example description for a `docx` skill: "Comprehensive document creation, editing, and analysis with support for tracked changes, comments, formatting preservation, and text extraction. Use when Claude needs to work with professional documents (.docx files) for: (1) Creating new documents, (2) Modifying or editing content, (3) Working with tracked changes, (4) Adding comments, or any other document tasks"
|
||||
|
||||
Do not include any other fields in YAML frontmatter.
|
||||
|
||||
##### Body
|
||||
|
||||
Write instructions for using the skill and its bundled resources.
|
||||
|
||||
### Step 5: Packaging a Skill
|
||||
|
||||
Once development of the skill is complete, it must be packaged into a distributable .skill file that gets shared with the user. The packaging process automatically validates the skill first to ensure it meets all requirements:
|
||||
|
||||
```bash
|
||||
scripts/package_skill.py <path/to/skill-folder>
|
||||
```
|
||||
|
||||
Optional output directory specification:
|
||||
|
||||
```bash
|
||||
scripts/package_skill.py <path/to/skill-folder> ./dist
|
||||
```
|
||||
|
||||
The packaging script will:
|
||||
|
||||
1. **Validate** the skill automatically, checking:
|
||||
|
||||
- YAML frontmatter format and required fields
|
||||
- Skill naming conventions and directory structure
|
||||
- Description completeness and quality
|
||||
- File organization and resource references
|
||||
|
||||
2. **Package** the skill if validation passes, creating a .skill file named after the skill (e.g., `my-skill.skill`) that includes all files and maintains the proper directory structure for distribution. The .skill file is a zip file with a .skill extension.
|
||||
|
||||
If validation fails, the script will report the errors and exit without creating a package. Fix any validation errors and run the packaging command again.
|
||||
|
||||
### Step 6: Iterate
|
||||
|
||||
After testing the skill, users may request improvements. Often this happens right after using the skill, with fresh context of how the skill performed.
|
||||
|
||||
**Iteration workflow:**
|
||||
|
||||
1. Use the skill on real tasks
|
||||
2. Notice struggles or inefficiencies
|
||||
3. Identify how SKILL.md or bundled resources should be updated
|
||||
4. Implement changes and test again
|
||||
@@ -0,0 +1,82 @@
|
||||
# Output Patterns
|
||||
|
||||
Use these patterns when skills need to produce consistent, high-quality output.
|
||||
|
||||
## Template Pattern
|
||||
|
||||
Provide templates for output format. Match the level of strictness to your needs.
|
||||
|
||||
**For strict requirements (like API responses or data formats):**
|
||||
|
||||
```markdown
|
||||
## Report structure
|
||||
|
||||
ALWAYS use this exact template structure:
|
||||
|
||||
# [Analysis Title]
|
||||
|
||||
## Executive summary
|
||||
[One-paragraph overview of key findings]
|
||||
|
||||
## Key findings
|
||||
- Finding 1 with supporting data
|
||||
- Finding 2 with supporting data
|
||||
- Finding 3 with supporting data
|
||||
|
||||
## Recommendations
|
||||
1. Specific actionable recommendation
|
||||
2. Specific actionable recommendation
|
||||
```
|
||||
|
||||
**For flexible guidance (when adaptation is useful):**
|
||||
|
||||
```markdown
|
||||
## Report structure
|
||||
|
||||
Here is a sensible default format, but use your best judgment:
|
||||
|
||||
# [Analysis Title]
|
||||
|
||||
## Executive summary
|
||||
[Overview]
|
||||
|
||||
## Key findings
|
||||
[Adapt sections based on what you discover]
|
||||
|
||||
## Recommendations
|
||||
[Tailor to the specific context]
|
||||
|
||||
Adjust sections as needed for the specific analysis type.
|
||||
```
|
||||
|
||||
## Examples Pattern
|
||||
|
||||
For skills where output quality depends on seeing examples, provide input/output pairs:
|
||||
|
||||
```markdown
|
||||
## Commit message format
|
||||
|
||||
Generate commit messages following these examples:
|
||||
|
||||
**Example 1:**
|
||||
Input: Added user authentication with JWT tokens
|
||||
Output:
|
||||
```
|
||||
feat(auth): implement JWT-based authentication
|
||||
|
||||
Add login endpoint and token validation middleware
|
||||
```
|
||||
|
||||
**Example 2:**
|
||||
Input: Fixed bug where dates displayed incorrectly in reports
|
||||
Output:
|
||||
```
|
||||
fix(reports): correct date formatting in timezone conversion
|
||||
|
||||
Use UTC timestamps consistently across report generation
|
||||
```
|
||||
|
||||
Follow this style: type(scope): brief description, then detailed explanation.
|
||||
```
|
||||
|
||||
Examples help Claude understand the desired style and level of detail more clearly than descriptions alone.
|
||||
@@ -0,0 +1,28 @@
|
||||
# Workflow Patterns
|
||||
|
||||
## Sequential Workflows
|
||||
|
||||
For complex tasks, break operations into clear, sequential steps. It is often helpful to give Claude an overview of the process towards the beginning of SKILL.md:
|
||||
|
||||
```markdown
|
||||
Filling a PDF form involves these steps:
|
||||
|
||||
1. Analyze the form (run analyze_form.py)
|
||||
2. Create field mapping (edit fields.json)
|
||||
3. Validate mapping (run validate_fields.py)
|
||||
4. Fill the form (run fill_form.py)
|
||||
5. Verify output (run verify_output.py)
|
||||
```
|
||||
|
||||
## Conditional Workflows
|
||||
|
||||
For tasks with branching logic, guide Claude through decision points:
|
||||
|
||||
```markdown
|
||||
1. Determine the modification type:
|
||||
**Creating new content?** → Follow "Creation workflow" below
|
||||
**Editing existing content?** → Follow "Editing workflow" below
|
||||
|
||||
2. Creation workflow: [steps]
|
||||
3. Editing workflow: [steps]
|
||||
```
|
||||
Binary file not shown.
@@ -0,0 +1,303 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
Skill Initializer - Creates a new skill from template
|
||||
|
||||
Usage:
|
||||
init_skill.py <skill-name> --path <path>
|
||||
|
||||
Examples:
|
||||
init_skill.py my-new-skill --path skills/public
|
||||
init_skill.py my-api-helper --path skills/private
|
||||
init_skill.py custom-skill --path /custom/location
|
||||
"""
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
SKILL_TEMPLATE = """---
|
||||
name: {skill_name}
|
||||
description: [TODO: Complete and informative explanation of what the skill does and when to use it. Include WHEN to use this skill - specific scenarios, file types, or tasks that trigger it.]
|
||||
---
|
||||
|
||||
# {skill_title}
|
||||
|
||||
## Overview
|
||||
|
||||
[TODO: 1-2 sentences explaining what this skill enables]
|
||||
|
||||
## Structuring This Skill
|
||||
|
||||
[TODO: Choose the structure that best fits this skill's purpose. Common patterns:
|
||||
|
||||
**1. Workflow-Based** (best for sequential processes)
|
||||
- Works well when there are clear step-by-step procedures
|
||||
- Example: DOCX skill with "Workflow Decision Tree" → "Reading" → "Creating" → "Editing"
|
||||
- Structure: ## Overview → ## Workflow Decision Tree → ## Step 1 → ## Step 2...
|
||||
|
||||
**2. Task-Based** (best for tool collections)
|
||||
- Works well when the skill offers different operations/capabilities
|
||||
- Example: PDF skill with "Quick Start" → "Merge PDFs" → "Split PDFs" → "Extract Text"
|
||||
- Structure: ## Overview → ## Quick Start → ## Task Category 1 → ## Task Category 2...
|
||||
|
||||
**3. Reference/Guidelines** (best for standards or specifications)
|
||||
- Works well for brand guidelines, coding standards, or requirements
|
||||
- Example: Brand styling with "Brand Guidelines" → "Colors" → "Typography" → "Features"
|
||||
- Structure: ## Overview → ## Guidelines → ## Specifications → ## Usage...
|
||||
|
||||
**4. Capabilities-Based** (best for integrated systems)
|
||||
- Works well when the skill provides multiple interrelated features
|
||||
- Example: Product Management with "Core Capabilities" → numbered capability list
|
||||
- Structure: ## Overview → ## Core Capabilities → ### 1. Feature → ### 2. Feature...
|
||||
|
||||
Patterns can be mixed and matched as needed. Most skills combine patterns (e.g., start with task-based, add workflow for complex operations).
|
||||
|
||||
Delete this entire "Structuring This Skill" section when done - it's just guidance.]
|
||||
|
||||
## [TODO: Replace with the first main section based on chosen structure]
|
||||
|
||||
[TODO: Add content here. See examples in existing skills:
|
||||
- Code samples for technical skills
|
||||
- Decision trees for complex workflows
|
||||
- Concrete examples with realistic user requests
|
||||
- References to scripts/templates/references as needed]
|
||||
|
||||
## Resources
|
||||
|
||||
This skill includes example resource directories that demonstrate how to organize different types of bundled resources:
|
||||
|
||||
### scripts/
|
||||
Executable code (Python/Bash/etc.) that can be run directly to perform specific operations.
|
||||
|
||||
**Examples from other skills:**
|
||||
- PDF skill: `fill_fillable_fields.py`, `extract_form_field_info.py` - utilities for PDF manipulation
|
||||
- DOCX skill: `document.py`, `utilities.py` - Python modules for document processing
|
||||
|
||||
**Appropriate for:** Python scripts, shell scripts, or any executable code that performs automation, data processing, or specific operations.
|
||||
|
||||
**Note:** Scripts may be executed without loading into context, but can still be read by Claude for patching or environment adjustments.
|
||||
|
||||
### references/
|
||||
Documentation and reference material intended to be loaded into context to inform Claude's process and thinking.
|
||||
|
||||
**Examples from other skills:**
|
||||
- Product management: `communication.md`, `context_building.md` - detailed workflow guides
|
||||
- BigQuery: API reference documentation and query examples
|
||||
- Finance: Schema documentation, company policies
|
||||
|
||||
**Appropriate for:** In-depth documentation, API references, database schemas, comprehensive guides, or any detailed information that Claude should reference while working.
|
||||
|
||||
### assets/
|
||||
Files not intended to be loaded into context, but rather used within the output Claude produces.
|
||||
|
||||
**Examples from other skills:**
|
||||
- Brand styling: PowerPoint template files (.pptx), logo files
|
||||
- Frontend builder: HTML/React boilerplate project directories
|
||||
- Typography: Font files (.ttf, .woff2)
|
||||
|
||||
**Appropriate for:** Templates, boilerplate code, document templates, images, icons, fonts, or any files meant to be copied or used in the final output.
|
||||
|
||||
---
|
||||
|
||||
**Any unneeded directories can be deleted.** Not every skill requires all three types of resources.
|
||||
"""
|
||||
|
||||
EXAMPLE_SCRIPT = '''#!/usr/bin/env python3
|
||||
"""
|
||||
Example helper script for {skill_name}
|
||||
|
||||
This is a placeholder script that can be executed directly.
|
||||
Replace with actual implementation or delete if not needed.
|
||||
|
||||
Example real scripts from other skills:
|
||||
- pdf/scripts/fill_fillable_fields.py - Fills PDF form fields
|
||||
- pdf/scripts/convert_pdf_to_images.py - Converts PDF pages to images
|
||||
"""
|
||||
|
||||
def main():
|
||||
print("This is an example script for {skill_name}")
|
||||
# TODO: Add actual script logic here
|
||||
# This could be data processing, file conversion, API calls, etc.
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
'''
|
||||
|
||||
EXAMPLE_REFERENCE = """# Reference Documentation for {skill_title}
|
||||
|
||||
This is a placeholder for detailed reference documentation.
|
||||
Replace with actual reference content or delete if not needed.
|
||||
|
||||
Example real reference docs from other skills:
|
||||
- product-management/references/communication.md - Comprehensive guide for status updates
|
||||
- product-management/references/context_building.md - Deep-dive on gathering context
|
||||
- bigquery/references/ - API references and query examples
|
||||
|
||||
## When Reference Docs Are Useful
|
||||
|
||||
Reference docs are ideal for:
|
||||
- Comprehensive API documentation
|
||||
- Detailed workflow guides
|
||||
- Complex multi-step processes
|
||||
- Information too lengthy for main SKILL.md
|
||||
- Content that's only needed for specific use cases
|
||||
|
||||
## Structure Suggestions
|
||||
|
||||
### API Reference Example
|
||||
- Overview
|
||||
- Authentication
|
||||
- Endpoints with examples
|
||||
- Error codes
|
||||
- Rate limits
|
||||
|
||||
### Workflow Guide Example
|
||||
- Prerequisites
|
||||
- Step-by-step instructions
|
||||
- Common patterns
|
||||
- Troubleshooting
|
||||
- Best practices
|
||||
"""
|
||||
|
||||
EXAMPLE_ASSET = """# Example Asset File
|
||||
|
||||
This placeholder represents where asset files would be stored.
|
||||
Replace with actual asset files (templates, images, fonts, etc.) or delete if not needed.
|
||||
|
||||
Asset files are NOT intended to be loaded into context, but rather used within
|
||||
the output Claude produces.
|
||||
|
||||
Example asset files from other skills:
|
||||
- Brand guidelines: logo.png, slides_template.pptx
|
||||
- Frontend builder: hello-world/ directory with HTML/React boilerplate
|
||||
- Typography: custom-font.ttf, font-family.woff2
|
||||
- Data: sample_data.csv, test_dataset.json
|
||||
|
||||
## Common Asset Types
|
||||
|
||||
- Templates: .pptx, .docx, boilerplate directories
|
||||
- Images: .png, .jpg, .svg, .gif
|
||||
- Fonts: .ttf, .otf, .woff, .woff2
|
||||
- Boilerplate code: Project directories, starter files
|
||||
- Icons: .ico, .svg
|
||||
- Data files: .csv, .json, .xml, .yaml
|
||||
|
||||
Note: This is a text placeholder. Actual assets can be any file type.
|
||||
"""
|
||||
|
||||
|
||||
def title_case_skill_name(skill_name):
|
||||
"""Convert hyphenated skill name to Title Case for display."""
|
||||
return ' '.join(word.capitalize() for word in skill_name.split('-'))
|
||||
|
||||
|
||||
def init_skill(skill_name, path):
|
||||
"""
|
||||
Initialize a new skill directory with template SKILL.md.
|
||||
|
||||
Args:
|
||||
skill_name: Name of the skill
|
||||
path: Path where the skill directory should be created
|
||||
|
||||
Returns:
|
||||
Path to created skill directory, or None if error
|
||||
"""
|
||||
# Determine skill directory path
|
||||
skill_dir = Path(path).resolve() / skill_name
|
||||
|
||||
# Check if directory already exists
|
||||
if skill_dir.exists():
|
||||
print(f"❌ Error: Skill directory already exists: {skill_dir}")
|
||||
return None
|
||||
|
||||
# Create skill directory
|
||||
try:
|
||||
skill_dir.mkdir(parents=True, exist_ok=False)
|
||||
print(f"✅ Created skill directory: {skill_dir}")
|
||||
except Exception as e:
|
||||
print(f"❌ Error creating directory: {e}")
|
||||
return None
|
||||
|
||||
# Create SKILL.md from template
|
||||
skill_title = title_case_skill_name(skill_name)
|
||||
skill_content = SKILL_TEMPLATE.format(
|
||||
skill_name=skill_name,
|
||||
skill_title=skill_title
|
||||
)
|
||||
|
||||
skill_md_path = skill_dir / 'SKILL.md'
|
||||
try:
|
||||
skill_md_path.write_text(skill_content)
|
||||
print("✅ Created SKILL.md")
|
||||
except Exception as e:
|
||||
print(f"❌ Error creating SKILL.md: {e}")
|
||||
return None
|
||||
|
||||
# Create resource directories with example files
|
||||
try:
|
||||
# Create scripts/ directory with example script
|
||||
scripts_dir = skill_dir / 'scripts'
|
||||
scripts_dir.mkdir(exist_ok=True)
|
||||
example_script = scripts_dir / 'example.py'
|
||||
example_script.write_text(EXAMPLE_SCRIPT.format(skill_name=skill_name))
|
||||
example_script.chmod(0o755)
|
||||
print("✅ Created scripts/example.py")
|
||||
|
||||
# Create references/ directory with example reference doc
|
||||
references_dir = skill_dir / 'references'
|
||||
references_dir.mkdir(exist_ok=True)
|
||||
example_reference = references_dir / 'api_reference.md'
|
||||
example_reference.write_text(EXAMPLE_REFERENCE.format(skill_title=skill_title))
|
||||
print("✅ Created references/api_reference.md")
|
||||
|
||||
# Create assets/ directory with example asset placeholder
|
||||
assets_dir = skill_dir / 'assets'
|
||||
assets_dir.mkdir(exist_ok=True)
|
||||
example_asset = assets_dir / 'example_asset.txt'
|
||||
example_asset.write_text(EXAMPLE_ASSET)
|
||||
print("✅ Created assets/example_asset.txt")
|
||||
except Exception as e:
|
||||
print(f"❌ Error creating resource directories: {e}")
|
||||
return None
|
||||
|
||||
# Print next steps
|
||||
print(f"\n✅ Skill '{skill_name}' initialized successfully at {skill_dir}")
|
||||
print("\nNext steps:")
|
||||
print("1. Edit SKILL.md to complete the TODO items and update the description")
|
||||
print("2. Customize or delete the example files in scripts/, references/, and assets/")
|
||||
print("3. Run the validator when ready to check the skill structure")
|
||||
|
||||
return skill_dir
|
||||
|
||||
|
||||
def main():
|
||||
if len(sys.argv) < 4 or sys.argv[2] != '--path':
|
||||
print("Usage: init_skill.py <skill-name> --path <path>")
|
||||
print("\nSkill name requirements:")
|
||||
print(" - Hyphen-case identifier (e.g., 'data-analyzer')")
|
||||
print(" - Lowercase letters, digits, and hyphens only")
|
||||
print(" - Max 40 characters")
|
||||
print(" - Must match directory name exactly")
|
||||
print("\nExamples:")
|
||||
print(" init_skill.py my-new-skill --path skills/public")
|
||||
print(" init_skill.py my-api-helper --path skills/private")
|
||||
print(" init_skill.py custom-skill --path /custom/location")
|
||||
sys.exit(1)
|
||||
|
||||
skill_name = sys.argv[1]
|
||||
path = sys.argv[3]
|
||||
|
||||
print(f"🚀 Initializing skill: {skill_name}")
|
||||
print(f" Location: {path}")
|
||||
print()
|
||||
|
||||
result = init_skill(skill_name, path)
|
||||
|
||||
if result:
|
||||
sys.exit(0)
|
||||
else:
|
||||
sys.exit(1)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,110 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
Skill Packager - Creates a distributable .skill file of a skill folder
|
||||
|
||||
Usage:
|
||||
python utils/package_skill.py <path/to/skill-folder> [output-directory]
|
||||
|
||||
Example:
|
||||
python utils/package_skill.py skills/public/my-skill
|
||||
python utils/package_skill.py skills/public/my-skill ./dist
|
||||
"""
|
||||
|
||||
import sys
|
||||
import zipfile
|
||||
from pathlib import Path
|
||||
from quick_validate import validate_skill
|
||||
|
||||
|
||||
def package_skill(skill_path, output_dir=None):
|
||||
"""
|
||||
Package a skill folder into a .skill file.
|
||||
|
||||
Args:
|
||||
skill_path: Path to the skill folder
|
||||
output_dir: Optional output directory for the .skill file (defaults to current directory)
|
||||
|
||||
Returns:
|
||||
Path to the created .skill file, or None if error
|
||||
"""
|
||||
skill_path = Path(skill_path).resolve()
|
||||
|
||||
# Validate skill folder exists
|
||||
if not skill_path.exists():
|
||||
print(f"❌ Error: Skill folder not found: {skill_path}")
|
||||
return None
|
||||
|
||||
if not skill_path.is_dir():
|
||||
print(f"❌ Error: Path is not a directory: {skill_path}")
|
||||
return None
|
||||
|
||||
# Validate SKILL.md exists
|
||||
skill_md = skill_path / "SKILL.md"
|
||||
if not skill_md.exists():
|
||||
print(f"❌ Error: SKILL.md not found in {skill_path}")
|
||||
return None
|
||||
|
||||
# Run validation before packaging
|
||||
print("🔍 Validating skill...")
|
||||
valid, message = validate_skill(skill_path)
|
||||
if not valid:
|
||||
print(f"❌ Validation failed: {message}")
|
||||
print(" Please fix the validation errors before packaging.")
|
||||
return None
|
||||
print(f"✅ {message}\n")
|
||||
|
||||
# Determine output location
|
||||
skill_name = skill_path.name
|
||||
if output_dir:
|
||||
output_path = Path(output_dir).resolve()
|
||||
output_path.mkdir(parents=True, exist_ok=True)
|
||||
else:
|
||||
output_path = Path.cwd()
|
||||
|
||||
skill_filename = output_path / f"{skill_name}.skill"
|
||||
|
||||
# Create the .skill file (zip format)
|
||||
try:
|
||||
with zipfile.ZipFile(skill_filename, 'w', zipfile.ZIP_DEFLATED) as zipf:
|
||||
# Walk through the skill directory
|
||||
for file_path in skill_path.rglob('*'):
|
||||
if file_path.is_file():
|
||||
# Calculate the relative path within the zip
|
||||
arcname = file_path.relative_to(skill_path.parent)
|
||||
zipf.write(file_path, arcname)
|
||||
print(f" Added: {arcname}")
|
||||
|
||||
print(f"\n✅ Successfully packaged skill to: {skill_filename}")
|
||||
return skill_filename
|
||||
|
||||
except Exception as e:
|
||||
print(f"❌ Error creating .skill file: {e}")
|
||||
return None
|
||||
|
||||
|
||||
def main():
|
||||
if len(sys.argv) < 2:
|
||||
print("Usage: python utils/package_skill.py <path/to/skill-folder> [output-directory]")
|
||||
print("\nExample:")
|
||||
print(" python utils/package_skill.py skills/public/my-skill")
|
||||
print(" python utils/package_skill.py skills/public/my-skill ./dist")
|
||||
sys.exit(1)
|
||||
|
||||
skill_path = sys.argv[1]
|
||||
output_dir = sys.argv[2] if len(sys.argv) > 2 else None
|
||||
|
||||
print(f"📦 Packaging skill: {skill_path}")
|
||||
if output_dir:
|
||||
print(f" Output directory: {output_dir}")
|
||||
print()
|
||||
|
||||
result = package_skill(skill_path, output_dir)
|
||||
|
||||
if result:
|
||||
sys.exit(0)
|
||||
else:
|
||||
sys.exit(1)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,95 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
Quick validation script for skills - minimal version
|
||||
"""
|
||||
|
||||
import sys
|
||||
import os
|
||||
import re
|
||||
import yaml
|
||||
from pathlib import Path
|
||||
|
||||
def validate_skill(skill_path):
|
||||
"""Basic validation of a skill"""
|
||||
skill_path = Path(skill_path)
|
||||
|
||||
# Check SKILL.md exists
|
||||
skill_md = skill_path / 'SKILL.md'
|
||||
if not skill_md.exists():
|
||||
return False, "SKILL.md not found"
|
||||
|
||||
# Read and validate frontmatter
|
||||
content = skill_md.read_text()
|
||||
if not content.startswith('---'):
|
||||
return False, "No YAML frontmatter found"
|
||||
|
||||
# Extract frontmatter
|
||||
match = re.match(r'^---\n(.*?)\n---', content, re.DOTALL)
|
||||
if not match:
|
||||
return False, "Invalid frontmatter format"
|
||||
|
||||
frontmatter_text = match.group(1)
|
||||
|
||||
# Parse YAML frontmatter
|
||||
try:
|
||||
frontmatter = yaml.safe_load(frontmatter_text)
|
||||
if not isinstance(frontmatter, dict):
|
||||
return False, "Frontmatter must be a YAML dictionary"
|
||||
except yaml.YAMLError as e:
|
||||
return False, f"Invalid YAML in frontmatter: {e}"
|
||||
|
||||
# Define allowed properties
|
||||
ALLOWED_PROPERTIES = {'name', 'description', 'license', 'allowed-tools', 'metadata'}
|
||||
|
||||
# Check for unexpected properties (excluding nested keys under metadata)
|
||||
unexpected_keys = set(frontmatter.keys()) - ALLOWED_PROPERTIES
|
||||
if unexpected_keys:
|
||||
return False, (
|
||||
f"Unexpected key(s) in SKILL.md frontmatter: {', '.join(sorted(unexpected_keys))}. "
|
||||
f"Allowed properties are: {', '.join(sorted(ALLOWED_PROPERTIES))}"
|
||||
)
|
||||
|
||||
# Check required fields
|
||||
if 'name' not in frontmatter:
|
||||
return False, "Missing 'name' in frontmatter"
|
||||
if 'description' not in frontmatter:
|
||||
return False, "Missing 'description' in frontmatter"
|
||||
|
||||
# Extract name for validation
|
||||
name = frontmatter.get('name', '')
|
||||
if not isinstance(name, str):
|
||||
return False, f"Name must be a string, got {type(name).__name__}"
|
||||
name = name.strip()
|
||||
if name:
|
||||
# Check naming convention (hyphen-case: lowercase with hyphens)
|
||||
if not re.match(r'^[a-z0-9-]+$', name):
|
||||
return False, f"Name '{name}' should be hyphen-case (lowercase letters, digits, and hyphens only)"
|
||||
if name.startswith('-') or name.endswith('-') or '--' in name:
|
||||
return False, f"Name '{name}' cannot start/end with hyphen or contain consecutive hyphens"
|
||||
# Check name length (max 64 characters per spec)
|
||||
if len(name) > 64:
|
||||
return False, f"Name is too long ({len(name)} characters). Maximum is 64 characters."
|
||||
|
||||
# Extract and validate description
|
||||
description = frontmatter.get('description', '')
|
||||
if not isinstance(description, str):
|
||||
return False, f"Description must be a string, got {type(description).__name__}"
|
||||
description = description.strip()
|
||||
if description:
|
||||
# Check for angle brackets
|
||||
if '<' in description or '>' in description:
|
||||
return False, "Description cannot contain angle brackets (< or >)"
|
||||
# Check description length (max 1024 characters per spec)
|
||||
if len(description) > 1024:
|
||||
return False, f"Description is too long ({len(description)} characters). Maximum is 1024 characters."
|
||||
|
||||
return True, "Skill is valid!"
|
||||
|
||||
if __name__ == "__main__":
|
||||
if len(sys.argv) != 2:
|
||||
print("Usage: python quick_validate.py <skill_directory>")
|
||||
sys.exit(1)
|
||||
|
||||
valid, message = validate_skill(sys.argv[1])
|
||||
print(message)
|
||||
sys.exit(0 if valid else 1)
|
||||
@@ -0,0 +1,21 @@
|
||||
---
|
||||
name: skill-sync
|
||||
description: Synchronize skills from the current project to the global Claude skills directory (~/.claude/skills). Use this when you have modified skills in the project and want to update the global installation. Supports Windows 11.
|
||||
---
|
||||
|
||||
# Skill Sync
|
||||
|
||||
Synchronize skills from the project's `skills/` directory to the global Claude environment.
|
||||
|
||||
## Workflow
|
||||
|
||||
### 1. Synchronize All Skills
|
||||
When changes are made to skills within this project, run the synchronization script to update the global installation.
|
||||
|
||||
- **Operation**: Run `powershell -ExecutionPolicy Bypass -File skills/skill-sync/scripts/sync.ps1`
|
||||
- **Effect**: All skill directories in `skills/` will be copied to `~\.claude\skills`. Existing skills in the target directory will be updated, while additional skills in the target that are not in the project remain untouched.
|
||||
|
||||
## Resources
|
||||
|
||||
### scripts/
|
||||
- `sync.ps1`: PowerShell script for synchronizing skill directories to the global Claude skills folder.
|
||||
@@ -0,0 +1,28 @@
|
||||
# Get the current skill directory (the parent of the scripts folder)
|
||||
$ScriptDir = Split-Path -Parent $MyInvocation.MyCommand.Path
|
||||
$SkillDir = Split-Path -Parent $ScriptDir
|
||||
$ProjectSkillsDir = Split-Path -Parent $SkillDir
|
||||
$GlobalSkillsDir = Join-Path $HOME ".claude\skills"
|
||||
|
||||
Write-Host "Syncing skills from $ProjectSkillsDir to $GlobalSkillsDir..." -ForegroundColor Cyan
|
||||
|
||||
if (-not (Test-Path $GlobalSkillsDir)) {
|
||||
Write-Host "Global skills directory does not exist. Creating it..." -ForegroundColor Yellow
|
||||
New-Item -ItemType Directory -Path $GlobalSkillsDir -Force
|
||||
}
|
||||
|
||||
# Get all skill directories in the project
|
||||
$Skills = Get-ChildItem -Path $ProjectSkillsDir -Directory
|
||||
|
||||
foreach ($Skill in $Skills) {
|
||||
$SkillName = $Skill.Name
|
||||
$TargetDir = Join-Path $GlobalSkillsDir $SkillName
|
||||
|
||||
Write-Host "Syncing skill: $SkillName..." -ForegroundColor Green
|
||||
|
||||
# Sync the skill directory
|
||||
# Note: Copy-Item -Recurse -Force will overwrite existing files
|
||||
Copy-Item -Path $Skill.FullName -Destination $GlobalSkillsDir -Recurse -Force
|
||||
}
|
||||
|
||||
Write-Host "Sync completed successfully." -ForegroundColor Green
|
||||
Reference in New Issue
Block a user