Files
agent-skills/skills/g3fo-bug-finder/SKILL.md
T
ken.li b88e934f86 Add Agent Skills documentation and implement g3fo-bug-finder skill
- Created README.md to introduce the Agent Skills repository and its functionalities.
- Added g3fo-bug-finder skill with detailed SKILL.md for automated bug diagnosis on remote servers.
- Included service inventory reference for mapping services to server configurations.
- Implemented scripts for fetching logs and configurations from remote servers using SSH and Nacos API.
- Provided usage instructions and environment setup for Python dependencies.
2026-01-16 17:56:29 +08:00

75 lines
4.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
name: g3fo-bug-finder
description: 自动登录远程 Ubuntu 服务器,提取 g3fo 项目日志,并结合本地代码进行 Bug 诊断和修复建议。支持多节点日志提取和 Nacos 配置分析。
---
# g3fo Bug Finder (g3fo 故障排查专家)
本 Skill 用于自动化排查部署在远程服务器上的 g3fo 项目应用问题。它能够连接服务器、抓取日志、解析堆栈、定位源码并提供修复方案。
## 触发场景
- 用户要求排查特定服务的错误(如:“看看 order-service 为什么报错”)。
- 用户提到特定的错误码(如:“错误码 5002 是怎么回事?”)。
- 用户需要分析远程服务器上的 Java 异常堆栈。
## 工作流
### 1. 识别服务与服务器
当收到指令后,首先读取 `references/service_inventory.md`:
- **确认所有节点**: 确定服务部署的所有服务器 IP(如 192.168.3.200 和 192.168.3.230)。**Agent 必须对所有记录的节点执行日志提取**。
- 如果用户提供了**服务名**(如 gateway),查找对应行。
- 如果用户提供了**错误码**(如 12005),根据 `错误码起始` 判定所属服务(12000-12999 为 user)。
- 确认目标 `Docker 服务名`(如 g3fo-gateway-service)或 `日志文件路径`(如 /data/logs/g3fo-user/info/info.log)。
- 确认目标 `Server IP`、`Log Method` 以及相关的 `Path/YML`。
### 2. 获取 SSH 凭据
- 检查环境中是否存在凭据(如环境变量)。
- 如果没有,以交互方式询问用户:**SSH 用户名** 和 **密码**(或提醒用户配置私钥)。
- **注意**:不要在对话中存储密码,仅用于当前会话运行脚本。
### 3. 提取远程日志
使用 Python 脚本 `scripts/fetch_logs.py` 抓取日志。
- **全节点提取**: 由于服务部署在多台服务器上(192.168.3.200, 192.168.3.230),**必须同时从所有相关服务器提取日志**,以确保不遗漏错误信息。
- **日志级别选择**:
- 默认抓取 `info` 级别。
- 如果涉及报错排查,优先抓取 `error` 级别。
- **路径拼装**: 严格遵循 `/data/logs/g3fo-{service}/{level}/{level}.log`。
命令模版示例 (需对两台机器分别运行):
```bash
# 对 Server A 运行
python scripts/fetch_logs.py --host 192.168.3.200 --username root --password afe1234 --mode docker --service [SERVICE] --yml-path [YML]
# 对 Server B 运行
python scripts/fetch_logs.py --host 192.168.3.230 --username root --password afe1234 --mode docker --service [SERVICE] --yml-path [YML]
```
### 4. 异常分析与代码定位
- **日志解析**:从获取的日志中提取 `Exception`、`Error` 或 `Caused by` 附近的堆栈信息(Stack Trace)。
- **网络问题判断**:如果日志分析结论为网络连通性问题(如 `Connection refused`, `ConnectTimeoutException`, `UnknownHostException` 等):
- **直接说明**:在报告中直接说明是网络连通性问题。
- **连通性测试**:使用 `nc -zv {IP} {Port}` 或 `telnet {IP} {Port}` 到目标服务器进行测试。
- **无需深度分析**:这种情况下不需要进一步分析源码或 Nacos 配置。
- **源码检索**:提取堆栈中的全限定类名(如 `com.example.service.OrderService`)和行号。
- **配置获取**:如需分析配置(如数据库地址、中间件端口等),使用 `scripts/fetch_configs.py`。
- 示例:`python scripts/fetch_configs.py --data-id redis.yml`
- **本地比对**:使用 `read_file` 或 `codebase_search` 查看本地对应的源码逻辑。
### 5. 输出诊断报告
报告应包含:
- **错误类型**:Java 异常类名。
- **根本原因**:根据日志和代码逻辑分析得出的结论(如果是网络问题,请务必指出)。
- **连通性测试结果**(如适用):展示服务器间网络测试的输出。
- **关联代码**:引用本地源码的相关片段。
- **修复方案**:具体的代码修改建议或配置调整建议。
## 依赖要求
- 本地 Python 环境。
- 安装 `paramiko` 库:`pip install paramiko`。
- 目标服务器支持 SSH 登录。
## 注意事项
- 如果日志文件过大,建议优先使用脚本自带的 `grep` 过滤功能。
- 始终确保本地代码分支与服务器部署版本尽量一致。