Files
RAG-CUT/README.md
T
2026-07-16 11:12:17 +08:00

508 lines
21 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# RAG-cut
自研 AI 助手知识库 Demo —— 以**高质量文档切片**为核心,召回为次要目标。
重建文档 ingestion / 切片流水线,解决现有 AFE 知识库的两类问题:
1. **文档解析不足**:Word / PDF / Excel 中的图片、表格无法有效提取
2. **切片规则不佳**:同一逻辑段落被硬切到不同 chunk,图片丢失或错位
需求来源:`[docx/自研搭建AI助手知识库.pdf](docx/自研搭建AI助手知识库.pdf)`
---
## 功能概览
| 能力 | 说明 |
| --------- | --------------------------------------------------------------------------------------------- |
| 多格式解析 | pdf / doc / docx / ppt / pptx / ppsx / xlsx / csv / md / txt / html / json / xml / log / 常见图片 |
| PDF 双引擎 | MinerU(优先)+ PyMuPDF 回退;统一噪声过滤 |
| 自动切分策略 | 前端可选自动,或手动指定 mode |
| 语义切分 | PDF / Word 走 `heading_layout_multimodal`;超长章节仅按长度/子标题拆分(不再产父片) |
| 父子标识符 | 独立模式 `parent_child`:父/子标识符双层切分,子片检索、父片上下文 |
| 表格切分 | 自动识别表头 / 说明行;Q&A 表按行 1 片;普通表按行数自适应 |
| Demo UI | 原文预览 · 切片列表 · Markdown 预览 · **历史切片回看** · BM25 召回测试 |
| CLI / API | `scripts/chunk_cli.py` · FastAPI `/api/chunk` · `/api/results` · `/api/recall` |
未支持:`.wps`(可先转为 docx/pdf)。独立图片暂不做 OCR。
---
## 技术架构
```
上传文件
│
▼
格式路由 (registry)
│
▼
解析器 (parsers) ──► 有序 Block 流(heading / paragraph / table / image …)
│ PDF:MinerU(优先)或 PyMuPDF → 噪声过滤 → 跨页表合并
▼
版面增强 (layout_meta) ──► 标题层级、图片/表格上下文、order_index
│
▼
切分策略 (split_policy / pdf_strategy)
│
▼
切分器 (splitters) ──► Block 分组(含父子切片关联)
│
▼
渲染器 (renderer) ──► Markdown Chunk + embedding_text / retrieval 标记
```
**核心设计**:先解析为不可拆分的 Block,再按结构切分。`image` / `table` 为原子块,不会被拆到不同 chunk。
### 目录结构
```
RAG-cut/
├── backend/
│ ├── rag_cut/
│ │ ├── models.py / pipeline.py / renderer.py / layout_meta.py
│ │ ├── split_policy.py / retrieval.py
│ │ ├── parsers/ # word / ppt / pdf / xlsx / text / image + MinerU
│ │ │ └── pdf/ # PyMuPDF 流水线、噪声过滤、跨页表合并
│ │ └── splitters/ # default / heading / pdf_semantic / by_row / delimiter
│ ├── api/main.py # FastAPI:/api/chunk · /api/results · /api/recall · 静态前端
│ ├── tests/
│ └── run.py
├── frontend/ # Demo(index.html + css + js)
├── scripts/ # chunk_cli · run_all_samples
├── storage/ # uploads / assets / results(运行时生成)
├── docx/ # 需求 PDF、MinerU 说明
└── requirements.txt
```
---
## 快速开始
```bash
pip install -r requirements.txt
cd backend
python run.py
```
浏览器打开:**[http://127.0.0.1:8000](http://127.0.0.1:8000)**
前后端分离开发时,另开终端:
```bash
cd frontend
python run.py # http://127.0.0.1:5173 ,API 需自行指向 :8000
```
可选参数:`--host 0.0.0.0 --port 8000 --no-reload`(后端)· `--port 5173`(前端)
### 环境变量(PDF)
| 变量 | 默认 | 说明 |
| -------------------- | ------ | ------------------------------------------------------ |
| `RAG_CUT_PDF_ENGINE` | `auto` | `auto` 优先 MinerU,失败回退 PyMuPDF;可强制 `mineru` / `pymupdf` |
| `RAG_CUT_MINERU_CMD` | — | MinerU 可执行文件路径(不在 PATH 时) |
| `RAG_CUT_MINERU_API_URL` | — | 复用已启动的 MinerU API,避免每次切分重复加载模型 |
| `RAG_CUT_MINERU_TIMEOUT` | `540` | 单次 MinerU 解析超时秒数;超时后会清理其临时服务和子进程 |
| `RAG_CUT_MINERU_OCR` | `1` | MinerU 图片无 OCR 时用 Tesseract 回填;`0` 关闭以加快切分 |
详见 `[docx/mineru-integration.md](docx/mineru-integration.md)`。
### 依赖与运行环境
| 类别 | 选型 |
| ------------ | -------------------------------------------------------------- |
| Python | 3.12+ |
| Web | FastAPI + Uvicorn |
| PDF | MinerU(`requirements.txt` 已含)/ PyMuPDF + pdfplumber |
| Excel | openpyxl |
| Office 转 PDF | 本机 **LibreOffice** 或 **Microsoft Word/PowerPoint**(doc/ppt 必需) |
| 召回 | 内置 BM25 词法评分(无向量库) |
---
## 各格式处理
所有格式统一:**解析 → Block → 切分 → Markdown Chunk**。
| 扩展名 | 解析器 | 要点 | 默认切分 |
| ---------------------- | ------------------- | ------------------------------- | ------------------------------ |
| `.pdf` | `PdfParser` | MinerU / PyMuPDF → 噪声过滤 → 跨页表合并 | 语义切分(`max=2600`) |
| `.doc` `.docx` | `WordParser` | 转 PDF 后复用 PDF 流水线 | 同上 |
| `.ppt` `.pptx` `.ppsx` | `PptParser` | 转 PDF;页码 ≈ 幻灯片序号 | `default`(`max=2600`) |
| `.xlsx` `.xls` `.csv` | `SpreadsheetParser` | 自动检测 preamble / 表头 / 数据起始行 | `by_row`(Q&A 表 1 行/片,否则按行数自适应) |
| `.md` `.html` | `TextParser` | 按标题 / 段落拆块 | `default`(`max=1800`) |
| `.json` | `TextParser` | 顶层 key / 数组元素 → `code` 块 | `default` |
| `.txt` `.xml` `.log` | `TextParser` | 按空行分段(不解析 XML 结构) | `default` |
| 图片 | `ImageParser` | 整图 1 个 Block | 整图 1 片 |
### 统一产出
| 项目 | 说明 |
| -------- | ---------------------------------------------------------------------------------------------- |
| Block 类型 | `heading` / `paragraph` / `table` / `image` / `list` / `code` |
| Chunk | Markdown + `meta`(`heading` / `pages` / `embedding_text` / `retrieval` / 可选 `parent_chunk_id`) |
| 图片 | `storage/assets/{doc_id}/`;JSON 内为 `assets/{doc_id}/…` |
| 结果缓存 | `storage/results/{doc_id}.json`(供召回) |
| 中间产物 | `_conversion/`(Office→PDF)、`_mineru/`(MinerU 输出) |
### Word / PPT 转 PDF
优先 LibreOffice(`soffice --headless`),Windows 可回退 Office COM。转换失败时 API 返回 500 并附带各后端错误信息。
### Excel / CSV 布局检测
`detect_spreadsheet_layout` 会识别常见模板(如第 1 行字段说明、第 2 行列表头、第 3 行起数据),跳过 preamble;Q&A 评测表(列名含 query/answer 等)自动 `rows_per_chunk=1`。
> `.xls` 依赖 openpyxl,老二进制格式可能失败,建议先转 xlsx。
---
## PDF 切割方法
`.pdf` / `.doc` / `.docx` 在 `mode=default` 时走专用语义切分器
`splitters/pdf_semantic.py`(策略名:`heading_layout_multimodal`)。
Word 先转 PDF,再与 PDF 共用同一套规则。默认 `max_chunk_size=2600`、`overlap=120`。
### 流水线
```
PDF/Word
│
├─ 解析:MinerU(优先)或 PyMuPDF → Block 流
├─ 噪声过滤:页眉/页脚/页码、**整段文档目录**、装饰小图
├─ 跨页表格合并
├─ 版面增强:标题层级、图片/表格上下文、order_index
├─ pdf_strategy:打标签(操作手册 vs 年报,见下)
└─ heading_layout_multimodal 切分 → Chunk
```
### 核心规则:`heading_layout_multimodal`
目标:同一逻辑小节(标题 + 正文 + 同节截图/表格)尽量落在同一 chunk,避免把步骤截图和说明文字拆开。
| 步骤 | 行为 |
|------|------|
| 1. 再过滤噪声 | 跳过页码、running header、**文档目录(目录/Contents 及条目)**、过小装饰图 |
| 2. 识别标题 | 解析器已标的 `heading`,或正则/字号启发式补识别 |
| 3. 维护标题路径 | 用标题栈记录章节层级,如 `1 Getting Started → A. Log in` |
| 4. 按标题开新片 | 遇到同级或更高级标题时 flush 当前组,开新 chunk |
| 5. 图文同节 | 正文、image、table 跟在当前标题路径下;绑定前后文到 `meta` |
| 6. 超长二次切 | 超过 `max_chunk_size` 时按子标题/内容标记/长度拆成多片;**不单独保留父片** |
**可识别的标题形态**(示例):
- 编号:`1.` / `1.2.3` / `A.` / `Chapter 2`(操作步骤 `Step N …` 作为正文保留,不单独开章)
- 中文章节:`一、` / `(一)` 等
- 字号显著大于正文的短文本(PyMuPDF 路径)
封面 Logo 等无语义组会标 `retrieval=false`,不进召回候选。
### 父子标识符切分(`mode=parent_child`)
适合「细粒度检索、粗粒度召回」:子片用于 BM25 检索,命中后通过 `parent_chunk_id` 关联父片上下文。
> 默认切分**不会**再因超长章节自动产出父片+子片;需要父子结构时请显式选择本模式。
流程:
1. 按 **父级标识符**(`parent_delimiter`)切开,得到父片;超过父级最大长度时再按长度二次切
2. 每个父片再按 **子级标识符**(`child_delimiter`,可选)切开,得到子片;超过子级最大长度时再按长度二次切
3. 父片:`is_section_parent=true`,`retrieval=false`
4. 子片:`is_sub_chunk=true`,`retrieval=true`,`parent_chunk_id` 指向父片
参数约束:
| 参数 | 说明 |
|------|------|
| `parent_delimiter` | 必填;父级切开标记(不会写入切片正文) |
| `child_delimiter` | 可选;缺省则仅按子级最大长度拆子片 |
| `max_chunk_size` | 父级最大长度 |
| `child_max_size` | 子级最大长度,且 ≤ 父级、硬上限 1500 |
| `overlap` | 超长二次切分时的重叠 |
Demo / CLI / API 均可显式选择此模式。
### 文档类型标签:`pdf_strategy`
`splitters/pdf_strategy.py` 根据文件名与正文特征打标签,写入
`split_config.pdf_chunk_strategy`(Demo 自动策略面板会显示):
| 标签 | 典型特征 | 说明 |
|------|----------|------|
| `pdf_feature_step_screenshot` | 步骤标题多、「点击/输入/选择」等操作词多、截图密度高 | 操作手册型 |
| `pdf_outline_report` | 「年报/财务报表/董事会」等词多,或文件名含 annual/report/年报 | 年报/报告型 |
> 当前两类**切分算法相同**(都走 `heading_layout_multimodal`);标签用于结果标注,并为后续分叉策略预留。
### 其他切分模式(PDF 也可用)
| 模式 | 何时用 | 行为 |
|------|--------|------|
| `default` | Demo / 自动策略(推荐) | 上文语义切分 |
| `delimiter` | CLI/API/前端显式指定 | 按自定义标识符切,再按长度二次切 |
| `parent_child` | CLI/API/前端显式指定 | 父子标识符切分(检索用子片、父片作上下文) |
| `by_row` | 一般不用于 PDF | 面向表格文档 |
相关代码:`pdf_semantic.py` · `pdf_strategy.py` · `parsers/pdf/noise_filter.py` · `pipeline.py`。
---
## 切分策略(全格式)
后端共 **4** 种切分模式(`SplitMode`):`default` / `delimiter` / `parent_child` / `by_row`。
前端「自动」(`auto`) 不是独立模式:不传 `mode` 时由 `split_policy.choose_split_config` 按扩展名与表结构选型。
API **不会按扩展名拦截 mode**——凡解析器支持的格式均可手动指定任意模式;下表区分设计意图与有效行为。
### 可解析格式(21 种)
| 类别 | 扩展名 |
| ---- | ------ |
| Office | `.pdf` `.doc` `.docx` `.ppt` `.pptx` `.ppsx` |
| 表格 | `.xlsx` `.xls` `.csv` |
| 文本 | `.md` `.txt` `.html` `.htm` `.json` `.xml` `.log` |
| 图片 | `.jpg` `.jpeg` `.png` `.bmp` `.gif` |
> 需求中的 `.wps` 尚未接入解析器。
### 四种模式与文档格式
| 模式 | 适用(设计意图) | 实际覆盖的格式 | 行为 |
| ---- | ---------------- | -------------- | ---- |
| `default` | 非表格 | 除表格类外全部;表格也可手动指定 | 结构/标题感知切分;`.pdf`/`.doc`/`.docx` 走语义切分(`heading_layout_multimodal`);单章超长时按子标题或长度拆开(**不产父片**) |
| `delimiter` | 非表格(正文含可匹配标识符) | 全部可解析格式 | 按自定义标识符(如 `###`)切开;超长片段再按 `max_chunk_size` + `overlap`;无标识符时效果差 |
| `parent_child` | 非表格(层级标识符) | 全部可解析格式 | 先按父标识符切父片,再按子标识符切子片;检索用子片,父片作上下文;典型如 `.md`/文本 |
| `by_row` | 表格 | **有效**:`.xlsx` `.xls` `.csv`,或解析后含带 `rows` 的 TABLE block | 每片 = 表头 + N 行 Markdown;无表格行数据时退化为整篇一片;一般不用于 PDF/Word |
### 自动策略(不传 mode)
| 文档 | 自动选的模式 |
| ---- | ------------ |
| `.xlsx` `.xls` `.csv`,或整篇就一张表(含 `rows`) | `by_row` |
| `.pdf` `.ppt` `.pptx` `.ppsx` | `default` |
| 图片类 | `default` |
| 文本类(`.md` `.txt` `.html` `.htm` `.json` `.xml` `.log`) | `default` |
| `.doc` `.docx` 等其余 | `default` |
**一句话**:`by_row` 专吃表格;另外三种面向正文结构,其中 `default` 对 PDF/Word 有专用路径,`delimiter` / `parent_child` 对有标识符的文本最有用,但格式本身不限制。
### `default` 决策树
```
.pdf / .doc / .docx
└─ 见上方「PDF 切割方法」(heading_layout_multimodal)
单 table 且含 rows
└─ 自动 by_row
其他 + 足够标题
└─ 标题大纲切;若某章超过 max_chunk_size
→ 优先按子标题 / 内容标记拆分
→ 否则按长度拆分(不保留整章父片)
有 page 元数据、标题不足
└─ 按页;单页超限再按长度
否则
└─ 按 max_chunk_size 打包;image/table 不拆
```
`meta.retrieval = false` 的切片(如封面 Logo,或 `parent_child` 模式的父片)不参与召回。
### SplitConfig
| 参数 | 默认 | 说明 |
| ------------------------------------- | --------- | ------------------------------ |
| `mode` | `default` | `default` / `delimiter` / `parent_child` / `by_row` |
| `delimiter` | — | `delimiter` 模式必填 |
| `parent_delimiter` / `child_delimiter` | — | `parent_child` 模式:父标识符必填,子标识符可选 |
| `max_chunk_size` | 1500 | 单 chunk 上限(父级长度;自动策略常覆盖为 1800–2600) |
| `child_max_size` | 512 | `parent_child` 子级最大长度(≤ 父级,且 ≤ 1500) |
| `overlap` | 150 | 超长二次切分重叠 |
| `header_row_start` / `header_row_end` | 1 | 表头行(1-based) |
| `start_row` | 2 | 数据起始行 |
| `rows_per_chunk` | 1 | 每片数据行数 |
---
## Demo 前端
三栏工作区:**原文** · **切片列表** · **切片预览**;左侧可选手动切分模式,或保持「自动」由系统选型。
| 区域 | 说明 |
| ------ | ---------------------------------------- |
| 上传 | 拖拽 / 选择文件 |
| 切分模式 | 自动 / 默认(结构感知)/ 通用标识符 / **父子标识符** / 按行;手动模式可调对应参数 |
| 策略面板 | 切分后回显 mode / max_size / overlap / PDF 策略 |
| 原文 | PDF iframe / docx(mammoth)/ 文本 / 图片 |
| 切片列表 | 类型标签、父子切片标记、字符数 |
| 召回测试 | query + top_k;展示得分与命中切片 |
| API 面板 | 可折叠请求/响应 JSON |
---
## CLI / API
### CLI
```bash
# 显式传 SplitConfig(不会走前端那种「全自动」空配置)
python scripts/chunk_cli.py "path/to/file.pdf" --preview 3
python scripts/chunk_cli.py "path/to/file.xlsx" --mode by_row --rows-per-chunk 5 -o storage/result.json
python scripts/chunk_cli.py "readme.md" --mode delimiter --delimiter "##"
python scripts/chunk_cli.py "readme.md" --mode parent_child \
--parent-delimiter "##" --child-delimiter "###" \
--max-chunk-size 2000 --child-max-size 512
```
批量样例(若有 `data/`):`python scripts/run_all_samples.py`
### API
```bash
# 健康检查
curl http://127.0.0.1:8000/health
# 切分(不传 mode 等参数 → 自动策略,与 Demo 一致)
curl -X POST http://127.0.0.1:8000/api/chunk -F "file=@./your.pdf"
# 切分(显式参数)
curl -X POST http://127.0.0.1:8000/api/chunk \
-F "file=@./your.xlsx" \
-F "mode=by_row" \
-F "rows_per_chunk=5"
# 召回(先切分拿到 doc_id)
curl -X POST http://127.0.0.1:8000/api/recall \
-H "Content-Type: application/json" \
-d '{"doc_id":"abcdef123456","query":"如何修改交易密码","top_k":5}'
```
也可直接:`cd backend && uvicorn api.main:app --reload --port 8000`
### Python
```python
from pathlib import Path
from rag_cut import chunk_document, SplitConfig, SplitMode
from rag_cut.retrieval import recall_chunks
# 自动策略
result = chunk_document(Path("doc.pdf"))
# 显式配置
result = chunk_document(
Path("doc.docx"),
config=SplitConfig(mode=SplitMode.DEFAULT, max_chunk_size=1500),
)
hits, n = recall_chunks("login password", result.chunks, top_k=5)
```
将 `backend/` 加入 `PYTHONPATH`,或通过 `scripts/chunk_cli.py` 调用。
### 输出字段(节选)
```json
{
"filename": "guide.docx",
"doc_id": "61c2a152cd5e",
"split_mode": "default",
"split_config": {
"mode": "default",
"max_chunk_size": 2600,
"pdf_chunk_strategy": "pdf_feature_step_screenshot",
"chunk_strategy": "heading_layout_multimodal"
},
"block_count": 269,
"chunk_count": 48,
"chunks": [
{
"index": 0,
"content": "# …\n\n![image3](assets/61c2a152cd5e/image3.png)",
"char_count": 431,
"block_types": ["heading", "paragraph", "image"],
"meta": {
"heading": "1 Getting Started",
"pages": [3],
"embedding_text": "…",
"retrieval": true
}
}
],
"assets_dir": "storage/assets/61c2a152cd5e"
}
```
---
## 期望效果对照
| 需求 | 状态 |
| ----------------- | ------------------------------------------- |
| 图片保留原文位置 | ✅ PDF / Word / PPT(转 PDF) |
| 表格提取为 Markdown | ✅ PDF / Excel / CSV |
| Word 同标题内容同 chunk | ✅ 语义切分 |
| PDF 多栏阅读顺序 | ✅ 双栏检测 + 行合并 |
| 页眉页脚 / 页码噪声过滤 | ✅ MinerU 类型跳过 + 位置启发式 |
| 手册截图文字不混入正文 | ✅ 大图区域过滤 |
| 父子切片(超长章节) | ❌ 已从默认切分移除;超长仅按长度/子标题拆分 |
| 父子标识符切分 | ✅ 独立模式 `parent_child`(父/子标识符 + 双长度) |
| 召回入参/出参可检视 | ✅ `/api/recall` + Demo |
| wps | ⏳ 未实现 |
---
## 后续计划
- [ ] PDF OCR / 视觉描述增强
- [ ] 向量召回(Embedding)
- [ ] wps 格式支持
---
## 参考
- `[docx/自研搭建AI助手知识库.pdf](docx/自研搭建AI助手知识库.pdf)` — 需求
- `[docx/mineru-integration.md](docx/mineru-integration.md)` — MinerU 接入
- `[docx/tencent-cloud-document-splitting-settings.md](docx/tencent-cloud-document-splitting-settings.md)` — 腾讯云切分参考
- `[CLAUDE.md](CLAUDE.md)` / `[AGENTS.md](AGENTS.md)` — AI 协作说明