diff --git a/.gitignore b/.gitignore index 379b370..5335383 100644 --- a/.gitignore +++ b/.gitignore @@ -14,11 +14,15 @@ storage/results/ storage/assets/ output/ -# IDE / local agent +# IDE / local agent / task notes (keep local only) .idea/ .vscode/ .claude/ .learnings/ +AGENTS.md +CLAUDE.md +IMPACT_ANALYSIS.md +TASK_SUMMARY.md # OS Thumbs.db diff --git a/AGENTS.md b/AGENTS.md deleted file mode 100644 index 32d525f..0000000 --- a/AGENTS.md +++ /dev/null @@ -1,48 +0,0 @@ -# AGENTS.md - -This file provides guidance to Codex (Codex.ai/code) when working with code in this repository. - -## Project status - -This is a **greenfield project** — as of 2026-07-06 there is no code, no build system, and no git history yet. The only artifact is the requirements document `docx/自研搭建AI助手知识库.pdf` ("Self-built AI Assistant Knowledge Base"). Treat this file as the source of truth for requirements until code supersedes it. - -## What this project is - -A self-hosted **AI assistant knowledge base demo** (RAG system) whose primary goal is **high-quality document chunking/slicing**; recall quality is secondary. The driving problem: an existing knowledge-base feature ("AFE") slices poorly — it drops images and tables from word/pdf/excel documents, and its rules forcibly split a single passage across separate chunks. This project rebuilds the ingestion/slicing pipeline and exposes a demo UI to inspect the results. - -## Requirements (from the PDF) - -### Demo capabilities -- Upload files and **view the chunking/slicing result**. -- **Test recall** and inspect recall input parameters and return parameters. - -### Supported document types -- **Office docs:** pdf, doc, docx, ppt, pptx, wps, ppsx -- **Tabular / structured:** xlsx, xls, csv, md, txt, html, json, xml, log -- **Images:** jpg, png, jpeg, bmp, gif - -### Slicing rules -- **Table documents:** default split **or** split-by-row. -- **Other documents:** default split **or** universal-identifier split. -- **Tables inside a slice:** rendered as **Markdown**. - -### Expected effects — basic -1. Images stay in their original position in the rendered text — not lost or relocated to another chunk. -2. Table content is extracted correctly and converted to Markdown. -3. Slicing is structurally aware: for Word, all content under the same heading stays within one chunk. - -### Expected effects — advanced -1. Images/graphics are extracted as images **and** their text is OCR-recognized and converted to text. -2. PDF text laid out in horizontal columns is read back in the correct reading order (not jumbled). - -## Environment notes - -- **OS:** Windows 11. Shell is **bash** (Git Bash / MSYS2) — use Unix syntax (`/dev/null`, forward slashes), not PowerShell/CMD. -- **Python:** available at `/d/conda/python` (a conda environment). Relevant libraries **already installed** and useful for this project: `PyMuPDF` (fitz), `pdfplumber`, `pdfminer.six`, `pypdf`/`PyPDF2`, `pypdfium2`, `pikepdf`, `pdf2image`. -- **Reading the requirements PDF:** the file is image-heavy. `pdftotext -enc UTF-8` (at `/mingw64/bin/pdftotext`) extracts the little body text present but **misses the embedded diagrams** that the PDF references with "如下图" ("as shown below"). `pdftoppm` (image rendering) is **not** installed, so to view the diagrams use PyMuPDF from Python instead, e.g. `fitz.open(path)[page].get_pixmap()`. -- **Not a git repository.** Do not assume `git` workflows; initialize one only if asked. - -## Working in this repo - -- Before adding a chunking/slicing behavior, re-check the PDF's "期望效果" section above — the heading-awareness rule for Word and the image-position-preservation rule are the two most likely to be violated by naive splitters. -- The folder is named "RAG-cut" — chunking quality is the headline deliverable, recall is secondary. Prioritize ingestion/parsing fidelity over retrieval sophistication. diff --git a/CLAUDE.md b/CLAUDE.md deleted file mode 100644 index 04b668f..0000000 --- a/CLAUDE.md +++ /dev/null @@ -1,48 +0,0 @@ -# CLAUDE.md - -This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository. - -## Project status - -This is a **greenfield project** — as of 2026-07-06 there is no code, no build system, and no git history yet. The only artifact is the requirements document `docx/自研搭建AI助手知识库.pdf` ("Self-built AI Assistant Knowledge Base"). Treat this file as the source of truth for requirements until code supersedes it. - -## What this project is - -A self-hosted **AI assistant knowledge base demo** (RAG system) whose primary goal is **high-quality document chunking/slicing**; recall quality is secondary. The driving problem: an existing knowledge-base feature ("AFE") slices poorly — it drops images and tables from word/pdf/excel documents, and its rules forcibly split a single passage across separate chunks. This project rebuilds the ingestion/slicing pipeline and exposes a demo UI to inspect the results. - -## Requirements (from the PDF) - -### Demo capabilities -- Upload files and **view the chunking/slicing result**. -- **Test recall** and inspect recall input parameters and return parameters. - -### Supported document types -- **Office docs:** pdf, doc, docx, ppt, pptx, wps, ppsx -- **Tabular / structured:** xlsx, xls, csv, md, txt, html, json, xml, log -- **Images:** jpg, png, jpeg, bmp, gif - -### Slicing rules -- **Table documents:** default split **or** split-by-row. -- **Other documents:** default split **or** universal-identifier split. -- **Tables inside a slice:** rendered as **Markdown**. - -### Expected effects — basic -1. Images stay in their original position in the rendered text — not lost or relocated to another chunk. -2. Table content is extracted correctly and converted to Markdown. -3. Slicing is structurally aware: for Word, all content under the same heading stays within one chunk. - -### Expected effects — advanced -1. Images/graphics are extracted as images **and** their text is OCR-recognized and converted to text. -2. PDF text laid out in horizontal columns is read back in the correct reading order (not jumbled). - -## Environment notes - -- **OS:** Windows 11. Shell is **bash** (Git Bash / MSYS2) — use Unix syntax (`/dev/null`, forward slashes), not PowerShell/CMD. -- **Python:** available at `/d/conda/python` (a conda environment). Relevant libraries **already installed** and useful for this project: `PyMuPDF` (fitz), `pdfplumber`, `pdfminer.six`, `pypdf`/`PyPDF2`, `pypdfium2`, `pikepdf`, `pdf2image`. -- **Reading the requirements PDF:** the file is image-heavy. `pdftotext -enc UTF-8` (at `/mingw64/bin/pdftotext`) extracts the little body text present but **misses the embedded diagrams** that the PDF references with "如下图" ("as shown below"). `pdftoppm` (image rendering) is **not** installed, so to view the diagrams use PyMuPDF from Python instead, e.g. `fitz.open(path)[page].get_pixmap()`. -- **Not a git repository.** Do not assume `git` workflows; initialize one only if asked. - -## Working in this repo - -- Before adding a chunking/slicing behavior, re-check the PDF's "期望效果" section above — the heading-awareness rule for Word and the image-position-preservation rule are the two most likely to be violated by naive splitters. -- The folder is named "RAG-cut" — chunking quality is the headline deliverable, recall is secondary. Prioritize ingestion/parsing fidelity over retrieval sophistication. diff --git a/IMPACT_ANALYSIS.md b/IMPACT_ANALYSIS.md deleted file mode 100644 index c48e546..0000000 --- a/IMPACT_ANALYSIS.md +++ /dev/null @@ -1,31 +0,0 @@ -# Impact Analysis Report — 目录不解析、不切片 - -## 1. 改动概览 - -- **背景与目标**:所有文件跳过文档目录(目录 / Contents / TOC),不进入解析结果与切片。 -- **涉及模块**:`noise_filter.py`、`pipeline.py`、`pdf_semantic.py`、`mineru_adapter.py`、相关测试、`README.md` -- **改动类型**:功能策略 / 缺陷修复(噪声过滤增强) - -## 2. 方法级改动分析 - -| 位置 | 差异 | -| --- | --- | -| `is_toc_title_text` / `is_toc_entry_line` / `is_toc_noise_text` | 统一识别目录标题与带页码/引导线条目 | -| `filter_toc_blocks` | 删除目录页/目录区;保留无页码的正文内清单(如形態指標) | -| `filter_noise_blocks` | 串联 TOC 过滤 | -| `pipeline.chunk_document` | 全格式解析后强制 `filter_toc_blocks` | -| MinerU `_SKIP_MINERU_TYPES` | 增加 `toc` / `contents` 等类型 | - -## 3. 调用方与影响范围 - -- PDF/Office(MinerU / PyMuPDF)及 md/txt 等所有走 `chunk_document` 的格式 -- **破坏性变更:否**(目录内容不再出现在切片中,属预期) - -## 4. 风险与回滚 - -- **风险级别**:低~中(极少数非目录但带引导线+页码的行可能被误删) -- **回滚方式是否简单:是** - -## 5. 验证与测试 - -- 新增/更新噪声与语义切分单测;请对含目录的样本重新切分确认 diff --git a/README.md b/README.md index ef32468..85023ab 100644 --- a/README.md +++ b/README.md @@ -1,36 +1,104 @@ # RAG-cut -自研 AI 助手知识库 Demo —— 以**高质量文档切片**为核心,召回为次要目标。 +自研 AI 助手知识库 Demo —— 以**高质量文档切片**为核心,召回为辅。 -重建文档 ingestion / 切片流水线,解决现有 AFE 知识库的两类问题: +重建文档 ingestion / 切片流水线,针对现有 AFE 知识库的两类问题: -1. **文档解析不足**:Word / PDF / Excel 中的图片、表格无法有效提取 -2. **切片规则不佳**:同一逻辑段落被硬切到不同 chunk,图片丢失或错位 +1. **解析不足**:Word / PDF / Excel 中的图片、表格无法有效提取 +2. **切分不佳**:同一逻辑段落被硬切到不同 chunk,图片丢失或错位 -需求来源:`[docx/自研搭建AI助手知识库.pdf](docx/自研搭建AI助手知识库.pdf)` +| 项 | 说明 | +| --- | --- | +| 仓库 | `http://192.168.3.110:3000/AFE_RAG/RAG-CUT.git` | +| 需求文档 | [`docx/自研搭建AI助手知识库.pdf`](docx/自研搭建AI助手知识库.pdf) | +| 样例文档 | `自研RAG代表文档/` | --- ## 功能概览 - -| 能力 | 说明 | -| --------- | --------------------------------------------------------------------------------------------- | -| 多格式解析 | pdf / doc / docx / ppt / pptx / ppsx / xlsx / csv / md / txt / html / json / xml / log / 常见图片 | -| PDF 双引擎 | MinerU(优先)+ PyMuPDF 回退;统一噪声过滤 | -| 自动切分策略 | 前端可选自动,或手动指定 mode | -| 语义切分 | PDF / Word 走 `heading_layout_multimodal`;超长章节仅按长度/子标题拆分(不再产父片) | -| 父子标识符 | 独立模式 `parent_child`:父/子标识符双层切分,子片检索、父片上下文 | -| 表格切分 | 自动识别表头 / 说明行;Q&A 表按行 1 片;普通表按行数自适应 | -| Demo UI | 原文预览 · 切片列表 · Markdown 预览 · **历史切片回看** · BM25 召回测试 | -| CLI / API | `scripts/chunk_cli.py` · FastAPI `/api/chunk` · `/api/results` · `/api/recall` | - +| 能力 | 说明 | +| --- | --- | +| 多格式解析 | pdf / doc / docx / ppt / pptx / ppsx / xlsx / csv / md / txt / html / json / xml / log / 常见图片 | +| PDF 双引擎 | MinerU(优先)+ PyMuPDF 回退;统一噪声过滤(页眉页脚、页码、**目录 TOC**) | +| 四种切分模式 | `default` · `delimiter` · `parent_child` · `by_row`(Demo 默认选结构感知) | +| 语义切分 | PDF / Word 在 `default` 下走 `heading_layout_multimodal`;超长章节按子标题/长度拆分(不产父片) | +| 父子标识符 | 独立模式 `parent_child`:子片检索、父片上下文 | +| 表格切分 | `by_row`:自动识别表头 / 说明行;Q&A 表可按行 1 片 | +| Demo UI | 原文预览 · 切片列表 · Markdown 预览 · **历史切片回看** · BM25 召回测试 | +| CLI / API | `scripts/chunk_cli.py` · `/api/chunk` · `/api/results` · `/api/recall` | 未支持:`.wps`(可先转为 docx/pdf)。独立图片暂不做 OCR。 --- +## 快速开始 +### 1. 克隆与依赖 + +```bash +git clone http://192.168.3.110:3000/AFE_RAG/RAG-CUT.git +cd RAG-CUT + +pip install -r requirements.txt +``` + +另需本机工具(按文档类型): + +| 场景 | 依赖 | +| --- | --- | +| `.doc` / `.ppt` 等 | **LibreOffice**(`soffice`)或 Windows 上的 **Microsoft Word/PowerPoint** | +| PDF 高精度解析 | MinerU(`requirements.txt` 已含 `mineru[all]`) | +| MinerU 图内文字回填 | 本机 Tesseract(可选,见下方环境变量) | + +### 2. 启动 Demo + +```bash +cd backend +python run.py +``` + +浏览器打开:**http://127.0.0.1:8000**(后端同时托管前端静态页)。 + +前后端分离开发时另开终端: + +```bash +cd frontend +python run.py # http://127.0.0.1:5173 ,API 默认指向 :8000 +``` + +可选参数:`--host 0.0.0.0 --port 8000 --no-reload`(后端)· `--port 5173`(前端) + +### 3. 运行测试 + +```bash +cd backend +python -m pytest tests/ -q +``` + +### 环境变量(PDF) + +| 变量 | 默认 | 说明 | +| --- | --- | --- | +| `RAG_CUT_PDF_ENGINE` | `auto` | `auto` 优先 MinerU,失败回退 PyMuPDF;可强制 `mineru` / `pymupdf` | +| `RAG_CUT_MINERU_CMD` | — | MinerU 可执行路径(不在 PATH 时) | +| `RAG_CUT_MINERU_API_URL` | — | 复用已启动的 MinerU API,避免每次切分重复加载模型 | +| `RAG_CUT_MINERU_TIMEOUT` | `540` | 单次 MinerU 解析超时(秒) | +| `RAG_CUT_MINERU_OCR` | `1` | MinerU 图片无 OCR 时用 Tesseract 回填;`0` 关闭以加快切分 | + +详见 [`docx/mineru-integration.md`](docx/mineru-integration.md)。 + +### 运行栈 + +| 类别 | 选型 | +| --- | --- | +| Python | 3.12+ | +| Web | FastAPI + Uvicorn | +| PDF | MinerU / PyMuPDF + pdfplumber | +| Excel | openpyxl | +| 召回 | 内置 BM25(无向量库) | + +--- ## 技术架构 @@ -47,6 +115,9 @@ 版面增强 (layout_meta) ──► 标题层级、图片/表格上下文、order_index │ ▼ +目录过滤 (filter_toc_blocks) ──► 全格式剔除「目录 / Contents」区 + │ + ▼ 切分策略 (split_policy / pdf_strategy) │ ▼ @@ -68,124 +139,61 @@ RAG-cut/ │ │ ├── split_policy.py / retrieval.py │ │ ├── parsers/ # word / ppt / pdf / xlsx / text / image + MinerU │ │ │ └── pdf/ # PyMuPDF 流水线、噪声过滤、跨页表合并 -│ │ └── splitters/ # default / heading / pdf_semantic / by_row / delimiter -│ ├── api/main.py # FastAPI:/api/chunk · /api/results · /api/recall · 静态前端 +│ │ └── splitters/ # default / heading / pdf_semantic / by_row / delimiter / parent_child +│ ├── api/main.py # FastAPI + 静态前端 │ ├── tests/ │ └── run.py ├── frontend/ # Demo(index.html + css + js) ├── scripts/ # chunk_cli · run_all_samples -├── storage/ # uploads / assets / results(运行时生成) -├── docx/ # 需求 PDF、MinerU 说明 +├── storage/ # uploads / assets / results(运行时,已 gitignore) +├── docx/ # 需求 PDF、MinerU / 腾讯云切分说明 +├── 自研RAG代表文档/ # 样例语料 └── requirements.txt ``` --- - - -## 快速开始 - -```bash -pip install -r requirements.txt - -cd backend -python run.py -``` - -浏览器打开:**[http://127.0.0.1:8000](http://127.0.0.1:8000)** - -前后端分离开发时,另开终端: - -```bash -cd frontend -python run.py # http://127.0.0.1:5173 ,API 需自行指向 :8000 -``` - -可选参数:`--host 0.0.0.0 --port 8000 --no-reload`(后端)· `--port 5173`(前端) - -### 环境变量(PDF) - - -| 变量 | 默认 | 说明 | -| -------------------- | ------ | ------------------------------------------------------ | -| `RAG_CUT_PDF_ENGINE` | `auto` | `auto` 优先 MinerU,失败回退 PyMuPDF;可强制 `mineru` / `pymupdf` | -| `RAG_CUT_MINERU_CMD` | — | MinerU 可执行文件路径(不在 PATH 时) | -| `RAG_CUT_MINERU_API_URL` | — | 复用已启动的 MinerU API,避免每次切分重复加载模型 | -| `RAG_CUT_MINERU_TIMEOUT` | `540` | 单次 MinerU 解析超时秒数;超时后会清理其临时服务和子进程 | -| `RAG_CUT_MINERU_OCR` | `1` | MinerU 图片无 OCR 时用 Tesseract 回填;`0` 关闭以加快切分 | - - -详见 `[docx/mineru-integration.md](docx/mineru-integration.md)`。 - -### 依赖与运行环境 - - -| 类别 | 选型 | -| ------------ | -------------------------------------------------------------- | -| Python | 3.12+ | -| Web | FastAPI + Uvicorn | -| PDF | MinerU(`requirements.txt` 已含)/ PyMuPDF + pdfplumber | -| Excel | openpyxl | -| Office 转 PDF | 本机 **LibreOffice** 或 **Microsoft Word/PowerPoint**(doc/ppt 必需) | -| 召回 | 内置 BM25 词法评分(无向量库) | - - ---- - - - ## 各格式处理 -所有格式统一:**解析 → Block → 切分 → Markdown Chunk**。 - - -| 扩展名 | 解析器 | 要点 | 默认切分 | -| ---------------------- | ------------------- | ------------------------------- | ------------------------------ | -| `.pdf` | `PdfParser` | MinerU / PyMuPDF → 噪声过滤 → 跨页表合并 | 语义切分(`max=2600`) | -| `.doc` `.docx` | `WordParser` | 转 PDF 后复用 PDF 流水线 | 同上 | -| `.ppt` `.pptx` `.ppsx` | `PptParser` | 转 PDF;页码 ≈ 幻灯片序号 | `default`(`max=2600`) | -| `.xlsx` `.xls` `.csv` | `SpreadsheetParser` | 自动检测 preamble / 表头 / 数据起始行 | `by_row`(Q&A 表 1 行/片,否则按行数自适应) | -| `.md` `.html` | `TextParser` | 按标题 / 段落拆块 | `default`(`max=1800`) | -| `.json` | `TextParser` | 顶层 key / 数组元素 → `code` 块 | `default` | -| `.txt` `.xml` `.log` | `TextParser` | 按空行分段(不解析 XML 结构) | `default` | -| 图片 | `ImageParser` | 整图 1 个 Block | 整图 1 片 | - - +统一流水线:**解析 → Block →(目录过滤)→ 切分 → Markdown Chunk**。 +| 扩展名 | 解析器 | 要点 | 推荐切分 | +| --- | --- | --- | --- | +| `.pdf` | `PdfParser` | MinerU / PyMuPDF → 噪声过滤 → 跨页表合并 | `default`(语义,`max≈2600`) | +| `.doc` `.docx` | `WordParser` | 转 PDF 后复用 PDF 流水线 | 同上 | +| `.ppt` `.pptx` `.ppsx` | `PptParser` | 转 PDF;页码 ≈ 幻灯片序号 | `default` | +| `.xlsx` `.xls` `.csv` | `SpreadsheetParser` | 检测 preamble / 表头 / 数据起始行 | **`by_row`**(Demo 需手动选) | +| `.md` `.html` | `TextParser` | 按标题 / 段落拆块 | `default` / `delimiter` / `parent_child` | +| `.json` | `TextParser` | 顶层 key / 数组元素 → `code` 块 | `default` | +| `.txt` `.xml` `.log` | `TextParser` | 按空行分段 | `default` | +| 图片 | `ImageParser` | 整图 1 个 Block | 整图 1 片 | ### 统一产出 - -| 项目 | 说明 | -| -------- | ---------------------------------------------------------------------------------------------- | -| Block 类型 | `heading` / `paragraph` / `table` / `image` / `list` / `code` | -| Chunk | Markdown + `meta`(`heading` / `pages` / `embedding_text` / `retrieval` / 可选 `parent_chunk_id`) | -| 图片 | `storage/assets/{doc_id}/`;JSON 内为 `assets/{doc_id}/…` | -| 结果缓存 | `storage/results/{doc_id}.json`(供召回) | -| 中间产物 | `_conversion/`(Office→PDF)、`_mineru/`(MinerU 输出) | - - - +| 项目 | 说明 | +| --- | --- | +| Block 类型 | `heading` / `paragraph` / `table` / `image` / `list` / `code` | +| Chunk | Markdown + `meta`(`heading` / `pages` / `embedding_text` / `retrieval` / 可选 `parent_chunk_id`) | +| 图片 | `storage/assets/{doc_id}/`;JSON 内为 `assets/{doc_id}/…` | +| 结果缓存 | `storage/results/{doc_id}.json`(历史回看 / 召回) | +| 中间产物 | `_conversion/`(Office→PDF)、`_mineru/`(MinerU 输出) | ### Word / PPT 转 PDF 优先 LibreOffice(`soffice --headless`),Windows 可回退 Office COM。转换失败时 API 返回 500 并附带各后端错误信息。 -### Excel / CSV 布局检测 +### Excel / CSV -`detect_spreadsheet_layout` 会识别常见模板(如第 1 行字段说明、第 2 行列表头、第 3 行起数据),跳过 preamble;Q&A 评测表(列名含 query/answer 等)自动 `rows_per_chunk=1`。 +`detect_spreadsheet_layout` 识别常见模板(说明行 + 表头 + 数据区);Q&A 评测表(列名含 query/answer 等)适合 `rows_per_chunk=1`。 > `.xls` 依赖 openpyxl,老二进制格式可能失败,建议先转 xlsx。 --- +## PDF / Word 语义切分 - -## PDF 切割方法 - -`.pdf` / `.doc` / `.docx` 在 `mode=default` 时走专用语义切分器 -`splitters/pdf_semantic.py`(策略名:`heading_layout_multimodal`)。 -Word 先转 PDF,再与 PDF 共用同一套规则。默认 `max_chunk_size=2600`、`overlap=120`。 +`.pdf` / `.doc` / `.docx` 在 `mode=default` 时走 `splitters/pdf_semantic.py`(策略名:`heading_layout_multimodal`)。 +Word 先转 PDF,再与 PDF 共用规则。自动策略下默认 `max_chunk_size=2600`、`overlap=120`;Demo 默认表单约为 `2200` / `120`。 ### 流水线 @@ -193,190 +201,121 @@ Word 先转 PDF,再与 PDF 共用同一套规则。默认 `max_chunk_size=2600 PDF/Word │ ├─ 解析:MinerU(优先)或 PyMuPDF → Block 流 - ├─ 噪声过滤:页眉/页脚/页码、**整段文档目录**、装饰小图 + ├─ 噪声过滤:页眉/页脚/页码、文档目录、装饰小图 ├─ 跨页表格合并 - ├─ 版面增强:标题层级、图片/表格上下文、order_index - ├─ pdf_strategy:打标签(操作手册 vs 年报,见下) - └─ heading_layout_multimodal 切分 → Chunk + ├─ 版面增强:标题层级、图片/表格上下文 + ├─ pdf_strategy:打标签(操作手册 vs 年报) + └─ heading_layout_multimodal → Chunk ``` ### 核心规则:`heading_layout_multimodal` -目标:同一逻辑小节(标题 + 正文 + 同节截图/表格)尽量落在同一 chunk,避免把步骤截图和说明文字拆开。 +目标:同一逻辑小节(标题 + 正文 + 同节截图/表格)尽量落在同一 chunk。 | 步骤 | 行为 | -|------|------| -| 1. 再过滤噪声 | 跳过页码、running header、**文档目录(目录/Contents 及条目)**、过小装饰图 | +| --- | --- | +| 1. 再过滤噪声 | 跳过页码、running header、目录区、过小装饰图 | +| 2. 识别标题 | 解析器已标 `heading`,或正则/字号启发式补识别 | +| 3. 维护标题路径 | 标题栈记录章节层级 | +| 4. 按标题开新片 | 同级或更高级标题时 flush,开新 chunk | +| 5. 图文同节 | 正文、image、table 跟在当前标题路径下 | +| 6. 超长二次切 | 超过 `max_chunk_size` 时按子标题/长度拆分;**不单独保留父片** | -| 2. 识别标题 | 解析器已标的 `heading`,或正则/字号启发式补识别 | -| 3. 维护标题路径 | 用标题栈记录章节层级,如 `1 Getting Started → A. Log in` | -| 4. 按标题开新片 | 遇到同级或更高级标题时 flush 当前组,开新 chunk | -| 5. 图文同节 | 正文、image、table 跟在当前标题路径下;绑定前后文到 `meta` | -| 6. 超长二次切 | 超过 `max_chunk_size` 时按子标题/内容标记/长度拆成多片;**不单独保留父片** | +可识别标题形态示例:`1.` / `1.2.3` / `A.` / `Chapter 2` / `一、` / `(一)`;操作步骤 `Step N …` 作正文保留。封面 Logo 等无语义组标 `retrieval=false`。 -**可识别的标题形态**(示例): +### 文档类型标签:`pdf_strategy` -- 编号:`1.` / `1.2.3` / `A.` / `Chapter 2`(操作步骤 `Step N …` 作为正文保留,不单独开章) -- 中文章节:`一、` / `(一)` 等 -- 字号显著大于正文的短文本(PyMuPDF 路径) +写入 `split_config.pdf_chunk_strategy`(Demo 策略面板可显示): -封面 Logo 等无语义组会标 `retrieval=false`,不进召回候选。 +| 标签 | 典型特征 | 说明 | +| --- | --- | --- | +| `pdf_feature_step_screenshot` | 步骤标题多、操作词多、截图密度高 | 操作手册型 | +| `pdf_outline_report` | 「年报/财务报表」等词,或文件名含 annual/report/年报 | 年报/报告型 | -### 父子标识符切分(`mode=parent_child`) +当前两类切分算法相同,标签用于结果标注与后续策略分叉预留。 -适合「细粒度检索、粗粒度召回」:子片用于 BM25 检索,命中后通过 `parent_chunk_id` 关联父片上下文。 +--- -> 默认切分**不会**再因超长章节自动产出父片+子片;需要父子结构时请显式选择本模式。 +## 切分模式(全格式) -流程: -1. 按 **父级标识符**(`parent_delimiter`)切开,得到父片;超过父级最大长度时再按长度二次切 -2. 每个父片再按 **子级标识符**(`child_delimiter`,可选)切开,得到子片;超过子级最大长度时再按长度二次切 +后端 **4** 种模式(`SplitMode`):`default` / `delimiter` / `parent_child` / `by_row`。 + +| 模式 | 设计意图 | 行为 | +| --- | --- | --- | +| `default` | 非表格正文 | PDF/Word 走语义切分;其他按标题/页/长度打包;`image`/`table` 不拆 | +| `delimiter` | 正文含可匹配标识符 | 按自定义标识符切开,再按长度二次切 | +| `parent_child` | 层级标识符 | 父标识符切父片,子标识符切子片;检索用子片 | +| `by_row` | 表格 | 每片 = 表头 + N 行 Markdown | + +**Demo**:始终显式传 `mode`(默认 `default`)。切 Excel/CSV 时请选手动「按行切分」。 + +**自动策略**(`choose_split_config`):仅当 API/Python **不传** `mode`/`config` 时生效——表格类选 `by_row`,PDF/PPT 等选 `default` 并覆盖 `max_chunk_size`。CLI 默认带 `--mode default`,不会走该自动选型。 + +### 父子标识符(`mode=parent_child`) + +1. 按 `parent_delimiter` 切父片;超长再按长度二次切 +2. 再按 `child_delimiter`(可选)切子片 3. 父片:`is_section_parent=true`,`retrieval=false` 4. 子片:`is_sub_chunk=true`,`retrieval=true`,`parent_chunk_id` 指向父片 -参数约束: - | 参数 | 说明 | -|------|------| -| `parent_delimiter` | 必填;父级切开标记(不会写入切片正文) | -| `child_delimiter` | 可选;缺省则仅按子级最大长度拆子片 | +| --- | --- | +| `parent_delimiter` | 必填;不写入切片正文 | +| `child_delimiter` | 可选;缺省则仅按子级最大长度拆 | | `max_chunk_size` | 父级最大长度 | -| `child_max_size` | 子级最大长度,且 ≤ 父级、硬上限 1500 | -| `overlap` | 超长二次切分时的重叠 | +| `child_max_size` | 子级最大长度(≤ 父级,硬上限 1500) | +| `overlap` | 超长二次切分重叠 | -Demo / CLI / API 均可显式选择此模式。 -### 文档类型标签:`pdf_strategy` - -`splitters/pdf_strategy.py` 根据文件名与正文特征打标签,写入 -`split_config.pdf_chunk_strategy`(Demo 自动策略面板会显示): - -| 标签 | 典型特征 | 说明 | -|------|----------|------| -| `pdf_feature_step_screenshot` | 步骤标题多、「点击/输入/选择」等操作词多、截图密度高 | 操作手册型 | -| `pdf_outline_report` | 「年报/财务报表/董事会」等词多,或文件名含 annual/report/年报 | 年报/报告型 | - -> 当前两类**切分算法相同**(都走 `heading_layout_multimodal`);标签用于结果标注,并为后续分叉策略预留。 - -### 其他切分模式(PDF 也可用) - -| 模式 | 何时用 | 行为 | -|------|--------|------| -| `default` | Demo / 自动策略(推荐) | 上文语义切分 | -| `delimiter` | CLI/API/前端显式指定 | 按自定义标识符切,再按长度二次切 | -| `parent_child` | CLI/API/前端显式指定 | 父子标识符切分(检索用子片、父片作上下文) | -| `by_row` | 一般不用于 PDF | 面向表格文档 | - -相关代码:`pdf_semantic.py` · `pdf_strategy.py` · `parsers/pdf/noise_filter.py` · `pipeline.py`。 - ---- - -## 切分策略(全格式) - -后端共 **4** 种切分模式(`SplitMode`):`default` / `delimiter` / `parent_child` / `by_row`。 -前端「自动」(`auto`) 不是独立模式:不传 `mode` 时由 `split_policy.choose_split_config` 按扩展名与表结构选型。 - -API **不会按扩展名拦截 mode**——凡解析器支持的格式均可手动指定任意模式;下表区分设计意图与有效行为。 - -### 可解析格式(21 种) - -| 类别 | 扩展名 | -| ---- | ------ | -| Office | `.pdf` `.doc` `.docx` `.ppt` `.pptx` `.ppsx` | -| 表格 | `.xlsx` `.xls` `.csv` | -| 文本 | `.md` `.txt` `.html` `.htm` `.json` `.xml` `.log` | -| 图片 | `.jpg` `.jpeg` `.png` `.bmp` `.gif` | - -> 需求中的 `.wps` 尚未接入解析器。 - -### 四种模式与文档格式 - -| 模式 | 适用(设计意图) | 实际覆盖的格式 | 行为 | -| ---- | ---------------- | -------------- | ---- | -| `default` | 非表格 | 除表格类外全部;表格也可手动指定 | 结构/标题感知切分;`.pdf`/`.doc`/`.docx` 走语义切分(`heading_layout_multimodal`);单章超长时按子标题或长度拆开(**不产父片**) | -| `delimiter` | 非表格(正文含可匹配标识符) | 全部可解析格式 | 按自定义标识符(如 `###`)切开;超长片段再按 `max_chunk_size` + `overlap`;无标识符时效果差 | -| `parent_child` | 非表格(层级标识符) | 全部可解析格式 | 先按父标识符切父片,再按子标识符切子片;检索用子片,父片作上下文;典型如 `.md`/文本 | -| `by_row` | 表格 | **有效**:`.xlsx` `.xls` `.csv`,或解析后含带 `rows` 的 TABLE block | 每片 = 表头 + N 行 Markdown;无表格行数据时退化为整篇一片;一般不用于 PDF/Word | - -### 自动策略(不传 mode) - -| 文档 | 自动选的模式 | -| ---- | ------------ | -| `.xlsx` `.xls` `.csv`,或整篇就一张表(含 `rows`) | `by_row` | -| `.pdf` `.ppt` `.pptx` `.ppsx` | `default` | -| 图片类 | `default` | -| 文本类(`.md` `.txt` `.html` `.htm` `.json` `.xml` `.log`) | `default` | -| `.doc` `.docx` 等其余 | `default` | - -**一句话**:`by_row` 专吃表格;另外三种面向正文结构,其中 `default` 对 PDF/Word 有专用路径,`delimiter` / `parent_child` 对有标识符的文本最有用,但格式本身不限制。 - - -### `default` 决策树 +### `default` 决策树(摘要) ``` -.pdf / .doc / .docx - └─ 见上方「PDF 切割方法」(heading_layout_multimodal) -单 table 且含 rows - └─ 自动 by_row -其他 + 足够标题 - └─ 标题大纲切;若某章超过 max_chunk_size - → 优先按子标题 / 内容标记拆分 - → 否则按长度拆分(不保留整章父片) -有 page 元数据、标题不足 - └─ 按页;单页超限再按长度 -否则 - └─ 按 max_chunk_size 打包;image/table 不拆 +.pdf / .doc / .docx → heading_layout_multimodal +单 table 且含 rows → 仅「不传 mode」时自动 by_row;显式 default 则走通用切分 +有足够标题 → 标题大纲切;超长按子标题/长度拆(不产父片) +有 page、标题不足 → 按页;单页超限再按长度 +否则 → 按 max_chunk_size 打包 ``` -`meta.retrieval = false` 的切片(如封面 Logo,或 `parent_child` 模式的父片)不参与召回。 - ### SplitConfig +| 参数 | 默认 | 说明 | +| --- | --- | --- | +| `mode` | `default` | 见上表 | +| `delimiter` | — | `delimiter` 模式必填 | +| `parent_delimiter` / `child_delimiter` | — | `parent_child`:父必填,子可选 | +| `max_chunk_size` | 1500 | 自动策略常覆盖为 1800–2600 | +| `child_max_size` | 512 | `parent_child` 子级上限 | +| `overlap` | 150 | 超长二次切分重叠 | +| `header_row_start` / `header_row_end` | 1 | 表头行(1-based) | +| `start_row` | 2 | 数据起始行 | +| `rows_per_chunk` | 1 | 每片数据行数 | -| 参数 | 默认 | 说明 | -| ------------------------------------- | --------- | ------------------------------ | -| `mode` | `default` | `default` / `delimiter` / `parent_child` / `by_row` | -| `delimiter` | — | `delimiter` 模式必填 | -| `parent_delimiter` / `child_delimiter` | — | `parent_child` 模式:父标识符必填,子标识符可选 | -| `max_chunk_size` | 1500 | 单 chunk 上限(父级长度;自动策略常覆盖为 1800–2600) | -| `child_max_size` | 512 | `parent_child` 子级最大长度(≤ 父级,且 ≤ 1500) | -| `overlap` | 150 | 超长二次切分重叠 | -| `header_row_start` / `header_row_end` | 1 | 表头行(1-based) | -| `start_row` | 2 | 数据起始行 | -| `rows_per_chunk` | 1 | 每片数据行数 | - +可解析扩展名(21 种):`.pdf` `.doc` `.docx` `.ppt` `.pptx` `.ppsx` `.xlsx` `.xls` `.csv` `.md` `.txt` `.html` `.htm` `.json` `.xml` `.log` `.jpg` `.jpeg` `.png` `.bmp` `.gif`。需求中的 `.wps` 尚未接入。 --- - - ## Demo 前端 -三栏工作区:**原文** · **切片列表** · **切片预览**;左侧可选手动切分模式,或保持「自动」由系统选型。 - - -| 区域 | 说明 | -| ------ | ---------------------------------------- | -| 上传 | 拖拽 / 选择文件 | -| 切分模式 | 自动 / 默认(结构感知)/ 通用标识符 / **父子标识符** / 按行;手动模式可调对应参数 | -| 策略面板 | 切分后回显 mode / max_size / overlap / PDF 策略 | -| 原文 | PDF iframe / docx(mammoth)/ 文本 / 图片 | -| 切片列表 | 类型标签、父子切片标记、字符数 | -| 召回测试 | query + top_k;展示得分与命中切片 | -| API 面板 | 可折叠请求/响应 JSON | +三栏:**原文** · **切片列表** · **切片预览**;左侧配置切分模式与参数;下方可测召回。 +| 区域 | 说明 | +| --- | --- | +| 上传 | 拖拽 / 选择文件 | +| 切分模式 | 默认(结构感知)/ 通用标识符 / 父子标识符 / 按行 | +| 策略面板 | 切分后回显 mode / max_size / overlap / PDF 策略 | +| 历史切片 | 列出 `storage/results` 中已切文档,可回看 / 删除 | +| 原文 | PDF iframe / docx(mammoth)/ 文本 / 图片 | +| 切片列表 | 类型标签、父子标记、字符数 | +| 召回测试 | query + top_k;`retrieval=false` 的片不进候选 | +| API 面板 | 可折叠请求/响应 JSON | --- - - ## CLI / API - - ### CLI ```bash -# 显式传 SplitConfig(不会走前端那种「全自动」空配置) python scripts/chunk_cli.py "path/to/file.pdf" --preview 3 python scripts/chunk_cli.py "path/to/file.xlsx" --mode by_row --rows-per-chunk 5 -o storage/result.json python scripts/chunk_cli.py "readme.md" --mode delimiter --delimiter "##" @@ -393,22 +332,25 @@ python scripts/chunk_cli.py "readme.md" --mode parent_child \ # 健康检查 curl http://127.0.0.1:8000/health -# 切分(不传 mode 等参数 → 自动策略,与 Demo 一致) +# 切分(不传 mode → 服务端自动策略) curl -X POST http://127.0.0.1:8000/api/chunk -F "file=@./your.pdf" -# 切分(显式参数) +# 切分(显式参数,与 Demo 一致) curl -X POST http://127.0.0.1:8000/api/chunk \ -F "file=@./your.xlsx" \ -F "mode=by_row" \ -F "rows_per_chunk=5" +# 历史结果 +curl http://127.0.0.1:8000/api/results + # 召回(先切分拿到 doc_id) curl -X POST http://127.0.0.1:8000/api/recall \ -H "Content-Type: application/json" \ -d '{"doc_id":"abcdef123456","query":"如何修改交易密码","top_k":5}' ``` -也可直接:`cd backend && uvicorn api.main:app --reload --port 8000` +主要端点:`GET /health` · `POST /api/chunk` · `GET/DELETE /api/results…` · `POST /api/recall`。 ### Python @@ -417,13 +359,12 @@ from pathlib import Path from rag_cut import chunk_document, SplitConfig, SplitMode from rag_cut.retrieval import recall_chunks -# 自动策略 +# 不传 config → 自动策略 result = chunk_document(Path("doc.pdf")) -# 显式配置 result = chunk_document( Path("doc.docx"), - config=SplitConfig(mode=SplitMode.DEFAULT, max_chunk_size=1500), + config=SplitConfig(mode=SplitMode.DEFAULT, max_chunk_size=2600, overlap=120), ) hits, n = recall_chunks("login password", result.chunks, top_k=5) @@ -466,29 +407,25 @@ hits, n = recall_chunks("login password", result.chunks, top_k=5) --- - - ## 期望效果对照 - -| 需求 | 状态 | -| ----------------- | ------------------------------------------- | -| 图片保留原文位置 | ✅ PDF / Word / PPT(转 PDF) | -| 表格提取为 Markdown | ✅ PDF / Excel / CSV | -| Word 同标题内容同 chunk | ✅ 语义切分 | -| PDF 多栏阅读顺序 | ✅ 双栏检测 + 行合并 | -| 页眉页脚 / 页码噪声过滤 | ✅ MinerU 类型跳过 + 位置启发式 | -| 手册截图文字不混入正文 | ✅ 大图区域过滤 | -| 父子切片(超长章节) | ❌ 已从默认切分移除;超长仅按长度/子标题拆分 | -| 父子标识符切分 | ✅ 独立模式 `parent_child`(父/子标识符 + 双长度) | -| 召回入参/出参可检视 | ✅ `/api/recall` + Demo | -| wps | ⏳ 未实现 | - +| 需求 | 状态 | +| --- | --- | +| 图片保留原文位置 | ✅ PDF / Word / PPT(转 PDF) | +| 表格提取为 Markdown | ✅ PDF / Excel / CSV | +| Word 同标题内容同 chunk | ✅ 语义切分 | +| PDF 多栏阅读顺序 | ✅ 双栏检测 + 行合并 | +| 页眉页脚 / 页码噪声过滤 | ✅ MinerU 类型跳过 + 位置启发式 | +| 文档目录不进切片 | ✅ 全格式 `filter_toc_blocks` | +| 手册截图文字不混入正文 | ✅ 大图区域过滤 | +| 父子切片(超长章节自动) | ❌ 已从默认切分移除 | +| 父子标识符切分 | ✅ `parent_child` | +| 召回入参/出参可检视 | ✅ `/api/recall` + Demo | +| 历史切片回看 | ✅ `/api/results` + Demo | +| wps | ⏳ 未实现 | --- - - ## 后续计划 - [ ] PDF OCR / 视觉描述增强 @@ -497,11 +434,9 @@ hits, n = recall_chunks("login password", result.chunks, top_k=5) --- - - ## 参考 -- `[docx/自研搭建AI助手知识库.pdf](docx/自研搭建AI助手知识库.pdf)` — 需求 -- `[docx/mineru-integration.md](docx/mineru-integration.md)` — MinerU 接入 -- `[docx/tencent-cloud-document-splitting-settings.md](docx/tencent-cloud-document-splitting-settings.md)` — 腾讯云切分参考 -- `[CLAUDE.md](CLAUDE.md)` / `[AGENTS.md](AGENTS.md)` — AI 协作说明 +- [`docx/自研搭建AI助手知识库.pdf`](docx/自研搭建AI助手知识库.pdf) — 需求 +- [`docx/mineru-integration.md`](docx/mineru-integration.md) — MinerU 接入 +- [`docx/tencent-cloud-document-splitting-settings.md`](docx/tencent-cloud-document-splitting-settings.md) — 腾讯云切分参考 +- [`CLAUDE.md`](CLAUDE.md) / [`AGENTS.md`](AGENTS.md) — AI 协作说明 diff --git a/TASK_SUMMARY.md b/TASK_SUMMARY.md deleted file mode 100644 index a35b5b3..0000000 --- a/TASK_SUMMARY.md +++ /dev/null @@ -1,22 +0,0 @@ -# TASK_SUMMARY — 目录不解析、不切片 - -## 1. 任务基本信息 - -- **任务名称**:所有文件跳过文档目录 -- **相关项目**:RAG-cut - -## 2. 改动说明 - -解析与切分链路统一剔除「目录 / Contents / TOC」标题及带页码引导线的目录条目;正文内无页码清单(如形態指標列表)保留。 - -## 3. 影响与风险 - -- 见 `IMPACT_ANALYSIS.md`;**破坏性变更:否** - -## 4. 测试与验证 - -- 运行 `test_pdf_noise_filter`、`test_heading_layout_multimodal` 等相关单测 - -## 5. 后续事项 - -- 若遇无引导线、无页码的「伪目录」仍进切片,可再扩展 TOC 区检测