first commit

Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
陈辅元
2026-07-16 11:12:17 +08:00
co-authored by Cursor
commit 4003624b8c
80 changed files with 9990 additions and 0 deletions
+48
View File
@@ -0,0 +1,48 @@
# AGENTS.md
This file provides guidance to Codex (Codex.ai/code) when working with code in this repository.
## Project status
This is a **greenfield project** — as of 2026-07-06 there is no code, no build system, and no git history yet. The only artifact is the requirements document `docx/自研搭建AI助手知识库.pdf` ("Self-built AI Assistant Knowledge Base"). Treat this file as the source of truth for requirements until code supersedes it.
## What this project is
A self-hosted **AI assistant knowledge base demo** (RAG system) whose primary goal is **high-quality document chunking/slicing**; recall quality is secondary. The driving problem: an existing knowledge-base feature ("AFE") slices poorly — it drops images and tables from word/pdf/excel documents, and its rules forcibly split a single passage across separate chunks. This project rebuilds the ingestion/slicing pipeline and exposes a demo UI to inspect the results.
## Requirements (from the PDF)
### Demo capabilities
- Upload files and **view the chunking/slicing result**.
- **Test recall** and inspect recall input parameters and return parameters.
### Supported document types
- **Office docs:** pdf, doc, docx, ppt, pptx, wps, ppsx
- **Tabular / structured:** xlsx, xls, csv, md, txt, html, json, xml, log
- **Images:** jpg, png, jpeg, bmp, gif
### Slicing rules
- **Table documents:** default split **or** split-by-row.
- **Other documents:** default split **or** universal-identifier split.
- **Tables inside a slice:** rendered as **Markdown**.
### Expected effects — basic
1. Images stay in their original position in the rendered text — not lost or relocated to another chunk.
2. Table content is extracted correctly and converted to Markdown.
3. Slicing is structurally aware: for Word, all content under the same heading stays within one chunk.
### Expected effects — advanced
1. Images/graphics are extracted as images **and** their text is OCR-recognized and converted to text.
2. PDF text laid out in horizontal columns is read back in the correct reading order (not jumbled).
## Environment notes
- **OS:** Windows 11. Shell is **bash** (Git Bash / MSYS2) — use Unix syntax (`/dev/null`, forward slashes), not PowerShell/CMD.
- **Python:** available at `/d/conda/python` (a conda environment). Relevant libraries **already installed** and useful for this project: `PyMuPDF` (fitz), `pdfplumber`, `pdfminer.six`, `pypdf`/`PyPDF2`, `pypdfium2`, `pikepdf`, `pdf2image`.
- **Reading the requirements PDF:** the file is image-heavy. `pdftotext -enc UTF-8` (at `/mingw64/bin/pdftotext`) extracts the little body text present but **misses the embedded diagrams** that the PDF references with "如下图" ("as shown below"). `pdftoppm` (image rendering) is **not** installed, so to view the diagrams use PyMuPDF from Python instead, e.g. `fitz.open(path)[page].get_pixmap()`.
- **Not a git repository.** Do not assume `git` workflows; initialize one only if asked.
## Working in this repo
- Before adding a chunking/slicing behavior, re-check the PDF's "期望效果" section above — the heading-awareness rule for Word and the image-position-preservation rule are the two most likely to be violated by naive splitters.
- The folder is named "RAG-cut" — chunking quality is the headline deliverable, recall is secondary. Prioritize ingestion/parsing fidelity over retrieval sophistication.