3.3 KiB
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Project status
This is a greenfield project — as of 2026-07-06 there is no code, no build system, and no git history yet. The only artifact is the requirements document docx/自研搭建AI助手知识库.pdf ("Self-built AI Assistant Knowledge Base"). Treat this file as the source of truth for requirements until code supersedes it.
What this project is
A self-hosted AI assistant knowledge base demo (RAG system) whose primary goal is high-quality document chunking/slicing; recall quality is secondary. The driving problem: an existing knowledge-base feature ("AFE") slices poorly — it drops images and tables from word/pdf/excel documents, and its rules forcibly split a single passage across separate chunks. This project rebuilds the ingestion/slicing pipeline and exposes a demo UI to inspect the results.
Requirements (from the PDF)
Demo capabilities
- Upload files and view the chunking/slicing result.
- Test recall and inspect recall input parameters and return parameters.
Supported document types
- Office docs: pdf, doc, docx, ppt, pptx, wps, ppsx
- Tabular / structured: xlsx, xls, csv, md, txt, html, json, xml, log
- Images: jpg, png, jpeg, bmp, gif
Slicing rules
- Table documents: default split or split-by-row.
- Other documents: default split or universal-identifier split.
- Tables inside a slice: rendered as Markdown.
Expected effects — basic
- Images stay in their original position in the rendered text — not lost or relocated to another chunk.
- Table content is extracted correctly and converted to Markdown.
- Slicing is structurally aware: for Word, all content under the same heading stays within one chunk.
Expected effects — advanced
- Images/graphics are extracted as images and their text is OCR-recognized and converted to text.
- PDF text laid out in horizontal columns is read back in the correct reading order (not jumbled).
Environment notes
- OS: Windows 11. Shell is bash (Git Bash / MSYS2) — use Unix syntax (
/dev/null, forward slashes), not PowerShell/CMD. - Python: available at
/d/conda/python(a conda environment). Relevant libraries already installed and useful for this project:PyMuPDF(fitz),pdfplumber,pdfminer.six,pypdf/PyPDF2,pypdfium2,pikepdf,pdf2image. - Reading the requirements PDF: the file is image-heavy.
pdftotext -enc UTF-8(at/mingw64/bin/pdftotext) extracts the little body text present but misses the embedded diagrams that the PDF references with "如下图" ("as shown below").pdftoppm(image rendering) is not installed, so to view the diagrams use PyMuPDF from Python instead, e.g.fitz.open(path)[page].get_pixmap(). - Not a git repository. Do not assume
gitworkflows; initialize one only if asked.
Working in this repo
- Before adding a chunking/slicing behavior, re-check the PDF's "期望效果" section above — the heading-awareness rule for Word and the image-position-preservation rule are the two most likely to be violated by naive splitters.
- The folder is named "RAG-cut" — chunking quality is the headline deliverable, recall is secondary. Prioritize ingestion/parsing fidelity over retrieval sophistication.