# AGENTS.md This file provides guidance to Codex (Codex.ai/code) when working with code in this repository. ## Project status This is a **greenfield project** — as of 2026-07-06 there is no code, no build system, and no git history yet. The only artifact is the requirements document `docx/自研搭建AI助手知识库.pdf` ("Self-built AI Assistant Knowledge Base"). Treat this file as the source of truth for requirements until code supersedes it. ## What this project is A self-hosted **AI assistant knowledge base demo** (RAG system) whose primary goal is **high-quality document chunking/slicing**; recall quality is secondary. The driving problem: an existing knowledge-base feature ("AFE") slices poorly — it drops images and tables from word/pdf/excel documents, and its rules forcibly split a single passage across separate chunks. This project rebuilds the ingestion/slicing pipeline and exposes a demo UI to inspect the results. ## Requirements (from the PDF) ### Demo capabilities - Upload files and **view the chunking/slicing result**. - **Test recall** and inspect recall input parameters and return parameters. ### Supported document types - **Office docs:** pdf, doc, docx, ppt, pptx, wps, ppsx - **Tabular / structured:** xlsx, xls, csv, md, txt, html, json, xml, log - **Images:** jpg, png, jpeg, bmp, gif ### Slicing rules - **Table documents:** default split **or** split-by-row. - **Other documents:** default split **or** universal-identifier split. - **Tables inside a slice:** rendered as **Markdown**. ### Expected effects — basic 1. Images stay in their original position in the rendered text — not lost or relocated to another chunk. 2. Table content is extracted correctly and converted to Markdown. 3. Slicing is structurally aware: for Word, all content under the same heading stays within one chunk. ### Expected effects — advanced 1. Images/graphics are extracted as images **and** their text is OCR-recognized and converted to text. 2. PDF text laid out in horizontal columns is read back in the correct reading order (not jumbled). ## Environment notes - **OS:** Windows 11. Shell is **bash** (Git Bash / MSYS2) — use Unix syntax (`/dev/null`, forward slashes), not PowerShell/CMD. - **Python:** available at `/d/conda/python` (a conda environment). Relevant libraries **already installed** and useful for this project: `PyMuPDF` (fitz), `pdfplumber`, `pdfminer.six`, `pypdf`/`PyPDF2`, `pypdfium2`, `pikepdf`, `pdf2image`. - **Reading the requirements PDF:** the file is image-heavy. `pdftotext -enc UTF-8` (at `/mingw64/bin/pdftotext`) extracts the little body text present but **misses the embedded diagrams** that the PDF references with "如下图" ("as shown below"). `pdftoppm` (image rendering) is **not** installed, so to view the diagrams use PyMuPDF from Python instead, e.g. `fitz.open(path)[page].get_pixmap()`. - **Not a git repository.** Do not assume `git` workflows; initialize one only if asked. ## Working in this repo - Before adding a chunking/slicing behavior, re-check the PDF's "期望效果" section above — the heading-awareness rule for Word and the image-position-preservation rule are the two most likely to be violated by naive splitters. - The folder is named "RAG-cut" — chunking quality is the headline deliverable, recall is secondary. Prioritize ingestion/parsing fidelity over retrieval sophistication.