49 lines
3.3 KiB
Markdown
49 lines
3.3 KiB
Markdown
# AGENTS.md
|
|||
|
|
|
||
|
|
This file provides guidance to Codex (Codex.ai/code) when working with code in this repository.
|
||
|
|
|
||
|
|
## Project status
|
||
|
|
|
||
|
|
This is a **greenfield project** — as of 2026-07-06 there is no code, no build system, and no git history yet. The only artifact is the requirements document `docx/自研搭建AI助手知识库.pdf` ("Self-built AI Assistant Knowledge Base"). Treat this file as the source of truth for requirements until code supersedes it.
|
||
|
|
|
||
|
|
## What this project is
|
||
|
|
|
||
|
|
A self-hosted **AI assistant knowledge base demo** (RAG system) whose primary goal is **high-quality document chunking/slicing**; recall quality is secondary. The driving problem: an existing knowledge-base feature ("AFE") slices poorly — it drops images and tables from word/pdf/excel documents, and its rules forcibly split a single passage across separate chunks. This project rebuilds the ingestion/slicing pipeline and exposes a demo UI to inspect the results.
|
||
|
|
|
||
|
|
## Requirements (from the PDF)
|
||
|
|
|
||
|
|
### Demo capabilities
|
||
|
|
- Upload files and **view the chunking/slicing result**.
|
||
|
|
- **Test recall** and inspect recall input parameters and return parameters.
|
||
|
|
|
||
|
|
### Supported document types
|
||
|
|
- **Office docs:** pdf, doc, docx, ppt, pptx, wps, ppsx
|
||
|
|
- **Tabular / structured:** xlsx, xls, csv, md, txt, html, json, xml, log
|
||
|
|
- **Images:** jpg, png, jpeg, bmp, gif
|
||
|
|
|
||
|
|
### Slicing rules
|
||
|
|
- **Table documents:** default split **or** split-by-row.
|
||
|
|
- **Other documents:** default split **or** universal-identifier split.
|
||
|
|
- **Tables inside a slice:** rendered as **Markdown**.
|
||
|
|
|
||
|
|
### Expected effects — basic
|
||
|
|
1. Images stay in their original position in the rendered text — not lost or relocated to another chunk.
|
||
|
|
2. Table content is extracted correctly and converted to Markdown.
|
||
|
|
3. Slicing is structurally aware: for Word, all content under the same heading stays within one chunk.
|
||
|
|
|
||
|
|
### Expected effects — advanced
|
||
|
|
1. Images/graphics are extracted as images **and** their text is OCR-recognized and converted to text.
|
||
|
|
2. PDF text laid out in horizontal columns is read back in the correct reading order (not jumbled).
|
||
|
|
|
||
|
|
## Environment notes
|
||
|
|
|
||
|
|
- **OS:** Windows 11. Shell is **bash** (Git Bash / MSYS2) — use Unix syntax (`/dev/null`, forward slashes), not PowerShell/CMD.
|
||
|
|
- **Python:** available at `/d/conda/python` (a conda environment). Relevant libraries **already installed** and useful for this project: `PyMuPDF` (fitz), `pdfplumber`, `pdfminer.six`, `pypdf`/`PyPDF2`, `pypdfium2`, `pikepdf`, `pdf2image`.
|
||
|
|
- **Reading the requirements PDF:** the file is image-heavy. `pdftotext -enc UTF-8` (at `/mingw64/bin/pdftotext`) extracts the little body text present but **misses the embedded diagrams** that the PDF references with "如下图" ("as shown below"). `pdftoppm` (image rendering) is **not** installed, so to view the diagrams use PyMuPDF from Python instead, e.g. `fitz.open(path)[page].get_pixmap()`.
|
||
|
|
- **Not a git repository.** Do not assume `git` workflows; initialize one only if asked.
|
||
|
|
|
||
|
|
## Working in this repo
|
||
|
|
|
||
|
|
- Before adding a chunking/slicing behavior, re-check the PDF's "期望效果" section above — the heading-awareness rule for Word and the image-position-preservation rule are the two most likely to be violated by naive splitters.
|
||
|
|
- The folder is named "RAG-cut" — chunking quality is the headline deliverable, recall is secondary. Prioritize ingestion/parsing fidelity over retrieval sophistication.
|