diff --git a/IMPACT_ANALYSIS.md b/IMPACT_ANALYSIS.md index b2ae197..6c638d2 100644 --- a/IMPACT_ANALYSIS.md +++ b/IMPACT_ANALYSIS.md @@ -1,3 +1,79 @@ +# Impact Analysis Report — CJK 字面量校验(CASE 展示标签) + +## 1. 改动概览 + +- **背景与目标**:修复「CASE … THEN/ELSE 中使用中文状态标签」被 `check_no_cjk_in_sql_string_literals` 全文扫描误判为非法,导致合法查询反复重试仍失败的问题。规则本意是禁止在 **WHERE/HAVING/ON/CASE 条件** 中用中文与代码列比对。 +- **涉及模块**:`backend/utils/validators.py`(CJK 校验改为基于 sqlglot AST + 解析失败时回退全文扫描)、`backend/agents/orchestrator.py`(传入 `dialect`、T-SQL 硬性说明与注释对齐)、`backend/config/prompts.py`(2b 条款与校验语义一致)。 +- **改动类型**:缺陷修复(验证逻辑与提示词澄清)。 + +## 2. 方法级改动 + +| 位置 | 变更 | +|------|------| +| `check_no_cjk_in_sql_string_literals` | 解析成功时仅当 CJK 出现在 `Where`/`Having`/`Join.on` 或 `Case` 分支的 **WHEN 条件**(`If.this`)中报错;`Case` 的 **THEN/ELSE** 与纯 `SELECT` 展示字面量允许。支持 `Literal` 与 `National`(`N'…'`)。解析失败时回退旧版全文扫描(偏严)。 | +| `orchestrator._validate_sql` | 调用 `check_no_cjk_in_sql_string_literals(sql, dialect=dialect)`。 | + +## 3. 调用方与影响范围 + +- **调用方**:`_validate_sql` 内唯一调用点。 +- **破坏性变更**:否。行为变更:此前会失败的「仅 CASE/SELECT 含中文标签」现为通过;`WHERE col = N'中文'` 等仍失败。 + +## 4. 风险与回滚 + +- **风险级别**:低。解析异常时仍走严格全文扫描。 +- **回滚**:还原 `validators.py` / 相关 prompt 与 orchestrator 片段。 + +**回滚方式是否简单**:是。 + +## 5. 验证与测试 + +- 已执行:脚本用例 — `CASE … THEN '\u8d85\u989d'` 通过;`WHERE x = N'\u8d85\u989d'` 失败。 + +## 6. 配置变更 + +- 无。 + +--- + +# Impact Analysis Report — Prompt:无时间表述则不擅自加日期条件 + +## 1. 改动概览 + +- **背景与目标**:强化 Text2SQL 提示词,避免用户未提任何时间时模型在 `WHERE` 中私自添加日期过滤;Few-shot 黄金适配时若当前问题比范例少时间条件,应去掉多余日期条件。 +- **涉及模块**:`backend/config/prompts.py`(`SQL_GENERATOR_*`、`GOLDEN_SQL_ADAPT_*`)、`backend/agents/orchestrator.py`(T-SQL 用户侧硬性要求追加一句)。 +- **改动类型**:配置/提示词优化(行为变更:模型被约束为少加臆测日期条件;**非** API 签名变更)。 + +## 2. 方法级改动 + +| 位置 | 变更 | +|------|------| +| `SQL_GENERATOR_SYSTEM` | 新增硬性条款 **1b**(无时间表述则不加日期条件);收紧第 **2** 条中「未给日期」与日期列的说明;推断指南增加**前提**行;黄金范例括号内补充「完全未提时间勿照抄日期」。 | +| `SQL_GENERATOR_USER` | 末尾一句指向 **1b**。 | +| `GOLDEN_SQL_ADAPT_SYSTEM` | 第 **2** 条补充:当前问题无时间要求时去掉标准答案中无关日期过滤。 | +| `orchestrator` `_generate_sql` / `_generate_sql_golden_adapt` | T-SQL 硬性要求字符串追加与 **1b** 一致的一句提醒。 | + +## 3. 调用方与影响范围 + +- **调用方**:所有经 `Text2SQLOrchestrator` 走 SQL 生成的路径(含黄金适配分支)。 +- **破坏性变更**:否(仅 prompt 与 user 附加说明文本变化)。 + +## 4. 风险与回滚 + +- **风险级别**:低。可能使「未说时间」类问题返回更宽结果集(符合预期);若业务依赖模型以前「默认加近期」的隐式行为,需改为在用户问题或后端注入默认时间窗。 +- **回滚**:还原上述文件相关 diff。 + +**回滚方式是否简单**:是。 + +## 5. 验证与测试 + +- 建议手工:问题不含任何时间词 → 生成 SQL 应无新增日期列条件(除非 Schema/问题语义强制);含「今天」→ 仍可用 `GETDATE()`。 + +## 6. 配置变更 + +- 无。 + +--- + # Impact Analysis Report — 会话上文注入 Text2SQL ## 1. 改动概览 @@ -672,3 +748,109 @@ ## 6. 验证与测试 - 已执行:`python -m py_compile api_server.py`。 + +--- + +# Impact Analysis Report — API 回包增加 data.sql / data.explanation 便捷字段 + +## 1. 改动概览 + +- **背景与目标**:前端需要在响应 `data` 内直接读取 `{"sql":"..."}`(SQL 可含 `--` 注释且无 ``` 围栏),并单独读取自然语言解释字段。 +- **涉及模块**:`api_server.py`(`NLChatSuccessData`、`_nl_dict_from_generation`)。 +- **改动类型**:向后兼容的字段补充(不移除/不改名原有字段)。 + +## 2. 方法级改动 + +| 位置 | 变更 | +|------|------| +| `NLChatSuccessData` | 新增可选字段 `sql`、`explanation`(便捷读取)。 | +| `_nl_dict_from_generation` | 在原有 `intent/branch_result` 之外补充 `payload["sql"]` 与 `payload["explanation"]`;`explanation` 优先取 `sql_delivery_message`,否则回退到 `sql_explain`。 | + +## 3. 调用方与影响范围 + +- **调用方**:`/g3sb/api/nl/chat` 与 `/g3sb/api/nl/chat/stream` 的结束包 `data`。 +- **破坏性变更**:否(仅新增字段;原 `branch_result.sql` 等保持不变)。 + +## 4. 风险与回滚 + +- **风险级别**:低(新增字段)。 +- **回滚**:回退本改动提交即可。 + +**回滚方式是否简单**:是。 + +## 5. 验证与测试 + +- 已执行:`python -m py_compile api_server.py`。 + +--- + +# Impact Analysis Report — SSE 流式 SQL 说明与可选 `{"sql":...}` 片段(可开关) + +## 1. 改动概览 + +- **背景与目标**:流式 `/g3sb/api/nl/chat/stream` 需要给客户端展示 SQL(LLM 原生 delta)并在必要时提供结构化 SQL。由于当前前端会直接拼接展示 delta 且保留 ``,因此结构化 `` 片段默认不补发,避免同一次流里出现两段可见 SQL;如有需要可通过环境变量开关启用。 +- **涉及模块**:`api_server.py`(`_chat_stream_events`)。 +- **改动类型**:向后兼容增强(流式末尾追加内容;并过滤 LLM 可能自行输出的 `...`,避免前端出现重复 SQL/重复结构化片段)。 + +## 2. 方法级改动分析 + +| 位置 | 变更 | +|------|------| +| `_chat_stream_events` | 透传 LLM 的 `sql_gen` delta 用于展示;过滤 LLM 自吐的 `...` 段。流结束时结构化 `{"sql":...}` **默认不补发**,仅在 `SSE_APPEND_SQL_DATA_TAG=true` 时启用补发,避免前端拼接展示时出现重复 SQL。 | + +## 3. 调用方与影响范围分析 + +- **调用方**:前端 SSE 消费逻辑(`stage="sql_gen"`、`stream_kind="content"`)。 +- **影响**: + - **展示**:若前端直接展示原始流文本,将额外看到「SQL 说明」与 `...`;但通常前端会对 `...` 做隐藏/抽取,不影响页面展示。 + - **解析**:前端可在流末尾稳定抽取 SQL(而不依赖临时拼接或额外请求非流式接口)。 +- **破坏性变更**:否(仅追加;不更改原有字段与 SSE 事件结构)。 + +## 4. 风险与回滚 + +- **风险级别**:低(开关打开时,若前端未过滤 ``,可能展示出标签;默认关闭避免此问题)。 +- **回滚**:默认即不补发;如需恢复补发仅需设置 `SSE_APPEND_SQL_DATA_TAG=true`,或回退代码提交。 + +**回滚方式是否简单**:是。 + +## 5. 验证与测试 + +- 已验证:`python -m py_compile api_server.py`;默认不再补发 `{"sql":...}`,且不会透传 LLM 自行生成的 `...` 段;需要结构化输出时可设置 `SSE_APPEND_SQL_DATA_TAG=true` 验证补发行为。 + +--- + +# Impact Analysis Report — api_server 精简未使用请求/查询参数 + +## 1. 改动概览 + +- **背景与目标**:删除 `api_server.py` 中未被业务逻辑使用的请求字段与路由形参,减少噪音;stub 接口不再声明从不读取的请求体模型。 +- **涉及模块**:`api_server.py`。 +- **改动类型**:重构(对外行为基本不变)。 + +## 2. 方法级改动 + +| 位置 | 变更 | +|------|------| +| `NLChatRequest` | 移除字段 `taskId`(仓库内无读取)。 | +| `SqlExecuteBody` | 删除模型;`POST /g3sb/api/nl/sql/execute` 无请求体形参(仍返回 501)。 | +| `POST /g3sb/api/nl/operation-logs` | 移除未使用的 `Body` 形参(仍返回 200)。 | +| `GET /g3sb/api/nl/operation-logs` | 仅保留 `limit`、`offset`;其余查询参数从签名中移除。 | +| `GET /g3sb/api/nl/knowledge-docs/uploads` | 移除未使用的 `doc_type`,保留 `limit`、`offset`。 | +| `lifespan`、全局异常处理 | 未使用的 `app`/`_req` 形参改为 `_app` / `_`(语义不变)。 | + +## 3. 调用方与影响范围 + +- **调用方**:`ai-g3sb-backman-web` 的 `nlClient.ts` 仍可对上述 GET 附带更多 query、对 POST 仍发送 JSON;FastAPI 对未声明的 query 通常忽略,不声明的 body 仍会被接收但不解析。 +- **OpenAPI**:`/sql/execute` 与 `POST /operation-logs` 的文档中可能不再展示请求体 schema(与当前「不读 body」实现一致)。 +- **破坏性变更**:否(若客户端依赖 OpenAPI 生成且严格要求 `taskId` 出现在 schema,仅文档层面变化;运行时多传 `taskId` 仍被 Pydantic 忽略)。 + +## 4. 风险与回滚 + +- **风险级别**:低。 +- **回滚**:回退本改动对应提交即可。 + +**回滚方式是否简单**:是。 + +## 5. 验证与测试 + +- 已执行:`python -m py_compile api_server.py`。 diff --git a/TASK_SUMMARY.md b/TASK_SUMMARY.md index e00b69d..0337010 100644 --- a/TASK_SUMMARY.md +++ b/TASK_SUMMARY.md @@ -45,6 +45,9 @@ ## 流式输出(补充) - **`/g3sb/api/nl/chat/stream`**:`chat` 与 `sql_gen` 阶段的正文改为多段 SSE 分片推送(默认每段约 64 字符,可用环境变量 `SSE_STREAM_CHUNK_CHARS` 调整),与前端 `onDelta` 累加逻辑一致。 +- **重复 SQL 修复(补充)**:为避免同一次 SSE 响应中出现两段 SQL(LLM 自吐 `...` + 服务端末尾补发 `{"sql":...}`),后端在透传 LLM delta 时会过滤掉任何 `...` 段,仅保留流结束时的标准结构化 `` 片段供前端解析。 + - **现状约束**:由于当前前端会直接拼接展示 delta 且保留 ``,流式接口默认**不再补发**结构化 `{"sql":...}`,避免“SQL + 可见 ``”造成重复展示。 + - **如需补发**:设置环境变量 `SSE_APPEND_SQL_DATA_TAG=true` 才会在流末尾补发结构化 ``(仅建议给会过滤 `` 的客户端使用)。 ## 后续事项 diff --git a/__pycache__/api_server.cpython-312.pyc b/__pycache__/api_server.cpython-312.pyc index b4a19b4..1b9dc1d 100644 Binary files a/__pycache__/api_server.cpython-312.pyc and b/__pycache__/api_server.cpython-312.pyc differ diff --git a/api_server.py b/api_server.py index 1736769..8047c84 100644 --- a/api_server.py +++ b/api_server.py @@ -32,7 +32,7 @@ from utils.repo_logging import configure_text2sql_api_logging configure_text2sql_api_logging(_REPO_DIR) -from main import setup_environment, load_schema, create_orchestrator, resolve_sql_dialect +from bootstrap import setup_environment, load_schema, create_orchestrator, resolve_sql_dialect from agents.orchestrator import GenerationResult from nl_lite_store import lite_nl_store from utils.dialog_classifier import DialogIntent, classify_dialog diff --git a/backend/__pycache__/__init__.cpython-312.pyc b/backend/__pycache__/__init__.cpython-312.pyc new file mode 100644 index 0000000..51ac340 Binary files /dev/null and b/backend/__pycache__/__init__.cpython-312.pyc differ diff --git a/backend/__pycache__/bootstrap.cpython-312.pyc b/backend/__pycache__/bootstrap.cpython-312.pyc new file mode 100644 index 0000000..6fcca77 Binary files /dev/null and b/backend/__pycache__/bootstrap.cpython-312.pyc differ diff --git a/backend/__pycache__/main.cpython-312.pyc b/backend/__pycache__/main.cpython-312.pyc index f4d4c9a..ee056d4 100644 Binary files a/backend/__pycache__/main.cpython-312.pyc and b/backend/__pycache__/main.cpython-312.pyc differ diff --git a/backend/agents/__pycache__/orchestrator.cpython-312.pyc b/backend/agents/__pycache__/orchestrator.cpython-312.pyc index 081ed23..ef66838 100644 Binary files a/backend/agents/__pycache__/orchestrator.cpython-312.pyc and b/backend/agents/__pycache__/orchestrator.cpython-312.pyc differ diff --git a/backend/agents/orchestrator.py b/backend/agents/orchestrator.py index 99eef6d..16cf346 100644 --- a/backend/agents/orchestrator.py +++ b/backend/agents/orchestrator.py @@ -302,6 +302,28 @@ class Text2SQLOrchestrator: logger.info("对手方/经纪商问题:优先纳入 %s,调整后选表:%s", present, merged) return merged[:max_tables] + _VC_USER_ACCESSIBLE_FUNCTION = "VCUserAccessibleFunction" + + def _prioritize_vc_user_accessible_function( + self, relevant_tables: List[str], max_tables: int = 5 + ) -> List[str]: + """ + 若 Schema 中存在 VCUserAccessibleFunction,则置于选表列表最前,便于模型先根据 + FunctionID / DatabaseView 等列定位业务视图,再关联其余表生成 SQL。 + """ + vc = self._VC_USER_ACCESSIBLE_FUNCTION + if not self.schema_manager.get_table(vc): + return relevant_tables[:max_tables] + + rest = [t for t in relevant_tables if t != vc] + merged = [vc] + rest + logger.info( + "已优先纳入目录视图 %s(置于选表前列),当前选表:%s", + vc, + merged[:max_tables], + ) + return merged[:max_tables] + def _expand_relations(self, table_names: List[str]) -> List[str]: """ 外键扩展:自动添加关联表 @@ -413,8 +435,9 @@ class Text2SQLOrchestrator: "**禁止** `CURDATE()`、`NOW()`、`CURRENT_DATE`(MySQL)。" "条件请使用 T-SQL 惯用写法(例如 IS NOT NULL)。" "排版仍须遵守:关键字大写、SELECT 每列一行缩进、WHERE 续行以 AND 开头、PascalCase 英文别名。" - "\n**禁止**在单引号字符串字面量或 `N'…'` 中出现任何中日韩文字;" - "业务中文须映射为 Schema 注释中的代码或通过维表 JOIN,勿写 `= '过户费'` 这类比对。" + "\n在 `WHERE`/`HAVING`/`JOIN ON` 及 `CASE/WHEN` 的**条件**中,**禁止**用含中日韩的 `'`/`N'…'` 字面量与代码型列比对;" + "业务中文须映射为 Schema 注释中的代码或通过维表 JOIN(勿写 `= '过户费'`)。`SELECT` 或 `CASE … THEN/ELSE` 的展示用中文标签允许。" + "若用户问题与对话上文均未要求按时间筛选,不得在 WHERE 中擅自添加日期列条件(与系统提示 1b 一致)。" ) if validation_feedback: @@ -577,8 +600,9 @@ class Text2SQLOrchestrator: "**禁止** `CURDATE()`、`NOW()`、`CURRENT_DATE`(MySQL)。" "条件请使用 T-SQL 惯用写法(例如 IS NOT NULL)。" "排版仍须遵守:关键字大写、SELECT 每列一行缩进、WHERE 续行以 AND 开头、PascalCase 英文别名。" - "\n**禁止**在单引号字符串字面量或 `N'…'` 中出现任何中日韩文字;" - "业务中文须映射为 Schema 注释中的代码或通过维表 JOIN,勿写 `= '过户费'` 这类比对。" + "\n在 `WHERE`/`HAVING`/`JOIN ON` 及 `CASE/WHEN` 的**条件**中,**禁止**用含中日韩的 `'`/`N'…'` 字面量与代码型列比对;" + "业务中文须映射为 Schema 注释中的代码或通过维表 JOIN(勿写 `= '过户费'`)。`SELECT` 或 `CASE … THEN/ELSE` 的展示用中文标签允许。" + "若用户问题与对话上文均未要求按时间筛选,不得在 WHERE 中擅自添加日期列条件(与系统提示 1b 一致)。" ) if validation_feedback: @@ -659,9 +683,9 @@ class Text2SQLOrchestrator: if not danger_ok: errors.extend(danger_errors) - # T-SQL:禁止中文等业务词出现在字符串字面量(如 FeeNatureID = '过户费') + # T-SQL:禁止中文出现在 WHERE/HAVING/ON/CASE 条件等比对语境(展示用 CASE THEN/ELSE 允许) if dialect == "tsql": - cjk_ok, cjk_errors = check_no_cjk_in_sql_string_literals(sql) + cjk_ok, cjk_errors = check_no_cjk_in_sql_string_literals(sql, dialect=dialect) if not cjk_ok: errors.extend(cjk_errors) @@ -845,6 +869,9 @@ class Text2SQLOrchestrator: relevant_tables = self._prioritize_broker_tables( linker_question, relevant_tables ) + relevant_tables = self._prioritize_vc_user_accessible_function( + relevant_tables + ) # 1.3 外键扩展 expanded_tables = self._expand_relations(relevant_tables) @@ -856,6 +883,14 @@ class Text2SQLOrchestrator: include_columns=True, max_columns_per_table=20 ) + if self._VC_USER_ACCESSIBLE_FUNCTION in expanded_tables: + filtered_schema_str += ( + "\n\n【选表提示】已包含视图 " + + self._VC_USER_ACCESSIBLE_FUNCTION + + "(列含 UserID、FunctionID、Category、Name、DatabaseView)。" + "生成 SQL 时可先通过该视图用 DatabaseView / FunctionID 等定位目标业务视图或功能," + "再与 Schema 中其余表做 JOIN 或子查询;若问题已明确具体表名,可直接查询该表。" + ) logger.info(f" 选中表:{relevant_tables},扩展后:{expanded_tables}") else: # 重试:在子 Schema 中并入「上次失败 SQL」实际引用到的表,并对齐程序校验与生成上下文 diff --git a/backend/bootstrap.py b/backend/bootstrap.py new file mode 100644 index 0000000..63005a1 --- /dev/null +++ b/backend/bootstrap.py @@ -0,0 +1,215 @@ +""" +Text2SQL 共享启动逻辑:环境检查、Schema 加载、Orchestrator 构造。 + +供 `main` CLI 与 `api_server` 复用,避免 API 层依赖 CLI 入口模块。 +""" + +from __future__ import annotations + +import logging +import os +import sys +from pathlib import Path +from typing import Any, Optional + +from dotenv import load_dotenv + +logger = logging.getLogger(__name__) + + +def _repo_root() -> Path: + """仓库根目录(含 data/、.env、api_server.py 的目录)。""" + return Path(__file__).resolve().parent.parent + + +def load_project_env() -> None: + """加载项目根目录 .env,供后续 os.getenv 使用。""" + load_dotenv(_repo_root() / ".env") + + +def resolve_sql_dialect(name: str) -> str: + """CLI / 配置中的方言别名统一为 sqlglot 方言名(SQL Server -> tsql)。""" + n = (name or "sqlserver").lower().strip() + if n in ("sqlserver", "mssql"): + return "tsql" + return n + + +def setup_environment() -> bool: + """环境检查""" + load_project_env() + ms_key = os.getenv("MODELSCOPE_API_KEY", "").strip() + oa_key = os.getenv("OPENAI_API_KEY", "").strip() + oa_key_ok = oa_key and not ( + oa_key.startswith("http://") or oa_key.startswith("https://") + ) + if ms_key: + pass # ModelScope:BASE_URL / MODEL 有默认值,仅需 KEY + elif oa_key_ok: + if not os.getenv("OPENAI_EMBEDDING_MODEL", "").strip(): + logger.warning("未设置 OPENAI_EMBEDDING_MODEL(远程 Embedding 模型名)") + return False + else: + if not os.getenv("DASHSCOPE_API_KEY", "").strip(): + logger.warning( + "未设置 MODELSCOPE_API_KEY、" + "OPENAI_API_KEY+OPENAI_EMBEDDING_MODEL 或 DASHSCOPE_API_KEY" + ) + return False + base = ( + os.getenv("DASHSCOPE_BASE_URL") or os.getenv("DASHSCOPE_base_url", "") + ).strip() + if not base: + logger.warning("未设置 DASHSCOPE_BASE_URL(或 DASHSCOPE_base_url)") + return False + if not os.getenv("DASHSCOPE_MODEL", "").strip(): + logger.warning("未设置 DASHSCOPE_MODEL") + return False + + # 检查Schema文件(支持相对路径和绝对路径) + schema_path_str = os.getenv( + "SCHEMA_PATH", "./data/schemas/G3SB_MCDataDictionary_table_structure.json" + ) + schema_path = Path(schema_path_str) + + # 如果是相对路径,尝试从多个位置查找 + if not schema_path.is_absolute(): + # 尝试1: PyInstaller 临时目录(单文件模式) + if getattr(sys, "frozen", False) and hasattr(sys, "_MEIPASS"): + meipass_schema = Path(sys._MEIPASS) / schema_path_str + if meipass_schema.exists(): + schema_path = meipass_schema + + # 尝试2: 当前工作目录 + if not schema_path.exists(): + schema_path = Path.cwd() / schema_path_str + + # 尝试3: 仓库根目录(本文件位于 backend/) + if not schema_path.exists(): + script_dir = Path(__file__).resolve().parent.parent + schema_path = script_dir / schema_path_str + + # 尝试4: 可执行文件所在目录 + if not schema_path.exists() and getattr(sys, "frozen", False): + exe_dir = Path(sys.executable).parent + schema_path = exe_dir / schema_path_str + + if not schema_path.exists(): + logger.warning(f"Schema文件不存在: {schema_path}") + logger.info("请将G3SB Schema文件放置在 ./data/schemas/ 目录") + return False + + # 检查 LLM Key(DeepSeek / OpenAI 可切换) + llm_sc = (os.getenv("LLM_SERVICE_CODE") or "").strip().lower() + if llm_sc and llm_sc not in ("deepseek", "openai"): + logger.warning("未知 LLM_SERVICE_CODE=%r(仅支持 deepseek/openai)", llm_sc) + return False + if llm_sc == "openai": + if not (os.getenv("OPENAI_API_KEY") or "").strip(): + logger.warning("LLM_SERVICE_CODE=openai 但 OPENAI_API_KEY 未设置") + return False + else: + if not (os.getenv("DEEPSEEK_API_KEY") or "").strip(): + logger.warning("环境变量 DEEPSEEK_API_KEY 未设置(默认 LLM_SERVICE_CODE=deepseek)") + logger.info("请在 .env 文件中配置,或 export DEEPSEEK_API_KEY=your_key") + return False + + logger.info(f"[OK] 环境检查通过") + logger.info(f" - Schema: {schema_path}") + logger.info( + " - LLM: %s", + (llm_sc or "deepseek(auto)"), + ) + + return True + + +def _default_g3sb_meta_path(structure_path: str) -> Optional[str]: + """若存在与 table_structure 同名的 table_meta 文件则返回其路径。""" + p = Path(structure_path) + if "table_structure" not in p.name: + return None + cand = p.parent / p.name.replace("table_structure", "table_meta") + return str(cand) if cand.is_file() else None + + +def load_schema(schema_path: str, schema_meta_path: Optional[str] = None): + """加载 Schema;G3SB structure JSON 会自动尝试配对 table_meta(可用 --schema-meta 指定)。""" + from schema.manager import SchemaManager + + meta = ( + schema_meta_path + if schema_meta_path is not None + else _default_g3sb_meta_path(schema_path) + ) + logger.info(f"加载Schema: {schema_path}") + if meta: + logger.info(f" 表注释(meta): {meta}") + + schema_mgr = SchemaManager.load_from_json( + schema_path, g3sb_meta_path=meta + ) + + stats = schema_mgr.get_statistics() + logger.info( + f"[OK] Schema加载完成: {stats['database']}, " + f"共{stats['total_tables']}张表, {stats['total_columns']}个字段" + ) + + return schema_mgr + + +def create_orchestrator(schema_mgr: Any, args: Any): + """创建编排器""" + from agents.orchestrator import Text2SQLOrchestrator + from llm.router import create_llm_client, resolve_llm_service_code + from llm.deepseek_client import DeepSeekConfig + + translate_en = os.getenv("TRANSLATE_EN_TO_ZH", "true").strip().lower() not in ( + "0", + "false", + "no", + "off", + ) + if getattr(args, "no_translate_en", False): + translate_en = False + + sc = resolve_llm_service_code() + if sc == "deepseek": + api_key = (args.api_key or os.getenv("DEEPSEEK_API_KEY") or "").strip() + if not api_key: + raise ValueError( + "未配置 DeepSeek API Key:请在 .env 中设置 DEEPSEEK_API_KEY," + "或使用命令行参数 --api-key" + ) + base_url = (os.getenv("DEEPSEEK_BASE_URL") or "https://api.deepseek.com").strip() + cfg = DeepSeekConfig( + api_key=api_key, + base_url=base_url, + model_name=args.model or "deepseek-chat", + temperature=args.temperature, + max_tokens=args.max_tokens, + ) + llm_client = create_llm_client("deepseek", **cfg.__dict__) + else: + # openai:完全由 OPENAI_* 决定;同时沿用 temperature/max_tokens 作为默认值覆盖 + llm_client = create_llm_client( + "openai", + temperature=args.temperature, + max_tokens=args.max_tokens, + ) + + orchestrator = Text2SQLOrchestrator( + schema_manager=schema_mgr, + llm_client=llm_client, + vector_db_path=args.vector_db, + max_retry=args.max_retry, + use_vector_search=not args.no_vector_search, + # Few-shot配置 + fewshot_enabled=not args.no_fewshot, + fewshot_top_k=args.fewshot_top_k, + fewshot_min_rating=args.fewshot_min_rating, + translate_english_to_zh=translate_en, + ) + + return orchestrator diff --git a/backend/config/__pycache__/prompts.cpython-312.pyc b/backend/config/__pycache__/prompts.cpython-312.pyc index 6d87c73..8cadc67 100644 Binary files a/backend/config/__pycache__/prompts.cpython-312.pyc and b/backend/config/__pycache__/prompts.cpython-312.pyc differ diff --git a/backend/config/prompts.py b/backend/config/prompts.py index 7d42f0b..ef073e3 100644 --- a/backend/config/prompts.py +++ b/backend/config/prompts.py @@ -55,8 +55,9 @@ SQL_GENERATOR_SYSTEM = """你是一个精通SQL的数据库专家,有10年以 **硬性约束(必须遵守)**: 1. **合理推断业务语义**:用户问题中的时间范围(如"2024年1月")、状态含义(如"活跃"对应Active)、常见业务默认值(如"当前"指近期),应根据Schema中的字段注释和常见业务逻辑进行合理推断并转化为WHERE条件;但禁止编造问题中未提及的过滤维度或指标。 -2. **禁止虚构值与占位符**:不得使用 `'[日期]'`、`TODO`、`xxx`、空泛占位等冒充具体字面量。若用户未给出具体日期、代码或 ID,应根据问题上下文推断合理值(如"2024年1月" → `ValueDate >= '2024-01-01' AND ValueDate < '2024-02-01'`),或使用Schema中常见的枚举值(如状态字段的`A/D/X`),**不要**留空或写占位符。 -2b. **禁止在 SQL 字符串字面量中出现中文(CJK)**:用户问题里的中文业务词(如「过户费」「未结算」「活跃」)**禁止**写成 `'…中文…'` 或 `N'…中文…'` 去和代码型列(如 `FeeNatureID`、`SettleStatus`、`State`)比较。必须根据 **Schema 字段注释** 写成库内真实**代码/单字母/数字**(如 `State = 'A'`、`SettleStatus = 'U'`);若业务词对应维表或码表,应 **JOIN 维表** 用其键列或英文名列过滤,**不得**用中文当字面量。 +1b. **无时间表述则不加日期条件(强制)**:若用户问题及对话上文**均未**出现任何可映射为**按时间筛选**的表述——包括但不限于:具体日历日期、年月/季度区间、「今天/昨日/本周/本月/本年/本季度」「最近N天/过去一周/过去一年」等相对时间——则 **不得**在 `WHERE`/`HAVING` 中**新增**对日期/时间类型列的过滤(例如 `TradeDate >= '...'`、`BETWEEN ... AND ...`、与 `CAST(GETDATE() AS DATE)` / `DATEADD` 结合的日期条件)。**禁止**以「防止结果集过大」「报表通常只看近期」「默认只查当年」等理由擅加日期窗。仅当用户**明确**提出时间要求、或问题语义**显式**指向某时段(如「2024年1月的销售额」「今天的成交」)时,才写对应日期条件;完全未提时间时,查询在日期维度上可为全表/全历史(仅受问题中**已出现**的非时间条件约束)。 +2. **禁止虚构值与占位符**:不得使用 `'[日期]'`、`TODO`、`xxx`、空泛占位等冒充具体字面量。若用户未给出具体日期、代码或 ID:**日期类**仅当问题里**已经**出现可映射的时间表述时,才按上文与「常见时间推断指南」写出具体区间或 `GETDATE()` 条件;若全文无任何时间表述,**不得**为凑条件而编造日期过滤(与 **1b** 一致)。**非日期类**(状态码、ID 等)仍可根据问题上下文与 Schema 枚举填写合理值,**不要**留空或写占位符。 +2b. **禁止在过滤/比对条件中使用中文(CJK)字面量**:在 `WHERE`/`HAVING`/`JOIN … ON` 以及 `CASE WHEN` 的**条件部分**,用户问题里的中文业务词(如「过户费」「未结算」「活跃」)**禁止**写成 `'…中文…'` 或 `N'…中文…'` 去和代码型列(如 `FeeNatureID`、`SettleStatus`、`State`)比较。必须根据 **Schema 字段注释** 写成库内真实**代码/单字母/数字**(如 `State = 'A'`、`SettleStatus = 'U'`);若业务词对应维表或码表,应 **JOIN 维表** 用其键列或英文名列过滤。**允许**在 `SELECT` 列表达式或 `CASE … THEN … ELSE …` 的**展示结果**中使用中文标签字符串(如状态说明),此类不属于「与代码列比对」。 3. **输出版式与别名风格(统一规范)**:除遵守目标方言语法外,SQL **排版与命名**须与下方「标准版式范例」一致: - **关键字**:`SELECT`、`FROM`、`JOIN`/`LEFT JOIN`、`ON`、`WHERE`、`AND`、`GROUP BY`、`ORDER BY`、`HAVING` 等使用**大写**。 - **换行与缩进**:`SELECT` 后换行;每个输出列**独占一行**,行首 **4 个空格**,列表达式之间用**行尾逗号**分隔(最后一列无逗号)。 @@ -86,6 +87,7 @@ SQL_GENERATOR_SYSTEM = """你是一个精通SQL的数据库专家,有10年以 - 负债/负数含义:许多余额字段负值表示负债(如LoanBalance) **常见时间/状态推断指南**(需结合Schema字段注释): +- **前提**:下列日期规则**仅当**用户问题或对话中**已出现**对应时间表述时适用;若完全未提时间,**不要**套用下列规则去加日期条件(见 **1b**)。 - "2024年1月" → `WHERE date_col >= '2024-01-01' AND date_col < '2024-02-01'` - "今天" / "当日" → `WHERE date_col >= CAST(GETDATE() AS DATE) AND date_col < DATEADD(DAY,1,CAST(GETDATE() AS DATE))` - "最近N天" → `WHERE date_col >= DATEADD(DAY, -N, CAST(GETDATE() AS DATE))` @@ -103,7 +105,7 @@ SQL_GENERATOR_SYSTEM = """你是一个精通SQL的数据库专家,有10年以 用户问题(示例):按对手方列出截至 2026-04-02 的所有未结算交易。 -(日期规则:用户问题里**已写出具体日期**时,在 WHERE 中写入相同字面量,例如 `CashSettleDate <= '2026-04-02'`;**未给出具体日期**但包含时间范围描述(如"2024年1月")时,应合理推断为日期区间,例如 `OrderDate >= '2024-01-01' AND OrderDate < '2024-02-01'`,禁止使用 `'[日期]'` 等占位符。) +(日期规则:用户问题里**已写出具体日期**时,在 WHERE 中写入相同字面量,例如 `CashSettleDate <= '2026-04-02'`;**未给出具体日期**但包含时间范围描述(如"2024年1月")时,应合理推断为日期区间,例如 `OrderDate >= '2024-01-01' AND OrderDate < '2024-02-01'`,禁止使用 `'[日期]'` 等占位符。**若用户完全未提及任何时间与时段**,则不要添加日期列条件,勿因本范例含日期而照抄日期过滤。) SQL(表名、字段名须与当前 Schema 一致;**以下版式、别名、JOIN/WHERE/GROUP BY/ORDER BY 结构为强制模板**): @@ -130,30 +132,41 @@ ORDER BY TotalUnsettledAmount DESC; **示例(版式与范例一致)**: 示例1 - 单表查询: -问题:查询账户ID为'ACC001'的账户余额 -Schema: MCAccount(AccountID, Name, AvailableBalance, MarketValue, MarginValue) +问题:查询账户ID为'M050013'的账户余额 +Schema: MCAccount(AccountID, AvailableBalance, AssetBalance, LiabilityBalance, MaximumAvailableBalance) SQL: SELECT - AccountID, - Name AS AccountName, + RTRIM(AccountID) AS AccountID, AvailableBalance, - MarketValue, - MarginValue -FROM MCAccount -WHERE AccountID = 'ACC001'; + AssetBalance, + LiabilityBalance, + MaximumAvailableBalance +FROM dbo.MCAccount +WHERE RTRIM(AccountID) = N'M050013'; 示例2 - 多表 INNER JOIN: -问题:查询账户'ACC001'持有的所有股票及数量 -Schema: MCAccount(AccountID), MCAccountInstrument(AccountID, MarketID, InstrumentID, Settled) +问题:查询账户'M050013'持有的所有股票及数量 +Schema: BCAccountInstrument(AccountID, MarketID, InstrumentID, DailyOpenLedgerQuantity, DailyOpenSettledQuantity), MCInstrument(MarketID, InstrumentID, Name, InstrumentTypeID), MCInstrumentType(InstrumentTypeID, Name) SQL: SELECT - a.AccountID, - i.MarketID, - i.InstrumentID AS InstrumentCode, - i.Settled AS HoldingQty -FROM MCAccount a -JOIN MCAccountInstrument i ON a.AccountID = i.AccountID -WHERE a.AccountID = 'ACC001'; + RTRIM(bai.AccountID) AS AccountID, + RTRIM(bai.MarketID) AS MarketID, + RTRIM(bai.InstrumentID) AS InstrumentID, + RTRIM(mi.Name) AS InstrumentName, + RTRIM(mi.InstrumentTypeID) AS InstrumentTypeID, + it.Name AS InstrumentTypeName, + bai.DailyOpenLedgerQuantity AS LedgerQuantity, + bai.DailyOpenSettledQuantity AS SettledQuantity +FROM dbo.BCAccountInstrument AS bai +INNER JOIN dbo.MCInstrument AS mi + ON bai.MarketID = mi.MarketID + AND bai.InstrumentID = mi.InstrumentID +LEFT JOIN dbo.MCInstrumentType AS it + ON mi.InstrumentTypeID = it.InstrumentTypeID +WHERE RTRIM(bai.AccountID) = N'M050013' + AND bai.DailyOpenLedgerQuantity <> 0 +ORDER BY bai.MarketID, bai.InstrumentID; + 示例3 - 时间范围推断(关键!): 问题:查询2024年1月的总销售额 @@ -223,7 +236,7 @@ SQL_GENERATOR_USER = """Schema信息: 数据库方言:{dialect} -请生成**有用 SQL**(见系统提示定义):必须与「业务级黄金范例」**同构**——大写关键字、多行缩进版式、PascalCase 别名、该展示对手方/账户等名称时须 LEFT JOIN 维表;禁止输出挤成一行的「极简 SQL」。""" +请生成**有用 SQL**(见系统提示定义):必须与「业务级黄金范例」**同构**——大写关键字、多行缩进版式、PascalCase 别名、该展示对手方/账户等名称时须 LEFT JOIN 维表;禁止输出挤成一行的「极简 SQL」。**若当前问题未要求按时间筛选,不得在 WHERE 中擅自添加日期条件(系统提示 1b)。**""" # ========== Few-shot 黄金 SQL 条件适配(Chroma 库内为已校验正确答案)========== @@ -233,7 +246,7 @@ GOLDEN_SQL_ADAPT_SYSTEM = """你是精通 Microsoft SQL Server (T-SQL) 的数据 **你必须遵守**: 1. **以标准答案为主干**:优先保留其 `FROM`/`JOIN`/`ON`、主 `SELECT` 列清单与聚合/分组逻辑;**不要随意更换主表、不要拆掉必要 JOIN**,除非当前 Schema 片段中已不存在该表(此时在 Schema 内做最小替换并说明等价关系仅在脑中完成)。 -2. **只改「条件类」内容**:重点调整 `WHERE`/`HAVING`/`ORDER BY`/`TOP` 中的字面量、日期区间、状态码、账户/合约/代码等过滤;将用户问题中的时间范围、业务对象、筛选口径反映到这些条件中。 +2. **只改「条件类」内容**:重点调整 `WHERE`/`HAVING`/`ORDER BY`/`TOP` 中的字面量、日期区间、状态码、账户/合约/代码等过滤;将用户问题中的时间范围、业务对象、筛选口径反映到这些条件中。**若当前用户问题相较范例问题「少了」时间要求**(完全未提时间或时段),应**去掉**标准答案中仅因范例日期而存在的日期过滤,**禁止**保留与当前问题无关的日期条件(与 SQL_GENERATOR 的 **1b** 一致)。 3. **Schema 绝对优先**:表名、列名必须来自下方「当前 Schema 片段」;禁止臆造字段。若标准答案中某列在片段中不存在,按片段改写为合法列。 4. **T-SQL 与版式**:与常规生成一致——关键字大写、多行缩进、`WHERE` 续行以 `AND` 开头、需要时 PascalCase 英文别名;禁止 MySQL 反引号与 `CURDATE()` 等。 5. **禁止在字符串字面量中写中日韩文字**去匹配代码列;须用 Schema 注释中的代码或 JOIN 维表(与系统提示 SQL_GENERATOR 一致)。 diff --git a/backend/main.py b/backend/main.py index 2cbecde..73824cf 100644 --- a/backend/main.py +++ b/backend/main.py @@ -10,10 +10,6 @@ import sys import argparse import logging from pathlib import Path -from typing import Optional - -from dotenv import load_dotenv - # 从仓库根目录运行 python backend/main.py 时,将 backend 加入模块搜索路径 _backend_dir = Path(__file__).resolve().parent if str(_backend_dir) not in sys.path: @@ -27,201 +23,13 @@ logging.basicConfig( ) logger = logging.getLogger(__name__) -def _repo_root() -> Path: - """仓库根目录(含 data/、.env、api_server.py 的目录)。""" - return Path(__file__).resolve().parent.parent - - -def _load_project_env(): - """加载项目根目录 .env,供后续 os.getenv 使用。""" - load_dotenv(_repo_root() / ".env") - - -def resolve_sql_dialect(name: str) -> str: - """CLI / 配置中的方言别名统一为 sqlglot 方言名(SQL Server -> tsql)。""" - n = (name or "sqlserver").lower().strip() - if n in ("sqlserver", "mssql"): - return "tsql" - return n - - -def setup_environment(): - """环境检查""" - _load_project_env() - ms_key = os.getenv("MODELSCOPE_API_KEY", "").strip() - oa_key = os.getenv("OPENAI_API_KEY", "").strip() - oa_key_ok = oa_key and not ( - oa_key.startswith("http://") or oa_key.startswith("https://") - ) - if ms_key: - pass # ModelScope:BASE_URL / MODEL 有默认值,仅需 KEY - elif oa_key_ok: - if not os.getenv("OPENAI_EMBEDDING_MODEL", "").strip(): - logger.warning("未设置 OPENAI_EMBEDDING_MODEL(远程 Embedding 模型名)") - return False - else: - if not os.getenv("DASHSCOPE_API_KEY", "").strip(): - logger.warning( - "未设置 MODELSCOPE_API_KEY、" - "OPENAI_API_KEY+OPENAI_EMBEDDING_MODEL 或 DASHSCOPE_API_KEY" - ) - return False - base = ( - os.getenv("DASHSCOPE_BASE_URL") or os.getenv("DASHSCOPE_base_url", "") - ).strip() - if not base: - logger.warning("未设置 DASHSCOPE_BASE_URL(或 DASHSCOPE_base_url)") - return False - if not os.getenv("DASHSCOPE_MODEL", "").strip(): - logger.warning("未设置 DASHSCOPE_MODEL") - return False - - # 检查Schema文件(支持相对路径和绝对路径) - schema_path_str = os.getenv("SCHEMA_PATH", "./data/schemas/G3SB_MCDataDictionary_table_structure.json") - schema_path = Path(schema_path_str) - - # 如果是相对路径,尝试从多个位置查找 - if not schema_path.is_absolute(): - # 尝试1: PyInstaller 临时目录(单文件模式) - import sys - if getattr(sys, 'frozen', False) and hasattr(sys, '_MEIPASS'): - meipass_schema = Path(sys._MEIPASS) / schema_path_str - if meipass_schema.exists(): - schema_path = meipass_schema - - # 尝试2: 当前工作目录 - if not schema_path.exists(): - schema_path = Path.cwd() / schema_path_str - - # 尝试3: 脚本所在目录 - if not schema_path.exists(): - script_dir = Path(__file__).resolve().parent.parent - schema_path = script_dir / schema_path_str - - # 尝试4: 可执行文件所在目录 - if not schema_path.exists() and getattr(sys, 'frozen', False): - exe_dir = Path(sys.executable).parent - schema_path = exe_dir / schema_path_str - - if not schema_path.exists(): - logger.warning(f"Schema文件不存在: {schema_path}") - logger.info("请将G3SB Schema文件放置在 ./data/schemas/ 目录") - return False - - # 检查 LLM Key(DeepSeek / OpenAI 可切换) - llm_sc = (os.getenv("LLM_SERVICE_CODE") or "").strip().lower() - if llm_sc and llm_sc not in ("deepseek", "openai"): - logger.warning("未知 LLM_SERVICE_CODE=%r(仅支持 deepseek/openai)", llm_sc) - return False - if llm_sc == "openai": - if not (os.getenv("OPENAI_API_KEY") or "").strip(): - logger.warning("LLM_SERVICE_CODE=openai 但 OPENAI_API_KEY 未设置") - return False - else: - if not (os.getenv("DEEPSEEK_API_KEY") or "").strip(): - logger.warning("环境变量 DEEPSEEK_API_KEY 未设置(默认 LLM_SERVICE_CODE=deepseek)") - logger.info("请在 .env 文件中配置,或 export DEEPSEEK_API_KEY=your_key") - return False - - logger.info(f"[OK] 环境检查通过") - logger.info(f" - Schema: {schema_path}") - logger.info( - " - LLM: %s", - (llm_sc or "deepseek(auto)"), - ) - - return True - - -def _default_g3sb_meta_path(structure_path: str) -> Optional[str]: - """若存在与 table_structure 同名的 table_meta 文件则返回其路径。""" - p = Path(structure_path) - if "table_structure" not in p.name: - return None - cand = p.parent / p.name.replace("table_structure", "table_meta") - return str(cand) if cand.is_file() else None - - -def load_schema(schema_path: str, schema_meta_path: Optional[str] = None): - """加载 Schema;G3SB structure JSON 会自动尝试配对 table_meta(可用 --schema-meta 指定)。""" - from schema.manager import SchemaManager - - meta = ( - schema_meta_path - if schema_meta_path is not None - else _default_g3sb_meta_path(schema_path) - ) - logger.info(f"加载Schema: {schema_path}") - if meta: - logger.info(f" 表注释(meta): {meta}") - - schema_mgr = SchemaManager.load_from_json( - schema_path, g3sb_meta_path=meta - ) - - stats = schema_mgr.get_statistics() - logger.info( - f"[OK] Schema加载完成: {stats['database']}, " - f"共{stats['total_tables']}张表, {stats['total_columns']}个字段" - ) - - return schema_mgr - - -def create_orchestrator(schema_mgr, args): - """创建编排器""" - from agents.orchestrator import Text2SQLOrchestrator - from llm.router import create_llm_client, resolve_llm_service_code - from llm.deepseek_client import DeepSeekConfig - - translate_en = os.getenv("TRANSLATE_EN_TO_ZH", "true").strip().lower() not in ( - "0", - "false", - "no", - "off", - ) - if getattr(args, "no_translate_en", False): - translate_en = False - - sc = resolve_llm_service_code() - if sc == "deepseek": - api_key = (args.api_key or os.getenv("DEEPSEEK_API_KEY") or "").strip() - if not api_key: - raise ValueError( - "未配置 DeepSeek API Key:请在 .env 中设置 DEEPSEEK_API_KEY," - "或使用命令行参数 --api-key" - ) - base_url = (os.getenv("DEEPSEEK_BASE_URL") or "https://api.deepseek.com").strip() - cfg = DeepSeekConfig( - api_key=api_key, - base_url=base_url, - model_name=args.model or "deepseek-chat", - temperature=args.temperature, - max_tokens=args.max_tokens, - ) - llm_client = create_llm_client("deepseek", **cfg.__dict__) - else: - # openai:完全由 OPENAI_* 决定;同时沿用 temperature/max_tokens 作为默认值覆盖 - llm_client = create_llm_client( - "openai", - temperature=args.temperature, - max_tokens=args.max_tokens, - ) - - orchestrator = Text2SQLOrchestrator( - schema_manager=schema_mgr, - llm_client=llm_client, - vector_db_path=args.vector_db, - max_retry=args.max_retry, - use_vector_search=not args.no_vector_search, - # Few-shot配置 - fewshot_enabled=not args.no_fewshot, - fewshot_top_k=args.fewshot_top_k, - fewshot_min_rating=args.fewshot_min_rating, - translate_english_to_zh=translate_en, - ) - - return orchestrator +from bootstrap import ( + create_orchestrator, + load_project_env, + load_schema, + resolve_sql_dialect, + setup_environment, +) def single_query(orchestrator, question: str, dialect: str = "tsql"): @@ -361,7 +169,7 @@ def main(): default=2, help="最大重试次数(默认: 2)" ) - _load_project_env() + load_project_env() parser.add_argument( "--vector-db", default=os.getenv("VECTOR_DB_PATH", "./data/embeddings/chroma").strip(), diff --git a/backend/utils/__pycache__/fewshot_chroma_store.cpython-312.pyc b/backend/utils/__pycache__/fewshot_chroma_store.cpython-312.pyc index bfe9137..5c4ffda 100644 Binary files a/backend/utils/__pycache__/fewshot_chroma_store.cpython-312.pyc and b/backend/utils/__pycache__/fewshot_chroma_store.cpython-312.pyc differ diff --git a/backend/utils/__pycache__/fewshot_selector.cpython-312.pyc b/backend/utils/__pycache__/fewshot_selector.cpython-312.pyc index 078654d..5109071 100644 Binary files a/backend/utils/__pycache__/fewshot_selector.cpython-312.pyc and b/backend/utils/__pycache__/fewshot_selector.cpython-312.pyc differ diff --git a/backend/utils/__pycache__/validators.cpython-312.pyc b/backend/utils/__pycache__/validators.cpython-312.pyc index f77392b..10a8e0e 100644 Binary files a/backend/utils/__pycache__/validators.cpython-312.pyc and b/backend/utils/__pycache__/validators.cpython-312.pyc differ diff --git a/backend/utils/validators.py b/backend/utils/validators.py index f683a2b..b7b1238 100644 --- a/backend/utils/validators.py +++ b/backend/utils/validators.py @@ -4,7 +4,7 @@ SQL 验证工具集 import re import logging -from typing import Tuple, List, Dict +from typing import Tuple, List, Dict, Optional logger = logging.getLogger(__name__) @@ -14,11 +14,56 @@ _CJK_IN_STRING_RE = re.compile( ) -def check_no_cjk_in_sql_string_literals(sql: str) -> Tuple[bool, List[str]]: - """ - 扫描 SQL 中单引号字符串(含 T-SQL N'…'),若字面量内出现 CJK 则判失败。 +def _cjk_text_in_string_literal(text: str) -> bool: + return bool(_CJK_IN_STRING_RE.search(text)) - 跳过 ``--`` 行注释与 ``/* */`` 块注释内的文本,避免误报。 + +def _node_in_subtree(root, target) -> bool: + if root is None: + return False + for n in root.walk(): + if n is target: + return True + return False + + +def _cjk_string_in_forbidden_context(node, exp) -> bool: + """ + 禁止含 CJK 的字面量出现在「比对/过滤」语境:WHERE、HAVING、JOIN ON、 + 以及 CASE 分支的 WHEN 条件(含简单 CASE 的 WHEN 值), + 但允许出现在 SELECT 投影、CASE 的 THEN/ELSE 结果等纯展示位置。 + """ + if node.find_ancestor(exp.Where): + return True + if node.find_ancestor(exp.Having): + return True + join = node.find_ancestor(exp.Join) + if join is not None: + on = join.args.get("on") + if on is not None and _node_in_subtree(on, node): + return True + + case = node.find_ancestor(exp.Case) + while case is not None: + default = case.args.get("default") + if default is not None and _node_in_subtree(default, node): + return False + for br in case.args.get("ifs") or []: + then_expr = br.args.get("true") + cond = br.this + if then_expr is not None and _node_in_subtree(then_expr, node): + return False + if cond is not None and _node_in_subtree(cond, node): + return True + case = case.find_ancestor(exp.Case) + + return False + + +def _check_no_cjk_legacy_text_scan(sql: str) -> Tuple[bool, List[str]]: + """ + 解析失败时的回退:扫描单引号字符串(含 N'…'),字面量内出现 CJK 即失败。 + 跳过 ``--`` 行注释与 ``/* */`` 块注释内的文本。 """ errors: List[str] = [] i = 0 @@ -27,7 +72,6 @@ def check_no_cjk_in_sql_string_literals(sql: str) -> Tuple[bool, List[str]]: in_block_comment = False def _read_single_quoted_string(start: int) -> Tuple[str, int]: - """从 start 指向的 opening `'` 之后开始读,返回 (内容, 闭合引号后下标)。""" j = start parts: List[str] = [] while j < n: @@ -66,10 +110,9 @@ def check_no_cjk_in_sql_string_literals(sql: str) -> Tuple[bool, List[str]]: i += 2 continue - # N' 或 n' 前缀的 Unicode 字面量 if i + 1 < n and sql[i] in "Nn" and sql[i + 1] == "'": body, i = _read_single_quoted_string(i + 2) - if _CJK_IN_STRING_RE.search(body): + if _cjk_text_in_string_literal(body): prev = body[:48] + ("…" if len(body) > 48 else "") errors.append( "SQL 字符串字面量中含中文或与业务中文直接作为比对值(禁止)。" @@ -79,7 +122,7 @@ def check_no_cjk_in_sql_string_literals(sql: str) -> Tuple[bool, List[str]]: if sql[i] == "'": body, i = _read_single_quoted_string(i + 1) - if _CJK_IN_STRING_RE.search(body): + if _cjk_text_in_string_literal(body): prev = body[:48] + ("…" if len(body) > 48 else "") errors.append( "SQL 字符串字面量中含中文或与业务中文直接作为比对值(禁止)。" @@ -92,6 +135,41 @@ def check_no_cjk_in_sql_string_literals(sql: str) -> Tuple[bool, List[str]]: return len(errors) == 0, errors +def check_no_cjk_in_sql_string_literals(sql: str, dialect: str = "tsql") -> Tuple[bool, List[str]]: + """ + 禁止在「过滤/比对」语境使用含中日韩字符的字符串字面量(含 T-SQL ``N'…'``)。 + + 允许在 SELECT 投影、CASE 的 THEN/ELSE 结果等展示用字面量中使用中文标签。 + 解析失败时回退为全文扫描(与旧版一致,偏严)。 + """ + from sqlglot import parse_one, exp + + errors: List[str] = [] + + try: + parsed = parse_one(sql, dialect=dialect) + except Exception as e: + logger.debug("CJK 校验回退为全文扫描(SQL 解析失败): %s", e) + return _check_no_cjk_legacy_text_scan(sql) + + for node in parsed.walk(): + text: Optional[str] = None + if isinstance(node, exp.Literal) and node.is_string: + text = str(node.this) + elif isinstance(node, exp.National): + text = str(node.this) + if text is None or not _cjk_text_in_string_literal(text): + continue + if _cjk_string_in_forbidden_context(node, exp): + prev = text[:48] + ("…" if len(text) > 48 else "") + errors.append( + "SQL 在 WHERE/HAVING/JOIN ON 或 CASE/WHEN 条件中出现含中文的字符串字面量(禁止)。" + f"片段近似: …'{prev}'… — 请改用 Schema 注释中的代码/枚举,或通过维表 JOIN 用键列过滤。" + ) + + return len(errors) == 0, errors + + # 危险操作关键词(除非明确允许) DANGEROUS_KEYWORDS = [ "DROP", "DELETE", "UPDATE", "INSERT", "ALTER", "TRUNCATE", diff --git a/data/embeddings/chroma_fewshot/9756e6ea-6d51-4e24-a490-4a58f0a585b9/length.bin b/data/embeddings/chroma_fewshot/9756e6ea-6d51-4e24-a490-4a58f0a585b9/length.bin index d1fa5fe..16406af 100644 Binary files a/data/embeddings/chroma_fewshot/9756e6ea-6d51-4e24-a490-4a58f0a585b9/length.bin and b/data/embeddings/chroma_fewshot/9756e6ea-6d51-4e24-a490-4a58f0a585b9/length.bin differ diff --git a/logs/text2sql_api.log b/logs/text2sql_api.log index 0a01ac8..a02e2bd 100644 --- a/logs/text2sql_api.log +++ b/logs/text2sql_api.log @@ -1,68 +1,135 @@ -2026-04-16 16:22:05 INFO [utils.repo_logging] repo_logging.py:79 configure_text2sql_api_logging() | 日志文件: C:\Users\24019\Desktop\backman-camel\logs\text2sql_api.log -2026-04-16 16:22:10 INFO [__main__] api_server.py:85 () | [OK] 已加载配置文件: C:\Users\24019\Desktop\backman-camel\.env -2026-04-16 16:22:10 INFO [__main__] api_server.py:1108 () | 启动服务: http://0.0.0.0:8041 -2026-04-16 16:22:10 INFO [__main__] api_server.py:1109 () | API文档: http://0.0.0.0:8041/docs -2026-04-16 16:22:10 INFO [uvicorn.error] server.py:92 _serve() | Started server process [33080] -2026-04-16 16:22:10 INFO [uvicorn.error] on.py:48 startup() | Waiting for application startup. -2026-04-16 16:22:10 INFO [__main__] api_server.py:799 lifespan() | ============================================================ -2026-04-16 16:22:10 INFO [__main__] api_server.py:800 lifespan() | Text2SQL API Server 启动中... -2026-04-16 16:22:10 INFO [__main__] api_server.py:801 lifespan() | ============================================================ -2026-04-16 16:22:10 INFO [main] main.py:126 setup_environment() | [OK] 环境检查通过 -2026-04-16 16:22:10 INFO [main] main.py:127 setup_environment() | - Schema: data\schemas\G3SB_MCDataDictionary_table_structure.json -2026-04-16 16:22:10 INFO [main] main.py:128 setup_environment() | - LLM: openai -2026-04-16 16:22:10 INFO [main] main.py:154 load_schema() | 加载Schema: ./data/schemas/G3SB_MCDataDictionary_table_structure.json -2026-04-16 16:22:10 INFO [main] main.py:156 load_schema() | 表注释(meta): ./data/schemas/G3SB_MCDataDictionary_table_meta.json -2026-04-16 16:22:10 INFO [schema.loader] loader.py:107 load_from_json() | [OK] 加载Schema完成(G3SB schemas): G3SB_MCDataDictionary_table_structure, 共2516张表 -2026-04-16 16:22:10 INFO [main] main.py:163 load_schema() | [OK] Schema加载完成: G3SB_MCDataDictionary_table_structure, 共2516张表, 59196个字段 -2026-04-16 16:22:14 INFO [llm.openai_client] openai_client.py:55 __init__() | [OK] OpenAIClient初始化: model=gpt-4o-mini, base_url=http://113.192.49.54:9080/v1 -2026-04-16 16:22:17 INFO [utils.embedding] embedding.py:144 __init__() | 使用 OpenAI Embedding API:model=text-embedding-ada-002,base_url=http://113.192.49.54:9080/v1,max_batch=100 -2026-04-16 16:22:19 INFO [utils.fewshot_chroma_store] fewshot_chroma_store.py:86 __init__() | [OK] FewShotChromaStore: 磁盘 data\embeddings\chroma_fewshot collection=fewshot_samples count=50 -2026-04-16 16:22:19 INFO [utils.fewshot_selector] fewshot_selector.py:128 _init_chroma_mode() | [Few-shot] 已从向量库加载(Chroma 50 条,data\embeddings\chroma_fewshot) -2026-04-16 16:22:19 INFO [utils.fewshot_selector] fewshot_selector.py:159 _init_chroma_mode() | [OK] Few-shot 使用 Chroma(50 条) -2026-04-16 16:22:19 INFO [agents.orchestrator] orchestrator.py:116 __init__() | Few-shot已启用: top_k=3, min_rating=7 -2026-04-16 16:22:19 INFO [agents.orchestrator] orchestrator.py:127 __init__() | [OK] Text2SQLOrchestrator初始化完成: max_retry=2, use_vector_search=True, fewshot=on, nl→zh_norm=on -2026-04-16 16:22:19 INFO [__main__] api_server.py:132 get_orchestrator() | [OK] Orchestrator 初始化完成 -2026-04-16 16:22:19 INFO [__main__] api_server.py:805 lifespan() | [OK] 服务已就绪 -2026-04-16 16:22:19 INFO [uvicorn.error] on.py:62 startup() | Application startup complete. -2026-04-16 16:22:19 INFO [uvicorn.error] server.py:224 _log_started_message() | Uvicorn running on http://0.0.0.0:8041 (Press CTRL+C to quit) -2026-04-16 16:22:20 INFO [uvicorn.access] httptools_impl.py:483 send() | 127.0.0.1:14651 - "POST /g3sb/api/nl/chat/stream HTTP/1.1" 200 -2026-04-16 16:22:20 INFO [__main__] api_server.py:669 _chat_stream_events() | [API/stream] 开始: user_id='anonymous' visitor_biz_id=None session_id=None service_code=None model='gpt-4o-mini' lang_code='auto' msg_chars=15 preview='所有客户账户之间的股票转移记录' dialog_context_chars=0 last_turn_was_data_query=False -2026-04-16 16:22:23 INFO [llm.openai_client] openai_client.py:55 __init__() | [OK] OpenAIClient初始化: model=gpt-4o-mini, base_url=http://113.192.49.54:9080/v1 -2026-04-16 16:22:23 INFO [utils.dialog_classifier] dialog_classifier.py:261 classify_dialog() | [dialog] intent=text2sql (hybrid fast: query hint) preview='所有客户账户之间的股票转移记录' -2026-04-16 16:22:23 INFO [__main__] api_server.py:712 _chat_stream_events() | [GEN/API/stream] dialect=tsql top_k=20 dialog_context_chars=0 question_len=15 preview='所有客户账户之间的股票转移记录' -2026-04-16 16:22:26 INFO [llm.openai_client] openai_client.py:55 __init__() | [OK] OpenAIClient初始化: model=gpt-4o-mini, base_url=http://113.192.49.54:9080/v1 -2026-04-16 16:22:29 INFO [agents.orchestrator] orchestrator.py:803 generate() | [GEN] 问句已归一中文:所有客户账户的股票转移记录 -2026-04-16 16:22:29 INFO [agents.orchestrator] orchestrator.py:816 generate() | [GEN] 开始生成SQL: question_chars=13 preview='所有客户账户的股票转移记录' dialog_context_chars=0 -2026-04-16 16:22:29 INFO [agents.orchestrator] orchestrator.py:831 generate() | 尝试 #1 -2026-04-16 16:22:29 INFO [schema.indexer] indexer.py:82 __init__() | [OK] 初始化SchemaIndexer(Chroma磁盘 path=data\embeddings\chroma): collection=schema_tables, count=2516 -2026-04-16 16:22:29 INFO [schema.indexer] indexer.py:106 ensure_index_for_schema() | Schema 向量索引已就绪(2516 张表),跳过向量化 -2026-04-16 16:22:29 INFO [agents.orchestrator] orchestrator.py:188 _coarse_filter() | [Orchestrator] 开始向量检索: query_chars=13 query_preview='所有客户账户的股票转移记录' -2026-04-16 16:22:30 INFO [utils.embedding] embedding.py:163 _set_dim_from_vector() | [OK] Embedding 向量维度:1536 -2026-04-16 16:22:31 INFO [schema.indexer] indexer.py:244 search() | Schema 向量检索: query_chars=13 命中=20(阈值=0.1)top=[('VSBHKRpt0430', 0.8264), ('VSBHKRpt0431', 0.8188), ('TSBTransferInstruction', 0.8147), ('VSBTransferInstruction', 0.8134), ('VSBHKRpt0397', 0.8126), ('VSBHKRpt0672', 0.8117), ('VSBHKRpt0090C', 0.81), ('VSBHKRpt0570', 0.8084), ('VSBHKRpt1030A', 0.8077), ('VSBRpt0999E', 0.8065), ('VSBRpt0999C', 0.8064), ('TSBAccountEntitlementRelease', 0.8063), ('VSBRpt0999A', 0.8059), ('TSBAccountInstrumentMovement', 0.8038), ('VSBTransferInstructionGenerationByAccountContract', 0.8033)] -2026-04-16 16:22:31 INFO [agents.orchestrator] orchestrator.py:204 _coarse_filter() | [Orchestrator] 向量粗筛: 命中=20 张(阈值内),表名+分: [('VSBHKRpt0430', 0.8264), ('VSBHKRpt0431', 0.8188), ('TSBTransferInstruction', 0.8147), ('VSBTransferInstruction', 0.8134), ('VSBHKRpt0397', 0.8126), ('VSBHKRpt0672', 0.8117), ('VSBHKRpt0090C', 0.81), ('VSBHKRpt0570', 0.8084), ('VSBHKRpt1030A', 0.8077), ('VSBRpt0999E', 0.8065), ('VSBRpt0999C', 0.8064), ('TSBAccountEntitlementRelease', 0.8063), ('VSBRpt0999A', 0.8059), ('TSBAccountInstrumentMovement', 0.8038), ('VSBTransferInstructionGenerationByAccountContract', 0.8033), ('VSBRpt0060', 0.8021), ('XCGatewayStockReconciliationReport', 0.802), ('TSBAccountEntitlementHold', 0.8013), ('WSBBatchLocationTransferDetail', 0.8012), ('VCAccountCashMovement', 0.8012)] -2026-04-16 16:22:35 INFO [agents.orchestrator] orchestrator.py:258 _llm_select_tables() | LLM精筛选中表:['VSBHKRpt0430', 'VSBHKRpt0431', 'TSBTransferInstruction', 'VSBTransferInstruction', 'TSBAccountInstrumentMovement'] | reasoning_chars=192 reasoning_preview='问题涉及客户账户的股票转移记录,VSBHKRpt0430和VSBHKRpt0431提供了客户股票的日常转移报告和账户工具移动报告,TSBTransferInstruction和VSBTransferInstruction包含账户转移指令的详细信息,TSBAccountInstrumentMovement记录账户工具的移动交易。这些表共同涵盖了股票转移的各个方面,确保了信息的完整性。' -2026-04-16 16:22:35 INFO [agents.orchestrator] orchestrator.py:859 generate() | 选中表:['VSBHKRpt0430', 'VSBHKRpt0431', 'TSBTransferInstruction', 'VSBTransferInstruction', 'TSBAccountInstrumentMovement'],扩展后:['TSBAccountInstrumentMovement', 'TSBTransferInstruction', 'VSBHKRpt0431', 'VSBHKRpt0430', 'VSBTransferInstruction'] -2026-04-16 16:22:37 INFO [utils.fewshot_selector] fewshot_selector.py:209 select_best_with_score() | [Few-shot] best_with_score: qid=Q2 score=0.9575 preview='列出今日所有客户账户之间的股票转移记录。' -2026-04-16 16:22:37 INFO [agents.orchestrator] orchestrator.py:432 _generate_sql_golden_adapt() | [GEN] 黄金 few-shot 条件适配: qid=Q2 score=0.9575 -2026-04-16 16:22:41 INFO [agents.orchestrator] orchestrator.py:450 _generate_sql_golden_adapt() | 生成的SQL(黄金适配,chars=494): --- 查询所有客户账户间股票转移记录 --- 使用 TSBAccountInstrumentMovement,MovementType='T' 表示账户间转移 -SELECT - m.MovementID, - m.AccountID AS FromAccountID, - m.TransferToAccountID, - m.InstrumentID, - i.Name AS InstrumentSymbol, - m.MovementType, - m.Quantity AS TransferQuantity, - m.ValueDate AS TransferDate -FROM TSBAccountInstrumentMovement m -LEFT JOIN MCInstrument i ON m.InstrumentID = i.InstrumentID -WHERE m.MovementType = 'T' - AND m.ValueDate = CAST(GETDATE() AS DATE) -ORDER BY m.ValueDate; -2026-04-16 16:22:41 INFO [db.engine] engine.py:45 get_engine() | SQLAlchemy engine initialized from database_url -2026-04-16 16:22:41 INFO [db.dbhub_tools] dbhub_tools.py:627 _execute_sql() | _execute_sql 执行语句数=1 readonly=True max_rows=1 -2026-04-16 16:22:44 INFO [agents.orchestrator] orchestrator.py:754 _validate_sql() | [validate] 程序+探针+LLM 汇总: valid=True err_count=0 warn_count=0 db_execution_status=0 sql_chars=494 -2026-04-16 16:22:44 INFO [agents.orchestrator] orchestrator.py:921 generate() | [OK] SQL生成与验证通过(1次尝试) -2026-04-16 16:22:44 INFO [__main__] api_server.py:765 _chat_stream_events() | [API/stream] Text2SQL 完成: valid=True attempts=1 tables_used=['TSBAccountInstrumentMovement', 'TSBTransferInstruction', 'VSBHKRpt0431', 'VSBHKRpt0430', 'VSBTransferInstruction'] sql_chars=494 sql_head="-- 查询所有客户账户间股票转移记录\n-- 使用 TSBAccountInstrumentMovement,MovementType='T' 表示账户间转移\nSELECT \n m.MovementID,\n m.AccountID AS FromAccountID,\n m.TransferToAccountID,\n m.InstrumentID,\n i.Name AS InstrumentSymbol,\n m.MovementType,\n m.Quantity AS TransferQuantity,\n m.ValueDate AS TransferDate\nFROM TSBAccountInstrumentMovement m\nLEFT JOIN MCInstrument i ON m.InstrumentID = i.InstrumentID\nWHERE m.MovementType = 'T'\n AND m.ValueDate = CAST(GETDATE() AS DATE)\nORDER BY m.ValueDate;" +2026-04-17 10:48:50 INFO [utils.repo_logging] repo_logging.py:79 configure_text2sql_api_logging() | 日志文件: C:\Users\24019\Desktop\backman-camel\logs\text2sql_api.log +2026-04-17 10:48:55 INFO [__main__] api_server.py:86 () | [OK] 已加载配置文件: C:\Users\24019\Desktop\backman-camel\.env +2026-04-17 10:48:55 INFO [__main__] api_server.py:1299 () | 启动服务: http://0.0.0.0:8041 +2026-04-17 10:48:55 INFO [__main__] api_server.py:1300 () | API文档: http://0.0.0.0:8041/docs +2026-04-17 10:48:55 INFO [uvicorn.error] server.py:92 _serve() | Started server process [38600] +2026-04-17 10:48:55 INFO [uvicorn.error] on.py:48 startup() | Waiting for application startup. +2026-04-17 10:48:55 INFO [__main__] api_server.py:969 lifespan() | ============================================================ +2026-04-17 10:48:55 INFO [__main__] api_server.py:970 lifespan() | Text2SQL API Server 启动中... +2026-04-17 10:48:55 INFO [__main__] api_server.py:971 lifespan() | ============================================================ +2026-04-17 10:48:55 INFO [main] main.py:126 setup_environment() | [OK] 环境检查通过 +2026-04-17 10:48:55 INFO [main] main.py:127 setup_environment() | - Schema: data\schemas\G3SB_MCDataDictionary_table_structure.json +2026-04-17 10:48:55 INFO [main] main.py:128 setup_environment() | - LLM: openai +2026-04-17 10:48:55 INFO [main] main.py:154 load_schema() | 加载Schema: ./data/schemas/G3SB_MCDataDictionary_table_structure.json +2026-04-17 10:48:55 INFO [main] main.py:156 load_schema() | 表注释(meta): ./data/schemas/G3SB_MCDataDictionary_table_meta.json +2026-04-17 10:48:55 INFO [schema.loader] loader.py:107 load_from_json() | [OK] 加载Schema完成(G3SB schemas): G3SB_MCDataDictionary_table_structure, 共2516张表 +2026-04-17 10:48:55 INFO [main] main.py:163 load_schema() | [OK] Schema加载完成: G3SB_MCDataDictionary_table_structure, 共2516张表, 59196个字段 +2026-04-17 10:48:59 INFO [llm.openai_client] openai_client.py:55 __init__() | [OK] OpenAIClient初始化: model=gpt-4o-mini, base_url=http://113.192.49.54:9080/v1 +2026-04-17 10:49:01 INFO [utils.embedding] embedding.py:144 __init__() | 使用 OpenAI Embedding API:model=text-embedding-ada-002,base_url=http://113.192.49.54:9080/v1,max_batch=100 +2026-04-17 10:49:03 INFO [utils.fewshot_chroma_store] fewshot_chroma_store.py:86 __init__() | [OK] FewShotChromaStore: 磁盘 data\embeddings\chroma_fewshot collection=fewshot_samples count=50 +2026-04-17 10:49:03 INFO [utils.fewshot_selector] fewshot_selector.py:128 _init_chroma_mode() | [Few-shot] 已从向量库加载(Chroma 50 条,data\embeddings\chroma_fewshot) +2026-04-17 10:49:03 INFO [utils.fewshot_selector] fewshot_selector.py:159 _init_chroma_mode() | [OK] Few-shot 使用 Chroma(50 条) +2026-04-17 10:49:03 INFO [agents.orchestrator] orchestrator.py:116 __init__() | Few-shot已启用: top_k=3, min_rating=7 +2026-04-17 10:49:03 INFO [agents.orchestrator] orchestrator.py:127 __init__() | [OK] Text2SQLOrchestrator初始化完成: max_retry=2, use_vector_search=True, fewshot=on, nl→zh_norm=on +2026-04-17 10:49:03 INFO [__main__] api_server.py:133 get_orchestrator() | [OK] Orchestrator 初始化完成 +2026-04-17 10:49:03 INFO [__main__] api_server.py:975 lifespan() | [OK] 服务已就绪 +2026-04-17 10:49:03 INFO [uvicorn.error] on.py:62 startup() | Application startup complete. +2026-04-17 10:49:03 INFO [uvicorn.error] server.py:224 _log_started_message() | Uvicorn running on http://0.0.0.0:8041 (Press CTRL+C to quit) +2026-04-17 10:49:16 INFO [uvicorn.access] httptools_impl.py:483 send() | 127.0.0.1:29292 - "POST /g3sb/api/nl/chat/stream HTTP/1.1" 200 +2026-04-17 10:49:16 INFO [__main__] api_server.py:735 _chat_stream_events() | [API/stream] 开始: user_id='anonymous' visitor_biz_id=None session_id=None service_code=None model='gpt-4o-mini' lang_code='auto' msg_chars=13 preview='查看目前余额最高的账号名称' dialog_context_chars=0 last_turn_was_data_query=False +2026-04-17 10:49:19 INFO [llm.openai_client] openai_client.py:55 __init__() | [OK] OpenAIClient初始化: model=gpt-4o-mini, base_url=http://113.192.49.54:9080/v1 +2026-04-17 10:49:19 INFO [utils.dialog_classifier] dialog_classifier.py:261 classify_dialog() | [dialog] intent=text2sql (hybrid fast: query hint) preview='查看目前余额最高的账号名称' +2026-04-17 10:49:19 INFO [__main__] api_server.py:801 _chat_stream_events() | [GEN/API/stream] dialect=tsql top_k=20 dialog_context_chars=0 question_len=13 preview='查看目前余额最高的账号名称' +2026-04-17 10:49:22 INFO [llm.openai_client] openai_client.py:55 __init__() | [OK] OpenAIClient初始化: model=gpt-4o-mini, base_url=http://113.192.49.54:9080/v1 +2026-04-17 10:49:24 INFO [agents.orchestrator] orchestrator.py:827 generate() | [GEN] 问句已归一中文:查看余额最高的账号名称 +2026-04-17 10:49:24 INFO [agents.orchestrator] orchestrator.py:840 generate() | [GEN] 开始生成SQL: question_chars=11 preview='查看余额最高的账号名称' dialog_context_chars=0 +2026-04-17 10:49:24 INFO [agents.orchestrator] orchestrator.py:855 generate() | 尝试 #1 +2026-04-17 10:49:24 INFO [schema.indexer] indexer.py:82 __init__() | [OK] 初始化SchemaIndexer(Chroma磁盘 path=data\embeddings\chroma): collection=schema_tables, count=2516 +2026-04-17 10:49:24 INFO [schema.indexer] indexer.py:106 ensure_index_for_schema() | Schema 向量索引已就绪(2516 张表),跳过向量化 +2026-04-17 10:49:24 INFO [agents.orchestrator] orchestrator.py:188 _coarse_filter() | [Orchestrator] 开始向量检索: query_chars=11 query_preview='查看余额最高的账号名称' +2026-04-17 10:49:27 INFO [utils.embedding] embedding.py:163 _set_dim_from_vector() | [OK] Embedding 向量维度:1536 +2026-04-17 10:49:27 INFO [schema.indexer] indexer.py:244 search() | Schema 向量检索: query_chars=11 命中=20(阈值=0.1)top=[('VSBFrontOfficeBalanceAccountCash', 0.7854), ('VSBRpt9009', 0.7824), ('BCAccountCash', 0.7824), ('VXVNSyncBalanceAccountCash', 0.7812), ('VCAccountOnlineCashWithdrawOrTransferredAmount', 0.781), ('VCAccountBalance', 0.7798), ('VXPhilipsB2BInterfaceAccountCashBalance', 0.779), ('VCBAccountCash', 0.7788), ('VSBAccountContractOnlineAvailableCashInAdvanceAmount', 0.7775), ('VCAccountDailyCreditLimit', 0.7765), ('VSBAyersFrontOfficeBalanceAccountCash', 0.776), ('VSBRpt0060', 0.7756), ('VSBRpt0055D_CIBalance', 0.7745), ('VSBRpt9008', 0.7745), ('VSBVNFrontOfficeMasterAccount', 0.7745)] +2026-04-17 10:49:27 INFO [agents.orchestrator] orchestrator.py:204 _coarse_filter() | [Orchestrator] 向量粗筛: 命中=20 张(阈值内),表名+分: [('VSBFrontOfficeBalanceAccountCash', 0.7854), ('VSBRpt9009', 0.7824), ('BCAccountCash', 0.7824), ('VXVNSyncBalanceAccountCash', 0.7812), ('VCAccountOnlineCashWithdrawOrTransferredAmount', 0.781), ('VCAccountBalance', 0.7798), ('VXPhilipsB2BInterfaceAccountCashBalance', 0.779), ('VCBAccountCash', 0.7788), ('VSBAccountContractOnlineAvailableCashInAdvanceAmount', 0.7775), ('VCAccountDailyCreditLimit', 0.7765), ('VSBAyersFrontOfficeBalanceAccountCash', 0.776), ('VSBRpt0060', 0.7756), ('VSBRpt0055D_CIBalance', 0.7745), ('VSBRpt9008', 0.7745), ('VSBVNFrontOfficeMasterAccount', 0.7745), ('VSBRpt0061', 0.7744), ('VSBRpt0055N_CIBalance', 0.7743), ('VCBAccountCashForBatchCashWithdrawal', 0.7739), ('VSBRPT0290I', 0.7737), ('VCBAccountSummary', 0.7736)] +2026-04-17 10:49:31 INFO [agents.orchestrator] orchestrator.py:258 _llm_select_tables() | LLM精筛选中表:['BCAccountCash', 'VCAccountBalance', 'VCBAccountCash'] | reasoning_chars=101 reasoning_preview='问题涉及查看余额,BCAccountCash存储账户现金余额信息,VCAccountBalance提供账户余额视图,VCBAccountCash是账户现金视图,这些表能够提供账户名称及其对应的余额信息。' +2026-04-17 10:49:31 INFO [agents.orchestrator] orchestrator.py:320 _prioritize_vc_user_accessible_function() | 已优先纳入目录视图 VCUserAccessibleFunction(置于选表前列),当前选表:['VCUserAccessibleFunction', 'BCAccountCash', 'VCAccountBalance', 'VCBAccountCash'] +2026-04-17 10:49:31 INFO [agents.orchestrator] orchestrator.py:894 generate() | 选中表:['VCUserAccessibleFunction', 'BCAccountCash', 'VCAccountBalance', 'VCBAccountCash'],扩展后:['VCBAccountCash', 'BCAccountCash', 'VCAccountBalance', 'VCUserAccessibleFunction'] +2026-04-17 10:49:32 INFO [utils.fewshot_selector] fewshot_selector.py:209 select_best_with_score() | [Few-shot] best_with_score: qid=Q17 score=0.8266 preview='哪些客户已超过其信用额度?' +2026-04-17 10:49:34 INFO [utils.fewshot_selector] fewshot_selector.py:357 _select_chroma() | Few-shot(Chroma)选择: 问题='查看余额最高的账号名称...' → 选中3个示例 (top_k=3, min_rating=7) +2026-04-17 10:49:34 INFO [agents.orchestrator] orchestrator.py:568 _generate_sql() | 已注入 3 个 few-shot 示例: qid=['Q17', 'Q27', 'Q23'] question_zh_preview=['哪些客户已超过其信用额度?', '列出所有超过3天的暂记账户分录。', '列出所有逾期未付保证金利息的账户。'] +2026-04-17 10:49:37 INFO [agents.orchestrator] orchestrator.py:635 _generate_sql() | 生成的SQL(chars=178): +SELECT + a.AccountID, + a.AccountName AS AccountName, + a.AvailableBalance, + a.MarketValue, + a.LedgerBalance +FROM VCAccountBalance a +ORDER BY a.AvailableBalance DESC; +2026-04-17 10:49:37 INFO [db.engine] engine.py:45 get_engine() | SQLAlchemy engine initialized from database_url +2026-04-17 10:49:37 INFO [db.dbhub_tools] dbhub_tools.py:627 _execute_sql() | _execute_sql 执行语句数=1 readonly=True max_rows=1 +2026-04-17 10:49:38 INFO [agents.orchestrator] orchestrator.py:778 _validate_sql() | [validate] 程序+探针+LLM 汇总: valid=True err_count=0 warn_count=0 db_execution_status=1 sql_chars=178 +2026-04-17 10:49:38 INFO [agents.orchestrator] orchestrator.py:956 generate() | [OK] SQL生成与验证通过(1次尝试) +2026-04-17 10:49:41 INFO [__main__] api_server.py:916 _chat_stream_events() | [API/stream] Text2SQL 完成: valid=True attempts=1 tables_used=['VCBAccountCash', 'BCAccountCash', 'VCAccountBalance', 'VCUserAccessibleFunction'] sql_chars=178 sql_head='SELECT\n a.AccountID,\n a.AccountName AS AccountName,\n a.AvailableBalance,\n a.MarketValue,\n a.LedgerBalance\nFROM VCAccountBalance a\nORDER BY a.AvailableBalance DESC;' +2026-04-17 10:53:01 INFO [uvicorn.access] httptools_impl.py:483 send() | 127.0.0.1:24658 - "POST /g3sb/api/nl/chat/stream HTTP/1.1" 200 +2026-04-17 10:53:01 INFO [__main__] api_server.py:735 _chat_stream_events() | [API/stream] 开始: user_id='anonymous' visitor_biz_id=None session_id=None service_code=None model='gpt-4o-mini' lang_code='auto' msg_chars=23 preview="查询账户'M050013'持有的所有股票及数量" dialog_context_chars=0 last_turn_was_data_query=False +2026-04-17 10:53:04 INFO [llm.openai_client] openai_client.py:55 __init__() | [OK] OpenAIClient初始化: model=gpt-4o-mini, base_url=http://113.192.49.54:9080/v1 +2026-04-17 10:53:04 INFO [utils.dialog_classifier] dialog_classifier.py:261 classify_dialog() | [dialog] intent=text2sql (hybrid fast: query hint) preview="查询账户'M050013'持有的所有股票及数量" +2026-04-17 10:53:04 INFO [__main__] api_server.py:801 _chat_stream_events() | [GEN/API/stream] dialect=tsql top_k=20 dialog_context_chars=0 question_len=23 preview="查询账户'M050013'持有的所有股票及数量" +2026-04-17 10:53:07 INFO [llm.openai_client] openai_client.py:55 __init__() | [OK] OpenAIClient初始化: model=gpt-4o-mini, base_url=http://113.192.49.54:9080/v1 +2026-04-17 10:53:08 INFO [agents.orchestrator] orchestrator.py:827 generate() | [GEN] 问句已归一中文:查询账户'M050013'持有的所有股票及其数量 +2026-04-17 10:53:08 INFO [agents.orchestrator] orchestrator.py:840 generate() | [GEN] 开始生成SQL: question_chars=24 preview="查询账户'M050013'持有的所有股票及其数量" dialog_context_chars=0 +2026-04-17 10:53:08 INFO [agents.orchestrator] orchestrator.py:855 generate() | 尝试 #1 +2026-04-17 10:53:09 INFO [schema.indexer] indexer.py:106 ensure_index_for_schema() | Schema 向量索引已就绪(2516 张表),跳过向量化 +2026-04-17 10:53:09 INFO [agents.orchestrator] orchestrator.py:188 _coarse_filter() | [Orchestrator] 开始向量检索: query_chars=24 query_preview="查询账户'M050013'持有的所有股票及其数量" +2026-04-17 10:53:10 INFO [schema.indexer] indexer.py:244 search() | Schema 向量检索: query_chars=24 命中=20(阈值=0.1)top=[('VSBHKRpt0695A', 0.8153), ('VSBHKRpt0365A', 0.8114), ('VSBHKRpt1330', 0.809), ('VSBHKRPT0156', 0.8061), ('VSBHKRpt0672', 0.8053), ('VSBHKRpt0390D', 0.8051), ('VSBHKRPT0590', 0.8046), ('VSBRpt0055U_AccountInformation', 0.8038), ('VSBHKRpt4000A', 0.8028), ('VSBHKRptMSTATInvestmentPortfolioNoHoldings', 0.8026), ('VSBHKRpt1150', 0.8022), ('VSBRpt0055M_AccountInformation', 0.8019), ('VSBRpt0055U_Stock', 0.8017), ('VSBRpt0302V_AccountInformation', 0.8017), ('TSBYellowFormShareAllotment', 0.8016)] +2026-04-17 10:53:10 INFO [agents.orchestrator] orchestrator.py:204 _coarse_filter() | [Orchestrator] 向量粗筛: 命中=20 张(阈值内),表名+分: [('VSBHKRpt0695A', 0.8153), ('VSBHKRpt0365A', 0.8114), ('VSBHKRpt1330', 0.809), ('VSBHKRPT0156', 0.8061), ('VSBHKRpt0672', 0.8053), ('VSBHKRpt0390D', 0.8051), ('VSBHKRPT0590', 0.8046), ('VSBRpt0055U_AccountInformation', 0.8038), ('VSBHKRpt4000A', 0.8028), ('VSBHKRptMSTATInvestmentPortfolioNoHoldings', 0.8026), ('VSBHKRpt1150', 0.8022), ('VSBRpt0055M_AccountInformation', 0.8019), ('VSBRpt0055U_Stock', 0.8017), ('VSBRpt0302V_AccountInformation', 0.8017), ('TSBYellowFormShareAllotment', 0.8016), ('VSBHKRpt0450B', 0.8012), ('VSBHKRpt0390', 0.8009), ('VSBRpt0060', 0.8006), ('TSBAccountEntitlement', 0.8006), ('VSBHKRPT0157', 0.8004)] +2026-04-17 10:53:14 INFO [agents.orchestrator] orchestrator.py:258 _llm_select_tables() | LLM精筛选中表:['VSBHKRpt0390D', 'VSBHKRpt0390', 'VSBHKRpt0672'] | reasoning_chars=119 reasoning_preview="用户查询账户'M050013'持有的所有股票及其数量,VSBHKRpt0390D和VSBHKRpt0390均为账户持仓相关表,包含账户及其持有的股票信息,VSBHKRpt0672提供账户的基本信息,三者可以通过账户ID关联,满足查询需求。" +2026-04-17 10:53:14 INFO [agents.orchestrator] orchestrator.py:320 _prioritize_vc_user_accessible_function() | 已优先纳入目录视图 VCUserAccessibleFunction(置于选表前列),当前选表:['VCUserAccessibleFunction', 'VSBHKRpt0390D', 'VSBHKRpt0390', 'VSBHKRpt0672'] +2026-04-17 10:53:14 INFO [agents.orchestrator] orchestrator.py:894 generate() | 选中表:['VCUserAccessibleFunction', 'VSBHKRpt0390D', 'VSBHKRpt0390', 'VSBHKRpt0672'],扩展后:['VSBHKRpt0390D', 'VSBHKRpt0390', 'VSBHKRpt0672', 'VCUserAccessibleFunction'] +2026-04-17 10:53:15 INFO [utils.fewshot_selector] fewshot_selector.py:209 select_best_with_score() | [Few-shot] best_with_score: qid=Q2 score=0.8459 preview='列出今日所有客户账户之间的股票转移记录。' +2026-04-17 10:53:16 INFO [utils.fewshot_selector] fewshot_selector.py:357 _select_chroma() | Few-shot(Chroma)选择: 问题='查询账户'M050013'持有的所有股票及其数量...' → 选中3个示例 (top_k=3, min_rating=7) +2026-04-17 10:53:16 INFO [agents.orchestrator] orchestrator.py:568 _generate_sql() | 已注入 3 个 few-shot 示例: qid=['Q2', 'Q32', 'Q40'] question_zh_preview=['列出今日所有客户账户之间的股票转移记录。', '各市场客户不同资产(如股票、债券、基金、虚拟资产)的总持仓是多少?', '列出今日执行的所有股票转移/移动。'] +2026-04-17 10:53:18 INFO [agents.orchestrator] orchestrator.py:635 _generate_sql() | 生成的SQL(chars=169): +SELECT + h.AccountID, + h.InstrumentID, + h.InstrumentName AS InstrumentSymbol, + h.HoldQuantity AS HoldingQty +FROM VSBHKRpt0390 h +WHERE h.AccountID = 'M050013'; +2026-04-17 10:53:18 INFO [db.dbhub_tools] dbhub_tools.py:627 _execute_sql() | _execute_sql 执行语句数=1 readonly=True max_rows=1 +2026-04-17 10:53:18 INFO [agents.orchestrator] orchestrator.py:778 _validate_sql() | [validate] 程序+探针+LLM 汇总: valid=True err_count=0 warn_count=0 db_execution_status=1 sql_chars=169 +2026-04-17 10:53:18 INFO [agents.orchestrator] orchestrator.py:956 generate() | [OK] SQL生成与验证通过(1次尝试) +2026-04-17 10:53:20 INFO [__main__] api_server.py:916 _chat_stream_events() | [API/stream] Text2SQL 完成: valid=True attempts=1 tables_used=['VSBHKRpt0390D', 'VSBHKRpt0390', 'VSBHKRpt0672', 'VCUserAccessibleFunction'] sql_chars=169 sql_head="SELECT\n h.AccountID,\n h.InstrumentID,\n h.InstrumentName AS InstrumentSymbol,\n h.HoldQuantity AS HoldingQty\nFROM VSBHKRpt0390 h\nWHERE h.AccountID = 'M050013';" +2026-04-17 15:07:22 INFO [uvicorn.error] server.py:272 shutdown() | Shutting down +2026-04-17 15:07:23 INFO [uvicorn.error] server.py:102 _serve() | Finished server process [38600] +2026-04-17 15:07:23 ERROR [uvicorn.error] on.py:134 send() | Traceback (most recent call last): + File "D:\conda\Lib\asyncio\runners.py", line 194, in run + return runner.run(main) + ^^^^^^^^^^^^^^^^ + File "D:\conda\Lib\asyncio\runners.py", line 118, in run + return self._loop.run_until_complete(task) + ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + File "D:\conda\Lib\asyncio\base_events.py", line 674, in run_until_complete + self.run_forever() + File "D:\conda\Lib\asyncio\windows_events.py", line 322, in run_forever + super().run_forever() + File "D:\conda\Lib\asyncio\base_events.py", line 641, in run_forever + self._run_once() + File "D:\conda\Lib\asyncio\base_events.py", line 1986, in _run_once + handle._run() + File "D:\conda\Lib\asyncio\events.py", line 88, in _run + self._context.run(self._callback, *self._args) + File "C:\Users\24019\AppData\Roaming\Python\Python312\site-packages\uvicorn\server.py", line 78, in serve + with self.capture_signals(): + ^^^^^^^^^^^^^^^^^^^^^^ + File "D:\conda\Lib\contextlib.py", line 144, in __exit__ + next(self.gen) + File "C:\Users\24019\AppData\Roaming\Python\Python312\site-packages\uvicorn\server.py", line 339, in capture_signals + signal.raise_signal(captured_signal) + File "D:\conda\Lib\asyncio\runners.py", line 157, in _on_sigint + raise KeyboardInterrupt() +KeyboardInterrupt + +During handling of the above exception, another exception occurred: + +Traceback (most recent call last): + File "C:\Users\24019\AppData\Roaming\Python\Python312\site-packages\starlette\routing.py", line 645, in lifespan + await receive() + File "C:\Users\24019\AppData\Roaming\Python\Python312\site-packages\uvicorn\lifespan\on.py", line 137, in receive + return await self.receive_queue.get() + ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + File "D:\conda\Lib\asyncio\queues.py", line 158, in get + await getter +asyncio.exceptions.CancelledError +