feat: 优化 pdf skill

This commit is contained in:
hp0912 2026-07-25 20:19:02 +08:00
parent 06df60d706
commit afb487260a
5 changed files with 1230 additions and 67 deletions

View File

@ -1,6 +1,6 @@
---
name: pdf
description: "处理本地 PDF 文件或远程 HTTPS PDF 链接,包括安全下载、元数据与页面检查、分段文本和表格提取、页面 PNG 渲染、从文本创建 PDF、合并、拆分、旋转及最终质量校验。当用户提供 .pdf 文件或 HTTPS PDF 地址,或要求总结、读取、生成、编辑、转换或审阅 PDF 时使用。"
description: "处理本地 PDF 文件或远程 HTTPS PDF 链接,包括安全下载、元数据与页面检查、多引擎分段文本提取质量检测、表格提取、按需页面 PNG 渲染、从文本创建 PDF、合并、拆分、旋转及最终质量校验。当用户提供 .pdf 文件或 HTTPS PDF 地址,或要求总结、读取、生成、编辑、转换或审阅 PDF 时使用。"
---
# PDF 处理
@ -15,6 +15,8 @@ description: "处理本地 PDF 文件或远程 HTTPS PDF 链接,包括安全
- 不创建或猜测脚本清单以外的文件。
- 每次检查脚本返回的 JSON只有 `ok``true` 时才继续。
- 收到 `ok: false` 时,依据 `error` 调整合法参数或向用户说明失败原因,不要把参数改传给其他脚本碰运气。
- 阅读或总结时只使用 `pages[]``usable_for_summary: true` 的文本。`needs_ocr: false` 时不得为了“常规检查”继续渲染或调用图片识别。
- 不把 PDF 密码作为脚本参数;工具调用参数可能进入运行日志。
## 脚本清单
@ -22,7 +24,7 @@ description: "处理本地 PDF 文件或远程 HTTPS PDF 链接,包括安全
| --- | --- | --- |
| `scripts/download_pdf.py` | 下载并校验远程 HTTPS PDF | `urllib`、`pypdf` |
| `scripts/inspect_pdf.py` | 检查页数、加密、元数据、页面尺寸和表单数量 | `pypdf` |
| `scripts/extract_text.py` | 按页、按字符分段提取正文 | `pdfplumber` |
| `scripts/extract_text.py` | 多引擎提取、质量检测并分段返回正文 | Poppler `pdftotext`、`pdfplumber``pypdf` 校验 |
| `scripts/extract_tables.py` | 按页提取表格 | `pdfplumber` |
| `scripts/render_pdf.py` | 把指定页面渲染为 PNG | Poppler `pdftoppm` |
| `scripts/create_pdf.py` | 从 UTF-8 文本或 Markdown 创建 PDF | `reportlab`、`pypdf` |
@ -35,10 +37,10 @@ description: "处理本地 PDF 文件或远程 HTTPS PDF 链接,包括安全
1. 为任务选择简短目录名,把中间文件放在 `tmp/pdfs/<任务名>/`
2. 远程 HTTPS 链接先调用 `download_pdf.py`;本地文件直接进入下一步。
3. 调用 `inspect_pdf.py` 检查文件。遇到加密 PDF 且任务需要正文时,向用户索取密码,不要绕过加密
4. 阅读或总结时,循环调用 `extract_text.py` 直至 `has_more``false`;需要表格时再调用 `extract_tables.py`
5. 调用 `render_pdf.py` 检查首页、复杂页面、无文本页面和用户关心的页面。创建或修改 PDF 时分批渲染全部页面
6. 使用提取文本和渲染结果完成回答;明确区分原文内容与基于版式或图表的推断
3. 调用 `inspect_pdf.py` 检查文件。遇到加密 PDF 时停止处理,请用户提供已解密副本;当前固定脚本不接收密码
4. 阅读或总结时调用 `extract_text.py`。结果为 `usable_for_summary: true` 时使用可靠页文本并根据游标继续;同时为 `needs_ocr: false` 时直接回答,不调用渲染或图片识别
5. 仅在文本结果为 `needs_ocr: true`、用户明确要求检查视觉版式,或任务涉及创建/修改 PDF 时调用 `render_pdf.py`
6. 文本兜底时先只渲染 1 个有代表性的页面确认是否为扫描内容;确认后才按任务所需页面继续渲染和识别,不默认渲染前 5 页
7. 创建或修改后的最终 PDF 写入 `output/pdf/`,重新执行检查、文本提取和全部页面渲染。
8. 最终产物位于临时目录之外且不再需要缓存时,调用 `cleanup_pdf_temp.py` 清理本次任务目录。
@ -78,16 +80,24 @@ description: "处理本地 PDF 文件或远程 HTTPS PDF 链接,包括安全
--input 'tmp/pdfs/<任务名>/source.pdf'
```
默认使用 `auto` 引擎:先由 Poppler `pdftotext` 提取;结果不可用或命令不可用时自动尝试 `pdfplumber`,并可逐页选择质量更好的结果。脚本使用 `pypdf` 获取标准页数,并拒绝把页数不一致的提取结果当作成功。不要直接执行 `pdftotext`
默认单次最多处理 8 页、返回 24000 个字符。可使用:
- `--start-page <页码>`、`--end-page <页码>`:页码从 `1` 开始。
- `--start-offset <字符偏移>`:继续读取被字符上限截断的同一页。
- `--start-offset <字符偏移>`:继续读取被字符上限截断的同一页;大于 `0` 时同时传入上次返回的 `next_engine`
- `--max-pages <页数>`、`--max-chars <字符数>`:控制单次输出。
- `--layout`:仅在需要尽量保留版面空格时使用。
- `--engine <auto|poppler|pdfplumber>`:首次及跨页提取保持 `auto`;同页字符续读时传入上次返回的 `next_engine`
- `--timeout <秒>`Poppler 提取超时,默认 `120`
如果结果中 `has_more: true`,继续调用同一脚本,并把返回的 `next_page` 传给 `--start-page`、`next_offset` 传给 `--start-offset`。后续调用必须保留首次调用的 `--end-page`(如果指定)及其他提取选项。重复直到 `has_more: false`。不要只读取第一批文本就总结长文档。
先检查 `usable_for_summary``text_quality.status`
提取文本为空或很少时,将页面视为可能的扫描件,改用渲染图读取;不要误判为空白 PDF。
- `usable_for_summary: true`:只使用 `pages[]` 中同样标为 `usable_for_summary: true``text`;可疑页的文本会被置空。如果 `has_more: true`,始终传回 `next_page``next_offset`。仅当 `next_offset` 大于 `0` 时,把非空的 `next_engine` 传给 `--engine` 以固定同页字符游标;这种调用只续读当前页。当前页完成后返回的 `next_offset``0`,此时不要传 `--engine`,让下一页重新使用 `auto`。保留首次调用的 `--end-page`(如果指定)及其他选项,直至 `has_more: false`
- `usable_for_summary: false`:本批次没有可靠文本,不要使用返回内容。查看 `engine_attempts`、`text_quality.reasons`、`text_quality.suspect_pages` 和 `needs_ocr`;若 `has_more: true`,仍按跨页游标继续检查后续批次,避免漏掉后续可搜索文本。
- `needs_ocr: true`:一个或多个页面未得到可靠文本。优先选择 `text_quality.suspect_pages` 中的首个页码,只渲染该页诊断;确认确为扫描内容后才调用图片识别。
`complete_text_coverage: true` 表示本批次所有页面均有可靠文本。`text_quality` 按页检测空白或过少文本、页面实际可见图像覆盖过大但文字不足、`(cid:...)`、Unicode 替换字符、异常控制字符及外观像汉字的部首字符;`pages[].extractor` 表示该页最终采用的引擎。`status: mixed` 表示同一批次同时包含可靠页和可疑页:可先使用可靠页文本,同时只核验 `suspect_pages`。不要只根据“肉眼看起来能读”判定提取结果可靠。
## 提取表格
@ -101,7 +111,13 @@ description: "处理本地 PDF 文件或远程 HTTPS PDF 链接,包括安全
## 渲染页面
调用 `scripts/render_pdf.py`,不要直接执行 `pdftoppm`
只有满足以下任一条件时才调用 `scripts/render_pdf.py`
- `extract_text.py` 返回 `needs_ocr: true`
- 用户明确要求审阅版式、图表、印章、公式或页面外观;
- 创建或修改 PDF 后进行最终视觉检查。
不要因为输入是 PDF、需要总结或需要检查首页就自动渲染。调用脚本时不要直接执行 `pdftoppm`
```text
--input 'tmp/pdfs/<任务名>/source.pdf' --output-dir 'tmp/pdfs/<任务名>/rendered' --start-page 1
@ -109,7 +125,7 @@ description: "处理本地 PDF 文件或远程 HTTPS PDF 链接,包括安全
默认 150 DPI、单次最多 10 页。可使用 `--end-page`、`--max-pages`、`--dpi`、`--timeout` 和 `--overwrite`。若 `has_more: true`,使用 `next_page` 继续,并保留首次调用的 `--end-page`(如果指定)、输出目录及其他渲染选项。脚本返回标准化的 `page-0001.png` 文件路径。
对文字较小或图表密集的页面提高 DPI。使用可用的图像查看或识别工具检查返回的 PNG不要尝试把图片路径交给下载脚本。
文本兜底首次只传一个页码,例如同时设置 `--start-page 2 --end-page 2 --max-pages 1`对文字较小或图表密集的页面提高 DPI。使用可用的图像查看或识别工具检查返回的 PNG不要尝试把图片路径交给下载脚本。
## 创建 PDF

View File

@ -1,4 +1,4 @@
interface:
display_name: "PDF 处理"
short_description: "读取、创建、审阅与可视化检查本地或远程 PDF 文件"
default_prompt: "使用 $pdf 下载并审阅这个 PDF提取要点并检查页面布局。"
short_description: "读取、创建和审阅本地或远程 PDF文本优先并按需检查版式"
default_prompt: "使用 $pdf 下载或读取这个 PDF优先提取可靠文本并按我的要求处理内容。"

View File

@ -4,12 +4,22 @@ from __future__ import annotations
import argparse
import json
import logging
import os
import tempfile
from pathlib import Path
from typing import Any, Callable, NoReturn, Optional
def quiet_pdf_library_logs() -> None:
"""Keep third-party recovery warnings out of the JSON tool response."""
logging.getLogger("pdfminer").setLevel(logging.ERROR)
logging.getLogger("pypdf").setLevel(logging.ERROR)
quiet_pdf_library_logs()
class SkillArgumentParser(argparse.ArgumentParser):
def error(self, message: str) -> NoReturn:
raise ValueError(f"参数错误:{message}")

View File

@ -63,10 +63,21 @@ def _clean_table(table) -> list[list[str]]:
]
def _page_count(path) -> int:
from pypdf import PdfReader
with path.open("rb") as stream:
reader = PdfReader(stream, strict=False)
if reader.is_encrypted:
raise ValueError("PDF 已加密,无法提取表格")
return len(reader.pages)
def _extract(args) -> dict[str, Any]:
import pdfplumber
path = input_pdf(args.input)
page_count = _page_count(path)
result_pages: list[dict[str, Any]] = []
table_count = 0
cell_count = 0
@ -74,7 +85,13 @@ def _extract(args) -> dict[str, Any]:
next_table = 0
with pdfplumber.open(path) as pdf:
page_count = len(pdf.pages)
observed_page_count = len(pdf.pages)
if observed_page_count != page_count:
raise RuntimeError(
"pdfplumber 页数与 PDF 检查结果不一致:"
f"{observed_page_count} != {page_count}"
"无法可靠提取表格"
)
start_page, actual_end, next_page = selected_page_window(
page_count,
args.start_page,

File diff suppressed because it is too large Load Diff