feat: 增强 pdf 技能
This commit is contained in:
parent
33809a79e4
commit
67351ccbdd
@ -1,233 +1,55 @@
|
||||
---
|
||||
name: pdf
|
||||
description: "处理本地 PDF 文件或远程 HTTPS PDF 链接,并下载 PDF 任务所需且不超过 25 MiB 的图片、音视频、压缩包和其他 HTTPS 附件;包括源 PDF 安全下载、元数据与页面检查、多引擎分段文本提取和质量检测、扫描页本地 OCR、表格提取、按需页面 PNG 渲染、从文本创建 PDF、合并、拆分、旋转及最终质量校验。当用户提供 .pdf 文件或 HTTPS PDF 地址,或要求总结、读取、识别扫描件、生成、编辑、转换或审阅 PDF 时使用。"
|
||||
description: "读取、OCR、创建、设计排版、转换和处理 PDF。支持本地文件与 HTTPS PDF、扫描件和任务图像,以原生文本提取及本地 OCR 为主,模型辅助疑难复核;提供 HTML/CSS 出版排版、文本生成、Office/LaTeX 导出、表单、页面与元数据操作。用户要求阅读或总结 PDF、识别扫描件、制作报告/简历/提案 PDF 或编辑现有 PDF 时使用;Office 主文档编辑由对应 skill 处理。"
|
||||
---
|
||||
|
||||
# PDF 处理
|
||||
|
||||
## 强制执行规则
|
||||
|
||||
当前智能体不能执行 shell、任意 Python 代码或系统命令。只能通过 `execute_skill_script` 调用本 Skill 中真实存在的固定脚本。
|
||||
|
||||
- 只调用下表列出的可执行脚本。
|
||||
- 不执行 `scripts/` 目录,不执行内部模块 `scripts/_pdf_common.py`。
|
||||
- 不传 `-c`、`python3`、`ls`、`pdftoppm` 或其他 shell/系统命令作为脚本参数。
|
||||
- 不创建或猜测脚本清单以外的文件。
|
||||
- 每次检查脚本返回的 JSON;只有 `ok` 为 `true` 时才继续。
|
||||
- 收到 `ok: false` 时,依据 `error` 调整合法参数或向用户说明失败原因,不要把参数改传给其他脚本碰运气。
|
||||
- 阅读或总结时只使用 `pages[]` 中 `usable_for_summary: true` 的文本。`needs_ocr: false` 时不得为了“常规检查”继续 OCR、渲染或调用图片识别。
|
||||
- 远程 PDF 源文件只交给 `download_pdf.py`;任务所需的远程图片、视频、音频、压缩包或其他附件只交给 `download_attachment.py`。不要在回复、日志摘要或文件名中复述可能含敏感查询参数的完整 URL。
|
||||
- PDF 最终文件一律写入 `/usr/local/src/pdf/`,下载缓存、中间文件和渲染结果一律写入 `/usr/local/src/pdf/tmp/<任务名>/`。始终传绝对路径;固定脚本会自动创建目录并拒绝该根目录之外的输出。
|
||||
- 不把 PDF 密码作为脚本参数;工具调用参数可能进入运行日志。
|
||||
|
||||
## 脚本清单
|
||||
|
||||
| 脚本 | 用途 | 底层能力 |
|
||||
| --- | --- | --- |
|
||||
| `scripts/download_pdf.py` | 下载并校验远程 HTTPS PDF | `urllib`、`pypdf` |
|
||||
| `scripts/download_attachment.py` | 下载图片、音视频、压缩包等通用 HTTPS 附件 | `urllib`、HEAD 大小探测、流式硬限制 |
|
||||
| `scripts/inspect_pdf.py` | 检查页数、加密、元数据、页面尺寸和表单数量 | `pypdf` |
|
||||
| `scripts/extract_text.py` | 多引擎提取、质量检测并分段返回正文 | Poppler `pdftotext`、`pdfplumber`;`pypdf` 校验 |
|
||||
| `scripts/ocr_text.py` | 对指定扫描页执行离线 OCR 并返回可靠文字 | RapidOCR、ONNX Runtime、Poppler `pdftoppm` |
|
||||
| `scripts/extract_tables.py` | 按页提取表格 | `pdfplumber` |
|
||||
| `scripts/render_pdf.py` | 把指定页面渲染为 PNG | Poppler `pdftoppm` |
|
||||
| `scripts/create_pdf.py` | 从 UTF-8 文本或 Markdown 创建 PDF | `reportlab`、`pypdf` |
|
||||
| `scripts/manage_pdf.py` | 合并、拆分或旋转 PDF | `pypdf` |
|
||||
| `scripts/cleanup_pdf_temp.py` | 安全删除本次任务临时目录 | Python 文件 API |
|
||||
|
||||
环境已预置所有依赖。不要安装依赖,也不要提示用户安装依赖。
|
||||
|
||||
## 标准流程
|
||||
|
||||
1. 为任务选择简短目录名,把中间文件放在 `/usr/local/src/pdf/tmp/<任务名>/`。
|
||||
2. 远程 HTTPS 链接先调用 `download_pdf.py`;本地文件直接进入下一步。
|
||||
3. 调用 `inspect_pdf.py` 检查文件。遇到加密 PDF 时停止处理,请用户提供已解密副本;当前固定脚本不接收密码。
|
||||
4. 阅读或总结时调用 `extract_text.py`。结果为 `usable_for_summary: true` 时使用可靠页文本并根据游标继续;同时为 `needs_ocr: false` 时直接回答,不调用 OCR、渲染或图片识别。
|
||||
5. 只有 `extract_text.py` 返回 `needs_ocr: true` 时,才对 `text_quality.suspect_pages` 调用 `ocr_text.py`。原生可靠文本优先,OCR 只补齐可疑页,不重复识别正常页。
|
||||
6. `ocr_text.py` 会在脚本内部临时渲染指定页面并交给本地 RapidOCR,完成后自动删除 PNG;普通扫描件解析不调用大模型识图,也不需要先调用 `render_pdf.py`。
|
||||
7. 仅在用户明确要求检查视觉版式,或任务涉及创建/修改 PDF 时调用 `render_pdf.py`。
|
||||
8. 创建或修改后的最终 PDF 写入 `/usr/local/src/pdf/`,重新执行检查、文本提取和全部页面渲染。
|
||||
9. 最终产物位于临时目录之外且不再需要缓存时,调用 `cleanup_pdf_temp.py` 清理本次任务目录。
|
||||
|
||||
## 下载远程 PDF
|
||||
|
||||
只接受 HTTPS 地址。完整保留 URL 及查询参数,不在回复、日志摘要或文件名中复述敏感参数。
|
||||
|
||||
调用 `scripts/download_pdf.py`:
|
||||
|
||||
```text
|
||||
--url 'https://example.com/document.pdf' --output '/usr/local/src/pdf/tmp/<任务名>/source.pdf'
|
||||
```
|
||||
|
||||
可选参数:
|
||||
|
||||
- `--timeout <秒>`:默认 `60`。
|
||||
- `--max-bytes <字节数>`:默认且最高 `26214400`(25 MiB),只允许设置更小的限制。
|
||||
- `--overwrite`:仅在目标是本次任务生成的缓存时使用。
|
||||
|
||||
脚本会创建父目录、流式下载、阻止 HTTPS 重定向降级到 HTTP,并验证 PDF。成功结果包含 `path`、`size_bytes`、`page_count` 和 `encrypted`。
|
||||
|
||||
## 下载通用附件
|
||||
|
||||
需要下载作为 PDF 任务素材的图片、视频、音频、压缩包或其他文件时,调用 `scripts/download_attachment.py`:
|
||||
|
||||
```text
|
||||
--url 'https://example.com/asset.bin?signature=...' --output '/usr/local/src/pdf/tmp/<任务名>/asset.bin'
|
||||
```
|
||||
|
||||
只接受 HTTPS 地址,`output` 可使用任意附件扩展名。可选参数只有 `--timeout <1-600>`(默认 `60`)和 `--overwrite`。附件上限固定为 25 MiB(26214400 字节),不可调高:脚本先用 HEAD 探测远端声明大小,再检查 GET 响应声明,并在流式接收时持续兜底计数;任一阶段发现超限都会返回 `ok: false` 和明确的“已拒绝下载”错误,且不会发布部分文件。
|
||||
|
||||
成功结果包含 `path`、实际 `size_bytes`、`declared_size_bytes`、`size_limit_bytes`、`size_probe` 和 `content_type`。本脚本不校验文件业务格式;远程 PDF 源文件仍使用 `download_pdf.py`。
|
||||
|
||||
## 检查 PDF
|
||||
|
||||
调用 `scripts/inspect_pdf.py`:
|
||||
|
||||
```text
|
||||
--input '/usr/local/src/pdf/tmp/<任务名>/source.pdf'
|
||||
```
|
||||
|
||||
使用返回的 `page_count`、`encrypted`、`metadata`、`page_layouts` 和 `form_field_count` 判断后续处理方式。不要直接调用 `pdfinfo`。
|
||||
|
||||
## 提取正文
|
||||
|
||||
首次调用 `scripts/extract_text.py`:
|
||||
|
||||
```text
|
||||
--input '/usr/local/src/pdf/tmp/<任务名>/source.pdf'
|
||||
```
|
||||
|
||||
默认使用 `auto` 引擎:先由 Poppler `pdftotext` 提取;结果不可用或命令不可用时自动尝试 `pdfplumber`,并可逐页选择质量更好的结果。脚本使用 `pypdf` 获取标准页数,并拒绝把页数不一致的提取结果当作成功。不要直接执行 `pdftotext`。
|
||||
|
||||
默认单次最多处理 8 页、返回 24000 个字符。可使用:
|
||||
|
||||
- `--start-page <页码>`、`--end-page <页码>`:页码从 `1` 开始。
|
||||
- `--start-offset <字符偏移>`:继续读取被字符上限截断的同一页;大于 `0` 时同时传入上次返回的 `next_engine`。
|
||||
- `--max-pages <页数>`、`--max-chars <字符数>`:控制单次输出。
|
||||
- `--layout`:仅在需要尽量保留版面空格时使用。
|
||||
- `--engine <auto|poppler|pdfplumber>`:首次及跨页提取保持 `auto`;同页字符续读时传入上次返回的 `next_engine`。
|
||||
- `--timeout <秒>`:Poppler 提取超时,默认 `120`。
|
||||
|
||||
先检查 `usable_for_summary` 和 `text_quality.status`:
|
||||
|
||||
- `usable_for_summary: true`:只使用 `pages[]` 中同样标为 `usable_for_summary: true` 的 `text`;可疑页的文本会被置空。如果 `has_more: true`,始终传回 `next_page` 和 `next_offset`。仅当 `next_offset` 大于 `0` 时,把非空的 `next_engine` 传给 `--engine` 以固定同页字符游标;这种调用只续读当前页。当前页完成后返回的 `next_offset` 为 `0`,此时不要传 `--engine`,让下一页重新使用 `auto`。保留首次调用的 `--end-page`(如果指定)及其他选项,直至 `has_more: false`。
|
||||
- `usable_for_summary: false`:本批次没有可靠文本,不要使用返回内容。查看 `engine_attempts`、`text_quality.reasons`、`text_quality.suspect_pages` 和 `needs_ocr`;若 `has_more: true`,仍按跨页游标继续检查后续批次,避免漏掉后续可搜索文本。
|
||||
- `needs_ocr: true`:一个或多个页面未得到可靠文本。把 `text_quality.suspect_pages` 中实际需要阅读的页码传给 `ocr_text.py`;不要先调用 `render_pdf.py`,也不要把临时图片交给大模型。
|
||||
|
||||
`complete_text_coverage: true` 表示本批次所有页面均有可靠文本。`text_quality` 按页检测空白或过少文本、页面实际可见图像覆盖过大但文字不足、`(cid:...)`、Unicode 替换字符、异常控制字符及外观像汉字的部首字符;`pages[].extractor` 表示该页最终采用的引擎。`status: mixed` 表示同一批次同时包含可靠页和可疑页:可先使用可靠页文本,同时只核验 `suspect_pages`。不要只根据“肉眼看起来能读”判定提取结果可靠。
|
||||
|
||||
## 本地 OCR 扫描页
|
||||
|
||||
仅当 `extract_text.py` 返回 `needs_ocr: true` 时调用 `scripts/ocr_text.py`。`--pages` 必须明确指定 `text_quality.suspect_pages` 中要读取的页,单次最多 4 页:
|
||||
|
||||
```text
|
||||
--input '/usr/local/src/pdf/tmp/<任务名>/source.pdf' --pages '2,5-6'
|
||||
```
|
||||
|
||||
默认以 260 DPI 临时渲染,并使用镜像中预置的 RapidOCR 与 ONNX Runtime 在本地识别。脚本不会联网下载模型,不会保留渲染图片,也不会调用大模型视觉能力。可选参数:
|
||||
|
||||
- `--dpi <150-400>`:文字过小或识别质量不足时适度提高,默认 `260`。
|
||||
- `--max-chars <字符数>`:默认 `24000`,最大 `60000`。
|
||||
- `--timeout <秒>`:每页 Poppler 渲染超时,默认 `180`。
|
||||
- `--start-offset <字符偏移>`:续读被字符上限截断的单页;使用时 `--pages` 只能包含该页。
|
||||
|
||||
只使用 `pages[]` 中 `usable_for_summary: true` 的 `text`。`status: empty`、`sparse` 或 `low_confidence` 的页面文本会被置空,并通过 `needs_review: true` 提醒人工检查。
|
||||
|
||||
如果 `has_more: true`:
|
||||
|
||||
- `next_offset > 0`:用 `--pages <next_page> --start-offset <next_offset>` 续读同一页。
|
||||
- `next_offset = 0`:用返回的 `remaining_pages` 继续下一批。
|
||||
- 同页续读完成后,再处理先前返回的其他 `remaining_pages`。
|
||||
|
||||
OCR 结果中的 `mean_confidence`、`line_count`、`render_seconds` 和 `ocr_seconds` 仅用于判断质量与性能。原生提取成功的页面始终采用 `extract_text.py` 结果,不用 OCR 覆盖。
|
||||
|
||||
## 提取表格
|
||||
|
||||
调用 `scripts/extract_tables.py`:
|
||||
|
||||
```text
|
||||
--input '/usr/local/src/pdf/tmp/<任务名>/source.pdf' --start-page 1
|
||||
```
|
||||
|
||||
默认单次最多处理 5 页、20 个表格和 2000 个单元格。可用 `--end-page`、`--start-table`、`--max-pages`、`--max-tables`、`--max-cells` 调整。若 `has_more: true`,把 `next_page` 传给 `--start-page`、`next_table` 传给 `--start-table` 后继续,并保留首次调用的 `--end-page`(如果指定)及其他提取选项。
|
||||
|
||||
## 渲染页面
|
||||
|
||||
只有满足以下任一条件时才调用 `scripts/render_pdf.py`:
|
||||
|
||||
- 用户明确要求审阅版式、图表、印章、公式或页面外观;
|
||||
- 创建或修改 PDF 后进行最终视觉检查。
|
||||
|
||||
不要因为输入是 PDF、需要总结、需要 OCR 或需要检查首页就自动调用本脚本;OCR 的临时渲染由 `ocr_text.py` 内部完成。调用脚本时不要直接执行 `pdftoppm`:
|
||||
|
||||
```text
|
||||
--input '/usr/local/src/pdf/tmp/<任务名>/source.pdf' --output-dir '/usr/local/src/pdf/tmp/<任务名>/rendered' --start-page 1
|
||||
```
|
||||
|
||||
默认 150 DPI、单次最多 10 页。可使用 `--end-page`、`--max-pages`、`--dpi`、`--timeout` 和 `--overwrite`。若 `has_more: true`,使用 `next_page` 继续,并保留首次调用的 `--end-page`(如果指定)、输出目录及其他渲染选项。脚本返回标准化的 `page-0001.png` 文件路径。
|
||||
|
||||
对文字较小或图表密集的页面提高 DPI。使用可用的图像查看工具检查返回的 PNG,不要尝试把图片路径交给下载脚本。
|
||||
|
||||
## 创建 PDF
|
||||
|
||||
先使用 `write_file` 把内容写为 UTF-8 `.txt` 或 `.md` 文件,再调用 `scripts/create_pdf.py`:
|
||||
|
||||
```text
|
||||
--input '/usr/local/src/pdf/tmp/<任务名>/content.md' --output '/usr/local/src/pdf/<文件名>.pdf' --title '文档标题'
|
||||
```
|
||||
|
||||
脚本支持 Markdown 标题、项目符号和简单表格,自动选择可嵌入的 Unicode 字体并添加页码。可选参数:
|
||||
|
||||
- `--page-size <A4|LETTER>`
|
||||
- `--font-path <TTF或TTC路径>`
|
||||
- `--font-size <字号>`
|
||||
- `--margin <points>`
|
||||
- `--overwrite`
|
||||
|
||||
输入内容只使用 ASCII 连字符 `-`;脚本也会把常见 Unicode 横线规范化为 ASCII 连字符。
|
||||
|
||||
## 合并、拆分与旋转
|
||||
|
||||
调用 `scripts/manage_pdf.py`,第一个参数必须是操作名。
|
||||
|
||||
合并:
|
||||
|
||||
```text
|
||||
merge --input 'a.pdf' --input 'b.pdf' --output '/usr/local/src/pdf/merged.pdf'
|
||||
```
|
||||
|
||||
拆分指定范围:
|
||||
|
||||
```text
|
||||
split --input 'source.pdf' --output-dir '/usr/local/src/pdf/split' --range 1-3 --range 4-6
|
||||
```
|
||||
|
||||
不传 `--range` 时每页生成一个 PDF。
|
||||
|
||||
旋转指定页面:
|
||||
|
||||
```text
|
||||
rotate --input 'source.pdf' --output '/usr/local/src/pdf/rotated.pdf' --pages '1,3-5' --degrees 90
|
||||
```
|
||||
|
||||
`--degrees` 只能是 `90`、`180` 或 `270`;不传 `--pages` 时旋转全部页面。目标已存在且确认可覆盖时添加 `--overwrite`。
|
||||
|
||||
## 清理临时目录
|
||||
|
||||
调用 `scripts/cleanup_pdf_temp.py`:
|
||||
|
||||
```text
|
||||
--task-dir '/usr/local/src/pdf/tmp/<任务名>'
|
||||
```
|
||||
|
||||
脚本只允许删除 `/usr/local/src/pdf/tmp/` 下一级任务目录,拒绝删除根目录、仓库目录或其他路径。
|
||||
|
||||
## 质量要求
|
||||
|
||||
- 不覆盖用户提供的源文件。
|
||||
- 创建或修改后重新检查页数、页面尺寸、加密状态和文本可读性。
|
||||
- 扫描件先做原生文字检测,再只 OCR 可疑页;不得把低置信度 OCR 文本当作可靠正文。
|
||||
- 逐页确认没有裁切、重叠、溢出、乱码、黑方块、错误分页或异常空白页。
|
||||
- 检查标题层级、段落间距、页边距、表格、图表、图片、页码及章节衔接。
|
||||
- 引用和参考文献必须可读,不得残留工具令牌、占位符或临时路径。
|
||||
- 只有最新渲染结果不存在可见缺陷时才交付创建或修改后的 PDF。
|
||||
# PDF 读取、设计与处理
|
||||
|
||||
## 运行约定
|
||||
|
||||
- 当前机器人不能直接运行 Bash、Python、Node.js 或系统命令。只通过 `execute_skill_script` 调用下列真实存在的固定脚本;参数是文件、文本、页码或受控 JSON,不能传 shell 命令、`-c`、`eval`、代码片段或解释器命令。固定脚本可在内部调用预置引擎,调用者不直接执行底层程序。
|
||||
- `scripts/_pdf_common.py`、`scripts/_render_html.cjs` 是内部实现,不直接执行;不调用原 pdf-skill 的 shell 安装器、通用 Python/Node CLI,不创建临时可执行脚本。
|
||||
- 依赖只在基础镜像构建时安装。任务中不运行 pip/npm/apt、不下载浏览器、TeX 包或 OCR 模型;缺失时报告具体依赖和镜像需更新,不能假装处理成功。
|
||||
- 最终 PDF 写入 `/usr/local/src/pdf/`;缓存、源 HTML/JSON、预览放在 `/usr/local/src/pdf/tmp/<任务名>/`。传绝对路径,保留用户源文件。`--overwrite` 仅用于本任务已生成的旧产物。
|
||||
- 每次检查返回 JSON;`ok: false` 先处理原因。`requires_visual_review`、`needs_review` 或 warnings 需要实际核验,程序运行成功不代表内容和版式通过。
|
||||
- HTTPS PDF 用 `download_pdf.py`,其他素材用 `download_attachment.py`。不在回复中复述敏感 URL 查询参数;加密 PDF 请用户提供已解密副本,密码不进入工具参数。
|
||||
|
||||
## 按任务读取
|
||||
|
||||
| 任务 | 指南 |
|
||||
| --- | --- |
|
||||
| 阅读、总结、扫描页、图片文字、图表辅助识别 | [读取与 OCR](references/reading-and-ocr.md) |
|
||||
| 创建报告、提案、简历、学术或品牌 PDF | [设计规范](references/design.md) + [HTML 创建接口](references/creation.md) |
|
||||
| Office/LaTeX 导出,PDF 内容重建为 Office | [转换](references/conversion.md) |
|
||||
| 表单填写、裁剪、嵌入图片、元数据 | [编辑](references/editing.md) |
|
||||
| 下载、分段提取、表格、简单文本 PDF、合并/拆分/旋转、渲染与清理 | [原有操作接口](references/operations.md) |
|
||||
| 镜像缺包或能力边界 | [依赖说明](references/dependencies.md) |
|
||||
|
||||
## 固定脚本
|
||||
|
||||
| 脚本 | 用途 |
|
||||
| --- | --- |
|
||||
| `scripts/download_pdf.py` | 安全下载并校验不超过 25 MiB 的 HTTPS PDF |
|
||||
| `scripts/download_attachment.py` | 下载不超过 25 MiB 的任务附件 |
|
||||
| `scripts/inspect_pdf.py` | 页数、尺寸、加密、元数据与表单数量 |
|
||||
| `scripts/extract_text.py` | 多引擎正文提取、逐页质量检测和字符游标 |
|
||||
| `scripts/ocr_text.py` | 对指定 PDF 页执行离线 OCR,标记局部疑难区域 |
|
||||
| `scripts/ocr_image.py` | 对单张任务图片执行离线 OCR |
|
||||
| `scripts/extract_tables.py` | 原生 PDF 表格分页提取 |
|
||||
| `scripts/render_pdf.py` | 按页输出 PNG,供版式检查或疑难辅助复核 |
|
||||
| `scripts/create_pdf.py` | 简单文本/Markdown 生成 PDF |
|
||||
| `scripts/create_design_pdf.py` | 静态 HTML/CSS、图表和公式设计排版 |
|
||||
| `scripts/convert_to_pdf.py` | Office 文件导出 PDF |
|
||||
| `scripts/compile_latex.py` | 仅用镜像缓存资源编译 LaTeX |
|
||||
| `scripts/edit_pdf.py` | 表单、裁剪、元数据和嵌入图片 |
|
||||
| `scripts/manage_pdf.py` | 合并、拆分、旋转 |
|
||||
| `scripts/cleanup_pdf_temp.py` | 清理本任务临时目录 |
|
||||
|
||||
## 工作原则
|
||||
|
||||
1. 阅读先检查 PDF,再提取可靠原生文本;扫描页或图片文字以本地 OCR 为主。只有低置信度、手写、复杂表格/公式、阅读顺序冲突或非文本图形理解需要时,才用大模型复核相关页/区域。不要把整个扫描件直接交给大模型代替 OCR。
|
||||
2. 创建设计先确认读者、用途、内容与输出限制,按需选择封面、配色和字体层级。用户模板、品牌、大纲、语言与篇幅优先;不强加独立封面,不为凑页数填充或删除内容。
|
||||
3. 简单文字选 `create_pdf.py`;需要封面、图文、页眉页脚、交叉引用、数学公式时选 `create_design_pdf.py`。通过 `write_file` 写静态内容文件,不写可执行代码。HTML 禁止脚本、事件处理程序和外部资源;公式、流程图由固定引擎本地处理。
|
||||
4. 编辑现有 PDF 保留内容与结构。裁剪不等于脱敏;表单字段值写入不等于外观正确;PDF 转 Office 应按提取/OCR 后重建来规划,不能承诺无损逆转换。
|
||||
5. 创建、转换或修改后重新检查页数、文本与关键数字,再渲染全部相关页逐页核验封面、字体、表格、公式、图表、页码、裁切和空白页。要求精确页数时使用 `--expected-pages`。修正后检查最新产物,才交付。
|
||||
6. 内容引用可核验。用户提供的材料可直接引用;新增时效、专业或不确定事实使用当前可用搜索工具查证,不编造统计、论文或参考文献。图片识别的猜测与原文分开标记。
|
||||
|
||||
@ -1,4 +1,4 @@
|
||||
interface:
|
||||
display_name: "PDF 处理"
|
||||
short_description: "读取、创建和审阅本地或远程 PDF,按需执行本地 OCR"
|
||||
default_prompt: "使用 $pdf 下载或读取这个 PDF,优先提取可靠文本,并只对扫描页执行本地 OCR。"
|
||||
display_name: "PDF 读取与设计"
|
||||
short_description: "本地 OCR 优先识别,设计排版、转换与编辑 PDF,并完成逐页校验"
|
||||
default_prompt: "使用 $pdf 读取或制作 PDF;扫描图像先做本地 OCR,疑难内容辅助复核,按内容设计版式并检查最终页面。"
|
||||
|
||||
74
skills/pdf/assets/design.css
Normal file
74
skills/pdf/assets/design.css
Normal file
@ -0,0 +1,74 @@
|
||||
/* Publication defaults. Author CSS overrides these tokens and styles. */
|
||||
:root {
|
||||
--page-width: 210mm; --page-height: 297mm;
|
||||
--accent: #8a3a2a; --accent-light: #f5ece7; --cover-bg: #30231f; --cover-text: #faf5ef;
|
||||
--ink: #202124; --muted: #62666a; --rule: #d8dadd;
|
||||
--font-display: 'PingFang SC', 'Microsoft YaHei', 'Noto Sans CJK SC', Arial, sans-serif;
|
||||
--font-body: 'SimSun', 'Noto Serif CJK SC', 'Times New Roman', serif;
|
||||
--font-sans: 'PingFang SC', 'Microsoft YaHei', 'Noto Sans CJK SC', Arial, sans-serif;
|
||||
}
|
||||
@page { size: A4; margin: 24mm 24mm 22mm;
|
||||
@top-left { content: string(chapter); font: 8pt var(--font-sans); color: #62666a; }
|
||||
@bottom-right { content: counter(page); font: 8pt var(--font-sans); color: #62666a; }
|
||||
}
|
||||
@page cover { size: A4; margin: 0;
|
||||
@top-left { content: none; } @bottom-right { content: none; }
|
||||
}
|
||||
* { box-sizing: border-box; }
|
||||
html, body { margin: 0; padding: 0; }
|
||||
body { color: var(--ink); font: 10.5pt/1.65 var(--font-body); overflow-wrap: break-word; }
|
||||
h1, h2, h3 { font-family: var(--font-display); color: var(--accent); break-after: avoid; line-height: 1.3; }
|
||||
h1 { font-size: 22pt; margin: 24pt 0 12pt; string-set: chapter content(text); }
|
||||
h2 { font-size: 15pt; margin: 18pt 0 8pt; }
|
||||
h3 { font-size: 11.5pt; margin: 12pt 0 6pt; }
|
||||
p { margin: 0 0 8pt; orphans: 3; widows: 3; }
|
||||
a { color: var(--accent); text-decoration: underline; }
|
||||
strong { font-family: var(--font-sans); }
|
||||
.section-start { break-before: page; }
|
||||
.eyebrow { font: 8.5pt/1.4 var(--font-sans); letter-spacing: .1em; }
|
||||
.lead { font-size: 13pt; line-height: 1.7; }
|
||||
.cover { page: cover; break-after: page; width: var(--page-width); height: var(--page-height);
|
||||
padding: 30mm 27mm; position: relative; display: flex; flex-direction: column; justify-content: center;
|
||||
background: var(--cover-bg); color: var(--cover-text); font-family: var(--font-sans); }
|
||||
.cover h1 { margin: 16pt 0; color: inherit; font: 700 40pt/1.2 var(--font-display); string-set: none; }
|
||||
.cover .subtitle { font-size: 14pt; line-height: 1.6; max-width: 140mm; }
|
||||
.cover .metadata { margin-top: 25mm; font-size: 10pt; line-height: 1.8; }
|
||||
.cover .accent-line { width: 24mm; border-top: 3pt solid var(--accent); margin: 14pt 0; }
|
||||
.cover-fullbleed::before { content: ''; position: absolute; top: 0; right: 20mm; width: 12mm; height: 28mm; background: var(--accent); }
|
||||
.cover-split { background: #f7f5f1; color: var(--ink); padding-left: 101mm; padding-right: 18mm; }
|
||||
.cover-split::before { content: ''; position: absolute; inset: 0 auto 0 0; width: 42%; background: var(--cover-bg); border-right: 2mm solid var(--accent); }
|
||||
.cover-split h1 { font-size: 28pt; }
|
||||
.cover-typographic, .cover-minimal, .cover-frame { background: #fafaf7; color: var(--ink); }
|
||||
.cover-typographic h1 { font-size: 48pt; }
|
||||
.cover-minimal { border-left: 3mm solid var(--accent); }
|
||||
.cover-minimal h1 { font-weight: 400; }
|
||||
.cover-frame::before { content: ''; position: absolute; inset: 11mm; border: 1pt solid var(--accent); pointer-events: none; }
|
||||
.cover-frame { text-align: center; align-items: center; }
|
||||
.cover-editorial::before { content: attr(data-mark); position: absolute; top: 12mm; right: 10mm; opacity: .07; font: 180pt/1 var(--font-display); }
|
||||
.cover-editorial h1 { font-size: 48pt; }
|
||||
figure { margin: 14pt 0; break-inside: avoid; }
|
||||
img, svg { max-width: 100%; }
|
||||
img { height: auto; }
|
||||
figcaption, .caption { font: 8.5pt/1.5 var(--font-sans); color: var(--muted); margin-top: 6pt; }
|
||||
.mermaid { text-align: center; margin: 12pt 0; break-inside: avoid; }
|
||||
.mermaid svg { max-height: 180mm; }
|
||||
.math-display { margin: 12pt 0; text-align: center; break-inside: avoid; }
|
||||
.katex { font-size: 1.05em; }
|
||||
table { width: 100%; border-collapse: collapse; margin: 12pt 0; font: 9.5pt/1.5 var(--font-sans); }
|
||||
thead { display: table-header-group; }
|
||||
tr { break-inside: avoid; }
|
||||
th { text-align: left; background: var(--accent); color: white; border-top: 1.5pt solid var(--accent); }
|
||||
th, td { padding: 7pt 8pt; overflow-wrap: anywhere; vertical-align: top; }
|
||||
td { border-bottom: .5pt solid var(--rule); }
|
||||
tbody tr:nth-child(even) { background: var(--accent-light); }
|
||||
tbody tr:last-child td { border-bottom: 1.5pt solid var(--accent); }
|
||||
.three-line th { background: transparent; color: var(--ink); border-top: 1.5pt solid var(--ink); border-bottom: .8pt solid var(--ink); }
|
||||
.three-line td { border: 0; } .three-line tbody tr { background: transparent; }
|
||||
.three-line tbody tr:last-child td { border-bottom: 1.5pt solid var(--ink); }
|
||||
.numeric { text-align: right; font-variant-numeric: tabular-nums; }
|
||||
blockquote, .callout, .theorem { margin: 12pt 0; padding: 8pt 12pt; border-left: 2pt solid var(--accent); background: var(--accent-light); }
|
||||
pre { font: 9pt/1.5 'Courier New', monospace; padding: 10pt; background: #f5f5f5; white-space: pre-wrap; overflow-wrap: anywhere; }
|
||||
ul, ol { padding-left: 20pt; } li { margin-bottom: 4pt; }
|
||||
.references { font-size: 9pt; line-height: 1.6; }
|
||||
.toc a { display: block; margin: 6pt 0; text-decoration: none; }
|
||||
.toc a::after { content: ' ' target-counter(attr(href), page); float: right; }
|
||||
26
skills/pdf/assets/report.html
Normal file
26
skills/pdf/assets/report.html
Normal file
@ -0,0 +1,26 @@
|
||||
<!doctype html>
|
||||
<html lang="zh-CN">
|
||||
<head><meta charset="UTF-8"><title>报告标题</title>
|
||||
<style>
|
||||
:root { --accent: #8a3a2a; --accent-light: #f5ece7; --cover-bg: #30231f; --cover-text: #faf5ef; }
|
||||
</style></head>
|
||||
<body>
|
||||
<!-- 本文件是结构示例:根据实际内容替换文字,按需要保留封面、目录和章节。默认排版 CSS 由固定脚本注入。 -->
|
||||
<section class="cover cover-fullbleed">
|
||||
<p class="eyebrow">报告类型 · 年份</p>
|
||||
<h1>报告标题</h1>
|
||||
<div class="accent-line"></div>
|
||||
<p class="subtitle">一句说明报告对象、范围和阅读目的的副标题。</p>
|
||||
<p class="metadata">作者或机构<br>发布日期</p>
|
||||
</section>
|
||||
<section>
|
||||
<h1 id="summary">核心结论</h1>
|
||||
<p class="lead">用可核对的证据说明主要结论,并区分事实与判断。</p>
|
||||
<h2 id="evidence">数据与依据</h2>
|
||||
<table><thead><tr><th>指标</th><th>观察</th><th>来源</th></tr></thead>
|
||||
<tbody><tr><td>示例指标</td><td>替换为真实数据或明确标注的示例</td><td>填写可验证来源</td></tr></tbody></table>
|
||||
<div class="callout"><strong>适用范围</strong><p>说明时间范围、样本、计算口径与不确定性。</p></div>
|
||||
<h2 id="actions">建议与下一步</h2>
|
||||
<p>建议应能追溯到证据;需要决策的事项写清楚条件与影响。</p>
|
||||
</section>
|
||||
</body></html>
|
||||
37
skills/pdf/references/conversion.md
Normal file
37
skills/pdf/references/conversion.md
Normal file
@ -0,0 +1,37 @@
|
||||
# Office、LaTeX 与 PDF 转换
|
||||
|
||||
只调用固定脚本。不能直接运行 LibreOffice、Tectonic、Bash、Python 或 Node,也不在任务中安装依赖。
|
||||
|
||||
## Office → PDF
|
||||
|
||||
Office 主文档的编辑、公式计算、图表和版式由 docx/xlsx/pptx 等对应 skill 完成。已有完成的源文件可调用 `scripts/convert_to_pdf.py`:
|
||||
|
||||
```text
|
||||
--input '/usr/local/src/pdf/tmp/task/report.docx' --output '/usr/local/src/pdf/report.pdf'
|
||||
```
|
||||
|
||||
支持 DOCX/DOC/ODT/RTF、PPTX/PPT/ODP、XLSX/XLS/ODS,最大 25 MiB。`--timeout` 默认 180 秒,1–600;可用 `--overwrite` 覆盖本任务旧产物。每次转换使用独立 LibreOffice profile 和临时副本,不修改源文件。
|
||||
|
||||
转换前确认 Excel 公式已重算、打印范围正确;核对字体替换、图表、分页和页数。转换可能有版式差异,必须渲染检查。CSV 和 HTML 分别先走 xlsx 或 HTML 设计接口,不能含糊地自动推断编码和布局。
|
||||
|
||||
## PDF → Office
|
||||
|
||||
PDF 是固定版面,不能承诺用 LibreOffice 直接得到结构完整的 Word、Excel 或 PPT。先用原生提取/OCR 得到可靠文本和表格,再由对应 skill 的固定写入接口重建;明确哪些结构可编辑、哪些需要保留图片。扫描件不会因为换扩展名就变成可编辑文本。
|
||||
|
||||
需要高保真还原时先确认重点是视觉一致还是编辑结构;保留原 PDF 对照。不得把每页截图贴入 Word 后声称正文可编辑。
|
||||
|
||||
## LaTeX → PDF
|
||||
|
||||
用户明确提供 LaTeX 模板或源文件时调用 `scripts/compile_latex.py`:
|
||||
|
||||
```text
|
||||
--input '/usr/local/src/pdf/tmp/task/main.tex' --output '/usr/local/src/pdf/paper.pdf'
|
||||
```
|
||||
|
||||
输入 `.tex` 最大 2 MiB,相关图片和被引用的 `.tex` 放在本任务目录;超时默认 180 秒,范围 1–600。固定脚本内部调用 Tectonic,禁用 shell escape,只使用镜像构建时缓存的 TeX 资源。缺失包、未缓存模板、编译失败或超时都明确报错,不能安装或偷偷切换到联网编译。
|
||||
|
||||
基础镜像预热 ctex、常用数学、表格、图片、几何和超链接包;特殊模板或包可能还需维护者补充镜像。Tectonic 自行处理常见重跑与引用,不把同一编译无意义重复多次。
|
||||
|
||||
保留用户模板的语言、字体、章节与引用规范。中文模板可使用 `ctexart` 和已缓存的 Fandol 字体;新的模板需确认所用字体实际存在。若只有少量数学公式而没有 LaTeX 模板,HTML + KaTeX 即可,无需编写完整 LaTeX 文档。
|
||||
|
||||
返回页数、警告和视觉复核标记。检查 undefined references、缺字、Overfull 等具体问题;输出成功后仍需提取与渲染检查,不把“有 PDF 文件”作为通过标准。
|
||||
48
skills/pdf/references/creation.md
Normal file
48
skills/pdf/references/creation.md
Normal file
@ -0,0 +1,48 @@
|
||||
# HTML/CSS 设计排版接口
|
||||
|
||||
先读 [设计规范](design.md)。简单文本使用 `create_pdf.py`;图文报告、封面、页眉页脚、目录、公式和流程图使用 `create_design_pdf.py`。
|
||||
|
||||
## 调用
|
||||
|
||||
通过 `write_file` 将内容写为 UTF-8 `.html`,以 `<!doctype html>` 声明标准模式,并包含 `<meta charset="utf-8">`。通过 `execute_skill_script` 调用 `scripts/create_design_pdf.py`,不直接调用 Node 或 Chromium:
|
||||
|
||||
```text
|
||||
--input '/usr/local/src/pdf/tmp/task/report.html' --output '/usr/local/src/pdf/report.pdf'
|
||||
```
|
||||
|
||||
可选 `--css <本地CSS>`、`--page-size A4|LETTER`、`--expected-pages <1–200>`、`--timeout <10–600>`(默认 180)和 `--overwrite`。HTML/CSS 单文件不超过 2 MiB。先读取 `assets/report.html` 了解结构,按实际内容编写;模板中的示例文字不可直接交付。
|
||||
|
||||
`assets/design.css` 由固定渲染器自动注入,提供变量、标题、六种封面、三线表、代码、引用、图注与目录。用户 HTML 内的 CSS 可覆盖默认值,`--css` 最后应用。不需要手动加载该样式或任何 JS 库。
|
||||
|
||||
## 静态内容与本地资源
|
||||
|
||||
- 只写静态 HTML/CSS/SVG。禁止 `<script>`、事件属性、iframe/object/embed、`javascript:` URL 和运行任意 JS。不能提供 Python、Bash 或 Node 代码让脚本代为执行。
|
||||
- 图片、CSS 和自备字体放在 HTML 同目录及其子目录,用相对路径引用;上级目录、符号链接越界、`file:` 和外部资源会被拒绝。资源单个不超过 25 MiB。
|
||||
- HTTP(S) 引用链接可以保留为可点击参考来源,但不能用作图片、CSS、字体或脚本下载入口。所需远程素材先由 `download_attachment.py` 下载。
|
||||
- Paged.js、KaTeX、Mermaid 和 KaTeX 字体从镜像读取,没有 CDN 例外。渲染器使用独立浏览器上下文并限制资源访问。
|
||||
- 使用语义化 `h1/h2/p/table/figure/figcaption`,对引用设置真实 `id` 和 `href="#..."`。自定义计数器跨页需核验;普通页码、目录目标页码可使用 Paged.js 支持的 counter/target-counter,不笼统禁止 CSS counter。
|
||||
|
||||
## 公式与流程图
|
||||
|
||||
不加载脚本、不写初始化代码。固定接口识别以下结构:
|
||||
|
||||
```html
|
||||
<p>公式为 <span class="math-inline">E=mc^2</span>。</p>
|
||||
<div class="math-display">\sum_{i=1}^{n} x_i</div>
|
||||
<div class="mermaid">flowchart LR
|
||||
A[资料] --> B[本地 OCR]
|
||||
B --> C[疑难复核]
|
||||
C --> D[PDF]</div>
|
||||
```
|
||||
|
||||
公式内容为 KaTeX 数学标记,Mermaid 为图形描述语言,不是任意程序入口。Mermaid 使用 neutral 主题与 strict 安全模式;不支持在图中插入配置指令。公式/图形解析失败则返回错误,不带着未渲染内容输出成功。
|
||||
|
||||
数据图表可用静态 SVG 或已生成图片。长公式应分行;大表/大图应调整结构、命名横向页或拆分,不能一律缩成不可读的小图。
|
||||
|
||||
## 分页与检查
|
||||
|
||||
- 默认 A4 内页与全页封面都使用实际物理尺寸;不采用原脚本固定 `scale: 1.5` 的补偿。
|
||||
- 先等待图片、字体和图形完成,再等待 Paged.js 的完成 Promise。超时、外部资源、图片失败、横向溢出或页数不符时不发布结果。
|
||||
- PDF 实际页数必须与分页引擎相同。`--expected-pages` 不匹配时明确失败,不裁掉页面或缩短内容冒充满足要求。
|
||||
- 返回 `page_count`、每页文字/图形统计、warnings 和 `requires_visual_review`。空白页提示结合上下文判断;少字的图表页不直接判坏。
|
||||
- 用 `inspect_pdf.py` 和 `extract_text.py` 核对最终文件,再用 `render_pdf.py` 查看全部页。自动检测不能发现所有纵向裁切、语义错误、缺字或表单外观问题。
|
||||
40
skills/pdf/references/dependencies.md
Normal file
40
skills/pdf/references/dependencies.md
Normal file
@ -0,0 +1,40 @@
|
||||
# 容器依赖与能力边界
|
||||
|
||||
依赖由 `silk-base/Dockerfile` 预装。机器人只调用固定脚本,不能运行安装命令;需要更新依赖时由维护者重新构建和部署基础镜像。
|
||||
|
||||
## 复用原 PDF 工具链
|
||||
|
||||
| 能力 | 镜像依赖 |
|
||||
| --- | --- |
|
||||
| 原生文字、表格、页面和表单 | pdfplumber 0.11.9、pypdf 6.10.0、Poppler |
|
||||
| 简单文本/Markdown PDF | ReportLab 4.4.9、Pillow、已安装中文与拉丁字体 |
|
||||
| PDF 页及单张图片 OCR | RapidOCR 3.9.1、ONNX Runtime 1.27.0、OpenCV 4.12.0.88、OmegaConf 2.3.1 |
|
||||
| Office 导出 | LibreOffice Writer/Calc/Impress 与系统字体 |
|
||||
| HTML 渲染运行时 | 已有 Node.js 与系统 Chromium,`CHROME_BIN=/usr/bin/chromium` |
|
||||
|
||||
OCR 使用 RapidOCR 包内的检测、方向分类与识别 ONNX 模型。缺少模型就返回错误;不回退联网下载。OmegaConf 固定兼容版本,避免旧版不能处理 RapidOCR 的路径配置。
|
||||
|
||||
## 本次补充
|
||||
|
||||
| 依赖 | 用途 | 安装位置/方式 |
|
||||
| --- | --- | --- |
|
||||
| poppler-data | Adobe CJK 字符映射,补齐部分中文 PDF 的渲染与提取 | Debian 软件包,显式安装并检查 GB1 映射文件 |
|
||||
| playwright-core 1.63.0 | 固定浏览器渲染桥接 | 全局 Node 包,复用系统 Chromium |
|
||||
| Paged.js 0.4.3 | CSS 分页、命名页、页眉页脚、目录目标页码 | 全局 Node 包,本地加载 |
|
||||
| KaTeX 0.18.7 | 行内/展示公式及其字体 | 全局 Node 包,本地加载 |
|
||||
| Mermaid 11.17.2 | 流程、关系等示意图 | 全局 Node 包,固定 strict 模式 |
|
||||
| Tectonic 0.17.0 | 用户提供的 LaTeX 源文件编译 | amd64/arm64 官方静态发行包,逐架构 SHA-256 校验 |
|
||||
|
||||
不需要额外引入 pikepdf、另一套 Python 浏览器、Matplotlib 或 LaTeX 完整发行版。页面/表单/元数据/图像继续复用 pypdf;数据图用静态 SVG、图片或现有 xlsx 图表。数学公式优先 KaTeX,有 LaTeX 模板时才用 Tectonic。
|
||||
|
||||
`poppler-data` 提供编码映射,不能用字体包替代。没有映射时即使 PDF 带有中文字体,部分文件仍可能预览缺字;遇到 `Missing language pack` 应更新镜像,不能把有缺字的预览用于 OCR 或当作空白原文。[Debian 包说明](https://packages.debian.org/trixie/poppler-data)
|
||||
|
||||
## 构建与离线运行
|
||||
|
||||
- 浏览器只从任务目录、skill 素材和固定本地包读取资源;不使用 CDN、不自动下载 Chromium。`NODE_PATH` 保留镜像的全局包目录。
|
||||
- 构建时实际运行 HTML + KaTeX + Mermaid + Paged.js 自检,不能只验证包能 import。
|
||||
- `TECTONIC_CACHE_DIR=/opt/tectonic-cache` 在构建时预热 ctex/Fandol、amsmath/amssymb、booktabs/longtable、graphicx、geometry、hyperref,随后验证仅缓存编译和中文文字提取。首次构建需要网络下载 TeX 资源;任务编译始终禁用 shell escape 且只读缓存包。
|
||||
- 特殊模板、额外文献工具或未预热的 TeX 包可能仍不可用。维护者把真实模板所需资源加入构建预热,再部署镜像;不能在机器人任务里解除缓存限制。
|
||||
- 原有宋体、黑体、仿宋、楷体、方正小标宋、微软雅黑、苹方 SC、SF Pro、Noto 等字体继续复用。渲染后仍需检查所选字体和字形,安装完成不代表所有模板无差异。
|
||||
|
||||
修改 Dockerfile 后必须重新构建并部署,已运行的旧容器不会自动获得新依赖。
|
||||
67
skills/pdf/references/design.md
Normal file
67
skills/pdf/references/design.md
Normal file
@ -0,0 +1,67 @@
|
||||
# PDF 视觉设计规范
|
||||
|
||||
先确定文档用途、读者、交付媒介、品牌和内容结构,再决定视觉。继承用户模板时先保持其规范;阅读、裁剪或填表不触发整本重新设计。
|
||||
|
||||
## 视觉方向
|
||||
|
||||
选择一套与内容相符的排版语言,贯穿封面、标题、表格和页眉。下面是起点,不是行业强制配色:
|
||||
|
||||
| 内容与气质 | 主色 | 浅色 | 深色封面或文字 |
|
||||
| --- | --- | --- | --- |
|
||||
| 科学、工程、数据分析 | `#2D5F8A` | `#E8F0F8` | `#152C3E` |
|
||||
| 商业、策略、运营 | `#8A3A2A` | `#F5ECE7` | `#30231F` |
|
||||
| 医疗、生态、公共服务 | `#2A6B5A` | `#EBF3EF` | `#19382F` |
|
||||
| 创意、文化、作品集 | `#6B2A35` | `#F5ECEE` | `#29171D` |
|
||||
| 学术、法律、正式材料 | `#3D4C5E` | `#F1F3F5` | `#202124` |
|
||||
|
||||
通常一个主色、深色正文和少量浅色层次就够。强调色用于标题、线条和关键数据;正文保持深色。颜色须有含义,不能只依赖颜色区分类别。文字与背景保持清晰对比,小字尽量达到 7:1;不要把装饰图形的对比度要求与正文混为一谈。
|
||||
|
||||
不要把渐变、卡片或深色背景机械地当成专业设计,也不要因旧规范的偏好一概禁止用户明确要求的颜色。内页以打印友好、清晰的信息层级为主;大面积深色更适合少量封面或章节页。
|
||||
|
||||
## 字体与尺度
|
||||
|
||||
镜像已提供 Noto CJK、宋体、黑体、仿宋、楷体、方正小标宋、微软雅黑、苹方 SC、SF Pro、Arial、Times New Roman 等。选择通常不超过两套视觉字体系统;中文与拉丁字体的必要回退不算额外装饰字体。不要假定宿主机所有字体都存在容器中,也不要从字体 CDN 加载。
|
||||
|
||||
| 风格 | 中文正文/标题 | 英文与数字 |
|
||||
| --- | --- | --- |
|
||||
| 现代报告 | 苹方 SC / 微软雅黑 / Noto Sans CJK SC | SF Pro / Arial |
|
||||
| 学术长文 | 宋体 / Noto Serif CJK SC,标题可搭黑体 | Times New Roman |
|
||||
| 正式公文 | 依用户规范选仿宋、楷体、方正小标宋 | 依模板 |
|
||||
|
||||
字号起点:封面 32–48 pt,一级标题 20–24 pt,二级 14–16 pt,三级 11.5–12 pt,正文 10.5–12 pt,图注 8.5–9.5 pt,页眉页脚 8–9 pt。中文正文行高 1.6–1.8;正文过长时先调整结构与分页,不用极小字号硬塞。
|
||||
|
||||
A4 内页通常上下 22–28 mm、左右 22–28 mm。段后约 8 pt,章节前约 24–26 pt;选定间距节奏后保持一致。简历可用 15–18 mm 边距,仍需保证可读性和完整内容。
|
||||
|
||||
## 六种可复用封面
|
||||
|
||||
`assets/design.css` 提供对应 class,`assets/report.html` 提供结构起点。独立封面按需使用;简历、短备忘录、用户指定页数紧张或已有模板时,标题区即可。
|
||||
|
||||
| class | 设计方式 | 常见用途 |
|
||||
| --- | --- | --- |
|
||||
| `cover-fullbleed` | 整页深底、大标题、短色带、作者日期 | 年报、主题报告 |
|
||||
| `cover-split` | 42% 色块与 58% 浅色内容区,清晰分割 | 提案、方案 |
|
||||
| `cover-typographic` | 浅底、展示字体、尺度对比 | 作品、专题材料 |
|
||||
| `cover-minimal` | 竖线、轻量标题、充分留白 | 简洁文档 |
|
||||
| `cover-frame` | 细框、居中结构、克制装饰 | 正式报告 |
|
||||
| `cover-editorial` | 大字背景、强标题、编辑式构图 | 杂志、创意内容 |
|
||||
|
||||
封面用命名页 `@page cover`,整页尺寸与纸张一致,`margin: 0`,内容区采用 `border-box`。不要在有页边距的内容区再套 `100vh` 或完整 A4 高度;这会产生裁切和额外空白页。图案用静态 SVG/几何元素;不能让纹理干扰标题。
|
||||
|
||||
## 内页
|
||||
|
||||
- 标题层级用字号、字重、间距和少量细线区分;标题不能孤立在页底。长章节按语义分段,避免把整节设为不可分页。
|
||||
- 正文不堆砌仪表盘卡片。报告可用少量重点引言或行动框;学术内容可用定义/定理边线。别用带阴影的网页组件代替正文结构。
|
||||
- 表格有真实表头、单位和来源。数字右对齐,小数位一致;长表允许跨页并重复表头,避免整张长表 `break-inside: avoid`。学术三线表用 `.three-line`;业务表可用主色表头和浅色交替行。
|
||||
- 图表比例服从信息:柱线图常用横向,流程图可纵向,页面容纳不下时拆分、调整方向或用横向命名页。不要强制所有图表横向,也不要通过 `overflow:hidden` 掩盖溢出。
|
||||
- 图表必须有可辨认的标签、单位、图例与图注;图表文字按最终 PDF 尺寸检查。数值须来自实际数据,不能用装饰图冒充统计结果。
|
||||
- 图片保持比例,注明图注;数据图优先原有 xlsx 图表、静态 SVG 或已提供图片。不生成任意 Matplotlib/Python/Node 代码作为运行入口。
|
||||
- 代码示例用浅灰底、等宽字体与换行;引用框用细左边线;数学公式用 KaTeX。需要完整 LaTeX 模板时走编译接口。
|
||||
- 页眉使用章节名称,页脚保持统一页码。目录、图表引用和参考文献使用真实锚点,可点击;分页后核对链接落点。
|
||||
|
||||
## 内容与交付
|
||||
|
||||
用户的大纲、语言、数字口径和篇幅优先。精确页数通过实际生成与 `--expected-pages` 核对;字数要求按用户范围执行,不默认放宽 20%,不为填页伪造内容。没有必要时不添加封面、目录或参考文献页。
|
||||
|
||||
已有材料里的数据注明来源;新增研究、统计、政策等需验证,来源不足就披露不确定性。参考文献格式依用户/机构要求,中文 GB/T 7714 或英文 APA 可作为选项;不编造作者、年份和出处。
|
||||
|
||||
逐页核验:文字无缺字/黑块,字号清晰;页面尺寸与边距一致;图文无裁切、重叠;表格跨页合理;公式和引用正确;封面与正文衔接自然;不存在非预期空白页。自动溢出检测只是辅助,不替代实际看图。
|
||||
54
skills/pdf/references/editing.md
Normal file
54
skills/pdf/references/editing.md
Normal file
@ -0,0 +1,54 @@
|
||||
# 表单、裁剪、元数据和嵌入图片
|
||||
|
||||
通过 `execute_skill_script` 调用 `scripts/edit_pdf.py`。所有输出遵循 `/usr/local/src/pdf/`;不覆盖源文件。原有合并、拆分、旋转仍用 `manage_pdf.py`。
|
||||
|
||||
## 表单
|
||||
|
||||
先查看真实字段:
|
||||
|
||||
```text
|
||||
form-info --input '/usr/local/src/pdf/tmp/task/form.pdf'
|
||||
```
|
||||
|
||||
返回字段 id、类型、当前值、只读状态、选择项和按钮状态。默认 50 项,`--offset`/`--limit`(最高 200)支持继续。没有 AcroForm 字段的扫描表单不能直接套用字段填写;XFA 和数字签名不是本接口支持的填写类型。
|
||||
|
||||
填写明确要求的字段:
|
||||
|
||||
```text
|
||||
form-fill --input '/usr/local/src/pdf/tmp/task/form.pdf' --output '/usr/local/src/pdf/filled.pdf' --data '{"name":"Alice","agree":true,"country":"CN"}'
|
||||
```
|
||||
|
||||
较长数据用 `--data-file <JSON>` 替代 `--data`,二选一,上限 2 MiB、500 个字段。
|
||||
|
||||
- 文本值必须是字符串,遵守字段长度。复选框可用 JSON `true/false`,或实际状态名如 `/Yes`;不能把非空字符串一律当作选中。
|
||||
- 单选按钮使用实际状态值;下拉/列表使用真实选项值,多选需原字段支持。无效值、未知字段、只读字段或签名字段明确失败。
|
||||
- pypdf 更新字段值与外观,再回读校验。必须渲染核对文字位置、换行、复选框和单选按钮状态。原表单字体未包含中文等字形时,不能仅凭字段值正确就交付;需保留可显示的模板或按用户要求重建表单版式。
|
||||
- 不把填写文本当作签署数字签名,也不主动提交表单。
|
||||
|
||||
## 裁剪
|
||||
|
||||
```text
|
||||
crop --input '/usr/local/src/pdf/tmp/task/source.pdf' --output '/usr/local/src/pdf/cropped.pdf' --pages '1-2' --box '20,30,575,812'
|
||||
```
|
||||
|
||||
坐标单位 pt,按未旋转 PDF 坐标系的左、下、右、上填写;框必须在页面 MediaBox 内。省略 `--pages` 处理全部页。修改 CropBox,保留页面内容。**裁剪不删除不可见内容,不是脱敏。** 敏感信息移除需专门的真正删改流程,不能用白块或裁剪冒充。
|
||||
|
||||
## 元数据
|
||||
|
||||
读取用 `inspect_pdf.py`。更新明确字段:
|
||||
|
||||
```text
|
||||
metadata --input '/usr/local/src/pdf/tmp/task/source.pdf' --output '/usr/local/src/pdf/updated.pdf' --data '{"Title":"报告","Author":"机构"}'
|
||||
```
|
||||
|
||||
支持 Title/Author/Subject/Keywords/Creator/Producer;保留未指定字段。该接口修改文档信息字典,现有 XMP 保留且通过 `xmp_preserved` 提示,不能声称已清理所有元数据。
|
||||
|
||||
## 嵌入图片
|
||||
|
||||
```text
|
||||
extract-images --input '/usr/local/src/pdf/tmp/task/source.pdf' --pages '2-3' --output-dir '/usr/local/src/pdf/tmp/task/images'
|
||||
```
|
||||
|
||||
一次最多 4 页、默认 20 张图片(`--max-images` 最高 50),本批总大小不超过 25 MiB。用 `remaining_pages` 和 `--start-image <next_image>` 续读。需要覆盖本任务旧输出时加 `--overwrite`。
|
||||
|
||||
此操作提取 PDF 中的图像对象,不代表完整页面:不包含周围文字、矢量图形,也可能是被裁剪/复用的图片。阅读扫描页和视觉复核应使用 OCR 内部渲染或 `render_pdf.py`,不能把提取出的零散图片当成原 PDF 页。
|
||||
164
skills/pdf/references/operations.md
Normal file
164
skills/pdf/references/operations.md
Normal file
@ -0,0 +1,164 @@
|
||||
# 原有固定操作接口
|
||||
|
||||
## 下载远程 PDF
|
||||
|
||||
只接受 HTTPS 地址。完整保留 URL 及查询参数,不在回复、日志摘要或文件名中复述敏感参数。
|
||||
|
||||
调用 `scripts/download_pdf.py`:
|
||||
|
||||
```text
|
||||
--url 'https://example.com/document.pdf' --output '/usr/local/src/pdf/tmp/<任务名>/source.pdf'
|
||||
```
|
||||
|
||||
可选参数:
|
||||
|
||||
- `--timeout <秒>`:默认 `60`。
|
||||
- `--max-bytes <字节数>`:默认且最高 `26214400`(25 MiB),只允许设置更小的限制。
|
||||
- `--overwrite`:仅在目标是本次任务生成的缓存时使用。
|
||||
|
||||
脚本会创建父目录、流式下载、阻止 HTTPS 重定向降级到 HTTP,并验证 PDF。成功结果包含 `path`、`size_bytes`、`page_count` 和 `encrypted`。
|
||||
|
||||
## 下载通用附件
|
||||
|
||||
需要下载作为 PDF 任务素材的图片、视频、音频、压缩包或其他文件时,调用 `scripts/download_attachment.py`:
|
||||
|
||||
```text
|
||||
--url 'https://example.com/asset.bin?signature=...' --output '/usr/local/src/pdf/tmp/<任务名>/asset.bin'
|
||||
```
|
||||
|
||||
只接受 HTTPS 地址,`output` 可使用任意附件扩展名。可选参数只有 `--timeout <1-600>`(默认 `60`)和 `--overwrite`。附件上限固定为 25 MiB(26214400 字节),不可调高:脚本先用 HEAD 探测远端声明大小,再检查 GET 响应声明,并在流式接收时持续兜底计数;任一阶段发现超限都会返回 `ok: false` 和明确的“已拒绝下载”错误,且不会发布部分文件。
|
||||
|
||||
成功结果包含 `path`、实际 `size_bytes`、`declared_size_bytes`、`size_limit_bytes`、`size_probe` 和 `content_type`。本脚本不校验文件业务格式;远程 PDF 源文件仍使用 `download_pdf.py`。
|
||||
|
||||
## 检查 PDF
|
||||
|
||||
调用 `scripts/inspect_pdf.py`:
|
||||
|
||||
```text
|
||||
--input '/usr/local/src/pdf/tmp/<任务名>/source.pdf'
|
||||
```
|
||||
|
||||
使用返回的 `page_count`、`encrypted`、`metadata`、`page_layouts` 和 `form_field_count` 判断后续处理方式。不要直接调用 `pdfinfo`。
|
||||
|
||||
## 提取正文
|
||||
|
||||
首次调用 `scripts/extract_text.py`:
|
||||
|
||||
```text
|
||||
--input '/usr/local/src/pdf/tmp/<任务名>/source.pdf'
|
||||
```
|
||||
|
||||
默认使用 `auto` 引擎:先由 Poppler `pdftotext` 提取;结果不可用或命令不可用时自动尝试 `pdfplumber`,并可逐页选择质量更好的结果。脚本使用 `pypdf` 获取标准页数,并拒绝把页数不一致的提取结果当作成功。不要直接执行 `pdftotext`。
|
||||
|
||||
默认单次最多处理 8 页、返回 24000 个字符。可使用:
|
||||
|
||||
- `--start-page <页码>`、`--end-page <页码>`:页码从 `1` 开始。
|
||||
- `--start-offset <字符偏移>`:继续读取被字符上限截断的同一页;大于 `0` 时同时传入上次返回的 `next_engine`。
|
||||
- `--max-pages <页数>`、`--max-chars <字符数>`:控制单次输出。
|
||||
- `--layout`:仅在需要尽量保留版面空格时使用。
|
||||
- `--engine <auto|poppler|pdfplumber>`:首次及跨页提取保持 `auto`;同页字符续读时传入上次返回的 `next_engine`。
|
||||
- `--timeout <秒>`:Poppler 提取超时,默认 `120`。
|
||||
|
||||
先检查 `usable_for_summary` 和 `text_quality.status`:
|
||||
|
||||
- `usable_for_summary: true`:只使用 `pages[]` 中同样标为 `usable_for_summary: true` 的 `text`;可疑页的文本会被置空。如果 `has_more: true`,始终传回 `next_page` 和 `next_offset`。仅当 `next_offset` 大于 `0` 时,把非空的 `next_engine` 传给 `--engine` 以固定同页字符游标;这种调用只续读当前页。当前页完成后返回的 `next_offset` 为 `0`,此时不要传 `--engine`,让下一页重新使用 `auto`。保留首次调用的 `--end-page`(如果指定)及其他选项,直至 `has_more: false`。
|
||||
- `usable_for_summary: false`:本批次没有可靠文本,不要使用返回内容。查看 `engine_attempts`、`text_quality.reasons`、`text_quality.suspect_pages` 和 `needs_ocr`;若 `has_more: true`,仍按跨页游标继续检查后续批次,避免漏掉后续可搜索文本。
|
||||
- `needs_ocr: true`:一个或多个页面未得到可靠文本。把 `text_quality.suspect_pages` 中实际需要阅读的页码传给 `ocr_text.py`;先调用 OCR;只有 OCR 标记疑难、结果与页面结构冲突或需要理解非文本图形时,才渲染相关页供模型辅助复核。
|
||||
|
||||
`complete_text_coverage: true` 表示本批次所有页面均有可靠文本。`text_quality` 按页检测空白或过少文本、页面实际可见图像覆盖过大但文字不足、`(cid:...)`、Unicode 替换字符、异常控制字符及外观像汉字的部首字符;`pages[].extractor` 表示该页最终采用的引擎。`status: mixed` 表示同一批次同时包含可靠页和可疑页:可先使用可靠页文本,同时只核验 `suspect_pages`。不要只根据“肉眼看起来能读”判定提取结果可靠。
|
||||
|
||||
## OCR 与疑难复核
|
||||
|
||||
详见 [读取与 OCR](reading-and-ocr.md),使用固定 `ocr_text.py` 或 `ocr_image.py`。
|
||||
|
||||
## 提取表格
|
||||
|
||||
调用 `scripts/extract_tables.py`:
|
||||
|
||||
```text
|
||||
--input '/usr/local/src/pdf/tmp/<任务名>/source.pdf' --start-page 1
|
||||
```
|
||||
|
||||
默认单次最多处理 5 页、20 个表格和 2000 个单元格。可用 `--end-page`、`--start-table`、`--max-pages`、`--max-tables`、`--max-cells` 调整。若 `has_more: true`,把 `next_page` 传给 `--start-page`、`next_table` 传给 `--start-table` 后继续,并保留首次调用的 `--end-page`(如果指定)及其他提取选项。
|
||||
|
||||
## 渲染页面
|
||||
|
||||
只有满足以下任一条件时才调用 `scripts/render_pdf.py`:
|
||||
|
||||
- 用户明确要求审阅版式、图表、印章、公式或页面外观;
|
||||
- 创建或修改 PDF 后进行最终视觉检查;
|
||||
- OCR 标记低置信度、读取顺序冲突或无法识别区域,需要大模型辅助复核。
|
||||
|
||||
普通文本总结和扫描文字读取不自动增加整本视觉识别;OCR 的临时渲染由 `ocr_text.py` 内部完成。调用脚本时不要直接执行 `pdftoppm`:
|
||||
|
||||
```text
|
||||
--input '/usr/local/src/pdf/tmp/<任务名>/source.pdf' --output-dir '/usr/local/src/pdf/tmp/<任务名>/rendered' --start-page 1
|
||||
```
|
||||
|
||||
默认 150 DPI、单次最多 10 页。可使用 `--end-page`、`--max-pages`、`--dpi`、`--timeout` 和 `--overwrite`。若 `has_more: true`,使用 `next_page` 继续,并保留首次调用的 `--end-page`(如果指定)、输出目录及其他渲染选项。脚本返回标准化的 `page-0001.png` 文件路径。
|
||||
|
||||
对文字较小或图表密集的页面提高 DPI。使用可用的图像查看工具检查返回的 PNG,不要尝试把图片路径交给下载脚本。
|
||||
|
||||
## 简单文本创建 PDF
|
||||
|
||||
先使用 `write_file` 把内容写为 UTF-8 `.txt` 或 `.md` 文件,再调用 `scripts/create_pdf.py`:
|
||||
|
||||
```text
|
||||
--input '/usr/local/src/pdf/tmp/<任务名>/content.md' --output '/usr/local/src/pdf/<文件名>.pdf' --title '文档标题'
|
||||
```
|
||||
|
||||
脚本支持 Markdown 标题、项目符号和简单表格,自动选择可嵌入的 Unicode 字体并添加页码。可选参数:
|
||||
|
||||
- `--page-size <A4|LETTER>`
|
||||
- `--font-path <TTF或TTC路径>`
|
||||
- `--font-size <字号>`
|
||||
- `--margin <points>`
|
||||
- `--overwrite`
|
||||
|
||||
输入内容只使用 ASCII 连字符 `-`;脚本也会把常见 Unicode 横线规范化为 ASCII 连字符。
|
||||
|
||||
## 合并、拆分与旋转
|
||||
|
||||
调用 `scripts/manage_pdf.py`,第一个参数必须是操作名。
|
||||
|
||||
合并:
|
||||
|
||||
```text
|
||||
merge --input 'a.pdf' --input 'b.pdf' --output '/usr/local/src/pdf/merged.pdf'
|
||||
```
|
||||
|
||||
拆分指定范围:
|
||||
|
||||
```text
|
||||
split --input 'source.pdf' --output-dir '/usr/local/src/pdf/split' --range 1-3 --range 4-6
|
||||
```
|
||||
|
||||
不传 `--range` 时每页生成一个 PDF。
|
||||
|
||||
旋转指定页面:
|
||||
|
||||
```text
|
||||
rotate --input 'source.pdf' --output '/usr/local/src/pdf/rotated.pdf' --pages '1,3-5' --degrees 90
|
||||
```
|
||||
|
||||
`--degrees` 只能是 `90`、`180` 或 `270`;不传 `--pages` 时旋转全部页面。目标已存在且确认可覆盖时添加 `--overwrite`。
|
||||
|
||||
## 清理临时目录
|
||||
|
||||
调用 `scripts/cleanup_pdf_temp.py`:
|
||||
|
||||
```text
|
||||
--task-dir '/usr/local/src/pdf/tmp/<任务名>'
|
||||
```
|
||||
|
||||
脚本只允许删除 `/usr/local/src/pdf/tmp/` 下一级任务目录,拒绝删除根目录、仓库目录或其他路径。
|
||||
|
||||
## 质量要求
|
||||
|
||||
- 不覆盖用户提供的源文件。
|
||||
- 创建或修改后重新检查页数、页面尺寸、加密状态和文本可读性。
|
||||
- 扫描件先做原生文字检测,再只 OCR 可疑页;不得把低置信度 OCR 文本当作可靠正文。
|
||||
- 逐页确认没有裁切、重叠、溢出、乱码、黑方块、错误分页或异常空白页。
|
||||
- 检查标题层级、段落间距、页边距、表格、图表、图片、页码及章节衔接。
|
||||
- 引用和参考文献必须可读,不得残留工具令牌、占位符或临时路径。
|
||||
- 只有最新渲染结果不存在可见缺陷时才交付创建或修改后的 PDF。
|
||||
44
skills/pdf/references/reading-and-ocr.md
Normal file
44
skills/pdf/references/reading-and-ocr.md
Normal file
@ -0,0 +1,44 @@
|
||||
# 原生文本、OCR 与模型辅助复核
|
||||
|
||||
## 识别顺序
|
||||
|
||||
1. `inspect_pdf.py` 检查页数与加密状态;`extract_text.py` 分批检查原生文本。
|
||||
2. 可靠原生文字直接使用。对 `text_quality.suspect_pages` 调用本地 `ocr_text.py`,普通扫描件不先做整本图片识别。
|
||||
3. PDF 任务中的单张截图、扫描图片用 `ocr_image.py`。文字识别以 OCR 为主,大模型用于具体疑难内容的辅助判断。
|
||||
4. OCR 返回 `needs_review`、低置信度区域,或文字与表格结构/阅读顺序冲突时,先按原图质量合理提高 DPI(最多 400)。仍有疑问或涉及手写、复杂公式、图表、印章时,用 `render_pdf.py` 只渲染相关页,通过当前环境实际可用的图像查看/识别能力辅助复核。可使用已安装的 image-recognition skill;不能虚构工具或识别结果。
|
||||
5. 复核时提供页码、待判断问题与 OCR 候选,区分“确认原文”“模型推测”“无法辨认”。不要让模型补写看不见的金额、账号、日期或姓名。辅助结果不能覆盖其他页已可靠提取的文字。
|
||||
|
||||
`needs_ocr: false` 表示原生文字无明显问题,不代表图中的趋势、流程、版式和关系已经被理解。任务需要这些非文本信息时可以按需查看相关页,不重复 OCR 全文。
|
||||
|
||||
## PDF 页 OCR
|
||||
|
||||
通过 `execute_skill_script` 调用 `scripts/ocr_text.py`,仅传参数:
|
||||
|
||||
```text
|
||||
--input '/usr/local/src/pdf/tmp/task/source.pdf' --pages '2,5-6'
|
||||
```
|
||||
|
||||
- 页码从 1 开始,一次最多 4 页;默认 260 DPI,允许 150–400,单页渲染不超过 2000 万像素。
|
||||
- `--timeout` 是每页渲染超时,默认 180 秒,范围 1–600。
|
||||
- 默认 `--max-chars 24000`,最高 60000。
|
||||
- 字符游标:`next_offset > 0` 时用 `--pages <next_page> --start-offset <next_offset>` 续读该页;`next_offset = 0` 时用 `remaining_pages` 继续。单页续读结束后仍要继续先前未处理的页面。
|
||||
- 使用镜像内 RapidOCR/ONNX 模型。固定脚本明确指定本地模型路径,缺少模型立即失败,不触发下载。渲染 PNG 只在脚本内部临时使用,完成即清理。
|
||||
|
||||
## 单张图像 OCR
|
||||
|
||||
调用 `scripts/ocr_image.py`:
|
||||
|
||||
```text
|
||||
--input '/usr/local/src/pdf/tmp/task/scan.png'
|
||||
```
|
||||
|
||||
支持 PNG/JPEG/WebP/TIFF/BMP,最大 25 MiB、2000 万像素、单帧;自动按 EXIF 调整方向。`--max-chars` 和 `--start-offset` 用于续读可靠文本。多页 TIFF 不默默只读首帧,应先得到按页文件或 PDF。
|
||||
|
||||
## 结果解读
|
||||
|
||||
- 仅使用 `usable_for_summary: true` 的 `text` 作为正文。`empty/sparse/low_confidence` 不等于空白页面;需要核验。
|
||||
- 单行置信度低于 0.60 的文字不会混入可靠正文,即使整页平均置信度很高。`review_regions` 给出 `text_candidate`、置信度和像素框;这些是**待核验候选**,不能当事实。
|
||||
- `needs_review: true` 可能与 `usable_for_summary: true` 同时出现:表示该页有可用文字,也有未确认区域。`complete_ocr_coverage` 只有全部请求页处理完成且无需复核时才为真。
|
||||
- `review_regions` 最多返回 40 个,每段候选最多 500 字;截断时有显式标记。需要完整复核时查看对应原图,不把列表上限误认为没有其他疑点。
|
||||
- OCR 的阅读顺序按几何位置排序,多栏、跨栏标题或复杂表格不保证逻辑顺序。表格先尝试 `extract_tables.py`;扫描表格用 OCR 字框对齐并核验行列、表头、合计,不把平铺文字直接当结构化表格。
|
||||
- 引用使用真实 PDF 页码;记录实际阅读/识别范围,不能只处理前几页就声称已覆盖全文。
|
||||
@ -23,6 +23,15 @@ def quiet_pdf_library_logs() -> None:
|
||||
quiet_pdf_library_logs()
|
||||
|
||||
|
||||
def check_poppler_resources(stderr: str) -> None:
|
||||
"""A zero exit code can still leave every CJK glyph invisible."""
|
||||
if "Missing language pack" in stderr:
|
||||
raise RuntimeError(
|
||||
"Poppler 缺少中文/CJK 编码映射;基础镜像需安装 poppler-data。"
|
||||
"预览可能缺字,不能据此执行 OCR 或判断原文为空白。"
|
||||
)
|
||||
|
||||
|
||||
class SkillArgumentParser(argparse.ArgumentParser):
|
||||
def error(self, message: str) -> NoReturn:
|
||||
raise ValueError(f"参数错误:{message}")
|
||||
|
||||
139
skills/pdf/scripts/_render_html.cjs
Normal file
139
skills/pdf/scripts/_render_html.cjs
Normal file
@ -0,0 +1,139 @@
|
||||
'use strict';
|
||||
// Internal renderer. The agent invokes create_design_pdf.py, never Node or this file.
|
||||
const fs = require('fs');
|
||||
const path = require('path');
|
||||
const crypto = require('crypto');
|
||||
|
||||
function packageRoot(name) {
|
||||
for (const root of require.resolve.paths(name) || []) {
|
||||
const candidate = path.join(root, name, 'package.json');
|
||||
if (fs.existsSync(candidate)) return fs.realpathSync(path.dirname(candidate));
|
||||
}
|
||||
throw new Error(`基础镜像缺少 ${name};需要更新镜像,任务中不能安装`);
|
||||
}
|
||||
function within(root, filename) {
|
||||
const relative = path.relative(root, filename);
|
||||
return relative === '' || (!relative.startsWith('..' + path.sep) && relative !== '..' && !path.isAbsolute(relative));
|
||||
}
|
||||
const MIME = {'.html':'text/html; charset=utf-8','.css':'text/css; charset=utf-8','.js':'text/javascript',
|
||||
'.mjs':'text/javascript','.png':'image/png','.jpg':'image/jpeg','.jpeg':'image/jpeg','.webp':'image/webp',
|
||||
'.svg':'image/svg+xml','.woff':'font/woff','.woff2':'font/woff2','.ttf':'font/ttf','.otf':'font/otf'};
|
||||
|
||||
async function render(request) {
|
||||
const {chromium} = require('playwright-core');
|
||||
const libraries = Object.fromEntries(['pagedjs','katex','mermaid'].map(name => [name, packageRoot(name)]));
|
||||
const candidates = [process.env.CHROME_BIN, process.env.CHROME_PATH, '/usr/bin/chromium', '/usr/bin/chromium-browser'];
|
||||
const executablePath = candidates.find(value => value && fs.existsSync(value));
|
||||
if (!executablePath) throw new Error('基础镜像缺少可用 Chromium;检查 CHROME_BIN,任务中不能下载浏览器');
|
||||
const sourceRoot = fs.realpathSync(path.dirname(request.input));
|
||||
const assetRoot = fs.realpathSync(request.assets);
|
||||
const errors = [], blocked = [];
|
||||
const nonce = crypto.randomBytes(18).toString('base64');
|
||||
const policy = `default-src 'none'; script-src 'nonce-${nonce}'; style-src 'unsafe-inline' https://pdf.local; img-src https://pdf.local data: blob:; font-src https://pdf.local data:; connect-src https://pdf.local; object-src 'none'; base-uri 'none'; form-action 'none'; frame-src 'none'`;
|
||||
async function inject(page, filename) {
|
||||
await page.evaluate(({content, nonce}) => {const script=document.createElement('script'); script.nonce=nonce; script.textContent=content; document.head.append(script);}, {content:fs.readFileSync(filename,'utf8'), nonce});
|
||||
}
|
||||
const browser = await chromium.launch({headless:true, executablePath,
|
||||
args:['--no-sandbox','--disable-setuid-sandbox','--disable-dev-shm-usage'], timeout:request.timeout_ms});
|
||||
try {
|
||||
const context = await browser.newContext({serviceWorkers:'block', acceptDownloads:false});
|
||||
const page = await context.newPage();
|
||||
page.setDefaultTimeout(request.timeout_ms);
|
||||
page.on('pageerror', error => errors.push(error.message.slice(0, 500)));
|
||||
page.on('console', message => {
|
||||
if (message.type() === 'error') errors.push(message.text().slice(0, 500));
|
||||
});
|
||||
await context.route('**/*', async route => {
|
||||
const url = new URL(route.request().url());
|
||||
if (url.protocol === 'data:' || url.protocol === 'blob:') return route.continue();
|
||||
let root = sourceRoot, relative = decodeURIComponent(url.pathname).replace(/^\/+/, '');
|
||||
if (url.origin !== 'https://pdf.local') {
|
||||
blocked.push('外部资源:' + url.hostname); return route.abort();
|
||||
}
|
||||
if (relative.startsWith('__skill__/')) { root = assetRoot; relative = relative.slice(10); }
|
||||
else if (relative.startsWith('__lib__/')) {
|
||||
const parts = relative.split('/'); root = libraries[parts[1]]; relative = parts.slice(2).join('/');
|
||||
}
|
||||
try {
|
||||
if (!root) throw new Error('unknown root');
|
||||
const target = fs.realpathSync(path.resolve(root, relative));
|
||||
if (!within(root, target) || !MIME[path.extname(target).toLowerCase()]) throw new Error('forbidden asset');
|
||||
if (fs.statSync(target).size > 25 * 1024 * 1024) throw new Error('asset too large');
|
||||
return route.fulfill({body:fs.readFileSync(target), contentType:MIME[path.extname(target).toLowerCase()], headers:{'Content-Security-Policy':policy}});
|
||||
} catch (_) { blocked.push('缺失或不允许的本地资源:' + relative.slice(0, 200)); return route.abort(); }
|
||||
});
|
||||
await page.goto('https://pdf.local/' + encodeURIComponent(path.basename(request.input)), {waitUntil:'load'});
|
||||
if (await page.evaluate(() => document.compatMode !== 'CSS1Compat'))
|
||||
throw new Error('HTML 需以 <!doctype html> 声明标准模式,否则分页和数学公式无法正确排版');
|
||||
// Reject active content in the actual DOM too, after the Python preflight.
|
||||
const active = await page.evaluate(() => [...document.querySelectorAll('*')].some(el =>
|
||||
['SCRIPT','IFRAME','OBJECT','EMBED','BASE','FRAME'].includes(el.tagName) ||
|
||||
[...el.attributes].some(a => /^on/i.test(a.name) || a.name === 'srcdoc')));
|
||||
if (active) throw new Error('HTML 包含主动脚本内容;只允许静态 HTML/CSS/SVG');
|
||||
await page.emulateMedia({media:'print'});
|
||||
let css = fs.readFileSync(path.join(assetRoot, 'design.css'), 'utf8');
|
||||
if (request.page_size === 'LETTER') css = css.replaceAll('210mm','215.9mm').replaceAll('297mm','279.4mm').replaceAll('size: A4','size: Letter');
|
||||
// Defaults precede author styles so explicit user typography takes precedence.
|
||||
await page.evaluate(css => {const style=document.createElement('style'); style.textContent=css; document.head.prepend(style);}, css);
|
||||
if (request.css) await page.addStyleTag({content:fs.readFileSync(request.css,'utf8')});
|
||||
const stats = await page.evaluate(() => ({figures:document.querySelectorAll('figure').length,
|
||||
tables:document.querySelectorAll('table').length, mermaid:document.querySelectorAll('.mermaid').length,
|
||||
math:document.querySelectorAll('.math-inline,.math-display').length}));
|
||||
if (stats.mermaid) {
|
||||
const unsafe = await page.locator('.mermaid').evaluateAll(nodes => nodes.some(n => /%%\s*\{|^\s*---/m.test(n.textContent)));
|
||||
if (unsafe) throw new Error('Mermaid 不允许内嵌配置;主题和安全选项由固定渲染器设置');
|
||||
await inject(page, path.join(libraries.mermaid,'dist/mermaid.min.js'));
|
||||
await page.evaluate(async () => {
|
||||
mermaid.initialize({startOnLoad:false, securityLevel:'strict', theme:'neutral', maxTextSize:50000,
|
||||
flowchart:{htmlLabels:false}, suppressErrorRendering:true});
|
||||
await mermaid.run({querySelector:'.mermaid'});
|
||||
});
|
||||
}
|
||||
if (stats.math) {
|
||||
await page.addStyleTag({url:'https://pdf.local/__lib__/katex/dist/katex.min.css'});
|
||||
await inject(page, path.join(libraries.katex,'dist/katex.min.js'));
|
||||
await page.evaluate(() => {
|
||||
for (const el of document.querySelectorAll('.math-inline,.math-display')) {
|
||||
katex.render(el.textContent, el, {displayMode:el.classList.contains('math-display'), throwOnError:true,
|
||||
trust:false, maxExpand:1000, maxSize:30, strict:'warn'});
|
||||
}
|
||||
});
|
||||
}
|
||||
await page.evaluate(async () => {
|
||||
await document.fonts.ready;
|
||||
await Promise.all([...document.images].map(img => img.complete ? Promise.resolve() : new Promise(resolve => {img.onload=img.onerror=resolve;})));
|
||||
});
|
||||
const broken = await page.evaluate(() => [...document.images].filter(img => !img.naturalWidth).length);
|
||||
if (broken || blocked.length || errors.length) throw new Error('资源加载失败:' + [...new Set([...blocked,...errors]), ...(broken ? [`${broken} 张图片不可读`] : [])].join(';'));
|
||||
await page.evaluate(() => {window.PagedConfig={auto:false};});
|
||||
await inject(page, path.join(libraries.pagedjs,'dist/paged.polyfill.js'));
|
||||
const total = await page.evaluate(async timeout => {
|
||||
const flow = await Promise.race([window.PagedPolyfill.preview(), new Promise((_,reject) =>
|
||||
setTimeout(() => reject(new Error('分页超时,未输出 PDF')), timeout))]);
|
||||
await document.fonts.ready;
|
||||
return flow.total;
|
||||
}, request.timeout_ms);
|
||||
if (!Number.isInteger(total) || total < 1 || total > 200) throw new Error('PDF 页数需在 1–200 之间');
|
||||
if (errors.length || blocked.length) throw new Error([...errors,...blocked].slice(0,8).join(';'));
|
||||
const pages = await page.evaluate(() => [...document.querySelectorAll('.pagedjs_page')].map((page, index) => {
|
||||
const area = page.querySelector('.pagedjs_page_content');
|
||||
const rect = area.getBoundingClientRect(), issues=[];
|
||||
for (const el of area.querySelectorAll('table,figure,img,svg,pre,.math-display,h1,h2,h3,p')) {
|
||||
const box=el.getBoundingClientRect();
|
||||
if (box.width && (box.left < rect.left-2 || box.right > rect.right+2 || el.scrollWidth > el.clientWidth+3))
|
||||
issues.push({tag:el.tagName.toLowerCase(),text:el.textContent.trim().slice(0,60)});
|
||||
}
|
||||
return {page:index+1,characters:area.innerText.trim().length,visuals:area.querySelectorAll('img,svg').length,overflows:issues.slice(0,10)};
|
||||
}));
|
||||
if (pages.length !== total) throw new Error('分页尚未完成:DOM 页数不一致');
|
||||
if (pages.some(p => p.overflows.length)) throw new Error('检测到内容横向溢出:' + JSON.stringify(pages.filter(p=>p.overflows.length)));
|
||||
await page.pdf({path:request.output, printBackground:true, preferCSSPageSize:true, tagged:true, scale:1,
|
||||
timeout:request.timeout_ms});
|
||||
return {ok:true,engine:'chromium+pagedjs',offline:true,page_count:total,content:stats,pages,
|
||||
warnings:pages.filter(p => !p.characters && !p.visuals).map(p=>`第 ${p.page} 页可能为空白,需要视觉检查`)};
|
||||
} finally {await browser.close();}
|
||||
}
|
||||
(async () => {
|
||||
try {const result=await render(JSON.parse(fs.readFileSync(process.argv[2],'utf8'))); process.stdout.write(JSON.stringify(result));}
|
||||
catch(error) {process.stdout.write(JSON.stringify({ok:false,error:String(error.message||error)})); process.exitCode=1;}
|
||||
})();
|
||||
57
skills/pdf/scripts/compile_latex.py
Normal file
57
skills/pdf/scripts/compile_latex.py
Normal file
@ -0,0 +1,57 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Compile a LaTeX project with cached Tectonic resources and shell escape disabled."""
|
||||
from pathlib import Path
|
||||
import re
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
|
||||
from _pdf_common import SkillArgumentParser, output_pdf, new_temp_pdf, publish_temp_file, run_cli
|
||||
|
||||
|
||||
def compile_document(args):
|
||||
from pypdf import PdfReader
|
||||
source=Path(args.input).expanduser().resolve()
|
||||
if not source.is_file() or source.suffix.lower()!='.tex' or not 0 < source.stat().st_size <= 2*1024*1024:
|
||||
raise ValueError('输入需为不超过 2 MiB 的本地 .tex 文件')
|
||||
if not 1 <= args.timeout <= 600:
|
||||
raise ValueError('timeout 必须在 1–600 秒之间')
|
||||
executable=shutil.which('tectonic')
|
||||
if not executable:
|
||||
raise RuntimeError('基础镜像缺少预置 Tectonic,需要更新镜像')
|
||||
target=output_pdf(args.output,args.overwrite)
|
||||
temporary=new_temp_pdf(target)
|
||||
try:
|
||||
with tempfile.TemporaryDirectory(prefix='pdf-latex-') as folder:
|
||||
command=[executable,'--untrusted','--only-cached','--keep-logs','--outdir',folder,str(source)]
|
||||
completed=subprocess.run(command,cwd=source.parent,capture_output=True,text=True,timeout=args.timeout,check=False)
|
||||
result=Path(folder)/(source.stem+'.pdf')
|
||||
messages=(completed.stdout+'\n'+completed.stderr).splitlines()
|
||||
if completed.returncode or not result.is_file():
|
||||
detail='\n'.join(messages[-15:])
|
||||
raise RuntimeError('LaTeX 编译失败;缺失 TeX 包需在基础镜像构建时预置,任务中不下载:'+detail)
|
||||
count=len(PdfReader(result).pages)
|
||||
if not count:
|
||||
raise ValueError('LaTeX 没有生成有效 PDF')
|
||||
logfile=Path(folder)/(source.stem+'.log')
|
||||
if logfile.exists():
|
||||
messages += logfile.read_text(errors='replace').splitlines()
|
||||
warnings=list(dict.fromkeys(line.strip() for line in messages if re.search(r'warning:|Overfull|Missing character|undefined references',line,re.I)))[:30]
|
||||
shutil.copyfile(result,temporary)
|
||||
publish_temp_file(temporary,target,args.overwrite)
|
||||
finally:
|
||||
temporary.unlink(missing_ok=True)
|
||||
return {'source':str(source),'path':str(target),'page_count':count,'engine':'tectonic','dependency_mode':'cached-only',
|
||||
'warnings':warnings,'requires_visual_review':True}
|
||||
|
||||
|
||||
def main(argv=None):
|
||||
parser=SkillArgumentParser(description='固定 LaTeX 编译接口,禁用 shell escape,只使用镜像内缓存包')
|
||||
parser.add_argument('--input',required=True); parser.add_argument('--output',required=True)
|
||||
parser.add_argument('--timeout',type=int,default=180); parser.add_argument('--overwrite',action='store_true')
|
||||
return run_cli(lambda:compile_document(parser.parse_args(sys.argv[1:] if argv is None else argv)))
|
||||
|
||||
|
||||
if __name__=='__main__':
|
||||
raise SystemExit(main())
|
||||
57
skills/pdf/scripts/convert_to_pdf.py
Normal file
57
skills/pdf/scripts/convert_to_pdf.py
Normal file
@ -0,0 +1,57 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Office to PDF using the preinstalled LibreOffice, isolated per invocation."""
|
||||
from pathlib import Path
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
|
||||
from _pdf_common import SkillArgumentParser, output_pdf, new_temp_pdf, publish_temp_file, run_cli
|
||||
|
||||
FORMATS={'.docx','.doc','.odt','.rtf','.pptx','.ppt','.odp','.xlsx','.xls','.ods'}
|
||||
|
||||
|
||||
def convert(args):
|
||||
from pypdf import PdfReader
|
||||
source=Path(args.input).expanduser().resolve()
|
||||
if not source.is_file() or source.suffix.lower() not in FORMATS or not 0 < source.stat().st_size <= 25*1024*1024:
|
||||
raise ValueError('输入需为不超过 25 MiB 的本地 Office 文档;不支持把 PDF 直接反向转成可编辑 Office')
|
||||
if not 1 <= args.timeout <= 600:
|
||||
raise ValueError('timeout 必须在 1–600 秒之间')
|
||||
executable=shutil.which('soffice') or shutil.which('libreoffice')
|
||||
if not executable:
|
||||
raise RuntimeError('基础镜像缺少 LibreOffice,需要更新镜像')
|
||||
target=output_pdf(args.output,args.overwrite)
|
||||
temporary=new_temp_pdf(target)
|
||||
try:
|
||||
with tempfile.TemporaryDirectory(prefix='pdf-office-') as folder:
|
||||
root=Path(folder)
|
||||
incoming=root/'input'; outgoing=root/'output'; profile=root/'profile'
|
||||
incoming.mkdir(); outgoing.mkdir(); profile.mkdir()
|
||||
local=incoming/('source'+source.suffix.lower()); shutil.copyfile(source,local)
|
||||
command=[executable,'-env:UserInstallation='+profile.as_uri(),'--headless','--nologo','--nodefault',
|
||||
'--nofirststartwizard','--convert-to','pdf','--outdir',str(outgoing),str(local)]
|
||||
completed=subprocess.run(command,capture_output=True,text=True,timeout=args.timeout,check=False)
|
||||
result=outgoing/'source.pdf'
|
||||
if completed.returncode or not result.is_file():
|
||||
raise RuntimeError('Office 转 PDF 失败:'+(completed.stderr or completed.stdout)[-1500:])
|
||||
count=len(PdfReader(result).pages)
|
||||
if not count:
|
||||
raise ValueError('转换结果没有页面')
|
||||
shutil.copyfile(result,temporary)
|
||||
publish_temp_file(temporary,target,args.overwrite)
|
||||
finally:
|
||||
temporary.unlink(missing_ok=True)
|
||||
return {'source':str(source),'path':str(target),'page_count':count,'engine':'libreoffice','requires_visual_review':True,
|
||||
'note':'转换前需在对应 Office skill 中重算公式并核对字体、图表和打印范围。'}
|
||||
|
||||
|
||||
def main(argv=None):
|
||||
parser=SkillArgumentParser(description='Office 文档导出 PDF,源文件保持不变')
|
||||
parser.add_argument('--input',required=True); parser.add_argument('--output',required=True)
|
||||
parser.add_argument('--timeout',type=int,default=180); parser.add_argument('--overwrite',action='store_true')
|
||||
return run_cli(lambda:convert(parser.parse_args(sys.argv[1:] if argv is None else argv)))
|
||||
|
||||
|
||||
if __name__=='__main__':
|
||||
raise SystemExit(main())
|
||||
107
skills/pdf/scripts/create_design_pdf.py
Normal file
107
skills/pdf/scripts/create_design_pdf.py
Normal file
@ -0,0 +1,107 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Render static HTML/CSS through the fixed, offline publication renderer."""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
from html.parser import HTMLParser
|
||||
from pathlib import Path
|
||||
|
||||
from _pdf_common import SkillArgumentParser, new_temp_pdf, output_pdf, publish_temp_file, run_cli
|
||||
|
||||
MAX_SOURCE_BYTES = 2 * 1024 * 1024
|
||||
|
||||
|
||||
class StaticHTML(HTMLParser):
|
||||
def handle_starttag(self, tag, attributes):
|
||||
tag = tag.lower()
|
||||
attrs = {key.lower(): value or '' for key, value in attributes}
|
||||
if tag in {'script', 'iframe', 'object', 'embed', 'base', 'frame', 'frameset'}:
|
||||
raise ValueError(f'HTML 不允许 {tag};仅支持静态 HTML/CSS/SVG,公式和 Mermaid 由固定渲染器处理')
|
||||
if any(key.startswith('on') for key in attrs) or 'srcdoc' in attrs:
|
||||
raise ValueError('HTML 不允许事件处理程序或 srcdoc')
|
||||
if tag == 'meta' and 'http-equiv' in attrs:
|
||||
raise ValueError('HTML 不允许 http-equiv;网络和文档策略由固定渲染器设置')
|
||||
for key in ('href', 'src', 'xlink:href', 'action', 'formaction'):
|
||||
value = ''.join(attrs.get(key, '').split()).lower()
|
||||
if value.startswith(('javascript:', 'vbscript:', 'file:')):
|
||||
raise ValueError('HTML 不允许脚本 URL 或 file: 资源;使用任务目录内相对路径')
|
||||
|
||||
handle_startendtag = handle_starttag
|
||||
|
||||
|
||||
def read_source(value: str, suffixes: set[str]) -> Path:
|
||||
source = Path(value).expanduser().resolve()
|
||||
if not source.is_file() or source.suffix.lower() not in suffixes:
|
||||
raise ValueError(f'输入必须是本地 {sorted(suffixes)} 文件')
|
||||
if not 0 < source.stat().st_size <= MAX_SOURCE_BYTES:
|
||||
raise ValueError('HTML/CSS 文件必须非空且不超过 2 MiB')
|
||||
return source
|
||||
|
||||
|
||||
def create(args) -> dict:
|
||||
from pypdf import PdfReader
|
||||
|
||||
source = read_source(args.input, {'.html', '.htm'})
|
||||
text = source.read_text(encoding='utf-8-sig')
|
||||
validator = StaticHTML(convert_charrefs=True)
|
||||
validator.feed(text)
|
||||
validator.close()
|
||||
css = read_source(args.css, {'.css'}) if args.css else None
|
||||
if not 10 <= args.timeout <= 600:
|
||||
raise ValueError('timeout 必须在 10–600 秒之间')
|
||||
if args.expected_pages is not None and not 1 <= args.expected_pages <= 200:
|
||||
raise ValueError('expected-pages 必须在 1–200 之间')
|
||||
output = output_pdf(args.output, args.overwrite)
|
||||
node = shutil.which('node')
|
||||
if not node:
|
||||
raise RuntimeError('基础镜像缺少预置 Node 运行时;需要更新镜像,任务中不能安装')
|
||||
temporary = new_temp_pdf(output)
|
||||
try:
|
||||
with tempfile.TemporaryDirectory(prefix='pdf-design-') as folder:
|
||||
request = Path(folder) / 'request.json'
|
||||
request.write_text(json.dumps({'input': str(source), 'output': str(temporary), 'css': str(css) if css else None,
|
||||
'page_size': args.page_size, 'timeout_ms': args.timeout * 1000,
|
||||
'assets': str(Path(__file__).resolve().parents[1] / 'assets')}, ensure_ascii=False))
|
||||
result = subprocess.run([node, str(Path(__file__).with_name('_render_html.cjs')), str(request)],
|
||||
capture_output=True, text=True, timeout=args.timeout + 20, check=False)
|
||||
try:
|
||||
details = json.loads(result.stdout)
|
||||
except (ValueError, TypeError):
|
||||
raise RuntimeError('排版器未返回有效 JSON:' + (result.stderr or result.stdout)[-1500:])
|
||||
if result.returncode or not details.get('ok'):
|
||||
raise RuntimeError(details.get('error', 'HTML 排版失败'))
|
||||
with temporary.open('rb') as handle:
|
||||
reader = PdfReader(handle)
|
||||
page_count = len(reader.pages)
|
||||
if reader.is_encrypted or not page_count:
|
||||
raise ValueError('排版器未生成有效 PDF')
|
||||
if page_count != details['page_count']:
|
||||
raise ValueError(f'分页结果与 PDF 页数不一致:{details["page_count"]} / {page_count}')
|
||||
if args.expected_pages is not None and page_count != args.expected_pages:
|
||||
raise ValueError(f'实际 {page_count} 页,与要求的 {args.expected_pages} 页不符;请调整排版后重试')
|
||||
publish_temp_file(temporary, output, args.overwrite)
|
||||
finally:
|
||||
temporary.unlink(missing_ok=True)
|
||||
details.pop('ok', None)
|
||||
return {'path': str(output), 'source': str(source), 'size_bytes': output.stat().st_size, **details,
|
||||
'requires_visual_review': True}
|
||||
|
||||
|
||||
def main(argv=None) -> int:
|
||||
parser = SkillArgumentParser(description='静态 HTML/CSS 排版为 PDF;使用离线 Paged.js、KaTeX、Mermaid 和系统 Chromium')
|
||||
parser.add_argument('--input', required=True)
|
||||
parser.add_argument('--output', required=True)
|
||||
parser.add_argument('--css')
|
||||
parser.add_argument('--page-size', choices=('A4', 'LETTER'), default='A4')
|
||||
parser.add_argument('--expected-pages', type=int)
|
||||
parser.add_argument('--timeout', type=int, default=180)
|
||||
parser.add_argument('--overwrite', action='store_true')
|
||||
return run_cli(lambda: create(parser.parse_args(sys.argv[1:] if argv is None else argv)))
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
raise SystemExit(main())
|
||||
246
skills/pdf/scripts/edit_pdf.py
Normal file
246
skills/pdf/scripts/edit_pdf.py
Normal file
@ -0,0 +1,246 @@
|
||||
#!/usr/bin/env python3
|
||||
"""AcroForm, crop, metadata and embedded-image operations using existing pypdf."""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import math
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
from _pdf_common import (SkillArgumentParser, input_pdf, output_pdf, output_directory,
|
||||
new_temp_pdf, publish_temp_file, parse_page_spec, run_cli)
|
||||
|
||||
|
||||
def load_data(inline, filename):
|
||||
if bool(inline) == bool(filename):
|
||||
raise ValueError('必须且只能提供 --data 或 --data-file')
|
||||
if filename and Path(filename).stat().st_size > 2 * 1024 * 1024:
|
||||
raise ValueError('JSON 上限为 2 MiB')
|
||||
raw = Path(filename).read_bytes() if filename else inline.encode('utf-8')
|
||||
if len(raw) > 2 * 1024 * 1024:
|
||||
raise ValueError('JSON 上限为 2 MiB')
|
||||
data = json.loads(raw)
|
||||
if not isinstance(data, dict):
|
||||
raise ValueError('JSON 必须是对象')
|
||||
return data
|
||||
|
||||
|
||||
def inherited_value(field, key, default=None):
|
||||
"""Read an inheritable field attribute without following malformed cycles."""
|
||||
seen = set()
|
||||
while field is not None:
|
||||
field = field.get_object()
|
||||
if id(field) in seen:
|
||||
raise ValueError('PDF 表单字段的父级引用存在循环')
|
||||
seen.add(id(field))
|
||||
if key in field:
|
||||
return field[key]
|
||||
field = field.get('/Parent')
|
||||
return default
|
||||
|
||||
|
||||
def field_info(reader):
|
||||
result = {}
|
||||
for name, field in (reader.get_fields() or {}).items():
|
||||
# get_fields() is a summary: /AP and /MaxLen live on the full object.
|
||||
original = field.indirect_reference.get_object()
|
||||
ft = str(inherited_value(original, '/FT', ''))
|
||||
flags = int(inherited_value(original, '/Ff', 0))
|
||||
kind = {'/Tx':'text', '/Ch':'choice', '/Sig':'signature'}.get(ft, 'unknown')
|
||||
if ft == '/Btn':
|
||||
kind = 'pushbutton' if flags & (1 << 16) else ('radio' if flags & (1 << 15) else 'checkbox')
|
||||
widgets = [original, *[kid.get_object() for kid in original.get('/Kids', [])]]
|
||||
states = set()
|
||||
if kind in {'checkbox', 'radio'}:
|
||||
for widget in widgets:
|
||||
appearance = widget.get('/AP')
|
||||
normal = appearance.get_object().get('/N') if appearance is not None else None
|
||||
if normal is not None:
|
||||
states.update(str(key) for key in normal.get_object().keys())
|
||||
if kind == 'radio' and flags & (1 << 14):
|
||||
states.discard('/Off')
|
||||
options = [{'value':str(option[0]), 'label':str(option[1])} if isinstance(option, list) else {'value':str(option), 'label':str(option)}
|
||||
for option in inherited_value(original, '/Opt', [])]
|
||||
value = inherited_value(original, '/V')
|
||||
result[name] = {'id':name, 'type':kind, 'read_only':bool(flags & 1), 'flags':flags,
|
||||
'current_value':[str(v) for v in value] if isinstance(value, list) else (str(value) if value is not None else None),
|
||||
'states':sorted(states), 'options':options, 'max_length':inherited_value(original, '/MaxLen')}
|
||||
return result
|
||||
|
||||
|
||||
def validated_values(infos, data):
|
||||
from pypdf.generic import NameObject
|
||||
if not data or len(data) > 500:
|
||||
raise ValueError('填表数据需为 1–500 个字段')
|
||||
values = {}
|
||||
for name, value in data.items():
|
||||
if name not in infos:
|
||||
raise ValueError(f'表单字段不存在:{name}')
|
||||
field = infos[name]
|
||||
if field['read_only']:
|
||||
raise ValueError(f'字段为只读:{name}')
|
||||
kind = field['type']
|
||||
if kind in {'signature','pushbutton','unknown'}:
|
||||
raise ValueError(f'不支持填写 {kind} 字段:{name}')
|
||||
if kind in {'checkbox','radio'}:
|
||||
states = field['states']
|
||||
if type(value) is bool and kind == 'checkbox':
|
||||
choices = [s for s in states if s != '/Off']
|
||||
if value and len(choices) != 1:
|
||||
raise ValueError(f'{name} 的选中状态不唯一,请使用具体状态名')
|
||||
value = choices[0] if value else '/Off'
|
||||
elif isinstance(value, str):
|
||||
value = '/' + value.lstrip('/')
|
||||
else:
|
||||
raise ValueError(f'{name} 需布尔值或有效状态名')
|
||||
if value not in states:
|
||||
raise ValueError(f'{name} 状态无效;可选:{states}')
|
||||
values[name] = NameObject(value)
|
||||
elif kind == 'choice':
|
||||
choices = {entry['value'] for entry in field['options']}
|
||||
selected = value if isinstance(value, list) else [value]
|
||||
if isinstance(value, list) and not field['flags'] & (1 << 21):
|
||||
raise ValueError(f'{name} 不支持多选')
|
||||
editable = bool(field['flags'] & (1 << 18))
|
||||
if not all(isinstance(item, str) and (editable or item in choices) for item in selected):
|
||||
raise ValueError(f'{name} 选项无效;可选:{sorted(choices)}')
|
||||
values[name] = value
|
||||
else:
|
||||
if not isinstance(value, str) or (field['max_length'] and len(value) > field['max_length']):
|
||||
raise ValueError(f'{name} 需文本,且不得超过字段长度上限')
|
||||
values[name] = value
|
||||
return values
|
||||
|
||||
|
||||
def execute(args):
|
||||
from pypdf import PdfReader, PdfWriter
|
||||
from pypdf.generic import RectangleObject
|
||||
|
||||
source = input_pdf(args.input)
|
||||
with source.open('rb') as handle:
|
||||
reader = PdfReader(handle)
|
||||
if reader.is_encrypted:
|
||||
raise ValueError('PDF 已加密,请提供已解密副本')
|
||||
if args.operation == 'form-info':
|
||||
fields = list(field_info(reader).values())
|
||||
if args.offset < 0 or not 1 <= args.limit <= 200:
|
||||
raise ValueError('offset 不能小于 0,limit 需为 1–200')
|
||||
stop = min(len(fields), args.offset + args.limit)
|
||||
acroform = reader.trailer['/Root'].get('/AcroForm')
|
||||
return {'field_count':len(fields), 'fields':fields[args.offset:stop], 'has_more':stop < len(fields),
|
||||
'next_offset':stop if stop < len(fields) else None,
|
||||
'has_xfa':bool(acroform and '/XFA' in acroform.get_object())}
|
||||
if args.operation == 'extract-images':
|
||||
return extract_images(reader, args)
|
||||
output = output_pdf(args.output, args.overwrite)
|
||||
if output == source:
|
||||
raise ValueError('输出文件不能覆盖源 PDF')
|
||||
writer = PdfWriter(clone_from=reader)
|
||||
temporary = new_temp_pdf(output)
|
||||
extra = {}
|
||||
try:
|
||||
if args.operation == 'form-fill':
|
||||
acroform = reader.trailer['/Root'].get('/AcroForm')
|
||||
if acroform and '/XFA' in acroform.get_object():
|
||||
raise ValueError('XFA 表单不属于 AcroForm 固定接口,不能声称填写成功')
|
||||
values = validated_values(field_info(reader), load_data(args.data, args.data_file))
|
||||
writer.update_page_form_field_values(None, values, auto_regenerate=False)
|
||||
extra = {'fields_filled':list(values), 'requires_visual_review':True}
|
||||
elif args.operation == 'metadata':
|
||||
data = load_data(args.data, args.data_file)
|
||||
allowed = {'Title','Author','Subject','Keywords','Creator','Producer'}
|
||||
if set(data) - allowed or not all(isinstance(v,str) and len(v) <= 4096 for v in data.values()):
|
||||
raise ValueError('元数据仅支持 Title/Author/Subject/Keywords/Creator/Producer 文本字段')
|
||||
writer.add_metadata({'/'+key:value for key,value in data.items()})
|
||||
extra = {'updated_keys':list(data), 'xmp_preserved':'/Metadata' in reader.trailer['/Root']}
|
||||
elif args.operation == 'crop':
|
||||
box = [float(v) for v in args.box.split(',')]
|
||||
if len(box) != 4 or not all(math.isfinite(v) for v in box) or box[0] >= box[2] or box[1] >= box[3]:
|
||||
raise ValueError('box 需为左,下,右,上四个有限坐标,单位 pt')
|
||||
pages = parse_page_spec(args.pages, len(writer.pages))
|
||||
for number in pages:
|
||||
media = writer.pages[number-1].mediabox
|
||||
if box[0] < media.left or box[1] < media.bottom or box[2] > media.right or box[3] > media.top:
|
||||
raise ValueError(f'裁剪框超出第 {number} 页 MediaBox')
|
||||
writer.pages[number-1].cropbox = RectangleObject(box)
|
||||
extra = {'cropped_pages':pages, 'box':box, 'warning':'裁剪只改变可见范围,不删除隐藏内容,不能用于脱敏。'}
|
||||
with temporary.open('wb') as stream:
|
||||
writer.write(stream)
|
||||
check = PdfReader(temporary)
|
||||
if len(check.pages) != len(reader.pages):
|
||||
raise ValueError('编辑后页数异常')
|
||||
if args.operation == 'form-fill':
|
||||
actual = field_info(check)
|
||||
for name, value in values.items():
|
||||
wanted = [str(v) for v in value] if isinstance(value, list) else str(value)
|
||||
if name not in actual or actual[name]['current_value'] != wanted:
|
||||
raise ValueError(f'字段回读校验失败:{name}')
|
||||
publish_temp_file(temporary, output, args.overwrite)
|
||||
finally:
|
||||
writer.close()
|
||||
temporary.unlink(missing_ok=True)
|
||||
return {'path':str(output), 'source':str(source), 'page_count':len(reader.pages), 'operation':args.operation, **extra}
|
||||
|
||||
|
||||
def extract_images(reader, args):
|
||||
if not 1 <= args.max_images <= 50 or args.start_image < 0:
|
||||
raise ValueError('max-images 需为 1–50,start-image 不能小于 0')
|
||||
pages = parse_page_spec(args.pages, len(reader.pages))
|
||||
if len(pages) > 4:
|
||||
raise ValueError('一次最多处理 4 页,请指定 pages')
|
||||
destination = output_directory(args.output_dir)
|
||||
result, total_bytes = [], 0
|
||||
for page_index, number in enumerate(pages):
|
||||
images = reader.pages[number-1].images
|
||||
start = args.start_image if page_index == 0 else 0
|
||||
for index in range(start, len(images)):
|
||||
if len(result) >= args.max_images:
|
||||
return {'images':result, 'has_more':True, 'next_page':number, 'next_image':index,
|
||||
'remaining_pages':pages[page_index:], 'note':'嵌入图片不是整页截图;页面外观请用 render_pdf.py。'}
|
||||
item = images[index]
|
||||
if item.image.width * item.image.height > 20_000_000:
|
||||
raise ValueError('嵌入图像超过 2000 万像素,请改为限制 DPI 的页面渲染')
|
||||
total_bytes += len(item.data)
|
||||
if total_bytes > 25 * 1024 * 1024:
|
||||
raise ValueError('本批嵌入图片超过 25 MiB,请减少页数或图片数')
|
||||
extension = Path(item.name).suffix.lower()
|
||||
if extension not in {'.png','.jpg','.jpeg','.jp2','.tif','.tiff'}:
|
||||
raise ValueError(f'不支持的嵌入图片编码:{extension},请用页面渲染')
|
||||
output = destination / f'page-{number:04d}-image-{index:03d}{extension}'
|
||||
temporary = new_temp_pdf(output)
|
||||
try:
|
||||
temporary.write_bytes(item.data)
|
||||
publish_temp_file(temporary, output, args.overwrite)
|
||||
finally:
|
||||
temporary.unlink(missing_ok=True)
|
||||
result.append({'page':number, 'image_index':index, 'path':str(output)})
|
||||
return {'images':result, 'has_more':False, 'note':'嵌入图片不是整页截图;页面外观请用 render_pdf.py。'}
|
||||
|
||||
|
||||
def main(argv=None):
|
||||
parser = SkillArgumentParser(description='PDF 表单、裁剪、元数据和嵌入图片固定接口')
|
||||
commands = parser.add_subparsers(dest='operation', required=True)
|
||||
for name in ('form-info','form-fill','metadata','crop','extract-images'):
|
||||
command = commands.add_parser(name)
|
||||
command.add_argument('--input', required=True)
|
||||
if name == 'form-info':
|
||||
command.add_argument('--offset', type=int, default=0)
|
||||
command.add_argument('--limit', type=int, default=50)
|
||||
else:
|
||||
command.add_argument('--overwrite', action='store_true')
|
||||
command.add_argument('--output-dir' if name == 'extract-images' else '--output', required=True)
|
||||
if name in {'form-fill','metadata'}:
|
||||
command.add_argument('--data')
|
||||
command.add_argument('--data-file')
|
||||
if name in {'crop','extract-images'}:
|
||||
command.add_argument('--pages')
|
||||
if name == 'crop':
|
||||
command.add_argument('--box', required=True)
|
||||
if name == 'extract-images':
|
||||
command.add_argument('--start-image', type=int, default=0)
|
||||
command.add_argument('--max-images', type=int, default=20)
|
||||
return run_cli(lambda: execute(parser.parse_args(sys.argv[1:] if argv is None else argv)))
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
raise SystemExit(main())
|
||||
47
skills/pdf/scripts/ocr_image.py
Normal file
47
skills/pdf/scripts/ocr_image.py
Normal file
@ -0,0 +1,47 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Local OCR for a single image used in a PDF workflow."""
|
||||
from pathlib import Path
|
||||
import sys
|
||||
import tempfile
|
||||
|
||||
from _pdf_common import SkillArgumentParser, run_cli
|
||||
from ocr_text import _create_ocr_engine, _ocr_page, MAX_PIXELS_PER_PAGE
|
||||
|
||||
|
||||
def recognize(args):
|
||||
from PIL import Image, ImageOps
|
||||
source = Path(args.input).expanduser().resolve()
|
||||
if not source.is_file() or source.suffix.lower() not in {'.png','.jpg','.jpeg','.webp','.tif','.tiff','.bmp'}:
|
||||
raise ValueError('输入必须是本地图像文件')
|
||||
if source.stat().st_size > 25 * 1024 * 1024:
|
||||
raise ValueError('图像文件不能超过 25 MiB')
|
||||
if not 1 <= args.max_chars <= 60000 or args.start_offset < 0:
|
||||
raise ValueError('max-chars 需为 1–60000,start-offset 不能小于 0')
|
||||
with Image.open(source) as original:
|
||||
if original.width * original.height > MAX_PIXELS_PER_PAGE or getattr(original, 'n_frames', 1) != 1:
|
||||
raise ValueError('图像上限 2000 万像素,且必须为单帧;多页扫描件请按 PDF 分页 OCR')
|
||||
with tempfile.TemporaryDirectory(prefix='pdf-image-ocr-') as folder:
|
||||
normalized = Path(folder) / 'image.png'
|
||||
ImageOps.exif_transpose(original).convert('RGB').save(normalized)
|
||||
result = _ocr_page(_create_ocr_engine(), normalized)
|
||||
text = result.pop('text')
|
||||
if args.start_offset > len(text):
|
||||
raise ValueError('start-offset 超过可靠 OCR 文本长度')
|
||||
selection = text[args.start_offset:args.start_offset + args.max_chars]
|
||||
next_offset = args.start_offset + len(selection)
|
||||
return {'path': str(source), 'engine': 'rapidocr', 'offline': True, **result,
|
||||
'text': selection, 'char_count': len(text), 'offset_start': args.start_offset,
|
||||
'offset_end': next_offset, 'has_more': next_offset < len(text),
|
||||
'next_offset': next_offset if next_offset < len(text) else None}
|
||||
|
||||
|
||||
def main(argv=None):
|
||||
parser = SkillArgumentParser(description='PDF 任务图像的本地 OCR,疑难区域保留复核标记')
|
||||
parser.add_argument('--input', required=True)
|
||||
parser.add_argument('--max-chars', type=int, default=24000)
|
||||
parser.add_argument('--start-offset', type=int, default=0)
|
||||
return run_cli(lambda: recognize(parser.parse_args(sys.argv[1:] if argv is None else argv)))
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
raise SystemExit(main())
|
||||
@ -6,6 +6,7 @@ import contextlib
|
||||
import importlib.metadata
|
||||
import io
|
||||
import logging
|
||||
import math
|
||||
import os
|
||||
import re
|
||||
import shutil
|
||||
@ -18,6 +19,7 @@ from typing import Any
|
||||
|
||||
from _pdf_common import (
|
||||
SkillArgumentParser,
|
||||
check_poppler_resources,
|
||||
input_pdf,
|
||||
parse_page_spec,
|
||||
run_cli,
|
||||
@ -178,6 +180,7 @@ def _render_page(
|
||||
f"第 {page_number} 页渲染失败:"
|
||||
f"{detail or 'pdftoppm 返回错误'}"
|
||||
)
|
||||
check_poppler_resources(completed.stderr)
|
||||
if not output.is_file() or output.stat().st_size <= 0:
|
||||
raise RuntimeError(f"第 {page_number} 页没有生成有效 PNG")
|
||||
return output, elapsed
|
||||
@ -185,17 +188,29 @@ def _render_page(
|
||||
|
||||
def _create_ocr_engine():
|
||||
try:
|
||||
import rapidocr
|
||||
from rapidocr import RapidOCR
|
||||
except ImportError as exc:
|
||||
raise RuntimeError("环境预置的 rapidocr 模块不可用") from exc
|
||||
|
||||
model_dir = Path(rapidocr.__file__).resolve().parent / "models"
|
||||
models = {"Det": "PP-OCRv6_det_small.onnx", "Cls": "ch_ppocr_mobile_v2.0_cls_mobile.onnx", "Rec": "PP-OCRv6_rec_small.onnx"}
|
||||
missing = [filename for filename in models.values() if not (model_dir / filename).is_file()]
|
||||
if missing:
|
||||
raise RuntimeError(f"基础镜像缺少本地 OCR 模型:{missing};任务中不能下载")
|
||||
params = {
|
||||
"Global.log_level": "error",
|
||||
# Keep uncertain lines for explicit review instead of silently dropping them.
|
||||
"Global.text_score": 0.0,
|
||||
**{f"{name}.model_path": str(model_dir / filename) for name, filename in models.items()},
|
||||
}
|
||||
captured_stdout = io.StringIO()
|
||||
captured_stderr = io.StringIO()
|
||||
with (
|
||||
contextlib.redirect_stdout(captured_stdout),
|
||||
contextlib.redirect_stderr(captured_stderr),
|
||||
):
|
||||
return RapidOCR()
|
||||
return RapidOCR(params=params)
|
||||
|
||||
|
||||
def _clean_text(value: Any) -> str:
|
||||
@ -235,7 +250,7 @@ def _ordered_lines(result: Any) -> list[dict[str, Any]]:
|
||||
confidence = float(scores[index])
|
||||
except (IndexError, TypeError, ValueError):
|
||||
confidence = 0.0
|
||||
confidence = max(0.0, min(1.0, confidence))
|
||||
confidence = max(0.0, min(1.0, confidence)) if math.isfinite(confidence) else 0.0
|
||||
box = _box_points(boxes[index] if index < len(boxes) else None)
|
||||
if box:
|
||||
left = min(point[0] for point in box)
|
||||
@ -314,8 +329,17 @@ def _ocr_page(engine: Any, image_path: Path) -> dict[str, Any]:
|
||||
else:
|
||||
status = "good"
|
||||
|
||||
reliable_lines = [line for line in lines if line["confidence"] >= MIN_MEAN_CONFIDENCE]
|
||||
review_lines = [line for line in lines if line["confidence"] < MIN_MEAN_CONFIDENCE]
|
||||
# A high page average must not certify a low-confidence amount or identifier.
|
||||
reliable_text = "\n".join(line["text"] for line in reliable_lines)
|
||||
return {
|
||||
"text": text,
|
||||
"text": reliable_text if status == "good" else "",
|
||||
"raw_char_count": len(text),
|
||||
"needs_review": status != "good" or bool(review_lines),
|
||||
"review_regions": [{"text_candidate": line["text"][:500], "confidence": round(line["confidence"], 4), "box": line["box"]} for line in review_lines[:40]],
|
||||
"review_regions_truncated": len(review_lines) > 40,
|
||||
"reading_order": "geometric_top_to_bottom",
|
||||
"status": status,
|
||||
"usable_for_summary": status == "good",
|
||||
"line_count": len(lines),
|
||||
@ -436,9 +460,10 @@ def _extract(args) -> dict[str, Any]:
|
||||
),
|
||||
"complete_ocr_coverage": (
|
||||
all_processed and all_complete and all_usable
|
||||
and not any(page["needs_review"] for page in page_outputs)
|
||||
),
|
||||
"needs_review": any(
|
||||
not page["usable_for_summary"] for page in page_outputs
|
||||
page["needs_review"] for page in page_outputs
|
||||
),
|
||||
"has_more": has_more,
|
||||
"next_page": next_page,
|
||||
|
||||
@ -12,6 +12,7 @@ from typing import Any
|
||||
|
||||
from _pdf_common import (
|
||||
SkillArgumentParser,
|
||||
check_poppler_resources,
|
||||
input_pdf,
|
||||
output_directory,
|
||||
publish_temp_file,
|
||||
@ -155,6 +156,7 @@ def _render(args) -> dict[str, Any]:
|
||||
if completed.returncode != 0:
|
||||
detail = (completed.stderr or completed.stdout or "").strip()[-2000:]
|
||||
raise RuntimeError(f"PDF 渲染失败:{detail or 'pdftoppm 返回错误'}")
|
||||
check_poppler_resources(completed.stderr)
|
||||
|
||||
rendered = sorted(
|
||||
temp_dir.glob("page-*.png"),
|
||||
|
||||
246
skills/pdf/tests/test_pdf_workflows.py
Normal file
246
skills/pdf/tests/test_pdf_workflows.py
Normal file
@ -0,0 +1,246 @@
|
||||
"""Behavioral checks for OCR quality gates and non-destructive PDF operations.
|
||||
|
||||
Uses the existing pypdf, ReportLab and Pillow dependencies. Browser, OCR-model
|
||||
and Office integration checks run separately against real installed engines.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import json
|
||||
import sys
|
||||
import tempfile
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
from types import SimpleNamespace
|
||||
from unittest.mock import patch
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "scripts"))
|
||||
|
||||
from PIL import Image
|
||||
from pypdf import PdfReader, PdfWriter
|
||||
from pypdf.generic import DictionaryObject, NameObject, TextStringObject
|
||||
from reportlab.pdfgen import canvas
|
||||
|
||||
import _pdf_common as common
|
||||
import create_design_pdf
|
||||
import edit_pdf
|
||||
import ocr_text
|
||||
import render_pdf
|
||||
|
||||
|
||||
class PDFWorkflows(unittest.TestCase):
|
||||
def setUp(self):
|
||||
self.directory = tempfile.TemporaryDirectory(prefix="pdf-regression-")
|
||||
self.root = Path(self.directory.name).resolve()
|
||||
self.root_patch = patch.object(common, "PDF_OUTPUT_ROOT", self.root)
|
||||
self.root_patch.start()
|
||||
self.source = self.root / "source.pdf"
|
||||
c = canvas.Canvas(str(self.source), pagesize=(600, 800))
|
||||
c.drawString(30, 750, "Original content 123")
|
||||
form = c.acroForm
|
||||
form.textfield(name="name", x=30, y=650, width=200, height=25, maxlen=40)
|
||||
form.checkbox(name="agree", x=30, y=600, checked=False)
|
||||
form.radio(name="mode", value="A", x=30, y=550, selected=True)
|
||||
form.radio(name="mode", value="B", x=80, y=550, selected=False)
|
||||
form.choice(name="country", value="CN", options=["CN", "US"], x=30, y=480, width=150, height=30)
|
||||
image = Image.new("RGB", (24, 16), "#8a3a2a")
|
||||
c.drawInlineImage(image, 350, 500, width=96, height=64)
|
||||
c.showPage()
|
||||
c.save()
|
||||
self.original_hash = hashlib.sha256(self.source.read_bytes()).hexdigest()
|
||||
|
||||
def tearDown(self):
|
||||
self.assertEqual(hashlib.sha256(self.source.read_bytes()).hexdigest(), self.original_hash)
|
||||
self.root_patch.stop()
|
||||
self.directory.cleanup()
|
||||
|
||||
def args(self, operation, **values):
|
||||
return SimpleNamespace(operation=operation, input=str(self.source),
|
||||
output=str(self.root / "result.pdf"), overwrite=False,
|
||||
data=None, data_file=None, pages=None, **values)
|
||||
|
||||
def test_form_inventory_and_roundtrip_all_supported_types(self):
|
||||
fields = edit_pdf.execute(self.args("form-info", offset=0, limit=50))
|
||||
infos = {field["id"]: field for field in fields["fields"]}
|
||||
self.assertEqual(infos["agree"]["type"], "checkbox")
|
||||
self.assertEqual(set(infos["agree"]["states"]), {"/Off", "/Yes"})
|
||||
self.assertEqual(set(infos["mode"]["states"]), {"/A", "/B"})
|
||||
self.assertFalse(fields["has_xfa"])
|
||||
args = self.args("form-fill")
|
||||
args.data = json.dumps({"name": "Alice 123", "agree": True, "mode": "B", "country": "US"})
|
||||
result = edit_pdf.execute(args)
|
||||
updated = PdfReader(result["path"])
|
||||
fields = updated.get_fields()
|
||||
self.assertEqual(fields["name"]["/V"], "Alice 123")
|
||||
self.assertEqual(fields["agree"]["/V"], "/Yes")
|
||||
self.assertEqual(fields["mode"]["/V"], "/B")
|
||||
self.assertEqual(fields["country"]["/V"], "US")
|
||||
widgets = [ref.get_object() for ref in updated.pages[0]["/Annots"]]
|
||||
checkbox = next(widget for widget in widgets if widget.get("/T") == "agree")
|
||||
self.assertEqual(checkbox["/AS"], "/Yes")
|
||||
radio = [widget for widget in widgets if widget.get("/Parent")]
|
||||
self.assertEqual(sorted(str(widget["/AS"]) for widget in radio), ["/B", "/Off"])
|
||||
self.assertIn("Original content 123", updated.pages[0].extract_text())
|
||||
|
||||
def test_checkbox_false_is_off(self):
|
||||
args = self.args("form-fill")
|
||||
args.data = '{"agree":false}'
|
||||
result = edit_pdf.execute(args)
|
||||
self.assertEqual(PdfReader(result["path"]).get_fields()["agree"]["/V"], "/Off")
|
||||
|
||||
def test_invalid_form_input_does_not_publish(self):
|
||||
for data in [{"missing":"x"}, {"agree":"false"}, {"mode":True}, {"mode":"Off"},
|
||||
{"country":"ZZ"}, {"country":["CN","US"]}, {"name":"x"*41}]:
|
||||
with self.subTest(data=data):
|
||||
args = self.args("form-fill")
|
||||
args.data = json.dumps(data)
|
||||
with self.assertRaises(ValueError):
|
||||
edit_pdf.execute(args)
|
||||
self.assertFalse(Path(args.output).exists())
|
||||
|
||||
def test_form_inventory_handles_indirect_appearance(self):
|
||||
writer = PdfWriter(clone_from=self.source)
|
||||
for ref in writer.pages[0]["/Annots"]:
|
||||
widget = ref.get_object()
|
||||
if "/AP" in widget:
|
||||
widget[NameObject("/AP")] = writer._add_object(widget["/AP"])
|
||||
alternative = self.root / "indirect.pdf"
|
||||
writer.write(alternative)
|
||||
args = self.args("form-info", offset=0, limit=50)
|
||||
args.input = str(alternative)
|
||||
result = edit_pdf.execute(args)
|
||||
self.assertEqual(result["field_count"], 4)
|
||||
|
||||
def test_xfa_rejected_for_fill(self):
|
||||
writer = PdfWriter(clone_from=self.source)
|
||||
writer.root_object["/AcroForm"][NameObject("/XFA")] = TextStringObject("unsupported")
|
||||
alternative = self.root / "xfa.pdf"
|
||||
writer.write(alternative)
|
||||
args = self.args("form-fill")
|
||||
args.input = str(alternative)
|
||||
args.data = '{"name":"Alice"}'
|
||||
with self.assertRaisesRegex(ValueError, "XFA"):
|
||||
edit_pdf.execute(args)
|
||||
self.assertFalse(Path(args.output).exists())
|
||||
|
||||
def test_crop_keeps_content_and_forms(self):
|
||||
result = edit_pdf.execute(self.args("crop", box="50,50,500,700"))
|
||||
reader = PdfReader(result["path"])
|
||||
self.assertEqual(list(reader.pages[0].cropbox), [50, 50, 500, 700])
|
||||
self.assertIn("Original content 123", reader.pages[0].extract_text())
|
||||
self.assertEqual(len(reader.get_fields()), 4)
|
||||
|
||||
def test_crop_rejects_out_of_bounds_and_nonfinite_values(self):
|
||||
for box in ["-1,0,300,400", "0,0,601,800", "10,0,0,20", "0,0,nan,20"]:
|
||||
with self.subTest(box=box), self.assertRaises(ValueError):
|
||||
edit_pdf.execute(self.args("crop", box=box))
|
||||
self.assertFalse((self.root / "result.pdf").exists())
|
||||
|
||||
def test_metadata_preserves_forms_and_unspecified_fields(self):
|
||||
original = PdfReader(self.source).metadata
|
||||
args = self.args("metadata")
|
||||
args.data = '{"Title":"中文报告","Author":"Test"}'
|
||||
result = edit_pdf.execute(args)
|
||||
reader = PdfReader(result["path"])
|
||||
self.assertEqual(reader.metadata.title, "中文报告")
|
||||
self.assertEqual(reader.metadata.producer, original.producer)
|
||||
self.assertEqual(len(reader.get_fields()), 4)
|
||||
|
||||
def test_embedded_image_has_original_dimensions(self):
|
||||
args = self.args("extract-images", output_dir=str(self.root / "images"), start_image=0, max_images=20)
|
||||
result = edit_pdf.execute(args)
|
||||
self.assertEqual(len(result["images"]), 1)
|
||||
with Image.open(result["images"][0]["path"]) as embedded:
|
||||
self.assertEqual(embedded.size, (24, 16))
|
||||
|
||||
def test_existing_output_is_preserved_on_failure(self):
|
||||
target = self.root / "result.pdf"
|
||||
target.write_bytes(b"previous result")
|
||||
args = self.args("metadata")
|
||||
args.overwrite = True
|
||||
args.data = '{"Unsupported":"value"}'
|
||||
with self.assertRaises(ValueError):
|
||||
edit_pdf.execute(args)
|
||||
self.assertEqual(target.read_bytes(), b"previous result")
|
||||
|
||||
def test_output_cannot_replace_source(self):
|
||||
args = self.args("crop", box="0,0,100,100")
|
||||
args.output = str(self.source)
|
||||
args.overwrite = True
|
||||
with self.assertRaises(ValueError):
|
||||
edit_pdf.execute(args)
|
||||
|
||||
def test_output_directory_boundary(self):
|
||||
args = self.args("crop", box="0,0,100,100")
|
||||
args.output = str(self.root.parent / "outside.pdf")
|
||||
with self.assertRaises(ValueError):
|
||||
edit_pdf.execute(args)
|
||||
|
||||
def test_missing_cjk_maps_rejects_incomplete_preview(self):
|
||||
destination = self.root / 'preview'
|
||||
args = render_pdf._parse_args(['--input', str(self.source), '--output-dir', str(destination)])
|
||||
completed = SimpleNamespace(returncode=0, stdout='', stderr="Syntax Error: Missing language pack for 'Adobe-GB1' mapping")
|
||||
with patch.object(render_pdf.shutil, 'which', return_value=sys.executable), patch.object(render_pdf.subprocess, 'run', return_value=completed):
|
||||
with self.assertRaisesRegex(RuntimeError, 'poppler-data'):
|
||||
render_pdf._render(args)
|
||||
self.assertEqual(list(destination.glob('*.png')), [])
|
||||
|
||||
|
||||
class OCRQuality(unittest.TestCase):
|
||||
def recognize(self, texts, scores):
|
||||
boxes = [[[0, i*30], [100, i*30], [100, i*30+20], [0, i*30+20]] for i in range(len(texts))]
|
||||
engine = lambda _: SimpleNamespace(txts=texts, scores=scores, boxes=boxes)
|
||||
return ocr_text._ocr_page(engine, Path("fixture.png"))
|
||||
|
||||
def test_mixed_confidence_does_not_certify_uncertain_amount(self):
|
||||
result = self.recognize(["Reliable document heading and text", "9999.99"], [.99, .3])
|
||||
self.assertEqual(result["status"], "good")
|
||||
self.assertNotIn("9999.99", result["text"])
|
||||
self.assertTrue(result["needs_review"])
|
||||
self.assertEqual(result["review_regions"][0]["text_candidate"], "9999.99")
|
||||
self.assertIsNotNone(result["review_regions"][0]["box"])
|
||||
|
||||
def test_unusable_page_is_not_returned_as_reliable_text(self):
|
||||
for texts, scores in [(["uncertain document"], [.3]), (["Hi"], [.99]), ([], [])]:
|
||||
with self.subTest(texts=texts):
|
||||
result = self.recognize(texts, scores)
|
||||
self.assertEqual(result["text"], "")
|
||||
self.assertFalse(result["usable_for_summary"])
|
||||
self.assertTrue(result["needs_review"])
|
||||
|
||||
def test_good_ocr_needs_no_model_assistance(self):
|
||||
result = self.recognize(["中文识别测试 12345", "English document"], [.99, .98])
|
||||
self.assertFalse(result["needs_review"])
|
||||
self.assertIn("中文识别测试", result["text"])
|
||||
|
||||
def test_invalid_confidence_is_not_treated_as_certain(self):
|
||||
result = self.recognize(['Unknown confidence amount 9999'], [float('nan')])
|
||||
self.assertEqual(result['text'], '')
|
||||
self.assertTrue(result['needs_review'])
|
||||
|
||||
def test_missing_local_models_fails_before_engine_initialization(self):
|
||||
with tempfile.TemporaryDirectory() as folder:
|
||||
fake = SimpleNamespace(__file__=str(Path(folder) / "__init__.py"))
|
||||
fake.RapidOCR = lambda **_: self.fail("Engine must not download missing models")
|
||||
with patch.dict(sys.modules, {"rapidocr":fake}):
|
||||
with self.assertRaisesRegex(RuntimeError, "本地 OCR 模型"):
|
||||
ocr_text._create_ocr_engine()
|
||||
|
||||
|
||||
class StaticDocument(unittest.TestCase):
|
||||
def test_active_content_is_rejected(self):
|
||||
for html in ["<script>alert(1)</script>", '<svg onload="x()"></svg>',
|
||||
'<iframe src="https://example.com"></iframe>',
|
||||
'<a href="javascript:alert(1)">x</a>',
|
||||
'<meta http-equiv="refresh" content="0;url=https://example.com">',
|
||||
'<img src="file:///etc/passwd">']:
|
||||
with self.subTest(html=html), self.assertRaises(ValueError):
|
||||
create_design_pdf.StaticHTML().feed(html)
|
||||
|
||||
def test_static_math_diagram_and_svg_are_accepted(self):
|
||||
html = '<h1>报告</h1><p class="math-inline">E=mc^2</p><div class="mermaid">flowchart LR\nA-->B</div><svg><rect width="10" height="10" /></svg>'
|
||||
create_design_pdf.StaticHTML().feed(html)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
Loading…
Reference in New Issue
Block a user