feat: 增强 pdf 技能

This commit is contained in:
hp0912 2026-09-09 12:30:35 +08:00
parent 33809a79e4
commit 67351ccbdd
21 changed files with 1547 additions and 236 deletions

View File

@ -1,233 +1,55 @@
--- ---
name: pdf name: pdf
description: "处理本地 PDF 文件或远程 HTTPS PDF 链接,并下载 PDF 任务所需且不超过 25 MiB 的图片、音视频、压缩包和其他 HTTPS 附件;包括源 PDF 安全下载、元数据与页面检查、多引擎分段文本提取和质量检测、扫描页本地 OCR、表格提取、按需页面 PNG 渲染、从文本创建 PDF、合并、拆分、旋转及最终质量校验。当用户提供 .pdf 文件或 HTTPS PDF 地址,或要求总结、读取、识别扫描件、生成、编辑、转换或审阅 PDF 时使用。" description: "读取、OCR、创建、设计排版、转换和处理 PDF。支持本地文件与 HTTPS PDF、扫描件和任务图像,以原生文本提取及本地 OCR 为主,模型辅助疑难复核;提供 HTML/CSS 出版排版、文本生成、Office/LaTeX 导出、表单、页面与元数据操作。用户要求阅读或总结 PDF、识别扫描件、制作报告/简历/提案 PDF 或编辑现有 PDF 时使用;Office 主文档编辑由对应 skill 处理。"
--- ---
# PDF 处理 # PDF 读取、设计与处理
## 强制执行规则 ## 运行约定
当前智能体不能执行 shell、任意 Python 代码或系统命令。只能通过 `execute_skill_script` 调用本 Skill 中真实存在的固定脚本。 - 当前机器人不能直接运行 Bash、Python、Node.js 或系统命令。只通过 `execute_skill_script` 调用下列真实存在的固定脚本;参数是文件、文本、页码或受控 JSON,不能传 shell 命令、`-c`、`eval`、代码片段或解释器命令。固定脚本可在内部调用预置引擎,调用者不直接执行底层程序。
- `scripts/_pdf_common.py`、`scripts/_render_html.cjs` 是内部实现,不直接执行;不调用原 pdf-skill 的 shell 安装器、通用 Python/Node CLI,不创建临时可执行脚本。
- 只调用下表列出的可执行脚本。 - 依赖只在基础镜像构建时安装。任务中不运行 pip/npm/apt、不下载浏览器、TeX 包或 OCR 模型;缺失时报告具体依赖和镜像需更新,不能假装处理成功。
- 不执行 `scripts/` 目录,不执行内部模块 `scripts/_pdf_common.py`。 - 最终 PDF 写入 `/usr/local/src/pdf/`;缓存、源 HTML/JSON、预览放在 `/usr/local/src/pdf/tmp/<任务名>/`。传绝对路径,保留用户源文件。`--overwrite` 仅用于本任务已生成的旧产物。
- 不传 `-c`、`python3`、`ls`、`pdftoppm` 或其他 shell/系统命令作为脚本参数。 - 每次检查返回 JSON;`ok: false` 先处理原因。`requires_visual_review`、`needs_review` 或 warnings 需要实际核验,程序运行成功不代表内容和版式通过。
- 不创建或猜测脚本清单以外的文件。 - HTTPS PDF 用 `download_pdf.py`,其他素材用 `download_attachment.py`。不在回复中复述敏感 URL 查询参数;加密 PDF 请用户提供已解密副本,密码不进入工具参数。
- 每次检查脚本返回的 JSON;只有 `ok` 为 `true` 时才继续。
- 收到 `ok: false` 时,依据 `error` 调整合法参数或向用户说明失败原因,不要把参数改传给其他脚本碰运气。 ## 按任务读取
- 阅读或总结时只使用 `pages[]` 中 `usable_for_summary: true` 的文本。`needs_ocr: false` 时不得为了“常规检查”继续 OCR、渲染或调用图片识别。
- 远程 PDF 源文件只交给 `download_pdf.py`;任务所需的远程图片、视频、音频、压缩包或其他附件只交给 `download_attachment.py`。不要在回复、日志摘要或文件名中复述可能含敏感查询参数的完整 URL。 | 任务 | 指南 |
- PDF 最终文件一律写入 `/usr/local/src/pdf/`,下载缓存、中间文件和渲染结果一律写入 `/usr/local/src/pdf/tmp/<任务名>/`。始终传绝对路径;固定脚本会自动创建目录并拒绝该根目录之外的输出。 | --- | --- |
- 不把 PDF 密码作为脚本参数;工具调用参数可能进入运行日志。 | 阅读、总结、扫描页、图片文字、图表辅助识别 | [读取与 OCR](references/reading-and-ocr.md) |
| 创建报告、提案、简历、学术或品牌 PDF | [设计规范](references/design.md) + [HTML 创建接口](references/creation.md) |
## 脚本清单 | Office/LaTeX 导出,PDF 内容重建为 Office | [转换](references/conversion.md) |
| 表单填写、裁剪、嵌入图片、元数据 | [编辑](references/editing.md) |
| 脚本 | 用途 | 底层能力 | | 下载、分段提取、表格、简单文本 PDF、合并/拆分/旋转、渲染与清理 | [原有操作接口](references/operations.md) |
| --- | --- | --- | | 镜像缺包或能力边界 | [依赖说明](references/dependencies.md) |
| `scripts/download_pdf.py` | 下载并校验远程 HTTPS PDF | `urllib`、`pypdf` |
| `scripts/download_attachment.py` | 下载图片、音视频、压缩包等通用 HTTPS 附件 | `urllib`、HEAD 大小探测、流式硬限制 | ## 固定脚本
| `scripts/inspect_pdf.py` | 检查页数、加密、元数据、页面尺寸和表单数量 | `pypdf` |
| `scripts/extract_text.py` | 多引擎提取、质量检测并分段返回正文 | Poppler `pdftotext`、`pdfplumber`;`pypdf` 校验 | | 脚本 | 用途 |
| `scripts/ocr_text.py` | 对指定扫描页执行离线 OCR 并返回可靠文字 | RapidOCR、ONNX Runtime、Poppler `pdftoppm` | | --- | --- |
| `scripts/extract_tables.py` | 按页提取表格 | `pdfplumber` | | `scripts/download_pdf.py` | 安全下载并校验不超过 25 MiB 的 HTTPS PDF |
| `scripts/render_pdf.py` | 把指定页面渲染为 PNG | Poppler `pdftoppm` | | `scripts/download_attachment.py` | 下载不超过 25 MiB 的任务附件 |
| `scripts/create_pdf.py` | 从 UTF-8 文本或 Markdown 创建 PDF | `reportlab`、`pypdf` | | `scripts/inspect_pdf.py` | 页数、尺寸、加密、元数据与表单数量 |
| `scripts/manage_pdf.py` | 合并、拆分或旋转 PDF | `pypdf` | | `scripts/extract_text.py` | 多引擎正文提取、逐页质量检测和字符游标 |
| `scripts/cleanup_pdf_temp.py` | 安全删除本次任务临时目录 | Python 文件 API | | `scripts/ocr_text.py` | 对指定 PDF 页执行离线 OCR,标记局部疑难区域 |
| `scripts/ocr_image.py` | 对单张任务图片执行离线 OCR |
环境已预置所有依赖。不要安装依赖,也不要提示用户安装依赖。 | `scripts/extract_tables.py` | 原生 PDF 表格分页提取 |
| `scripts/render_pdf.py` | 按页输出 PNG,供版式检查或疑难辅助复核 |
## 标准流程 | `scripts/create_pdf.py` | 简单文本/Markdown 生成 PDF |
| `scripts/create_design_pdf.py` | 静态 HTML/CSS、图表和公式设计排版 |
1. 为任务选择简短目录名,把中间文件放在 `/usr/local/src/pdf/tmp/<任务名>/`。 | `scripts/convert_to_pdf.py` | Office 文件导出 PDF |
2. 远程 HTTPS 链接先调用 `download_pdf.py`;本地文件直接进入下一步。 | `scripts/compile_latex.py` | 仅用镜像缓存资源编译 LaTeX |
3. 调用 `inspect_pdf.py` 检查文件。遇到加密 PDF 时停止处理,请用户提供已解密副本;当前固定脚本不接收密码。 | `scripts/edit_pdf.py` | 表单、裁剪、元数据和嵌入图片 |
4. 阅读或总结时调用 `extract_text.py`。结果为 `usable_for_summary: true` 时使用可靠页文本并根据游标继续;同时为 `needs_ocr: false` 时直接回答,不调用 OCR、渲染或图片识别。 | `scripts/manage_pdf.py` | 合并、拆分、旋转 |
5. 只有 `extract_text.py` 返回 `needs_ocr: true` 时,才对 `text_quality.suspect_pages` 调用 `ocr_text.py`。原生可靠文本优先,OCR 只补齐可疑页,不重复识别正常页。 | `scripts/cleanup_pdf_temp.py` | 清理本任务临时目录 |
6. `ocr_text.py` 会在脚本内部临时渲染指定页面并交给本地 RapidOCR,完成后自动删除 PNG;普通扫描件解析不调用大模型识图,也不需要先调用 `render_pdf.py`。
7. 仅在用户明确要求检查视觉版式,或任务涉及创建/修改 PDF 时调用 `render_pdf.py`。 ## 工作原则
8. 创建或修改后的最终 PDF 写入 `/usr/local/src/pdf/`,重新执行检查、文本提取和全部页面渲染。
9. 最终产物位于临时目录之外且不再需要缓存时,调用 `cleanup_pdf_temp.py` 清理本次任务目录。 1. 阅读先检查 PDF,再提取可靠原生文本;扫描页或图片文字以本地 OCR 为主。只有低置信度、手写、复杂表格/公式、阅读顺序冲突或非文本图形理解需要时,才用大模型复核相关页/区域。不要把整个扫描件直接交给大模型代替 OCR。
2. 创建设计先确认读者、用途、内容与输出限制,按需选择封面、配色和字体层级。用户模板、品牌、大纲、语言与篇幅优先;不强加独立封面,不为凑页数填充或删除内容。
## 下载远程 PDF 3. 简单文字选 `create_pdf.py`;需要封面、图文、页眉页脚、交叉引用、数学公式时选 `create_design_pdf.py`。通过 `write_file` 写静态内容文件,不写可执行代码。HTML 禁止脚本、事件处理程序和外部资源;公式、流程图由固定引擎本地处理。
4. 编辑现有 PDF 保留内容与结构。裁剪不等于脱敏;表单字段值写入不等于外观正确;PDF 转 Office 应按提取/OCR 后重建来规划,不能承诺无损逆转换。
只接受 HTTPS 地址。完整保留 URL 及查询参数,不在回复、日志摘要或文件名中复述敏感参数。 5. 创建、转换或修改后重新检查页数、文本与关键数字,再渲染全部相关页逐页核验封面、字体、表格、公式、图表、页码、裁切和空白页。要求精确页数时使用 `--expected-pages`。修正后检查最新产物,才交付。
6. 内容引用可核验。用户提供的材料可直接引用;新增时效、专业或不确定事实使用当前可用搜索工具查证,不编造统计、论文或参考文献。图片识别的猜测与原文分开标记。
调用 `scripts/download_pdf.py`:
```text
--url 'https://example.com/document.pdf' --output '/usr/local/src/pdf/tmp/<任务名>/source.pdf'
```
可选参数:
- `--timeout <秒>`:默认 `60`。
- `--max-bytes <字节数>`:默认且最高 `26214400`(25 MiB),只允许设置更小的限制。
- `--overwrite`:仅在目标是本次任务生成的缓存时使用。
脚本会创建父目录、流式下载、阻止 HTTPS 重定向降级到 HTTP,并验证 PDF。成功结果包含 `path`、`size_bytes`、`page_count` 和 `encrypted`。
## 下载通用附件
需要下载作为 PDF 任务素材的图片、视频、音频、压缩包或其他文件时,调用 `scripts/download_attachment.py`:
```text
--url 'https://example.com/asset.bin?signature=...' --output '/usr/local/src/pdf/tmp/<任务名>/asset.bin'
```
只接受 HTTPS 地址,`output` 可使用任意附件扩展名。可选参数只有 `--timeout <1-600>`(默认 `60`)和 `--overwrite`。附件上限固定为 25 MiB(26214400 字节),不可调高:脚本先用 HEAD 探测远端声明大小,再检查 GET 响应声明,并在流式接收时持续兜底计数;任一阶段发现超限都会返回 `ok: false` 和明确的“已拒绝下载”错误,且不会发布部分文件。
成功结果包含 `path`、实际 `size_bytes`、`declared_size_bytes`、`size_limit_bytes`、`size_probe` 和 `content_type`。本脚本不校验文件业务格式;远程 PDF 源文件仍使用 `download_pdf.py`。
## 检查 PDF
调用 `scripts/inspect_pdf.py`:
```text
--input '/usr/local/src/pdf/tmp/<任务名>/source.pdf'
```
使用返回的 `page_count`、`encrypted`、`metadata`、`page_layouts` 和 `form_field_count` 判断后续处理方式。不要直接调用 `pdfinfo`。
## 提取正文
首次调用 `scripts/extract_text.py`:
```text
--input '/usr/local/src/pdf/tmp/<任务名>/source.pdf'
```
默认使用 `auto` 引擎:先由 Poppler `pdftotext` 提取;结果不可用或命令不可用时自动尝试 `pdfplumber`,并可逐页选择质量更好的结果。脚本使用 `pypdf` 获取标准页数,并拒绝把页数不一致的提取结果当作成功。不要直接执行 `pdftotext`。
默认单次最多处理 8 页、返回 24000 个字符。可使用:
- `--start-page <页码>`、`--end-page <页码>`:页码从 `1` 开始。
- `--start-offset <字符偏移>`:继续读取被字符上限截断的同一页;大于 `0` 时同时传入上次返回的 `next_engine`。
- `--max-pages <页数>`、`--max-chars <字符数>`:控制单次输出。
- `--layout`:仅在需要尽量保留版面空格时使用。
- `--engine <auto|poppler|pdfplumber>`:首次及跨页提取保持 `auto`;同页字符续读时传入上次返回的 `next_engine`。
- `--timeout <秒>`:Poppler 提取超时,默认 `120`。
先检查 `usable_for_summary` 和 `text_quality.status`:
- `usable_for_summary: true`:只使用 `pages[]` 中同样标为 `usable_for_summary: true` 的 `text`;可疑页的文本会被置空。如果 `has_more: true`,始终传回 `next_page` 和 `next_offset`。仅当 `next_offset` 大于 `0` 时,把非空的 `next_engine` 传给 `--engine` 以固定同页字符游标;这种调用只续读当前页。当前页完成后返回的 `next_offset` 为 `0`,此时不要传 `--engine`,让下一页重新使用 `auto`。保留首次调用的 `--end-page`(如果指定)及其他选项,直至 `has_more: false`。
- `usable_for_summary: false`:本批次没有可靠文本,不要使用返回内容。查看 `engine_attempts`、`text_quality.reasons`、`text_quality.suspect_pages` 和 `needs_ocr`;若 `has_more: true`,仍按跨页游标继续检查后续批次,避免漏掉后续可搜索文本。
- `needs_ocr: true`:一个或多个页面未得到可靠文本。把 `text_quality.suspect_pages` 中实际需要阅读的页码传给 `ocr_text.py`;不要先调用 `render_pdf.py`,也不要把临时图片交给大模型。
`complete_text_coverage: true` 表示本批次所有页面均有可靠文本。`text_quality` 按页检测空白或过少文本、页面实际可见图像覆盖过大但文字不足、`(cid:...)`、Unicode 替换字符、异常控制字符及外观像汉字的部首字符;`pages[].extractor` 表示该页最终采用的引擎。`status: mixed` 表示同一批次同时包含可靠页和可疑页:可先使用可靠页文本,同时只核验 `suspect_pages`。不要只根据“肉眼看起来能读”判定提取结果可靠。
## 本地 OCR 扫描页
仅当 `extract_text.py` 返回 `needs_ocr: true` 时调用 `scripts/ocr_text.py`。`--pages` 必须明确指定 `text_quality.suspect_pages` 中要读取的页,单次最多 4 页:
```text
--input '/usr/local/src/pdf/tmp/<任务名>/source.pdf' --pages '2,5-6'
```
默认以 260 DPI 临时渲染,并使用镜像中预置的 RapidOCR 与 ONNX Runtime 在本地识别。脚本不会联网下载模型,不会保留渲染图片,也不会调用大模型视觉能力。可选参数:
- `--dpi <150-400>`:文字过小或识别质量不足时适度提高,默认 `260`。
- `--max-chars <字符数>`:默认 `24000`,最大 `60000`。
- `--timeout <秒>`:每页 Poppler 渲染超时,默认 `180`。
- `--start-offset <字符偏移>`:续读被字符上限截断的单页;使用时 `--pages` 只能包含该页。
只使用 `pages[]` 中 `usable_for_summary: true` 的 `text`。`status: empty`、`sparse` 或 `low_confidence` 的页面文本会被置空,并通过 `needs_review: true` 提醒人工检查。
如果 `has_more: true`:
- `next_offset > 0`:用 `--pages <next_page> --start-offset <next_offset>` 续读同一页。
- `next_offset = 0`:用返回的 `remaining_pages` 继续下一批。
- 同页续读完成后,再处理先前返回的其他 `remaining_pages`。
OCR 结果中的 `mean_confidence`、`line_count`、`render_seconds` 和 `ocr_seconds` 仅用于判断质量与性能。原生提取成功的页面始终采用 `extract_text.py` 结果,不用 OCR 覆盖。
## 提取表格
调用 `scripts/extract_tables.py`:
```text
--input '/usr/local/src/pdf/tmp/<任务名>/source.pdf' --start-page 1
```
默认单次最多处理 5 页、20 个表格和 2000 个单元格。可用 `--end-page`、`--start-table`、`--max-pages`、`--max-tables`、`--max-cells` 调整。若 `has_more: true`,把 `next_page` 传给 `--start-page`、`next_table` 传给 `--start-table` 后继续,并保留首次调用的 `--end-page`(如果指定)及其他提取选项。
## 渲染页面
只有满足以下任一条件时才调用 `scripts/render_pdf.py`:
- 用户明确要求审阅版式、图表、印章、公式或页面外观;
- 创建或修改 PDF 后进行最终视觉检查。
不要因为输入是 PDF、需要总结、需要 OCR 或需要检查首页就自动调用本脚本;OCR 的临时渲染由 `ocr_text.py` 内部完成。调用脚本时不要直接执行 `pdftoppm`:
```text
--input '/usr/local/src/pdf/tmp/<任务名>/source.pdf' --output-dir '/usr/local/src/pdf/tmp/<任务名>/rendered' --start-page 1
```
默认 150 DPI、单次最多 10 页。可使用 `--end-page`、`--max-pages`、`--dpi`、`--timeout` 和 `--overwrite`。若 `has_more: true`,使用 `next_page` 继续,并保留首次调用的 `--end-page`(如果指定)、输出目录及其他渲染选项。脚本返回标准化的 `page-0001.png` 文件路径。
对文字较小或图表密集的页面提高 DPI。使用可用的图像查看工具检查返回的 PNG,不要尝试把图片路径交给下载脚本。
## 创建 PDF
先使用 `write_file` 把内容写为 UTF-8 `.txt` 或 `.md` 文件,再调用 `scripts/create_pdf.py`:
```text
--input '/usr/local/src/pdf/tmp/<任务名>/content.md' --output '/usr/local/src/pdf/<文件名>.pdf' --title '文档标题'
```
脚本支持 Markdown 标题、项目符号和简单表格,自动选择可嵌入的 Unicode 字体并添加页码。可选参数:
- `--page-size <A4|LETTER>`
- `--font-path <TTF或TTC路径>`
- `--font-size <字号>`
- `--margin <points>`
- `--overwrite`
输入内容只使用 ASCII 连字符 `-`;脚本也会把常见 Unicode 横线规范化为 ASCII 连字符。
## 合并、拆分与旋转
调用 `scripts/manage_pdf.py`,第一个参数必须是操作名。
合并:
```text
merge --input 'a.pdf' --input 'b.pdf' --output '/usr/local/src/pdf/merged.pdf'
```
拆分指定范围:
```text
split --input 'source.pdf' --output-dir '/usr/local/src/pdf/split' --range 1-3 --range 4-6
```
不传 `--range` 时每页生成一个 PDF。
旋转指定页面:
```text
rotate --input 'source.pdf' --output '/usr/local/src/pdf/rotated.pdf' --pages '1,3-5' --degrees 90
```
`--degrees` 只能是 `90`、`180` 或 `270`;不传 `--pages` 时旋转全部页面。目标已存在且确认可覆盖时添加 `--overwrite`。
## 清理临时目录
调用 `scripts/cleanup_pdf_temp.py`:
```text
--task-dir '/usr/local/src/pdf/tmp/<任务名>'
```
脚本只允许删除 `/usr/local/src/pdf/tmp/` 下一级任务目录,拒绝删除根目录、仓库目录或其他路径。
## 质量要求
- 不覆盖用户提供的源文件。
- 创建或修改后重新检查页数、页面尺寸、加密状态和文本可读性。
- 扫描件先做原生文字检测,再只 OCR 可疑页;不得把低置信度 OCR 文本当作可靠正文。
- 逐页确认没有裁切、重叠、溢出、乱码、黑方块、错误分页或异常空白页。
- 检查标题层级、段落间距、页边距、表格、图表、图片、页码及章节衔接。
- 引用和参考文献必须可读,不得残留工具令牌、占位符或临时路径。
- 只有最新渲染结果不存在可见缺陷时才交付创建或修改后的 PDF。

View File

@ -1,4 +1,4 @@
interface: interface:
display_name: "PDF 处理" display_name: "PDF 读取与设计"
short_description: "读取、创建和审阅本地或远程 PDF,按需执行本地 OCR" short_description: "本地 OCR 优先识别,设计排版、转换与编辑 PDF,并完成逐页校验"
default_prompt: "使用 $pdf 下载或读取这个 PDF,优先提取可靠文本,并只对扫描页执行本地 OCR。" default_prompt: "使用 $pdf 读取或制作 PDF;扫描图像先做本地 OCR,疑难内容辅助复核,按内容设计版式并检查最终页面。"

View File

@ -0,0 +1,74 @@
/* Publication defaults. Author CSS overrides these tokens and styles. */
:root {
--page-width: 210mm; --page-height: 297mm;
--accent: #8a3a2a; --accent-light: #f5ece7; --cover-bg: #30231f; --cover-text: #faf5ef;
--ink: #202124; --muted: #62666a; --rule: #d8dadd;
--font-display: 'PingFang SC', 'Microsoft YaHei', 'Noto Sans CJK SC', Arial, sans-serif;
--font-body: 'SimSun', 'Noto Serif CJK SC', 'Times New Roman', serif;
--font-sans: 'PingFang SC', 'Microsoft YaHei', 'Noto Sans CJK SC', Arial, sans-serif;
}
@page { size: A4; margin: 24mm 24mm 22mm;
@top-left { content: string(chapter); font: 8pt var(--font-sans); color: #62666a; }
@bottom-right { content: counter(page); font: 8pt var(--font-sans); color: #62666a; }
}
@page cover { size: A4; margin: 0;
@top-left { content: none; } @bottom-right { content: none; }
}
* { box-sizing: border-box; }
html, body { margin: 0; padding: 0; }
body { color: var(--ink); font: 10.5pt/1.65 var(--font-body); overflow-wrap: break-word; }
h1, h2, h3 { font-family: var(--font-display); color: var(--accent); break-after: avoid; line-height: 1.3; }
h1 { font-size: 22pt; margin: 24pt 0 12pt; string-set: chapter content(text); }
h2 { font-size: 15pt; margin: 18pt 0 8pt; }
h3 { font-size: 11.5pt; margin: 12pt 0 6pt; }
p { margin: 0 0 8pt; orphans: 3; widows: 3; }
a { color: var(--accent); text-decoration: underline; }
strong { font-family: var(--font-sans); }
.section-start { break-before: page; }
.eyebrow { font: 8.5pt/1.4 var(--font-sans); letter-spacing: .1em; }
.lead { font-size: 13pt; line-height: 1.7; }
.cover { page: cover; break-after: page; width: var(--page-width); height: var(--page-height);
padding: 30mm 27mm; position: relative; display: flex; flex-direction: column; justify-content: center;
background: var(--cover-bg); color: var(--cover-text); font-family: var(--font-sans); }
.cover h1 { margin: 16pt 0; color: inherit; font: 700 40pt/1.2 var(--font-display); string-set: none; }
.cover .subtitle { font-size: 14pt; line-height: 1.6; max-width: 140mm; }
.cover .metadata { margin-top: 25mm; font-size: 10pt; line-height: 1.8; }
.cover .accent-line { width: 24mm; border-top: 3pt solid var(--accent); margin: 14pt 0; }
.cover-fullbleed::before { content: ''; position: absolute; top: 0; right: 20mm; width: 12mm; height: 28mm; background: var(--accent); }
.cover-split { background: #f7f5f1; color: var(--ink); padding-left: 101mm; padding-right: 18mm; }
.cover-split::before { content: ''; position: absolute; inset: 0 auto 0 0; width: 42%; background: var(--cover-bg); border-right: 2mm solid var(--accent); }
.cover-split h1 { font-size: 28pt; }
.cover-typographic, .cover-minimal, .cover-frame { background: #fafaf7; color: var(--ink); }
.cover-typographic h1 { font-size: 48pt; }
.cover-minimal { border-left: 3mm solid var(--accent); }
.cover-minimal h1 { font-weight: 400; }
.cover-frame::before { content: ''; position: absolute; inset: 11mm; border: 1pt solid var(--accent); pointer-events: none; }
.cover-frame { text-align: center; align-items: center; }
.cover-editorial::before { content: attr(data-mark); position: absolute; top: 12mm; right: 10mm; opacity: .07; font: 180pt/1 var(--font-display); }
.cover-editorial h1 { font-size: 48pt; }
figure { margin: 14pt 0; break-inside: avoid; }
img, svg { max-width: 100%; }
img { height: auto; }
figcaption, .caption { font: 8.5pt/1.5 var(--font-sans); color: var(--muted); margin-top: 6pt; }
.mermaid { text-align: center; margin: 12pt 0; break-inside: avoid; }
.mermaid svg { max-height: 180mm; }
.math-display { margin: 12pt 0; text-align: center; break-inside: avoid; }
.katex { font-size: 1.05em; }
table { width: 100%; border-collapse: collapse; margin: 12pt 0; font: 9.5pt/1.5 var(--font-sans); }
thead { display: table-header-group; }
tr { break-inside: avoid; }
th { text-align: left; background: var(--accent); color: white; border-top: 1.5pt solid var(--accent); }
th, td { padding: 7pt 8pt; overflow-wrap: anywhere; vertical-align: top; }
td { border-bottom: .5pt solid var(--rule); }
tbody tr:nth-child(even) { background: var(--accent-light); }
tbody tr:last-child td { border-bottom: 1.5pt solid var(--accent); }
.three-line th { background: transparent; color: var(--ink); border-top: 1.5pt solid var(--ink); border-bottom: .8pt solid var(--ink); }
.three-line td { border: 0; } .three-line tbody tr { background: transparent; }
.three-line tbody tr:last-child td { border-bottom: 1.5pt solid var(--ink); }
.numeric { text-align: right; font-variant-numeric: tabular-nums; }
blockquote, .callout, .theorem { margin: 12pt 0; padding: 8pt 12pt; border-left: 2pt solid var(--accent); background: var(--accent-light); }
pre { font: 9pt/1.5 'Courier New', monospace; padding: 10pt; background: #f5f5f5; white-space: pre-wrap; overflow-wrap: anywhere; }
ul, ol { padding-left: 20pt; } li { margin-bottom: 4pt; }
.references { font-size: 9pt; line-height: 1.6; }
.toc a { display: block; margin: 6pt 0; text-decoration: none; }
.toc a::after { content: ' ' target-counter(attr(href), page); float: right; }

View File

@ -0,0 +1,26 @@
<!doctype html>
<html lang="zh-CN">
<head><meta charset="UTF-8"><title>报告标题</title>
<style>
:root { --accent: #8a3a2a; --accent-light: #f5ece7; --cover-bg: #30231f; --cover-text: #faf5ef; }
</style></head>
<body>
<!-- 本文件是结构示例:根据实际内容替换文字,按需要保留封面、目录和章节。默认排版 CSS 由固定脚本注入。 -->
<section class="cover cover-fullbleed">
<p class="eyebrow">报告类型 · 年份</p>
<h1>报告标题</h1>
<div class="accent-line"></div>
<p class="subtitle">一句说明报告对象、范围和阅读目的的副标题。</p>
<p class="metadata">作者或机构<br>发布日期</p>
</section>
<section>
<h1 id="summary">核心结论</h1>
<p class="lead">用可核对的证据说明主要结论,并区分事实与判断。</p>
<h2 id="evidence">数据与依据</h2>
<table><thead><tr><th>指标</th><th>观察</th><th>来源</th></tr></thead>
<tbody><tr><td>示例指标</td><td>替换为真实数据或明确标注的示例</td><td>填写可验证来源</td></tr></tbody></table>
<div class="callout"><strong>适用范围</strong><p>说明时间范围、样本、计算口径与不确定性。</p></div>
<h2 id="actions">建议与下一步</h2>
<p>建议应能追溯到证据;需要决策的事项写清楚条件与影响。</p>
</section>
</body></html>

View File

@ -0,0 +1,37 @@
# Office、LaTeX 与 PDF 转换
只调用固定脚本。不能直接运行 LibreOffice、Tectonic、Bash、Python 或 Node,也不在任务中安装依赖。
## Office → PDF
Office 主文档的编辑、公式计算、图表和版式由 docx/xlsx/pptx 等对应 skill 完成。已有完成的源文件可调用 `scripts/convert_to_pdf.py`:
```text
--input '/usr/local/src/pdf/tmp/task/report.docx' --output '/usr/local/src/pdf/report.pdf'
```
支持 DOCX/DOC/ODT/RTF、PPTX/PPT/ODP、XLSX/XLS/ODS,最大 25 MiB。`--timeout` 默认 180 秒,1–600;可用 `--overwrite` 覆盖本任务旧产物。每次转换使用独立 LibreOffice profile 和临时副本,不修改源文件。
转换前确认 Excel 公式已重算、打印范围正确;核对字体替换、图表、分页和页数。转换可能有版式差异,必须渲染检查。CSV 和 HTML 分别先走 xlsx 或 HTML 设计接口,不能含糊地自动推断编码和布局。
## PDF → Office
PDF 是固定版面,不能承诺用 LibreOffice 直接得到结构完整的 Word、Excel 或 PPT。先用原生提取/OCR 得到可靠文本和表格,再由对应 skill 的固定写入接口重建;明确哪些结构可编辑、哪些需要保留图片。扫描件不会因为换扩展名就变成可编辑文本。
需要高保真还原时先确认重点是视觉一致还是编辑结构;保留原 PDF 对照。不得把每页截图贴入 Word 后声称正文可编辑。
## LaTeX → PDF
用户明确提供 LaTeX 模板或源文件时调用 `scripts/compile_latex.py`:
```text
--input '/usr/local/src/pdf/tmp/task/main.tex' --output '/usr/local/src/pdf/paper.pdf'
```
输入 `.tex` 最大 2 MiB,相关图片和被引用的 `.tex` 放在本任务目录;超时默认 180 秒,范围 1–600。固定脚本内部调用 Tectonic,禁用 shell escape,只使用镜像构建时缓存的 TeX 资源。缺失包、未缓存模板、编译失败或超时都明确报错,不能安装或偷偷切换到联网编译。
基础镜像预热 ctex、常用数学、表格、图片、几何和超链接包;特殊模板或包可能还需维护者补充镜像。Tectonic 自行处理常见重跑与引用,不把同一编译无意义重复多次。
保留用户模板的语言、字体、章节与引用规范。中文模板可使用 `ctexart` 和已缓存的 Fandol 字体;新的模板需确认所用字体实际存在。若只有少量数学公式而没有 LaTeX 模板,HTML + KaTeX 即可,无需编写完整 LaTeX 文档。
返回页数、警告和视觉复核标记。检查 undefined references、缺字、Overfull 等具体问题;输出成功后仍需提取与渲染检查,不把“有 PDF 文件”作为通过标准。

View File

@ -0,0 +1,48 @@
# HTML/CSS 设计排版接口
先读 [设计规范](design.md)。简单文本使用 `create_pdf.py`;图文报告、封面、页眉页脚、目录、公式和流程图使用 `create_design_pdf.py`。
## 调用
通过 `write_file` 将内容写为 UTF-8 `.html`,以 `<!doctype html>` 声明标准模式,并包含 `<meta charset="utf-8">`。通过 `execute_skill_script` 调用 `scripts/create_design_pdf.py`,不直接调用 Node 或 Chromium:
```text
--input '/usr/local/src/pdf/tmp/task/report.html' --output '/usr/local/src/pdf/report.pdf'
```
可选 `--css <本地CSS>`、`--page-size A4|LETTER`、`--expected-pages <1–200>`、`--timeout <10–600>`(默认 180)和 `--overwrite`。HTML/CSS 单文件不超过 2 MiB。先读取 `assets/report.html` 了解结构,按实际内容编写;模板中的示例文字不可直接交付。
`assets/design.css` 由固定渲染器自动注入,提供变量、标题、六种封面、三线表、代码、引用、图注与目录。用户 HTML 内的 CSS 可覆盖默认值,`--css` 最后应用。不需要手动加载该样式或任何 JS 库。
## 静态内容与本地资源
- 只写静态 HTML/CSS/SVG。禁止 `<script>`、事件属性、iframe/object/embed、`javascript:` URL 和运行任意 JS。不能提供 Python、Bash 或 Node 代码让脚本代为执行。
- 图片、CSS 和自备字体放在 HTML 同目录及其子目录,用相对路径引用;上级目录、符号链接越界、`file:` 和外部资源会被拒绝。资源单个不超过 25 MiB。
- HTTP(S) 引用链接可以保留为可点击参考来源,但不能用作图片、CSS、字体或脚本下载入口。所需远程素材先由 `download_attachment.py` 下载。
- Paged.js、KaTeX、Mermaid 和 KaTeX 字体从镜像读取,没有 CDN 例外。渲染器使用独立浏览器上下文并限制资源访问。
- 使用语义化 `h1/h2/p/table/figure/figcaption`,对引用设置真实 `id` 和 `href="#..."`。自定义计数器跨页需核验;普通页码、目录目标页码可使用 Paged.js 支持的 counter/target-counter,不笼统禁止 CSS counter。
## 公式与流程图
不加载脚本、不写初始化代码。固定接口识别以下结构:
```html
<p>公式为 <span class="math-inline">E=mc^2</span>。</p>
<div class="math-display">\sum_{i=1}^{n} x_i</div>
<div class="mermaid">flowchart LR
A[资料] --> B[本地 OCR]
B --> C[疑难复核]
C --> D[PDF]</div>
```
公式内容为 KaTeX 数学标记,Mermaid 为图形描述语言,不是任意程序入口。Mermaid 使用 neutral 主题与 strict 安全模式;不支持在图中插入配置指令。公式/图形解析失败则返回错误,不带着未渲染内容输出成功。
数据图表可用静态 SVG 或已生成图片。长公式应分行;大表/大图应调整结构、命名横向页或拆分,不能一律缩成不可读的小图。
## 分页与检查
- 默认 A4 内页与全页封面都使用实际物理尺寸;不采用原脚本固定 `scale: 1.5` 的补偿。
- 先等待图片、字体和图形完成,再等待 Paged.js 的完成 Promise。超时、外部资源、图片失败、横向溢出或页数不符时不发布结果。
- PDF 实际页数必须与分页引擎相同。`--expected-pages` 不匹配时明确失败,不裁掉页面或缩短内容冒充满足要求。
- 返回 `page_count`、每页文字/图形统计、warnings 和 `requires_visual_review`。空白页提示结合上下文判断;少字的图表页不直接判坏。
- 用 `inspect_pdf.py` 和 `extract_text.py` 核对最终文件,再用 `render_pdf.py` 查看全部页。自动检测不能发现所有纵向裁切、语义错误、缺字或表单外观问题。

View File

@ -0,0 +1,40 @@
# 容器依赖与能力边界
依赖由 `silk-base/Dockerfile` 预装。机器人只调用固定脚本,不能运行安装命令;需要更新依赖时由维护者重新构建和部署基础镜像。
## 复用原 PDF 工具链
| 能力 | 镜像依赖 |
| --- | --- |
| 原生文字、表格、页面和表单 | pdfplumber 0.11.9、pypdf 6.10.0、Poppler |
| 简单文本/Markdown PDF | ReportLab 4.4.9、Pillow、已安装中文与拉丁字体 |
| PDF 页及单张图片 OCR | RapidOCR 3.9.1、ONNX Runtime 1.27.0、OpenCV 4.12.0.88、OmegaConf 2.3.1 |
| Office 导出 | LibreOffice Writer/Calc/Impress 与系统字体 |
| HTML 渲染运行时 | 已有 Node.js 与系统 Chromium,`CHROME_BIN=/usr/bin/chromium` |
OCR 使用 RapidOCR 包内的检测、方向分类与识别 ONNX 模型。缺少模型就返回错误;不回退联网下载。OmegaConf 固定兼容版本,避免旧版不能处理 RapidOCR 的路径配置。
## 本次补充
| 依赖 | 用途 | 安装位置/方式 |
| --- | --- | --- |
| poppler-data | Adobe CJK 字符映射,补齐部分中文 PDF 的渲染与提取 | Debian 软件包,显式安装并检查 GB1 映射文件 |
| playwright-core 1.63.0 | 固定浏览器渲染桥接 | 全局 Node 包,复用系统 Chromium |
| Paged.js 0.4.3 | CSS 分页、命名页、页眉页脚、目录目标页码 | 全局 Node 包,本地加载 |
| KaTeX 0.18.7 | 行内/展示公式及其字体 | 全局 Node 包,本地加载 |
| Mermaid 11.17.2 | 流程、关系等示意图 | 全局 Node 包,固定 strict 模式 |
| Tectonic 0.17.0 | 用户提供的 LaTeX 源文件编译 | amd64/arm64 官方静态发行包,逐架构 SHA-256 校验 |
不需要额外引入 pikepdf、另一套 Python 浏览器、Matplotlib 或 LaTeX 完整发行版。页面/表单/元数据/图像继续复用 pypdf;数据图用静态 SVG、图片或现有 xlsx 图表。数学公式优先 KaTeX,有 LaTeX 模板时才用 Tectonic。
`poppler-data` 提供编码映射,不能用字体包替代。没有映射时即使 PDF 带有中文字体,部分文件仍可能预览缺字;遇到 `Missing language pack` 应更新镜像,不能把有缺字的预览用于 OCR 或当作空白原文。[Debian 包说明](https://packages.debian.org/trixie/poppler-data)
## 构建与离线运行
- 浏览器只从任务目录、skill 素材和固定本地包读取资源;不使用 CDN、不自动下载 Chromium。`NODE_PATH` 保留镜像的全局包目录。
- 构建时实际运行 HTML + KaTeX + Mermaid + Paged.js 自检,不能只验证包能 import。
- `TECTONIC_CACHE_DIR=/opt/tectonic-cache` 在构建时预热 ctex/Fandol、amsmath/amssymb、booktabs/longtable、graphicx、geometry、hyperref,随后验证仅缓存编译和中文文字提取。首次构建需要网络下载 TeX 资源;任务编译始终禁用 shell escape 且只读缓存包。
- 特殊模板、额外文献工具或未预热的 TeX 包可能仍不可用。维护者把真实模板所需资源加入构建预热,再部署镜像;不能在机器人任务里解除缓存限制。
- 原有宋体、黑体、仿宋、楷体、方正小标宋、微软雅黑、苹方 SC、SF Pro、Noto 等字体继续复用。渲染后仍需检查所选字体和字形,安装完成不代表所有模板无差异。
修改 Dockerfile 后必须重新构建并部署,已运行的旧容器不会自动获得新依赖。

View File

@ -0,0 +1,67 @@
# PDF 视觉设计规范
先确定文档用途、读者、交付媒介、品牌和内容结构,再决定视觉。继承用户模板时先保持其规范;阅读、裁剪或填表不触发整本重新设计。
## 视觉方向
选择一套与内容相符的排版语言,贯穿封面、标题、表格和页眉。下面是起点,不是行业强制配色:
| 内容与气质 | 主色 | 浅色 | 深色封面或文字 |
| --- | --- | --- | --- |
| 科学、工程、数据分析 | `#2D5F8A` | `#E8F0F8` | `#152C3E` |
| 商业、策略、运营 | `#8A3A2A` | `#F5ECE7` | `#30231F` |
| 医疗、生态、公共服务 | `#2A6B5A` | `#EBF3EF` | `#19382F` |
| 创意、文化、作品集 | `#6B2A35` | `#F5ECEE` | `#29171D` |
| 学术、法律、正式材料 | `#3D4C5E` | `#F1F3F5` | `#202124` |
通常一个主色、深色正文和少量浅色层次就够。强调色用于标题、线条和关键数据;正文保持深色。颜色须有含义,不能只依赖颜色区分类别。文字与背景保持清晰对比,小字尽量达到 7:1;不要把装饰图形的对比度要求与正文混为一谈。
不要把渐变、卡片或深色背景机械地当成专业设计,也不要因旧规范的偏好一概禁止用户明确要求的颜色。内页以打印友好、清晰的信息层级为主;大面积深色更适合少量封面或章节页。
## 字体与尺度
镜像已提供 Noto CJK、宋体、黑体、仿宋、楷体、方正小标宋、微软雅黑、苹方 SC、SF Pro、Arial、Times New Roman 等。选择通常不超过两套视觉字体系统;中文与拉丁字体的必要回退不算额外装饰字体。不要假定宿主机所有字体都存在容器中,也不要从字体 CDN 加载。
| 风格 | 中文正文/标题 | 英文与数字 |
| --- | --- | --- |
| 现代报告 | 苹方 SC / 微软雅黑 / Noto Sans CJK SC | SF Pro / Arial |
| 学术长文 | 宋体 / Noto Serif CJK SC,标题可搭黑体 | Times New Roman |
| 正式公文 | 依用户规范选仿宋、楷体、方正小标宋 | 依模板 |
字号起点:封面 32–48 pt,一级标题 20–24 pt,二级 14–16 pt,三级 11.5–12 pt,正文 10.5–12 pt,图注 8.5–9.5 pt,页眉页脚 8–9 pt。中文正文行高 1.6–1.8;正文过长时先调整结构与分页,不用极小字号硬塞。
A4 内页通常上下 22–28 mm、左右 22–28 mm。段后约 8 pt,章节前约 24–26 pt;选定间距节奏后保持一致。简历可用 15–18 mm 边距,仍需保证可读性和完整内容。
## 六种可复用封面
`assets/design.css` 提供对应 class,`assets/report.html` 提供结构起点。独立封面按需使用;简历、短备忘录、用户指定页数紧张或已有模板时,标题区即可。
| class | 设计方式 | 常见用途 |
| --- | --- | --- |
| `cover-fullbleed` | 整页深底、大标题、短色带、作者日期 | 年报、主题报告 |
| `cover-split` | 42% 色块与 58% 浅色内容区,清晰分割 | 提案、方案 |
| `cover-typographic` | 浅底、展示字体、尺度对比 | 作品、专题材料 |
| `cover-minimal` | 竖线、轻量标题、充分留白 | 简洁文档 |
| `cover-frame` | 细框、居中结构、克制装饰 | 正式报告 |
| `cover-editorial` | 大字背景、强标题、编辑式构图 | 杂志、创意内容 |
封面用命名页 `@page cover`,整页尺寸与纸张一致,`margin: 0`,内容区采用 `border-box`。不要在有页边距的内容区再套 `100vh` 或完整 A4 高度;这会产生裁切和额外空白页。图案用静态 SVG/几何元素;不能让纹理干扰标题。
## 内页
- 标题层级用字号、字重、间距和少量细线区分;标题不能孤立在页底。长章节按语义分段,避免把整节设为不可分页。
- 正文不堆砌仪表盘卡片。报告可用少量重点引言或行动框;学术内容可用定义/定理边线。别用带阴影的网页组件代替正文结构。
- 表格有真实表头、单位和来源。数字右对齐,小数位一致;长表允许跨页并重复表头,避免整张长表 `break-inside: avoid`。学术三线表用 `.three-line`;业务表可用主色表头和浅色交替行。
- 图表比例服从信息:柱线图常用横向,流程图可纵向,页面容纳不下时拆分、调整方向或用横向命名页。不要强制所有图表横向,也不要通过 `overflow:hidden` 掩盖溢出。
- 图表必须有可辨认的标签、单位、图例与图注;图表文字按最终 PDF 尺寸检查。数值须来自实际数据,不能用装饰图冒充统计结果。
- 图片保持比例,注明图注;数据图优先原有 xlsx 图表、静态 SVG 或已提供图片。不生成任意 Matplotlib/Python/Node 代码作为运行入口。
- 代码示例用浅灰底、等宽字体与换行;引用框用细左边线;数学公式用 KaTeX。需要完整 LaTeX 模板时走编译接口。
- 页眉使用章节名称,页脚保持统一页码。目录、图表引用和参考文献使用真实锚点,可点击;分页后核对链接落点。
## 内容与交付
用户的大纲、语言、数字口径和篇幅优先。精确页数通过实际生成与 `--expected-pages` 核对;字数要求按用户范围执行,不默认放宽 20%,不为填页伪造内容。没有必要时不添加封面、目录或参考文献页。
已有材料里的数据注明来源;新增研究、统计、政策等需验证,来源不足就披露不确定性。参考文献格式依用户/机构要求,中文 GB/T 7714 或英文 APA 可作为选项;不编造作者、年份和出处。
逐页核验:文字无缺字/黑块,字号清晰;页面尺寸与边距一致;图文无裁切、重叠;表格跨页合理;公式和引用正确;封面与正文衔接自然;不存在非预期空白页。自动溢出检测只是辅助,不替代实际看图。

View File

@ -0,0 +1,54 @@
# 表单、裁剪、元数据和嵌入图片
通过 `execute_skill_script` 调用 `scripts/edit_pdf.py`。所有输出遵循 `/usr/local/src/pdf/`;不覆盖源文件。原有合并、拆分、旋转仍用 `manage_pdf.py`。
## 表单
先查看真实字段:
```text
form-info --input '/usr/local/src/pdf/tmp/task/form.pdf'
```
返回字段 id、类型、当前值、只读状态、选择项和按钮状态。默认 50 项,`--offset`/`--limit`(最高 200)支持继续。没有 AcroForm 字段的扫描表单不能直接套用字段填写;XFA 和数字签名不是本接口支持的填写类型。
填写明确要求的字段:
```text
form-fill --input '/usr/local/src/pdf/tmp/task/form.pdf' --output '/usr/local/src/pdf/filled.pdf' --data '{"name":"Alice","agree":true,"country":"CN"}'
```
较长数据用 `--data-file <JSON>` 替代 `--data`,二选一,上限 2 MiB、500 个字段。
- 文本值必须是字符串,遵守字段长度。复选框可用 JSON `true/false`,或实际状态名如 `/Yes`;不能把非空字符串一律当作选中。
- 单选按钮使用实际状态值;下拉/列表使用真实选项值,多选需原字段支持。无效值、未知字段、只读字段或签名字段明确失败。
- pypdf 更新字段值与外观,再回读校验。必须渲染核对文字位置、换行、复选框和单选按钮状态。原表单字体未包含中文等字形时,不能仅凭字段值正确就交付;需保留可显示的模板或按用户要求重建表单版式。
- 不把填写文本当作签署数字签名,也不主动提交表单。
## 裁剪
```text
crop --input '/usr/local/src/pdf/tmp/task/source.pdf' --output '/usr/local/src/pdf/cropped.pdf' --pages '1-2' --box '20,30,575,812'
```
坐标单位 pt,按未旋转 PDF 坐标系的左、下、右、上填写;框必须在页面 MediaBox 内。省略 `--pages` 处理全部页。修改 CropBox,保留页面内容。**裁剪不删除不可见内容,不是脱敏。** 敏感信息移除需专门的真正删改流程,不能用白块或裁剪冒充。
## 元数据
读取用 `inspect_pdf.py`。更新明确字段:
```text
metadata --input '/usr/local/src/pdf/tmp/task/source.pdf' --output '/usr/local/src/pdf/updated.pdf' --data '{"Title":"报告","Author":"机构"}'
```
支持 Title/Author/Subject/Keywords/Creator/Producer;保留未指定字段。该接口修改文档信息字典,现有 XMP 保留且通过 `xmp_preserved` 提示,不能声称已清理所有元数据。
## 嵌入图片
```text
extract-images --input '/usr/local/src/pdf/tmp/task/source.pdf' --pages '2-3' --output-dir '/usr/local/src/pdf/tmp/task/images'
```
一次最多 4 页、默认 20 张图片(`--max-images` 最高 50),本批总大小不超过 25 MiB。用 `remaining_pages` 和 `--start-image <next_image>` 续读。需要覆盖本任务旧输出时加 `--overwrite`。
此操作提取 PDF 中的图像对象,不代表完整页面:不包含周围文字、矢量图形,也可能是被裁剪/复用的图片。阅读扫描页和视觉复核应使用 OCR 内部渲染或 `render_pdf.py`,不能把提取出的零散图片当成原 PDF 页。

View File

@ -0,0 +1,164 @@
# 原有固定操作接口
## 下载远程 PDF
只接受 HTTPS 地址。完整保留 URL 及查询参数,不在回复、日志摘要或文件名中复述敏感参数。
调用 `scripts/download_pdf.py`:
```text
--url 'https://example.com/document.pdf' --output '/usr/local/src/pdf/tmp/<任务名>/source.pdf'
```
可选参数:
- `--timeout <秒>`:默认 `60`。
- `--max-bytes <字节数>`:默认且最高 `26214400`(25 MiB),只允许设置更小的限制。
- `--overwrite`:仅在目标是本次任务生成的缓存时使用。
脚本会创建父目录、流式下载、阻止 HTTPS 重定向降级到 HTTP,并验证 PDF。成功结果包含 `path`、`size_bytes`、`page_count` 和 `encrypted`。
## 下载通用附件
需要下载作为 PDF 任务素材的图片、视频、音频、压缩包或其他文件时,调用 `scripts/download_attachment.py`:
```text
--url 'https://example.com/asset.bin?signature=...' --output '/usr/local/src/pdf/tmp/<任务名>/asset.bin'
```
只接受 HTTPS 地址,`output` 可使用任意附件扩展名。可选参数只有 `--timeout <1-600>`(默认 `60`)和 `--overwrite`。附件上限固定为 25 MiB(26214400 字节),不可调高:脚本先用 HEAD 探测远端声明大小,再检查 GET 响应声明,并在流式接收时持续兜底计数;任一阶段发现超限都会返回 `ok: false` 和明确的“已拒绝下载”错误,且不会发布部分文件。
成功结果包含 `path`、实际 `size_bytes`、`declared_size_bytes`、`size_limit_bytes`、`size_probe` 和 `content_type`。本脚本不校验文件业务格式;远程 PDF 源文件仍使用 `download_pdf.py`。
## 检查 PDF
调用 `scripts/inspect_pdf.py`:
```text
--input '/usr/local/src/pdf/tmp/<任务名>/source.pdf'
```
使用返回的 `page_count`、`encrypted`、`metadata`、`page_layouts` 和 `form_field_count` 判断后续处理方式。不要直接调用 `pdfinfo`。
## 提取正文
首次调用 `scripts/extract_text.py`:
```text
--input '/usr/local/src/pdf/tmp/<任务名>/source.pdf'
```
默认使用 `auto` 引擎:先由 Poppler `pdftotext` 提取;结果不可用或命令不可用时自动尝试 `pdfplumber`,并可逐页选择质量更好的结果。脚本使用 `pypdf` 获取标准页数,并拒绝把页数不一致的提取结果当作成功。不要直接执行 `pdftotext`。
默认单次最多处理 8 页、返回 24000 个字符。可使用:
- `--start-page <页码>`、`--end-page <页码>`:页码从 `1` 开始。
- `--start-offset <字符偏移>`:继续读取被字符上限截断的同一页;大于 `0` 时同时传入上次返回的 `next_engine`。
- `--max-pages <页数>`、`--max-chars <字符数>`:控制单次输出。
- `--layout`:仅在需要尽量保留版面空格时使用。
- `--engine <auto|poppler|pdfplumber>`:首次及跨页提取保持 `auto`;同页字符续读时传入上次返回的 `next_engine`。
- `--timeout <秒>`:Poppler 提取超时,默认 `120`。
先检查 `usable_for_summary` 和 `text_quality.status`:
- `usable_for_summary: true`:只使用 `pages[]` 中同样标为 `usable_for_summary: true` 的 `text`;可疑页的文本会被置空。如果 `has_more: true`,始终传回 `next_page` 和 `next_offset`。仅当 `next_offset` 大于 `0` 时,把非空的 `next_engine` 传给 `--engine` 以固定同页字符游标;这种调用只续读当前页。当前页完成后返回的 `next_offset` 为 `0`,此时不要传 `--engine`,让下一页重新使用 `auto`。保留首次调用的 `--end-page`(如果指定)及其他选项,直至 `has_more: false`。
- `usable_for_summary: false`:本批次没有可靠文本,不要使用返回内容。查看 `engine_attempts`、`text_quality.reasons`、`text_quality.suspect_pages` 和 `needs_ocr`;若 `has_more: true`,仍按跨页游标继续检查后续批次,避免漏掉后续可搜索文本。
- `needs_ocr: true`:一个或多个页面未得到可靠文本。把 `text_quality.suspect_pages` 中实际需要阅读的页码传给 `ocr_text.py`;先调用 OCR;只有 OCR 标记疑难、结果与页面结构冲突或需要理解非文本图形时,才渲染相关页供模型辅助复核。
`complete_text_coverage: true` 表示本批次所有页面均有可靠文本。`text_quality` 按页检测空白或过少文本、页面实际可见图像覆盖过大但文字不足、`(cid:...)`、Unicode 替换字符、异常控制字符及外观像汉字的部首字符;`pages[].extractor` 表示该页最终采用的引擎。`status: mixed` 表示同一批次同时包含可靠页和可疑页:可先使用可靠页文本,同时只核验 `suspect_pages`。不要只根据“肉眼看起来能读”判定提取结果可靠。
## OCR 与疑难复核
详见 [读取与 OCR](reading-and-ocr.md),使用固定 `ocr_text.py` 或 `ocr_image.py`。
## 提取表格
调用 `scripts/extract_tables.py`:
```text
--input '/usr/local/src/pdf/tmp/<任务名>/source.pdf' --start-page 1
```
默认单次最多处理 5 页、20 个表格和 2000 个单元格。可用 `--end-page`、`--start-table`、`--max-pages`、`--max-tables`、`--max-cells` 调整。若 `has_more: true`,把 `next_page` 传给 `--start-page`、`next_table` 传给 `--start-table` 后继续,并保留首次调用的 `--end-page`(如果指定)及其他提取选项。
## 渲染页面
只有满足以下任一条件时才调用 `scripts/render_pdf.py`:
- 用户明确要求审阅版式、图表、印章、公式或页面外观;
- 创建或修改 PDF 后进行最终视觉检查;
- OCR 标记低置信度、读取顺序冲突或无法识别区域,需要大模型辅助复核。
普通文本总结和扫描文字读取不自动增加整本视觉识别;OCR 的临时渲染由 `ocr_text.py` 内部完成。调用脚本时不要直接执行 `pdftoppm`:
```text
--input '/usr/local/src/pdf/tmp/<任务名>/source.pdf' --output-dir '/usr/local/src/pdf/tmp/<任务名>/rendered' --start-page 1
```
默认 150 DPI、单次最多 10 页。可使用 `--end-page`、`--max-pages`、`--dpi`、`--timeout` 和 `--overwrite`。若 `has_more: true`,使用 `next_page` 继续,并保留首次调用的 `--end-page`(如果指定)、输出目录及其他渲染选项。脚本返回标准化的 `page-0001.png` 文件路径。
对文字较小或图表密集的页面提高 DPI。使用可用的图像查看工具检查返回的 PNG,不要尝试把图片路径交给下载脚本。
## 简单文本创建 PDF
先使用 `write_file` 把内容写为 UTF-8 `.txt` 或 `.md` 文件,再调用 `scripts/create_pdf.py`:
```text
--input '/usr/local/src/pdf/tmp/<任务名>/content.md' --output '/usr/local/src/pdf/<文件名>.pdf' --title '文档标题'
```
脚本支持 Markdown 标题、项目符号和简单表格,自动选择可嵌入的 Unicode 字体并添加页码。可选参数:
- `--page-size <A4|LETTER>`
- `--font-path <TTF或TTC路径>`
- `--font-size <字号>`
- `--margin <points>`
- `--overwrite`
输入内容只使用 ASCII 连字符 `-`;脚本也会把常见 Unicode 横线规范化为 ASCII 连字符。
## 合并、拆分与旋转
调用 `scripts/manage_pdf.py`,第一个参数必须是操作名。
合并:
```text
merge --input 'a.pdf' --input 'b.pdf' --output '/usr/local/src/pdf/merged.pdf'
```
拆分指定范围:
```text
split --input 'source.pdf' --output-dir '/usr/local/src/pdf/split' --range 1-3 --range 4-6
```
不传 `--range` 时每页生成一个 PDF。
旋转指定页面:
```text
rotate --input 'source.pdf' --output '/usr/local/src/pdf/rotated.pdf' --pages '1,3-5' --degrees 90
```
`--degrees` 只能是 `90`、`180` 或 `270`;不传 `--pages` 时旋转全部页面。目标已存在且确认可覆盖时添加 `--overwrite`。
## 清理临时目录
调用 `scripts/cleanup_pdf_temp.py`:
```text
--task-dir '/usr/local/src/pdf/tmp/<任务名>'
```
脚本只允许删除 `/usr/local/src/pdf/tmp/` 下一级任务目录,拒绝删除根目录、仓库目录或其他路径。
## 质量要求
- 不覆盖用户提供的源文件。
- 创建或修改后重新检查页数、页面尺寸、加密状态和文本可读性。
- 扫描件先做原生文字检测,再只 OCR 可疑页;不得把低置信度 OCR 文本当作可靠正文。
- 逐页确认没有裁切、重叠、溢出、乱码、黑方块、错误分页或异常空白页。
- 检查标题层级、段落间距、页边距、表格、图表、图片、页码及章节衔接。
- 引用和参考文献必须可读,不得残留工具令牌、占位符或临时路径。
- 只有最新渲染结果不存在可见缺陷时才交付创建或修改后的 PDF。

View File

@ -0,0 +1,44 @@
# 原生文本、OCR 与模型辅助复核
## 识别顺序
1. `inspect_pdf.py` 检查页数与加密状态;`extract_text.py` 分批检查原生文本。
2. 可靠原生文字直接使用。对 `text_quality.suspect_pages` 调用本地 `ocr_text.py`,普通扫描件不先做整本图片识别。
3. PDF 任务中的单张截图、扫描图片用 `ocr_image.py`。文字识别以 OCR 为主,大模型用于具体疑难内容的辅助判断。
4. OCR 返回 `needs_review`、低置信度区域,或文字与表格结构/阅读顺序冲突时,先按原图质量合理提高 DPI(最多 400)。仍有疑问或涉及手写、复杂公式、图表、印章时,用 `render_pdf.py` 只渲染相关页,通过当前环境实际可用的图像查看/识别能力辅助复核。可使用已安装的 image-recognition skill;不能虚构工具或识别结果。
5. 复核时提供页码、待判断问题与 OCR 候选,区分“确认原文”“模型推测”“无法辨认”。不要让模型补写看不见的金额、账号、日期或姓名。辅助结果不能覆盖其他页已可靠提取的文字。
`needs_ocr: false` 表示原生文字无明显问题,不代表图中的趋势、流程、版式和关系已经被理解。任务需要这些非文本信息时可以按需查看相关页,不重复 OCR 全文。
## PDF 页 OCR
通过 `execute_skill_script` 调用 `scripts/ocr_text.py`,仅传参数:
```text
--input '/usr/local/src/pdf/tmp/task/source.pdf' --pages '2,5-6'
```
- 页码从 1 开始,一次最多 4 页;默认 260 DPI,允许 150–400,单页渲染不超过 2000 万像素。
- `--timeout` 是每页渲染超时,默认 180 秒,范围 1–600。
- 默认 `--max-chars 24000`,最高 60000。
- 字符游标:`next_offset > 0` 时用 `--pages <next_page> --start-offset <next_offset>` 续读该页;`next_offset = 0` 时用 `remaining_pages` 继续。单页续读结束后仍要继续先前未处理的页面。
- 使用镜像内 RapidOCR/ONNX 模型。固定脚本明确指定本地模型路径,缺少模型立即失败,不触发下载。渲染 PNG 只在脚本内部临时使用,完成即清理。
## 单张图像 OCR
调用 `scripts/ocr_image.py`:
```text
--input '/usr/local/src/pdf/tmp/task/scan.png'
```
支持 PNG/JPEG/WebP/TIFF/BMP,最大 25 MiB、2000 万像素、单帧;自动按 EXIF 调整方向。`--max-chars` 和 `--start-offset` 用于续读可靠文本。多页 TIFF 不默默只读首帧,应先得到按页文件或 PDF。
## 结果解读
- 仅使用 `usable_for_summary: true` 的 `text` 作为正文。`empty/sparse/low_confidence` 不等于空白页面;需要核验。
- 单行置信度低于 0.60 的文字不会混入可靠正文,即使整页平均置信度很高。`review_regions` 给出 `text_candidate`、置信度和像素框;这些是**待核验候选**,不能当事实。
- `needs_review: true` 可能与 `usable_for_summary: true` 同时出现:表示该页有可用文字,也有未确认区域。`complete_ocr_coverage` 只有全部请求页处理完成且无需复核时才为真。
- `review_regions` 最多返回 40 个,每段候选最多 500 字;截断时有显式标记。需要完整复核时查看对应原图,不把列表上限误认为没有其他疑点。
- OCR 的阅读顺序按几何位置排序,多栏、跨栏标题或复杂表格不保证逻辑顺序。表格先尝试 `extract_tables.py`;扫描表格用 OCR 字框对齐并核验行列、表头、合计,不把平铺文字直接当结构化表格。
- 引用使用真实 PDF 页码;记录实际阅读/识别范围,不能只处理前几页就声称已覆盖全文。

View File

@ -23,6 +23,15 @@ def quiet_pdf_library_logs() -> None:
quiet_pdf_library_logs() quiet_pdf_library_logs()
def check_poppler_resources(stderr: str) -> None:
"""A zero exit code can still leave every CJK glyph invisible."""
if "Missing language pack" in stderr:
raise RuntimeError(
"Poppler 缺少中文/CJK 编码映射;基础镜像需安装 poppler-data。"
"预览可能缺字,不能据此执行 OCR 或判断原文为空白。"
)
class SkillArgumentParser(argparse.ArgumentParser): class SkillArgumentParser(argparse.ArgumentParser):
def error(self, message: str) -> NoReturn: def error(self, message: str) -> NoReturn:
raise ValueError(f"参数错误:{message}") raise ValueError(f"参数错误:{message}")

View File

@ -0,0 +1,139 @@
'use strict';
// Internal renderer. The agent invokes create_design_pdf.py, never Node or this file.
const fs = require('fs');
const path = require('path');
const crypto = require('crypto');
function packageRoot(name) {
for (const root of require.resolve.paths(name) || []) {
const candidate = path.join(root, name, 'package.json');
if (fs.existsSync(candidate)) return fs.realpathSync(path.dirname(candidate));
}
throw new Error(`基础镜像缺少 ${name};需要更新镜像,任务中不能安装`);
}
function within(root, filename) {
const relative = path.relative(root, filename);
return relative === '' || (!relative.startsWith('..' + path.sep) && relative !== '..' && !path.isAbsolute(relative));
}
const MIME = {'.html':'text/html; charset=utf-8','.css':'text/css; charset=utf-8','.js':'text/javascript',
'.mjs':'text/javascript','.png':'image/png','.jpg':'image/jpeg','.jpeg':'image/jpeg','.webp':'image/webp',
'.svg':'image/svg+xml','.woff':'font/woff','.woff2':'font/woff2','.ttf':'font/ttf','.otf':'font/otf'};
async function render(request) {
const {chromium} = require('playwright-core');
const libraries = Object.fromEntries(['pagedjs','katex','mermaid'].map(name => [name, packageRoot(name)]));
const candidates = [process.env.CHROME_BIN, process.env.CHROME_PATH, '/usr/bin/chromium', '/usr/bin/chromium-browser'];
const executablePath = candidates.find(value => value && fs.existsSync(value));
if (!executablePath) throw new Error('基础镜像缺少可用 Chromium;检查 CHROME_BIN,任务中不能下载浏览器');
const sourceRoot = fs.realpathSync(path.dirname(request.input));
const assetRoot = fs.realpathSync(request.assets);
const errors = [], blocked = [];
const nonce = crypto.randomBytes(18).toString('base64');
const policy = `default-src 'none'; script-src 'nonce-${nonce}'; style-src 'unsafe-inline' https://pdf.local; img-src https://pdf.local data: blob:; font-src https://pdf.local data:; connect-src https://pdf.local; object-src 'none'; base-uri 'none'; form-action 'none'; frame-src 'none'`;
async function inject(page, filename) {
await page.evaluate(({content, nonce}) => {const script=document.createElement('script'); script.nonce=nonce; script.textContent=content; document.head.append(script);}, {content:fs.readFileSync(filename,'utf8'), nonce});
}
const browser = await chromium.launch({headless:true, executablePath,
args:['--no-sandbox','--disable-setuid-sandbox','--disable-dev-shm-usage'], timeout:request.timeout_ms});
try {
const context = await browser.newContext({serviceWorkers:'block', acceptDownloads:false});
const page = await context.newPage();
page.setDefaultTimeout(request.timeout_ms);
page.on('pageerror', error => errors.push(error.message.slice(0, 500)));
page.on('console', message => {
if (message.type() === 'error') errors.push(message.text().slice(0, 500));
});
await context.route('**/*', async route => {
const url = new URL(route.request().url());
if (url.protocol === 'data:' || url.protocol === 'blob:') return route.continue();
let root = sourceRoot, relative = decodeURIComponent(url.pathname).replace(/^\/+/, '');
if (url.origin !== 'https://pdf.local') {
blocked.push('外部资源:' + url.hostname); return route.abort();
}
if (relative.startsWith('__skill__/')) { root = assetRoot; relative = relative.slice(10); }
else if (relative.startsWith('__lib__/')) {
const parts = relative.split('/'); root = libraries[parts[1]]; relative = parts.slice(2).join('/');
}
try {
if (!root) throw new Error('unknown root');
const target = fs.realpathSync(path.resolve(root, relative));
if (!within(root, target) || !MIME[path.extname(target).toLowerCase()]) throw new Error('forbidden asset');
if (fs.statSync(target).size > 25 * 1024 * 1024) throw new Error('asset too large');
return route.fulfill({body:fs.readFileSync(target), contentType:MIME[path.extname(target).toLowerCase()], headers:{'Content-Security-Policy':policy}});
} catch (_) { blocked.push('缺失或不允许的本地资源:' + relative.slice(0, 200)); return route.abort(); }
});
await page.goto('https://pdf.local/' + encodeURIComponent(path.basename(request.input)), {waitUntil:'load'});
if (await page.evaluate(() => document.compatMode !== 'CSS1Compat'))
throw new Error('HTML 需以 <!doctype html> 声明标准模式,否则分页和数学公式无法正确排版');
// Reject active content in the actual DOM too, after the Python preflight.
const active = await page.evaluate(() => [...document.querySelectorAll('*')].some(el =>
['SCRIPT','IFRAME','OBJECT','EMBED','BASE','FRAME'].includes(el.tagName) ||
[...el.attributes].some(a => /^on/i.test(a.name) || a.name === 'srcdoc')));
if (active) throw new Error('HTML 包含主动脚本内容;只允许静态 HTML/CSS/SVG');
await page.emulateMedia({media:'print'});
let css = fs.readFileSync(path.join(assetRoot, 'design.css'), 'utf8');
if (request.page_size === 'LETTER') css = css.replaceAll('210mm','215.9mm').replaceAll('297mm','279.4mm').replaceAll('size: A4','size: Letter');
// Defaults precede author styles so explicit user typography takes precedence.
await page.evaluate(css => {const style=document.createElement('style'); style.textContent=css; document.head.prepend(style);}, css);
if (request.css) await page.addStyleTag({content:fs.readFileSync(request.css,'utf8')});
const stats = await page.evaluate(() => ({figures:document.querySelectorAll('figure').length,
tables:document.querySelectorAll('table').length, mermaid:document.querySelectorAll('.mermaid').length,
math:document.querySelectorAll('.math-inline,.math-display').length}));
if (stats.mermaid) {
const unsafe = await page.locator('.mermaid').evaluateAll(nodes => nodes.some(n => /%%\s*\{|^\s*---/m.test(n.textContent)));
if (unsafe) throw new Error('Mermaid 不允许内嵌配置;主题和安全选项由固定渲染器设置');
await inject(page, path.join(libraries.mermaid,'dist/mermaid.min.js'));
await page.evaluate(async () => {
mermaid.initialize({startOnLoad:false, securityLevel:'strict', theme:'neutral', maxTextSize:50000,
flowchart:{htmlLabels:false}, suppressErrorRendering:true});
await mermaid.run({querySelector:'.mermaid'});
});
}
if (stats.math) {
await page.addStyleTag({url:'https://pdf.local/__lib__/katex/dist/katex.min.css'});
await inject(page, path.join(libraries.katex,'dist/katex.min.js'));
await page.evaluate(() => {
for (const el of document.querySelectorAll('.math-inline,.math-display')) {
katex.render(el.textContent, el, {displayMode:el.classList.contains('math-display'), throwOnError:true,
trust:false, maxExpand:1000, maxSize:30, strict:'warn'});
}
});
}
await page.evaluate(async () => {
await document.fonts.ready;
await Promise.all([...document.images].map(img => img.complete ? Promise.resolve() : new Promise(resolve => {img.onload=img.onerror=resolve;})));
});
const broken = await page.evaluate(() => [...document.images].filter(img => !img.naturalWidth).length);
if (broken || blocked.length || errors.length) throw new Error('资源加载失败:' + [...new Set([...blocked,...errors]), ...(broken ? [`${broken} 张图片不可读`] : [])].join(';'));
await page.evaluate(() => {window.PagedConfig={auto:false};});
await inject(page, path.join(libraries.pagedjs,'dist/paged.polyfill.js'));
const total = await page.evaluate(async timeout => {
const flow = await Promise.race([window.PagedPolyfill.preview(), new Promise((_,reject) =>
setTimeout(() => reject(new Error('分页超时,未输出 PDF')), timeout))]);
await document.fonts.ready;
return flow.total;
}, request.timeout_ms);
if (!Number.isInteger(total) || total < 1 || total > 200) throw new Error('PDF 页数需在 1–200 之间');
if (errors.length || blocked.length) throw new Error([...errors,...blocked].slice(0,8).join(';'));
const pages = await page.evaluate(() => [...document.querySelectorAll('.pagedjs_page')].map((page, index) => {
const area = page.querySelector('.pagedjs_page_content');
const rect = area.getBoundingClientRect(), issues=[];
for (const el of area.querySelectorAll('table,figure,img,svg,pre,.math-display,h1,h2,h3,p')) {
const box=el.getBoundingClientRect();
if (box.width && (box.left < rect.left-2 || box.right > rect.right+2 || el.scrollWidth > el.clientWidth+3))
issues.push({tag:el.tagName.toLowerCase(),text:el.textContent.trim().slice(0,60)});
}
return {page:index+1,characters:area.innerText.trim().length,visuals:area.querySelectorAll('img,svg').length,overflows:issues.slice(0,10)};
}));
if (pages.length !== total) throw new Error('分页尚未完成:DOM 页数不一致');
if (pages.some(p => p.overflows.length)) throw new Error('检测到内容横向溢出:' + JSON.stringify(pages.filter(p=>p.overflows.length)));
await page.pdf({path:request.output, printBackground:true, preferCSSPageSize:true, tagged:true, scale:1,
timeout:request.timeout_ms});
return {ok:true,engine:'chromium+pagedjs',offline:true,page_count:total,content:stats,pages,
warnings:pages.filter(p => !p.characters && !p.visuals).map(p=>`第 ${p.page} 页可能为空白,需要视觉检查`)};
} finally {await browser.close();}
}
(async () => {
try {const result=await render(JSON.parse(fs.readFileSync(process.argv[2],'utf8'))); process.stdout.write(JSON.stringify(result));}
catch(error) {process.stdout.write(JSON.stringify({ok:false,error:String(error.message||error)})); process.exitCode=1;}
})();

View File

@ -0,0 +1,57 @@
#!/usr/bin/env python3
"""Compile a LaTeX project with cached Tectonic resources and shell escape disabled."""
from pathlib import Path
import re
import shutil
import subprocess
import sys
import tempfile
from _pdf_common import SkillArgumentParser, output_pdf, new_temp_pdf, publish_temp_file, run_cli
def compile_document(args):
from pypdf import PdfReader
source=Path(args.input).expanduser().resolve()
if not source.is_file() or source.suffix.lower()!='.tex' or not 0 < source.stat().st_size <= 2*1024*1024:
raise ValueError('输入需为不超过 2 MiB 的本地 .tex 文件')
if not 1 <= args.timeout <= 600:
raise ValueError('timeout 必须在 1–600 秒之间')
executable=shutil.which('tectonic')
if not executable:
raise RuntimeError('基础镜像缺少预置 Tectonic,需要更新镜像')
target=output_pdf(args.output,args.overwrite)
temporary=new_temp_pdf(target)
try:
with tempfile.TemporaryDirectory(prefix='pdf-latex-') as folder:
command=[executable,'--untrusted','--only-cached','--keep-logs','--outdir',folder,str(source)]
completed=subprocess.run(command,cwd=source.parent,capture_output=True,text=True,timeout=args.timeout,check=False)
result=Path(folder)/(source.stem+'.pdf')
messages=(completed.stdout+'\n'+completed.stderr).splitlines()
if completed.returncode or not result.is_file():
detail='\n'.join(messages[-15:])
raise RuntimeError('LaTeX 编译失败;缺失 TeX 包需在基础镜像构建时预置,任务中不下载:'+detail)
count=len(PdfReader(result).pages)
if not count:
raise ValueError('LaTeX 没有生成有效 PDF')
logfile=Path(folder)/(source.stem+'.log')
if logfile.exists():
messages += logfile.read_text(errors='replace').splitlines()
warnings=list(dict.fromkeys(line.strip() for line in messages if re.search(r'warning:|Overfull|Missing character|undefined references',line,re.I)))[:30]
shutil.copyfile(result,temporary)
publish_temp_file(temporary,target,args.overwrite)
finally:
temporary.unlink(missing_ok=True)
return {'source':str(source),'path':str(target),'page_count':count,'engine':'tectonic','dependency_mode':'cached-only',
'warnings':warnings,'requires_visual_review':True}
def main(argv=None):
parser=SkillArgumentParser(description='固定 LaTeX 编译接口,禁用 shell escape,只使用镜像内缓存包')
parser.add_argument('--input',required=True); parser.add_argument('--output',required=True)
parser.add_argument('--timeout',type=int,default=180); parser.add_argument('--overwrite',action='store_true')
return run_cli(lambda:compile_document(parser.parse_args(sys.argv[1:] if argv is None else argv)))
if __name__=='__main__':
raise SystemExit(main())

View File

@ -0,0 +1,57 @@
#!/usr/bin/env python3
"""Office to PDF using the preinstalled LibreOffice, isolated per invocation."""
from pathlib import Path
import shutil
import subprocess
import sys
import tempfile
from _pdf_common import SkillArgumentParser, output_pdf, new_temp_pdf, publish_temp_file, run_cli
FORMATS={'.docx','.doc','.odt','.rtf','.pptx','.ppt','.odp','.xlsx','.xls','.ods'}
def convert(args):
from pypdf import PdfReader
source=Path(args.input).expanduser().resolve()
if not source.is_file() or source.suffix.lower() not in FORMATS or not 0 < source.stat().st_size <= 25*1024*1024:
raise ValueError('输入需为不超过 25 MiB 的本地 Office 文档;不支持把 PDF 直接反向转成可编辑 Office')
if not 1 <= args.timeout <= 600:
raise ValueError('timeout 必须在 1–600 秒之间')
executable=shutil.which('soffice') or shutil.which('libreoffice')
if not executable:
raise RuntimeError('基础镜像缺少 LibreOffice,需要更新镜像')
target=output_pdf(args.output,args.overwrite)
temporary=new_temp_pdf(target)
try:
with tempfile.TemporaryDirectory(prefix='pdf-office-') as folder:
root=Path(folder)
incoming=root/'input'; outgoing=root/'output'; profile=root/'profile'
incoming.mkdir(); outgoing.mkdir(); profile.mkdir()
local=incoming/('source'+source.suffix.lower()); shutil.copyfile(source,local)
command=[executable,'-env:UserInstallation='+profile.as_uri(),'--headless','--nologo','--nodefault',
'--nofirststartwizard','--convert-to','pdf','--outdir',str(outgoing),str(local)]
completed=subprocess.run(command,capture_output=True,text=True,timeout=args.timeout,check=False)
result=outgoing/'source.pdf'
if completed.returncode or not result.is_file():
raise RuntimeError('Office 转 PDF 失败:'+(completed.stderr or completed.stdout)[-1500:])
count=len(PdfReader(result).pages)
if not count:
raise ValueError('转换结果没有页面')
shutil.copyfile(result,temporary)
publish_temp_file(temporary,target,args.overwrite)
finally:
temporary.unlink(missing_ok=True)
return {'source':str(source),'path':str(target),'page_count':count,'engine':'libreoffice','requires_visual_review':True,
'note':'转换前需在对应 Office skill 中重算公式并核对字体、图表和打印范围。'}
def main(argv=None):
parser=SkillArgumentParser(description='Office 文档导出 PDF,源文件保持不变')
parser.add_argument('--input',required=True); parser.add_argument('--output',required=True)
parser.add_argument('--timeout',type=int,default=180); parser.add_argument('--overwrite',action='store_true')
return run_cli(lambda:convert(parser.parse_args(sys.argv[1:] if argv is None else argv)))
if __name__=='__main__':
raise SystemExit(main())

View File

@ -0,0 +1,107 @@
#!/usr/bin/env python3
"""Render static HTML/CSS through the fixed, offline publication renderer."""
from __future__ import annotations
import json
import shutil
import subprocess
import sys
import tempfile
from html.parser import HTMLParser
from pathlib import Path
from _pdf_common import SkillArgumentParser, new_temp_pdf, output_pdf, publish_temp_file, run_cli
MAX_SOURCE_BYTES = 2 * 1024 * 1024
class StaticHTML(HTMLParser):
def handle_starttag(self, tag, attributes):
tag = tag.lower()
attrs = {key.lower(): value or '' for key, value in attributes}
if tag in {'script', 'iframe', 'object', 'embed', 'base', 'frame', 'frameset'}:
raise ValueError(f'HTML 不允许 {tag};仅支持静态 HTML/CSS/SVG,公式和 Mermaid 由固定渲染器处理')
if any(key.startswith('on') for key in attrs) or 'srcdoc' in attrs:
raise ValueError('HTML 不允许事件处理程序或 srcdoc')
if tag == 'meta' and 'http-equiv' in attrs:
raise ValueError('HTML 不允许 http-equiv;网络和文档策略由固定渲染器设置')
for key in ('href', 'src', 'xlink:href', 'action', 'formaction'):
value = ''.join(attrs.get(key, '').split()).lower()
if value.startswith(('javascript:', 'vbscript:', 'file:')):
raise ValueError('HTML 不允许脚本 URL 或 file: 资源;使用任务目录内相对路径')
handle_startendtag = handle_starttag
def read_source(value: str, suffixes: set[str]) -> Path:
source = Path(value).expanduser().resolve()
if not source.is_file() or source.suffix.lower() not in suffixes:
raise ValueError(f'输入必须是本地 {sorted(suffixes)} 文件')
if not 0 < source.stat().st_size <= MAX_SOURCE_BYTES:
raise ValueError('HTML/CSS 文件必须非空且不超过 2 MiB')
return source
def create(args) -> dict:
from pypdf import PdfReader
source = read_source(args.input, {'.html', '.htm'})
text = source.read_text(encoding='utf-8-sig')
validator = StaticHTML(convert_charrefs=True)
validator.feed(text)
validator.close()
css = read_source(args.css, {'.css'}) if args.css else None
if not 10 <= args.timeout <= 600:
raise ValueError('timeout 必须在 10–600 秒之间')
if args.expected_pages is not None and not 1 <= args.expected_pages <= 200:
raise ValueError('expected-pages 必须在 1–200 之间')
output = output_pdf(args.output, args.overwrite)
node = shutil.which('node')
if not node:
raise RuntimeError('基础镜像缺少预置 Node 运行时;需要更新镜像,任务中不能安装')
temporary = new_temp_pdf(output)
try:
with tempfile.TemporaryDirectory(prefix='pdf-design-') as folder:
request = Path(folder) / 'request.json'
request.write_text(json.dumps({'input': str(source), 'output': str(temporary), 'css': str(css) if css else None,
'page_size': args.page_size, 'timeout_ms': args.timeout * 1000,
'assets': str(Path(__file__).resolve().parents[1] / 'assets')}, ensure_ascii=False))
result = subprocess.run([node, str(Path(__file__).with_name('_render_html.cjs')), str(request)],
capture_output=True, text=True, timeout=args.timeout + 20, check=False)
try:
details = json.loads(result.stdout)
except (ValueError, TypeError):
raise RuntimeError('排版器未返回有效 JSON:' + (result.stderr or result.stdout)[-1500:])
if result.returncode or not details.get('ok'):
raise RuntimeError(details.get('error', 'HTML 排版失败'))
with temporary.open('rb') as handle:
reader = PdfReader(handle)
page_count = len(reader.pages)
if reader.is_encrypted or not page_count:
raise ValueError('排版器未生成有效 PDF')
if page_count != details['page_count']:
raise ValueError(f'分页结果与 PDF 页数不一致:{details["page_count"]} / {page_count}')
if args.expected_pages is not None and page_count != args.expected_pages:
raise ValueError(f'实际 {page_count} 页,与要求的 {args.expected_pages} 页不符;请调整排版后重试')
publish_temp_file(temporary, output, args.overwrite)
finally:
temporary.unlink(missing_ok=True)
details.pop('ok', None)
return {'path': str(output), 'source': str(source), 'size_bytes': output.stat().st_size, **details,
'requires_visual_review': True}
def main(argv=None) -> int:
parser = SkillArgumentParser(description='静态 HTML/CSS 排版为 PDF;使用离线 Paged.js、KaTeX、Mermaid 和系统 Chromium')
parser.add_argument('--input', required=True)
parser.add_argument('--output', required=True)
parser.add_argument('--css')
parser.add_argument('--page-size', choices=('A4', 'LETTER'), default='A4')
parser.add_argument('--expected-pages', type=int)
parser.add_argument('--timeout', type=int, default=180)
parser.add_argument('--overwrite', action='store_true')
return run_cli(lambda: create(parser.parse_args(sys.argv[1:] if argv is None else argv)))
if __name__ == '__main__':
raise SystemExit(main())

View File

@ -0,0 +1,246 @@
#!/usr/bin/env python3
"""AcroForm, crop, metadata and embedded-image operations using existing pypdf."""
from __future__ import annotations
import json
import math
import sys
from pathlib import Path
from _pdf_common import (SkillArgumentParser, input_pdf, output_pdf, output_directory,
new_temp_pdf, publish_temp_file, parse_page_spec, run_cli)
def load_data(inline, filename):
if bool(inline) == bool(filename):
raise ValueError('必须且只能提供 --data 或 --data-file')
if filename and Path(filename).stat().st_size > 2 * 1024 * 1024:
raise ValueError('JSON 上限为 2 MiB')
raw = Path(filename).read_bytes() if filename else inline.encode('utf-8')
if len(raw) > 2 * 1024 * 1024:
raise ValueError('JSON 上限为 2 MiB')
data = json.loads(raw)
if not isinstance(data, dict):
raise ValueError('JSON 必须是对象')
return data
def inherited_value(field, key, default=None):
"""Read an inheritable field attribute without following malformed cycles."""
seen = set()
while field is not None:
field = field.get_object()
if id(field) in seen:
raise ValueError('PDF 表单字段的父级引用存在循环')
seen.add(id(field))
if key in field:
return field[key]
field = field.get('/Parent')
return default
def field_info(reader):
result = {}
for name, field in (reader.get_fields() or {}).items():
# get_fields() is a summary: /AP and /MaxLen live on the full object.
original = field.indirect_reference.get_object()
ft = str(inherited_value(original, '/FT', ''))
flags = int(inherited_value(original, '/Ff', 0))
kind = {'/Tx':'text', '/Ch':'choice', '/Sig':'signature'}.get(ft, 'unknown')
if ft == '/Btn':
kind = 'pushbutton' if flags & (1 << 16) else ('radio' if flags & (1 << 15) else 'checkbox')
widgets = [original, *[kid.get_object() for kid in original.get('/Kids', [])]]
states = set()
if kind in {'checkbox', 'radio'}:
for widget in widgets:
appearance = widget.get('/AP')
normal = appearance.get_object().get('/N') if appearance is not None else None
if normal is not None:
states.update(str(key) for key in normal.get_object().keys())
if kind == 'radio' and flags & (1 << 14):
states.discard('/Off')
options = [{'value':str(option[0]), 'label':str(option[1])} if isinstance(option, list) else {'value':str(option), 'label':str(option)}
for option in inherited_value(original, '/Opt', [])]
value = inherited_value(original, '/V')
result[name] = {'id':name, 'type':kind, 'read_only':bool(flags & 1), 'flags':flags,
'current_value':[str(v) for v in value] if isinstance(value, list) else (str(value) if value is not None else None),
'states':sorted(states), 'options':options, 'max_length':inherited_value(original, '/MaxLen')}
return result
def validated_values(infos, data):
from pypdf.generic import NameObject
if not data or len(data) > 500:
raise ValueError('填表数据需为 1–500 个字段')
values = {}
for name, value in data.items():
if name not in infos:
raise ValueError(f'表单字段不存在:{name}')
field = infos[name]
if field['read_only']:
raise ValueError(f'字段为只读:{name}')
kind = field['type']
if kind in {'signature','pushbutton','unknown'}:
raise ValueError(f'不支持填写 {kind} 字段:{name}')
if kind in {'checkbox','radio'}:
states = field['states']
if type(value) is bool and kind == 'checkbox':
choices = [s for s in states if s != '/Off']
if value and len(choices) != 1:
raise ValueError(f'{name} 的选中状态不唯一,请使用具体状态名')
value = choices[0] if value else '/Off'
elif isinstance(value, str):
value = '/' + value.lstrip('/')
else:
raise ValueError(f'{name} 需布尔值或有效状态名')
if value not in states:
raise ValueError(f'{name} 状态无效;可选:{states}')
values[name] = NameObject(value)
elif kind == 'choice':
choices = {entry['value'] for entry in field['options']}
selected = value if isinstance(value, list) else [value]
if isinstance(value, list) and not field['flags'] & (1 << 21):
raise ValueError(f'{name} 不支持多选')
editable = bool(field['flags'] & (1 << 18))
if not all(isinstance(item, str) and (editable or item in choices) for item in selected):
raise ValueError(f'{name} 选项无效;可选:{sorted(choices)}')
values[name] = value
else:
if not isinstance(value, str) or (field['max_length'] and len(value) > field['max_length']):
raise ValueError(f'{name} 需文本,且不得超过字段长度上限')
values[name] = value
return values
def execute(args):
from pypdf import PdfReader, PdfWriter
from pypdf.generic import RectangleObject
source = input_pdf(args.input)
with source.open('rb') as handle:
reader = PdfReader(handle)
if reader.is_encrypted:
raise ValueError('PDF 已加密,请提供已解密副本')
if args.operation == 'form-info':
fields = list(field_info(reader).values())
if args.offset < 0 or not 1 <= args.limit <= 200:
raise ValueError('offset 不能小于 0,limit 需为 1–200')
stop = min(len(fields), args.offset + args.limit)
acroform = reader.trailer['/Root'].get('/AcroForm')
return {'field_count':len(fields), 'fields':fields[args.offset:stop], 'has_more':stop < len(fields),
'next_offset':stop if stop < len(fields) else None,
'has_xfa':bool(acroform and '/XFA' in acroform.get_object())}
if args.operation == 'extract-images':
return extract_images(reader, args)
output = output_pdf(args.output, args.overwrite)
if output == source:
raise ValueError('输出文件不能覆盖源 PDF')
writer = PdfWriter(clone_from=reader)
temporary = new_temp_pdf(output)
extra = {}
try:
if args.operation == 'form-fill':
acroform = reader.trailer['/Root'].get('/AcroForm')
if acroform and '/XFA' in acroform.get_object():
raise ValueError('XFA 表单不属于 AcroForm 固定接口,不能声称填写成功')
values = validated_values(field_info(reader), load_data(args.data, args.data_file))
writer.update_page_form_field_values(None, values, auto_regenerate=False)
extra = {'fields_filled':list(values), 'requires_visual_review':True}
elif args.operation == 'metadata':
data = load_data(args.data, args.data_file)
allowed = {'Title','Author','Subject','Keywords','Creator','Producer'}
if set(data) - allowed or not all(isinstance(v,str) and len(v) <= 4096 for v in data.values()):
raise ValueError('元数据仅支持 Title/Author/Subject/Keywords/Creator/Producer 文本字段')
writer.add_metadata({'/'+key:value for key,value in data.items()})
extra = {'updated_keys':list(data), 'xmp_preserved':'/Metadata' in reader.trailer['/Root']}
elif args.operation == 'crop':
box = [float(v) for v in args.box.split(',')]
if len(box) != 4 or not all(math.isfinite(v) for v in box) or box[0] >= box[2] or box[1] >= box[3]:
raise ValueError('box 需为左,下,右,上四个有限坐标,单位 pt')
pages = parse_page_spec(args.pages, len(writer.pages))
for number in pages:
media = writer.pages[number-1].mediabox
if box[0] < media.left or box[1] < media.bottom or box[2] > media.right or box[3] > media.top:
raise ValueError(f'裁剪框超出第 {number} 页 MediaBox')
writer.pages[number-1].cropbox = RectangleObject(box)
extra = {'cropped_pages':pages, 'box':box, 'warning':'裁剪只改变可见范围,不删除隐藏内容,不能用于脱敏。'}
with temporary.open('wb') as stream:
writer.write(stream)
check = PdfReader(temporary)
if len(check.pages) != len(reader.pages):
raise ValueError('编辑后页数异常')
if args.operation == 'form-fill':
actual = field_info(check)
for name, value in values.items():
wanted = [str(v) for v in value] if isinstance(value, list) else str(value)
if name not in actual or actual[name]['current_value'] != wanted:
raise ValueError(f'字段回读校验失败:{name}')
publish_temp_file(temporary, output, args.overwrite)
finally:
writer.close()
temporary.unlink(missing_ok=True)
return {'path':str(output), 'source':str(source), 'page_count':len(reader.pages), 'operation':args.operation, **extra}
def extract_images(reader, args):
if not 1 <= args.max_images <= 50 or args.start_image < 0:
raise ValueError('max-images 需为 1–50,start-image 不能小于 0')
pages = parse_page_spec(args.pages, len(reader.pages))
if len(pages) > 4:
raise ValueError('一次最多处理 4 页,请指定 pages')
destination = output_directory(args.output_dir)
result, total_bytes = [], 0
for page_index, number in enumerate(pages):
images = reader.pages[number-1].images
start = args.start_image if page_index == 0 else 0
for index in range(start, len(images)):
if len(result) >= args.max_images:
return {'images':result, 'has_more':True, 'next_page':number, 'next_image':index,
'remaining_pages':pages[page_index:], 'note':'嵌入图片不是整页截图;页面外观请用 render_pdf.py。'}
item = images[index]
if item.image.width * item.image.height > 20_000_000:
raise ValueError('嵌入图像超过 2000 万像素,请改为限制 DPI 的页面渲染')
total_bytes += len(item.data)
if total_bytes > 25 * 1024 * 1024:
raise ValueError('本批嵌入图片超过 25 MiB,请减少页数或图片数')
extension = Path(item.name).suffix.lower()
if extension not in {'.png','.jpg','.jpeg','.jp2','.tif','.tiff'}:
raise ValueError(f'不支持的嵌入图片编码:{extension},请用页面渲染')
output = destination / f'page-{number:04d}-image-{index:03d}{extension}'
temporary = new_temp_pdf(output)
try:
temporary.write_bytes(item.data)
publish_temp_file(temporary, output, args.overwrite)
finally:
temporary.unlink(missing_ok=True)
result.append({'page':number, 'image_index':index, 'path':str(output)})
return {'images':result, 'has_more':False, 'note':'嵌入图片不是整页截图;页面外观请用 render_pdf.py。'}
def main(argv=None):
parser = SkillArgumentParser(description='PDF 表单、裁剪、元数据和嵌入图片固定接口')
commands = parser.add_subparsers(dest='operation', required=True)
for name in ('form-info','form-fill','metadata','crop','extract-images'):
command = commands.add_parser(name)
command.add_argument('--input', required=True)
if name == 'form-info':
command.add_argument('--offset', type=int, default=0)
command.add_argument('--limit', type=int, default=50)
else:
command.add_argument('--overwrite', action='store_true')
command.add_argument('--output-dir' if name == 'extract-images' else '--output', required=True)
if name in {'form-fill','metadata'}:
command.add_argument('--data')
command.add_argument('--data-file')
if name in {'crop','extract-images'}:
command.add_argument('--pages')
if name == 'crop':
command.add_argument('--box', required=True)
if name == 'extract-images':
command.add_argument('--start-image', type=int, default=0)
command.add_argument('--max-images', type=int, default=20)
return run_cli(lambda: execute(parser.parse_args(sys.argv[1:] if argv is None else argv)))
if __name__ == '__main__':
raise SystemExit(main())

View File

@ -0,0 +1,47 @@
#!/usr/bin/env python3
"""Local OCR for a single image used in a PDF workflow."""
from pathlib import Path
import sys
import tempfile
from _pdf_common import SkillArgumentParser, run_cli
from ocr_text import _create_ocr_engine, _ocr_page, MAX_PIXELS_PER_PAGE
def recognize(args):
from PIL import Image, ImageOps
source = Path(args.input).expanduser().resolve()
if not source.is_file() or source.suffix.lower() not in {'.png','.jpg','.jpeg','.webp','.tif','.tiff','.bmp'}:
raise ValueError('输入必须是本地图像文件')
if source.stat().st_size > 25 * 1024 * 1024:
raise ValueError('图像文件不能超过 25 MiB')
if not 1 <= args.max_chars <= 60000 or args.start_offset < 0:
raise ValueError('max-chars 需为 1–60000,start-offset 不能小于 0')
with Image.open(source) as original:
if original.width * original.height > MAX_PIXELS_PER_PAGE or getattr(original, 'n_frames', 1) != 1:
raise ValueError('图像上限 2000 万像素,且必须为单帧;多页扫描件请按 PDF 分页 OCR')
with tempfile.TemporaryDirectory(prefix='pdf-image-ocr-') as folder:
normalized = Path(folder) / 'image.png'
ImageOps.exif_transpose(original).convert('RGB').save(normalized)
result = _ocr_page(_create_ocr_engine(), normalized)
text = result.pop('text')
if args.start_offset > len(text):
raise ValueError('start-offset 超过可靠 OCR 文本长度')
selection = text[args.start_offset:args.start_offset + args.max_chars]
next_offset = args.start_offset + len(selection)
return {'path': str(source), 'engine': 'rapidocr', 'offline': True, **result,
'text': selection, 'char_count': len(text), 'offset_start': args.start_offset,
'offset_end': next_offset, 'has_more': next_offset < len(text),
'next_offset': next_offset if next_offset < len(text) else None}
def main(argv=None):
parser = SkillArgumentParser(description='PDF 任务图像的本地 OCR,疑难区域保留复核标记')
parser.add_argument('--input', required=True)
parser.add_argument('--max-chars', type=int, default=24000)
parser.add_argument('--start-offset', type=int, default=0)
return run_cli(lambda: recognize(parser.parse_args(sys.argv[1:] if argv is None else argv)))
if __name__ == '__main__':
raise SystemExit(main())

View File

@ -6,6 +6,7 @@ import contextlib
import importlib.metadata import importlib.metadata
import io import io
import logging import logging
import math
import os import os
import re import re
import shutil import shutil
@ -18,6 +19,7 @@ from typing import Any
from _pdf_common import ( from _pdf_common import (
SkillArgumentParser, SkillArgumentParser,
check_poppler_resources,
input_pdf, input_pdf,
parse_page_spec, parse_page_spec,
run_cli, run_cli,
@ -178,6 +180,7 @@ def _render_page(
f"第 {page_number} 页渲染失败:" f"第 {page_number} 页渲染失败:"
f"{detail or 'pdftoppm 返回错误'}" f"{detail or 'pdftoppm 返回错误'}"
) )
check_poppler_resources(completed.stderr)
if not output.is_file() or output.stat().st_size <= 0: if not output.is_file() or output.stat().st_size <= 0:
raise RuntimeError(f"第 {page_number} 页没有生成有效 PNG") raise RuntimeError(f"第 {page_number} 页没有生成有效 PNG")
return output, elapsed return output, elapsed
@ -185,17 +188,29 @@ def _render_page(
def _create_ocr_engine(): def _create_ocr_engine():
try: try:
import rapidocr
from rapidocr import RapidOCR from rapidocr import RapidOCR
except ImportError as exc: except ImportError as exc:
raise RuntimeError("环境预置的 rapidocr 模块不可用") from exc raise RuntimeError("环境预置的 rapidocr 模块不可用") from exc
model_dir = Path(rapidocr.__file__).resolve().parent / "models"
models = {"Det": "PP-OCRv6_det_small.onnx", "Cls": "ch_ppocr_mobile_v2.0_cls_mobile.onnx", "Rec": "PP-OCRv6_rec_small.onnx"}
missing = [filename for filename in models.values() if not (model_dir / filename).is_file()]
if missing:
raise RuntimeError(f"基础镜像缺少本地 OCR 模型:{missing};任务中不能下载")
params = {
"Global.log_level": "error",
# Keep uncertain lines for explicit review instead of silently dropping them.
"Global.text_score": 0.0,
**{f"{name}.model_path": str(model_dir / filename) for name, filename in models.items()},
}
captured_stdout = io.StringIO() captured_stdout = io.StringIO()
captured_stderr = io.StringIO() captured_stderr = io.StringIO()
with ( with (
contextlib.redirect_stdout(captured_stdout), contextlib.redirect_stdout(captured_stdout),
contextlib.redirect_stderr(captured_stderr), contextlib.redirect_stderr(captured_stderr),
): ):
return RapidOCR() return RapidOCR(params=params)
def _clean_text(value: Any) -> str: def _clean_text(value: Any) -> str:
@ -235,7 +250,7 @@ def _ordered_lines(result: Any) -> list[dict[str, Any]]:
confidence = float(scores[index]) confidence = float(scores[index])
except (IndexError, TypeError, ValueError): except (IndexError, TypeError, ValueError):
confidence = 0.0 confidence = 0.0
confidence = max(0.0, min(1.0, confidence)) confidence = max(0.0, min(1.0, confidence)) if math.isfinite(confidence) else 0.0
box = _box_points(boxes[index] if index < len(boxes) else None) box = _box_points(boxes[index] if index < len(boxes) else None)
if box: if box:
left = min(point[0] for point in box) left = min(point[0] for point in box)
@ -314,8 +329,17 @@ def _ocr_page(engine: Any, image_path: Path) -> dict[str, Any]:
else: else:
status = "good" status = "good"
reliable_lines = [line for line in lines if line["confidence"] >= MIN_MEAN_CONFIDENCE]
review_lines = [line for line in lines if line["confidence"] < MIN_MEAN_CONFIDENCE]
# A high page average must not certify a low-confidence amount or identifier.
reliable_text = "\n".join(line["text"] for line in reliable_lines)
return { return {
"text": text, "text": reliable_text if status == "good" else "",
"raw_char_count": len(text),
"needs_review": status != "good" or bool(review_lines),
"review_regions": [{"text_candidate": line["text"][:500], "confidence": round(line["confidence"], 4), "box": line["box"]} for line in review_lines[:40]],
"review_regions_truncated": len(review_lines) > 40,
"reading_order": "geometric_top_to_bottom",
"status": status, "status": status,
"usable_for_summary": status == "good", "usable_for_summary": status == "good",
"line_count": len(lines), "line_count": len(lines),
@ -436,9 +460,10 @@ def _extract(args) -> dict[str, Any]:
), ),
"complete_ocr_coverage": ( "complete_ocr_coverage": (
all_processed and all_complete and all_usable all_processed and all_complete and all_usable
and not any(page["needs_review"] for page in page_outputs)
), ),
"needs_review": any( "needs_review": any(
not page["usable_for_summary"] for page in page_outputs page["needs_review"] for page in page_outputs
), ),
"has_more": has_more, "has_more": has_more,
"next_page": next_page, "next_page": next_page,

View File

@ -12,6 +12,7 @@ from typing import Any
from _pdf_common import ( from _pdf_common import (
SkillArgumentParser, SkillArgumentParser,
check_poppler_resources,
input_pdf, input_pdf,
output_directory, output_directory,
publish_temp_file, publish_temp_file,
@ -155,6 +156,7 @@ def _render(args) -> dict[str, Any]:
if completed.returncode != 0: if completed.returncode != 0:
detail = (completed.stderr or completed.stdout or "").strip()[-2000:] detail = (completed.stderr or completed.stdout or "").strip()[-2000:]
raise RuntimeError(f"PDF 渲染失败:{detail or 'pdftoppm 返回错误'}") raise RuntimeError(f"PDF 渲染失败:{detail or 'pdftoppm 返回错误'}")
check_poppler_resources(completed.stderr)
rendered = sorted( rendered = sorted(
temp_dir.glob("page-*.png"), temp_dir.glob("page-*.png"),

View File

@ -0,0 +1,246 @@
"""Behavioral checks for OCR quality gates and non-destructive PDF operations.
Uses the existing pypdf, ReportLab and Pillow dependencies. Browser, OCR-model
and Office integration checks run separately against real installed engines.
"""
from __future__ import annotations
import hashlib
import json
import sys
import tempfile
import unittest
from pathlib import Path
from types import SimpleNamespace
from unittest.mock import patch
sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "scripts"))
from PIL import Image
from pypdf import PdfReader, PdfWriter
from pypdf.generic import DictionaryObject, NameObject, TextStringObject
from reportlab.pdfgen import canvas
import _pdf_common as common
import create_design_pdf
import edit_pdf
import ocr_text
import render_pdf
class PDFWorkflows(unittest.TestCase):
def setUp(self):
self.directory = tempfile.TemporaryDirectory(prefix="pdf-regression-")
self.root = Path(self.directory.name).resolve()
self.root_patch = patch.object(common, "PDF_OUTPUT_ROOT", self.root)
self.root_patch.start()
self.source = self.root / "source.pdf"
c = canvas.Canvas(str(self.source), pagesize=(600, 800))
c.drawString(30, 750, "Original content 123")
form = c.acroForm
form.textfield(name="name", x=30, y=650, width=200, height=25, maxlen=40)
form.checkbox(name="agree", x=30, y=600, checked=False)
form.radio(name="mode", value="A", x=30, y=550, selected=True)
form.radio(name="mode", value="B", x=80, y=550, selected=False)
form.choice(name="country", value="CN", options=["CN", "US"], x=30, y=480, width=150, height=30)
image = Image.new("RGB", (24, 16), "#8a3a2a")
c.drawInlineImage(image, 350, 500, width=96, height=64)
c.showPage()
c.save()
self.original_hash = hashlib.sha256(self.source.read_bytes()).hexdigest()
def tearDown(self):
self.assertEqual(hashlib.sha256(self.source.read_bytes()).hexdigest(), self.original_hash)
self.root_patch.stop()
self.directory.cleanup()
def args(self, operation, **values):
return SimpleNamespace(operation=operation, input=str(self.source),
output=str(self.root / "result.pdf"), overwrite=False,
data=None, data_file=None, pages=None, **values)
def test_form_inventory_and_roundtrip_all_supported_types(self):
fields = edit_pdf.execute(self.args("form-info", offset=0, limit=50))
infos = {field["id"]: field for field in fields["fields"]}
self.assertEqual(infos["agree"]["type"], "checkbox")
self.assertEqual(set(infos["agree"]["states"]), {"/Off", "/Yes"})
self.assertEqual(set(infos["mode"]["states"]), {"/A", "/B"})
self.assertFalse(fields["has_xfa"])
args = self.args("form-fill")
args.data = json.dumps({"name": "Alice 123", "agree": True, "mode": "B", "country": "US"})
result = edit_pdf.execute(args)
updated = PdfReader(result["path"])
fields = updated.get_fields()
self.assertEqual(fields["name"]["/V"], "Alice 123")
self.assertEqual(fields["agree"]["/V"], "/Yes")
self.assertEqual(fields["mode"]["/V"], "/B")
self.assertEqual(fields["country"]["/V"], "US")
widgets = [ref.get_object() for ref in updated.pages[0]["/Annots"]]
checkbox = next(widget for widget in widgets if widget.get("/T") == "agree")
self.assertEqual(checkbox["/AS"], "/Yes")
radio = [widget for widget in widgets if widget.get("/Parent")]
self.assertEqual(sorted(str(widget["/AS"]) for widget in radio), ["/B", "/Off"])
self.assertIn("Original content 123", updated.pages[0].extract_text())
def test_checkbox_false_is_off(self):
args = self.args("form-fill")
args.data = '{"agree":false}'
result = edit_pdf.execute(args)
self.assertEqual(PdfReader(result["path"]).get_fields()["agree"]["/V"], "/Off")
def test_invalid_form_input_does_not_publish(self):
for data in [{"missing":"x"}, {"agree":"false"}, {"mode":True}, {"mode":"Off"},
{"country":"ZZ"}, {"country":["CN","US"]}, {"name":"x"*41}]:
with self.subTest(data=data):
args = self.args("form-fill")
args.data = json.dumps(data)
with self.assertRaises(ValueError):
edit_pdf.execute(args)
self.assertFalse(Path(args.output).exists())
def test_form_inventory_handles_indirect_appearance(self):
writer = PdfWriter(clone_from=self.source)
for ref in writer.pages[0]["/Annots"]:
widget = ref.get_object()
if "/AP" in widget:
widget[NameObject("/AP")] = writer._add_object(widget["/AP"])
alternative = self.root / "indirect.pdf"
writer.write(alternative)
args = self.args("form-info", offset=0, limit=50)
args.input = str(alternative)
result = edit_pdf.execute(args)
self.assertEqual(result["field_count"], 4)
def test_xfa_rejected_for_fill(self):
writer = PdfWriter(clone_from=self.source)
writer.root_object["/AcroForm"][NameObject("/XFA")] = TextStringObject("unsupported")
alternative = self.root / "xfa.pdf"
writer.write(alternative)
args = self.args("form-fill")
args.input = str(alternative)
args.data = '{"name":"Alice"}'
with self.assertRaisesRegex(ValueError, "XFA"):
edit_pdf.execute(args)
self.assertFalse(Path(args.output).exists())
def test_crop_keeps_content_and_forms(self):
result = edit_pdf.execute(self.args("crop", box="50,50,500,700"))
reader = PdfReader(result["path"])
self.assertEqual(list(reader.pages[0].cropbox), [50, 50, 500, 700])
self.assertIn("Original content 123", reader.pages[0].extract_text())
self.assertEqual(len(reader.get_fields()), 4)
def test_crop_rejects_out_of_bounds_and_nonfinite_values(self):
for box in ["-1,0,300,400", "0,0,601,800", "10,0,0,20", "0,0,nan,20"]:
with self.subTest(box=box), self.assertRaises(ValueError):
edit_pdf.execute(self.args("crop", box=box))
self.assertFalse((self.root / "result.pdf").exists())
def test_metadata_preserves_forms_and_unspecified_fields(self):
original = PdfReader(self.source).metadata
args = self.args("metadata")
args.data = '{"Title":"中文报告","Author":"Test"}'
result = edit_pdf.execute(args)
reader = PdfReader(result["path"])
self.assertEqual(reader.metadata.title, "中文报告")
self.assertEqual(reader.metadata.producer, original.producer)
self.assertEqual(len(reader.get_fields()), 4)
def test_embedded_image_has_original_dimensions(self):
args = self.args("extract-images", output_dir=str(self.root / "images"), start_image=0, max_images=20)
result = edit_pdf.execute(args)
self.assertEqual(len(result["images"]), 1)
with Image.open(result["images"][0]["path"]) as embedded:
self.assertEqual(embedded.size, (24, 16))
def test_existing_output_is_preserved_on_failure(self):
target = self.root / "result.pdf"
target.write_bytes(b"previous result")
args = self.args("metadata")
args.overwrite = True
args.data = '{"Unsupported":"value"}'
with self.assertRaises(ValueError):
edit_pdf.execute(args)
self.assertEqual(target.read_bytes(), b"previous result")
def test_output_cannot_replace_source(self):
args = self.args("crop", box="0,0,100,100")
args.output = str(self.source)
args.overwrite = True
with self.assertRaises(ValueError):
edit_pdf.execute(args)
def test_output_directory_boundary(self):
args = self.args("crop", box="0,0,100,100")
args.output = str(self.root.parent / "outside.pdf")
with self.assertRaises(ValueError):
edit_pdf.execute(args)
def test_missing_cjk_maps_rejects_incomplete_preview(self):
destination = self.root / 'preview'
args = render_pdf._parse_args(['--input', str(self.source), '--output-dir', str(destination)])
completed = SimpleNamespace(returncode=0, stdout='', stderr="Syntax Error: Missing language pack for 'Adobe-GB1' mapping")
with patch.object(render_pdf.shutil, 'which', return_value=sys.executable), patch.object(render_pdf.subprocess, 'run', return_value=completed):
with self.assertRaisesRegex(RuntimeError, 'poppler-data'):
render_pdf._render(args)
self.assertEqual(list(destination.glob('*.png')), [])
class OCRQuality(unittest.TestCase):
def recognize(self, texts, scores):
boxes = [[[0, i*30], [100, i*30], [100, i*30+20], [0, i*30+20]] for i in range(len(texts))]
engine = lambda _: SimpleNamespace(txts=texts, scores=scores, boxes=boxes)
return ocr_text._ocr_page(engine, Path("fixture.png"))
def test_mixed_confidence_does_not_certify_uncertain_amount(self):
result = self.recognize(["Reliable document heading and text", "9999.99"], [.99, .3])
self.assertEqual(result["status"], "good")
self.assertNotIn("9999.99", result["text"])
self.assertTrue(result["needs_review"])
self.assertEqual(result["review_regions"][0]["text_candidate"], "9999.99")
self.assertIsNotNone(result["review_regions"][0]["box"])
def test_unusable_page_is_not_returned_as_reliable_text(self):
for texts, scores in [(["uncertain document"], [.3]), (["Hi"], [.99]), ([], [])]:
with self.subTest(texts=texts):
result = self.recognize(texts, scores)
self.assertEqual(result["text"], "")
self.assertFalse(result["usable_for_summary"])
self.assertTrue(result["needs_review"])
def test_good_ocr_needs_no_model_assistance(self):
result = self.recognize(["中文识别测试 12345", "English document"], [.99, .98])
self.assertFalse(result["needs_review"])
self.assertIn("中文识别测试", result["text"])
def test_invalid_confidence_is_not_treated_as_certain(self):
result = self.recognize(['Unknown confidence amount 9999'], [float('nan')])
self.assertEqual(result['text'], '')
self.assertTrue(result['needs_review'])
def test_missing_local_models_fails_before_engine_initialization(self):
with tempfile.TemporaryDirectory() as folder:
fake = SimpleNamespace(__file__=str(Path(folder) / "__init__.py"))
fake.RapidOCR = lambda **_: self.fail("Engine must not download missing models")
with patch.dict(sys.modules, {"rapidocr":fake}):
with self.assertRaisesRegex(RuntimeError, "本地 OCR 模型"):
ocr_text._create_ocr_engine()
class StaticDocument(unittest.TestCase):
def test_active_content_is_rejected(self):
for html in ["<script>alert(1)</script>", '<svg onload="x()"></svg>',
'<iframe src="https://example.com"></iframe>',
'<a href="java&#x73;cript:alert(1)">x</a>',
'<meta http-equiv="refresh" content="0;url=https://example.com">',
'<img src="file:///etc/passwd">']:
with self.subTest(html=html), self.assertRaises(ValueError):
create_design_pdf.StaticHTML().feed(html)
def test_static_math_diagram_and_svg_are_accepted(self):
html = '<h1>报告</h1><p class="math-inline">E=mc^2</p><div class="mermaid">flowchart LR\nA-->B</div><svg><rect width="10" height="10" /></svg>'
create_design_pdf.StaticHTML().feed(html)
if __name__ == "__main__":
unittest.main()