feat: word ocr, ppt skills
This commit is contained in:
parent
b81403f199
commit
b3193cb0c7
@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
name: docx
|
name: docx
|
||||||
description: "创建、读取、编辑、转换、批注、接受修订、校验和渲染本地或远程 HTTPS Microsoft Word 文档。用户提到 Word、文档、报告、备忘录、合同、信函、模板、目录、页眉页脚、页码、表格、图片、批注或修订,或提供 HTTPS Word 地址、.docx、.dotx、.doc 文件时使用;支持安全下载、结构化创建、跨 Run 查找替换、安全 OOXML 解包/打包、旧格式转换、关系与 XML 校验及逐页视觉检查。若主要交付物是 PDF、电子表格、Google Docs 或普通代码,则不要使用。"
|
description: "创建、读取、编辑、转换、批注、接受修订、校验和渲染本地或远程 HTTPS Microsoft Word 文档,并按需识别文档图片、截图和扫描页中的文字。用户提到 Word、文档、报告、备忘录、合同、信函、模板、目录、页眉页脚、页码、表格、图片、图片文字 OCR、批注或修订,或提供 HTTPS Word 地址、.docx、.dotx、.doc 文件时使用;支持安全下载、结构化创建、跨 Run 查找替换、本地 RapidOCR、安全 OOXML 解包/打包、旧格式转换、关系与 XML 校验及逐页视觉检查。若主要交付物是 PDF、电子表格、Google Docs 或普通代码,则不要使用。"
|
||||||
---
|
---
|
||||||
|
|
||||||
# Word 文档处理
|
# Word 文档处理
|
||||||
@ -13,6 +13,7 @@ description: "创建、读取、编辑、转换、批注、接受修订、校验
|
|||||||
- 不把 `python3`、`soffice`、`libreoffice`、`pandoc`、`pdftoppm`、`zip`、`unzip`、`find`、`rm` 或其他系统命令作为脚本参数。
|
- 不把 `python3`、`soffice`、`libreoffice`、`pandoc`、`pdftoppm`、`zip`、`unzip`、`find`、`rm` 或其他系统命令作为脚本参数。
|
||||||
- LibreOffice、Pandoc、Poppler 和 ZIP 操作只允许由固定 Python 脚本在内部调用。
|
- LibreOffice、Pandoc、Poppler 和 ZIP 操作只允许由固定 Python 脚本在内部调用。
|
||||||
- 每次检查脚本返回 JSON;只有 `ok` 为 `true` 时才继续。`validate_document.py` 还必须返回 `status: valid`。
|
- 每次检查脚本返回 JSON;只有 `ok` 为 `true` 时才继续。`validate_document.py` 还必须返回 `status: valid`。
|
||||||
|
- 只在需要读取图片、截图或扫描页中的文字时调用 `ocr_document.py`。只使用 `pages[]` 中 `usable_for_summary: true` 的 `text`;低置信度结果不得作为可靠正文。
|
||||||
- 远程地址只交给 `download_document.py`;不要在回复、日志摘要或文件名中复述可能含敏感查询参数的完整 URL。
|
- 远程地址只交给 `download_document.py`;不要在回复、日志摘要或文件名中复述可能含敏感查询参数的完整 URL。
|
||||||
- 不覆盖用户提供的源文件。最终结果写入 `output/docx/`,中间产物写入 `tmp/docx/<任务名>/`。
|
- 不覆盖用户提供的源文件。最终结果写入 `output/docx/`,中间产物写入 `tmp/docx/<任务名>/`。
|
||||||
- 环境已预置依赖,不安装软件包,也不提示用户安装依赖。
|
- 环境已预置依赖,不安装软件包,也不提示用户安装依赖。
|
||||||
@ -23,6 +24,7 @@ description: "创建、读取、编辑、转换、批注、接受修订、校验
|
|||||||
| --- | --- | --- |
|
| --- | --- | --- |
|
||||||
| `scripts/download_document.py` | 下载并校验远程 HTTPS Word 文档 | Python `urllib`、安全 OOXML 解析、`python-docx` |
|
| `scripts/download_document.py` | 下载并校验远程 HTTPS Word 文档 | Python `urllib`、安全 OOXML 解析、`python-docx` |
|
||||||
| `scripts/inspect_document.py` | 分段读取正文、表格、样式、批注和修订 | `python-docx`、安全 OOXML 解析 |
|
| `scripts/inspect_document.py` | 分段读取正文、表格、样式、批注和修订 | `python-docx`、安全 OOXML 解析 |
|
||||||
|
| `scripts/ocr_document.py` | 按页识别图片、截图和扫描页中的文字 | RapidOCR、ONNX Runtime、LibreOffice、Poppler、`pdfplumber` |
|
||||||
| `scripts/create_document.py` | 按受控 JSON 创建专业 DOCX | `python-docx`、Pillow |
|
| `scripts/create_document.py` | 按受控 JSON 创建专业 DOCX | `python-docx`、Pillow |
|
||||||
| `scripts/edit_document.py` | 查找替换、追加/插入内容、调整样式和页面 | `python-docx` |
|
| `scripts/edit_document.py` | 查找替换、追加/插入内容、调整样式和页面 | `python-docx` |
|
||||||
| `scripts/add_comment.py` | 给精确文本范围添加批注 | `python-docx` |
|
| `scripts/add_comment.py` | 给精确文本范围添加批注 | `python-docx` |
|
||||||
@ -38,12 +40,13 @@ description: "创建、读取、编辑、转换、批注、接受修订、校验
|
|||||||
1. 输入是 HTTPS 地址时,先调用 `download_document.py` 下载到本次任务临时目录;本地文件直接进入下一步。
|
1. 输入是 HTTPS 地址时,先调用 `download_document.py` 下载到本次任务临时目录;本地文件直接进入下一步。
|
||||||
2. 旧版 `.doc` 或模板 `.dotx` 先调用 `convert_document.py` 转为 `.docx`;保留原文件。
|
2. 旧版 `.doc` 或模板 `.dotx` 先调用 `convert_document.py` 转为 `.docx`;保留原文件。
|
||||||
3. 编辑、总结或重组现有文档前调用 `inspect_document.py`,确认段落、表格、章节、页眉页脚、批注和修订状态。
|
3. 编辑、总结或重组现有文档前调用 `inspect_document.py`,确认段落、表格、章节、页眉页脚、批注和修订状态。
|
||||||
4. 新建文档使用 `create_document.py`;常规编辑使用 `edit_document.py`;添加批注使用 `add_comment.py`。
|
4. 需要读取截图、扫描页或图片中的文字时调用 `ocr_document.py`。省略 `--pages` 可自动选择含有效图片的页面;不要默认 OCR 没有图片的普通正文页,也不要用 OCR 覆盖可靠的原生文本。
|
||||||
5. 输入有修订时,先确认用户希望保留还是接受。普通编辑脚本默认拒绝含修订的文档,避免把修订静默损坏。
|
5. 新建文档使用 `create_document.py`;常规编辑使用 `edit_document.py`;添加批注使用 `add_comment.py`。
|
||||||
6. 只有固定编辑脚本不能完成的 OOXML 高级需求,才使用 `unpack_document.py` → 编辑 XML 文件 → `pack_document.py`;不得直接运行 ZIP 或 shell 命令。
|
6. 输入有修订时,先确认用户希望保留还是接受。普通编辑脚本默认拒绝含修订的文档,避免把修订静默损坏。
|
||||||
7. 所有创建或修改结果必须调用 `validate_document.py --check-convert`,确保 `status: valid`、`issue_count: 0`。
|
7. 只有固定编辑脚本不能完成的 OOXML 高级需求,才使用 `unpack_document.py` → 编辑 XML 文件 → `pack_document.py`;不得直接运行 ZIP 或 shell 命令。
|
||||||
8. 再调用 `render_document.py` 渲染全部页面,逐页检查版式;有游标时继续到 `has_more: false`。
|
8. 所有创建或修改结果必须调用 `validate_document.py --check-convert`,确保 `status: valid`、`issue_count: 0`。
|
||||||
9. 结构、内容、修订/批注和视觉检查都通过后才交付。
|
9. 再调用 `render_document.py` 渲染全部页面,逐页检查版式;有游标时继续到 `has_more: false`。
|
||||||
|
10. 结构、内容、修订/批注和视觉检查都通过后才交付。
|
||||||
|
|
||||||
## 下载远程文档
|
## 下载远程文档
|
||||||
|
|
||||||
@ -85,9 +88,36 @@ description: "创建、读取、编辑、转换、批注、接受修订、校验
|
|||||||
- `tracked_changes.total` 和 `authors`:是否存在修订及修订作者。
|
- `tracked_changes.total` 和 `authors`:是否存在修订及修订作者。
|
||||||
- `comments`:批注正文和作者。
|
- `comments`:批注正文和作者。
|
||||||
- `sections`:纸张、方向、页边距、页眉、页脚。
|
- `sections`:纸张、方向、页边距、页眉、页脚。
|
||||||
|
- `has_images`、`inline_image_count`、`media_part_count`:是否需要进一步读取图片文字;浮动图片可能只计入媒体部件。
|
||||||
- `archive.missing_required_parts`、`duplicate_members`:结构异常。
|
- `archive.missing_required_parts`、`duplicate_members`:结构异常。
|
||||||
- `has_more`、`next_paragraph`、`next_table`:继续读取长文档。
|
- `has_more`、`next_paragraph`、`next_table`:继续读取长文档。
|
||||||
|
|
||||||
|
## 识别图片中的文字
|
||||||
|
|
||||||
|
需要读取图片、截图或扫描页中的文字时调用:
|
||||||
|
|
||||||
|
```text
|
||||||
|
--input 'source.docx'
|
||||||
|
```
|
||||||
|
|
||||||
|
省略 `--pages` 时,脚本会把 Word 临时转换为 PDF,自动选择包含足够大图片的页面,每次最多处理 4 页。需要识别较小图片或指定页面时传:
|
||||||
|
|
||||||
|
```text
|
||||||
|
--input 'source.docx' --pages '2,5-6'
|
||||||
|
```
|
||||||
|
|
||||||
|
脚本通过 LibreOffice 和 Poppler 临时渲染页面,使用本地 RapidOCR 识别图片区域;临时 PDF 和 PNG 会自动删除,不联网,也不调用大模型识图。PDF 原生文本层用于过滤正文、页眉、页脚和页码产生的重复 OCR,因此 `pages[].text` 只返回可靠的额外图片文字。
|
||||||
|
|
||||||
|
检查:
|
||||||
|
|
||||||
|
- `candidate_pages`:自动检测到的图片页;`selection_mode` 表示自动或显式选页。
|
||||||
|
- `status: good` 且 `usable_for_summary: true`:可以把 `text` 补充到原生文档内容中。
|
||||||
|
- `status: no_image_text`:图片区域没有识别到额外文字,不是错误。
|
||||||
|
- `status: sparse` 或 `low_confidence`:不要使用返回文字;根据 `needs_review` 人工核验。
|
||||||
|
- `filtered_native_line_count` 和 `filtered_outside_image_line_count`:被当作原生文字或图片区域外文字过滤的 OCR 行数。
|
||||||
|
|
||||||
|
默认 260 DPI,可用 `--dpi 150-400` 调整。若 `has_more: true`:`next_offset > 0` 时传 `--pages <next_page> --start-offset <next_offset>`;`next_offset = 0` 时把 `remaining_pages` 作为下一次 `--pages`。普通小徽标和面积不足页面约 1.5% 的图片不会进入自动候选,但仍可用 `--pages` 显式识别。
|
||||||
|
|
||||||
## 创建文档
|
## 创建文档
|
||||||
|
|
||||||
调用 `scripts/create_document.py`:
|
调用 `scripts/create_document.py`:
|
||||||
|
|||||||
@ -1,4 +1,4 @@
|
|||||||
interface:
|
interface:
|
||||||
display_name: "Word 文档"
|
display_name: "Word 文档"
|
||||||
short_description: "安全下载、创建、编辑、校验并渲染专业 Word 文档"
|
short_description: "安全读取、编辑、图片 OCR、校验并渲染专业 Word 文档"
|
||||||
default_prompt: "使用 $docx 创建或处理本地文件或 HTTPS 链接中的 Word 文档,并完成内容与版式校验。"
|
default_prompt: "使用 $docx 读取或创建 Word 文档,按需识别图片文字,并完成内容与版式校验。"
|
||||||
|
|||||||
@ -333,6 +333,9 @@ def main() -> dict[str, Any]:
|
|||||||
"paragraph_count": len(all_paragraphs),
|
"paragraph_count": len(all_paragraphs),
|
||||||
"table_count": len(document.tables),
|
"table_count": len(document.tables),
|
||||||
"section_count": len(document.sections),
|
"section_count": len(document.sections),
|
||||||
|
"inline_image_count": len(document.inline_shapes),
|
||||||
|
"media_part_count": archive["media_count"],
|
||||||
|
"has_images": archive["media_count"] > 0,
|
||||||
"character_count": len(total_text),
|
"character_count": len(total_text),
|
||||||
"word_count_estimate": len(total_text.split()),
|
"word_count_estimate": len(total_text.split()),
|
||||||
"sections": sections,
|
"sections": sections,
|
||||||
|
|||||||
804
skills/docx/scripts/ocr_document.py
Normal file
804
skills/docx/scripts/ocr_document.py
Normal file
@ -0,0 +1,804 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import contextlib
|
||||||
|
import difflib
|
||||||
|
import importlib.metadata
|
||||||
|
import io
|
||||||
|
import logging
|
||||||
|
import os
|
||||||
|
import re
|
||||||
|
import shutil
|
||||||
|
import tempfile
|
||||||
|
import time
|
||||||
|
import unicodedata
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from _docx_common import (
|
||||||
|
WORD_INPUT_SUFFIXES,
|
||||||
|
SkillArgumentParser,
|
||||||
|
find_program,
|
||||||
|
input_file,
|
||||||
|
run_cli,
|
||||||
|
run_program,
|
||||||
|
run_soffice_convert,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
DEFAULT_DPI = 260
|
||||||
|
DEFAULT_MAX_CHARS = 24000
|
||||||
|
DEFAULT_TIMEOUT_SECONDS = 180
|
||||||
|
MAX_PAGES_PER_CALL = 4
|
||||||
|
MAX_PIXELS_PER_PAGE = 20_000_000
|
||||||
|
MIN_MEAN_CONFIDENCE = 0.60
|
||||||
|
MIN_MEANINGFUL_CHARS = 5
|
||||||
|
MIN_OCR_IMAGE_AREA_RATIO = 0.015
|
||||||
|
WHITESPACE_PATTERN = re.compile(r"[ \t]+")
|
||||||
|
|
||||||
|
|
||||||
|
for variable, value in (
|
||||||
|
("OMP_NUM_THREADS", "2"),
|
||||||
|
("OPENBLAS_NUM_THREADS", "1"),
|
||||||
|
("MKL_NUM_THREADS", "1"),
|
||||||
|
("NUMEXPR_NUM_THREADS", "1"),
|
||||||
|
):
|
||||||
|
os.environ.setdefault(variable, value)
|
||||||
|
|
||||||
|
for logger_name in ("rapidocr", "RapidOCR", "onnxruntime"):
|
||||||
|
logging.getLogger(logger_name).setLevel(logging.ERROR)
|
||||||
|
|
||||||
|
|
||||||
|
def build_parser():
|
||||||
|
parser = SkillArgumentParser(
|
||||||
|
description="渲染 Word 页面并用本地 OCR 提取图片中的文字。"
|
||||||
|
)
|
||||||
|
parser.add_argument("--input", required=True)
|
||||||
|
parser.add_argument(
|
||||||
|
"--pages",
|
||||||
|
help=(
|
||||||
|
"要识别的页码,例如 2 或 2,5-6;省略时自动选择含可读图片的页面"
|
||||||
|
),
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--start-offset",
|
||||||
|
type=int,
|
||||||
|
default=0,
|
||||||
|
help="续读单页图片文字时的字符偏移量",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--max-chars",
|
||||||
|
type=int,
|
||||||
|
default=DEFAULT_MAX_CHARS,
|
||||||
|
help=f"单次最多返回字符数,默认 {DEFAULT_MAX_CHARS}",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--dpi",
|
||||||
|
type=int,
|
||||||
|
default=DEFAULT_DPI,
|
||||||
|
help=f"OCR 渲染分辨率,默认 {DEFAULT_DPI} DPI",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--timeout",
|
||||||
|
type=int,
|
||||||
|
default=DEFAULT_TIMEOUT_SECONDS,
|
||||||
|
help=f"转换和单页渲染超时秒数,默认 {DEFAULT_TIMEOUT_SECONDS}",
|
||||||
|
)
|
||||||
|
return parser
|
||||||
|
|
||||||
|
|
||||||
|
def _parse_page_spec(value: str, page_count: int) -> list[int]:
|
||||||
|
if not value.strip():
|
||||||
|
raise ValueError("pages 不能为空")
|
||||||
|
pages: set[int] = set()
|
||||||
|
for raw_part in value.split(","):
|
||||||
|
part = raw_part.strip()
|
||||||
|
if not part:
|
||||||
|
continue
|
||||||
|
if "-" in part:
|
||||||
|
pieces = part.split("-", 1)
|
||||||
|
try:
|
||||||
|
start = int(pieces[0])
|
||||||
|
end = int(pieces[1])
|
||||||
|
except ValueError as exc:
|
||||||
|
raise ValueError(f"页码范围格式错误:{part}") from exc
|
||||||
|
if start > end:
|
||||||
|
raise ValueError(f"页码范围起始值不能大于结束值:{part}")
|
||||||
|
else:
|
||||||
|
try:
|
||||||
|
start = end = int(part)
|
||||||
|
except ValueError as exc:
|
||||||
|
raise ValueError(f"页码格式错误:{part}") from exc
|
||||||
|
if start < 1 or end > page_count:
|
||||||
|
raise ValueError(f"页码必须在 1 到 {page_count} 之间:{part}")
|
||||||
|
pages.update(range(start, end + 1))
|
||||||
|
if not pages:
|
||||||
|
raise ValueError("pages 不能为空")
|
||||||
|
return sorted(pages)
|
||||||
|
|
||||||
|
|
||||||
|
def _clean_text(value: Any) -> str:
|
||||||
|
text = str(value or "").replace("\x00", "").strip()
|
||||||
|
return "\n".join(
|
||||||
|
WHITESPACE_PATTERN.sub(" ", line).strip()
|
||||||
|
for line in text.replace("\r\n", "\n").replace("\r", "\n").split("\n")
|
||||||
|
if line.strip()
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _comparison_key(value: str) -> str:
|
||||||
|
normalized = unicodedata.normalize("NFKC", value).casefold()
|
||||||
|
return "".join(character for character in normalized if character.isalnum())
|
||||||
|
|
||||||
|
|
||||||
|
def _normalized_box(
|
||||||
|
x0: Any,
|
||||||
|
top: Any,
|
||||||
|
x1: Any,
|
||||||
|
bottom: Any,
|
||||||
|
page_width: float,
|
||||||
|
page_height: float,
|
||||||
|
) -> tuple[float, float, float, float] | None:
|
||||||
|
try:
|
||||||
|
left = max(0.0, min(page_width, float(x0)))
|
||||||
|
upper = max(0.0, min(page_height, float(top)))
|
||||||
|
right = max(0.0, min(page_width, float(x1)))
|
||||||
|
lower = max(0.0, min(page_height, float(bottom)))
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
return None
|
||||||
|
if (
|
||||||
|
page_width <= 0
|
||||||
|
or page_height <= 0
|
||||||
|
or right <= left
|
||||||
|
or lower <= upper
|
||||||
|
):
|
||||||
|
return None
|
||||||
|
return (
|
||||||
|
left / page_width,
|
||||||
|
upper / page_height,
|
||||||
|
right / page_width,
|
||||||
|
lower / page_height,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _group_native_words(
|
||||||
|
words: list[dict[str, Any]],
|
||||||
|
page_width: float,
|
||||||
|
page_height: float,
|
||||||
|
) -> list[dict[str, Any]]:
|
||||||
|
sorted_words = sorted(
|
||||||
|
words,
|
||||||
|
key=lambda word: (
|
||||||
|
round(float(word.get("top", 0.0)) / 2.5),
|
||||||
|
float(word.get("x0", 0.0)),
|
||||||
|
),
|
||||||
|
)
|
||||||
|
groups: list[list[dict[str, Any]]] = []
|
||||||
|
group_top: float | None = None
|
||||||
|
for word in sorted_words:
|
||||||
|
try:
|
||||||
|
top = float(word.get("top", 0.0))
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
continue
|
||||||
|
if not groups or group_top is None or abs(top - group_top) > 3.0:
|
||||||
|
groups.append([word])
|
||||||
|
group_top = top
|
||||||
|
else:
|
||||||
|
groups[-1].append(word)
|
||||||
|
group_top = sum(
|
||||||
|
float(item.get("top", 0.0)) for item in groups[-1]
|
||||||
|
) / len(groups[-1])
|
||||||
|
|
||||||
|
lines: list[dict[str, Any]] = []
|
||||||
|
for group in groups:
|
||||||
|
group.sort(key=lambda word: float(word.get("x0", 0.0)))
|
||||||
|
text = _clean_text(
|
||||||
|
" ".join(str(word.get("text", "")) for word in group)
|
||||||
|
)
|
||||||
|
key = _comparison_key(text)
|
||||||
|
if not key:
|
||||||
|
continue
|
||||||
|
box = _normalized_box(
|
||||||
|
min(float(word.get("x0", 0.0)) for word in group),
|
||||||
|
min(float(word.get("top", 0.0)) for word in group),
|
||||||
|
max(float(word.get("x1", 0.0)) for word in group),
|
||||||
|
max(float(word.get("bottom", 0.0)) for word in group),
|
||||||
|
page_width,
|
||||||
|
page_height,
|
||||||
|
)
|
||||||
|
if box is not None:
|
||||||
|
lines.append({"key": key, "box": box})
|
||||||
|
return lines
|
||||||
|
|
||||||
|
|
||||||
|
def _page_profiles(
|
||||||
|
pdf_path: Path,
|
||||||
|
) -> tuple[dict[int, dict[str, Any]], list[str]]:
|
||||||
|
import pdfplumber
|
||||||
|
|
||||||
|
profiles: dict[int, dict[str, Any]] = {}
|
||||||
|
warnings: list[str] = []
|
||||||
|
with pdfplumber.open(pdf_path) as pdf:
|
||||||
|
for page_number, page in enumerate(pdf.pages, start=1):
|
||||||
|
page_width = float(page.width)
|
||||||
|
page_height = float(page.height)
|
||||||
|
try:
|
||||||
|
words = page.extract_words(
|
||||||
|
x_tolerance=2,
|
||||||
|
y_tolerance=3,
|
||||||
|
keep_blank_chars=False,
|
||||||
|
use_text_flow=False,
|
||||||
|
)
|
||||||
|
except Exception as exc:
|
||||||
|
words = []
|
||||||
|
warnings.append(
|
||||||
|
f"第 {page_number} 页原生文本层读取失败:{exc}"
|
||||||
|
)
|
||||||
|
native_lines = _group_native_words(
|
||||||
|
list(words or []),
|
||||||
|
page_width,
|
||||||
|
page_height,
|
||||||
|
)
|
||||||
|
native_keys = [line["key"] for line in native_lines]
|
||||||
|
|
||||||
|
all_image_boxes: list[
|
||||||
|
tuple[float, float, float, float]
|
||||||
|
] = []
|
||||||
|
ocr_image_boxes: list[
|
||||||
|
tuple[float, float, float, float]
|
||||||
|
] = []
|
||||||
|
image_area = 0.0
|
||||||
|
image_count = 0
|
||||||
|
for image in page.images:
|
||||||
|
box = _normalized_box(
|
||||||
|
image.get("x0"),
|
||||||
|
image.get("top"),
|
||||||
|
image.get("x1"),
|
||||||
|
image.get("bottom"),
|
||||||
|
page_width,
|
||||||
|
page_height,
|
||||||
|
)
|
||||||
|
if box is None:
|
||||||
|
continue
|
||||||
|
image_count += 1
|
||||||
|
all_image_boxes.append(box)
|
||||||
|
area = (box[2] - box[0]) * (box[3] - box[1])
|
||||||
|
image_area += area
|
||||||
|
if area >= MIN_OCR_IMAGE_AREA_RATIO:
|
||||||
|
ocr_image_boxes.append(box)
|
||||||
|
|
||||||
|
profiles[page_number] = {
|
||||||
|
"image_count": image_count,
|
||||||
|
"ocr_image_count": len(ocr_image_boxes),
|
||||||
|
"image_area_ratio": round(min(1.0, image_area), 4),
|
||||||
|
"native_text_char_count": sum(
|
||||||
|
len(key) for key in native_keys
|
||||||
|
),
|
||||||
|
"_native_keys": native_keys,
|
||||||
|
"_native_boxes": native_lines,
|
||||||
|
"_image_boxes": all_image_boxes,
|
||||||
|
}
|
||||||
|
return profiles, warnings
|
||||||
|
|
||||||
|
|
||||||
|
def _box_points(value: Any) -> list[list[float]] | None:
|
||||||
|
if value is None:
|
||||||
|
return None
|
||||||
|
try:
|
||||||
|
points = [
|
||||||
|
[round(float(point[0]), 2), round(float(point[1]), 2)]
|
||||||
|
for point in value
|
||||||
|
]
|
||||||
|
except (IndexError, TypeError, ValueError):
|
||||||
|
return None
|
||||||
|
return points if len(points) == 4 else None
|
||||||
|
|
||||||
|
|
||||||
|
def _ordered_lines(result: Any) -> list[dict[str, Any]]:
|
||||||
|
texts = list(getattr(result, "txts", None) or ())
|
||||||
|
scores = list(getattr(result, "scores", None) or ())
|
||||||
|
raw_boxes = getattr(result, "boxes", None)
|
||||||
|
boxes = list(raw_boxes) if raw_boxes is not None else []
|
||||||
|
|
||||||
|
lines: list[dict[str, Any]] = []
|
||||||
|
for index, raw_text in enumerate(texts):
|
||||||
|
text = _clean_text(raw_text)
|
||||||
|
if not text:
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
confidence = float(scores[index])
|
||||||
|
except (IndexError, TypeError, ValueError):
|
||||||
|
confidence = 0.0
|
||||||
|
confidence = max(0.0, min(1.0, confidence))
|
||||||
|
box = _box_points(boxes[index] if index < len(boxes) else None)
|
||||||
|
if box:
|
||||||
|
left = min(point[0] for point in box)
|
||||||
|
top = min(point[1] for point in box)
|
||||||
|
else:
|
||||||
|
left = float(index)
|
||||||
|
top = float(index)
|
||||||
|
lines.append(
|
||||||
|
{
|
||||||
|
"text": text,
|
||||||
|
"confidence": confidence,
|
||||||
|
"box": box,
|
||||||
|
"_left": left,
|
||||||
|
"_top": top,
|
||||||
|
"_index": index,
|
||||||
|
}
|
||||||
|
)
|
||||||
|
lines.sort(
|
||||||
|
key=lambda line: (
|
||||||
|
round(line["_top"] / 10.0),
|
||||||
|
line["_left"],
|
||||||
|
line["_index"],
|
||||||
|
)
|
||||||
|
)
|
||||||
|
return lines
|
||||||
|
|
||||||
|
|
||||||
|
def _similar_to_any(
|
||||||
|
candidate: str,
|
||||||
|
references: list[str],
|
||||||
|
*,
|
||||||
|
threshold: float,
|
||||||
|
) -> bool:
|
||||||
|
if not candidate:
|
||||||
|
return False
|
||||||
|
for reference in references:
|
||||||
|
if not reference:
|
||||||
|
continue
|
||||||
|
if candidate == reference:
|
||||||
|
return True
|
||||||
|
shorter = min(len(candidate), len(reference))
|
||||||
|
longer = max(len(candidate), len(reference))
|
||||||
|
if shorter >= 3 and candidate in reference:
|
||||||
|
return True
|
||||||
|
if (
|
||||||
|
shorter >= 3
|
||||||
|
and reference in candidate
|
||||||
|
and longer <= round(shorter * 1.25)
|
||||||
|
):
|
||||||
|
return True
|
||||||
|
if shorter >= 3 and difflib.SequenceMatcher(
|
||||||
|
None,
|
||||||
|
candidate,
|
||||||
|
reference,
|
||||||
|
).ratio() >= threshold:
|
||||||
|
return True
|
||||||
|
return False
|
||||||
|
|
||||||
|
|
||||||
|
def _line_center(
|
||||||
|
box: list[list[float]] | None,
|
||||||
|
image_width: int,
|
||||||
|
image_height: int,
|
||||||
|
) -> tuple[float, float] | None:
|
||||||
|
if not box or image_width <= 0 or image_height <= 0:
|
||||||
|
return None
|
||||||
|
return (
|
||||||
|
sum(point[0] for point in box) / len(box) / image_width,
|
||||||
|
sum(point[1] for point in box) / len(box) / image_height,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _point_in_boxes(
|
||||||
|
point: tuple[float, float] | None,
|
||||||
|
boxes: list[tuple[float, float, float, float]],
|
||||||
|
*,
|
||||||
|
padding: float = 0.005,
|
||||||
|
) -> bool:
|
||||||
|
if point is None:
|
||||||
|
return False
|
||||||
|
x, y = point
|
||||||
|
return any(
|
||||||
|
x0 - padding <= x <= x1 + padding
|
||||||
|
and y0 - padding <= y <= y1 + padding
|
||||||
|
for x0, y0, x1, y1 in boxes
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _line_is_native(
|
||||||
|
line: dict[str, Any],
|
||||||
|
profile: dict[str, Any],
|
||||||
|
image_width: int,
|
||||||
|
image_height: int,
|
||||||
|
) -> bool:
|
||||||
|
candidate = _comparison_key(line["text"])
|
||||||
|
if _similar_to_any(
|
||||||
|
candidate,
|
||||||
|
profile["_native_keys"],
|
||||||
|
threshold=0.82,
|
||||||
|
):
|
||||||
|
return True
|
||||||
|
|
||||||
|
center = _line_center(line["box"], image_width, image_height)
|
||||||
|
if center is None:
|
||||||
|
return False
|
||||||
|
x, y = center
|
||||||
|
padding = 0.01
|
||||||
|
for native_line in profile["_native_boxes"]:
|
||||||
|
x0, y0, x1, y1 = native_line["box"]
|
||||||
|
if (
|
||||||
|
x0 - padding <= x <= x1 + padding
|
||||||
|
and y0 - padding <= y <= y1 + padding
|
||||||
|
and _similar_to_any(
|
||||||
|
candidate,
|
||||||
|
[native_line["key"]],
|
||||||
|
threshold=0.68,
|
||||||
|
)
|
||||||
|
):
|
||||||
|
return True
|
||||||
|
return False
|
||||||
|
|
||||||
|
|
||||||
|
def _create_ocr_engine():
|
||||||
|
try:
|
||||||
|
from rapidocr import RapidOCR
|
||||||
|
except ImportError as exc:
|
||||||
|
raise RuntimeError("环境预置的 rapidocr 模块不可用") from exc
|
||||||
|
|
||||||
|
captured_stdout = io.StringIO()
|
||||||
|
captured_stderr = io.StringIO()
|
||||||
|
with (
|
||||||
|
contextlib.redirect_stdout(captured_stdout),
|
||||||
|
contextlib.redirect_stderr(captured_stderr),
|
||||||
|
):
|
||||||
|
return RapidOCR()
|
||||||
|
|
||||||
|
|
||||||
|
def _ocr_page(
|
||||||
|
engine: Any,
|
||||||
|
image_path: Path,
|
||||||
|
profile: dict[str, Any],
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
from PIL import Image
|
||||||
|
|
||||||
|
with Image.open(image_path) as image:
|
||||||
|
image_width, image_height = image.size
|
||||||
|
|
||||||
|
captured_stdout = io.StringIO()
|
||||||
|
captured_stderr = io.StringIO()
|
||||||
|
started = time.monotonic()
|
||||||
|
with (
|
||||||
|
contextlib.redirect_stdout(captured_stdout),
|
||||||
|
contextlib.redirect_stderr(captured_stderr),
|
||||||
|
):
|
||||||
|
result = engine(str(image_path))
|
||||||
|
elapsed = time.monotonic() - started
|
||||||
|
|
||||||
|
raw_lines = _ordered_lines(result)
|
||||||
|
image_lines: list[dict[str, Any]] = []
|
||||||
|
seen: set[str] = set()
|
||||||
|
filtered_native = 0
|
||||||
|
filtered_outside_images = 0
|
||||||
|
filtered_duplicates = 0
|
||||||
|
image_boxes = profile["_image_boxes"]
|
||||||
|
for line in raw_lines:
|
||||||
|
if _line_is_native(line, profile, image_width, image_height):
|
||||||
|
filtered_native += 1
|
||||||
|
continue
|
||||||
|
center = _line_center(line["box"], image_width, image_height)
|
||||||
|
if image_boxes and not _point_in_boxes(center, image_boxes):
|
||||||
|
filtered_outside_images += 1
|
||||||
|
continue
|
||||||
|
key = _comparison_key(line["text"])
|
||||||
|
if key and key in seen:
|
||||||
|
filtered_duplicates += 1
|
||||||
|
continue
|
||||||
|
if key:
|
||||||
|
seen.add(key)
|
||||||
|
image_lines.append(line)
|
||||||
|
|
||||||
|
text = "\n".join(line["text"] for line in image_lines)
|
||||||
|
weighted_chars = [
|
||||||
|
max(1, sum(1 for character in line["text"] if not character.isspace()))
|
||||||
|
for line in image_lines
|
||||||
|
]
|
||||||
|
total_weight = sum(weighted_chars)
|
||||||
|
mean_confidence = (
|
||||||
|
sum(
|
||||||
|
line["confidence"] * weight
|
||||||
|
for line, weight in zip(image_lines, weighted_chars)
|
||||||
|
)
|
||||||
|
/ total_weight
|
||||||
|
if total_weight
|
||||||
|
else 0.0
|
||||||
|
)
|
||||||
|
meaningful_chars = sum(1 for character in text if character.isalnum())
|
||||||
|
low_confidence_lines = sum(
|
||||||
|
1
|
||||||
|
for line in image_lines
|
||||||
|
if line["confidence"] < MIN_MEAN_CONFIDENCE
|
||||||
|
)
|
||||||
|
|
||||||
|
reasons: list[str] = []
|
||||||
|
if not text:
|
||||||
|
status = "no_image_text"
|
||||||
|
reasons.append("未识别到原生文本之外的图片文字")
|
||||||
|
elif meaningful_chars < MIN_MEANINGFUL_CHARS:
|
||||||
|
status = "sparse"
|
||||||
|
reasons.append(
|
||||||
|
f"图片中的有效文字少于 {MIN_MEANINGFUL_CHARS} 个字符"
|
||||||
|
)
|
||||||
|
elif mean_confidence < MIN_MEAN_CONFIDENCE:
|
||||||
|
status = "low_confidence"
|
||||||
|
reasons.append(
|
||||||
|
"图片文字 OCR 平均置信度低于 "
|
||||||
|
f"{round(MIN_MEAN_CONFIDENCE * 100)}%"
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
status = "good"
|
||||||
|
|
||||||
|
return {
|
||||||
|
"text": text,
|
||||||
|
"status": status,
|
||||||
|
"usable_for_summary": status == "good",
|
||||||
|
"needs_review": status in {"sparse", "low_confidence"},
|
||||||
|
"raw_ocr_line_count": len(raw_lines),
|
||||||
|
"image_line_count": len(image_lines),
|
||||||
|
"filtered_native_line_count": filtered_native,
|
||||||
|
"filtered_outside_image_line_count": filtered_outside_images,
|
||||||
|
"filtered_duplicate_line_count": filtered_duplicates,
|
||||||
|
"low_confidence_line_count": low_confidence_lines,
|
||||||
|
"mean_confidence": round(mean_confidence, 4),
|
||||||
|
"meaningful_chars": meaningful_chars,
|
||||||
|
"reasons": reasons,
|
||||||
|
"ocr_seconds": round(elapsed, 3),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _pdf_pages(path: Path) -> tuple[int, dict[int, tuple[float, float]]]:
|
||||||
|
from pypdf import PdfReader
|
||||||
|
|
||||||
|
page_sizes: dict[int, tuple[float, float]] = {}
|
||||||
|
with path.open("rb") as stream:
|
||||||
|
reader = PdfReader(stream, strict=False)
|
||||||
|
if reader.is_encrypted:
|
||||||
|
raise ValueError("LibreOffice 生成了加密 PDF,无法执行 OCR")
|
||||||
|
page_count = len(reader.pages)
|
||||||
|
for page_number, page in enumerate(reader.pages, start=1):
|
||||||
|
page_sizes[page_number] = (
|
||||||
|
abs(float(page.cropbox.width)),
|
||||||
|
abs(float(page.cropbox.height)),
|
||||||
|
)
|
||||||
|
return page_count, page_sizes
|
||||||
|
|
||||||
|
|
||||||
|
def _render_page(
|
||||||
|
pdf_path: Path,
|
||||||
|
page_number: int,
|
||||||
|
page_size: tuple[float, float],
|
||||||
|
dpi: int,
|
||||||
|
timeout: int,
|
||||||
|
temp_dir: Path,
|
||||||
|
) -> tuple[Path, float]:
|
||||||
|
width_points, height_points = page_size
|
||||||
|
estimated_pixels = (
|
||||||
|
width_points * dpi / 72.0
|
||||||
|
* height_points * dpi / 72.0
|
||||||
|
)
|
||||||
|
if estimated_pixels > MAX_PIXELS_PER_PAGE:
|
||||||
|
raise ValueError(
|
||||||
|
f"第 {page_number} 页按 {dpi} DPI 渲染预计超过 "
|
||||||
|
f"{MAX_PIXELS_PER_PAGE} 像素,请降低 dpi"
|
||||||
|
)
|
||||||
|
|
||||||
|
prefix = temp_dir / f"page-{page_number:04d}"
|
||||||
|
output = prefix.with_suffix(".png")
|
||||||
|
started = time.monotonic()
|
||||||
|
run_program(
|
||||||
|
[
|
||||||
|
find_program("pdftoppm"),
|
||||||
|
"-f",
|
||||||
|
str(page_number),
|
||||||
|
"-l",
|
||||||
|
str(page_number),
|
||||||
|
"-singlefile",
|
||||||
|
"-png",
|
||||||
|
"-r",
|
||||||
|
str(dpi),
|
||||||
|
str(pdf_path),
|
||||||
|
str(prefix),
|
||||||
|
],
|
||||||
|
timeout=timeout,
|
||||||
|
)
|
||||||
|
elapsed = time.monotonic() - started
|
||||||
|
if not output.is_file() or output.stat().st_size <= 0:
|
||||||
|
raise RuntimeError(f"第 {page_number} 页没有生成有效 PNG")
|
||||||
|
return output, elapsed
|
||||||
|
|
||||||
|
|
||||||
|
def _package_version(name: str) -> str | None:
|
||||||
|
try:
|
||||||
|
return importlib.metadata.version(name)
|
||||||
|
except importlib.metadata.PackageNotFoundError:
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> dict[str, Any]:
|
||||||
|
args = build_parser().parse_args()
|
||||||
|
if args.start_offset < 0:
|
||||||
|
raise ValueError("start-offset 不能小于 0")
|
||||||
|
if args.max_chars < 1 or args.max_chars > 60000:
|
||||||
|
raise ValueError("max-chars 必须在 1 到 60000 之间")
|
||||||
|
if args.dpi < 150 or args.dpi > 400:
|
||||||
|
raise ValueError("dpi 必须在 150 到 400 之间")
|
||||||
|
if args.timeout < 1 or args.timeout > 600:
|
||||||
|
raise ValueError("timeout 必须在 1 到 600 秒之间")
|
||||||
|
|
||||||
|
source = input_file(args.input, WORD_INPUT_SUFFIXES)
|
||||||
|
page_outputs: list[dict[str, Any]] = []
|
||||||
|
returned_chars = 0
|
||||||
|
next_page: int | None = None
|
||||||
|
next_offset = 0
|
||||||
|
remaining_pages: list[int] = []
|
||||||
|
office_output = {"stdout": "", "stderr": ""}
|
||||||
|
|
||||||
|
with tempfile.TemporaryDirectory(prefix="docx-ocr-") as temp_name:
|
||||||
|
temp_dir = Path(temp_name)
|
||||||
|
staged_input = temp_dir / f"document{source.suffix.lower()}"
|
||||||
|
shutil.copy2(source, staged_input)
|
||||||
|
pdf_path, office_output = run_soffice_convert(
|
||||||
|
staged_input,
|
||||||
|
target_format="pdf",
|
||||||
|
output_dir=temp_dir / "pdf",
|
||||||
|
timeout=args.timeout,
|
||||||
|
)
|
||||||
|
page_count, page_sizes = _pdf_pages(pdf_path)
|
||||||
|
if page_count < 1:
|
||||||
|
raise ValueError("Word 文档没有可执行 OCR 的页面")
|
||||||
|
profiles, profile_warnings = _page_profiles(pdf_path)
|
||||||
|
if len(profiles) != page_count:
|
||||||
|
raise RuntimeError(
|
||||||
|
f"渲染结果有 {page_count} 页,但只检查到 "
|
||||||
|
f"{len(profiles)} 页"
|
||||||
|
)
|
||||||
|
candidate_pages = [
|
||||||
|
page_number
|
||||||
|
for page_number, profile in profiles.items()
|
||||||
|
if profile["ocr_image_count"] > 0
|
||||||
|
]
|
||||||
|
|
||||||
|
selection_mode = "explicit" if args.pages else "auto"
|
||||||
|
if args.pages:
|
||||||
|
target_pages = _parse_page_spec(args.pages, page_count)
|
||||||
|
if len(target_pages) > MAX_PAGES_PER_CALL:
|
||||||
|
raise ValueError(
|
||||||
|
f"单次最多 OCR {MAX_PAGES_PER_CALL} 页,"
|
||||||
|
"请拆分 pages 后重试"
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
target_pages = candidate_pages
|
||||||
|
|
||||||
|
if args.start_offset > 0:
|
||||||
|
if not args.pages or len(target_pages) != 1:
|
||||||
|
raise ValueError(
|
||||||
|
"使用 start-offset 时必须显式指定且只指定一页"
|
||||||
|
)
|
||||||
|
|
||||||
|
selected_pages = target_pages[:MAX_PAGES_PER_CALL]
|
||||||
|
queued_pages = target_pages[MAX_PAGES_PER_CALL:]
|
||||||
|
if selected_pages:
|
||||||
|
engine = _create_ocr_engine()
|
||||||
|
for index, page_number in enumerate(selected_pages):
|
||||||
|
budget = args.max_chars - returned_chars
|
||||||
|
if budget <= 0:
|
||||||
|
next_page = page_number
|
||||||
|
remaining_pages = (
|
||||||
|
selected_pages[index:] + queued_pages
|
||||||
|
)
|
||||||
|
break
|
||||||
|
|
||||||
|
image_path, render_seconds = _render_page(
|
||||||
|
pdf_path,
|
||||||
|
page_number,
|
||||||
|
page_sizes[page_number],
|
||||||
|
args.dpi,
|
||||||
|
args.timeout,
|
||||||
|
temp_dir,
|
||||||
|
)
|
||||||
|
result = _ocr_page(
|
||||||
|
engine,
|
||||||
|
image_path,
|
||||||
|
profiles[page_number],
|
||||||
|
)
|
||||||
|
full_text = result.pop("text")
|
||||||
|
offset = args.start_offset if index == 0 else 0
|
||||||
|
if offset > len(full_text):
|
||||||
|
raise ValueError(
|
||||||
|
f"start-offset 超过第 {page_number} 页图片文字长度 "
|
||||||
|
f"{len(full_text)}"
|
||||||
|
)
|
||||||
|
|
||||||
|
usable = bool(result["usable_for_summary"])
|
||||||
|
if not usable:
|
||||||
|
page_text = ""
|
||||||
|
complete = True
|
||||||
|
else:
|
||||||
|
remaining_text = full_text[offset:]
|
||||||
|
page_text = remaining_text[:budget]
|
||||||
|
complete = len(page_text) == len(remaining_text)
|
||||||
|
|
||||||
|
profile = profiles[page_number]
|
||||||
|
page_outputs.append(
|
||||||
|
{
|
||||||
|
"page": page_number,
|
||||||
|
"text": page_text,
|
||||||
|
"char_count": len(full_text),
|
||||||
|
"offset_start": offset if usable else 0,
|
||||||
|
"offset_end": (
|
||||||
|
offset + len(page_text) if usable else 0
|
||||||
|
),
|
||||||
|
"complete": complete,
|
||||||
|
"render_seconds": round(render_seconds, 3),
|
||||||
|
"image_count": profile["image_count"],
|
||||||
|
"ocr_image_count": profile["ocr_image_count"],
|
||||||
|
"image_area_ratio": profile[
|
||||||
|
"image_area_ratio"
|
||||||
|
],
|
||||||
|
"native_text_char_count": profile[
|
||||||
|
"native_text_char_count"
|
||||||
|
],
|
||||||
|
**result,
|
||||||
|
}
|
||||||
|
)
|
||||||
|
returned_chars += len(page_text)
|
||||||
|
|
||||||
|
if not complete:
|
||||||
|
next_page = page_number
|
||||||
|
next_offset = offset + len(page_text)
|
||||||
|
remaining_pages = (
|
||||||
|
selected_pages[index + 1 :] + queued_pages
|
||||||
|
)
|
||||||
|
break
|
||||||
|
|
||||||
|
if next_page is None and queued_pages:
|
||||||
|
next_page = queued_pages[0]
|
||||||
|
remaining_pages = queued_pages
|
||||||
|
|
||||||
|
all_processed = len(page_outputs) == len(selected_pages)
|
||||||
|
all_complete = all(page["complete"] for page in page_outputs)
|
||||||
|
all_safe = all(
|
||||||
|
page["status"] in {"good", "no_image_text"}
|
||||||
|
for page in page_outputs
|
||||||
|
)
|
||||||
|
has_more = next_page is not None
|
||||||
|
return {
|
||||||
|
"source": str(source),
|
||||||
|
"page_count": page_count,
|
||||||
|
"selection_mode": selection_mode,
|
||||||
|
"candidate_pages": candidate_pages,
|
||||||
|
"selected_pages": selected_pages,
|
||||||
|
"processed_pages": [page["page"] for page in page_outputs],
|
||||||
|
"engine": "rapidocr",
|
||||||
|
"engine_version": _package_version("rapidocr"),
|
||||||
|
"runtime": "onnxruntime",
|
||||||
|
"runtime_version": _package_version("onnxruntime"),
|
||||||
|
"offline": True,
|
||||||
|
"dpi": args.dpi,
|
||||||
|
"returned_chars": returned_chars,
|
||||||
|
"pages": page_outputs,
|
||||||
|
"usable_for_summary": any(
|
||||||
|
page["usable_for_summary"] for page in page_outputs
|
||||||
|
),
|
||||||
|
"complete_ocr_coverage": (
|
||||||
|
not has_more
|
||||||
|
and all_processed
|
||||||
|
and all_complete
|
||||||
|
and all_safe
|
||||||
|
),
|
||||||
|
"needs_review": any(page["needs_review"] for page in page_outputs),
|
||||||
|
"has_more": has_more,
|
||||||
|
"next_page": next_page,
|
||||||
|
"next_offset": next_offset,
|
||||||
|
"remaining_pages": remaining_pages,
|
||||||
|
"profile_warnings": profile_warnings,
|
||||||
|
"office_stdout": office_output["stdout"],
|
||||||
|
"office_stderr": office_output["stderr"],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(run_cli(main))
|
||||||
419
skills/pptx/SKILL.md
Normal file
419
skills/pptx/SKILL.md
Normal file
@ -0,0 +1,419 @@
|
|||||||
|
---
|
||||||
|
name: pptx
|
||||||
|
description: "创建、读取、编辑、复制页面、转换、校验和渲染本地或远程 HTTPS PowerPoint 演示文稿与模板,并按需识别页面图片、截图和图表中的文字。用户提到 PPT、PPTX、PowerPoint、演示文稿、幻灯片、路演稿、汇报材料、演讲者备注、模板、版式或图片文字 OCR,或提供 HTTPS 演示文稿地址、.pptx、.potx、.ppsx、.ppt 文件时使用;支持安全下载、结构化创建、保留 Run 格式的文本替换、页面删除/重排/复制、Markdown 提取、本地 RapidOCR、旧格式转换、受控 OOXML 解包/打包、关系与图表校验及逐页视觉检查。若主要交付物不是演示文稿且不需要读取或修改 PPT 内容,则不要使用。"
|
||||||
|
---
|
||||||
|
|
||||||
|
# PowerPoint 演示文稿处理
|
||||||
|
|
||||||
|
## 强制执行规则
|
||||||
|
|
||||||
|
当前智能体不能直接执行 shell、任意 Python、Node.js 或系统命令。只能通过 `execute_skill_script` 调用本 Skill 中真实存在的固定 Python 脚本。
|
||||||
|
|
||||||
|
- 只调用下表列出的可执行脚本;不执行 `scripts/` 目录、`scripts/_pptx_common.py`、`scripts/_presentation_builder.js` 或 `scripts/_icon_renderer.js`。
|
||||||
|
- 不把 `python3`、`node`、`soffice`、`libreoffice`、`pdftoppm`、`zip`、`unzip`、`rm` 或其他系统命令作为脚本参数。
|
||||||
|
- PptxGenJS、React Icons、Sharp、LibreOffice、Poppler 和 ZIP 操作只允许由固定脚本在内部调用。
|
||||||
|
- 每次检查脚本返回的 JSON;只有 `ok` 为 `true` 时才继续。`validate_presentation.py` 还必须返回 `status: valid`、`issue_count: 0`。
|
||||||
|
- 只在需要读取图片、截图或视觉图表中的文字时调用 `ocr_presentation.py`。只使用 `slides[]` 中 `usable_for_summary: true` 的 `text`;低置信度结果不得作为可靠正文。
|
||||||
|
- 不覆盖用户提供的源文件。最终结果写入 `output/pptx/`,中间产物写入 `tmp/pptx/<任务名>/`。
|
||||||
|
- 远程地址只交给 `download_presentation.py`;不要在回复、日志摘要或文件名中复述可能含敏感查询参数的完整 URL。
|
||||||
|
- 环境已预置全部依赖,不安装软件包,也不提示用户安装依赖。
|
||||||
|
|
||||||
|
## 脚本清单
|
||||||
|
|
||||||
|
| 脚本 | 用途 | 底层能力 |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| `scripts/download_presentation.py` | 下载并校验远程 HTTPS 演示文稿 | Python `urllib`、安全 OOXML 解析、`python-pptx` |
|
||||||
|
| `scripts/inspect_presentation.py` | 分段读取页面、文本、表格、图表、图片和备注 | `python-pptx`、安全 OOXML 解析 |
|
||||||
|
| `scripts/ocr_presentation.py` | 按页识别图片、截图和视觉图表中的文字 | RapidOCR、ONNX Runtime、LibreOffice、Poppler |
|
||||||
|
| `scripts/extract_presentation.py` | 提取整份演示文稿为 Markdown | `markitdown[pptx]` |
|
||||||
|
| `scripts/create_presentation.py` | 按受控 JSON 创建专业 PPTX | PptxGenJS |
|
||||||
|
| `scripts/edit_presentation.py` | 文本替换、删除/重排页面、修改属性 | `python-pptx` |
|
||||||
|
| `scripts/duplicate_slide.py` | 复制现有 PPTX 页面并维护包关系 | `lxml`、Python `zipfile` |
|
||||||
|
| `scripts/render_icon.py` | 把允许的 React Icons 图标渲染为 PNG | React、React DOM、React Icons、Sharp |
|
||||||
|
| `scripts/convert_presentation.py` | `.ppt/.potx/.ppsx/.pptx` 转 PPTX 或 PDF | LibreOffice |
|
||||||
|
| `scripts/unpack_presentation.py` | 安全解包 OOXML 供高级编辑 | Python `zipfile` |
|
||||||
|
| `scripts/pack_presentation.py` | 把 OOXML 目录安全打包为演示文稿 | Python `zipfile`、`python-pptx` |
|
||||||
|
| `scripts/validate_presentation.py` | 校验 ZIP、XML、关系、页面、图表、边界和可渲染性 | `defusedxml`、`lxml`、`python-pptx`、LibreOffice |
|
||||||
|
| `scripts/render_presentation.py` | 把演示文稿渲染为逐页 PNG、联系表和 PDF | LibreOffice、Poppler、Pillow |
|
||||||
|
|
||||||
|
## 标准流程
|
||||||
|
|
||||||
|
1. 输入是 HTTPS 地址时,先调用 `download_presentation.py` 下载到本次任务临时目录;本地文件直接进入下一步。
|
||||||
|
2. 旧版 `.ppt` 先调用 `convert_presentation.py` 转为 `.pptx`。需要把 `.potx/.ppsx` 当普通演示文稿编辑时,也先转为 `.pptx`。
|
||||||
|
3. 编辑、总结或套用模板前先调用 `inspect_presentation.py`;内容较多时再调用 `extract_presentation.py` 获取完整 Markdown。
|
||||||
|
4. 需要读取截图、扫描页或图片中的文字时,从 `image_slides` 选择相关页调用 `ocr_presentation.py`。不要默认 OCR 全部页面,也不要用 OCR 覆盖可靠的原生文本。
|
||||||
|
5. 从零创建使用 `create_presentation.py`;常规文本与页面编辑使用 `edit_presentation.py`;复用模板页面使用 `duplicate_slide.py`。
|
||||||
|
6. 只有固定脚本不能完成的精细模板编辑,才使用 `unpack_presentation.py` → 最小化编辑 OOXML → `pack_presentation.py`。
|
||||||
|
7. 创建或修改后必须调用 `validate_presentation.py --check-render`。基于模板制作时同时传 `--original <原模板>`。
|
||||||
|
8. 再调用 `render_presentation.py --contact-sheet` 渲染全部页面;若 `has_more: true`,用 `next_slide` 继续,直到检查完整份演示文稿。
|
||||||
|
9. 逐页检查内容、顺序、溢出、重叠、间距、对齐、对比度、占位文本和备注。发现问题后修复并重新校验、重新渲染受影响页面。
|
||||||
|
10. 文件结构、内容和视觉检查全部通过后才交付。
|
||||||
|
|
||||||
|
## 下载远程演示文稿
|
||||||
|
|
||||||
|
只接受 HTTPS 地址。完整保留 URL 及查询参数传给脚本,但不要在回复或输出文件名中暴露查询参数。
|
||||||
|
|
||||||
|
```text
|
||||||
|
--url 'https://example.com/deck.pptx?signature=...' --output 'tmp/pptx/<任务名>/source.pptx'
|
||||||
|
```
|
||||||
|
|
||||||
|
可选参数:
|
||||||
|
|
||||||
|
- `--timeout <1-600>`:连接和读取超时秒数,默认 `60`。
|
||||||
|
- `--max-bytes <字节数>`:默认 100 MiB,最高 512 MiB。
|
||||||
|
- `--overwrite`:只覆盖本次任务生成的旧缓存。
|
||||||
|
|
||||||
|
`output` 扩展名必须与远程内容的真实格式一致,支持 `.pptx/.potx/.ppsx/.ppt`。脚本限制重定向只能继续使用 HTTPS,流式限制大小,先写临时文件,再原子发布。
|
||||||
|
|
||||||
|
## 检查和提取内容
|
||||||
|
|
||||||
|
调用结构检查:
|
||||||
|
|
||||||
|
```text
|
||||||
|
--input 'source.pptx' --start-slide 1 --max-slides 30
|
||||||
|
```
|
||||||
|
|
||||||
|
常用参数:
|
||||||
|
|
||||||
|
- `--include-runs`:需要检查局部字体、粗体、字号或跨 Run 文本替换时使用。
|
||||||
|
- `--max-shapes <1-1000>`、`--max-table-cells <1-10000>`:限制单次结构输出。
|
||||||
|
- `--max-chars <1000-1000000>`:限制 JSON 中返回的正文量。
|
||||||
|
|
||||||
|
重点检查 `slide_size`、`slides[].layout`、形状边界、表格、图表系列、图片类型、`slides[].media`、`image_slides`、`speaker_notes`、批注摘要和 `external_relationships`。`media.image_count` 表示页面中的图片数量,`media.image_area_ratio` 是图片大致占页比例,`media.native_text_char_count` 是可直接提取的原生文字量。长演示文稿根据 `selection.has_more` 和 `next_slide` 分段读取。
|
||||||
|
|
||||||
|
需要连续正文时调用:
|
||||||
|
|
||||||
|
```text
|
||||||
|
--input 'source.pptx' --output 'tmp/pptx/<任务名>/content.md'
|
||||||
|
```
|
||||||
|
|
||||||
|
Markdown 适合检查遗漏、错字和顺序,不代表页面版式。
|
||||||
|
|
||||||
|
## 识别图片中的文字
|
||||||
|
|
||||||
|
只有图片、截图、扫描页或视觉图表中的文字对任务有意义时才调用:
|
||||||
|
|
||||||
|
```text
|
||||||
|
--input 'source.pptx' --slides '2,5-6'
|
||||||
|
```
|
||||||
|
|
||||||
|
`slides` 必须明确指定页码,单次最多 4 页。脚本先通过 LibreOffice 和 Poppler 临时渲染选定页面,再使用本地 RapidOCR 识别;渲染图片会自动删除,不联网,也不调用大模型识图。
|
||||||
|
|
||||||
|
脚本会根据原生文本框的位置和文字相似度过滤重复内容,因此 `slides[].text` 只返回原生文本之外的可靠图片文字。检查:
|
||||||
|
|
||||||
|
- `status: good` 且 `usable_for_summary: true`:可以把 `text` 补充到同页原生文本中。
|
||||||
|
- `status: no_image_text`:未发现额外图片文字,不是错误。
|
||||||
|
- `status: sparse` 或 `low_confidence`:不要使用返回文字;根据 `needs_review` 人工核验。
|
||||||
|
- `filtered_native_line_count`:被识别为原生文本并去重的 OCR 行数。
|
||||||
|
- `picture_count`、`chart_count` 和 `image_area_ratio`:用于理解本页视觉内容规模。
|
||||||
|
|
||||||
|
默认 260 DPI。小字可使用 `--dpi 300-400`;单页预计像素过大时降低 DPI。可用 `--max-chars` 控制输出;若 `has_more: true`,根据 `next_slide`、`next_offset` 和 `remaining_slides` 继续。只有 `next_offset > 0` 时才传 `--start-offset`,且此时 `slides` 只能包含该页。
|
||||||
|
|
||||||
|
## 创建演示文稿
|
||||||
|
|
||||||
|
调用:
|
||||||
|
|
||||||
|
```text
|
||||||
|
--output 'output/pptx/result.pptx' --spec '<JSON对象>'
|
||||||
|
```
|
||||||
|
|
||||||
|
内容较长时先把 JSON 写入任务临时目录,再传 `--spec-file`。目标是本次任务旧产物且确认可覆盖时才传 `--overwrite`。
|
||||||
|
|
||||||
|
顶层结构:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"layout": "LAYOUT_WIDE",
|
||||||
|
"properties": {
|
||||||
|
"title": "2026 年产品路线图",
|
||||||
|
"author": "示例公司",
|
||||||
|
"subject": "产品规划"
|
||||||
|
},
|
||||||
|
"theme": {
|
||||||
|
"head_font": "Noto Sans CJK SC",
|
||||||
|
"body_font": "Noto Sans CJK SC",
|
||||||
|
"language": "zh-CN"
|
||||||
|
},
|
||||||
|
"slides": []
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
支持布局:
|
||||||
|
|
||||||
|
| `layout` | 画布尺寸 |
|
||||||
|
| --- | --- |
|
||||||
|
| `LAYOUT_WIDE` | 13.333 × 7.5 英寸 |
|
||||||
|
| `LAYOUT_16X9` | 10 × 5.625 英寸 |
|
||||||
|
| `LAYOUT_4X3` | 10 × 7.5 英寸 |
|
||||||
|
|
||||||
|
每页使用:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"background": "0F172A",
|
||||||
|
"speaker_notes": "本页讲解约 45 秒。",
|
||||||
|
"elements": []
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
坐标和尺寸 `x/y/w/h` 均使用英寸。颜色必须是不带 `#` 的 6 位十六进制值;不要把透明度拼入 8 位颜色,透明度使用 `transparency: 0-100`。所有可见元素必须位于画布内。
|
||||||
|
|
||||||
|
### 文本
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"type": "text",
|
||||||
|
"text": "从洞察到增长",
|
||||||
|
"options": {
|
||||||
|
"x": 0.7,
|
||||||
|
"y": 0.6,
|
||||||
|
"w": 8.8,
|
||||||
|
"h": 0.7,
|
||||||
|
"fontFace": "Noto Sans CJK SC",
|
||||||
|
"fontSize": 30,
|
||||||
|
"bold": true,
|
||||||
|
"color": "F8FAFC",
|
||||||
|
"margin": 0,
|
||||||
|
"breakLine": false
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
局部格式使用 `runs`:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"type": "text",
|
||||||
|
"runs": [
|
||||||
|
{"text": "收入 ", "options": {"bold": true}},
|
||||||
|
{"text": "+28%", "options": {"bold": true, "color": "22C55E"}}
|
||||||
|
],
|
||||||
|
"options": {
|
||||||
|
"x": 0.8,
|
||||||
|
"y": 2.0,
|
||||||
|
"w": 4.0,
|
||||||
|
"h": 0.6,
|
||||||
|
"fontSize": 24,
|
||||||
|
"margin": 0
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
列表不要输入字面量 `•`。每个列表项使用独立 Run,并设置 `bullet: true`;除最后一项外设置 `breakLine: true`,项目间距用 `paraSpaceAfter`。
|
||||||
|
|
||||||
|
### 形状、图片和图标
|
||||||
|
|
||||||
|
形状:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"type": "shape",
|
||||||
|
"shape": "roundRect",
|
||||||
|
"options": {
|
||||||
|
"x": 0.8,
|
||||||
|
"y": 1.7,
|
||||||
|
"w": 3.6,
|
||||||
|
"h": 2.2,
|
||||||
|
"fill": {"color": "E0F2FE"},
|
||||||
|
"line": {"color": "BAE6FD", "width": 1},
|
||||||
|
"shadow": {"type": "outer", "color": "0F172A", "opacity": 0.15, "blur": 2, "angle": 45, "distance": 1}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
常用形状名包括 `rect`、`roundRect`、`ellipse`、`line`、`chevron`、`triangle` 和 `hexagon`。阴影 `offset/distance` 不得为负数;向上投影时改变角度。
|
||||||
|
|
||||||
|
图片:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"type": "image",
|
||||||
|
"path": "/absolute/path/chart.png",
|
||||||
|
"options": {"x": 7.2, "y": 1.4, "w": 5.2, "h": 4.8}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
需要图标时先调用:
|
||||||
|
|
||||||
|
```text
|
||||||
|
--library fi --name FiTrendingUp --color 2563EB --size 256 --output 'tmp/pptx/<任务名>/trend.png'
|
||||||
|
```
|
||||||
|
|
||||||
|
允许的图标库:`fa6`、`fi`、`hi2`、`io5`、`lu`、`md`、`ri`、`tb`。把生成的 PNG 作为普通图片元素插入。
|
||||||
|
|
||||||
|
### 图表
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"type": "chart",
|
||||||
|
"chart_type": "bar",
|
||||||
|
"data": [
|
||||||
|
{
|
||||||
|
"name": "收入",
|
||||||
|
"labels": ["Q1", "Q2", "Q3", "Q4"],
|
||||||
|
"values": [120, 148, 176, 215]
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"options": {
|
||||||
|
"x": 0.8,
|
||||||
|
"y": 1.6,
|
||||||
|
"w": 6.0,
|
||||||
|
"h": 4.7,
|
||||||
|
"catAxisLabelFontSize": 12,
|
||||||
|
"valAxisLabelFontSize": 11,
|
||||||
|
"showLegend": false,
|
||||||
|
"showValue": true,
|
||||||
|
"dataLabelPosition": "outEnd",
|
||||||
|
"chartColors": ["2563EB"]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
支持 `area/bar/bar3d/bubble/bubble3d/doughnut/line/pie/radar/scatter`。PowerPoint 原生支持的图表必须保留为可编辑图表;只有桑基图、网络图等没有对应原生类型的可视化才使用图片。堆积条形图或柱形图的数据标签只能使用 `ctr/inEnd/inBase`,不能使用 `outEnd`。
|
||||||
|
|
||||||
|
### 表格
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"type": "table",
|
||||||
|
"rows": [
|
||||||
|
[
|
||||||
|
{"text": "指标", "options": {"bold": true, "color": "FFFFFF", "fill": "1E3A8A"}},
|
||||||
|
{"text": "本期", "options": {"bold": true, "color": "FFFFFF", "fill": "1E3A8A"}}
|
||||||
|
],
|
||||||
|
["收入", "2,150 万元"],
|
||||||
|
["增长率", "28.0%"]
|
||||||
|
],
|
||||||
|
"options": {
|
||||||
|
"x": 0.8,
|
||||||
|
"y": 1.8,
|
||||||
|
"w": 6.2,
|
||||||
|
"h": 2.2,
|
||||||
|
"border": {"type": "solid", "color": "CBD5E1", "pt": 1},
|
||||||
|
"fontFace": "Noto Sans CJK SC",
|
||||||
|
"fontSize": 14,
|
||||||
|
"margin": 0.08
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## 编辑现有演示文稿
|
||||||
|
|
||||||
|
调用:
|
||||||
|
|
||||||
|
```text
|
||||||
|
--input 'source.pptx' --output 'output/pptx/edited.pptx' --spec '<JSON对象>'
|
||||||
|
```
|
||||||
|
|
||||||
|
JSON 顶层只有 `operations`,按数组顺序执行:
|
||||||
|
|
||||||
|
| `operations[].type` | 关键字段 |
|
||||||
|
| --- | --- |
|
||||||
|
| `replace_text` | `find`、`replace`;可选 `slides/match_case/whole_word/count/required/include_notes` |
|
||||||
|
| `set_properties` | `properties`;其中 `revision` 必须是正整数 |
|
||||||
|
| `delete_slides` | `slides[]`,页码从 1 开始 |
|
||||||
|
| `reorder_slides` | `order[]`,必须完整且不重复地列出当前全部页码 |
|
||||||
|
|
||||||
|
示例:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"operations": [
|
||||||
|
{
|
||||||
|
"type": "replace_text",
|
||||||
|
"find": "2025 年",
|
||||||
|
"replace": "2026 年",
|
||||||
|
"match_case": true,
|
||||||
|
"required": true
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"type": "reorder_slides",
|
||||||
|
"order": [1, 3, 2, 4]
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
文本替换会处理同一段落内跨多个 Run 的匹配,并尽量保留替换起点和结尾的格式。输入含批注、ActiveX、宏或嵌入对象时,常规编辑脚本会停止,改用受控 OOXML 流程,避免静默丢失内容。输入含外部链接时默认停止;只有用户明确接受风险后才传 `--allow-external-links`。
|
||||||
|
|
||||||
|
## 复制模板页面
|
||||||
|
|
||||||
|
只对 `.pptx` 使用:
|
||||||
|
|
||||||
|
```text
|
||||||
|
--input 'template.pptx' --output 'tmp/pptx/<任务名>/expanded.pptx' --slide 2 --after 4
|
||||||
|
```
|
||||||
|
|
||||||
|
`slide` 是复制来源,`after` 是插入位置;省略 `after` 时紧跟来源页插入。脚本会更新页面清单、关系和内容类型,并移除不能安全共享的备注/批注关系。
|
||||||
|
|
||||||
|
复制页仍可能与原页共享图表、SmartArt 或嵌入对象部件。若返回的 `shared_relationships` 非空,修改这些对象前先做 OOXML 级独立复制;否则改动一页可能同时影响另一页。
|
||||||
|
|
||||||
|
## 高级 OOXML 编辑
|
||||||
|
|
||||||
|
固定编辑脚本无法表达且确实需要精细模板操作时:
|
||||||
|
|
||||||
|
1. 调用 `unpack_presentation.py --input <pptx> --output-dir <空目录>`。
|
||||||
|
2. 先完成页面复制、删除和重排,再修改页面内容。
|
||||||
|
3. 只最小化编辑相关 XML;不要重排、格式化或重写无关部件。
|
||||||
|
4. 每个列表项保留独立的 `<a:p>`;保留相邻 `<a:pPr>` 以继承缩进与间距;不要在文本中写字面量项目符号。
|
||||||
|
5. 有前后空格的 `<a:t>` 设置 `xml:space="preserve"`。
|
||||||
|
6. 调用 `pack_presentation.py --input-dir <目录> --output <新pptx>`。
|
||||||
|
7. 立即调用 `validate_presentation.py --original <源pptx> --check-render`。
|
||||||
|
|
||||||
|
不得手工复制单个 `slideN.xml`。新页面必须同时登记到 `ppt/presentation.xml`、`ppt/_rels/presentation.xml.rels` 和 `[Content_Types].xml`,并处理页面关系。
|
||||||
|
|
||||||
|
## 设计和排版要求
|
||||||
|
|
||||||
|
- 先确定与主题相关的配色和一个贯穿全稿的视觉母题。一个主色承担约三分之二视觉权重,搭配 1–2 个辅助色和一个强调色。
|
||||||
|
- 封面、章节页和结尾页可以使用深色背景,内容页使用浅色背景;同一套演示文稿保持一致。
|
||||||
|
- 每页至少包含一种有信息作用的视觉元素:图片、图表、图标、流程、时间线或重点数字。不要只放标题和大段项目符号。
|
||||||
|
- 交替使用双栏、卡片网格、半幅图片、对比栏和流程布局;不要连续复用同一种版式。
|
||||||
|
- 中文正文优先使用 `Noto Sans CJK SC`,拉丁正文优先使用 Arial 或 Calibri。字体会在最终用户的 PowerPoint 中渲染,LibreOffice 预览可能发生替换;非预置字体至少保留约 10% 宽度余量。
|
||||||
|
- 标题通常为 32–44 pt,分区标题 20–26 pt,正文 14–18 pt,注释 10–12 pt。正文左对齐;只对短标题或数字做居中。
|
||||||
|
- 页面边缘至少留 0.5 英寸,内容块间距至少 0.3 英寸。同类元素使用一致的栅格、间距和对齐。
|
||||||
|
- 文本框需要与图形边缘精确对齐时设置 `margin: 0`;字符间距使用 `charSpacing`,不要使用无效的 `letterSpacing`。
|
||||||
|
- 不用标题下划线、整页装饰色条或卡片单侧色边充当“设计感”。优先使用留白、轻微底色、阴影、图片裁切和图标层次。
|
||||||
|
- 不默认使用与主题无关的蓝色或米黄色;不交付低对比、越界、截断或互相重叠的元素。
|
||||||
|
- 演讲者备注只写入 `speaker_notes`,不要伪装成页面上的隐藏文本。
|
||||||
|
|
||||||
|
## 校验与视觉检查
|
||||||
|
|
||||||
|
结构校验:
|
||||||
|
|
||||||
|
```text
|
||||||
|
--input 'output/pptx/result.pptx' --check-render
|
||||||
|
```
|
||||||
|
|
||||||
|
模板派生结果:
|
||||||
|
|
||||||
|
```text
|
||||||
|
--input 'output/pptx/result.pptx' --original 'template.pptx' --check-render
|
||||||
|
```
|
||||||
|
|
||||||
|
必须满足:
|
||||||
|
|
||||||
|
- `status: valid`
|
||||||
|
- `issue_count: 0`
|
||||||
|
- `archive.missing_required_parts` 和 `duplicate_members` 为空
|
||||||
|
- `render.pdf_pages` 与 `slide_count` 一致
|
||||||
|
|
||||||
|
警告也必须逐项评估,特别是 `placeholder_text`、空页面、孤立页面部件和外部关系。
|
||||||
|
|
||||||
|
渲染全部页面:
|
||||||
|
|
||||||
|
```text
|
||||||
|
--input 'output/pptx/result.pptx' --output-dir 'tmp/pptx/<任务名>/rendered' --contact-sheet --include-pdf
|
||||||
|
```
|
||||||
|
|
||||||
|
默认 150 DPI、单次最多 30 页。可用 `--start-slide/--end-slide/--max-slides` 分批,复杂图表或小字可把 `--dpi` 提高到 180–220。联系表用于快速检查整体节奏,逐页 PNG 用于最终 QA。
|
||||||
|
|
||||||
|
逐页检查:
|
||||||
|
|
||||||
|
- 文本是否被截断、溢出或因字体替换异常换行。
|
||||||
|
- 图形、文字、页脚和来源是否重叠。
|
||||||
|
- 页面边距、列宽、卡片尺寸、基线和间距是否一致。
|
||||||
|
- 图标与文字是否有足够对比度,图表标签是否可读。
|
||||||
|
- 模板占位内容、示例数据和多余装饰是否全部清除。
|
||||||
|
- 页面顺序、标题层级、图表数值、演讲者备注和来源是否正确。
|
||||||
|
|
||||||
|
第一次渲染发现问题是正常的;修复后必须重新运行结构校验,并重新生成受影响页面的预览。
|
||||||
4
skills/pptx/agents/openai.yaml
Normal file
4
skills/pptx/agents/openai.yaml
Normal file
@ -0,0 +1,4 @@
|
|||||||
|
interface:
|
||||||
|
display_name: "PowerPoint 演示文稿"
|
||||||
|
short_description: "创建、编辑、图片 OCR、校验并渲染 PowerPoint 演示文稿"
|
||||||
|
default_prompt: "使用 $pptx 读取或创建专业演示文稿,按需识别图片文字,并完成结构与视觉校验。"
|
||||||
106
skills/pptx/scripts/_icon_renderer.js
Normal file
106
skills/pptx/scripts/_icon_renderer.js
Normal file
@ -0,0 +1,106 @@
|
|||||||
|
#!/usr/bin/env node
|
||||||
|
|
||||||
|
"use strict";
|
||||||
|
|
||||||
|
const fs = require("fs");
|
||||||
|
const path = require("path");
|
||||||
|
const React = require("react");
|
||||||
|
const ReactDOMServer = require("react-dom/server");
|
||||||
|
const sharp = require("sharp");
|
||||||
|
|
||||||
|
const LIBRARIES = {
|
||||||
|
fa6: "react-icons/fa6",
|
||||||
|
fi: "react-icons/fi",
|
||||||
|
hi2: "react-icons/hi2",
|
||||||
|
io5: "react-icons/io5",
|
||||||
|
lu: "react-icons/lu",
|
||||||
|
md: "react-icons/md",
|
||||||
|
ri: "react-icons/ri",
|
||||||
|
tb: "react-icons/tb",
|
||||||
|
};
|
||||||
|
|
||||||
|
function fail(message) {
|
||||||
|
process.stderr.write(`${message}\n`);
|
||||||
|
process.exit(1);
|
||||||
|
}
|
||||||
|
|
||||||
|
function parseArgs(argv) {
|
||||||
|
const values = {};
|
||||||
|
for (let index = 0; index < argv.length; index += 2) {
|
||||||
|
if (!argv[index]?.startsWith("--") || argv[index + 1] === undefined) {
|
||||||
|
fail("参数必须按 --name value 成对提供");
|
||||||
|
}
|
||||||
|
values[argv[index].slice(2)] = argv[index + 1];
|
||||||
|
}
|
||||||
|
if (!values.spec || !values.output) {
|
||||||
|
fail("缺少 --spec 或 --output");
|
||||||
|
}
|
||||||
|
return values;
|
||||||
|
}
|
||||||
|
|
||||||
|
function color(value, label) {
|
||||||
|
if (typeof value !== "string" || !/^[0-9A-Fa-f]{6}$/.test(value)) {
|
||||||
|
throw new Error(`${label} 必须是不带 # 的 6 位十六进制颜色`);
|
||||||
|
}
|
||||||
|
return `#${value.toUpperCase()}`;
|
||||||
|
}
|
||||||
|
|
||||||
|
async function main() {
|
||||||
|
const args = parseArgs(process.argv.slice(2));
|
||||||
|
const specPath = path.resolve(args.spec);
|
||||||
|
const outputPath = path.resolve(args.output);
|
||||||
|
if (path.extname(outputPath).toLowerCase() !== ".png") {
|
||||||
|
fail("output 必须使用 .png 扩展名");
|
||||||
|
}
|
||||||
|
let spec;
|
||||||
|
try {
|
||||||
|
spec = JSON.parse(fs.readFileSync(specPath, "utf8"));
|
||||||
|
} catch (error) {
|
||||||
|
fail(`读取 spec 失败:${error.message}`);
|
||||||
|
}
|
||||||
|
try {
|
||||||
|
const libraryPath = LIBRARIES[spec.library];
|
||||||
|
if (!libraryPath) {
|
||||||
|
throw new Error(`不支持的图标库:${spec.library}`);
|
||||||
|
}
|
||||||
|
if (typeof spec.name !== "string" || !/^[A-Za-z][A-Za-z0-9]*$/.test(spec.name)) {
|
||||||
|
throw new Error("name 必须是合法的 React Icons 导出名称");
|
||||||
|
}
|
||||||
|
const size = spec.size ?? 256;
|
||||||
|
if (!Number.isInteger(size) || size < 32 || size > 2048) {
|
||||||
|
throw new Error("size 必须是 32 到 2048 的整数");
|
||||||
|
}
|
||||||
|
const moduleExports = require(libraryPath);
|
||||||
|
const Icon = moduleExports[spec.name];
|
||||||
|
if (typeof Icon !== "function") {
|
||||||
|
throw new Error(`${spec.library} 中不存在图标 ${spec.name}`);
|
||||||
|
}
|
||||||
|
const foreground = color(spec.color || "111827", "color");
|
||||||
|
const markup = ReactDOMServer.renderToStaticMarkup(
|
||||||
|
React.createElement(Icon, {
|
||||||
|
color: foreground,
|
||||||
|
size,
|
||||||
|
title: typeof spec.title === "string" ? spec.title : undefined,
|
||||||
|
}),
|
||||||
|
);
|
||||||
|
let pipeline = sharp(Buffer.from(markup))
|
||||||
|
.resize(size, size, { fit: "contain" });
|
||||||
|
if (spec.background) {
|
||||||
|
pipeline = pipeline.flatten({ background: color(spec.background, "background") });
|
||||||
|
}
|
||||||
|
await pipeline.png().toFile(outputPath);
|
||||||
|
const metadata = await sharp(outputPath).metadata();
|
||||||
|
process.stdout.write(
|
||||||
|
`${JSON.stringify({
|
||||||
|
library: spec.library,
|
||||||
|
name: spec.name,
|
||||||
|
width: metadata.width,
|
||||||
|
height: metadata.height,
|
||||||
|
})}\n`,
|
||||||
|
);
|
||||||
|
} catch (error) {
|
||||||
|
fail(error instanceof Error ? error.message : String(error));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
main();
|
||||||
426
skills/pptx/scripts/_pptx_common.py
Normal file
426
skills/pptx/scripts/_pptx_common.py
Normal file
@ -0,0 +1,426 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import shutil
|
||||||
|
import stat
|
||||||
|
import subprocess
|
||||||
|
import tempfile
|
||||||
|
import zipfile
|
||||||
|
from datetime import date, datetime, time
|
||||||
|
from decimal import Decimal
|
||||||
|
from pathlib import Path, PurePosixPath
|
||||||
|
from typing import Any, Callable, NoReturn, Optional
|
||||||
|
from urllib.parse import quote
|
||||||
|
|
||||||
|
|
||||||
|
OOXML_PRESENTATION_SUFFIXES = {".pptx", ".potx", ".ppsx"}
|
||||||
|
PRESENTATION_INPUT_SUFFIXES = OOXML_PRESENTATION_SUFFIXES | {".ppt"}
|
||||||
|
PRESENTATION_OUTPUT_SUFFIXES = {".pptx", ".potx", ".pdf"}
|
||||||
|
IMAGE_SUFFIXES = {".png", ".jpg", ".jpeg", ".webp"}
|
||||||
|
|
||||||
|
MAX_ARCHIVE_MEMBERS = 20_000
|
||||||
|
MAX_ARCHIVE_UNCOMPRESSED_BYTES = 1_073_741_824
|
||||||
|
MAX_MEMBER_BYTES = 268_435_456
|
||||||
|
|
||||||
|
REL_NS = "http://schemas.openxmlformats.org/package/2006/relationships"
|
||||||
|
P_NS = "http://schemas.openxmlformats.org/presentationml/2006/main"
|
||||||
|
R_NS = "http://schemas.openxmlformats.org/officeDocument/2006/relationships"
|
||||||
|
A_NS = "http://schemas.openxmlformats.org/drawingml/2006/main"
|
||||||
|
|
||||||
|
|
||||||
|
class SkillArgumentParser(argparse.ArgumentParser):
|
||||||
|
def error(self, message: str) -> NoReturn:
|
||||||
|
raise ValueError(f"参数错误:{message}")
|
||||||
|
|
||||||
|
|
||||||
|
def json_default(value: Any) -> Any:
|
||||||
|
if isinstance(value, (datetime, date, time)):
|
||||||
|
return value.isoformat()
|
||||||
|
if isinstance(value, Decimal):
|
||||||
|
return float(value)
|
||||||
|
if isinstance(value, Path):
|
||||||
|
return str(value)
|
||||||
|
return str(value)
|
||||||
|
|
||||||
|
|
||||||
|
def emit(payload: dict[str, Any]) -> None:
|
||||||
|
print(json.dumps(payload, ensure_ascii=False, default=json_default))
|
||||||
|
|
||||||
|
|
||||||
|
def failure_message(exc: Exception) -> str:
|
||||||
|
if isinstance(exc, (FileExistsError, FileNotFoundError)):
|
||||||
|
return str(exc)
|
||||||
|
if isinstance(exc, PermissionError):
|
||||||
|
return "文件处理失败:没有目标路径的访问权限"
|
||||||
|
if isinstance(exc, subprocess.TimeoutExpired):
|
||||||
|
return f"外部程序执行超时({exc.timeout} 秒)"
|
||||||
|
if isinstance(exc, subprocess.CalledProcessError):
|
||||||
|
stderr = (exc.stderr or "").strip()
|
||||||
|
return f"外部程序执行失败:{stderr or exc}"
|
||||||
|
if isinstance(exc, zipfile.BadZipFile):
|
||||||
|
return "文件不是有效的 OOXML 压缩包"
|
||||||
|
if isinstance(exc, OSError):
|
||||||
|
return f"文件处理失败:{exc}"
|
||||||
|
return str(exc)
|
||||||
|
|
||||||
|
|
||||||
|
def run_cli(action: Callable[[], dict[str, Any]]) -> int:
|
||||||
|
try:
|
||||||
|
result = action()
|
||||||
|
except Exception as exc:
|
||||||
|
emit({"ok": False, "error": failure_message(exc)})
|
||||||
|
return 1
|
||||||
|
emit({"ok": True, **result})
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
def input_file(value: str, suffixes: Optional[set[str]] = None) -> Path:
|
||||||
|
path = Path(value).expanduser().resolve()
|
||||||
|
if not path.exists():
|
||||||
|
raise FileNotFoundError(f"输入文件不存在:{path}")
|
||||||
|
if not path.is_file():
|
||||||
|
raise ValueError(f"输入路径不是文件:{path}")
|
||||||
|
if path.stat().st_size <= 0:
|
||||||
|
raise ValueError(f"输入文件为空:{path}")
|
||||||
|
if suffixes is not None and path.suffix.lower() not in suffixes:
|
||||||
|
expected = "、".join(sorted(suffixes))
|
||||||
|
raise ValueError(f"不支持的输入格式 {path.suffix};允许:{expected}")
|
||||||
|
return path
|
||||||
|
|
||||||
|
|
||||||
|
def output_file(
|
||||||
|
value: str,
|
||||||
|
suffixes: Optional[set[str]] = None,
|
||||||
|
*,
|
||||||
|
overwrite: bool = False,
|
||||||
|
) -> Path:
|
||||||
|
path = Path(value).expanduser().resolve()
|
||||||
|
if suffixes is not None and path.suffix.lower() not in suffixes:
|
||||||
|
expected = "、".join(sorted(suffixes))
|
||||||
|
raise ValueError(f"不支持的输出格式 {path.suffix};允许:{expected}")
|
||||||
|
if path.exists() and not path.is_file():
|
||||||
|
raise ValueError(f"目标路径不是文件:{path}")
|
||||||
|
if path.exists() and not overwrite:
|
||||||
|
raise FileExistsError(f"目标文件已存在:{path}")
|
||||||
|
path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
return path
|
||||||
|
|
||||||
|
|
||||||
|
def output_directory(value: str) -> Path:
|
||||||
|
path = Path(value).expanduser().resolve()
|
||||||
|
if path.exists() and not path.is_dir():
|
||||||
|
raise ValueError(f"输出路径不是目录:{path}")
|
||||||
|
path.mkdir(parents=True, exist_ok=True)
|
||||||
|
return path
|
||||||
|
|
||||||
|
|
||||||
|
def publish_file(source: Path, destination: Path, *, overwrite: bool) -> None:
|
||||||
|
destination.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
if overwrite:
|
||||||
|
os.replace(source, destination)
|
||||||
|
return
|
||||||
|
try:
|
||||||
|
os.link(source, destination)
|
||||||
|
except FileExistsError as exc:
|
||||||
|
raise FileExistsError(f"目标文件已存在:{destination}") from exc
|
||||||
|
except OSError:
|
||||||
|
if destination.exists():
|
||||||
|
raise FileExistsError(f"目标文件已存在:{destination}")
|
||||||
|
shutil.copy2(source, destination)
|
||||||
|
finally:
|
||||||
|
if source.exists():
|
||||||
|
source.unlink()
|
||||||
|
|
||||||
|
|
||||||
|
def load_json_argument(
|
||||||
|
inline_value: Optional[str],
|
||||||
|
file_value: Optional[str],
|
||||||
|
*,
|
||||||
|
label: str,
|
||||||
|
max_bytes: int = 4 * 1024 * 1024,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
if bool(inline_value) == bool(file_value):
|
||||||
|
raise ValueError(f"{label} 必须且只能通过内联 JSON 或 JSON 文件提供一次")
|
||||||
|
if file_value:
|
||||||
|
source = input_file(file_value, {".json"})
|
||||||
|
if source.stat().st_size > max_bytes:
|
||||||
|
raise ValueError(f"{label} JSON 文件超过 {max_bytes} 字节限制")
|
||||||
|
raw = source.read_text(encoding="utf-8")
|
||||||
|
else:
|
||||||
|
raw = inline_value or ""
|
||||||
|
if len(raw.encode("utf-8")) > max_bytes:
|
||||||
|
raise ValueError(f"{label} 内联 JSON 超过 {max_bytes} 字节限制")
|
||||||
|
try:
|
||||||
|
parsed = json.loads(raw)
|
||||||
|
except json.JSONDecodeError as exc:
|
||||||
|
raise ValueError(
|
||||||
|
f"{label} JSON 无效:第 {exc.lineno} 行第 {exc.colno} 列,{exc.msg}"
|
||||||
|
) from exc
|
||||||
|
if not isinstance(parsed, dict):
|
||||||
|
raise ValueError(f"{label} JSON 顶层必须是对象")
|
||||||
|
return parsed
|
||||||
|
|
||||||
|
|
||||||
|
def find_program(*names: str) -> str:
|
||||||
|
for name in names:
|
||||||
|
resolved = shutil.which(name)
|
||||||
|
if resolved:
|
||||||
|
return resolved
|
||||||
|
raise FileNotFoundError(f"运行环境缺少命令:{' / '.join(names)}")
|
||||||
|
|
||||||
|
|
||||||
|
def run_program(
|
||||||
|
args: list[str],
|
||||||
|
*,
|
||||||
|
timeout: int,
|
||||||
|
cwd: Optional[Path] = None,
|
||||||
|
env: Optional[dict[str, str]] = None,
|
||||||
|
) -> subprocess.CompletedProcess[str]:
|
||||||
|
if timeout < 1 or timeout > 900:
|
||||||
|
raise ValueError("timeout 必须在 1 到 900 秒之间")
|
||||||
|
completed = subprocess.run(
|
||||||
|
args,
|
||||||
|
cwd=str(cwd) if cwd else None,
|
||||||
|
env=env,
|
||||||
|
stdin=subprocess.DEVNULL,
|
||||||
|
stdout=subprocess.PIPE,
|
||||||
|
stderr=subprocess.PIPE,
|
||||||
|
text=True,
|
||||||
|
timeout=timeout,
|
||||||
|
check=False,
|
||||||
|
)
|
||||||
|
if completed.returncode != 0:
|
||||||
|
raise subprocess.CalledProcessError(
|
||||||
|
completed.returncode,
|
||||||
|
args,
|
||||||
|
output=completed.stdout,
|
||||||
|
stderr=completed.stderr,
|
||||||
|
)
|
||||||
|
return completed
|
||||||
|
|
||||||
|
|
||||||
|
def office_profile_uri(profile: Path) -> str:
|
||||||
|
return "file://" + quote(str(profile.resolve()), safe="/:")
|
||||||
|
|
||||||
|
|
||||||
|
def run_soffice_convert(
|
||||||
|
source: Path,
|
||||||
|
*,
|
||||||
|
target_format: str,
|
||||||
|
output_dir: Path,
|
||||||
|
timeout: int,
|
||||||
|
filter_name: Optional[str] = None,
|
||||||
|
) -> tuple[Path, dict[str, str]]:
|
||||||
|
soffice = find_program("soffice", "libreoffice")
|
||||||
|
output_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
profile = Path(tempfile.mkdtemp(prefix="pptx-soffice-profile-"))
|
||||||
|
try:
|
||||||
|
cache_dir = profile / "cache"
|
||||||
|
cache_dir.mkdir()
|
||||||
|
process_env = os.environ.copy()
|
||||||
|
process_env["XDG_CACHE_HOME"] = str(cache_dir)
|
||||||
|
convert_arg = target_format
|
||||||
|
if filter_name:
|
||||||
|
convert_arg = f"{target_format}:{filter_name}"
|
||||||
|
completed = run_program(
|
||||||
|
[
|
||||||
|
soffice,
|
||||||
|
f"-env:UserInstallation={office_profile_uri(profile)}",
|
||||||
|
"--headless",
|
||||||
|
"--nologo",
|
||||||
|
"--nodefault",
|
||||||
|
"--nolockcheck",
|
||||||
|
"--nofirststartwizard",
|
||||||
|
"--convert-to",
|
||||||
|
convert_arg,
|
||||||
|
"--outdir",
|
||||||
|
str(output_dir),
|
||||||
|
str(source),
|
||||||
|
],
|
||||||
|
timeout=timeout,
|
||||||
|
env=process_env,
|
||||||
|
)
|
||||||
|
expected = output_dir / f"{source.stem}.{target_format}"
|
||||||
|
if not expected.exists():
|
||||||
|
candidates = sorted(output_dir.glob(f"{source.stem}.*"))
|
||||||
|
if len(candidates) == 1:
|
||||||
|
expected = candidates[0]
|
||||||
|
else:
|
||||||
|
raise RuntimeError(
|
||||||
|
"LibreOffice 未生成预期文件;"
|
||||||
|
f"stdout={completed.stdout[-1000:]!r} "
|
||||||
|
f"stderr={completed.stderr[-1000:]!r}"
|
||||||
|
)
|
||||||
|
return expected, {
|
||||||
|
"stdout": completed.stdout[-2000:],
|
||||||
|
"stderr": completed.stderr[-2000:],
|
||||||
|
}
|
||||||
|
finally:
|
||||||
|
shutil.rmtree(profile, ignore_errors=True)
|
||||||
|
|
||||||
|
|
||||||
|
def _safe_archive_name(name: str) -> PurePosixPath:
|
||||||
|
candidate = PurePosixPath(name)
|
||||||
|
if (
|
||||||
|
candidate.is_absolute()
|
||||||
|
or not candidate.parts
|
||||||
|
or ".." in candidate.parts
|
||||||
|
or candidate.parts[0].endswith(":")
|
||||||
|
):
|
||||||
|
raise ValueError(f"OOXML 压缩包含不安全路径:{name}")
|
||||||
|
return candidate
|
||||||
|
|
||||||
|
|
||||||
|
def _zip_member_is_symlink(info: zipfile.ZipInfo) -> bool:
|
||||||
|
mode = (info.external_attr >> 16) & 0xFFFF
|
||||||
|
return stat.S_ISLNK(mode)
|
||||||
|
|
||||||
|
|
||||||
|
def inspect_archive(path: Path) -> dict[str, Any]:
|
||||||
|
total_size = 0
|
||||||
|
xml_count = 0
|
||||||
|
media_count = 0
|
||||||
|
slide_count = 0
|
||||||
|
with zipfile.ZipFile(path) as archive:
|
||||||
|
infos = archive.infolist()
|
||||||
|
if len(infos) > MAX_ARCHIVE_MEMBERS:
|
||||||
|
raise ValueError(
|
||||||
|
f"OOXML 压缩包成员过多:{len(infos)} > {MAX_ARCHIVE_MEMBERS}"
|
||||||
|
)
|
||||||
|
seen: set[str] = set()
|
||||||
|
duplicates: list[str] = []
|
||||||
|
for info in infos:
|
||||||
|
_safe_archive_name(info.filename)
|
||||||
|
if _zip_member_is_symlink(info):
|
||||||
|
raise ValueError(f"OOXML 压缩包含符号链接:{info.filename}")
|
||||||
|
if info.file_size > MAX_MEMBER_BYTES:
|
||||||
|
raise ValueError(f"OOXML 成员过大:{info.filename}")
|
||||||
|
total_size += info.file_size
|
||||||
|
if total_size > MAX_ARCHIVE_UNCOMPRESSED_BYTES:
|
||||||
|
raise ValueError("OOXML 解压后总大小超过安全限制")
|
||||||
|
if info.filename in seen:
|
||||||
|
duplicates.append(info.filename)
|
||||||
|
seen.add(info.filename)
|
||||||
|
if info.filename.endswith((".xml", ".rels")):
|
||||||
|
xml_count += 1
|
||||||
|
if info.filename.startswith("ppt/media/") and not info.is_dir():
|
||||||
|
media_count += 1
|
||||||
|
if (
|
||||||
|
info.filename.startswith("ppt/slides/slide")
|
||||||
|
and info.filename.endswith(".xml")
|
||||||
|
and "/_rels/" not in info.filename
|
||||||
|
):
|
||||||
|
slide_count += 1
|
||||||
|
required = {
|
||||||
|
"[Content_Types].xml",
|
||||||
|
"_rels/.rels",
|
||||||
|
"ppt/presentation.xml",
|
||||||
|
"ppt/_rels/presentation.xml.rels",
|
||||||
|
}
|
||||||
|
missing = sorted(required - seen)
|
||||||
|
return {
|
||||||
|
"member_count": len(infos),
|
||||||
|
"uncompressed_bytes": total_size,
|
||||||
|
"xml_part_count": xml_count,
|
||||||
|
"media_count": media_count,
|
||||||
|
"slide_part_count": slide_count,
|
||||||
|
"duplicate_members": duplicates,
|
||||||
|
"missing_required_parts": missing,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def safe_extract_presentation(source: Path, destination: Path) -> dict[str, Any]:
|
||||||
|
archive_info = inspect_archive(source)
|
||||||
|
if archive_info["missing_required_parts"]:
|
||||||
|
raise ValueError(
|
||||||
|
"演示文稿缺少必要部件:"
|
||||||
|
+ "、".join(archive_info["missing_required_parts"])
|
||||||
|
)
|
||||||
|
destination.mkdir(parents=True, exist_ok=True)
|
||||||
|
root = destination.resolve()
|
||||||
|
with zipfile.ZipFile(source) as archive:
|
||||||
|
for info in archive.infolist():
|
||||||
|
relative = _safe_archive_name(info.filename)
|
||||||
|
target = destination.joinpath(*relative.parts)
|
||||||
|
resolved = target.resolve()
|
||||||
|
if root not in resolved.parents and resolved != root:
|
||||||
|
raise ValueError(f"OOXML 成员逃逸输出目录:{info.filename}")
|
||||||
|
if info.is_dir():
|
||||||
|
target.mkdir(parents=True, exist_ok=True)
|
||||||
|
continue
|
||||||
|
target.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
with archive.open(info, "r") as source_handle, target.open("wb") as target_handle:
|
||||||
|
shutil.copyfileobj(source_handle, target_handle)
|
||||||
|
return archive_info
|
||||||
|
|
||||||
|
|
||||||
|
def pack_presentation_directory(source_dir: Path, destination: Path) -> dict[str, Any]:
|
||||||
|
source_dir = source_dir.expanduser().resolve()
|
||||||
|
if not source_dir.exists() or not source_dir.is_dir():
|
||||||
|
raise ValueError(f"待打包目录不存在或不是目录:{source_dir}")
|
||||||
|
required = {
|
||||||
|
source_dir / "[Content_Types].xml",
|
||||||
|
source_dir / "_rels" / ".rels",
|
||||||
|
source_dir / "ppt" / "presentation.xml",
|
||||||
|
source_dir / "ppt" / "_rels" / "presentation.xml.rels",
|
||||||
|
}
|
||||||
|
missing = sorted(
|
||||||
|
str(path.relative_to(source_dir)) for path in required if not path.is_file()
|
||||||
|
)
|
||||||
|
if missing:
|
||||||
|
raise ValueError("待打包目录缺少必要部件:" + "、".join(missing))
|
||||||
|
|
||||||
|
files: list[Path] = []
|
||||||
|
total_size = 0
|
||||||
|
for path in sorted(source_dir.rglob("*")):
|
||||||
|
if path.is_symlink():
|
||||||
|
raise ValueError(f"待打包目录包含符号链接:{path}")
|
||||||
|
if path.is_file():
|
||||||
|
size = path.stat().st_size
|
||||||
|
if size > MAX_MEMBER_BYTES:
|
||||||
|
raise ValueError(f"待打包文件过大:{path}")
|
||||||
|
total_size += size
|
||||||
|
if total_size > MAX_ARCHIVE_UNCOMPRESSED_BYTES:
|
||||||
|
raise ValueError("待打包文件总大小超过安全限制")
|
||||||
|
files.append(path)
|
||||||
|
if len(files) > MAX_ARCHIVE_MEMBERS:
|
||||||
|
raise ValueError("待打包文件数量超过安全限制")
|
||||||
|
|
||||||
|
with zipfile.ZipFile(
|
||||||
|
destination,
|
||||||
|
"w",
|
||||||
|
compression=zipfile.ZIP_DEFLATED,
|
||||||
|
compresslevel=6,
|
||||||
|
) as archive:
|
||||||
|
for path in files:
|
||||||
|
archive.write(path, path.relative_to(source_dir).as_posix())
|
||||||
|
return inspect_archive(destination)
|
||||||
|
|
||||||
|
|
||||||
|
def parse_xml_bytes(payload: bytes, *, label: str = "XML") -> Any:
|
||||||
|
try:
|
||||||
|
from defusedxml import ElementTree as ET
|
||||||
|
from defusedxml.common import DefusedXmlException
|
||||||
|
|
||||||
|
try:
|
||||||
|
return ET.fromstring(payload)
|
||||||
|
except (ET.ParseError, DefusedXmlException, ValueError) as exc:
|
||||||
|
raise ValueError(f"{label} 解析失败:{exc}") from exc
|
||||||
|
except ModuleNotFoundError:
|
||||||
|
from lxml import etree
|
||||||
|
|
||||||
|
parser = etree.XMLParser(
|
||||||
|
resolve_entities=False,
|
||||||
|
no_network=True,
|
||||||
|
recover=False,
|
||||||
|
huge_tree=False,
|
||||||
|
remove_comments=False,
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
return etree.fromstring(payload, parser=parser)
|
||||||
|
except (etree.XMLSyntaxError, ValueError) as exc:
|
||||||
|
raise ValueError(f"{label} 解析失败:{exc}") from exc
|
||||||
431
skills/pptx/scripts/_presentation_builder.js
Normal file
431
skills/pptx/scripts/_presentation_builder.js
Normal file
@ -0,0 +1,431 @@
|
|||||||
|
#!/usr/bin/env node
|
||||||
|
|
||||||
|
"use strict";
|
||||||
|
|
||||||
|
const fs = require("fs");
|
||||||
|
const path = require("path");
|
||||||
|
const PptxGenJS = require("pptxgenjs");
|
||||||
|
|
||||||
|
const LAYOUTS = {
|
||||||
|
LAYOUT_WIDE: { width: 13.333, height: 7.5 },
|
||||||
|
LAYOUT_16X9: { width: 10, height: 5.625 },
|
||||||
|
LAYOUT_4X3: { width: 10, height: 7.5 },
|
||||||
|
};
|
||||||
|
|
||||||
|
function fail(message) {
|
||||||
|
process.stderr.write(`${message}\n`);
|
||||||
|
process.exit(1);
|
||||||
|
}
|
||||||
|
|
||||||
|
function parseArgs(argv) {
|
||||||
|
const values = {};
|
||||||
|
for (let index = 0; index < argv.length; index += 2) {
|
||||||
|
const key = argv[index];
|
||||||
|
const value = argv[index + 1];
|
||||||
|
if (!key || !key.startsWith("--") || value === undefined) {
|
||||||
|
fail("参数必须按 --name value 成对提供");
|
||||||
|
}
|
||||||
|
values[key.slice(2)] = value;
|
||||||
|
}
|
||||||
|
if (!values.spec || !values.output) {
|
||||||
|
fail("缺少 --spec 或 --output");
|
||||||
|
}
|
||||||
|
return values;
|
||||||
|
}
|
||||||
|
|
||||||
|
function isPlainObject(value) {
|
||||||
|
return value !== null && typeof value === "object" && !Array.isArray(value);
|
||||||
|
}
|
||||||
|
|
||||||
|
function clone(value) {
|
||||||
|
return JSON.parse(JSON.stringify(value));
|
||||||
|
}
|
||||||
|
|
||||||
|
function requireObject(value, label) {
|
||||||
|
if (!isPlainObject(value)) {
|
||||||
|
throw new Error(`${label} 必须是对象`);
|
||||||
|
}
|
||||||
|
return value;
|
||||||
|
}
|
||||||
|
|
||||||
|
function requireArray(value, label, maxLength = 1000) {
|
||||||
|
if (!Array.isArray(value)) {
|
||||||
|
throw new Error(`${label} 必须是数组`);
|
||||||
|
}
|
||||||
|
if (value.length > maxLength) {
|
||||||
|
throw new Error(`${label} 超过 ${maxLength} 项限制`);
|
||||||
|
}
|
||||||
|
return value;
|
||||||
|
}
|
||||||
|
|
||||||
|
function requireString(value, label, maxLength = 100000) {
|
||||||
|
if (typeof value !== "string") {
|
||||||
|
throw new Error(`${label} 必须是字符串`);
|
||||||
|
}
|
||||||
|
if (value.length > maxLength) {
|
||||||
|
throw new Error(`${label} 超过 ${maxLength} 字符限制`);
|
||||||
|
}
|
||||||
|
return value;
|
||||||
|
}
|
||||||
|
|
||||||
|
function cleanColor(value, label) {
|
||||||
|
if (typeof value !== "string" || !/^[0-9A-Fa-f]{6}$/.test(value)) {
|
||||||
|
throw new Error(`${label} 必须是不带 # 的 6 位十六进制颜色`);
|
||||||
|
}
|
||||||
|
return value.toUpperCase();
|
||||||
|
}
|
||||||
|
|
||||||
|
function validateTree(value, label = "options", depth = 0) {
|
||||||
|
if (depth > 12) {
|
||||||
|
throw new Error(`${label} 嵌套层级过深`);
|
||||||
|
}
|
||||||
|
if (value === null || typeof value === "boolean" || typeof value === "number") {
|
||||||
|
if (typeof value === "number" && !Number.isFinite(value)) {
|
||||||
|
throw new Error(`${label} 包含非有限数值`);
|
||||||
|
}
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (typeof value === "string") {
|
||||||
|
if (value.length > 200000) {
|
||||||
|
throw new Error(`${label} 字符串过长`);
|
||||||
|
}
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (Array.isArray(value)) {
|
||||||
|
if (value.length > 5000) {
|
||||||
|
throw new Error(`${label} 数组过长`);
|
||||||
|
}
|
||||||
|
value.forEach((item, index) => validateTree(item, `${label}[${index}]`, depth + 1));
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (!isPlainObject(value)) {
|
||||||
|
throw new Error(`${label} 包含不支持的数据类型`);
|
||||||
|
}
|
||||||
|
const entries = Object.entries(value);
|
||||||
|
if (entries.length > 500) {
|
||||||
|
throw new Error(`${label} 字段过多`);
|
||||||
|
}
|
||||||
|
for (const [key, item] of entries) {
|
||||||
|
if (["__proto__", "prototype", "constructor"].includes(key)) {
|
||||||
|
throw new Error(`${label} 包含禁止字段 ${key}`);
|
||||||
|
}
|
||||||
|
if (key === "color" && typeof item === "string") {
|
||||||
|
cleanColor(item, `${label}.${key}`);
|
||||||
|
}
|
||||||
|
if (key === "chartColors" && Array.isArray(item)) {
|
||||||
|
item.forEach((color, index) => cleanColor(color, `${label}.chartColors[${index}]`));
|
||||||
|
}
|
||||||
|
if (key === "offset" && label.toLowerCase().includes("shadow")) {
|
||||||
|
if (typeof item !== "number" || item < 0) {
|
||||||
|
throw new Error(`${label}.offset 必须是非负数`);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
validateTree(item, `${label}.${key}`, depth + 1);
|
||||||
|
}
|
||||||
|
if (isPlainObject(value.fill) && value.fill.type === "gradient") {
|
||||||
|
throw new Error(`${label}.fill 不支持渐变;请使用渐变背景图片`);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function requireBox(options, label) {
|
||||||
|
for (const key of ["x", "y", "w", "h"]) {
|
||||||
|
if (typeof options[key] !== "number" || !Number.isFinite(options[key])) {
|
||||||
|
throw new Error(`${label}.${key} 必须是有限数值`);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function normalizeRuns(runs, label) {
|
||||||
|
return requireArray(runs, label, 2000).map((run, index) => {
|
||||||
|
requireObject(run, `${label}[${index}]`);
|
||||||
|
const text = requireString(run.text ?? "", `${label}[${index}].text`);
|
||||||
|
const options = clone(run.options || {});
|
||||||
|
requireObject(options, `${label}[${index}].options`);
|
||||||
|
validateTree(options, `${label}[${index}].options`);
|
||||||
|
return { text, options };
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
function normalizeTableRows(rows, label) {
|
||||||
|
return requireArray(rows, label, 500).map((row, rowIndex) =>
|
||||||
|
requireArray(row, `${label}[${rowIndex}]`, 100).map((cell, columnIndex) => {
|
||||||
|
if (typeof cell === "string" || typeof cell === "number") {
|
||||||
|
return String(cell);
|
||||||
|
}
|
||||||
|
requireObject(cell, `${label}[${rowIndex}][${columnIndex}]`);
|
||||||
|
const normalized = {
|
||||||
|
text: requireString(
|
||||||
|
String(cell.text ?? ""),
|
||||||
|
`${label}[${rowIndex}][${columnIndex}].text`,
|
||||||
|
),
|
||||||
|
options: clone(cell.options || {}),
|
||||||
|
};
|
||||||
|
requireObject(
|
||||||
|
normalized.options,
|
||||||
|
`${label}[${rowIndex}][${columnIndex}].options`,
|
||||||
|
);
|
||||||
|
validateTree(
|
||||||
|
normalized.options,
|
||||||
|
`${label}[${rowIndex}][${columnIndex}].options`,
|
||||||
|
);
|
||||||
|
return normalized;
|
||||||
|
}),
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
function normalizeChartData(data, label) {
|
||||||
|
return requireArray(data, label, 100).map((series, index) => {
|
||||||
|
requireObject(series, `${label}[${index}]`);
|
||||||
|
const labels = requireArray(series.labels, `${label}[${index}].labels`, 10000);
|
||||||
|
const values = requireArray(series.values, `${label}[${index}].values`, 10000);
|
||||||
|
if (labels.length !== values.length) {
|
||||||
|
throw new Error(`${label}[${index}] 的 labels 与 values 长度必须一致`);
|
||||||
|
}
|
||||||
|
return {
|
||||||
|
name: requireString(String(series.name ?? ""), `${label}[${index}].name`, 500),
|
||||||
|
labels: labels.map((item) => String(item)),
|
||||||
|
values: values.map((item, valueIndex) => {
|
||||||
|
if (typeof item !== "number" || !Number.isFinite(item)) {
|
||||||
|
throw new Error(`${label}[${index}].values[${valueIndex}] 必须是有限数值`);
|
||||||
|
}
|
||||||
|
return item;
|
||||||
|
}),
|
||||||
|
};
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
function resolveImagePath(value, label) {
|
||||||
|
const raw = requireString(value, label, 4096);
|
||||||
|
const resolved = path.resolve(raw);
|
||||||
|
if (!fs.existsSync(resolved) || !fs.statSync(resolved).isFile()) {
|
||||||
|
throw new Error(`${label} 文件不存在:${resolved}`);
|
||||||
|
}
|
||||||
|
return resolved;
|
||||||
|
}
|
||||||
|
|
||||||
|
function elementBounds(element, options, slideNumber, index, dimensions, warnings) {
|
||||||
|
if (!["text", "shape", "image", "chart", "table"].includes(element.type)) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
requireBox(options, `slides[${slideNumber - 1}].elements[${index}].options`);
|
||||||
|
const tolerance = 0.002;
|
||||||
|
if (
|
||||||
|
options.x < -tolerance ||
|
||||||
|
options.y < -tolerance ||
|
||||||
|
options.x + options.w > dimensions.width + tolerance ||
|
||||||
|
options.y + options.h > dimensions.height + tolerance
|
||||||
|
) {
|
||||||
|
warnings.push({
|
||||||
|
slide: slideNumber,
|
||||||
|
element: index + 1,
|
||||||
|
code: "out_of_bounds",
|
||||||
|
message: "元素超出演示文稿画布",
|
||||||
|
});
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
async function build(spec, outputPath) {
|
||||||
|
requireObject(spec, "spec");
|
||||||
|
const slides = requireArray(spec.slides, "slides", 300);
|
||||||
|
if (slides.length === 0) {
|
||||||
|
throw new Error("slides 至少需要一页");
|
||||||
|
}
|
||||||
|
|
||||||
|
const layout = spec.layout || "LAYOUT_WIDE";
|
||||||
|
if (!Object.prototype.hasOwnProperty.call(LAYOUTS, layout)) {
|
||||||
|
throw new Error(`layout 不支持:${layout}`);
|
||||||
|
}
|
||||||
|
const dimensions = LAYOUTS[layout];
|
||||||
|
const pptx = new PptxGenJS();
|
||||||
|
pptx.layout = layout;
|
||||||
|
|
||||||
|
const properties = spec.properties || {};
|
||||||
|
requireObject(properties, "properties");
|
||||||
|
const propertyMap = {
|
||||||
|
author: "author",
|
||||||
|
company: "company",
|
||||||
|
subject: "subject",
|
||||||
|
title: "title",
|
||||||
|
comments: "comments",
|
||||||
|
revision: "revision",
|
||||||
|
};
|
||||||
|
for (const [sourceKey, targetKey] of Object.entries(propertyMap)) {
|
||||||
|
if (properties[sourceKey] !== undefined) {
|
||||||
|
pptx[targetKey] = requireString(
|
||||||
|
String(properties[sourceKey]),
|
||||||
|
`properties.${sourceKey}`,
|
||||||
|
4000,
|
||||||
|
);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
const theme = spec.theme || {};
|
||||||
|
requireObject(theme, "theme");
|
||||||
|
const headFont = requireString(
|
||||||
|
String(theme.head_font || "Noto Sans CJK SC"),
|
||||||
|
"theme.head_font",
|
||||||
|
200,
|
||||||
|
);
|
||||||
|
const bodyFont = requireString(
|
||||||
|
String(theme.body_font || "Noto Sans CJK SC"),
|
||||||
|
"theme.body_font",
|
||||||
|
200,
|
||||||
|
);
|
||||||
|
pptx.theme = {
|
||||||
|
headFontFace: headFont,
|
||||||
|
bodyFontFace: bodyFont,
|
||||||
|
lang: requireString(String(theme.language || "zh-CN"), "theme.language", 40),
|
||||||
|
};
|
||||||
|
pptx.lang = theme.language || "zh-CN";
|
||||||
|
|
||||||
|
const warnings = [];
|
||||||
|
const elementCounts = {};
|
||||||
|
for (let slideIndex = 0; slideIndex < slides.length; slideIndex += 1) {
|
||||||
|
const slideSpec = requireObject(slides[slideIndex], `slides[${slideIndex}]`);
|
||||||
|
const slide = pptx.addSlide();
|
||||||
|
if (slideSpec.background !== undefined) {
|
||||||
|
slide.background = {
|
||||||
|
color: cleanColor(slideSpec.background, `slides[${slideIndex}].background`),
|
||||||
|
};
|
||||||
|
}
|
||||||
|
const elements = requireArray(
|
||||||
|
slideSpec.elements || [],
|
||||||
|
`slides[${slideIndex}].elements`,
|
||||||
|
1000,
|
||||||
|
);
|
||||||
|
if (elements.length === 0) {
|
||||||
|
warnings.push({
|
||||||
|
slide: slideIndex + 1,
|
||||||
|
code: "empty_slide",
|
||||||
|
message: "页面没有可见元素",
|
||||||
|
});
|
||||||
|
}
|
||||||
|
for (let elementIndex = 0; elementIndex < elements.length; elementIndex += 1) {
|
||||||
|
const label = `slides[${slideIndex}].elements[${elementIndex}]`;
|
||||||
|
const element = requireObject(elements[elementIndex], label);
|
||||||
|
const type = requireString(element.type, `${label}.type`, 50);
|
||||||
|
const options = clone(element.options || {});
|
||||||
|
requireObject(options, `${label}.options`);
|
||||||
|
validateTree(options, `${label}.options`);
|
||||||
|
elementBounds(
|
||||||
|
element,
|
||||||
|
options,
|
||||||
|
slideIndex + 1,
|
||||||
|
elementIndex,
|
||||||
|
dimensions,
|
||||||
|
warnings,
|
||||||
|
);
|
||||||
|
elementCounts[type] = (elementCounts[type] || 0) + 1;
|
||||||
|
|
||||||
|
if (type === "text") {
|
||||||
|
if (options.fontFace === undefined) {
|
||||||
|
options.fontFace = bodyFont;
|
||||||
|
}
|
||||||
|
const content =
|
||||||
|
element.runs !== undefined
|
||||||
|
? normalizeRuns(element.runs, `${label}.runs`)
|
||||||
|
: requireString(String(element.text ?? ""), `${label}.text`);
|
||||||
|
slide.addText(content, options);
|
||||||
|
} else if (type === "shape") {
|
||||||
|
const shapeName = requireString(element.shape || "rect", `${label}.shape`, 100);
|
||||||
|
const shapeType = pptx.ShapeType[shapeName];
|
||||||
|
if (!shapeType) {
|
||||||
|
throw new Error(`${label}.shape 不支持:${shapeName}`);
|
||||||
|
}
|
||||||
|
slide.addShape(shapeType, options);
|
||||||
|
} else if (type === "image") {
|
||||||
|
if (element.path !== undefined) {
|
||||||
|
options.path = resolveImagePath(element.path, `${label}.path`);
|
||||||
|
} else if (element.data !== undefined) {
|
||||||
|
const data = requireString(element.data, `${label}.data`, 20_000_000);
|
||||||
|
if (!/^data:image\/(?:png|jpeg|jpg|webp);base64,/.test(data)) {
|
||||||
|
throw new Error(`${label}.data 必须是受支持图片的 base64 data URL`);
|
||||||
|
}
|
||||||
|
options.data = data;
|
||||||
|
} else {
|
||||||
|
throw new Error(`${label} 必须提供 path 或 data`);
|
||||||
|
}
|
||||||
|
slide.addImage(options);
|
||||||
|
} else if (type === "chart") {
|
||||||
|
const chartName = requireString(element.chart_type, `${label}.chart_type`, 50);
|
||||||
|
const chartType = pptx.ChartType[chartName];
|
||||||
|
if (!chartType) {
|
||||||
|
throw new Error(`${label}.chart_type 不支持:${chartName}`);
|
||||||
|
}
|
||||||
|
const chartData = normalizeChartData(element.data, `${label}.data`);
|
||||||
|
const chartOptions = {
|
||||||
|
showLegend: chartData.length > 1,
|
||||||
|
showTitle: Boolean(options.title),
|
||||||
|
showValue: true,
|
||||||
|
chartColors: ["2563EB", "14B8A6", "F97316", "8B5CF6", "E11D48"],
|
||||||
|
catAxisLabelColor: "475569",
|
||||||
|
valAxisLabelColor: "475569",
|
||||||
|
valGridLine: { color: "E2E8F0", size: 1 },
|
||||||
|
catAxisLabelFontFace: bodyFont,
|
||||||
|
valAxisLabelFontFace: bodyFont,
|
||||||
|
dataLabelFontFace: bodyFont,
|
||||||
|
legendFontFace: bodyFont,
|
||||||
|
titleFontFace: headFont,
|
||||||
|
...options,
|
||||||
|
};
|
||||||
|
validateTree(chartOptions, `${label}.options`);
|
||||||
|
slide.addChart(chartType, chartData, chartOptions);
|
||||||
|
} else if (type === "table") {
|
||||||
|
if (options.fontFace === undefined) {
|
||||||
|
options.fontFace = bodyFont;
|
||||||
|
}
|
||||||
|
const rows = normalizeTableRows(element.rows, `${label}.rows`);
|
||||||
|
slide.addTable(rows, options);
|
||||||
|
} else {
|
||||||
|
throw new Error(`${label}.type 不支持:${type}`);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if (slideSpec.speaker_notes !== undefined) {
|
||||||
|
const notes = requireString(
|
||||||
|
String(slideSpec.speaker_notes),
|
||||||
|
`slides[${slideIndex}].speaker_notes`,
|
||||||
|
100000,
|
||||||
|
);
|
||||||
|
slide.addNotes(notes);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
await pptx.writeFile({ fileName: outputPath });
|
||||||
|
if (!fs.existsSync(outputPath) || fs.statSync(outputPath).size === 0) {
|
||||||
|
throw new Error("PptxGenJS 未生成有效输出文件");
|
||||||
|
}
|
||||||
|
return {
|
||||||
|
slide_count: slides.length,
|
||||||
|
element_counts: elementCounts,
|
||||||
|
layout,
|
||||||
|
width_inches: dimensions.width,
|
||||||
|
height_inches: dimensions.height,
|
||||||
|
warnings,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
async function main() {
|
||||||
|
const args = parseArgs(process.argv.slice(2));
|
||||||
|
const specPath = path.resolve(args.spec);
|
||||||
|
const outputPath = path.resolve(args.output);
|
||||||
|
if (!fs.existsSync(specPath) || !fs.statSync(specPath).isFile()) {
|
||||||
|
fail(`spec 文件不存在:${specPath}`);
|
||||||
|
}
|
||||||
|
if (path.extname(outputPath).toLowerCase() !== ".pptx") {
|
||||||
|
fail("output 必须使用 .pptx 扩展名");
|
||||||
|
}
|
||||||
|
let spec;
|
||||||
|
try {
|
||||||
|
spec = JSON.parse(fs.readFileSync(specPath, "utf8"));
|
||||||
|
} catch (error) {
|
||||||
|
fail(`读取 spec 失败:${error.message}`);
|
||||||
|
}
|
||||||
|
try {
|
||||||
|
const result = await build(spec, outputPath);
|
||||||
|
process.stdout.write(`${JSON.stringify(result)}\n`);
|
||||||
|
} catch (error) {
|
||||||
|
fail(error instanceof Error ? error.message : String(error));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
main();
|
||||||
91
skills/pptx/scripts/convert_presentation.py
Normal file
91
skills/pptx/scripts/convert_presentation.py
Normal file
@ -0,0 +1,91 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import shutil
|
||||||
|
import tempfile
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from _pptx_common import (
|
||||||
|
PRESENTATION_INPUT_SUFFIXES,
|
||||||
|
SkillArgumentParser,
|
||||||
|
input_file,
|
||||||
|
output_file,
|
||||||
|
publish_file,
|
||||||
|
run_cli,
|
||||||
|
run_soffice_convert,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def build_parser() -> argparse.ArgumentParser:
|
||||||
|
parser = SkillArgumentParser(description="转换 PowerPoint 演示文稿格式。")
|
||||||
|
parser.add_argument("--input", required=True)
|
||||||
|
parser.add_argument("--output", required=True)
|
||||||
|
parser.add_argument("--timeout", type=int, default=180)
|
||||||
|
parser.add_argument("--overwrite", action="store_true")
|
||||||
|
return parser
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> dict[str, Any]:
|
||||||
|
from pptx import Presentation
|
||||||
|
|
||||||
|
args = build_parser().parse_args()
|
||||||
|
source = input_file(args.input, PRESENTATION_INPUT_SUFFIXES)
|
||||||
|
destination = output_file(
|
||||||
|
args.output,
|
||||||
|
{".pptx", ".pdf"},
|
||||||
|
overwrite=args.overwrite,
|
||||||
|
)
|
||||||
|
target_suffix = destination.suffix.lower()
|
||||||
|
if source.suffix.lower() == ".ppt" and target_suffix != ".pptx":
|
||||||
|
raise ValueError("旧版 .ppt 必须先转换为 .pptx,再转换为 PDF")
|
||||||
|
|
||||||
|
with tempfile.TemporaryDirectory(prefix="pptx-convert-") as temp_name:
|
||||||
|
temp_dir = Path(temp_name)
|
||||||
|
staged_source = temp_dir / f"source{source.suffix.lower()}"
|
||||||
|
shutil.copy2(source, staged_source)
|
||||||
|
office_output = {"stdout": "", "stderr": ""}
|
||||||
|
if target_suffix == ".pptx" and source.suffix.lower() == ".pptx":
|
||||||
|
converted = temp_dir / "converted.pptx"
|
||||||
|
shutil.copy2(staged_source, converted)
|
||||||
|
else:
|
||||||
|
converted, office_output = run_soffice_convert(
|
||||||
|
staged_source,
|
||||||
|
target_format=target_suffix.lstrip("."),
|
||||||
|
output_dir=temp_dir / "converted",
|
||||||
|
timeout=args.timeout,
|
||||||
|
filter_name=(
|
||||||
|
"Impress MS PowerPoint 2007 XML"
|
||||||
|
if target_suffix == ".pptx"
|
||||||
|
else None
|
||||||
|
),
|
||||||
|
)
|
||||||
|
if target_suffix == ".pptx":
|
||||||
|
presentation = Presentation(str(converted))
|
||||||
|
page_count = len(presentation.slides)
|
||||||
|
if page_count < 1:
|
||||||
|
raise ValueError("转换后的演示文稿不包含页面")
|
||||||
|
else:
|
||||||
|
from pypdf import PdfReader
|
||||||
|
|
||||||
|
page_count = len(PdfReader(str(converted)).pages)
|
||||||
|
if page_count < 1:
|
||||||
|
raise ValueError("转换后的 PDF 不包含页面")
|
||||||
|
staged_output = temp_dir / f"publish{target_suffix}"
|
||||||
|
shutil.copy2(converted, staged_output)
|
||||||
|
publish_file(staged_output, destination, overwrite=args.overwrite)
|
||||||
|
return {
|
||||||
|
"source": str(source),
|
||||||
|
"path": str(destination),
|
||||||
|
"format": target_suffix.lstrip("."),
|
||||||
|
"page_count": page_count,
|
||||||
|
"size_bytes": destination.stat().st_size,
|
||||||
|
"office_stdout": office_output["stdout"],
|
||||||
|
"office_stderr": office_output["stderr"],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(run_cli(main))
|
||||||
95
skills/pptx/scripts/create_presentation.py
Normal file
95
skills/pptx/scripts/create_presentation.py
Normal file
@ -0,0 +1,95 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import tempfile
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from _pptx_common import (
|
||||||
|
SkillArgumentParser,
|
||||||
|
find_program,
|
||||||
|
load_json_argument,
|
||||||
|
output_file,
|
||||||
|
publish_file,
|
||||||
|
run_cli,
|
||||||
|
run_program,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def build_parser() -> argparse.ArgumentParser:
|
||||||
|
parser = SkillArgumentParser(description="按受控 JSON 说明创建 PowerPoint 演示文稿。")
|
||||||
|
parser.add_argument("--output", required=True)
|
||||||
|
parser.add_argument("--spec")
|
||||||
|
parser.add_argument("--spec-file")
|
||||||
|
parser.add_argument("--timeout", type=int, default=180)
|
||||||
|
parser.add_argument("--overwrite", action="store_true")
|
||||||
|
return parser
|
||||||
|
|
||||||
|
|
||||||
|
def _validate_spec(spec: dict[str, Any]) -> None:
|
||||||
|
slides = spec.get("slides")
|
||||||
|
if not isinstance(slides, list) or not slides:
|
||||||
|
raise ValueError("spec.slides 必须是非空数组")
|
||||||
|
if len(slides) > 300:
|
||||||
|
raise ValueError("单个演示文稿最多支持 300 页")
|
||||||
|
for index, slide in enumerate(slides):
|
||||||
|
if not isinstance(slide, dict):
|
||||||
|
raise ValueError(f"spec.slides[{index}] 必须是对象")
|
||||||
|
elements = slide.get("elements", [])
|
||||||
|
if not isinstance(elements, list):
|
||||||
|
raise ValueError(f"spec.slides[{index}].elements 必须是数组")
|
||||||
|
if len(elements) > 1000:
|
||||||
|
raise ValueError(f"spec.slides[{index}].elements 超过 1000 项限制")
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> dict[str, Any]:
|
||||||
|
args = build_parser().parse_args()
|
||||||
|
destination = output_file(args.output, {".pptx"}, overwrite=args.overwrite)
|
||||||
|
spec = load_json_argument(args.spec, args.spec_file, label="演示文稿说明")
|
||||||
|
_validate_spec(spec)
|
||||||
|
builder = Path(__file__).with_name("_presentation_builder.js").resolve()
|
||||||
|
if not builder.is_file():
|
||||||
|
raise FileNotFoundError(f"内部构建器不存在:{builder}")
|
||||||
|
|
||||||
|
with tempfile.TemporaryDirectory(prefix="pptx-create-") as temp_name:
|
||||||
|
temp_dir = Path(temp_name)
|
||||||
|
spec_path = temp_dir / "spec.json"
|
||||||
|
staged_output = temp_dir / "presentation.pptx"
|
||||||
|
spec_path.write_text(
|
||||||
|
json.dumps(spec, ensure_ascii=False),
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
completed = run_program(
|
||||||
|
[
|
||||||
|
find_program("node"),
|
||||||
|
str(builder),
|
||||||
|
"--spec",
|
||||||
|
str(spec_path),
|
||||||
|
"--output",
|
||||||
|
str(staged_output),
|
||||||
|
],
|
||||||
|
timeout=args.timeout,
|
||||||
|
cwd=Path.cwd(),
|
||||||
|
env=os.environ.copy(),
|
||||||
|
)
|
||||||
|
if not staged_output.is_file() or staged_output.stat().st_size <= 0:
|
||||||
|
raise RuntimeError("PptxGenJS 未生成输出文件")
|
||||||
|
try:
|
||||||
|
builder_result = json.loads(completed.stdout.strip())
|
||||||
|
except json.JSONDecodeError as exc:
|
||||||
|
raise RuntimeError("内部构建器返回了无效结果") from exc
|
||||||
|
publish_file(staged_output, destination, overwrite=args.overwrite)
|
||||||
|
|
||||||
|
return {
|
||||||
|
"path": str(destination),
|
||||||
|
"size_bytes": destination.stat().st_size,
|
||||||
|
**builder_result,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(run_cli(main))
|
||||||
271
skills/pptx/scripts/download_presentation.py
Normal file
271
skills/pptx/scripts/download_presentation.py
Normal file
@ -0,0 +1,271 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import os
|
||||||
|
import socket
|
||||||
|
import sys
|
||||||
|
import tempfile
|
||||||
|
import urllib.error
|
||||||
|
import urllib.parse
|
||||||
|
import urllib.request
|
||||||
|
import zipfile
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any, NoReturn, Optional
|
||||||
|
|
||||||
|
from _pptx_common import (
|
||||||
|
PRESENTATION_INPUT_SUFFIXES,
|
||||||
|
emit,
|
||||||
|
failure_message,
|
||||||
|
inspect_archive,
|
||||||
|
output_file,
|
||||||
|
parse_xml_bytes,
|
||||||
|
publish_file,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
DEFAULT_TIMEOUT_SECONDS = 60
|
||||||
|
DEFAULT_MAX_BYTES = 100 * 1024 * 1024
|
||||||
|
MAX_ALLOWED_BYTES = 512 * 1024 * 1024
|
||||||
|
CHUNK_SIZE = 1024 * 1024
|
||||||
|
USER_AGENT = "wechat-robot-pptx-skill/1.0"
|
||||||
|
OLE_COMPOUND_MAGIC = bytes.fromhex("D0CF11E0A1B11AE1")
|
||||||
|
CONTENT_TYPES_NS = "http://schemas.openxmlformats.org/package/2006/content-types"
|
||||||
|
PRESENTATION_CONTENT_TYPES = {
|
||||||
|
(
|
||||||
|
"application/vnd.openxmlformats-officedocument."
|
||||||
|
"presentationml.presentation.main+xml"
|
||||||
|
): ".pptx",
|
||||||
|
(
|
||||||
|
"application/vnd.openxmlformats-officedocument."
|
||||||
|
"presentationml.template.main+xml"
|
||||||
|
): ".potx",
|
||||||
|
(
|
||||||
|
"application/vnd.openxmlformats-officedocument."
|
||||||
|
"presentationml.slideshow.main+xml"
|
||||||
|
): ".ppsx",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
class SkillArgumentParser(argparse.ArgumentParser):
|
||||||
|
def error(self, message: str) -> NoReturn:
|
||||||
|
raise ValueError(f"参数错误:{message}")
|
||||||
|
|
||||||
|
|
||||||
|
def _validate_https_url(value: str) -> str:
|
||||||
|
url = value.strip()
|
||||||
|
if not url:
|
||||||
|
raise ValueError("演示文稿 URL 不能为空")
|
||||||
|
if any(character.isspace() or ord(character) < 32 for character in url):
|
||||||
|
raise ValueError("演示文稿 URL 不能包含空白字符或控制字符")
|
||||||
|
parsed = urllib.parse.urlsplit(url)
|
||||||
|
if parsed.scheme.lower() != "https" or not parsed.hostname:
|
||||||
|
raise ValueError("演示文稿 URL 必须是有效的 HTTPS 地址")
|
||||||
|
if parsed.username is not None or parsed.password is not None:
|
||||||
|
raise ValueError("演示文稿 URL 不允许包含用户名或密码")
|
||||||
|
try:
|
||||||
|
parsed.port
|
||||||
|
except ValueError as exc:
|
||||||
|
raise ValueError("演示文稿 URL 端口格式不正确") from exc
|
||||||
|
return url
|
||||||
|
|
||||||
|
|
||||||
|
class HTTPSOnlyRedirectHandler(urllib.request.HTTPRedirectHandler):
|
||||||
|
def redirect_request(self, req, fp, code, msg, headers, newurl):
|
||||||
|
return super().redirect_request(
|
||||||
|
req,
|
||||||
|
fp,
|
||||||
|
code,
|
||||||
|
msg,
|
||||||
|
headers,
|
||||||
|
_validate_https_url(newurl),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _parse_args(argv: list[str]) -> argparse.Namespace:
|
||||||
|
parser = SkillArgumentParser(description="下载并校验远程 HTTPS 演示文稿")
|
||||||
|
parser.add_argument("--url", required=True)
|
||||||
|
parser.add_argument("--output", required=True)
|
||||||
|
parser.add_argument("--timeout", type=int, default=DEFAULT_TIMEOUT_SECONDS)
|
||||||
|
parser.add_argument("--max-bytes", type=int, default=DEFAULT_MAX_BYTES)
|
||||||
|
parser.add_argument("--overwrite", action="store_true")
|
||||||
|
args = parser.parse_args(argv)
|
||||||
|
args.url = _validate_https_url(args.url)
|
||||||
|
if args.timeout < 1 or args.timeout > 600:
|
||||||
|
raise ValueError("timeout 必须在 1 到 600 秒之间")
|
||||||
|
if args.max_bytes < 1 or args.max_bytes > MAX_ALLOWED_BYTES:
|
||||||
|
raise ValueError(f"max-bytes 必须在 1 到 {MAX_ALLOWED_BYTES} 之间")
|
||||||
|
args.output = output_file(
|
||||||
|
args.output,
|
||||||
|
PRESENTATION_INPUT_SUFFIXES,
|
||||||
|
overwrite=args.overwrite,
|
||||||
|
)
|
||||||
|
return args
|
||||||
|
|
||||||
|
|
||||||
|
def _detect_ooxml_suffix(path: Path) -> str:
|
||||||
|
with zipfile.ZipFile(path) as archive:
|
||||||
|
root = parse_xml_bytes(
|
||||||
|
archive.read("[Content_Types].xml"),
|
||||||
|
label="[Content_Types].xml",
|
||||||
|
)
|
||||||
|
for element in root.iter(f"{{{CONTENT_TYPES_NS}}}Override"):
|
||||||
|
if element.attrib.get("PartName") != "/ppt/presentation.xml":
|
||||||
|
continue
|
||||||
|
detected = PRESENTATION_CONTENT_TYPES.get(
|
||||||
|
element.attrib.get("ContentType", "")
|
||||||
|
)
|
||||||
|
if detected:
|
||||||
|
return detected
|
||||||
|
raise ValueError("下载内容不是受支持的 PowerPoint OOXML 文件")
|
||||||
|
|
||||||
|
|
||||||
|
def _validate_ooxml(path: Path, expected_suffix: str) -> dict[str, Any]:
|
||||||
|
archive = inspect_archive(path)
|
||||||
|
if archive["missing_required_parts"]:
|
||||||
|
raise ValueError(
|
||||||
|
"下载内容不是有效的 PowerPoint OOXML 文件;缺少:"
|
||||||
|
+ "、".join(archive["missing_required_parts"])
|
||||||
|
)
|
||||||
|
if archive["duplicate_members"]:
|
||||||
|
raise ValueError(
|
||||||
|
"PowerPoint 压缩包含重复成员:"
|
||||||
|
+ "、".join(archive["duplicate_members"][:10])
|
||||||
|
)
|
||||||
|
detected_suffix = _detect_ooxml_suffix(path)
|
||||||
|
if detected_suffix != expected_suffix:
|
||||||
|
raise ValueError(
|
||||||
|
"下载内容的实际格式为 "
|
||||||
|
f"{detected_suffix},但 output 使用了 {expected_suffix}"
|
||||||
|
)
|
||||||
|
|
||||||
|
try:
|
||||||
|
from pptx import Presentation
|
||||||
|
except ImportError as exc:
|
||||||
|
raise RuntimeError("当前 Python 未加载环境预置的 python-pptx 模块") from exc
|
||||||
|
try:
|
||||||
|
presentation = Presentation(str(path))
|
||||||
|
slide_count = len(presentation.slides)
|
||||||
|
except Exception as exc:
|
||||||
|
raise ValueError("下载内容不是可解析的 PowerPoint 演示文稿") from exc
|
||||||
|
return {
|
||||||
|
"format": detected_suffix.lstrip("."),
|
||||||
|
"slide_count": slide_count,
|
||||||
|
"validation": "ooxml-and-python-pptx",
|
||||||
|
"archive": archive,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _validate_legacy_ppt(path: Path) -> dict[str, Any]:
|
||||||
|
with path.open("rb") as stream:
|
||||||
|
magic = stream.read(len(OLE_COMPOUND_MAGIC))
|
||||||
|
if magic != OLE_COMPOUND_MAGIC:
|
||||||
|
raise ValueError("下载内容不是有效的旧版 PowerPoint 复合文件")
|
||||||
|
return {
|
||||||
|
"format": "ppt",
|
||||||
|
"validation": "ole-compound-signature",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _validate_presentation(path: Path, suffix: str) -> dict[str, Any]:
|
||||||
|
if suffix == ".ppt":
|
||||||
|
return _validate_legacy_ppt(path)
|
||||||
|
return _validate_ooxml(path, suffix)
|
||||||
|
|
||||||
|
|
||||||
|
def _download(args: argparse.Namespace) -> dict[str, Any]:
|
||||||
|
output: Path = args.output
|
||||||
|
request = urllib.request.Request(
|
||||||
|
args.url,
|
||||||
|
headers={
|
||||||
|
"Accept": (
|
||||||
|
"application/vnd.openxmlformats-officedocument."
|
||||||
|
"presentationml.presentation,"
|
||||||
|
"application/vnd.ms-powerpoint,"
|
||||||
|
"application/octet-stream;q=0.9,*/*;q=0.1"
|
||||||
|
),
|
||||||
|
"Accept-Encoding": "identity",
|
||||||
|
"User-Agent": USER_AGENT,
|
||||||
|
},
|
||||||
|
method="GET",
|
||||||
|
)
|
||||||
|
opener = urllib.request.build_opener(HTTPSOnlyRedirectHandler())
|
||||||
|
temp_path: Optional[Path] = None
|
||||||
|
downloaded_bytes = 0
|
||||||
|
try:
|
||||||
|
with tempfile.NamedTemporaryFile(
|
||||||
|
mode="wb",
|
||||||
|
prefix=f".{output.stem}.",
|
||||||
|
suffix=f".part{output.suffix}",
|
||||||
|
dir=str(output.parent),
|
||||||
|
delete=False,
|
||||||
|
) as temp_file:
|
||||||
|
temp_path = Path(temp_file.name)
|
||||||
|
with opener.open(request, timeout=args.timeout) as response:
|
||||||
|
_validate_https_url(response.geturl())
|
||||||
|
content_length = response.headers.get("Content-Length")
|
||||||
|
if content_length:
|
||||||
|
try:
|
||||||
|
expected_bytes = int(content_length)
|
||||||
|
except ValueError:
|
||||||
|
expected_bytes = 0
|
||||||
|
if expected_bytes > args.max_bytes:
|
||||||
|
raise ValueError(
|
||||||
|
f"远程文件超过大小限制:最多允许 {args.max_bytes} 字节"
|
||||||
|
)
|
||||||
|
while True:
|
||||||
|
chunk = response.read(CHUNK_SIZE)
|
||||||
|
if not chunk:
|
||||||
|
break
|
||||||
|
downloaded_bytes += len(chunk)
|
||||||
|
if downloaded_bytes > args.max_bytes:
|
||||||
|
raise ValueError(
|
||||||
|
f"远程文件超过大小限制:最多允许 {args.max_bytes} 字节"
|
||||||
|
)
|
||||||
|
temp_file.write(chunk)
|
||||||
|
temp_file.flush()
|
||||||
|
os.fsync(temp_file.fileno())
|
||||||
|
if downloaded_bytes == 0:
|
||||||
|
raise ValueError("远程服务器返回了空文件")
|
||||||
|
details = _validate_presentation(temp_path, output.suffix.lower())
|
||||||
|
publish_file(temp_path, output, overwrite=args.overwrite)
|
||||||
|
temp_path = None
|
||||||
|
return {
|
||||||
|
"path": str(output),
|
||||||
|
"size_bytes": downloaded_bytes,
|
||||||
|
**details,
|
||||||
|
}
|
||||||
|
finally:
|
||||||
|
if temp_path is not None:
|
||||||
|
try:
|
||||||
|
temp_path.unlink(missing_ok=True)
|
||||||
|
except OSError:
|
||||||
|
pass
|
||||||
|
|
||||||
|
|
||||||
|
def _failure_message(exc: Exception) -> str:
|
||||||
|
if isinstance(exc, urllib.error.HTTPError):
|
||||||
|
return f"下载失败:远程服务器返回 HTTP {exc.code}"
|
||||||
|
if isinstance(exc, (TimeoutError, socket.timeout)):
|
||||||
|
return "下载失败:连接或读取超时"
|
||||||
|
if isinstance(exc, urllib.error.URLError):
|
||||||
|
if isinstance(exc.reason, (TimeoutError, socket.timeout)):
|
||||||
|
return "下载失败:连接或读取超时"
|
||||||
|
return "下载失败:无法访问远程服务器"
|
||||||
|
return failure_message(exc)
|
||||||
|
|
||||||
|
|
||||||
|
def main(argv: Optional[list[str]] = None) -> int:
|
||||||
|
try:
|
||||||
|
args = _parse_args(sys.argv[1:] if argv is None else argv)
|
||||||
|
result = _download(args)
|
||||||
|
except Exception as exc:
|
||||||
|
emit({"ok": False, "error": _failure_message(exc)})
|
||||||
|
return 1
|
||||||
|
emit({"ok": True, **result})
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
292
skills/pptx/scripts/duplicate_slide.py
Normal file
292
skills/pptx/scripts/duplicate_slide.py
Normal file
@ -0,0 +1,292 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import posixpath
|
||||||
|
import re
|
||||||
|
import tempfile
|
||||||
|
import zipfile
|
||||||
|
from pathlib import Path, PurePosixPath
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from _pptx_common import (
|
||||||
|
P_NS,
|
||||||
|
R_NS,
|
||||||
|
REL_NS,
|
||||||
|
SkillArgumentParser,
|
||||||
|
input_file,
|
||||||
|
inspect_archive,
|
||||||
|
output_file,
|
||||||
|
publish_file,
|
||||||
|
run_cli,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
CONTENT_TYPES_NS = "http://schemas.openxmlformats.org/package/2006/content-types"
|
||||||
|
SLIDE_REL_TYPE_SUFFIX = "/slide"
|
||||||
|
SLIDE_CONTENT_TYPE = (
|
||||||
|
"application/vnd.openxmlformats-officedocument."
|
||||||
|
"presentationml.slide+xml"
|
||||||
|
)
|
||||||
|
DROP_RELATIONSHIP_SUFFIXES = ("/notesSlide", "/comments", "/comment")
|
||||||
|
|
||||||
|
|
||||||
|
def build_parser() -> argparse.ArgumentParser:
|
||||||
|
parser = SkillArgumentParser(
|
||||||
|
description="安全复制现有 PPTX 页面并更新 OOXML 包关系。"
|
||||||
|
)
|
||||||
|
parser.add_argument("--input", required=True)
|
||||||
|
parser.add_argument("--output", required=True)
|
||||||
|
parser.add_argument("--slide", type=int, required=True)
|
||||||
|
parser.add_argument("--after", type=int)
|
||||||
|
parser.add_argument("--overwrite", action="store_true")
|
||||||
|
return parser
|
||||||
|
|
||||||
|
|
||||||
|
def _parse_xml(payload: bytes, label: str) -> Any:
|
||||||
|
from lxml import etree
|
||||||
|
|
||||||
|
parser = etree.XMLParser(
|
||||||
|
resolve_entities=False,
|
||||||
|
no_network=True,
|
||||||
|
recover=False,
|
||||||
|
huge_tree=False,
|
||||||
|
remove_blank_text=False,
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
return etree.fromstring(payload, parser=parser)
|
||||||
|
except etree.XMLSyntaxError as exc:
|
||||||
|
raise ValueError(f"{label} 解析失败:{exc}") from exc
|
||||||
|
|
||||||
|
|
||||||
|
def _serialize_xml(root: Any) -> bytes:
|
||||||
|
from lxml import etree
|
||||||
|
|
||||||
|
return etree.tostring(
|
||||||
|
root,
|
||||||
|
encoding="UTF-8",
|
||||||
|
xml_declaration=True,
|
||||||
|
standalone=True,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _next_relationship_id(existing: set[str]) -> str:
|
||||||
|
number = 1
|
||||||
|
while f"rId{number}" in existing:
|
||||||
|
number += 1
|
||||||
|
return f"rId{number}"
|
||||||
|
|
||||||
|
|
||||||
|
def _slide_rels_name(slide_part: str) -> str:
|
||||||
|
path = PurePosixPath(slide_part)
|
||||||
|
return (path.parent / "_rels" / f"{path.name}.rels").as_posix()
|
||||||
|
|
||||||
|
|
||||||
|
def _resolve_part_target(owner_part: str, target: str) -> str:
|
||||||
|
if not target or target.startswith("#"):
|
||||||
|
raise ValueError(f"关系目标无效:{target!r}")
|
||||||
|
if target.startswith("/"):
|
||||||
|
normalized = posixpath.normpath(target).lstrip("/")
|
||||||
|
else:
|
||||||
|
normalized = posixpath.normpath(
|
||||||
|
posixpath.join(posixpath.dirname(owner_part), target)
|
||||||
|
)
|
||||||
|
if normalized == ".." or normalized.startswith("../"):
|
||||||
|
raise ValueError(f"关系目标逃逸 OOXML 包根目录:{target}")
|
||||||
|
return normalized.lstrip("/")
|
||||||
|
|
||||||
|
|
||||||
|
def _duplicate(
|
||||||
|
source: Path,
|
||||||
|
staged: Path,
|
||||||
|
*,
|
||||||
|
source_slide: int,
|
||||||
|
insert_after: int,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
with zipfile.ZipFile(source, "r") as incoming:
|
||||||
|
infos = incoming.infolist()
|
||||||
|
names = {info.filename for info in infos}
|
||||||
|
presentation_root = _parse_xml(
|
||||||
|
incoming.read("ppt/presentation.xml"),
|
||||||
|
"ppt/presentation.xml",
|
||||||
|
)
|
||||||
|
presentation_rels_root = _parse_xml(
|
||||||
|
incoming.read("ppt/_rels/presentation.xml.rels"),
|
||||||
|
"ppt/_rels/presentation.xml.rels",
|
||||||
|
)
|
||||||
|
content_types_root = _parse_xml(
|
||||||
|
incoming.read("[Content_Types].xml"),
|
||||||
|
"[Content_Types].xml",
|
||||||
|
)
|
||||||
|
|
||||||
|
slide_id_list = presentation_root.find(f"{{{P_NS}}}sldIdLst")
|
||||||
|
if slide_id_list is None:
|
||||||
|
raise ValueError("presentation.xml 缺少 sldIdLst")
|
||||||
|
slide_ids = list(slide_id_list)
|
||||||
|
if source_slide < 1 or source_slide > len(slide_ids):
|
||||||
|
raise ValueError(f"slide 超出页面总数 {len(slide_ids)}")
|
||||||
|
if insert_after < 1 or insert_after > len(slide_ids):
|
||||||
|
raise ValueError(f"after 超出页面总数 {len(slide_ids)}")
|
||||||
|
|
||||||
|
rel_targets: dict[str, str] = {}
|
||||||
|
existing_rel_ids: set[str] = set()
|
||||||
|
slide_relationship_type = None
|
||||||
|
for relationship in presentation_rels_root:
|
||||||
|
relationship_id = relationship.attrib.get("Id", "")
|
||||||
|
existing_rel_ids.add(relationship_id)
|
||||||
|
target = relationship.attrib.get("Target", "")
|
||||||
|
if relationship.attrib.get("TargetMode") == "External":
|
||||||
|
continue
|
||||||
|
posix_target = _resolve_part_target(
|
||||||
|
"ppt/presentation.xml",
|
||||||
|
target,
|
||||||
|
)
|
||||||
|
rel_targets[relationship_id] = posix_target
|
||||||
|
if relationship.attrib.get("Type", "").endswith(SLIDE_REL_TYPE_SUFFIX):
|
||||||
|
slide_relationship_type = relationship.attrib.get("Type")
|
||||||
|
if not slide_relationship_type:
|
||||||
|
raise ValueError("presentation.xml.rels 不包含 slide 关系类型")
|
||||||
|
|
||||||
|
source_rel_id = slide_ids[source_slide - 1].attrib.get(f"{{{R_NS}}}id")
|
||||||
|
source_part = rel_targets.get(source_rel_id or "")
|
||||||
|
if not source_part or source_part not in names:
|
||||||
|
raise ValueError("无法解析待复制页面的 slide 部件")
|
||||||
|
|
||||||
|
existing_numbers = [
|
||||||
|
int(match.group(1))
|
||||||
|
for name in names
|
||||||
|
if (
|
||||||
|
match := re.fullmatch(r"ppt/slides/slide(\d+)\.xml", name)
|
||||||
|
)
|
||||||
|
]
|
||||||
|
new_part_number = max(existing_numbers, default=0) + 1
|
||||||
|
new_part = f"ppt/slides/slide{new_part_number}.xml"
|
||||||
|
new_rels_part = _slide_rels_name(new_part)
|
||||||
|
new_rel_id = _next_relationship_id(existing_rel_ids)
|
||||||
|
numeric_ids = [
|
||||||
|
int(item.attrib["id"])
|
||||||
|
for item in slide_ids
|
||||||
|
if item.attrib.get("id", "").isdigit()
|
||||||
|
]
|
||||||
|
new_slide_id = str(max(numeric_ids, default=255) + 1)
|
||||||
|
|
||||||
|
from lxml import etree
|
||||||
|
|
||||||
|
new_relationship = etree.Element(
|
||||||
|
f"{{{REL_NS}}}Relationship",
|
||||||
|
Id=new_rel_id,
|
||||||
|
Type=slide_relationship_type,
|
||||||
|
Target=f"slides/slide{new_part_number}.xml",
|
||||||
|
)
|
||||||
|
presentation_rels_root.append(new_relationship)
|
||||||
|
new_slide_id_element = etree.Element(
|
||||||
|
f"{{{P_NS}}}sldId",
|
||||||
|
id=new_slide_id,
|
||||||
|
)
|
||||||
|
new_slide_id_element.set(f"{{{R_NS}}}id", new_rel_id)
|
||||||
|
slide_id_list.insert(insert_after, new_slide_id_element)
|
||||||
|
|
||||||
|
override_exists = any(
|
||||||
|
node.attrib.get("PartName") == f"/{new_part}"
|
||||||
|
for node in content_types_root.findall(
|
||||||
|
f"{{{CONTENT_TYPES_NS}}}Override"
|
||||||
|
)
|
||||||
|
)
|
||||||
|
if not override_exists:
|
||||||
|
content_types_root.append(
|
||||||
|
etree.Element(
|
||||||
|
f"{{{CONTENT_TYPES_NS}}}Override",
|
||||||
|
PartName=f"/{new_part}",
|
||||||
|
ContentType=SLIDE_CONTENT_TYPE,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
replacements = {
|
||||||
|
"ppt/presentation.xml": _serialize_xml(presentation_root),
|
||||||
|
"ppt/_rels/presentation.xml.rels": _serialize_xml(
|
||||||
|
presentation_rels_root
|
||||||
|
),
|
||||||
|
"[Content_Types].xml": _serialize_xml(content_types_root),
|
||||||
|
}
|
||||||
|
additions = {new_part: incoming.read(source_part)}
|
||||||
|
source_rels_part = _slide_rels_name(source_part)
|
||||||
|
dropped_relationships: list[str] = []
|
||||||
|
shared_relationships: list[dict[str, str]] = []
|
||||||
|
if source_rels_part in names:
|
||||||
|
slide_rels_root = _parse_xml(
|
||||||
|
incoming.read(source_rels_part),
|
||||||
|
source_rels_part,
|
||||||
|
)
|
||||||
|
for relationship in list(slide_rels_root):
|
||||||
|
relationship_type = relationship.attrib.get("Type", "")
|
||||||
|
if relationship_type.endswith(DROP_RELATIONSHIP_SUFFIXES):
|
||||||
|
dropped_relationships.append(relationship_type)
|
||||||
|
slide_rels_root.remove(relationship)
|
||||||
|
continue
|
||||||
|
if relationship_type.endswith(
|
||||||
|
("/chart", "/diagramData", "/diagramDrawing", "/oleObject")
|
||||||
|
):
|
||||||
|
shared_relationships.append(
|
||||||
|
{
|
||||||
|
"type": relationship_type,
|
||||||
|
"target": relationship.attrib.get("Target", ""),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
additions[new_rels_part] = _serialize_xml(slide_rels_root)
|
||||||
|
|
||||||
|
with zipfile.ZipFile(
|
||||||
|
staged,
|
||||||
|
"w",
|
||||||
|
compression=zipfile.ZIP_DEFLATED,
|
||||||
|
compresslevel=6,
|
||||||
|
) as outgoing:
|
||||||
|
for info in infos:
|
||||||
|
payload = replacements.get(info.filename)
|
||||||
|
if payload is None:
|
||||||
|
payload = incoming.read(info.filename)
|
||||||
|
outgoing.writestr(info, payload)
|
||||||
|
for name, payload in additions.items():
|
||||||
|
outgoing.writestr(name, payload)
|
||||||
|
|
||||||
|
archive = inspect_archive(staged)
|
||||||
|
return {
|
||||||
|
"source_slide": source_slide,
|
||||||
|
"insert_after": insert_after,
|
||||||
|
"new_slide": insert_after + 1,
|
||||||
|
"new_slide_part": new_part,
|
||||||
|
"dropped_relationship_types": dropped_relationships,
|
||||||
|
"shared_relationships": shared_relationships,
|
||||||
|
"archive": archive,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> dict[str, Any]:
|
||||||
|
from pptx import Presentation
|
||||||
|
|
||||||
|
args = build_parser().parse_args()
|
||||||
|
source = input_file(args.input, {".pptx"})
|
||||||
|
destination = output_file(args.output, {".pptx"}, overwrite=args.overwrite)
|
||||||
|
if source == destination:
|
||||||
|
raise ValueError("不能覆盖输入演示文稿;请使用新的 output 路径")
|
||||||
|
insert_after = args.slide if args.after is None else args.after
|
||||||
|
with tempfile.TemporaryDirectory(prefix="pptx-duplicate-") as temp_name:
|
||||||
|
staged = Path(temp_name) / "duplicated.pptx"
|
||||||
|
result = _duplicate(
|
||||||
|
source,
|
||||||
|
staged,
|
||||||
|
source_slide=args.slide,
|
||||||
|
insert_after=insert_after,
|
||||||
|
)
|
||||||
|
presentation = Presentation(str(staged))
|
||||||
|
result["slide_count"] = len(presentation.slides)
|
||||||
|
publish_file(staged, destination, overwrite=args.overwrite)
|
||||||
|
return {
|
||||||
|
"source": str(source),
|
||||||
|
"path": str(destination),
|
||||||
|
**result,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(run_cli(main))
|
||||||
367
skills/pptx/scripts/edit_presentation.py
Normal file
367
skills/pptx/scripts/edit_presentation.py
Normal file
@ -0,0 +1,367 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import re
|
||||||
|
import tempfile
|
||||||
|
import zipfile
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any, Iterable, Optional
|
||||||
|
|
||||||
|
from _pptx_common import (
|
||||||
|
OOXML_PRESENTATION_SUFFIXES,
|
||||||
|
SkillArgumentParser,
|
||||||
|
input_file,
|
||||||
|
load_json_argument,
|
||||||
|
output_file,
|
||||||
|
parse_xml_bytes,
|
||||||
|
publish_file,
|
||||||
|
run_cli,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def build_parser() -> argparse.ArgumentParser:
|
||||||
|
parser = SkillArgumentParser(description="按受控操作编辑 PowerPoint 演示文稿。")
|
||||||
|
parser.add_argument("--input", required=True)
|
||||||
|
parser.add_argument("--output", required=True)
|
||||||
|
parser.add_argument("--spec")
|
||||||
|
parser.add_argument("--spec-file")
|
||||||
|
parser.add_argument("--allow-external-links", action="store_true")
|
||||||
|
parser.add_argument("--overwrite", action="store_true")
|
||||||
|
return parser
|
||||||
|
|
||||||
|
|
||||||
|
def _external_relationship_count(path: Path) -> int:
|
||||||
|
count = 0
|
||||||
|
with zipfile.ZipFile(path) as archive:
|
||||||
|
for name in archive.namelist():
|
||||||
|
if not name.endswith(".rels"):
|
||||||
|
continue
|
||||||
|
root = parse_xml_bytes(archive.read(name), label=name)
|
||||||
|
count += sum(
|
||||||
|
1
|
||||||
|
for node in root
|
||||||
|
if node.attrib.get("TargetMode") == "External"
|
||||||
|
)
|
||||||
|
return count
|
||||||
|
|
||||||
|
|
||||||
|
def _package_risks(path: Path) -> list[str]:
|
||||||
|
risks: list[str] = []
|
||||||
|
with zipfile.ZipFile(path) as archive:
|
||||||
|
names = archive.namelist()
|
||||||
|
if any(name.startswith("ppt/comments/") for name in names):
|
||||||
|
risks.append("comments")
|
||||||
|
if any(
|
||||||
|
name.endswith((".bin", ".vbaProject"))
|
||||||
|
for name in names
|
||||||
|
):
|
||||||
|
risks.append("binary_embedded_objects_or_macros")
|
||||||
|
if any(name.startswith("ppt/activeX/") for name in names):
|
||||||
|
risks.append("activex")
|
||||||
|
return risks
|
||||||
|
|
||||||
|
|
||||||
|
def _iter_shapes(shapes: Any) -> Iterable[Any]:
|
||||||
|
from pptx.enum.shapes import MSO_SHAPE_TYPE
|
||||||
|
|
||||||
|
for shape in shapes:
|
||||||
|
yield shape
|
||||||
|
if shape.shape_type == MSO_SHAPE_TYPE.GROUP:
|
||||||
|
yield from _iter_shapes(shape.shapes)
|
||||||
|
|
||||||
|
|
||||||
|
def _iter_text_frames(slide: Any, *, include_notes: bool) -> Iterable[Any]:
|
||||||
|
for shape in _iter_shapes(slide.shapes):
|
||||||
|
if getattr(shape, "has_text_frame", False):
|
||||||
|
yield shape.text_frame
|
||||||
|
if getattr(shape, "has_table", False):
|
||||||
|
for row in shape.table.rows:
|
||||||
|
for cell in row.cells:
|
||||||
|
yield cell.text_frame
|
||||||
|
if include_notes:
|
||||||
|
try:
|
||||||
|
text_frame = slide.notes_slide.notes_text_frame
|
||||||
|
if text_frame is not None:
|
||||||
|
yield text_frame
|
||||||
|
except (AttributeError, KeyError, ValueError):
|
||||||
|
pass
|
||||||
|
|
||||||
|
|
||||||
|
def _find_spans(
|
||||||
|
text: str,
|
||||||
|
needle: str,
|
||||||
|
*,
|
||||||
|
match_case: bool,
|
||||||
|
whole_word: bool,
|
||||||
|
limit: Optional[int],
|
||||||
|
) -> list[tuple[int, int]]:
|
||||||
|
flags = 0 if match_case else re.IGNORECASE
|
||||||
|
escaped = re.escape(needle)
|
||||||
|
if whole_word:
|
||||||
|
escaped = rf"(?<!\w){escaped}(?!\w)"
|
||||||
|
matches = list(re.finditer(escaped, text, flags))
|
||||||
|
if limit is not None:
|
||||||
|
matches = matches[:limit]
|
||||||
|
return [(match.start(), match.end()) for match in matches]
|
||||||
|
|
||||||
|
|
||||||
|
def _run_at_offset(runs: list[Any], offset: int, *, end: bool = False) -> tuple[int, int]:
|
||||||
|
cursor = 0
|
||||||
|
for index, run in enumerate(runs):
|
||||||
|
next_cursor = cursor + len(run.text)
|
||||||
|
if offset < next_cursor or (end and offset == next_cursor):
|
||||||
|
return index, offset - cursor
|
||||||
|
cursor = next_cursor
|
||||||
|
if not runs:
|
||||||
|
raise ValueError("段落没有可编辑的文本 Run")
|
||||||
|
return len(runs) - 1, len(runs[-1].text)
|
||||||
|
|
||||||
|
|
||||||
|
def _replace_in_paragraph(
|
||||||
|
paragraph: Any,
|
||||||
|
needle: str,
|
||||||
|
replacement: str,
|
||||||
|
*,
|
||||||
|
match_case: bool,
|
||||||
|
whole_word: bool,
|
||||||
|
limit: Optional[int],
|
||||||
|
) -> int:
|
||||||
|
runs = list(paragraph.runs)
|
||||||
|
if not runs:
|
||||||
|
return 0
|
||||||
|
text = "".join(run.text for run in runs)
|
||||||
|
spans = _find_spans(
|
||||||
|
text,
|
||||||
|
needle,
|
||||||
|
match_case=match_case,
|
||||||
|
whole_word=whole_word,
|
||||||
|
limit=limit,
|
||||||
|
)
|
||||||
|
for start, end in reversed(spans):
|
||||||
|
start_index, start_offset = _run_at_offset(runs, start)
|
||||||
|
end_index, end_offset = _run_at_offset(runs, end, end=True)
|
||||||
|
if start_index == end_index:
|
||||||
|
original = runs[start_index].text
|
||||||
|
runs[start_index].text = (
|
||||||
|
original[:start_offset] + replacement + original[end_offset:]
|
||||||
|
)
|
||||||
|
continue
|
||||||
|
prefix = runs[start_index].text[:start_offset]
|
||||||
|
suffix = runs[end_index].text[end_offset:]
|
||||||
|
runs[start_index].text = prefix + replacement
|
||||||
|
for index in range(start_index + 1, end_index):
|
||||||
|
runs[index].text = ""
|
||||||
|
runs[end_index].text = suffix
|
||||||
|
return len(spans)
|
||||||
|
|
||||||
|
|
||||||
|
def _selected_slides(presentation: Any, indexes: Optional[list[Any]]) -> list[Any]:
|
||||||
|
if indexes is None:
|
||||||
|
return list(presentation.slides)
|
||||||
|
if not isinstance(indexes, list) or not indexes:
|
||||||
|
raise ValueError("slides 必须是非空页码数组")
|
||||||
|
selected: list[Any] = []
|
||||||
|
for value in indexes:
|
||||||
|
if not isinstance(value, int) or isinstance(value, bool):
|
||||||
|
raise ValueError("slides 中的页码必须是整数")
|
||||||
|
if value < 1 or value > len(presentation.slides):
|
||||||
|
raise ValueError(f"页码超出范围:{value}")
|
||||||
|
selected.append(presentation.slides[value - 1])
|
||||||
|
return selected
|
||||||
|
|
||||||
|
|
||||||
|
def _replace_text(presentation: Any, operation: dict[str, Any]) -> dict[str, Any]:
|
||||||
|
needle = operation.get("find")
|
||||||
|
replacement = operation.get("replace")
|
||||||
|
if not isinstance(needle, str) or not needle:
|
||||||
|
raise ValueError("replace_text.find 必须是非空字符串")
|
||||||
|
if not isinstance(replacement, str):
|
||||||
|
raise ValueError("replace_text.replace 必须是字符串")
|
||||||
|
match_case = bool(operation.get("match_case", True))
|
||||||
|
whole_word = bool(operation.get("whole_word", False))
|
||||||
|
required = bool(operation.get("required", True))
|
||||||
|
include_notes = bool(operation.get("include_notes", False))
|
||||||
|
limit_value = operation.get("count")
|
||||||
|
if limit_value is not None:
|
||||||
|
if (
|
||||||
|
not isinstance(limit_value, int)
|
||||||
|
or isinstance(limit_value, bool)
|
||||||
|
or limit_value < 1
|
||||||
|
or limit_value > 10000
|
||||||
|
):
|
||||||
|
raise ValueError("replace_text.count 必须是 1 到 10000 的整数")
|
||||||
|
remaining = limit_value
|
||||||
|
changed = 0
|
||||||
|
for slide in _selected_slides(presentation, operation.get("slides")):
|
||||||
|
for text_frame in _iter_text_frames(slide, include_notes=include_notes):
|
||||||
|
for paragraph in text_frame.paragraphs:
|
||||||
|
per_paragraph_limit = remaining
|
||||||
|
replacements = _replace_in_paragraph(
|
||||||
|
paragraph,
|
||||||
|
needle,
|
||||||
|
replacement,
|
||||||
|
match_case=match_case,
|
||||||
|
whole_word=whole_word,
|
||||||
|
limit=per_paragraph_limit,
|
||||||
|
)
|
||||||
|
changed += replacements
|
||||||
|
if remaining is not None:
|
||||||
|
remaining -= replacements
|
||||||
|
if remaining <= 0:
|
||||||
|
break
|
||||||
|
if remaining is not None and remaining <= 0:
|
||||||
|
break
|
||||||
|
if remaining is not None and remaining <= 0:
|
||||||
|
break
|
||||||
|
if required and changed == 0:
|
||||||
|
raise ValueError(f"未找到必须替换的文本:{needle}")
|
||||||
|
return {"type": "replace_text", "replacement_count": changed}
|
||||||
|
|
||||||
|
|
||||||
|
def _set_properties(presentation: Any, operation: dict[str, Any]) -> dict[str, Any]:
|
||||||
|
properties = operation.get("properties")
|
||||||
|
if not isinstance(properties, dict):
|
||||||
|
raise ValueError("set_properties.properties 必须是对象")
|
||||||
|
allowed = {
|
||||||
|
"title",
|
||||||
|
"subject",
|
||||||
|
"author",
|
||||||
|
"keywords",
|
||||||
|
"comments",
|
||||||
|
"category",
|
||||||
|
"content_status",
|
||||||
|
"identifier",
|
||||||
|
"language",
|
||||||
|
"last_modified_by",
|
||||||
|
"revision",
|
||||||
|
"version",
|
||||||
|
}
|
||||||
|
changed: list[str] = []
|
||||||
|
for key, value in properties.items():
|
||||||
|
if key not in allowed:
|
||||||
|
raise ValueError(f"set_properties 不支持字段:{key}")
|
||||||
|
if key == "revision":
|
||||||
|
if (
|
||||||
|
not isinstance(value, int)
|
||||||
|
or isinstance(value, bool)
|
||||||
|
or value < 1
|
||||||
|
):
|
||||||
|
raise ValueError("set_properties.revision 必须是正整数")
|
||||||
|
elif value is not None and not isinstance(value, str):
|
||||||
|
raise ValueError(f"set_properties.{key} 必须是字符串或 null")
|
||||||
|
setattr(presentation.core_properties, key, value)
|
||||||
|
changed.append(key)
|
||||||
|
return {"type": "set_properties", "fields": changed}
|
||||||
|
|
||||||
|
|
||||||
|
def _delete_slides(presentation: Any, operation: dict[str, Any]) -> dict[str, Any]:
|
||||||
|
indexes = operation.get("slides")
|
||||||
|
if not isinstance(indexes, list) or not indexes:
|
||||||
|
raise ValueError("delete_slides.slides 必须是非空页码数组")
|
||||||
|
normalized: set[int] = set()
|
||||||
|
for value in indexes:
|
||||||
|
if not isinstance(value, int) or isinstance(value, bool):
|
||||||
|
raise ValueError("delete_slides.slides 中的页码必须是整数")
|
||||||
|
if value < 1 or value > len(presentation.slides):
|
||||||
|
raise ValueError(f"要删除的页码超出范围:{value}")
|
||||||
|
normalized.add(value)
|
||||||
|
if len(normalized) >= len(presentation.slides):
|
||||||
|
raise ValueError("不能删除演示文稿中的全部页面")
|
||||||
|
slide_ids = presentation.slides._sldIdLst
|
||||||
|
for index in sorted(normalized, reverse=True):
|
||||||
|
slide_id = slide_ids[index - 1]
|
||||||
|
relationship_id = slide_id.rId
|
||||||
|
slide_ids.remove(slide_id)
|
||||||
|
presentation.part.drop_rel(relationship_id)
|
||||||
|
return {"type": "delete_slides", "deleted": sorted(normalized)}
|
||||||
|
|
||||||
|
|
||||||
|
def _reorder_slides(presentation: Any, operation: dict[str, Any]) -> dict[str, Any]:
|
||||||
|
order = operation.get("order")
|
||||||
|
expected = list(range(1, len(presentation.slides) + 1))
|
||||||
|
if (
|
||||||
|
not isinstance(order, list)
|
||||||
|
or any(
|
||||||
|
not isinstance(value, int) or isinstance(value, bool)
|
||||||
|
for value in order
|
||||||
|
)
|
||||||
|
or sorted(order) != expected
|
||||||
|
):
|
||||||
|
raise ValueError(
|
||||||
|
"reorder_slides.order 必须完整且不重复地列出当前全部页码"
|
||||||
|
)
|
||||||
|
slide_ids = presentation.slides._sldIdLst
|
||||||
|
original = list(slide_ids)
|
||||||
|
for slide_id in original:
|
||||||
|
slide_ids.remove(slide_id)
|
||||||
|
for number in order:
|
||||||
|
slide_ids.append(original[number - 1])
|
||||||
|
return {"type": "reorder_slides", "order": order}
|
||||||
|
|
||||||
|
|
||||||
|
def _apply_operation(presentation: Any, operation: dict[str, Any]) -> dict[str, Any]:
|
||||||
|
operation_type = operation.get("type")
|
||||||
|
if operation_type == "replace_text":
|
||||||
|
return _replace_text(presentation, operation)
|
||||||
|
if operation_type == "set_properties":
|
||||||
|
return _set_properties(presentation, operation)
|
||||||
|
if operation_type == "delete_slides":
|
||||||
|
return _delete_slides(presentation, operation)
|
||||||
|
if operation_type == "reorder_slides":
|
||||||
|
return _reorder_slides(presentation, operation)
|
||||||
|
raise ValueError(f"不支持的编辑操作:{operation_type}")
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> dict[str, Any]:
|
||||||
|
from pptx import Presentation
|
||||||
|
|
||||||
|
args = build_parser().parse_args()
|
||||||
|
source = input_file(args.input, OOXML_PRESENTATION_SUFFIXES)
|
||||||
|
destination = output_file(args.output, {".pptx"}, overwrite=args.overwrite)
|
||||||
|
spec = load_json_argument(args.spec, args.spec_file, label="编辑说明")
|
||||||
|
operations = spec.get("operations")
|
||||||
|
if not isinstance(operations, list) or not operations:
|
||||||
|
raise ValueError("编辑说明的 operations 必须是非空数组")
|
||||||
|
if len(operations) > 1000:
|
||||||
|
raise ValueError("编辑操作不能超过 1000 项")
|
||||||
|
|
||||||
|
external_relationships = _external_relationship_count(source)
|
||||||
|
if external_relationships and not args.allow_external_links:
|
||||||
|
raise ValueError(
|
||||||
|
"输入演示文稿包含外部链接;如用户接受外部链接可能变化的风险,"
|
||||||
|
"请显式传 --allow-external-links"
|
||||||
|
)
|
||||||
|
risks = _package_risks(source)
|
||||||
|
if risks:
|
||||||
|
raise ValueError(
|
||||||
|
"输入演示文稿包含 python-pptx 不能可靠保留的内容:"
|
||||||
|
+ "、".join(risks)
|
||||||
|
+ ";请改用 unpack_presentation.py 做最小化 OOXML 编辑"
|
||||||
|
)
|
||||||
|
|
||||||
|
presentation = Presentation(str(source))
|
||||||
|
results: list[dict[str, Any]] = []
|
||||||
|
for index, operation in enumerate(operations):
|
||||||
|
if not isinstance(operation, dict):
|
||||||
|
raise ValueError(f"operations[{index}] 必须是对象")
|
||||||
|
results.append(_apply_operation(presentation, operation))
|
||||||
|
|
||||||
|
with tempfile.TemporaryDirectory(prefix="pptx-edit-") as temp_name:
|
||||||
|
staged = Path(temp_name) / "edited.pptx"
|
||||||
|
presentation.save(str(staged))
|
||||||
|
reopened = Presentation(str(staged))
|
||||||
|
slide_count = len(reopened.slides)
|
||||||
|
publish_file(staged, destination, overwrite=args.overwrite)
|
||||||
|
return {
|
||||||
|
"source": str(source),
|
||||||
|
"path": str(destination),
|
||||||
|
"slide_count": slide_count,
|
||||||
|
"external_relationship_count": external_relationships,
|
||||||
|
"operations": results,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(run_cli(main))
|
||||||
57
skills/pptx/scripts/extract_presentation.py
Normal file
57
skills/pptx/scripts/extract_presentation.py
Normal file
@ -0,0 +1,57 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import tempfile
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from _pptx_common import (
|
||||||
|
OOXML_PRESENTATION_SUFFIXES,
|
||||||
|
SkillArgumentParser,
|
||||||
|
input_file,
|
||||||
|
output_file,
|
||||||
|
publish_file,
|
||||||
|
run_cli,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def build_parser() -> argparse.ArgumentParser:
|
||||||
|
parser = SkillArgumentParser(description="把演示文稿正文提取为 Markdown。")
|
||||||
|
parser.add_argument("--input", required=True)
|
||||||
|
parser.add_argument("--output", required=True)
|
||||||
|
parser.add_argument("--max-chars", type=int, default=2_000_000)
|
||||||
|
parser.add_argument("--overwrite", action="store_true")
|
||||||
|
return parser
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> dict[str, Any]:
|
||||||
|
from markitdown import MarkItDown
|
||||||
|
|
||||||
|
args = build_parser().parse_args()
|
||||||
|
if args.max_chars < 1000 or args.max_chars > 10_000_000:
|
||||||
|
raise ValueError("max-chars 必须在 1000 到 10000000 之间")
|
||||||
|
source = input_file(args.input, OOXML_PRESENTATION_SUFFIXES)
|
||||||
|
destination = output_file(args.output, {".md"}, overwrite=args.overwrite)
|
||||||
|
|
||||||
|
result = MarkItDown(enable_plugins=False).convert(str(source))
|
||||||
|
markdown = result.text_content or ""
|
||||||
|
truncated = len(markdown) > args.max_chars
|
||||||
|
if truncated:
|
||||||
|
markdown = markdown[: args.max_chars]
|
||||||
|
markdown += "\n\n<!-- 内容因 max-chars 限制而截断 -->\n"
|
||||||
|
with tempfile.TemporaryDirectory(prefix="pptx-extract-") as temp_name:
|
||||||
|
staged = Path(temp_name) / "presentation.md"
|
||||||
|
staged.write_text(markdown, encoding="utf-8")
|
||||||
|
publish_file(staged, destination, overwrite=args.overwrite)
|
||||||
|
return {
|
||||||
|
"source": str(source),
|
||||||
|
"path": str(destination),
|
||||||
|
"characters": len(markdown),
|
||||||
|
"truncated": truncated,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(run_cli(main))
|
||||||
404
skills/pptx/scripts/inspect_presentation.py
Normal file
404
skills/pptx/scripts/inspect_presentation.py
Normal file
@ -0,0 +1,404 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import zipfile
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any, Optional
|
||||||
|
|
||||||
|
from _pptx_common import (
|
||||||
|
OOXML_PRESENTATION_SUFFIXES,
|
||||||
|
SkillArgumentParser,
|
||||||
|
input_file,
|
||||||
|
inspect_archive,
|
||||||
|
parse_xml_bytes,
|
||||||
|
run_cli,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
EMU_PER_INCH = 914400
|
||||||
|
|
||||||
|
|
||||||
|
def build_parser() -> argparse.ArgumentParser:
|
||||||
|
parser = SkillArgumentParser(description="分段检查演示文稿的结构、文本和媒体。")
|
||||||
|
parser.add_argument("--input", required=True)
|
||||||
|
parser.add_argument("--start-slide", type=int, default=1)
|
||||||
|
parser.add_argument("--max-slides", type=int, default=30)
|
||||||
|
parser.add_argument("--max-shapes", type=int, default=200)
|
||||||
|
parser.add_argument("--max-table-cells", type=int, default=500)
|
||||||
|
parser.add_argument("--max-chars", type=int, default=100000)
|
||||||
|
parser.add_argument("--include-runs", action="store_true")
|
||||||
|
return parser
|
||||||
|
|
||||||
|
|
||||||
|
def _inches(value: Any) -> Optional[float]:
|
||||||
|
try:
|
||||||
|
return round(int(value) / EMU_PER_INCH, 4)
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def _iter_shapes(shapes: Any):
|
||||||
|
from pptx.enum.shapes import MSO_SHAPE_TYPE
|
||||||
|
|
||||||
|
for shape in shapes:
|
||||||
|
yield shape
|
||||||
|
if shape.shape_type == MSO_SHAPE_TYPE.GROUP:
|
||||||
|
yield from _iter_shapes(shape.shapes)
|
||||||
|
|
||||||
|
|
||||||
|
def _shape_text_char_count(shape: Any) -> int:
|
||||||
|
texts: list[str] = []
|
||||||
|
if getattr(shape, "has_text_frame", False):
|
||||||
|
texts.append(shape.text_frame.text)
|
||||||
|
if getattr(shape, "has_table", False):
|
||||||
|
for row in shape.table.rows:
|
||||||
|
texts.extend(cell.text for cell in row.cells)
|
||||||
|
return sum(
|
||||||
|
1
|
||||||
|
for text in texts
|
||||||
|
for character in text
|
||||||
|
if character.isalnum()
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _contains_picture(shape: Any) -> bool:
|
||||||
|
from pptx.enum.shapes import MSO_SHAPE_TYPE
|
||||||
|
|
||||||
|
if shape.shape_type == MSO_SHAPE_TYPE.PICTURE:
|
||||||
|
return True
|
||||||
|
if shape.shape_type == MSO_SHAPE_TYPE.GROUP:
|
||||||
|
return any(_contains_picture(child) for child in shape.shapes)
|
||||||
|
return False
|
||||||
|
|
||||||
|
|
||||||
|
def _slide_media_profile(
|
||||||
|
slide: Any,
|
||||||
|
slide_width: int,
|
||||||
|
slide_height: int,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
from pptx.enum.shapes import MSO_SHAPE_TYPE
|
||||||
|
|
||||||
|
shapes = list(_iter_shapes(slide.shapes))
|
||||||
|
image_count = sum(
|
||||||
|
1 for shape in shapes
|
||||||
|
if shape.shape_type == MSO_SHAPE_TYPE.PICTURE
|
||||||
|
)
|
||||||
|
chart_count = sum(
|
||||||
|
1 for shape in shapes
|
||||||
|
if getattr(shape, "has_chart", False)
|
||||||
|
)
|
||||||
|
native_text_char_count = sum(
|
||||||
|
_shape_text_char_count(shape) for shape in shapes
|
||||||
|
)
|
||||||
|
image_area = 0.0
|
||||||
|
slide_area = slide_width * slide_height
|
||||||
|
if slide_area > 0:
|
||||||
|
for shape in slide.shapes:
|
||||||
|
if not _contains_picture(shape):
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
raw_left = int(shape.left)
|
||||||
|
raw_top = int(shape.top)
|
||||||
|
left = max(0, raw_left)
|
||||||
|
top = max(0, raw_top)
|
||||||
|
right = min(slide_width, raw_left + int(shape.width))
|
||||||
|
bottom = min(slide_height, raw_top + int(shape.height))
|
||||||
|
except (AttributeError, TypeError, ValueError):
|
||||||
|
continue
|
||||||
|
if right > left and bottom > top:
|
||||||
|
image_area += (right - left) * (bottom - top) / slide_area
|
||||||
|
return {
|
||||||
|
"image_count": image_count,
|
||||||
|
"chart_count": chart_count,
|
||||||
|
"image_area_ratio": round(min(1.0, image_area), 4),
|
||||||
|
"native_text_char_count": native_text_char_count,
|
||||||
|
"has_images": image_count > 0,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _run_payload(run: Any) -> dict[str, Any]:
|
||||||
|
font = run.font
|
||||||
|
hyperlink = None
|
||||||
|
try:
|
||||||
|
hyperlink = run.hyperlink.address
|
||||||
|
except (AttributeError, KeyError, ValueError):
|
||||||
|
pass
|
||||||
|
return {
|
||||||
|
"text": run.text,
|
||||||
|
"bold": font.bold,
|
||||||
|
"italic": font.italic,
|
||||||
|
"underline": font.underline,
|
||||||
|
"font_name": font.name,
|
||||||
|
"font_size_pt": round(font.size.pt, 2) if font.size is not None else None,
|
||||||
|
"hyperlink": hyperlink,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _paragraph_payload(paragraph: Any, *, include_runs: bool) -> dict[str, Any]:
|
||||||
|
payload: dict[str, Any] = {
|
||||||
|
"text": paragraph.text,
|
||||||
|
"level": paragraph.level,
|
||||||
|
"alignment": str(paragraph.alignment) if paragraph.alignment else None,
|
||||||
|
}
|
||||||
|
if include_runs:
|
||||||
|
payload["runs"] = [_run_payload(run) for run in paragraph.runs]
|
||||||
|
return payload
|
||||||
|
|
||||||
|
|
||||||
|
def _text_frame_payload(text_frame: Any, *, include_runs: bool) -> dict[str, Any]:
|
||||||
|
return {
|
||||||
|
"text": text_frame.text,
|
||||||
|
"paragraphs": [
|
||||||
|
_paragraph_payload(paragraph, include_runs=include_runs)
|
||||||
|
for paragraph in text_frame.paragraphs
|
||||||
|
],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _chart_payload(shape: Any, *, max_points: int = 100) -> dict[str, Any]:
|
||||||
|
chart = shape.chart
|
||||||
|
series_payload: list[dict[str, Any]] = []
|
||||||
|
for series in list(chart.series)[:50]:
|
||||||
|
values: list[Any] = []
|
||||||
|
try:
|
||||||
|
values = list(series.values)[:max_points]
|
||||||
|
except (AttributeError, TypeError, ValueError):
|
||||||
|
pass
|
||||||
|
series_payload.append(
|
||||||
|
{
|
||||||
|
"name": getattr(series, "name", None),
|
||||||
|
"point_count": len(values),
|
||||||
|
"values": values,
|
||||||
|
}
|
||||||
|
)
|
||||||
|
categories: list[str] = []
|
||||||
|
try:
|
||||||
|
categories = [str(item.label) for item in chart.plots[0].categories][:max_points]
|
||||||
|
except (AttributeError, IndexError, TypeError, ValueError):
|
||||||
|
pass
|
||||||
|
return {
|
||||||
|
"chart_type": str(chart.chart_type),
|
||||||
|
"has_title": chart.has_title,
|
||||||
|
"series": series_payload,
|
||||||
|
"categories": categories,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _shape_payload(
|
||||||
|
shape: Any,
|
||||||
|
*,
|
||||||
|
include_runs: bool,
|
||||||
|
max_table_cells: int,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
payload: dict[str, Any] = {
|
||||||
|
"name": shape.name,
|
||||||
|
"shape_type": str(shape.shape_type),
|
||||||
|
"x": _inches(shape.left),
|
||||||
|
"y": _inches(shape.top),
|
||||||
|
"w": _inches(shape.width),
|
||||||
|
"h": _inches(shape.height),
|
||||||
|
}
|
||||||
|
if getattr(shape, "has_text_frame", False):
|
||||||
|
payload["text_frame"] = _text_frame_payload(
|
||||||
|
shape.text_frame,
|
||||||
|
include_runs=include_runs,
|
||||||
|
)
|
||||||
|
if getattr(shape, "has_table", False):
|
||||||
|
rows: list[list[str]] = []
|
||||||
|
cell_count = 0
|
||||||
|
truncated = False
|
||||||
|
for row in shape.table.rows:
|
||||||
|
values: list[str] = []
|
||||||
|
for cell in row.cells:
|
||||||
|
if cell_count >= max_table_cells:
|
||||||
|
truncated = True
|
||||||
|
break
|
||||||
|
values.append(cell.text)
|
||||||
|
cell_count += 1
|
||||||
|
if values:
|
||||||
|
rows.append(values)
|
||||||
|
if truncated:
|
||||||
|
break
|
||||||
|
payload["table"] = {
|
||||||
|
"row_count": len(shape.table.rows),
|
||||||
|
"column_count": len(shape.table.columns),
|
||||||
|
"rows": rows,
|
||||||
|
"truncated": truncated,
|
||||||
|
}
|
||||||
|
if getattr(shape, "has_chart", False):
|
||||||
|
payload["chart"] = _chart_payload(shape)
|
||||||
|
if shape.shape_type == 13:
|
||||||
|
try:
|
||||||
|
payload["image"] = {
|
||||||
|
"content_type": shape.image.content_type,
|
||||||
|
"filename": shape.image.filename,
|
||||||
|
"size_bytes": len(shape.image.blob),
|
||||||
|
}
|
||||||
|
except (AttributeError, KeyError, ValueError):
|
||||||
|
payload["image"] = {"readable": False}
|
||||||
|
return payload
|
||||||
|
|
||||||
|
|
||||||
|
def _slide_notes(slide: Any) -> str:
|
||||||
|
try:
|
||||||
|
text_frame = slide.notes_slide.notes_text_frame
|
||||||
|
return text_frame.text if text_frame is not None else ""
|
||||||
|
except (AttributeError, KeyError, ValueError):
|
||||||
|
return ""
|
||||||
|
|
||||||
|
|
||||||
|
def _comment_summary(path: Path) -> dict[str, Any]:
|
||||||
|
comment_parts: list[str] = []
|
||||||
|
authors: set[str] = set()
|
||||||
|
with zipfile.ZipFile(path) as archive:
|
||||||
|
for name in archive.namelist():
|
||||||
|
if name.startswith("ppt/comments/comment") and name.endswith(".xml"):
|
||||||
|
comment_parts.append(name)
|
||||||
|
if name in {
|
||||||
|
"ppt/commentAuthors.xml",
|
||||||
|
"ppt/authors.xml",
|
||||||
|
}:
|
||||||
|
root = parse_xml_bytes(archive.read(name), label=name)
|
||||||
|
for node in root.iter():
|
||||||
|
author = node.attrib.get("name")
|
||||||
|
if author:
|
||||||
|
authors.add(author)
|
||||||
|
return {
|
||||||
|
"part_count": len(comment_parts),
|
||||||
|
"authors": sorted(authors),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _external_relationships(path: Path) -> list[dict[str, str]]:
|
||||||
|
relationships: list[dict[str, str]] = []
|
||||||
|
with zipfile.ZipFile(path) as archive:
|
||||||
|
for name in archive.namelist():
|
||||||
|
if not name.endswith(".rels"):
|
||||||
|
continue
|
||||||
|
root = parse_xml_bytes(archive.read(name), label=name)
|
||||||
|
for node in root:
|
||||||
|
if node.attrib.get("TargetMode") != "External":
|
||||||
|
continue
|
||||||
|
relationships.append(
|
||||||
|
{
|
||||||
|
"part": name,
|
||||||
|
"type": node.attrib.get("Type", ""),
|
||||||
|
"target": node.attrib.get("Target", ""),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
return relationships[:200]
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> dict[str, Any]:
|
||||||
|
from pptx import Presentation
|
||||||
|
|
||||||
|
args = build_parser().parse_args()
|
||||||
|
if args.start_slide < 1:
|
||||||
|
raise ValueError("start-slide 必须大于 0")
|
||||||
|
if args.max_slides < 1 or args.max_slides > 100:
|
||||||
|
raise ValueError("max-slides 必须在 1 到 100 之间")
|
||||||
|
if args.max_shapes < 1 or args.max_shapes > 1000:
|
||||||
|
raise ValueError("max-shapes 必须在 1 到 1000 之间")
|
||||||
|
if args.max_table_cells < 1 or args.max_table_cells > 10000:
|
||||||
|
raise ValueError("max-table-cells 必须在 1 到 10000 之间")
|
||||||
|
if args.max_chars < 1000 or args.max_chars > 1_000_000:
|
||||||
|
raise ValueError("max-chars 必须在 1000 到 1000000 之间")
|
||||||
|
|
||||||
|
source = input_file(args.input, OOXML_PRESENTATION_SUFFIXES)
|
||||||
|
archive = inspect_archive(source)
|
||||||
|
if archive["missing_required_parts"]:
|
||||||
|
raise ValueError(
|
||||||
|
"演示文稿缺少必要部件:"
|
||||||
|
+ "、".join(archive["missing_required_parts"])
|
||||||
|
)
|
||||||
|
presentation = Presentation(str(source))
|
||||||
|
slide_count = len(presentation.slides)
|
||||||
|
if args.start_slide > slide_count and slide_count > 0:
|
||||||
|
raise ValueError(f"start-slide 超出页面总数 {slide_count}")
|
||||||
|
|
||||||
|
end_slide = min(slide_count, args.start_slide + args.max_slides - 1)
|
||||||
|
payload_slides: list[dict[str, Any]] = []
|
||||||
|
char_count = 0
|
||||||
|
truncated_by_chars = False
|
||||||
|
for slide_number in range(args.start_slide, end_slide + 1):
|
||||||
|
slide = presentation.slides[slide_number - 1]
|
||||||
|
title = slide.shapes.title.text if slide.shapes.title is not None else ""
|
||||||
|
media_profile = _slide_media_profile(
|
||||||
|
slide,
|
||||||
|
int(presentation.slide_width),
|
||||||
|
int(presentation.slide_height),
|
||||||
|
)
|
||||||
|
shapes: list[dict[str, Any]] = []
|
||||||
|
for shape in list(slide.shapes)[: args.max_shapes]:
|
||||||
|
item = _shape_payload(
|
||||||
|
shape,
|
||||||
|
include_runs=args.include_runs,
|
||||||
|
max_table_cells=args.max_table_cells,
|
||||||
|
)
|
||||||
|
item_chars = len(str(item))
|
||||||
|
if char_count + item_chars > args.max_chars:
|
||||||
|
truncated_by_chars = True
|
||||||
|
break
|
||||||
|
shapes.append(item)
|
||||||
|
char_count += item_chars
|
||||||
|
notes = _slide_notes(slide)
|
||||||
|
if char_count + len(notes) > args.max_chars:
|
||||||
|
notes = notes[: max(0, args.max_chars - char_count)]
|
||||||
|
truncated_by_chars = True
|
||||||
|
char_count += len(notes)
|
||||||
|
payload_slides.append(
|
||||||
|
{
|
||||||
|
"number": slide_number,
|
||||||
|
"title": title,
|
||||||
|
"layout": getattr(slide.slide_layout, "name", None),
|
||||||
|
"shape_count": len(slide.shapes),
|
||||||
|
"shapes_truncated": len(slide.shapes) > args.max_shapes,
|
||||||
|
"shapes": shapes,
|
||||||
|
"speaker_notes": notes,
|
||||||
|
"media": media_profile,
|
||||||
|
}
|
||||||
|
)
|
||||||
|
if truncated_by_chars:
|
||||||
|
break
|
||||||
|
|
||||||
|
last_slide = payload_slides[-1]["number"] if payload_slides else args.start_slide - 1
|
||||||
|
next_slide = last_slide + 1 if last_slide < slide_count else None
|
||||||
|
properties = presentation.core_properties
|
||||||
|
return {
|
||||||
|
"source": str(source),
|
||||||
|
"slide_count": slide_count,
|
||||||
|
"slide_size": {
|
||||||
|
"width_inches": _inches(presentation.slide_width),
|
||||||
|
"height_inches": _inches(presentation.slide_height),
|
||||||
|
},
|
||||||
|
"properties": {
|
||||||
|
"title": properties.title,
|
||||||
|
"subject": properties.subject,
|
||||||
|
"author": properties.author,
|
||||||
|
"keywords": properties.keywords,
|
||||||
|
"comments": properties.comments,
|
||||||
|
"last_modified_by": properties.last_modified_by,
|
||||||
|
},
|
||||||
|
"selection": {
|
||||||
|
"start_slide": args.start_slide,
|
||||||
|
"end_slide": last_slide,
|
||||||
|
"has_more": next_slide is not None,
|
||||||
|
"next_slide": next_slide,
|
||||||
|
"truncated_by_chars": truncated_by_chars,
|
||||||
|
},
|
||||||
|
"slides": payload_slides,
|
||||||
|
"image_slides": [
|
||||||
|
slide["number"]
|
||||||
|
for slide in payload_slides
|
||||||
|
if slide["media"]["has_images"]
|
||||||
|
],
|
||||||
|
"comments": _comment_summary(source),
|
||||||
|
"external_relationships": _external_relationships(source),
|
||||||
|
"archive": archive,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(run_cli(main))
|
||||||
727
skills/pptx/scripts/ocr_presentation.py
Normal file
727
skills/pptx/scripts/ocr_presentation.py
Normal file
@ -0,0 +1,727 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import contextlib
|
||||||
|
import difflib
|
||||||
|
import importlib.metadata
|
||||||
|
import io
|
||||||
|
import logging
|
||||||
|
import os
|
||||||
|
import re
|
||||||
|
import shutil
|
||||||
|
import tempfile
|
||||||
|
import time
|
||||||
|
import unicodedata
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any, Iterable
|
||||||
|
|
||||||
|
from _pptx_common import (
|
||||||
|
OOXML_PRESENTATION_SUFFIXES,
|
||||||
|
SkillArgumentParser,
|
||||||
|
find_program,
|
||||||
|
input_file,
|
||||||
|
run_cli,
|
||||||
|
run_program,
|
||||||
|
run_soffice_convert,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
DEFAULT_DPI = 260
|
||||||
|
DEFAULT_MAX_CHARS = 24000
|
||||||
|
DEFAULT_TIMEOUT_SECONDS = 180
|
||||||
|
MAX_SLIDES_PER_CALL = 4
|
||||||
|
MAX_PIXELS_PER_SLIDE = 20_000_000
|
||||||
|
MIN_MEAN_CONFIDENCE = 0.60
|
||||||
|
MIN_MEANINGFUL_CHARS = 5
|
||||||
|
WHITESPACE_PATTERN = re.compile(r"[ \t]+")
|
||||||
|
|
||||||
|
|
||||||
|
for variable, value in (
|
||||||
|
("OMP_NUM_THREADS", "2"),
|
||||||
|
("OPENBLAS_NUM_THREADS", "1"),
|
||||||
|
("MKL_NUM_THREADS", "1"),
|
||||||
|
("NUMEXPR_NUM_THREADS", "1"),
|
||||||
|
):
|
||||||
|
os.environ.setdefault(variable, value)
|
||||||
|
|
||||||
|
for logger_name in ("rapidocr", "RapidOCR", "onnxruntime"):
|
||||||
|
logging.getLogger(logger_name).setLevel(logging.ERROR)
|
||||||
|
|
||||||
|
|
||||||
|
def build_parser():
|
||||||
|
parser = SkillArgumentParser(
|
||||||
|
description="渲染指定演示文稿页面并用本地 OCR 提取图片文字。"
|
||||||
|
)
|
||||||
|
parser.add_argument("--input", required=True)
|
||||||
|
parser.add_argument(
|
||||||
|
"--slides",
|
||||||
|
required=True,
|
||||||
|
help="要识别的页码,例如 2 或 2,5-6;单次最多 4 页",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--start-offset",
|
||||||
|
type=int,
|
||||||
|
default=0,
|
||||||
|
help="续读单页图片文字时的字符偏移量",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--max-chars",
|
||||||
|
type=int,
|
||||||
|
default=DEFAULT_MAX_CHARS,
|
||||||
|
help=f"单次最多返回字符数,默认 {DEFAULT_MAX_CHARS}",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--dpi",
|
||||||
|
type=int,
|
||||||
|
default=DEFAULT_DPI,
|
||||||
|
help=f"OCR 渲染分辨率,默认 {DEFAULT_DPI} DPI",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--timeout",
|
||||||
|
type=int,
|
||||||
|
default=DEFAULT_TIMEOUT_SECONDS,
|
||||||
|
help=f"转换和单页渲染超时秒数,默认 {DEFAULT_TIMEOUT_SECONDS}",
|
||||||
|
)
|
||||||
|
return parser
|
||||||
|
|
||||||
|
|
||||||
|
def _parse_slide_spec(value: str, slide_count: int) -> list[int]:
|
||||||
|
if not value.strip():
|
||||||
|
raise ValueError("slides 不能为空")
|
||||||
|
slides: set[int] = set()
|
||||||
|
for raw_part in value.split(","):
|
||||||
|
part = raw_part.strip()
|
||||||
|
if not part:
|
||||||
|
continue
|
||||||
|
if "-" in part:
|
||||||
|
pieces = part.split("-", 1)
|
||||||
|
try:
|
||||||
|
start = int(pieces[0])
|
||||||
|
end = int(pieces[1])
|
||||||
|
except ValueError as exc:
|
||||||
|
raise ValueError(f"页码范围格式错误:{part}") from exc
|
||||||
|
if start > end:
|
||||||
|
raise ValueError(f"页码范围起始值不能大于结束值:{part}")
|
||||||
|
else:
|
||||||
|
try:
|
||||||
|
start = end = int(part)
|
||||||
|
except ValueError as exc:
|
||||||
|
raise ValueError(f"页码格式错误:{part}") from exc
|
||||||
|
if start < 1 or end > slide_count:
|
||||||
|
raise ValueError(f"页码必须在 1 到 {slide_count} 之间:{part}")
|
||||||
|
slides.update(range(start, end + 1))
|
||||||
|
if not slides:
|
||||||
|
raise ValueError("slides 不能为空")
|
||||||
|
return sorted(slides)
|
||||||
|
|
||||||
|
|
||||||
|
def _clean_text(value: Any) -> str:
|
||||||
|
text = str(value or "").replace("\x00", "").strip()
|
||||||
|
return "\n".join(
|
||||||
|
WHITESPACE_PATTERN.sub(" ", line).strip()
|
||||||
|
for line in text.replace("\r\n", "\n").replace("\r", "\n").split("\n")
|
||||||
|
if line.strip()
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _comparison_key(value: str) -> str:
|
||||||
|
normalized = unicodedata.normalize("NFKC", value).casefold()
|
||||||
|
return "".join(character for character in normalized if character.isalnum())
|
||||||
|
|
||||||
|
|
||||||
|
def _iter_shapes(shapes: Any) -> Iterable[Any]:
|
||||||
|
from pptx.enum.shapes import MSO_SHAPE_TYPE
|
||||||
|
|
||||||
|
for shape in shapes:
|
||||||
|
yield shape
|
||||||
|
if shape.shape_type == MSO_SHAPE_TYPE.GROUP:
|
||||||
|
yield from _iter_shapes(shape.shapes)
|
||||||
|
|
||||||
|
|
||||||
|
def _shape_text_fragments(shape: Any) -> list[str]:
|
||||||
|
fragments: list[str] = []
|
||||||
|
if getattr(shape, "has_text_frame", False):
|
||||||
|
fragments.extend(
|
||||||
|
line
|
||||||
|
for line in _clean_text(shape.text_frame.text).splitlines()
|
||||||
|
if line
|
||||||
|
)
|
||||||
|
if getattr(shape, "has_table", False):
|
||||||
|
for row in shape.table.rows:
|
||||||
|
for cell in row.cells:
|
||||||
|
fragments.extend(
|
||||||
|
line
|
||||||
|
for line in _clean_text(cell.text).splitlines()
|
||||||
|
if line
|
||||||
|
)
|
||||||
|
return fragments
|
||||||
|
|
||||||
|
|
||||||
|
def _normalized_box(
|
||||||
|
shape: Any,
|
||||||
|
slide_width: int,
|
||||||
|
slide_height: int,
|
||||||
|
) -> tuple[float, float, float, float] | None:
|
||||||
|
try:
|
||||||
|
left = float(shape.left)
|
||||||
|
top = float(shape.top)
|
||||||
|
right = left + float(shape.width)
|
||||||
|
bottom = top + float(shape.height)
|
||||||
|
except (AttributeError, TypeError, ValueError):
|
||||||
|
return None
|
||||||
|
if slide_width <= 0 or slide_height <= 0:
|
||||||
|
return None
|
||||||
|
x0 = max(0.0, min(1.0, left / slide_width))
|
||||||
|
y0 = max(0.0, min(1.0, top / slide_height))
|
||||||
|
x1 = max(0.0, min(1.0, right / slide_width))
|
||||||
|
y1 = max(0.0, min(1.0, bottom / slide_height))
|
||||||
|
if x1 <= x0 or y1 <= y0:
|
||||||
|
return None
|
||||||
|
return (x0, y0, x1, y1)
|
||||||
|
|
||||||
|
|
||||||
|
def _contains_picture(shape: Any) -> bool:
|
||||||
|
from pptx.enum.shapes import MSO_SHAPE_TYPE
|
||||||
|
|
||||||
|
if shape.shape_type == MSO_SHAPE_TYPE.PICTURE:
|
||||||
|
return True
|
||||||
|
if shape.shape_type == MSO_SHAPE_TYPE.GROUP:
|
||||||
|
return any(_contains_picture(child) for child in shape.shapes)
|
||||||
|
return False
|
||||||
|
|
||||||
|
|
||||||
|
def _slide_profile(
|
||||||
|
slide: Any,
|
||||||
|
slide_width: int,
|
||||||
|
slide_height: int,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
from pptx.enum.shapes import MSO_SHAPE_TYPE
|
||||||
|
|
||||||
|
all_shapes = list(_iter_shapes(slide.shapes))
|
||||||
|
native_fragments: list[str] = []
|
||||||
|
picture_count = 0
|
||||||
|
chart_count = 0
|
||||||
|
for shape in all_shapes:
|
||||||
|
native_fragments.extend(_shape_text_fragments(shape))
|
||||||
|
if shape.shape_type == MSO_SHAPE_TYPE.PICTURE:
|
||||||
|
picture_count += 1
|
||||||
|
if getattr(shape, "has_chart", False):
|
||||||
|
chart_count += 1
|
||||||
|
|
||||||
|
unique_fragments = list(dict.fromkeys(native_fragments))
|
||||||
|
native_keys = [
|
||||||
|
key
|
||||||
|
for fragment in unique_fragments
|
||||||
|
if (key := _comparison_key(fragment))
|
||||||
|
]
|
||||||
|
native_boxes: list[dict[str, Any]] = []
|
||||||
|
picture_area = 0.0
|
||||||
|
for shape in slide.shapes:
|
||||||
|
box = _normalized_box(shape, slide_width, slide_height)
|
||||||
|
shape_fragments = _shape_text_fragments(shape)
|
||||||
|
if box and shape_fragments:
|
||||||
|
native_boxes.append(
|
||||||
|
{
|
||||||
|
"box": box,
|
||||||
|
"keys": [
|
||||||
|
key
|
||||||
|
for fragment in shape_fragments
|
||||||
|
if (key := _comparison_key(fragment))
|
||||||
|
],
|
||||||
|
}
|
||||||
|
)
|
||||||
|
if box and _contains_picture(shape):
|
||||||
|
picture_area += (box[2] - box[0]) * (box[3] - box[1])
|
||||||
|
|
||||||
|
return {
|
||||||
|
"picture_count": picture_count,
|
||||||
|
"chart_count": chart_count,
|
||||||
|
"image_area_ratio": round(min(1.0, picture_area), 4),
|
||||||
|
"native_text_char_count": sum(
|
||||||
|
1
|
||||||
|
for fragment in unique_fragments
|
||||||
|
for character in fragment
|
||||||
|
if character.isalnum()
|
||||||
|
),
|
||||||
|
"_native_keys": native_keys,
|
||||||
|
"_native_boxes": native_boxes,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _box_points(value: Any) -> list[list[float]] | None:
|
||||||
|
if value is None:
|
||||||
|
return None
|
||||||
|
try:
|
||||||
|
points = [
|
||||||
|
[round(float(point[0]), 2), round(float(point[1]), 2)]
|
||||||
|
for point in value
|
||||||
|
]
|
||||||
|
except (IndexError, TypeError, ValueError):
|
||||||
|
return None
|
||||||
|
return points if len(points) == 4 else None
|
||||||
|
|
||||||
|
|
||||||
|
def _ordered_lines(result: Any) -> list[dict[str, Any]]:
|
||||||
|
texts = list(getattr(result, "txts", None) or ())
|
||||||
|
scores = list(getattr(result, "scores", None) or ())
|
||||||
|
raw_boxes = getattr(result, "boxes", None)
|
||||||
|
boxes = list(raw_boxes) if raw_boxes is not None else []
|
||||||
|
|
||||||
|
lines: list[dict[str, Any]] = []
|
||||||
|
for index, raw_text in enumerate(texts):
|
||||||
|
text = _clean_text(raw_text)
|
||||||
|
if not text:
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
confidence = float(scores[index])
|
||||||
|
except (IndexError, TypeError, ValueError):
|
||||||
|
confidence = 0.0
|
||||||
|
confidence = max(0.0, min(1.0, confidence))
|
||||||
|
box = _box_points(boxes[index] if index < len(boxes) else None)
|
||||||
|
if box:
|
||||||
|
left = min(point[0] for point in box)
|
||||||
|
top = min(point[1] for point in box)
|
||||||
|
else:
|
||||||
|
left = float(index)
|
||||||
|
top = float(index)
|
||||||
|
lines.append(
|
||||||
|
{
|
||||||
|
"text": text,
|
||||||
|
"confidence": confidence,
|
||||||
|
"box": box,
|
||||||
|
"_left": left,
|
||||||
|
"_top": top,
|
||||||
|
"_index": index,
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
lines.sort(
|
||||||
|
key=lambda line: (
|
||||||
|
round(line["_top"] / 10.0),
|
||||||
|
line["_left"],
|
||||||
|
line["_index"],
|
||||||
|
)
|
||||||
|
)
|
||||||
|
return lines
|
||||||
|
|
||||||
|
|
||||||
|
def _similar_to_any(
|
||||||
|
candidate: str,
|
||||||
|
references: list[str],
|
||||||
|
*,
|
||||||
|
threshold: float,
|
||||||
|
) -> bool:
|
||||||
|
if not candidate:
|
||||||
|
return False
|
||||||
|
for reference in references:
|
||||||
|
if not reference:
|
||||||
|
continue
|
||||||
|
if candidate == reference:
|
||||||
|
return True
|
||||||
|
shorter = min(len(candidate), len(reference))
|
||||||
|
longer = max(len(candidate), len(reference))
|
||||||
|
if shorter >= 3 and candidate in reference:
|
||||||
|
return True
|
||||||
|
if (
|
||||||
|
shorter >= 3
|
||||||
|
and reference in candidate
|
||||||
|
and longer <= round(shorter * 1.25)
|
||||||
|
):
|
||||||
|
return True
|
||||||
|
if shorter >= 3 and difflib.SequenceMatcher(
|
||||||
|
None,
|
||||||
|
candidate,
|
||||||
|
reference,
|
||||||
|
).ratio() >= threshold:
|
||||||
|
return True
|
||||||
|
return False
|
||||||
|
|
||||||
|
|
||||||
|
def _line_center(
|
||||||
|
box: list[list[float]] | None,
|
||||||
|
image_width: int,
|
||||||
|
image_height: int,
|
||||||
|
) -> tuple[float, float] | None:
|
||||||
|
if not box or image_width <= 0 or image_height <= 0:
|
||||||
|
return None
|
||||||
|
return (
|
||||||
|
sum(point[0] for point in box) / len(box) / image_width,
|
||||||
|
sum(point[1] for point in box) / len(box) / image_height,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _line_is_native(
|
||||||
|
line: dict[str, Any],
|
||||||
|
profile: dict[str, Any],
|
||||||
|
image_width: int,
|
||||||
|
image_height: int,
|
||||||
|
) -> bool:
|
||||||
|
candidate = _comparison_key(line["text"])
|
||||||
|
if _similar_to_any(
|
||||||
|
candidate,
|
||||||
|
profile["_native_keys"],
|
||||||
|
threshold=0.82,
|
||||||
|
):
|
||||||
|
return True
|
||||||
|
|
||||||
|
center = _line_center(line["box"], image_width, image_height)
|
||||||
|
if center is None:
|
||||||
|
return False
|
||||||
|
x, y = center
|
||||||
|
padding = 0.01
|
||||||
|
for native_box in profile["_native_boxes"]:
|
||||||
|
x0, y0, x1, y1 = native_box["box"]
|
||||||
|
if (
|
||||||
|
x0 - padding <= x <= x1 + padding
|
||||||
|
and y0 - padding <= y <= y1 + padding
|
||||||
|
and _similar_to_any(
|
||||||
|
candidate,
|
||||||
|
native_box["keys"],
|
||||||
|
threshold=0.68,
|
||||||
|
)
|
||||||
|
):
|
||||||
|
return True
|
||||||
|
return False
|
||||||
|
|
||||||
|
|
||||||
|
def _create_ocr_engine():
|
||||||
|
try:
|
||||||
|
from rapidocr import RapidOCR
|
||||||
|
except ImportError as exc:
|
||||||
|
raise RuntimeError("环境预置的 rapidocr 模块不可用") from exc
|
||||||
|
|
||||||
|
captured_stdout = io.StringIO()
|
||||||
|
captured_stderr = io.StringIO()
|
||||||
|
with (
|
||||||
|
contextlib.redirect_stdout(captured_stdout),
|
||||||
|
contextlib.redirect_stderr(captured_stderr),
|
||||||
|
):
|
||||||
|
return RapidOCR()
|
||||||
|
|
||||||
|
|
||||||
|
def _ocr_slide(
|
||||||
|
engine: Any,
|
||||||
|
image_path: Path,
|
||||||
|
profile: dict[str, Any],
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
from PIL import Image
|
||||||
|
|
||||||
|
with Image.open(image_path) as image:
|
||||||
|
image_width, image_height = image.size
|
||||||
|
|
||||||
|
captured_stdout = io.StringIO()
|
||||||
|
captured_stderr = io.StringIO()
|
||||||
|
started = time.monotonic()
|
||||||
|
with (
|
||||||
|
contextlib.redirect_stdout(captured_stdout),
|
||||||
|
contextlib.redirect_stderr(captured_stderr),
|
||||||
|
):
|
||||||
|
result = engine(str(image_path))
|
||||||
|
elapsed = time.monotonic() - started
|
||||||
|
|
||||||
|
raw_lines = _ordered_lines(result)
|
||||||
|
image_lines: list[dict[str, Any]] = []
|
||||||
|
seen: set[str] = set()
|
||||||
|
filtered_native = 0
|
||||||
|
filtered_duplicates = 0
|
||||||
|
for line in raw_lines:
|
||||||
|
if _line_is_native(line, profile, image_width, image_height):
|
||||||
|
filtered_native += 1
|
||||||
|
continue
|
||||||
|
key = _comparison_key(line["text"])
|
||||||
|
if key and key in seen:
|
||||||
|
filtered_duplicates += 1
|
||||||
|
continue
|
||||||
|
if key:
|
||||||
|
seen.add(key)
|
||||||
|
image_lines.append(line)
|
||||||
|
|
||||||
|
text = "\n".join(line["text"] for line in image_lines)
|
||||||
|
weighted_chars = [
|
||||||
|
max(1, sum(1 for character in line["text"] if not character.isspace()))
|
||||||
|
for line in image_lines
|
||||||
|
]
|
||||||
|
total_weight = sum(weighted_chars)
|
||||||
|
mean_confidence = (
|
||||||
|
sum(
|
||||||
|
line["confidence"] * weight
|
||||||
|
for line, weight in zip(image_lines, weighted_chars)
|
||||||
|
)
|
||||||
|
/ total_weight
|
||||||
|
if total_weight
|
||||||
|
else 0.0
|
||||||
|
)
|
||||||
|
meaningful_chars = sum(1 for character in text if character.isalnum())
|
||||||
|
low_confidence_lines = sum(
|
||||||
|
1
|
||||||
|
for line in image_lines
|
||||||
|
if line["confidence"] < MIN_MEAN_CONFIDENCE
|
||||||
|
)
|
||||||
|
|
||||||
|
reasons: list[str] = []
|
||||||
|
if not text:
|
||||||
|
status = "no_image_text"
|
||||||
|
reasons.append("未识别到原生文本之外的图片文字")
|
||||||
|
elif meaningful_chars < MIN_MEANINGFUL_CHARS:
|
||||||
|
status = "sparse"
|
||||||
|
reasons.append(
|
||||||
|
f"图片中的有效文字少于 {MIN_MEANINGFUL_CHARS} 个字符"
|
||||||
|
)
|
||||||
|
elif mean_confidence < MIN_MEAN_CONFIDENCE:
|
||||||
|
status = "low_confidence"
|
||||||
|
reasons.append(
|
||||||
|
"图片文字 OCR 平均置信度低于 "
|
||||||
|
f"{round(MIN_MEAN_CONFIDENCE * 100)}%"
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
status = "good"
|
||||||
|
|
||||||
|
return {
|
||||||
|
"text": text,
|
||||||
|
"status": status,
|
||||||
|
"usable_for_summary": status == "good",
|
||||||
|
"needs_review": status in {"sparse", "low_confidence"},
|
||||||
|
"raw_ocr_line_count": len(raw_lines),
|
||||||
|
"image_line_count": len(image_lines),
|
||||||
|
"filtered_native_line_count": filtered_native,
|
||||||
|
"filtered_duplicate_line_count": filtered_duplicates,
|
||||||
|
"low_confidence_line_count": low_confidence_lines,
|
||||||
|
"mean_confidence": round(mean_confidence, 4),
|
||||||
|
"meaningful_chars": meaningful_chars,
|
||||||
|
"reasons": reasons,
|
||||||
|
"ocr_seconds": round(elapsed, 3),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _pdf_pages(path: Path) -> tuple[int, dict[int, tuple[float, float]]]:
|
||||||
|
from pypdf import PdfReader
|
||||||
|
|
||||||
|
page_sizes: dict[int, tuple[float, float]] = {}
|
||||||
|
with path.open("rb") as stream:
|
||||||
|
reader = PdfReader(stream, strict=False)
|
||||||
|
if reader.is_encrypted:
|
||||||
|
raise ValueError("LibreOffice 生成了加密 PDF,无法执行 OCR")
|
||||||
|
page_count = len(reader.pages)
|
||||||
|
for page_number, page in enumerate(reader.pages, start=1):
|
||||||
|
page_sizes[page_number] = (
|
||||||
|
abs(float(page.cropbox.width)),
|
||||||
|
abs(float(page.cropbox.height)),
|
||||||
|
)
|
||||||
|
return page_count, page_sizes
|
||||||
|
|
||||||
|
|
||||||
|
def _render_slide(
|
||||||
|
pdf_path: Path,
|
||||||
|
slide_number: int,
|
||||||
|
page_size: tuple[float, float],
|
||||||
|
dpi: int,
|
||||||
|
timeout: int,
|
||||||
|
temp_dir: Path,
|
||||||
|
) -> tuple[Path, float]:
|
||||||
|
width_points, height_points = page_size
|
||||||
|
estimated_pixels = (
|
||||||
|
width_points * dpi / 72.0
|
||||||
|
* height_points * dpi / 72.0
|
||||||
|
)
|
||||||
|
if estimated_pixels > MAX_PIXELS_PER_SLIDE:
|
||||||
|
raise ValueError(
|
||||||
|
f"第 {slide_number} 页按 {dpi} DPI 渲染预计超过 "
|
||||||
|
f"{MAX_PIXELS_PER_SLIDE} 像素,请降低 dpi"
|
||||||
|
)
|
||||||
|
|
||||||
|
prefix = temp_dir / f"slide-{slide_number:04d}"
|
||||||
|
output = prefix.with_suffix(".png")
|
||||||
|
started = time.monotonic()
|
||||||
|
run_program(
|
||||||
|
[
|
||||||
|
find_program("pdftoppm"),
|
||||||
|
"-f",
|
||||||
|
str(slide_number),
|
||||||
|
"-l",
|
||||||
|
str(slide_number),
|
||||||
|
"-singlefile",
|
||||||
|
"-png",
|
||||||
|
"-r",
|
||||||
|
str(dpi),
|
||||||
|
str(pdf_path),
|
||||||
|
str(prefix),
|
||||||
|
],
|
||||||
|
timeout=timeout,
|
||||||
|
)
|
||||||
|
elapsed = time.monotonic() - started
|
||||||
|
if not output.is_file() or output.stat().st_size <= 0:
|
||||||
|
raise RuntimeError(f"第 {slide_number} 页没有生成有效 PNG")
|
||||||
|
return output, elapsed
|
||||||
|
|
||||||
|
|
||||||
|
def _package_version(name: str) -> str | None:
|
||||||
|
try:
|
||||||
|
return importlib.metadata.version(name)
|
||||||
|
except importlib.metadata.PackageNotFoundError:
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> dict[str, Any]:
|
||||||
|
from pptx import Presentation
|
||||||
|
|
||||||
|
args = build_parser().parse_args()
|
||||||
|
if args.start_offset < 0:
|
||||||
|
raise ValueError("start-offset 不能小于 0")
|
||||||
|
if args.max_chars < 1 or args.max_chars > 60000:
|
||||||
|
raise ValueError("max-chars 必须在 1 到 60000 之间")
|
||||||
|
if args.dpi < 150 or args.dpi > 400:
|
||||||
|
raise ValueError("dpi 必须在 150 到 400 之间")
|
||||||
|
if args.timeout < 1 or args.timeout > 600:
|
||||||
|
raise ValueError("timeout 必须在 1 到 600 秒之间")
|
||||||
|
|
||||||
|
source = input_file(args.input, OOXML_PRESENTATION_SUFFIXES)
|
||||||
|
presentation = Presentation(str(source))
|
||||||
|
slide_count = len(presentation.slides)
|
||||||
|
if slide_count < 1:
|
||||||
|
raise ValueError("演示文稿没有可执行 OCR 的页面")
|
||||||
|
requested_slides = _parse_slide_spec(args.slides, slide_count)
|
||||||
|
if len(requested_slides) > MAX_SLIDES_PER_CALL:
|
||||||
|
raise ValueError(
|
||||||
|
f"单次最多 OCR {MAX_SLIDES_PER_CALL} 页,请拆分 slides 后重试"
|
||||||
|
)
|
||||||
|
if args.start_offset > 0 and len(requested_slides) != 1:
|
||||||
|
raise ValueError("使用 start-offset 时 slides 必须只包含一页")
|
||||||
|
|
||||||
|
slide_width = int(presentation.slide_width)
|
||||||
|
slide_height = int(presentation.slide_height)
|
||||||
|
profiles = {
|
||||||
|
slide_number: _slide_profile(
|
||||||
|
presentation.slides[slide_number - 1],
|
||||||
|
slide_width,
|
||||||
|
slide_height,
|
||||||
|
)
|
||||||
|
for slide_number in requested_slides
|
||||||
|
}
|
||||||
|
|
||||||
|
engine = _create_ocr_engine()
|
||||||
|
page_outputs: list[dict[str, Any]] = []
|
||||||
|
returned_chars = 0
|
||||||
|
next_slide: int | None = None
|
||||||
|
next_offset = 0
|
||||||
|
remaining_slides: list[int] = []
|
||||||
|
office_output = {"stdout": "", "stderr": ""}
|
||||||
|
|
||||||
|
with tempfile.TemporaryDirectory(prefix="pptx-ocr-") as temp_name:
|
||||||
|
temp_dir = Path(temp_name)
|
||||||
|
staged_input = temp_dir / f"presentation{source.suffix.lower()}"
|
||||||
|
shutil.copy2(source, staged_input)
|
||||||
|
pdf_path, office_output = run_soffice_convert(
|
||||||
|
staged_input,
|
||||||
|
target_format="pdf",
|
||||||
|
output_dir=temp_dir / "pdf",
|
||||||
|
timeout=args.timeout,
|
||||||
|
)
|
||||||
|
pdf_page_count, page_sizes = _pdf_pages(pdf_path)
|
||||||
|
if pdf_page_count != slide_count:
|
||||||
|
raise RuntimeError(
|
||||||
|
f"演示文稿有 {slide_count} 页,但渲染结果有 "
|
||||||
|
f"{pdf_page_count} 页"
|
||||||
|
)
|
||||||
|
|
||||||
|
for index, slide_number in enumerate(requested_slides):
|
||||||
|
budget = args.max_chars - returned_chars
|
||||||
|
if budget <= 0:
|
||||||
|
next_slide = slide_number
|
||||||
|
remaining_slides = requested_slides[index:]
|
||||||
|
break
|
||||||
|
|
||||||
|
image_path, render_seconds = _render_slide(
|
||||||
|
pdf_path,
|
||||||
|
slide_number,
|
||||||
|
page_sizes[slide_number],
|
||||||
|
args.dpi,
|
||||||
|
args.timeout,
|
||||||
|
temp_dir,
|
||||||
|
)
|
||||||
|
result = _ocr_slide(
|
||||||
|
engine,
|
||||||
|
image_path,
|
||||||
|
profiles[slide_number],
|
||||||
|
)
|
||||||
|
full_text = result.pop("text")
|
||||||
|
offset = args.start_offset if index == 0 else 0
|
||||||
|
if offset > len(full_text):
|
||||||
|
raise ValueError(
|
||||||
|
f"start-offset 超过第 {slide_number} 页图片文字长度 "
|
||||||
|
f"{len(full_text)}"
|
||||||
|
)
|
||||||
|
|
||||||
|
usable = bool(result["usable_for_summary"])
|
||||||
|
if not usable:
|
||||||
|
slide_text = ""
|
||||||
|
complete = True
|
||||||
|
else:
|
||||||
|
remaining_text = full_text[offset:]
|
||||||
|
slide_text = remaining_text[:budget]
|
||||||
|
complete = len(slide_text) == len(remaining_text)
|
||||||
|
|
||||||
|
profile = profiles[slide_number]
|
||||||
|
page_outputs.append(
|
||||||
|
{
|
||||||
|
"slide": slide_number,
|
||||||
|
"text": slide_text,
|
||||||
|
"char_count": len(full_text),
|
||||||
|
"offset_start": offset if usable else 0,
|
||||||
|
"offset_end": offset + len(slide_text) if usable else 0,
|
||||||
|
"complete": complete,
|
||||||
|
"render_seconds": round(render_seconds, 3),
|
||||||
|
"picture_count": profile["picture_count"],
|
||||||
|
"chart_count": profile["chart_count"],
|
||||||
|
"image_area_ratio": profile["image_area_ratio"],
|
||||||
|
"native_text_char_count": profile[
|
||||||
|
"native_text_char_count"
|
||||||
|
],
|
||||||
|
**result,
|
||||||
|
}
|
||||||
|
)
|
||||||
|
returned_chars += len(slide_text)
|
||||||
|
|
||||||
|
if not complete:
|
||||||
|
next_slide = slide_number
|
||||||
|
next_offset = offset + len(slide_text)
|
||||||
|
remaining_slides = requested_slides[index + 1 :]
|
||||||
|
break
|
||||||
|
|
||||||
|
all_processed = len(page_outputs) == len(requested_slides)
|
||||||
|
all_complete = all(page["complete"] for page in page_outputs)
|
||||||
|
all_safe = all(
|
||||||
|
page["status"] in {"good", "no_image_text"}
|
||||||
|
for page in page_outputs
|
||||||
|
)
|
||||||
|
return {
|
||||||
|
"source": str(source),
|
||||||
|
"slide_count": slide_count,
|
||||||
|
"engine": "rapidocr",
|
||||||
|
"engine_version": _package_version("rapidocr"),
|
||||||
|
"runtime": "onnxruntime",
|
||||||
|
"runtime_version": _package_version("onnxruntime"),
|
||||||
|
"offline": True,
|
||||||
|
"dpi": args.dpi,
|
||||||
|
"requested_slides": requested_slides,
|
||||||
|
"processed_slides": [page["slide"] for page in page_outputs],
|
||||||
|
"returned_chars": returned_chars,
|
||||||
|
"slides": page_outputs,
|
||||||
|
"usable_for_summary": any(
|
||||||
|
page["usable_for_summary"] for page in page_outputs
|
||||||
|
),
|
||||||
|
"complete_ocr_coverage": (
|
||||||
|
all_processed and all_complete and all_safe
|
||||||
|
),
|
||||||
|
"needs_review": any(page["needs_review"] for page in page_outputs),
|
||||||
|
"has_more": next_slide is not None,
|
||||||
|
"next_slide": next_slide,
|
||||||
|
"next_offset": next_offset,
|
||||||
|
"remaining_slides": remaining_slides,
|
||||||
|
"office_stdout": office_output["stdout"],
|
||||||
|
"office_stderr": office_output["stderr"],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(run_cli(main))
|
||||||
55
skills/pptx/scripts/pack_presentation.py
Normal file
55
skills/pptx/scripts/pack_presentation.py
Normal file
@ -0,0 +1,55 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import tempfile
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from _pptx_common import (
|
||||||
|
SkillArgumentParser,
|
||||||
|
output_file,
|
||||||
|
pack_presentation_directory,
|
||||||
|
publish_file,
|
||||||
|
run_cli,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def build_parser() -> argparse.ArgumentParser:
|
||||||
|
parser = SkillArgumentParser(description="把 OOXML 目录安全打包为 PowerPoint 文件。")
|
||||||
|
parser.add_argument("--input-dir", required=True)
|
||||||
|
parser.add_argument("--output", required=True)
|
||||||
|
parser.add_argument("--overwrite", action="store_true")
|
||||||
|
return parser
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> dict[str, Any]:
|
||||||
|
from pptx import Presentation
|
||||||
|
|
||||||
|
args = build_parser().parse_args()
|
||||||
|
source_dir = Path(args.input_dir).expanduser().resolve()
|
||||||
|
destination = output_file(
|
||||||
|
args.output,
|
||||||
|
{".pptx", ".potx", ".ppsx"},
|
||||||
|
overwrite=args.overwrite,
|
||||||
|
)
|
||||||
|
with tempfile.TemporaryDirectory(prefix="pptx-pack-") as temp_name:
|
||||||
|
staged = Path(temp_name) / destination.name
|
||||||
|
archive = pack_presentation_directory(source_dir, staged)
|
||||||
|
try:
|
||||||
|
presentation = Presentation(str(staged))
|
||||||
|
slide_count = len(presentation.slides)
|
||||||
|
except Exception as exc:
|
||||||
|
raise ValueError("打包结果不能被 python-pptx 打开") from exc
|
||||||
|
publish_file(staged, destination, overwrite=args.overwrite)
|
||||||
|
return {
|
||||||
|
"source_dir": str(source_dir),
|
||||||
|
"path": str(destination),
|
||||||
|
"slide_count": slide_count,
|
||||||
|
"archive": archive,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(run_cli(main))
|
||||||
81
skills/pptx/scripts/render_icon.py
Normal file
81
skills/pptx/scripts/render_icon.py
Normal file
@ -0,0 +1,81 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import tempfile
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from _pptx_common import (
|
||||||
|
SkillArgumentParser,
|
||||||
|
find_program,
|
||||||
|
output_file,
|
||||||
|
publish_file,
|
||||||
|
run_cli,
|
||||||
|
run_program,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def build_parser() -> argparse.ArgumentParser:
|
||||||
|
parser = SkillArgumentParser(description="把 React Icons 图标渲染为透明 PNG。")
|
||||||
|
parser.add_argument("--library", required=True)
|
||||||
|
parser.add_argument("--name", required=True)
|
||||||
|
parser.add_argument("--output", required=True)
|
||||||
|
parser.add_argument("--color", default="111827")
|
||||||
|
parser.add_argument("--background")
|
||||||
|
parser.add_argument("--size", type=int, default=256)
|
||||||
|
parser.add_argument("--title")
|
||||||
|
parser.add_argument("--timeout", type=int, default=60)
|
||||||
|
parser.add_argument("--overwrite", action="store_true")
|
||||||
|
return parser
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> dict[str, Any]:
|
||||||
|
args = build_parser().parse_args()
|
||||||
|
destination = output_file(args.output, {".png"}, overwrite=args.overwrite)
|
||||||
|
renderer = Path(__file__).with_name("_icon_renderer.js").resolve()
|
||||||
|
if not renderer.is_file():
|
||||||
|
raise FileNotFoundError(f"内部图标渲染器不存在:{renderer}")
|
||||||
|
spec = {
|
||||||
|
"library": args.library,
|
||||||
|
"name": args.name,
|
||||||
|
"color": args.color,
|
||||||
|
"background": args.background,
|
||||||
|
"size": args.size,
|
||||||
|
"title": args.title,
|
||||||
|
}
|
||||||
|
with tempfile.TemporaryDirectory(prefix="pptx-icon-") as temp_name:
|
||||||
|
temp_dir = Path(temp_name)
|
||||||
|
spec_path = temp_dir / "spec.json"
|
||||||
|
staged = temp_dir / "icon.png"
|
||||||
|
spec_path.write_text(json.dumps(spec, ensure_ascii=False), encoding="utf-8")
|
||||||
|
completed = run_program(
|
||||||
|
[
|
||||||
|
find_program("node"),
|
||||||
|
str(renderer),
|
||||||
|
"--spec",
|
||||||
|
str(spec_path),
|
||||||
|
"--output",
|
||||||
|
str(staged),
|
||||||
|
],
|
||||||
|
timeout=args.timeout,
|
||||||
|
cwd=Path.cwd(),
|
||||||
|
env=os.environ.copy(),
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
result = json.loads(completed.stdout.strip())
|
||||||
|
except json.JSONDecodeError as exc:
|
||||||
|
raise RuntimeError("内部图标渲染器返回了无效结果") from exc
|
||||||
|
publish_file(staged, destination, overwrite=args.overwrite)
|
||||||
|
return {
|
||||||
|
"path": str(destination),
|
||||||
|
"size_bytes": destination.stat().st_size,
|
||||||
|
**result,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(run_cli(main))
|
||||||
224
skills/pptx/scripts/render_presentation.py
Normal file
224
skills/pptx/scripts/render_presentation.py
Normal file
@ -0,0 +1,224 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import math
|
||||||
|
import shutil
|
||||||
|
import tempfile
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any, Optional
|
||||||
|
|
||||||
|
from _pptx_common import (
|
||||||
|
OOXML_PRESENTATION_SUFFIXES,
|
||||||
|
SkillArgumentParser,
|
||||||
|
find_program,
|
||||||
|
input_file,
|
||||||
|
output_directory,
|
||||||
|
publish_file,
|
||||||
|
run_cli,
|
||||||
|
run_program,
|
||||||
|
run_soffice_convert,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def build_parser() -> argparse.ArgumentParser:
|
||||||
|
parser = SkillArgumentParser(
|
||||||
|
description="通过 LibreOffice 和 Poppler 把演示文稿渲染为逐页 PNG。"
|
||||||
|
)
|
||||||
|
parser.add_argument("--input", required=True)
|
||||||
|
parser.add_argument("--output-dir", required=True)
|
||||||
|
parser.add_argument("--start-slide", type=int, default=1)
|
||||||
|
parser.add_argument("--end-slide", type=int)
|
||||||
|
parser.add_argument("--max-slides", type=int, default=30)
|
||||||
|
parser.add_argument("--dpi", type=int, default=150)
|
||||||
|
parser.add_argument("--timeout", type=int, default=180)
|
||||||
|
parser.add_argument("--include-pdf", action="store_true")
|
||||||
|
parser.add_argument("--contact-sheet", action="store_true")
|
||||||
|
parser.add_argument("--overwrite", action="store_true")
|
||||||
|
return parser
|
||||||
|
|
||||||
|
|
||||||
|
def _contact_sheet(
|
||||||
|
images: list[Path],
|
||||||
|
slide_numbers: list[int],
|
||||||
|
destination: Path,
|
||||||
|
) -> None:
|
||||||
|
from PIL import Image, ImageDraw, ImageFont
|
||||||
|
|
||||||
|
columns = min(4, max(1, len(images)))
|
||||||
|
rows = math.ceil(len(images) / columns)
|
||||||
|
thumb_width = 420
|
||||||
|
label_height = 34
|
||||||
|
gap = 18
|
||||||
|
opened: list[Image.Image] = []
|
||||||
|
try:
|
||||||
|
for path in images:
|
||||||
|
opened.append(Image.open(path).convert("RGB"))
|
||||||
|
aspect = opened[0].height / opened[0].width
|
||||||
|
thumb_height = max(1, round(thumb_width * aspect))
|
||||||
|
canvas_width = columns * thumb_width + (columns + 1) * gap
|
||||||
|
canvas_height = rows * (thumb_height + label_height) + (rows + 1) * gap
|
||||||
|
canvas = Image.new("RGB", (canvas_width, canvas_height), "white")
|
||||||
|
draw = ImageDraw.Draw(canvas)
|
||||||
|
try:
|
||||||
|
font = ImageFont.truetype(
|
||||||
|
"/usr/share/fonts/truetype/dejavu/DejaVuSans.ttf",
|
||||||
|
20,
|
||||||
|
)
|
||||||
|
except OSError:
|
||||||
|
font = ImageFont.load_default()
|
||||||
|
for index, (image, slide_number) in enumerate(zip(opened, slide_numbers)):
|
||||||
|
row, column = divmod(index, columns)
|
||||||
|
x = gap + column * (thumb_width + gap)
|
||||||
|
y = gap + row * (thumb_height + label_height)
|
||||||
|
thumbnail = image.copy()
|
||||||
|
thumbnail.thumbnail((thumb_width, thumb_height))
|
||||||
|
paste_x = x + (thumb_width - thumbnail.width) // 2
|
||||||
|
paste_y = y + (thumb_height - thumbnail.height) // 2
|
||||||
|
canvas.paste(thumbnail, (paste_x, paste_y))
|
||||||
|
label = f"Slide {slide_number}"
|
||||||
|
label_box = draw.textbbox((0, 0), label, font=font)
|
||||||
|
label_width = label_box[2] - label_box[0]
|
||||||
|
draw.text(
|
||||||
|
(x + (thumb_width - label_width) // 2, y + thumb_height + 6),
|
||||||
|
label,
|
||||||
|
fill="black",
|
||||||
|
font=font,
|
||||||
|
)
|
||||||
|
canvas.save(destination, format="PNG", optimize=True)
|
||||||
|
finally:
|
||||||
|
for image in opened:
|
||||||
|
image.close()
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> dict[str, Any]:
|
||||||
|
from pypdf import PdfReader
|
||||||
|
|
||||||
|
args = build_parser().parse_args()
|
||||||
|
if args.start_slide < 1:
|
||||||
|
raise ValueError("start-slide 必须大于 0")
|
||||||
|
if args.end_slide is not None and args.end_slide < args.start_slide:
|
||||||
|
raise ValueError("end-slide 不能小于 start-slide")
|
||||||
|
if args.max_slides < 1 or args.max_slides > 100:
|
||||||
|
raise ValueError("max-slides 必须在 1 到 100 之间")
|
||||||
|
if args.dpi < 72 or args.dpi > 300:
|
||||||
|
raise ValueError("dpi 必须在 72 到 300 之间")
|
||||||
|
|
||||||
|
source = input_file(args.input, OOXML_PRESENTATION_SUFFIXES)
|
||||||
|
destination_dir = output_directory(args.output_dir)
|
||||||
|
with tempfile.TemporaryDirectory(prefix="pptx-render-") as temp_name:
|
||||||
|
temp_dir = Path(temp_name)
|
||||||
|
staged_input = temp_dir / f"presentation{source.suffix.lower()}"
|
||||||
|
shutil.copy2(source, staged_input)
|
||||||
|
pdf_path, office_output = run_soffice_convert(
|
||||||
|
staged_input,
|
||||||
|
target_format="pdf",
|
||||||
|
output_dir=temp_dir / "pdf",
|
||||||
|
timeout=args.timeout,
|
||||||
|
)
|
||||||
|
slide_count = len(PdfReader(str(pdf_path)).pages)
|
||||||
|
if slide_count < 1:
|
||||||
|
raise ValueError("LibreOffice 生成的 PDF 没有页面")
|
||||||
|
if args.start_slide > slide_count:
|
||||||
|
raise ValueError(f"start-slide 超出页面总数 {slide_count}")
|
||||||
|
requested_end = slide_count if args.end_slide is None else args.end_slide
|
||||||
|
if requested_end > slide_count:
|
||||||
|
raise ValueError(f"end-slide 超出页面总数 {slide_count}")
|
||||||
|
actual_end = min(
|
||||||
|
requested_end,
|
||||||
|
args.start_slide + args.max_slides - 1,
|
||||||
|
)
|
||||||
|
|
||||||
|
raw_prefix = temp_dir / "raw-slide"
|
||||||
|
run_program(
|
||||||
|
[
|
||||||
|
find_program("pdftoppm"),
|
||||||
|
"-png",
|
||||||
|
"-r",
|
||||||
|
str(args.dpi),
|
||||||
|
"-f",
|
||||||
|
str(args.start_slide),
|
||||||
|
"-l",
|
||||||
|
str(actual_end),
|
||||||
|
str(pdf_path),
|
||||||
|
str(raw_prefix),
|
||||||
|
],
|
||||||
|
timeout=args.timeout,
|
||||||
|
)
|
||||||
|
raw_pages = sorted(
|
||||||
|
temp_dir.glob("raw-slide-*.png"),
|
||||||
|
key=lambda path: int(path.stem.rsplit("-", 1)[1]),
|
||||||
|
)
|
||||||
|
expected_count = actual_end - args.start_slide + 1
|
||||||
|
if len(raw_pages) != expected_count:
|
||||||
|
raise RuntimeError(
|
||||||
|
f"Poppler 应生成 {expected_count} 页,实际生成 {len(raw_pages)} 页"
|
||||||
|
)
|
||||||
|
slide_numbers = list(range(args.start_slide, actual_end + 1))
|
||||||
|
destinations = [
|
||||||
|
destination_dir / f"slide-{number:04d}.png"
|
||||||
|
for number in slide_numbers
|
||||||
|
]
|
||||||
|
contact_destination = destination_dir / (
|
||||||
|
f"contact-sheet-{args.start_slide:04d}-{actual_end:04d}.png"
|
||||||
|
)
|
||||||
|
if args.contact_sheet:
|
||||||
|
destinations.append(contact_destination)
|
||||||
|
if args.include_pdf:
|
||||||
|
destinations.append(destination_dir / "presentation.pdf")
|
||||||
|
if not args.overwrite:
|
||||||
|
existing = [str(path) for path in destinations if path.exists()]
|
||||||
|
if existing:
|
||||||
|
raise FileExistsError("以下渲染目标已存在:" + "、".join(existing))
|
||||||
|
|
||||||
|
output_paths: list[str] = []
|
||||||
|
published_pages: list[Path] = []
|
||||||
|
for slide_number, raw_page in zip(slide_numbers, raw_pages):
|
||||||
|
destination = destination_dir / f"slide-{slide_number:04d}.png"
|
||||||
|
publish_file(raw_page, destination, overwrite=args.overwrite)
|
||||||
|
output_paths.append(str(destination))
|
||||||
|
published_pages.append(destination)
|
||||||
|
|
||||||
|
contact_sheet_path: Optional[str] = None
|
||||||
|
if args.contact_sheet:
|
||||||
|
staged_contact = temp_dir / "contact-sheet.png"
|
||||||
|
_contact_sheet(
|
||||||
|
published_pages,
|
||||||
|
slide_numbers,
|
||||||
|
staged_contact,
|
||||||
|
)
|
||||||
|
publish_file(
|
||||||
|
staged_contact,
|
||||||
|
contact_destination,
|
||||||
|
overwrite=args.overwrite,
|
||||||
|
)
|
||||||
|
contact_sheet_path = str(contact_destination)
|
||||||
|
|
||||||
|
pdf_output: Optional[str] = None
|
||||||
|
if args.include_pdf:
|
||||||
|
destination_pdf = destination_dir / "presentation.pdf"
|
||||||
|
staged_pdf = temp_dir / "publish.pdf"
|
||||||
|
shutil.copy2(pdf_path, staged_pdf)
|
||||||
|
publish_file(staged_pdf, destination_pdf, overwrite=args.overwrite)
|
||||||
|
pdf_output = str(destination_pdf)
|
||||||
|
|
||||||
|
next_slide = actual_end + 1 if actual_end < requested_end else None
|
||||||
|
return {
|
||||||
|
"source": str(source),
|
||||||
|
"slide_count": slide_count,
|
||||||
|
"start_slide": args.start_slide,
|
||||||
|
"end_slide": actual_end,
|
||||||
|
"rendered_slides": output_paths,
|
||||||
|
"contact_sheet": contact_sheet_path,
|
||||||
|
"pdf": pdf_output,
|
||||||
|
"has_more": next_slide is not None,
|
||||||
|
"next_slide": next_slide,
|
||||||
|
"dpi": args.dpi,
|
||||||
|
"office_stdout": office_output["stdout"],
|
||||||
|
"office_stderr": office_output["stderr"],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(run_cli(main))
|
||||||
40
skills/pptx/scripts/unpack_presentation.py
Normal file
40
skills/pptx/scripts/unpack_presentation.py
Normal file
@ -0,0 +1,40 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from _pptx_common import (
|
||||||
|
OOXML_PRESENTATION_SUFFIXES,
|
||||||
|
SkillArgumentParser,
|
||||||
|
input_file,
|
||||||
|
output_directory,
|
||||||
|
run_cli,
|
||||||
|
safe_extract_presentation,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def build_parser() -> argparse.ArgumentParser:
|
||||||
|
parser = SkillArgumentParser(description="安全解包 PowerPoint OOXML 文件。")
|
||||||
|
parser.add_argument("--input", required=True)
|
||||||
|
parser.add_argument("--output-dir", required=True)
|
||||||
|
return parser
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> dict[str, Any]:
|
||||||
|
args = build_parser().parse_args()
|
||||||
|
source = input_file(args.input, OOXML_PRESENTATION_SUFFIXES)
|
||||||
|
destination = output_directory(args.output_dir)
|
||||||
|
if any(destination.iterdir()):
|
||||||
|
raise ValueError("output-dir 必须为空目录")
|
||||||
|
archive = safe_extract_presentation(source, destination)
|
||||||
|
return {
|
||||||
|
"source": str(source),
|
||||||
|
"path": str(destination),
|
||||||
|
"archive": archive,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(run_cli(main))
|
||||||
478
skills/pptx/scripts/validate_presentation.py
Normal file
478
skills/pptx/scripts/validate_presentation.py
Normal file
@ -0,0 +1,478 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import posixpath
|
||||||
|
import re
|
||||||
|
import shutil
|
||||||
|
import tempfile
|
||||||
|
import zipfile
|
||||||
|
from pathlib import Path, PurePosixPath
|
||||||
|
from typing import Any, Optional
|
||||||
|
|
||||||
|
from _pptx_common import (
|
||||||
|
OOXML_PRESENTATION_SUFFIXES,
|
||||||
|
P_NS,
|
||||||
|
R_NS,
|
||||||
|
SkillArgumentParser,
|
||||||
|
input_file,
|
||||||
|
inspect_archive,
|
||||||
|
parse_xml_bytes,
|
||||||
|
run_cli,
|
||||||
|
run_soffice_convert,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
REL_NS = "http://schemas.openxmlformats.org/package/2006/relationships"
|
||||||
|
C_NS = "http://schemas.openxmlformats.org/drawingml/2006/chart"
|
||||||
|
PLACEHOLDER_RE = re.compile(
|
||||||
|
r"\b(?:lorem|ipsum|todo|x{3,})\b|\[insert|this\s+(?:page|slide).+layout",
|
||||||
|
re.IGNORECASE,
|
||||||
|
)
|
||||||
|
EMU_PER_INCH = 914400
|
||||||
|
|
||||||
|
|
||||||
|
def build_parser() -> argparse.ArgumentParser:
|
||||||
|
parser = SkillArgumentParser(description="校验 PowerPoint 的结构、关系和可渲染性。")
|
||||||
|
parser.add_argument("--input", required=True)
|
||||||
|
parser.add_argument("--original")
|
||||||
|
parser.add_argument("--check-render", action="store_true")
|
||||||
|
parser.add_argument("--timeout", type=int, default=180)
|
||||||
|
return parser
|
||||||
|
|
||||||
|
|
||||||
|
def _issue(
|
||||||
|
code: str,
|
||||||
|
message: str,
|
||||||
|
*,
|
||||||
|
part: Optional[str] = None,
|
||||||
|
slide: Optional[int] = None,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
payload: dict[str, Any] = {"code": code, "message": message}
|
||||||
|
if part is not None:
|
||||||
|
payload["part"] = part
|
||||||
|
if slide is not None:
|
||||||
|
payload["slide"] = slide
|
||||||
|
return payload
|
||||||
|
|
||||||
|
|
||||||
|
def _owner_part_for_rels(name: str) -> str:
|
||||||
|
if name == "_rels/.rels":
|
||||||
|
return ""
|
||||||
|
path = PurePosixPath(name)
|
||||||
|
if path.parent.name != "_rels" or not path.name.endswith(".rels"):
|
||||||
|
raise ValueError(f"关系部件路径无效:{name}")
|
||||||
|
owner_name = path.name[: -len(".rels")]
|
||||||
|
owner_parent = path.parent.parent
|
||||||
|
return (owner_parent / owner_name).as_posix()
|
||||||
|
|
||||||
|
|
||||||
|
def _resolve_relationship_target(owner_part: str, target: str) -> Optional[str]:
|
||||||
|
if not target or target.startswith("#"):
|
||||||
|
return None
|
||||||
|
if target.startswith("/"):
|
||||||
|
normalized = posixpath.normpath(target).lstrip("/")
|
||||||
|
if normalized == ".." or normalized.startswith("../"):
|
||||||
|
return None
|
||||||
|
return normalized
|
||||||
|
base = posixpath.dirname(owner_part)
|
||||||
|
normalized = posixpath.normpath(posixpath.join(base, target))
|
||||||
|
if normalized == ".." or normalized.startswith("../"):
|
||||||
|
return None
|
||||||
|
return normalized.lstrip("/")
|
||||||
|
|
||||||
|
|
||||||
|
def _validate_relationships(
|
||||||
|
archive: zipfile.ZipFile,
|
||||||
|
names: set[str],
|
||||||
|
) -> tuple[list[dict[str, Any]], list[dict[str, Any]], dict[str, dict[str, str]]]:
|
||||||
|
errors: list[dict[str, Any]] = []
|
||||||
|
externals: list[dict[str, Any]] = []
|
||||||
|
maps: dict[str, dict[str, str]] = {}
|
||||||
|
for name in sorted(item for item in names if item.endswith(".rels")):
|
||||||
|
try:
|
||||||
|
root = parse_xml_bytes(archive.read(name), label=name)
|
||||||
|
except ValueError as exc:
|
||||||
|
errors.append(_issue("invalid_relationship_xml", str(exc), part=name))
|
||||||
|
continue
|
||||||
|
owner_part = _owner_part_for_rels(name)
|
||||||
|
relationships: dict[str, str] = {}
|
||||||
|
for node in root.findall(f"{{{REL_NS}}}Relationship"):
|
||||||
|
relationship_id = node.attrib.get("Id", "")
|
||||||
|
target = node.attrib.get("Target", "")
|
||||||
|
if not relationship_id or relationship_id in relationships:
|
||||||
|
errors.append(
|
||||||
|
_issue(
|
||||||
|
"duplicate_or_missing_relationship_id",
|
||||||
|
f"关系 Id 缺失或重复:{relationship_id!r}",
|
||||||
|
part=name,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
continue
|
||||||
|
if node.attrib.get("TargetMode") == "External":
|
||||||
|
externals.append(
|
||||||
|
{
|
||||||
|
"part": name,
|
||||||
|
"id": relationship_id,
|
||||||
|
"type": node.attrib.get("Type", ""),
|
||||||
|
"target": target,
|
||||||
|
}
|
||||||
|
)
|
||||||
|
relationships[relationship_id] = target
|
||||||
|
continue
|
||||||
|
resolved = _resolve_relationship_target(owner_part, target)
|
||||||
|
if resolved is None:
|
||||||
|
errors.append(
|
||||||
|
_issue(
|
||||||
|
"unsafe_relationship_target",
|
||||||
|
f"关系目标不安全:{target}",
|
||||||
|
part=name,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
continue
|
||||||
|
relationships[relationship_id] = resolved
|
||||||
|
if resolved not in names:
|
||||||
|
errors.append(
|
||||||
|
_issue(
|
||||||
|
"missing_relationship_target",
|
||||||
|
f"关系目标不存在:{resolved}",
|
||||||
|
part=name,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
maps[owner_part] = relationships
|
||||||
|
return errors, externals, maps
|
||||||
|
|
||||||
|
|
||||||
|
def _validate_slide_order(
|
||||||
|
archive: zipfile.ZipFile,
|
||||||
|
relationship_maps: dict[str, dict[str, str]],
|
||||||
|
) -> tuple[list[dict[str, Any]], list[str]]:
|
||||||
|
errors: list[dict[str, Any]] = []
|
||||||
|
ordered_parts: list[str] = []
|
||||||
|
root = parse_xml_bytes(
|
||||||
|
archive.read("ppt/presentation.xml"),
|
||||||
|
label="ppt/presentation.xml",
|
||||||
|
)
|
||||||
|
slide_ids: set[str] = set()
|
||||||
|
relationship_ids: set[str] = set()
|
||||||
|
presentation_relationships = relationship_maps.get(
|
||||||
|
"ppt/presentation.xml",
|
||||||
|
{},
|
||||||
|
)
|
||||||
|
for node in root.findall(f".//{{{P_NS}}}sldId"):
|
||||||
|
slide_id = node.attrib.get("id", "")
|
||||||
|
relationship_id = node.attrib.get(f"{{{R_NS}}}id", "")
|
||||||
|
if not slide_id or slide_id in slide_ids:
|
||||||
|
errors.append(
|
||||||
|
_issue(
|
||||||
|
"duplicate_or_missing_slide_id",
|
||||||
|
f"页面 id 缺失或重复:{slide_id!r}",
|
||||||
|
part="ppt/presentation.xml",
|
||||||
|
)
|
||||||
|
)
|
||||||
|
slide_ids.add(slide_id)
|
||||||
|
if not relationship_id or relationship_id in relationship_ids:
|
||||||
|
errors.append(
|
||||||
|
_issue(
|
||||||
|
"duplicate_or_missing_slide_relationship",
|
||||||
|
f"页面关系 id 缺失或重复:{relationship_id!r}",
|
||||||
|
part="ppt/presentation.xml",
|
||||||
|
)
|
||||||
|
)
|
||||||
|
relationship_ids.add(relationship_id)
|
||||||
|
target = presentation_relationships.get(relationship_id)
|
||||||
|
if not target or not target.startswith("ppt/slides/slide"):
|
||||||
|
errors.append(
|
||||||
|
_issue(
|
||||||
|
"invalid_slide_relationship",
|
||||||
|
f"页面关系 {relationship_id!r} 未指向有效 slide 部件",
|
||||||
|
part="ppt/presentation.xml",
|
||||||
|
)
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
ordered_parts.append(target)
|
||||||
|
if not ordered_parts:
|
||||||
|
errors.append(
|
||||||
|
_issue(
|
||||||
|
"no_slides",
|
||||||
|
"演示文稿不包含页面",
|
||||||
|
part="ppt/presentation.xml",
|
||||||
|
)
|
||||||
|
)
|
||||||
|
return errors, ordered_parts
|
||||||
|
|
||||||
|
|
||||||
|
def _validate_charts(
|
||||||
|
archive: zipfile.ZipFile,
|
||||||
|
names: set[str],
|
||||||
|
) -> list[dict[str, Any]]:
|
||||||
|
errors: list[dict[str, Any]] = []
|
||||||
|
for name in sorted(
|
||||||
|
item
|
||||||
|
for item in names
|
||||||
|
if item.startswith("ppt/charts/chart") and item.endswith(".xml")
|
||||||
|
):
|
||||||
|
root = parse_xml_bytes(archive.read(name), label=name)
|
||||||
|
declared_axis_ids: set[str] = set()
|
||||||
|
for axis_name in ("catAx", "valAx", "dateAx", "serAx"):
|
||||||
|
for axis in root.findall(f".//{{{C_NS}}}{axis_name}"):
|
||||||
|
node = axis.find(f"{{{C_NS}}}axId")
|
||||||
|
if node is not None and node.attrib.get("val"):
|
||||||
|
declared_axis_ids.add(node.attrib["val"])
|
||||||
|
referenced_axis_ids = {
|
||||||
|
node.attrib["val"]
|
||||||
|
for node in root.findall(f".//{{{C_NS}}}axId")
|
||||||
|
if node.attrib.get("val")
|
||||||
|
}
|
||||||
|
# PptxGenJS writes the standard primary series-axis id on ordinary
|
||||||
|
# two-dimensional charts even though no serAx element is required.
|
||||||
|
# Secondary value/category ids must still have real declarations.
|
||||||
|
tolerated_series_axis_ids = {"2094734556"}
|
||||||
|
missing_axis_ids = sorted(
|
||||||
|
referenced_axis_ids
|
||||||
|
- declared_axis_ids
|
||||||
|
- tolerated_series_axis_ids
|
||||||
|
)
|
||||||
|
if declared_axis_ids and missing_axis_ids:
|
||||||
|
errors.append(
|
||||||
|
_issue(
|
||||||
|
"undeclared_chart_axis",
|
||||||
|
"图表引用了未声明的坐标轴:"
|
||||||
|
+ "、".join(missing_axis_ids),
|
||||||
|
part=name,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
for bar_chart in root.findall(f".//{{{C_NS}}}barChart"):
|
||||||
|
grouping = bar_chart.find(f"{{{C_NS}}}grouping")
|
||||||
|
grouping_value = grouping.attrib.get("val") if grouping is not None else ""
|
||||||
|
if grouping_value not in {"stacked", "percentStacked"}:
|
||||||
|
continue
|
||||||
|
invalid_label = bar_chart.find(
|
||||||
|
f".//{{{C_NS}}}dLblPos[@val='outEnd']"
|
||||||
|
)
|
||||||
|
if invalid_label is not None:
|
||||||
|
errors.append(
|
||||||
|
_issue(
|
||||||
|
"invalid_stacked_chart_label_position",
|
||||||
|
"堆积柱形或条形图不能使用 outEnd 数据标签位置",
|
||||||
|
part=name,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
return errors
|
||||||
|
|
||||||
|
|
||||||
|
def _visual_structure(path: Path) -> tuple[list[dict[str, Any]], list[dict[str, Any]], int]:
|
||||||
|
from pptx import Presentation
|
||||||
|
|
||||||
|
errors: list[dict[str, Any]] = []
|
||||||
|
warnings: list[dict[str, Any]] = []
|
||||||
|
presentation = Presentation(str(path))
|
||||||
|
slide_width = int(presentation.slide_width)
|
||||||
|
slide_height = int(presentation.slide_height)
|
||||||
|
tolerance = 2000
|
||||||
|
for slide_number, slide in enumerate(presentation.slides, start=1):
|
||||||
|
if len(slide.shapes) == 0:
|
||||||
|
warnings.append(
|
||||||
|
_issue("empty_slide", "页面没有可见形状", slide=slide_number)
|
||||||
|
)
|
||||||
|
for shape in slide.shapes:
|
||||||
|
left = int(shape.left)
|
||||||
|
top = int(shape.top)
|
||||||
|
right = left + int(shape.width)
|
||||||
|
bottom = top + int(shape.height)
|
||||||
|
if (
|
||||||
|
left < -tolerance
|
||||||
|
or top < -tolerance
|
||||||
|
or right > slide_width + tolerance
|
||||||
|
or bottom > slide_height + tolerance
|
||||||
|
):
|
||||||
|
errors.append(
|
||||||
|
_issue(
|
||||||
|
"shape_out_of_bounds",
|
||||||
|
f"形状 {shape.name!r} 超出页面边界",
|
||||||
|
slide=slide_number,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
text = getattr(shape, "text", "")
|
||||||
|
if text and PLACEHOLDER_RE.search(text):
|
||||||
|
warnings.append(
|
||||||
|
_issue(
|
||||||
|
"placeholder_text",
|
||||||
|
f"形状 {shape.name!r} 可能残留占位文本",
|
||||||
|
slide=slide_number,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
return errors, warnings, len(presentation.slides)
|
||||||
|
|
||||||
|
|
||||||
|
def _validate_once(path: Path, *, check_render: bool, timeout: int) -> dict[str, Any]:
|
||||||
|
archive_info = inspect_archive(path)
|
||||||
|
errors: list[dict[str, Any]] = []
|
||||||
|
warnings: list[dict[str, Any]] = []
|
||||||
|
if archive_info["missing_required_parts"]:
|
||||||
|
errors.append(
|
||||||
|
_issue(
|
||||||
|
"missing_required_parts",
|
||||||
|
"缺少必要部件:"
|
||||||
|
+ "、".join(archive_info["missing_required_parts"]),
|
||||||
|
)
|
||||||
|
)
|
||||||
|
return {
|
||||||
|
"errors": errors,
|
||||||
|
"warnings": warnings,
|
||||||
|
"archive": archive_info,
|
||||||
|
"slide_count": 0,
|
||||||
|
"external_relationships": [],
|
||||||
|
}
|
||||||
|
if archive_info["duplicate_members"]:
|
||||||
|
errors.append(
|
||||||
|
_issue(
|
||||||
|
"duplicate_archive_members",
|
||||||
|
"压缩包存在重复成员:"
|
||||||
|
+ "、".join(archive_info["duplicate_members"][:20]),
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
with zipfile.ZipFile(path) as archive:
|
||||||
|
names = set(archive.namelist())
|
||||||
|
for name in sorted(
|
||||||
|
item for item in names if item.endswith((".xml", ".rels"))
|
||||||
|
):
|
||||||
|
try:
|
||||||
|
parse_xml_bytes(archive.read(name), label=name)
|
||||||
|
except ValueError as exc:
|
||||||
|
errors.append(_issue("invalid_xml", str(exc), part=name))
|
||||||
|
relationship_errors, externals, maps = _validate_relationships(
|
||||||
|
archive,
|
||||||
|
names,
|
||||||
|
)
|
||||||
|
errors.extend(relationship_errors)
|
||||||
|
slide_errors, ordered_parts = _validate_slide_order(archive, maps)
|
||||||
|
errors.extend(slide_errors)
|
||||||
|
orphaned_slides = sorted(
|
||||||
|
name
|
||||||
|
for name in names
|
||||||
|
if name.startswith("ppt/slides/slide")
|
||||||
|
and name.endswith(".xml")
|
||||||
|
and name not in set(ordered_parts)
|
||||||
|
)
|
||||||
|
if orphaned_slides:
|
||||||
|
warnings.append(
|
||||||
|
_issue(
|
||||||
|
"orphaned_slide_parts",
|
||||||
|
"存在未被 presentation.xml 引用的页面部件:"
|
||||||
|
+ "、".join(orphaned_slides[:20]),
|
||||||
|
)
|
||||||
|
)
|
||||||
|
errors.extend(_validate_charts(archive, names))
|
||||||
|
|
||||||
|
try:
|
||||||
|
shape_errors, shape_warnings, slide_count = _visual_structure(path)
|
||||||
|
errors.extend(shape_errors)
|
||||||
|
warnings.extend(shape_warnings)
|
||||||
|
except Exception as exc:
|
||||||
|
errors.append(
|
||||||
|
_issue(
|
||||||
|
"python_pptx_open_failed",
|
||||||
|
f"python-pptx 无法打开演示文稿:{exc}",
|
||||||
|
)
|
||||||
|
)
|
||||||
|
slide_count = 0
|
||||||
|
|
||||||
|
render_result: Optional[dict[str, Any]] = None
|
||||||
|
if check_render and not any(
|
||||||
|
item["code"] in {"missing_required_parts", "invalid_xml"}
|
||||||
|
for item in errors
|
||||||
|
):
|
||||||
|
from pypdf import PdfReader
|
||||||
|
|
||||||
|
with tempfile.TemporaryDirectory(prefix="pptx-validate-render-") as temp_name:
|
||||||
|
temp_dir = Path(temp_name)
|
||||||
|
staged = temp_dir / f"presentation{path.suffix.lower()}"
|
||||||
|
shutil.copy2(path, staged)
|
||||||
|
pdf, office_output = run_soffice_convert(
|
||||||
|
staged,
|
||||||
|
target_format="pdf",
|
||||||
|
output_dir=temp_dir / "pdf",
|
||||||
|
timeout=timeout,
|
||||||
|
)
|
||||||
|
pdf_pages = len(PdfReader(str(pdf)).pages)
|
||||||
|
render_result = {
|
||||||
|
"pdf_pages": pdf_pages,
|
||||||
|
"office_stdout": office_output["stdout"],
|
||||||
|
"office_stderr": office_output["stderr"],
|
||||||
|
}
|
||||||
|
if pdf_pages != slide_count:
|
||||||
|
errors.append(
|
||||||
|
_issue(
|
||||||
|
"render_page_count_mismatch",
|
||||||
|
f"结构页数为 {slide_count},渲染页数为 {pdf_pages}",
|
||||||
|
)
|
||||||
|
)
|
||||||
|
return {
|
||||||
|
"errors": errors,
|
||||||
|
"warnings": warnings,
|
||||||
|
"archive": archive_info,
|
||||||
|
"slide_count": slide_count,
|
||||||
|
"external_relationships": externals,
|
||||||
|
"render": render_result,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _signature(issue: dict[str, Any]) -> tuple[Any, ...]:
|
||||||
|
return (
|
||||||
|
issue.get("code"),
|
||||||
|
issue.get("part"),
|
||||||
|
issue.get("slide"),
|
||||||
|
issue.get("message"),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> dict[str, Any]:
|
||||||
|
args = build_parser().parse_args()
|
||||||
|
source = input_file(args.input, OOXML_PRESENTATION_SUFFIXES)
|
||||||
|
result = _validate_once(
|
||||||
|
source,
|
||||||
|
check_render=args.check_render,
|
||||||
|
timeout=args.timeout,
|
||||||
|
)
|
||||||
|
baseline_count = 0
|
||||||
|
if args.original:
|
||||||
|
original = input_file(args.original, OOXML_PRESENTATION_SUFFIXES)
|
||||||
|
baseline = _validate_once(
|
||||||
|
original,
|
||||||
|
check_render=False,
|
||||||
|
timeout=args.timeout,
|
||||||
|
)
|
||||||
|
baseline_signatures = {
|
||||||
|
_signature(item)
|
||||||
|
for item in baseline["errors"]
|
||||||
|
if item["code"]
|
||||||
|
in {
|
||||||
|
"invalid_xml",
|
||||||
|
"shape_out_of_bounds",
|
||||||
|
"python_pptx_open_failed",
|
||||||
|
}
|
||||||
|
}
|
||||||
|
before = len(result["errors"])
|
||||||
|
result["errors"] = [
|
||||||
|
item
|
||||||
|
for item in result["errors"]
|
||||||
|
if _signature(item) not in baseline_signatures
|
||||||
|
]
|
||||||
|
baseline_count = before - len(result["errors"])
|
||||||
|
result["original"] = str(original)
|
||||||
|
status = "valid" if not result["errors"] else "invalid"
|
||||||
|
return {
|
||||||
|
"source": str(source),
|
||||||
|
"status": status,
|
||||||
|
"issue_count": len(result["errors"]),
|
||||||
|
"warning_count": len(result["warnings"]),
|
||||||
|
"baselined_issue_count": baseline_count,
|
||||||
|
**result,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(run_cli(main))
|
||||||
Loading…
Reference in New Issue
Block a user