diff --git a/skills/docx/SKILL.md b/skills/docx/SKILL.md index a63b976..534df2a 100644 --- a/skills/docx/SKILL.md +++ b/skills/docx/SKILL.md @@ -1,6 +1,6 @@ --- name: docx -description: "创建、读取、编辑、转换、批注、接受修订、校验和渲染本地或远程 HTTPS Microsoft Word 文档。用户提到 Word、文档、报告、备忘录、合同、信函、模板、目录、页眉页脚、页码、表格、图片、批注或修订,或提供 HTTPS Word 地址、.docx、.dotx、.doc 文件时使用;支持安全下载、结构化创建、跨 Run 查找替换、安全 OOXML 解包/打包、旧格式转换、关系与 XML 校验及逐页视觉检查。若主要交付物是 PDF、电子表格、Google Docs 或普通代码,则不要使用。" +description: "创建、读取、编辑、转换、批注、接受修订、校验和渲染本地或远程 HTTPS Microsoft Word 文档,并按需识别文档图片、截图和扫描页中的文字。用户提到 Word、文档、报告、备忘录、合同、信函、模板、目录、页眉页脚、页码、表格、图片、图片文字 OCR、批注或修订,或提供 HTTPS Word 地址、.docx、.dotx、.doc 文件时使用;支持安全下载、结构化创建、跨 Run 查找替换、本地 RapidOCR、安全 OOXML 解包/打包、旧格式转换、关系与 XML 校验及逐页视觉检查。若主要交付物是 PDF、电子表格、Google Docs 或普通代码,则不要使用。" --- # Word 文档处理 @@ -13,6 +13,7 @@ description: "创建、读取、编辑、转换、批注、接受修订、校验 - 不把 `python3`、`soffice`、`libreoffice`、`pandoc`、`pdftoppm`、`zip`、`unzip`、`find`、`rm` 或其他系统命令作为脚本参数。 - LibreOffice、Pandoc、Poppler 和 ZIP 操作只允许由固定 Python 脚本在内部调用。 - 每次检查脚本返回 JSON;只有 `ok` 为 `true` 时才继续。`validate_document.py` 还必须返回 `status: valid`。 +- 只在需要读取图片、截图或扫描页中的文字时调用 `ocr_document.py`。只使用 `pages[]` 中 `usable_for_summary: true` 的 `text`;低置信度结果不得作为可靠正文。 - 远程地址只交给 `download_document.py`;不要在回复、日志摘要或文件名中复述可能含敏感查询参数的完整 URL。 - 不覆盖用户提供的源文件。最终结果写入 `output/docx/`,中间产物写入 `tmp/docx/<任务名>/`。 - 环境已预置依赖,不安装软件包,也不提示用户安装依赖。 @@ -23,6 +24,7 @@ description: "创建、读取、编辑、转换、批注、接受修订、校验 | --- | --- | --- | | `scripts/download_document.py` | 下载并校验远程 HTTPS Word 文档 | Python `urllib`、安全 OOXML 解析、`python-docx` | | `scripts/inspect_document.py` | 分段读取正文、表格、样式、批注和修订 | `python-docx`、安全 OOXML 解析 | +| `scripts/ocr_document.py` | 按页识别图片、截图和扫描页中的文字 | RapidOCR、ONNX Runtime、LibreOffice、Poppler、`pdfplumber` | | `scripts/create_document.py` | 按受控 JSON 创建专业 DOCX | `python-docx`、Pillow | | `scripts/edit_document.py` | 查找替换、追加/插入内容、调整样式和页面 | `python-docx` | | `scripts/add_comment.py` | 给精确文本范围添加批注 | `python-docx` | @@ -38,12 +40,13 @@ description: "创建、读取、编辑、转换、批注、接受修订、校验 1. 输入是 HTTPS 地址时,先调用 `download_document.py` 下载到本次任务临时目录;本地文件直接进入下一步。 2. 旧版 `.doc` 或模板 `.dotx` 先调用 `convert_document.py` 转为 `.docx`;保留原文件。 3. 编辑、总结或重组现有文档前调用 `inspect_document.py`,确认段落、表格、章节、页眉页脚、批注和修订状态。 -4. 新建文档使用 `create_document.py`;常规编辑使用 `edit_document.py`;添加批注使用 `add_comment.py`。 -5. 输入有修订时,先确认用户希望保留还是接受。普通编辑脚本默认拒绝含修订的文档,避免把修订静默损坏。 -6. 只有固定编辑脚本不能完成的 OOXML 高级需求,才使用 `unpack_document.py` → 编辑 XML 文件 → `pack_document.py`;不得直接运行 ZIP 或 shell 命令。 -7. 所有创建或修改结果必须调用 `validate_document.py --check-convert`,确保 `status: valid`、`issue_count: 0`。 -8. 再调用 `render_document.py` 渲染全部页面,逐页检查版式;有游标时继续到 `has_more: false`。 -9. 结构、内容、修订/批注和视觉检查都通过后才交付。 +4. 需要读取截图、扫描页或图片中的文字时调用 `ocr_document.py`。省略 `--pages` 可自动选择含有效图片的页面;不要默认 OCR 没有图片的普通正文页,也不要用 OCR 覆盖可靠的原生文本。 +5. 新建文档使用 `create_document.py`;常规编辑使用 `edit_document.py`;添加批注使用 `add_comment.py`。 +6. 输入有修订时,先确认用户希望保留还是接受。普通编辑脚本默认拒绝含修订的文档,避免把修订静默损坏。 +7. 只有固定编辑脚本不能完成的 OOXML 高级需求,才使用 `unpack_document.py` → 编辑 XML 文件 → `pack_document.py`;不得直接运行 ZIP 或 shell 命令。 +8. 所有创建或修改结果必须调用 `validate_document.py --check-convert`,确保 `status: valid`、`issue_count: 0`。 +9. 再调用 `render_document.py` 渲染全部页面,逐页检查版式;有游标时继续到 `has_more: false`。 +10. 结构、内容、修订/批注和视觉检查都通过后才交付。 ## 下载远程文档 @@ -85,9 +88,36 @@ description: "创建、读取、编辑、转换、批注、接受修订、校验 - `tracked_changes.total` 和 `authors`:是否存在修订及修订作者。 - `comments`:批注正文和作者。 - `sections`:纸张、方向、页边距、页眉、页脚。 +- `has_images`、`inline_image_count`、`media_part_count`:是否需要进一步读取图片文字;浮动图片可能只计入媒体部件。 - `archive.missing_required_parts`、`duplicate_members`:结构异常。 - `has_more`、`next_paragraph`、`next_table`:继续读取长文档。 +## 识别图片中的文字 + +需要读取图片、截图或扫描页中的文字时调用: + +```text +--input 'source.docx' +``` + +省略 `--pages` 时,脚本会把 Word 临时转换为 PDF,自动选择包含足够大图片的页面,每次最多处理 4 页。需要识别较小图片或指定页面时传: + +```text +--input 'source.docx' --pages '2,5-6' +``` + +脚本通过 LibreOffice 和 Poppler 临时渲染页面,使用本地 RapidOCR 识别图片区域;临时 PDF 和 PNG 会自动删除,不联网,也不调用大模型识图。PDF 原生文本层用于过滤正文、页眉、页脚和页码产生的重复 OCR,因此 `pages[].text` 只返回可靠的额外图片文字。 + +检查: + +- `candidate_pages`:自动检测到的图片页;`selection_mode` 表示自动或显式选页。 +- `status: good` 且 `usable_for_summary: true`:可以把 `text` 补充到原生文档内容中。 +- `status: no_image_text`:图片区域没有识别到额外文字,不是错误。 +- `status: sparse` 或 `low_confidence`:不要使用返回文字;根据 `needs_review` 人工核验。 +- `filtered_native_line_count` 和 `filtered_outside_image_line_count`:被当作原生文字或图片区域外文字过滤的 OCR 行数。 + +默认 260 DPI,可用 `--dpi 150-400` 调整。若 `has_more: true`:`next_offset > 0` 时传 `--pages --start-offset `;`next_offset = 0` 时把 `remaining_pages` 作为下一次 `--pages`。普通小徽标和面积不足页面约 1.5% 的图片不会进入自动候选,但仍可用 `--pages` 显式识别。 + ## 创建文档 调用 `scripts/create_document.py`: diff --git a/skills/docx/agents/openai.yaml b/skills/docx/agents/openai.yaml index eb4b46b..f3a2bd4 100644 --- a/skills/docx/agents/openai.yaml +++ b/skills/docx/agents/openai.yaml @@ -1,4 +1,4 @@ interface: display_name: "Word 文档" - short_description: "安全下载、创建、编辑、校验并渲染专业 Word 文档" - default_prompt: "使用 $docx 创建或处理本地文件或 HTTPS 链接中的 Word 文档,并完成内容与版式校验。" + short_description: "安全读取、编辑、图片 OCR、校验并渲染专业 Word 文档" + default_prompt: "使用 $docx 读取或创建 Word 文档,按需识别图片文字,并完成内容与版式校验。" diff --git a/skills/docx/scripts/inspect_document.py b/skills/docx/scripts/inspect_document.py index f439a6c..a28ffa0 100644 --- a/skills/docx/scripts/inspect_document.py +++ b/skills/docx/scripts/inspect_document.py @@ -333,6 +333,9 @@ def main() -> dict[str, Any]: "paragraph_count": len(all_paragraphs), "table_count": len(document.tables), "section_count": len(document.sections), + "inline_image_count": len(document.inline_shapes), + "media_part_count": archive["media_count"], + "has_images": archive["media_count"] > 0, "character_count": len(total_text), "word_count_estimate": len(total_text.split()), "sections": sections, diff --git a/skills/docx/scripts/ocr_document.py b/skills/docx/scripts/ocr_document.py new file mode 100644 index 0000000..2d73e93 --- /dev/null +++ b/skills/docx/scripts/ocr_document.py @@ -0,0 +1,804 @@ +#!/usr/bin/env python3 + +from __future__ import annotations + +import contextlib +import difflib +import importlib.metadata +import io +import logging +import os +import re +import shutil +import tempfile +import time +import unicodedata +from pathlib import Path +from typing import Any + +from _docx_common import ( + WORD_INPUT_SUFFIXES, + SkillArgumentParser, + find_program, + input_file, + run_cli, + run_program, + run_soffice_convert, +) + + +DEFAULT_DPI = 260 +DEFAULT_MAX_CHARS = 24000 +DEFAULT_TIMEOUT_SECONDS = 180 +MAX_PAGES_PER_CALL = 4 +MAX_PIXELS_PER_PAGE = 20_000_000 +MIN_MEAN_CONFIDENCE = 0.60 +MIN_MEANINGFUL_CHARS = 5 +MIN_OCR_IMAGE_AREA_RATIO = 0.015 +WHITESPACE_PATTERN = re.compile(r"[ \t]+") + + +for variable, value in ( + ("OMP_NUM_THREADS", "2"), + ("OPENBLAS_NUM_THREADS", "1"), + ("MKL_NUM_THREADS", "1"), + ("NUMEXPR_NUM_THREADS", "1"), +): + os.environ.setdefault(variable, value) + +for logger_name in ("rapidocr", "RapidOCR", "onnxruntime"): + logging.getLogger(logger_name).setLevel(logging.ERROR) + + +def build_parser(): + parser = SkillArgumentParser( + description="渲染 Word 页面并用本地 OCR 提取图片中的文字。" + ) + parser.add_argument("--input", required=True) + parser.add_argument( + "--pages", + help=( + "要识别的页码,例如 2 或 2,5-6;省略时自动选择含可读图片的页面" + ), + ) + parser.add_argument( + "--start-offset", + type=int, + default=0, + help="续读单页图片文字时的字符偏移量", + ) + parser.add_argument( + "--max-chars", + type=int, + default=DEFAULT_MAX_CHARS, + help=f"单次最多返回字符数,默认 {DEFAULT_MAX_CHARS}", + ) + parser.add_argument( + "--dpi", + type=int, + default=DEFAULT_DPI, + help=f"OCR 渲染分辨率,默认 {DEFAULT_DPI} DPI", + ) + parser.add_argument( + "--timeout", + type=int, + default=DEFAULT_TIMEOUT_SECONDS, + help=f"转换和单页渲染超时秒数,默认 {DEFAULT_TIMEOUT_SECONDS}", + ) + return parser + + +def _parse_page_spec(value: str, page_count: int) -> list[int]: + if not value.strip(): + raise ValueError("pages 不能为空") + pages: set[int] = set() + for raw_part in value.split(","): + part = raw_part.strip() + if not part: + continue + if "-" in part: + pieces = part.split("-", 1) + try: + start = int(pieces[0]) + end = int(pieces[1]) + except ValueError as exc: + raise ValueError(f"页码范围格式错误:{part}") from exc + if start > end: + raise ValueError(f"页码范围起始值不能大于结束值:{part}") + else: + try: + start = end = int(part) + except ValueError as exc: + raise ValueError(f"页码格式错误:{part}") from exc + if start < 1 or end > page_count: + raise ValueError(f"页码必须在 1 到 {page_count} 之间:{part}") + pages.update(range(start, end + 1)) + if not pages: + raise ValueError("pages 不能为空") + return sorted(pages) + + +def _clean_text(value: Any) -> str: + text = str(value or "").replace("\x00", "").strip() + return "\n".join( + WHITESPACE_PATTERN.sub(" ", line).strip() + for line in text.replace("\r\n", "\n").replace("\r", "\n").split("\n") + if line.strip() + ) + + +def _comparison_key(value: str) -> str: + normalized = unicodedata.normalize("NFKC", value).casefold() + return "".join(character for character in normalized if character.isalnum()) + + +def _normalized_box( + x0: Any, + top: Any, + x1: Any, + bottom: Any, + page_width: float, + page_height: float, +) -> tuple[float, float, float, float] | None: + try: + left = max(0.0, min(page_width, float(x0))) + upper = max(0.0, min(page_height, float(top))) + right = max(0.0, min(page_width, float(x1))) + lower = max(0.0, min(page_height, float(bottom))) + except (TypeError, ValueError): + return None + if ( + page_width <= 0 + or page_height <= 0 + or right <= left + or lower <= upper + ): + return None + return ( + left / page_width, + upper / page_height, + right / page_width, + lower / page_height, + ) + + +def _group_native_words( + words: list[dict[str, Any]], + page_width: float, + page_height: float, +) -> list[dict[str, Any]]: + sorted_words = sorted( + words, + key=lambda word: ( + round(float(word.get("top", 0.0)) / 2.5), + float(word.get("x0", 0.0)), + ), + ) + groups: list[list[dict[str, Any]]] = [] + group_top: float | None = None + for word in sorted_words: + try: + top = float(word.get("top", 0.0)) + except (TypeError, ValueError): + continue + if not groups or group_top is None or abs(top - group_top) > 3.0: + groups.append([word]) + group_top = top + else: + groups[-1].append(word) + group_top = sum( + float(item.get("top", 0.0)) for item in groups[-1] + ) / len(groups[-1]) + + lines: list[dict[str, Any]] = [] + for group in groups: + group.sort(key=lambda word: float(word.get("x0", 0.0))) + text = _clean_text( + " ".join(str(word.get("text", "")) for word in group) + ) + key = _comparison_key(text) + if not key: + continue + box = _normalized_box( + min(float(word.get("x0", 0.0)) for word in group), + min(float(word.get("top", 0.0)) for word in group), + max(float(word.get("x1", 0.0)) for word in group), + max(float(word.get("bottom", 0.0)) for word in group), + page_width, + page_height, + ) + if box is not None: + lines.append({"key": key, "box": box}) + return lines + + +def _page_profiles( + pdf_path: Path, +) -> tuple[dict[int, dict[str, Any]], list[str]]: + import pdfplumber + + profiles: dict[int, dict[str, Any]] = {} + warnings: list[str] = [] + with pdfplumber.open(pdf_path) as pdf: + for page_number, page in enumerate(pdf.pages, start=1): + page_width = float(page.width) + page_height = float(page.height) + try: + words = page.extract_words( + x_tolerance=2, + y_tolerance=3, + keep_blank_chars=False, + use_text_flow=False, + ) + except Exception as exc: + words = [] + warnings.append( + f"第 {page_number} 页原生文本层读取失败:{exc}" + ) + native_lines = _group_native_words( + list(words or []), + page_width, + page_height, + ) + native_keys = [line["key"] for line in native_lines] + + all_image_boxes: list[ + tuple[float, float, float, float] + ] = [] + ocr_image_boxes: list[ + tuple[float, float, float, float] + ] = [] + image_area = 0.0 + image_count = 0 + for image in page.images: + box = _normalized_box( + image.get("x0"), + image.get("top"), + image.get("x1"), + image.get("bottom"), + page_width, + page_height, + ) + if box is None: + continue + image_count += 1 + all_image_boxes.append(box) + area = (box[2] - box[0]) * (box[3] - box[1]) + image_area += area + if area >= MIN_OCR_IMAGE_AREA_RATIO: + ocr_image_boxes.append(box) + + profiles[page_number] = { + "image_count": image_count, + "ocr_image_count": len(ocr_image_boxes), + "image_area_ratio": round(min(1.0, image_area), 4), + "native_text_char_count": sum( + len(key) for key in native_keys + ), + "_native_keys": native_keys, + "_native_boxes": native_lines, + "_image_boxes": all_image_boxes, + } + return profiles, warnings + + +def _box_points(value: Any) -> list[list[float]] | None: + if value is None: + return None + try: + points = [ + [round(float(point[0]), 2), round(float(point[1]), 2)] + for point in value + ] + except (IndexError, TypeError, ValueError): + return None + return points if len(points) == 4 else None + + +def _ordered_lines(result: Any) -> list[dict[str, Any]]: + texts = list(getattr(result, "txts", None) or ()) + scores = list(getattr(result, "scores", None) or ()) + raw_boxes = getattr(result, "boxes", None) + boxes = list(raw_boxes) if raw_boxes is not None else [] + + lines: list[dict[str, Any]] = [] + for index, raw_text in enumerate(texts): + text = _clean_text(raw_text) + if not text: + continue + try: + confidence = float(scores[index]) + except (IndexError, TypeError, ValueError): + confidence = 0.0 + confidence = max(0.0, min(1.0, confidence)) + box = _box_points(boxes[index] if index < len(boxes) else None) + if box: + left = min(point[0] for point in box) + top = min(point[1] for point in box) + else: + left = float(index) + top = float(index) + lines.append( + { + "text": text, + "confidence": confidence, + "box": box, + "_left": left, + "_top": top, + "_index": index, + } + ) + lines.sort( + key=lambda line: ( + round(line["_top"] / 10.0), + line["_left"], + line["_index"], + ) + ) + return lines + + +def _similar_to_any( + candidate: str, + references: list[str], + *, + threshold: float, +) -> bool: + if not candidate: + return False + for reference in references: + if not reference: + continue + if candidate == reference: + return True + shorter = min(len(candidate), len(reference)) + longer = max(len(candidate), len(reference)) + if shorter >= 3 and candidate in reference: + return True + if ( + shorter >= 3 + and reference in candidate + and longer <= round(shorter * 1.25) + ): + return True + if shorter >= 3 and difflib.SequenceMatcher( + None, + candidate, + reference, + ).ratio() >= threshold: + return True + return False + + +def _line_center( + box: list[list[float]] | None, + image_width: int, + image_height: int, +) -> tuple[float, float] | None: + if not box or image_width <= 0 or image_height <= 0: + return None + return ( + sum(point[0] for point in box) / len(box) / image_width, + sum(point[1] for point in box) / len(box) / image_height, + ) + + +def _point_in_boxes( + point: tuple[float, float] | None, + boxes: list[tuple[float, float, float, float]], + *, + padding: float = 0.005, +) -> bool: + if point is None: + return False + x, y = point + return any( + x0 - padding <= x <= x1 + padding + and y0 - padding <= y <= y1 + padding + for x0, y0, x1, y1 in boxes + ) + + +def _line_is_native( + line: dict[str, Any], + profile: dict[str, Any], + image_width: int, + image_height: int, +) -> bool: + candidate = _comparison_key(line["text"]) + if _similar_to_any( + candidate, + profile["_native_keys"], + threshold=0.82, + ): + return True + + center = _line_center(line["box"], image_width, image_height) + if center is None: + return False + x, y = center + padding = 0.01 + for native_line in profile["_native_boxes"]: + x0, y0, x1, y1 = native_line["box"] + if ( + x0 - padding <= x <= x1 + padding + and y0 - padding <= y <= y1 + padding + and _similar_to_any( + candidate, + [native_line["key"]], + threshold=0.68, + ) + ): + return True + return False + + +def _create_ocr_engine(): + try: + from rapidocr import RapidOCR + except ImportError as exc: + raise RuntimeError("环境预置的 rapidocr 模块不可用") from exc + + captured_stdout = io.StringIO() + captured_stderr = io.StringIO() + with ( + contextlib.redirect_stdout(captured_stdout), + contextlib.redirect_stderr(captured_stderr), + ): + return RapidOCR() + + +def _ocr_page( + engine: Any, + image_path: Path, + profile: dict[str, Any], +) -> dict[str, Any]: + from PIL import Image + + with Image.open(image_path) as image: + image_width, image_height = image.size + + captured_stdout = io.StringIO() + captured_stderr = io.StringIO() + started = time.monotonic() + with ( + contextlib.redirect_stdout(captured_stdout), + contextlib.redirect_stderr(captured_stderr), + ): + result = engine(str(image_path)) + elapsed = time.monotonic() - started + + raw_lines = _ordered_lines(result) + image_lines: list[dict[str, Any]] = [] + seen: set[str] = set() + filtered_native = 0 + filtered_outside_images = 0 + filtered_duplicates = 0 + image_boxes = profile["_image_boxes"] + for line in raw_lines: + if _line_is_native(line, profile, image_width, image_height): + filtered_native += 1 + continue + center = _line_center(line["box"], image_width, image_height) + if image_boxes and not _point_in_boxes(center, image_boxes): + filtered_outside_images += 1 + continue + key = _comparison_key(line["text"]) + if key and key in seen: + filtered_duplicates += 1 + continue + if key: + seen.add(key) + image_lines.append(line) + + text = "\n".join(line["text"] for line in image_lines) + weighted_chars = [ + max(1, sum(1 for character in line["text"] if not character.isspace())) + for line in image_lines + ] + total_weight = sum(weighted_chars) + mean_confidence = ( + sum( + line["confidence"] * weight + for line, weight in zip(image_lines, weighted_chars) + ) + / total_weight + if total_weight + else 0.0 + ) + meaningful_chars = sum(1 for character in text if character.isalnum()) + low_confidence_lines = sum( + 1 + for line in image_lines + if line["confidence"] < MIN_MEAN_CONFIDENCE + ) + + reasons: list[str] = [] + if not text: + status = "no_image_text" + reasons.append("未识别到原生文本之外的图片文字") + elif meaningful_chars < MIN_MEANINGFUL_CHARS: + status = "sparse" + reasons.append( + f"图片中的有效文字少于 {MIN_MEANINGFUL_CHARS} 个字符" + ) + elif mean_confidence < MIN_MEAN_CONFIDENCE: + status = "low_confidence" + reasons.append( + "图片文字 OCR 平均置信度低于 " + f"{round(MIN_MEAN_CONFIDENCE * 100)}%" + ) + else: + status = "good" + + return { + "text": text, + "status": status, + "usable_for_summary": status == "good", + "needs_review": status in {"sparse", "low_confidence"}, + "raw_ocr_line_count": len(raw_lines), + "image_line_count": len(image_lines), + "filtered_native_line_count": filtered_native, + "filtered_outside_image_line_count": filtered_outside_images, + "filtered_duplicate_line_count": filtered_duplicates, + "low_confidence_line_count": low_confidence_lines, + "mean_confidence": round(mean_confidence, 4), + "meaningful_chars": meaningful_chars, + "reasons": reasons, + "ocr_seconds": round(elapsed, 3), + } + + +def _pdf_pages(path: Path) -> tuple[int, dict[int, tuple[float, float]]]: + from pypdf import PdfReader + + page_sizes: dict[int, tuple[float, float]] = {} + with path.open("rb") as stream: + reader = PdfReader(stream, strict=False) + if reader.is_encrypted: + raise ValueError("LibreOffice 生成了加密 PDF,无法执行 OCR") + page_count = len(reader.pages) + for page_number, page in enumerate(reader.pages, start=1): + page_sizes[page_number] = ( + abs(float(page.cropbox.width)), + abs(float(page.cropbox.height)), + ) + return page_count, page_sizes + + +def _render_page( + pdf_path: Path, + page_number: int, + page_size: tuple[float, float], + dpi: int, + timeout: int, + temp_dir: Path, +) -> tuple[Path, float]: + width_points, height_points = page_size + estimated_pixels = ( + width_points * dpi / 72.0 + * height_points * dpi / 72.0 + ) + if estimated_pixels > MAX_PIXELS_PER_PAGE: + raise ValueError( + f"第 {page_number} 页按 {dpi} DPI 渲染预计超过 " + f"{MAX_PIXELS_PER_PAGE} 像素,请降低 dpi" + ) + + prefix = temp_dir / f"page-{page_number:04d}" + output = prefix.with_suffix(".png") + started = time.monotonic() + run_program( + [ + find_program("pdftoppm"), + "-f", + str(page_number), + "-l", + str(page_number), + "-singlefile", + "-png", + "-r", + str(dpi), + str(pdf_path), + str(prefix), + ], + timeout=timeout, + ) + elapsed = time.monotonic() - started + if not output.is_file() or output.stat().st_size <= 0: + raise RuntimeError(f"第 {page_number} 页没有生成有效 PNG") + return output, elapsed + + +def _package_version(name: str) -> str | None: + try: + return importlib.metadata.version(name) + except importlib.metadata.PackageNotFoundError: + return None + + +def main() -> dict[str, Any]: + args = build_parser().parse_args() + if args.start_offset < 0: + raise ValueError("start-offset 不能小于 0") + if args.max_chars < 1 or args.max_chars > 60000: + raise ValueError("max-chars 必须在 1 到 60000 之间") + if args.dpi < 150 or args.dpi > 400: + raise ValueError("dpi 必须在 150 到 400 之间") + if args.timeout < 1 or args.timeout > 600: + raise ValueError("timeout 必须在 1 到 600 秒之间") + + source = input_file(args.input, WORD_INPUT_SUFFIXES) + page_outputs: list[dict[str, Any]] = [] + returned_chars = 0 + next_page: int | None = None + next_offset = 0 + remaining_pages: list[int] = [] + office_output = {"stdout": "", "stderr": ""} + + with tempfile.TemporaryDirectory(prefix="docx-ocr-") as temp_name: + temp_dir = Path(temp_name) + staged_input = temp_dir / f"document{source.suffix.lower()}" + shutil.copy2(source, staged_input) + pdf_path, office_output = run_soffice_convert( + staged_input, + target_format="pdf", + output_dir=temp_dir / "pdf", + timeout=args.timeout, + ) + page_count, page_sizes = _pdf_pages(pdf_path) + if page_count < 1: + raise ValueError("Word 文档没有可执行 OCR 的页面") + profiles, profile_warnings = _page_profiles(pdf_path) + if len(profiles) != page_count: + raise RuntimeError( + f"渲染结果有 {page_count} 页,但只检查到 " + f"{len(profiles)} 页" + ) + candidate_pages = [ + page_number + for page_number, profile in profiles.items() + if profile["ocr_image_count"] > 0 + ] + + selection_mode = "explicit" if args.pages else "auto" + if args.pages: + target_pages = _parse_page_spec(args.pages, page_count) + if len(target_pages) > MAX_PAGES_PER_CALL: + raise ValueError( + f"单次最多 OCR {MAX_PAGES_PER_CALL} 页," + "请拆分 pages 后重试" + ) + else: + target_pages = candidate_pages + + if args.start_offset > 0: + if not args.pages or len(target_pages) != 1: + raise ValueError( + "使用 start-offset 时必须显式指定且只指定一页" + ) + + selected_pages = target_pages[:MAX_PAGES_PER_CALL] + queued_pages = target_pages[MAX_PAGES_PER_CALL:] + if selected_pages: + engine = _create_ocr_engine() + for index, page_number in enumerate(selected_pages): + budget = args.max_chars - returned_chars + if budget <= 0: + next_page = page_number + remaining_pages = ( + selected_pages[index:] + queued_pages + ) + break + + image_path, render_seconds = _render_page( + pdf_path, + page_number, + page_sizes[page_number], + args.dpi, + args.timeout, + temp_dir, + ) + result = _ocr_page( + engine, + image_path, + profiles[page_number], + ) + full_text = result.pop("text") + offset = args.start_offset if index == 0 else 0 + if offset > len(full_text): + raise ValueError( + f"start-offset 超过第 {page_number} 页图片文字长度 " + f"{len(full_text)}" + ) + + usable = bool(result["usable_for_summary"]) + if not usable: + page_text = "" + complete = True + else: + remaining_text = full_text[offset:] + page_text = remaining_text[:budget] + complete = len(page_text) == len(remaining_text) + + profile = profiles[page_number] + page_outputs.append( + { + "page": page_number, + "text": page_text, + "char_count": len(full_text), + "offset_start": offset if usable else 0, + "offset_end": ( + offset + len(page_text) if usable else 0 + ), + "complete": complete, + "render_seconds": round(render_seconds, 3), + "image_count": profile["image_count"], + "ocr_image_count": profile["ocr_image_count"], + "image_area_ratio": profile[ + "image_area_ratio" + ], + "native_text_char_count": profile[ + "native_text_char_count" + ], + **result, + } + ) + returned_chars += len(page_text) + + if not complete: + next_page = page_number + next_offset = offset + len(page_text) + remaining_pages = ( + selected_pages[index + 1 :] + queued_pages + ) + break + + if next_page is None and queued_pages: + next_page = queued_pages[0] + remaining_pages = queued_pages + + all_processed = len(page_outputs) == len(selected_pages) + all_complete = all(page["complete"] for page in page_outputs) + all_safe = all( + page["status"] in {"good", "no_image_text"} + for page in page_outputs + ) + has_more = next_page is not None + return { + "source": str(source), + "page_count": page_count, + "selection_mode": selection_mode, + "candidate_pages": candidate_pages, + "selected_pages": selected_pages, + "processed_pages": [page["page"] for page in page_outputs], + "engine": "rapidocr", + "engine_version": _package_version("rapidocr"), + "runtime": "onnxruntime", + "runtime_version": _package_version("onnxruntime"), + "offline": True, + "dpi": args.dpi, + "returned_chars": returned_chars, + "pages": page_outputs, + "usable_for_summary": any( + page["usable_for_summary"] for page in page_outputs + ), + "complete_ocr_coverage": ( + not has_more + and all_processed + and all_complete + and all_safe + ), + "needs_review": any(page["needs_review"] for page in page_outputs), + "has_more": has_more, + "next_page": next_page, + "next_offset": next_offset, + "remaining_pages": remaining_pages, + "profile_warnings": profile_warnings, + "office_stdout": office_output["stdout"], + "office_stderr": office_output["stderr"], + } + + +if __name__ == "__main__": + raise SystemExit(run_cli(main)) diff --git a/skills/pptx/SKILL.md b/skills/pptx/SKILL.md new file mode 100644 index 0000000..60b8e7e --- /dev/null +++ b/skills/pptx/SKILL.md @@ -0,0 +1,419 @@ +--- +name: pptx +description: "创建、读取、编辑、复制页面、转换、校验和渲染本地或远程 HTTPS PowerPoint 演示文稿与模板,并按需识别页面图片、截图和图表中的文字。用户提到 PPT、PPTX、PowerPoint、演示文稿、幻灯片、路演稿、汇报材料、演讲者备注、模板、版式或图片文字 OCR,或提供 HTTPS 演示文稿地址、.pptx、.potx、.ppsx、.ppt 文件时使用;支持安全下载、结构化创建、保留 Run 格式的文本替换、页面删除/重排/复制、Markdown 提取、本地 RapidOCR、旧格式转换、受控 OOXML 解包/打包、关系与图表校验及逐页视觉检查。若主要交付物不是演示文稿且不需要读取或修改 PPT 内容,则不要使用。" +--- + +# PowerPoint 演示文稿处理 + +## 强制执行规则 + +当前智能体不能直接执行 shell、任意 Python、Node.js 或系统命令。只能通过 `execute_skill_script` 调用本 Skill 中真实存在的固定 Python 脚本。 + +- 只调用下表列出的可执行脚本;不执行 `scripts/` 目录、`scripts/_pptx_common.py`、`scripts/_presentation_builder.js` 或 `scripts/_icon_renderer.js`。 +- 不把 `python3`、`node`、`soffice`、`libreoffice`、`pdftoppm`、`zip`、`unzip`、`rm` 或其他系统命令作为脚本参数。 +- PptxGenJS、React Icons、Sharp、LibreOffice、Poppler 和 ZIP 操作只允许由固定脚本在内部调用。 +- 每次检查脚本返回的 JSON;只有 `ok` 为 `true` 时才继续。`validate_presentation.py` 还必须返回 `status: valid`、`issue_count: 0`。 +- 只在需要读取图片、截图或视觉图表中的文字时调用 `ocr_presentation.py`。只使用 `slides[]` 中 `usable_for_summary: true` 的 `text`;低置信度结果不得作为可靠正文。 +- 不覆盖用户提供的源文件。最终结果写入 `output/pptx/`,中间产物写入 `tmp/pptx/<任务名>/`。 +- 远程地址只交给 `download_presentation.py`;不要在回复、日志摘要或文件名中复述可能含敏感查询参数的完整 URL。 +- 环境已预置全部依赖,不安装软件包,也不提示用户安装依赖。 + +## 脚本清单 + +| 脚本 | 用途 | 底层能力 | +| --- | --- | --- | +| `scripts/download_presentation.py` | 下载并校验远程 HTTPS 演示文稿 | Python `urllib`、安全 OOXML 解析、`python-pptx` | +| `scripts/inspect_presentation.py` | 分段读取页面、文本、表格、图表、图片和备注 | `python-pptx`、安全 OOXML 解析 | +| `scripts/ocr_presentation.py` | 按页识别图片、截图和视觉图表中的文字 | RapidOCR、ONNX Runtime、LibreOffice、Poppler | +| `scripts/extract_presentation.py` | 提取整份演示文稿为 Markdown | `markitdown[pptx]` | +| `scripts/create_presentation.py` | 按受控 JSON 创建专业 PPTX | PptxGenJS | +| `scripts/edit_presentation.py` | 文本替换、删除/重排页面、修改属性 | `python-pptx` | +| `scripts/duplicate_slide.py` | 复制现有 PPTX 页面并维护包关系 | `lxml`、Python `zipfile` | +| `scripts/render_icon.py` | 把允许的 React Icons 图标渲染为 PNG | React、React DOM、React Icons、Sharp | +| `scripts/convert_presentation.py` | `.ppt/.potx/.ppsx/.pptx` 转 PPTX 或 PDF | LibreOffice | +| `scripts/unpack_presentation.py` | 安全解包 OOXML 供高级编辑 | Python `zipfile` | +| `scripts/pack_presentation.py` | 把 OOXML 目录安全打包为演示文稿 | Python `zipfile`、`python-pptx` | +| `scripts/validate_presentation.py` | 校验 ZIP、XML、关系、页面、图表、边界和可渲染性 | `defusedxml`、`lxml`、`python-pptx`、LibreOffice | +| `scripts/render_presentation.py` | 把演示文稿渲染为逐页 PNG、联系表和 PDF | LibreOffice、Poppler、Pillow | + +## 标准流程 + +1. 输入是 HTTPS 地址时,先调用 `download_presentation.py` 下载到本次任务临时目录;本地文件直接进入下一步。 +2. 旧版 `.ppt` 先调用 `convert_presentation.py` 转为 `.pptx`。需要把 `.potx/.ppsx` 当普通演示文稿编辑时,也先转为 `.pptx`。 +3. 编辑、总结或套用模板前先调用 `inspect_presentation.py`;内容较多时再调用 `extract_presentation.py` 获取完整 Markdown。 +4. 需要读取截图、扫描页或图片中的文字时,从 `image_slides` 选择相关页调用 `ocr_presentation.py`。不要默认 OCR 全部页面,也不要用 OCR 覆盖可靠的原生文本。 +5. 从零创建使用 `create_presentation.py`;常规文本与页面编辑使用 `edit_presentation.py`;复用模板页面使用 `duplicate_slide.py`。 +6. 只有固定脚本不能完成的精细模板编辑,才使用 `unpack_presentation.py` → 最小化编辑 OOXML → `pack_presentation.py`。 +7. 创建或修改后必须调用 `validate_presentation.py --check-render`。基于模板制作时同时传 `--original <原模板>`。 +8. 再调用 `render_presentation.py --contact-sheet` 渲染全部页面;若 `has_more: true`,用 `next_slide` 继续,直到检查完整份演示文稿。 +9. 逐页检查内容、顺序、溢出、重叠、间距、对齐、对比度、占位文本和备注。发现问题后修复并重新校验、重新渲染受影响页面。 +10. 文件结构、内容和视觉检查全部通过后才交付。 + +## 下载远程演示文稿 + +只接受 HTTPS 地址。完整保留 URL 及查询参数传给脚本,但不要在回复或输出文件名中暴露查询参数。 + +```text +--url 'https://example.com/deck.pptx?signature=...' --output 'tmp/pptx/<任务名>/source.pptx' +``` + +可选参数: + +- `--timeout <1-600>`:连接和读取超时秒数,默认 `60`。 +- `--max-bytes <字节数>`:默认 100 MiB,最高 512 MiB。 +- `--overwrite`:只覆盖本次任务生成的旧缓存。 + +`output` 扩展名必须与远程内容的真实格式一致,支持 `.pptx/.potx/.ppsx/.ppt`。脚本限制重定向只能继续使用 HTTPS,流式限制大小,先写临时文件,再原子发布。 + +## 检查和提取内容 + +调用结构检查: + +```text +--input 'source.pptx' --start-slide 1 --max-slides 30 +``` + +常用参数: + +- `--include-runs`:需要检查局部字体、粗体、字号或跨 Run 文本替换时使用。 +- `--max-shapes <1-1000>`、`--max-table-cells <1-10000>`:限制单次结构输出。 +- `--max-chars <1000-1000000>`:限制 JSON 中返回的正文量。 + +重点检查 `slide_size`、`slides[].layout`、形状边界、表格、图表系列、图片类型、`slides[].media`、`image_slides`、`speaker_notes`、批注摘要和 `external_relationships`。`media.image_count` 表示页面中的图片数量,`media.image_area_ratio` 是图片大致占页比例,`media.native_text_char_count` 是可直接提取的原生文字量。长演示文稿根据 `selection.has_more` 和 `next_slide` 分段读取。 + +需要连续正文时调用: + +```text +--input 'source.pptx' --output 'tmp/pptx/<任务名>/content.md' +``` + +Markdown 适合检查遗漏、错字和顺序,不代表页面版式。 + +## 识别图片中的文字 + +只有图片、截图、扫描页或视觉图表中的文字对任务有意义时才调用: + +```text +--input 'source.pptx' --slides '2,5-6' +``` + +`slides` 必须明确指定页码,单次最多 4 页。脚本先通过 LibreOffice 和 Poppler 临时渲染选定页面,再使用本地 RapidOCR 识别;渲染图片会自动删除,不联网,也不调用大模型识图。 + +脚本会根据原生文本框的位置和文字相似度过滤重复内容,因此 `slides[].text` 只返回原生文本之外的可靠图片文字。检查: + +- `status: good` 且 `usable_for_summary: true`:可以把 `text` 补充到同页原生文本中。 +- `status: no_image_text`:未发现额外图片文字,不是错误。 +- `status: sparse` 或 `low_confidence`:不要使用返回文字;根据 `needs_review` 人工核验。 +- `filtered_native_line_count`:被识别为原生文本并去重的 OCR 行数。 +- `picture_count`、`chart_count` 和 `image_area_ratio`:用于理解本页视觉内容规模。 + +默认 260 DPI。小字可使用 `--dpi 300-400`;单页预计像素过大时降低 DPI。可用 `--max-chars` 控制输出;若 `has_more: true`,根据 `next_slide`、`next_offset` 和 `remaining_slides` 继续。只有 `next_offset > 0` 时才传 `--start-offset`,且此时 `slides` 只能包含该页。 + +## 创建演示文稿 + +调用: + +```text +--output 'output/pptx/result.pptx' --spec '' +``` + +内容较长时先把 JSON 写入任务临时目录,再传 `--spec-file`。目标是本次任务旧产物且确认可覆盖时才传 `--overwrite`。 + +顶层结构: + +```json +{ + "layout": "LAYOUT_WIDE", + "properties": { + "title": "2026 年产品路线图", + "author": "示例公司", + "subject": "产品规划" + }, + "theme": { + "head_font": "Noto Sans CJK SC", + "body_font": "Noto Sans CJK SC", + "language": "zh-CN" + }, + "slides": [] +} +``` + +支持布局: + +| `layout` | 画布尺寸 | +| --- | --- | +| `LAYOUT_WIDE` | 13.333 × 7.5 英寸 | +| `LAYOUT_16X9` | 10 × 5.625 英寸 | +| `LAYOUT_4X3` | 10 × 7.5 英寸 | + +每页使用: + +```json +{ + "background": "0F172A", + "speaker_notes": "本页讲解约 45 秒。", + "elements": [] +} +``` + +坐标和尺寸 `x/y/w/h` 均使用英寸。颜色必须是不带 `#` 的 6 位十六进制值;不要把透明度拼入 8 位颜色,透明度使用 `transparency: 0-100`。所有可见元素必须位于画布内。 + +### 文本 + +```json +{ + "type": "text", + "text": "从洞察到增长", + "options": { + "x": 0.7, + "y": 0.6, + "w": 8.8, + "h": 0.7, + "fontFace": "Noto Sans CJK SC", + "fontSize": 30, + "bold": true, + "color": "F8FAFC", + "margin": 0, + "breakLine": false + } +} +``` + +局部格式使用 `runs`: + +```json +{ + "type": "text", + "runs": [ + {"text": "收入 ", "options": {"bold": true}}, + {"text": "+28%", "options": {"bold": true, "color": "22C55E"}} + ], + "options": { + "x": 0.8, + "y": 2.0, + "w": 4.0, + "h": 0.6, + "fontSize": 24, + "margin": 0 + } +} +``` + +列表不要输入字面量 `•`。每个列表项使用独立 Run,并设置 `bullet: true`;除最后一项外设置 `breakLine: true`,项目间距用 `paraSpaceAfter`。 + +### 形状、图片和图标 + +形状: + +```json +{ + "type": "shape", + "shape": "roundRect", + "options": { + "x": 0.8, + "y": 1.7, + "w": 3.6, + "h": 2.2, + "fill": {"color": "E0F2FE"}, + "line": {"color": "BAE6FD", "width": 1}, + "shadow": {"type": "outer", "color": "0F172A", "opacity": 0.15, "blur": 2, "angle": 45, "distance": 1} + } +} +``` + +常用形状名包括 `rect`、`roundRect`、`ellipse`、`line`、`chevron`、`triangle` 和 `hexagon`。阴影 `offset/distance` 不得为负数;向上投影时改变角度。 + +图片: + +```json +{ + "type": "image", + "path": "/absolute/path/chart.png", + "options": {"x": 7.2, "y": 1.4, "w": 5.2, "h": 4.8} +} +``` + +需要图标时先调用: + +```text +--library fi --name FiTrendingUp --color 2563EB --size 256 --output 'tmp/pptx/<任务名>/trend.png' +``` + +允许的图标库:`fa6`、`fi`、`hi2`、`io5`、`lu`、`md`、`ri`、`tb`。把生成的 PNG 作为普通图片元素插入。 + +### 图表 + +```json +{ + "type": "chart", + "chart_type": "bar", + "data": [ + { + "name": "收入", + "labels": ["Q1", "Q2", "Q3", "Q4"], + "values": [120, 148, 176, 215] + } + ], + "options": { + "x": 0.8, + "y": 1.6, + "w": 6.0, + "h": 4.7, + "catAxisLabelFontSize": 12, + "valAxisLabelFontSize": 11, + "showLegend": false, + "showValue": true, + "dataLabelPosition": "outEnd", + "chartColors": ["2563EB"] + } +} +``` + +支持 `area/bar/bar3d/bubble/bubble3d/doughnut/line/pie/radar/scatter`。PowerPoint 原生支持的图表必须保留为可编辑图表;只有桑基图、网络图等没有对应原生类型的可视化才使用图片。堆积条形图或柱形图的数据标签只能使用 `ctr/inEnd/inBase`,不能使用 `outEnd`。 + +### 表格 + +```json +{ + "type": "table", + "rows": [ + [ + {"text": "指标", "options": {"bold": true, "color": "FFFFFF", "fill": "1E3A8A"}}, + {"text": "本期", "options": {"bold": true, "color": "FFFFFF", "fill": "1E3A8A"}} + ], + ["收入", "2,150 万元"], + ["增长率", "28.0%"] + ], + "options": { + "x": 0.8, + "y": 1.8, + "w": 6.2, + "h": 2.2, + "border": {"type": "solid", "color": "CBD5E1", "pt": 1}, + "fontFace": "Noto Sans CJK SC", + "fontSize": 14, + "margin": 0.08 + } +} +``` + +## 编辑现有演示文稿 + +调用: + +```text +--input 'source.pptx' --output 'output/pptx/edited.pptx' --spec '' +``` + +JSON 顶层只有 `operations`,按数组顺序执行: + +| `operations[].type` | 关键字段 | +| --- | --- | +| `replace_text` | `find`、`replace`;可选 `slides/match_case/whole_word/count/required/include_notes` | +| `set_properties` | `properties`;其中 `revision` 必须是正整数 | +| `delete_slides` | `slides[]`,页码从 1 开始 | +| `reorder_slides` | `order[]`,必须完整且不重复地列出当前全部页码 | + +示例: + +```json +{ + "operations": [ + { + "type": "replace_text", + "find": "2025 年", + "replace": "2026 年", + "match_case": true, + "required": true + }, + { + "type": "reorder_slides", + "order": [1, 3, 2, 4] + } + ] +} +``` + +文本替换会处理同一段落内跨多个 Run 的匹配,并尽量保留替换起点和结尾的格式。输入含批注、ActiveX、宏或嵌入对象时,常规编辑脚本会停止,改用受控 OOXML 流程,避免静默丢失内容。输入含外部链接时默认停止;只有用户明确接受风险后才传 `--allow-external-links`。 + +## 复制模板页面 + +只对 `.pptx` 使用: + +```text +--input 'template.pptx' --output 'tmp/pptx/<任务名>/expanded.pptx' --slide 2 --after 4 +``` + +`slide` 是复制来源,`after` 是插入位置;省略 `after` 时紧跟来源页插入。脚本会更新页面清单、关系和内容类型,并移除不能安全共享的备注/批注关系。 + +复制页仍可能与原页共享图表、SmartArt 或嵌入对象部件。若返回的 `shared_relationships` 非空,修改这些对象前先做 OOXML 级独立复制;否则改动一页可能同时影响另一页。 + +## 高级 OOXML 编辑 + +固定编辑脚本无法表达且确实需要精细模板操作时: + +1. 调用 `unpack_presentation.py --input --output-dir <空目录>`。 +2. 先完成页面复制、删除和重排,再修改页面内容。 +3. 只最小化编辑相关 XML;不要重排、格式化或重写无关部件。 +4. 每个列表项保留独立的 ``;保留相邻 `` 以继承缩进与间距;不要在文本中写字面量项目符号。 +5. 有前后空格的 `` 设置 `xml:space="preserve"`。 +6. 调用 `pack_presentation.py --input-dir <目录> --output <新pptx>`。 +7. 立即调用 `validate_presentation.py --original <源pptx> --check-render`。 + +不得手工复制单个 `slideN.xml`。新页面必须同时登记到 `ppt/presentation.xml`、`ppt/_rels/presentation.xml.rels` 和 `[Content_Types].xml`,并处理页面关系。 + +## 设计和排版要求 + +- 先确定与主题相关的配色和一个贯穿全稿的视觉母题。一个主色承担约三分之二视觉权重,搭配 1–2 个辅助色和一个强调色。 +- 封面、章节页和结尾页可以使用深色背景,内容页使用浅色背景;同一套演示文稿保持一致。 +- 每页至少包含一种有信息作用的视觉元素:图片、图表、图标、流程、时间线或重点数字。不要只放标题和大段项目符号。 +- 交替使用双栏、卡片网格、半幅图片、对比栏和流程布局;不要连续复用同一种版式。 +- 中文正文优先使用 `Noto Sans CJK SC`,拉丁正文优先使用 Arial 或 Calibri。字体会在最终用户的 PowerPoint 中渲染,LibreOffice 预览可能发生替换;非预置字体至少保留约 10% 宽度余量。 +- 标题通常为 32–44 pt,分区标题 20–26 pt,正文 14–18 pt,注释 10–12 pt。正文左对齐;只对短标题或数字做居中。 +- 页面边缘至少留 0.5 英寸,内容块间距至少 0.3 英寸。同类元素使用一致的栅格、间距和对齐。 +- 文本框需要与图形边缘精确对齐时设置 `margin: 0`;字符间距使用 `charSpacing`,不要使用无效的 `letterSpacing`。 +- 不用标题下划线、整页装饰色条或卡片单侧色边充当“设计感”。优先使用留白、轻微底色、阴影、图片裁切和图标层次。 +- 不默认使用与主题无关的蓝色或米黄色;不交付低对比、越界、截断或互相重叠的元素。 +- 演讲者备注只写入 `speaker_notes`,不要伪装成页面上的隐藏文本。 + +## 校验与视觉检查 + +结构校验: + +```text +--input 'output/pptx/result.pptx' --check-render +``` + +模板派生结果: + +```text +--input 'output/pptx/result.pptx' --original 'template.pptx' --check-render +``` + +必须满足: + +- `status: valid` +- `issue_count: 0` +- `archive.missing_required_parts` 和 `duplicate_members` 为空 +- `render.pdf_pages` 与 `slide_count` 一致 + +警告也必须逐项评估,特别是 `placeholder_text`、空页面、孤立页面部件和外部关系。 + +渲染全部页面: + +```text +--input 'output/pptx/result.pptx' --output-dir 'tmp/pptx/<任务名>/rendered' --contact-sheet --include-pdf +``` + +默认 150 DPI、单次最多 30 页。可用 `--start-slide/--end-slide/--max-slides` 分批,复杂图表或小字可把 `--dpi` 提高到 180–220。联系表用于快速检查整体节奏,逐页 PNG 用于最终 QA。 + +逐页检查: + +- 文本是否被截断、溢出或因字体替换异常换行。 +- 图形、文字、页脚和来源是否重叠。 +- 页面边距、列宽、卡片尺寸、基线和间距是否一致。 +- 图标与文字是否有足够对比度,图表标签是否可读。 +- 模板占位内容、示例数据和多余装饰是否全部清除。 +- 页面顺序、标题层级、图表数值、演讲者备注和来源是否正确。 + +第一次渲染发现问题是正常的;修复后必须重新运行结构校验,并重新生成受影响页面的预览。 diff --git a/skills/pptx/agents/openai.yaml b/skills/pptx/agents/openai.yaml new file mode 100644 index 0000000..7c4d8b0 --- /dev/null +++ b/skills/pptx/agents/openai.yaml @@ -0,0 +1,4 @@ +interface: + display_name: "PowerPoint 演示文稿" + short_description: "创建、编辑、图片 OCR、校验并渲染 PowerPoint 演示文稿" + default_prompt: "使用 $pptx 读取或创建专业演示文稿,按需识别图片文字,并完成结构与视觉校验。" diff --git a/skills/pptx/scripts/_icon_renderer.js b/skills/pptx/scripts/_icon_renderer.js new file mode 100644 index 0000000..09d186c --- /dev/null +++ b/skills/pptx/scripts/_icon_renderer.js @@ -0,0 +1,106 @@ +#!/usr/bin/env node + +"use strict"; + +const fs = require("fs"); +const path = require("path"); +const React = require("react"); +const ReactDOMServer = require("react-dom/server"); +const sharp = require("sharp"); + +const LIBRARIES = { + fa6: "react-icons/fa6", + fi: "react-icons/fi", + hi2: "react-icons/hi2", + io5: "react-icons/io5", + lu: "react-icons/lu", + md: "react-icons/md", + ri: "react-icons/ri", + tb: "react-icons/tb", +}; + +function fail(message) { + process.stderr.write(`${message}\n`); + process.exit(1); +} + +function parseArgs(argv) { + const values = {}; + for (let index = 0; index < argv.length; index += 2) { + if (!argv[index]?.startsWith("--") || argv[index + 1] === undefined) { + fail("参数必须按 --name value 成对提供"); + } + values[argv[index].slice(2)] = argv[index + 1]; + } + if (!values.spec || !values.output) { + fail("缺少 --spec 或 --output"); + } + return values; +} + +function color(value, label) { + if (typeof value !== "string" || !/^[0-9A-Fa-f]{6}$/.test(value)) { + throw new Error(`${label} 必须是不带 # 的 6 位十六进制颜色`); + } + return `#${value.toUpperCase()}`; +} + +async function main() { + const args = parseArgs(process.argv.slice(2)); + const specPath = path.resolve(args.spec); + const outputPath = path.resolve(args.output); + if (path.extname(outputPath).toLowerCase() !== ".png") { + fail("output 必须使用 .png 扩展名"); + } + let spec; + try { + spec = JSON.parse(fs.readFileSync(specPath, "utf8")); + } catch (error) { + fail(`读取 spec 失败:${error.message}`); + } + try { + const libraryPath = LIBRARIES[spec.library]; + if (!libraryPath) { + throw new Error(`不支持的图标库:${spec.library}`); + } + if (typeof spec.name !== "string" || !/^[A-Za-z][A-Za-z0-9]*$/.test(spec.name)) { + throw new Error("name 必须是合法的 React Icons 导出名称"); + } + const size = spec.size ?? 256; + if (!Number.isInteger(size) || size < 32 || size > 2048) { + throw new Error("size 必须是 32 到 2048 的整数"); + } + const moduleExports = require(libraryPath); + const Icon = moduleExports[spec.name]; + if (typeof Icon !== "function") { + throw new Error(`${spec.library} 中不存在图标 ${spec.name}`); + } + const foreground = color(spec.color || "111827", "color"); + const markup = ReactDOMServer.renderToStaticMarkup( + React.createElement(Icon, { + color: foreground, + size, + title: typeof spec.title === "string" ? spec.title : undefined, + }), + ); + let pipeline = sharp(Buffer.from(markup)) + .resize(size, size, { fit: "contain" }); + if (spec.background) { + pipeline = pipeline.flatten({ background: color(spec.background, "background") }); + } + await pipeline.png().toFile(outputPath); + const metadata = await sharp(outputPath).metadata(); + process.stdout.write( + `${JSON.stringify({ + library: spec.library, + name: spec.name, + width: metadata.width, + height: metadata.height, + })}\n`, + ); + } catch (error) { + fail(error instanceof Error ? error.message : String(error)); + } +} + +main(); diff --git a/skills/pptx/scripts/_pptx_common.py b/skills/pptx/scripts/_pptx_common.py new file mode 100644 index 0000000..9d38592 --- /dev/null +++ b/skills/pptx/scripts/_pptx_common.py @@ -0,0 +1,426 @@ +#!/usr/bin/env python3 + +from __future__ import annotations + +import argparse +import json +import os +import shutil +import stat +import subprocess +import tempfile +import zipfile +from datetime import date, datetime, time +from decimal import Decimal +from pathlib import Path, PurePosixPath +from typing import Any, Callable, NoReturn, Optional +from urllib.parse import quote + + +OOXML_PRESENTATION_SUFFIXES = {".pptx", ".potx", ".ppsx"} +PRESENTATION_INPUT_SUFFIXES = OOXML_PRESENTATION_SUFFIXES | {".ppt"} +PRESENTATION_OUTPUT_SUFFIXES = {".pptx", ".potx", ".pdf"} +IMAGE_SUFFIXES = {".png", ".jpg", ".jpeg", ".webp"} + +MAX_ARCHIVE_MEMBERS = 20_000 +MAX_ARCHIVE_UNCOMPRESSED_BYTES = 1_073_741_824 +MAX_MEMBER_BYTES = 268_435_456 + +REL_NS = "http://schemas.openxmlformats.org/package/2006/relationships" +P_NS = "http://schemas.openxmlformats.org/presentationml/2006/main" +R_NS = "http://schemas.openxmlformats.org/officeDocument/2006/relationships" +A_NS = "http://schemas.openxmlformats.org/drawingml/2006/main" + + +class SkillArgumentParser(argparse.ArgumentParser): + def error(self, message: str) -> NoReturn: + raise ValueError(f"参数错误:{message}") + + +def json_default(value: Any) -> Any: + if isinstance(value, (datetime, date, time)): + return value.isoformat() + if isinstance(value, Decimal): + return float(value) + if isinstance(value, Path): + return str(value) + return str(value) + + +def emit(payload: dict[str, Any]) -> None: + print(json.dumps(payload, ensure_ascii=False, default=json_default)) + + +def failure_message(exc: Exception) -> str: + if isinstance(exc, (FileExistsError, FileNotFoundError)): + return str(exc) + if isinstance(exc, PermissionError): + return "文件处理失败:没有目标路径的访问权限" + if isinstance(exc, subprocess.TimeoutExpired): + return f"外部程序执行超时({exc.timeout} 秒)" + if isinstance(exc, subprocess.CalledProcessError): + stderr = (exc.stderr or "").strip() + return f"外部程序执行失败:{stderr or exc}" + if isinstance(exc, zipfile.BadZipFile): + return "文件不是有效的 OOXML 压缩包" + if isinstance(exc, OSError): + return f"文件处理失败:{exc}" + return str(exc) + + +def run_cli(action: Callable[[], dict[str, Any]]) -> int: + try: + result = action() + except Exception as exc: + emit({"ok": False, "error": failure_message(exc)}) + return 1 + emit({"ok": True, **result}) + return 0 + + +def input_file(value: str, suffixes: Optional[set[str]] = None) -> Path: + path = Path(value).expanduser().resolve() + if not path.exists(): + raise FileNotFoundError(f"输入文件不存在:{path}") + if not path.is_file(): + raise ValueError(f"输入路径不是文件:{path}") + if path.stat().st_size <= 0: + raise ValueError(f"输入文件为空:{path}") + if suffixes is not None and path.suffix.lower() not in suffixes: + expected = "、".join(sorted(suffixes)) + raise ValueError(f"不支持的输入格式 {path.suffix};允许:{expected}") + return path + + +def output_file( + value: str, + suffixes: Optional[set[str]] = None, + *, + overwrite: bool = False, +) -> Path: + path = Path(value).expanduser().resolve() + if suffixes is not None and path.suffix.lower() not in suffixes: + expected = "、".join(sorted(suffixes)) + raise ValueError(f"不支持的输出格式 {path.suffix};允许:{expected}") + if path.exists() and not path.is_file(): + raise ValueError(f"目标路径不是文件:{path}") + if path.exists() and not overwrite: + raise FileExistsError(f"目标文件已存在:{path}") + path.parent.mkdir(parents=True, exist_ok=True) + return path + + +def output_directory(value: str) -> Path: + path = Path(value).expanduser().resolve() + if path.exists() and not path.is_dir(): + raise ValueError(f"输出路径不是目录:{path}") + path.mkdir(parents=True, exist_ok=True) + return path + + +def publish_file(source: Path, destination: Path, *, overwrite: bool) -> None: + destination.parent.mkdir(parents=True, exist_ok=True) + if overwrite: + os.replace(source, destination) + return + try: + os.link(source, destination) + except FileExistsError as exc: + raise FileExistsError(f"目标文件已存在:{destination}") from exc + except OSError: + if destination.exists(): + raise FileExistsError(f"目标文件已存在:{destination}") + shutil.copy2(source, destination) + finally: + if source.exists(): + source.unlink() + + +def load_json_argument( + inline_value: Optional[str], + file_value: Optional[str], + *, + label: str, + max_bytes: int = 4 * 1024 * 1024, +) -> dict[str, Any]: + if bool(inline_value) == bool(file_value): + raise ValueError(f"{label} 必须且只能通过内联 JSON 或 JSON 文件提供一次") + if file_value: + source = input_file(file_value, {".json"}) + if source.stat().st_size > max_bytes: + raise ValueError(f"{label} JSON 文件超过 {max_bytes} 字节限制") + raw = source.read_text(encoding="utf-8") + else: + raw = inline_value or "" + if len(raw.encode("utf-8")) > max_bytes: + raise ValueError(f"{label} 内联 JSON 超过 {max_bytes} 字节限制") + try: + parsed = json.loads(raw) + except json.JSONDecodeError as exc: + raise ValueError( + f"{label} JSON 无效:第 {exc.lineno} 行第 {exc.colno} 列,{exc.msg}" + ) from exc + if not isinstance(parsed, dict): + raise ValueError(f"{label} JSON 顶层必须是对象") + return parsed + + +def find_program(*names: str) -> str: + for name in names: + resolved = shutil.which(name) + if resolved: + return resolved + raise FileNotFoundError(f"运行环境缺少命令:{' / '.join(names)}") + + +def run_program( + args: list[str], + *, + timeout: int, + cwd: Optional[Path] = None, + env: Optional[dict[str, str]] = None, +) -> subprocess.CompletedProcess[str]: + if timeout < 1 or timeout > 900: + raise ValueError("timeout 必须在 1 到 900 秒之间") + completed = subprocess.run( + args, + cwd=str(cwd) if cwd else None, + env=env, + stdin=subprocess.DEVNULL, + stdout=subprocess.PIPE, + stderr=subprocess.PIPE, + text=True, + timeout=timeout, + check=False, + ) + if completed.returncode != 0: + raise subprocess.CalledProcessError( + completed.returncode, + args, + output=completed.stdout, + stderr=completed.stderr, + ) + return completed + + +def office_profile_uri(profile: Path) -> str: + return "file://" + quote(str(profile.resolve()), safe="/:") + + +def run_soffice_convert( + source: Path, + *, + target_format: str, + output_dir: Path, + timeout: int, + filter_name: Optional[str] = None, +) -> tuple[Path, dict[str, str]]: + soffice = find_program("soffice", "libreoffice") + output_dir.mkdir(parents=True, exist_ok=True) + profile = Path(tempfile.mkdtemp(prefix="pptx-soffice-profile-")) + try: + cache_dir = profile / "cache" + cache_dir.mkdir() + process_env = os.environ.copy() + process_env["XDG_CACHE_HOME"] = str(cache_dir) + convert_arg = target_format + if filter_name: + convert_arg = f"{target_format}:{filter_name}" + completed = run_program( + [ + soffice, + f"-env:UserInstallation={office_profile_uri(profile)}", + "--headless", + "--nologo", + "--nodefault", + "--nolockcheck", + "--nofirststartwizard", + "--convert-to", + convert_arg, + "--outdir", + str(output_dir), + str(source), + ], + timeout=timeout, + env=process_env, + ) + expected = output_dir / f"{source.stem}.{target_format}" + if not expected.exists(): + candidates = sorted(output_dir.glob(f"{source.stem}.*")) + if len(candidates) == 1: + expected = candidates[0] + else: + raise RuntimeError( + "LibreOffice 未生成预期文件;" + f"stdout={completed.stdout[-1000:]!r} " + f"stderr={completed.stderr[-1000:]!r}" + ) + return expected, { + "stdout": completed.stdout[-2000:], + "stderr": completed.stderr[-2000:], + } + finally: + shutil.rmtree(profile, ignore_errors=True) + + +def _safe_archive_name(name: str) -> PurePosixPath: + candidate = PurePosixPath(name) + if ( + candidate.is_absolute() + or not candidate.parts + or ".." in candidate.parts + or candidate.parts[0].endswith(":") + ): + raise ValueError(f"OOXML 压缩包含不安全路径:{name}") + return candidate + + +def _zip_member_is_symlink(info: zipfile.ZipInfo) -> bool: + mode = (info.external_attr >> 16) & 0xFFFF + return stat.S_ISLNK(mode) + + +def inspect_archive(path: Path) -> dict[str, Any]: + total_size = 0 + xml_count = 0 + media_count = 0 + slide_count = 0 + with zipfile.ZipFile(path) as archive: + infos = archive.infolist() + if len(infos) > MAX_ARCHIVE_MEMBERS: + raise ValueError( + f"OOXML 压缩包成员过多:{len(infos)} > {MAX_ARCHIVE_MEMBERS}" + ) + seen: set[str] = set() + duplicates: list[str] = [] + for info in infos: + _safe_archive_name(info.filename) + if _zip_member_is_symlink(info): + raise ValueError(f"OOXML 压缩包含符号链接:{info.filename}") + if info.file_size > MAX_MEMBER_BYTES: + raise ValueError(f"OOXML 成员过大:{info.filename}") + total_size += info.file_size + if total_size > MAX_ARCHIVE_UNCOMPRESSED_BYTES: + raise ValueError("OOXML 解压后总大小超过安全限制") + if info.filename in seen: + duplicates.append(info.filename) + seen.add(info.filename) + if info.filename.endswith((".xml", ".rels")): + xml_count += 1 + if info.filename.startswith("ppt/media/") and not info.is_dir(): + media_count += 1 + if ( + info.filename.startswith("ppt/slides/slide") + and info.filename.endswith(".xml") + and "/_rels/" not in info.filename + ): + slide_count += 1 + required = { + "[Content_Types].xml", + "_rels/.rels", + "ppt/presentation.xml", + "ppt/_rels/presentation.xml.rels", + } + missing = sorted(required - seen) + return { + "member_count": len(infos), + "uncompressed_bytes": total_size, + "xml_part_count": xml_count, + "media_count": media_count, + "slide_part_count": slide_count, + "duplicate_members": duplicates, + "missing_required_parts": missing, + } + + +def safe_extract_presentation(source: Path, destination: Path) -> dict[str, Any]: + archive_info = inspect_archive(source) + if archive_info["missing_required_parts"]: + raise ValueError( + "演示文稿缺少必要部件:" + + "、".join(archive_info["missing_required_parts"]) + ) + destination.mkdir(parents=True, exist_ok=True) + root = destination.resolve() + with zipfile.ZipFile(source) as archive: + for info in archive.infolist(): + relative = _safe_archive_name(info.filename) + target = destination.joinpath(*relative.parts) + resolved = target.resolve() + if root not in resolved.parents and resolved != root: + raise ValueError(f"OOXML 成员逃逸输出目录:{info.filename}") + if info.is_dir(): + target.mkdir(parents=True, exist_ok=True) + continue + target.parent.mkdir(parents=True, exist_ok=True) + with archive.open(info, "r") as source_handle, target.open("wb") as target_handle: + shutil.copyfileobj(source_handle, target_handle) + return archive_info + + +def pack_presentation_directory(source_dir: Path, destination: Path) -> dict[str, Any]: + source_dir = source_dir.expanduser().resolve() + if not source_dir.exists() or not source_dir.is_dir(): + raise ValueError(f"待打包目录不存在或不是目录:{source_dir}") + required = { + source_dir / "[Content_Types].xml", + source_dir / "_rels" / ".rels", + source_dir / "ppt" / "presentation.xml", + source_dir / "ppt" / "_rels" / "presentation.xml.rels", + } + missing = sorted( + str(path.relative_to(source_dir)) for path in required if not path.is_file() + ) + if missing: + raise ValueError("待打包目录缺少必要部件:" + "、".join(missing)) + + files: list[Path] = [] + total_size = 0 + for path in sorted(source_dir.rglob("*")): + if path.is_symlink(): + raise ValueError(f"待打包目录包含符号链接:{path}") + if path.is_file(): + size = path.stat().st_size + if size > MAX_MEMBER_BYTES: + raise ValueError(f"待打包文件过大:{path}") + total_size += size + if total_size > MAX_ARCHIVE_UNCOMPRESSED_BYTES: + raise ValueError("待打包文件总大小超过安全限制") + files.append(path) + if len(files) > MAX_ARCHIVE_MEMBERS: + raise ValueError("待打包文件数量超过安全限制") + + with zipfile.ZipFile( + destination, + "w", + compression=zipfile.ZIP_DEFLATED, + compresslevel=6, + ) as archive: + for path in files: + archive.write(path, path.relative_to(source_dir).as_posix()) + return inspect_archive(destination) + + +def parse_xml_bytes(payload: bytes, *, label: str = "XML") -> Any: + try: + from defusedxml import ElementTree as ET + from defusedxml.common import DefusedXmlException + + try: + return ET.fromstring(payload) + except (ET.ParseError, DefusedXmlException, ValueError) as exc: + raise ValueError(f"{label} 解析失败:{exc}") from exc + except ModuleNotFoundError: + from lxml import etree + + parser = etree.XMLParser( + resolve_entities=False, + no_network=True, + recover=False, + huge_tree=False, + remove_comments=False, + ) + try: + return etree.fromstring(payload, parser=parser) + except (etree.XMLSyntaxError, ValueError) as exc: + raise ValueError(f"{label} 解析失败:{exc}") from exc diff --git a/skills/pptx/scripts/_presentation_builder.js b/skills/pptx/scripts/_presentation_builder.js new file mode 100644 index 0000000..30196e5 --- /dev/null +++ b/skills/pptx/scripts/_presentation_builder.js @@ -0,0 +1,431 @@ +#!/usr/bin/env node + +"use strict"; + +const fs = require("fs"); +const path = require("path"); +const PptxGenJS = require("pptxgenjs"); + +const LAYOUTS = { + LAYOUT_WIDE: { width: 13.333, height: 7.5 }, + LAYOUT_16X9: { width: 10, height: 5.625 }, + LAYOUT_4X3: { width: 10, height: 7.5 }, +}; + +function fail(message) { + process.stderr.write(`${message}\n`); + process.exit(1); +} + +function parseArgs(argv) { + const values = {}; + for (let index = 0; index < argv.length; index += 2) { + const key = argv[index]; + const value = argv[index + 1]; + if (!key || !key.startsWith("--") || value === undefined) { + fail("参数必须按 --name value 成对提供"); + } + values[key.slice(2)] = value; + } + if (!values.spec || !values.output) { + fail("缺少 --spec 或 --output"); + } + return values; +} + +function isPlainObject(value) { + return value !== null && typeof value === "object" && !Array.isArray(value); +} + +function clone(value) { + return JSON.parse(JSON.stringify(value)); +} + +function requireObject(value, label) { + if (!isPlainObject(value)) { + throw new Error(`${label} 必须是对象`); + } + return value; +} + +function requireArray(value, label, maxLength = 1000) { + if (!Array.isArray(value)) { + throw new Error(`${label} 必须是数组`); + } + if (value.length > maxLength) { + throw new Error(`${label} 超过 ${maxLength} 项限制`); + } + return value; +} + +function requireString(value, label, maxLength = 100000) { + if (typeof value !== "string") { + throw new Error(`${label} 必须是字符串`); + } + if (value.length > maxLength) { + throw new Error(`${label} 超过 ${maxLength} 字符限制`); + } + return value; +} + +function cleanColor(value, label) { + if (typeof value !== "string" || !/^[0-9A-Fa-f]{6}$/.test(value)) { + throw new Error(`${label} 必须是不带 # 的 6 位十六进制颜色`); + } + return value.toUpperCase(); +} + +function validateTree(value, label = "options", depth = 0) { + if (depth > 12) { + throw new Error(`${label} 嵌套层级过深`); + } + if (value === null || typeof value === "boolean" || typeof value === "number") { + if (typeof value === "number" && !Number.isFinite(value)) { + throw new Error(`${label} 包含非有限数值`); + } + return; + } + if (typeof value === "string") { + if (value.length > 200000) { + throw new Error(`${label} 字符串过长`); + } + return; + } + if (Array.isArray(value)) { + if (value.length > 5000) { + throw new Error(`${label} 数组过长`); + } + value.forEach((item, index) => validateTree(item, `${label}[${index}]`, depth + 1)); + return; + } + if (!isPlainObject(value)) { + throw new Error(`${label} 包含不支持的数据类型`); + } + const entries = Object.entries(value); + if (entries.length > 500) { + throw new Error(`${label} 字段过多`); + } + for (const [key, item] of entries) { + if (["__proto__", "prototype", "constructor"].includes(key)) { + throw new Error(`${label} 包含禁止字段 ${key}`); + } + if (key === "color" && typeof item === "string") { + cleanColor(item, `${label}.${key}`); + } + if (key === "chartColors" && Array.isArray(item)) { + item.forEach((color, index) => cleanColor(color, `${label}.chartColors[${index}]`)); + } + if (key === "offset" && label.toLowerCase().includes("shadow")) { + if (typeof item !== "number" || item < 0) { + throw new Error(`${label}.offset 必须是非负数`); + } + } + validateTree(item, `${label}.${key}`, depth + 1); + } + if (isPlainObject(value.fill) && value.fill.type === "gradient") { + throw new Error(`${label}.fill 不支持渐变;请使用渐变背景图片`); + } +} + +function requireBox(options, label) { + for (const key of ["x", "y", "w", "h"]) { + if (typeof options[key] !== "number" || !Number.isFinite(options[key])) { + throw new Error(`${label}.${key} 必须是有限数值`); + } + } +} + +function normalizeRuns(runs, label) { + return requireArray(runs, label, 2000).map((run, index) => { + requireObject(run, `${label}[${index}]`); + const text = requireString(run.text ?? "", `${label}[${index}].text`); + const options = clone(run.options || {}); + requireObject(options, `${label}[${index}].options`); + validateTree(options, `${label}[${index}].options`); + return { text, options }; + }); +} + +function normalizeTableRows(rows, label) { + return requireArray(rows, label, 500).map((row, rowIndex) => + requireArray(row, `${label}[${rowIndex}]`, 100).map((cell, columnIndex) => { + if (typeof cell === "string" || typeof cell === "number") { + return String(cell); + } + requireObject(cell, `${label}[${rowIndex}][${columnIndex}]`); + const normalized = { + text: requireString( + String(cell.text ?? ""), + `${label}[${rowIndex}][${columnIndex}].text`, + ), + options: clone(cell.options || {}), + }; + requireObject( + normalized.options, + `${label}[${rowIndex}][${columnIndex}].options`, + ); + validateTree( + normalized.options, + `${label}[${rowIndex}][${columnIndex}].options`, + ); + return normalized; + }), + ); +} + +function normalizeChartData(data, label) { + return requireArray(data, label, 100).map((series, index) => { + requireObject(series, `${label}[${index}]`); + const labels = requireArray(series.labels, `${label}[${index}].labels`, 10000); + const values = requireArray(series.values, `${label}[${index}].values`, 10000); + if (labels.length !== values.length) { + throw new Error(`${label}[${index}] 的 labels 与 values 长度必须一致`); + } + return { + name: requireString(String(series.name ?? ""), `${label}[${index}].name`, 500), + labels: labels.map((item) => String(item)), + values: values.map((item, valueIndex) => { + if (typeof item !== "number" || !Number.isFinite(item)) { + throw new Error(`${label}[${index}].values[${valueIndex}] 必须是有限数值`); + } + return item; + }), + }; + }); +} + +function resolveImagePath(value, label) { + const raw = requireString(value, label, 4096); + const resolved = path.resolve(raw); + if (!fs.existsSync(resolved) || !fs.statSync(resolved).isFile()) { + throw new Error(`${label} 文件不存在:${resolved}`); + } + return resolved; +} + +function elementBounds(element, options, slideNumber, index, dimensions, warnings) { + if (!["text", "shape", "image", "chart", "table"].includes(element.type)) { + return; + } + requireBox(options, `slides[${slideNumber - 1}].elements[${index}].options`); + const tolerance = 0.002; + if ( + options.x < -tolerance || + options.y < -tolerance || + options.x + options.w > dimensions.width + tolerance || + options.y + options.h > dimensions.height + tolerance + ) { + warnings.push({ + slide: slideNumber, + element: index + 1, + code: "out_of_bounds", + message: "元素超出演示文稿画布", + }); + } +} + +async function build(spec, outputPath) { + requireObject(spec, "spec"); + const slides = requireArray(spec.slides, "slides", 300); + if (slides.length === 0) { + throw new Error("slides 至少需要一页"); + } + + const layout = spec.layout || "LAYOUT_WIDE"; + if (!Object.prototype.hasOwnProperty.call(LAYOUTS, layout)) { + throw new Error(`layout 不支持:${layout}`); + } + const dimensions = LAYOUTS[layout]; + const pptx = new PptxGenJS(); + pptx.layout = layout; + + const properties = spec.properties || {}; + requireObject(properties, "properties"); + const propertyMap = { + author: "author", + company: "company", + subject: "subject", + title: "title", + comments: "comments", + revision: "revision", + }; + for (const [sourceKey, targetKey] of Object.entries(propertyMap)) { + if (properties[sourceKey] !== undefined) { + pptx[targetKey] = requireString( + String(properties[sourceKey]), + `properties.${sourceKey}`, + 4000, + ); + } + } + + const theme = spec.theme || {}; + requireObject(theme, "theme"); + const headFont = requireString( + String(theme.head_font || "Noto Sans CJK SC"), + "theme.head_font", + 200, + ); + const bodyFont = requireString( + String(theme.body_font || "Noto Sans CJK SC"), + "theme.body_font", + 200, + ); + pptx.theme = { + headFontFace: headFont, + bodyFontFace: bodyFont, + lang: requireString(String(theme.language || "zh-CN"), "theme.language", 40), + }; + pptx.lang = theme.language || "zh-CN"; + + const warnings = []; + const elementCounts = {}; + for (let slideIndex = 0; slideIndex < slides.length; slideIndex += 1) { + const slideSpec = requireObject(slides[slideIndex], `slides[${slideIndex}]`); + const slide = pptx.addSlide(); + if (slideSpec.background !== undefined) { + slide.background = { + color: cleanColor(slideSpec.background, `slides[${slideIndex}].background`), + }; + } + const elements = requireArray( + slideSpec.elements || [], + `slides[${slideIndex}].elements`, + 1000, + ); + if (elements.length === 0) { + warnings.push({ + slide: slideIndex + 1, + code: "empty_slide", + message: "页面没有可见元素", + }); + } + for (let elementIndex = 0; elementIndex < elements.length; elementIndex += 1) { + const label = `slides[${slideIndex}].elements[${elementIndex}]`; + const element = requireObject(elements[elementIndex], label); + const type = requireString(element.type, `${label}.type`, 50); + const options = clone(element.options || {}); + requireObject(options, `${label}.options`); + validateTree(options, `${label}.options`); + elementBounds( + element, + options, + slideIndex + 1, + elementIndex, + dimensions, + warnings, + ); + elementCounts[type] = (elementCounts[type] || 0) + 1; + + if (type === "text") { + if (options.fontFace === undefined) { + options.fontFace = bodyFont; + } + const content = + element.runs !== undefined + ? normalizeRuns(element.runs, `${label}.runs`) + : requireString(String(element.text ?? ""), `${label}.text`); + slide.addText(content, options); + } else if (type === "shape") { + const shapeName = requireString(element.shape || "rect", `${label}.shape`, 100); + const shapeType = pptx.ShapeType[shapeName]; + if (!shapeType) { + throw new Error(`${label}.shape 不支持:${shapeName}`); + } + slide.addShape(shapeType, options); + } else if (type === "image") { + if (element.path !== undefined) { + options.path = resolveImagePath(element.path, `${label}.path`); + } else if (element.data !== undefined) { + const data = requireString(element.data, `${label}.data`, 20_000_000); + if (!/^data:image\/(?:png|jpeg|jpg|webp);base64,/.test(data)) { + throw new Error(`${label}.data 必须是受支持图片的 base64 data URL`); + } + options.data = data; + } else { + throw new Error(`${label} 必须提供 path 或 data`); + } + slide.addImage(options); + } else if (type === "chart") { + const chartName = requireString(element.chart_type, `${label}.chart_type`, 50); + const chartType = pptx.ChartType[chartName]; + if (!chartType) { + throw new Error(`${label}.chart_type 不支持:${chartName}`); + } + const chartData = normalizeChartData(element.data, `${label}.data`); + const chartOptions = { + showLegend: chartData.length > 1, + showTitle: Boolean(options.title), + showValue: true, + chartColors: ["2563EB", "14B8A6", "F97316", "8B5CF6", "E11D48"], + catAxisLabelColor: "475569", + valAxisLabelColor: "475569", + valGridLine: { color: "E2E8F0", size: 1 }, + catAxisLabelFontFace: bodyFont, + valAxisLabelFontFace: bodyFont, + dataLabelFontFace: bodyFont, + legendFontFace: bodyFont, + titleFontFace: headFont, + ...options, + }; + validateTree(chartOptions, `${label}.options`); + slide.addChart(chartType, chartData, chartOptions); + } else if (type === "table") { + if (options.fontFace === undefined) { + options.fontFace = bodyFont; + } + const rows = normalizeTableRows(element.rows, `${label}.rows`); + slide.addTable(rows, options); + } else { + throw new Error(`${label}.type 不支持:${type}`); + } + } + if (slideSpec.speaker_notes !== undefined) { + const notes = requireString( + String(slideSpec.speaker_notes), + `slides[${slideIndex}].speaker_notes`, + 100000, + ); + slide.addNotes(notes); + } + } + + await pptx.writeFile({ fileName: outputPath }); + if (!fs.existsSync(outputPath) || fs.statSync(outputPath).size === 0) { + throw new Error("PptxGenJS 未生成有效输出文件"); + } + return { + slide_count: slides.length, + element_counts: elementCounts, + layout, + width_inches: dimensions.width, + height_inches: dimensions.height, + warnings, + }; +} + +async function main() { + const args = parseArgs(process.argv.slice(2)); + const specPath = path.resolve(args.spec); + const outputPath = path.resolve(args.output); + if (!fs.existsSync(specPath) || !fs.statSync(specPath).isFile()) { + fail(`spec 文件不存在:${specPath}`); + } + if (path.extname(outputPath).toLowerCase() !== ".pptx") { + fail("output 必须使用 .pptx 扩展名"); + } + let spec; + try { + spec = JSON.parse(fs.readFileSync(specPath, "utf8")); + } catch (error) { + fail(`读取 spec 失败:${error.message}`); + } + try { + const result = await build(spec, outputPath); + process.stdout.write(`${JSON.stringify(result)}\n`); + } catch (error) { + fail(error instanceof Error ? error.message : String(error)); + } +} + +main(); diff --git a/skills/pptx/scripts/convert_presentation.py b/skills/pptx/scripts/convert_presentation.py new file mode 100644 index 0000000..0cc996c --- /dev/null +++ b/skills/pptx/scripts/convert_presentation.py @@ -0,0 +1,91 @@ +#!/usr/bin/env python3 + +from __future__ import annotations + +import argparse +import shutil +import tempfile +from pathlib import Path +from typing import Any + +from _pptx_common import ( + PRESENTATION_INPUT_SUFFIXES, + SkillArgumentParser, + input_file, + output_file, + publish_file, + run_cli, + run_soffice_convert, +) + + +def build_parser() -> argparse.ArgumentParser: + parser = SkillArgumentParser(description="转换 PowerPoint 演示文稿格式。") + parser.add_argument("--input", required=True) + parser.add_argument("--output", required=True) + parser.add_argument("--timeout", type=int, default=180) + parser.add_argument("--overwrite", action="store_true") + return parser + + +def main() -> dict[str, Any]: + from pptx import Presentation + + args = build_parser().parse_args() + source = input_file(args.input, PRESENTATION_INPUT_SUFFIXES) + destination = output_file( + args.output, + {".pptx", ".pdf"}, + overwrite=args.overwrite, + ) + target_suffix = destination.suffix.lower() + if source.suffix.lower() == ".ppt" and target_suffix != ".pptx": + raise ValueError("旧版 .ppt 必须先转换为 .pptx,再转换为 PDF") + + with tempfile.TemporaryDirectory(prefix="pptx-convert-") as temp_name: + temp_dir = Path(temp_name) + staged_source = temp_dir / f"source{source.suffix.lower()}" + shutil.copy2(source, staged_source) + office_output = {"stdout": "", "stderr": ""} + if target_suffix == ".pptx" and source.suffix.lower() == ".pptx": + converted = temp_dir / "converted.pptx" + shutil.copy2(staged_source, converted) + else: + converted, office_output = run_soffice_convert( + staged_source, + target_format=target_suffix.lstrip("."), + output_dir=temp_dir / "converted", + timeout=args.timeout, + filter_name=( + "Impress MS PowerPoint 2007 XML" + if target_suffix == ".pptx" + else None + ), + ) + if target_suffix == ".pptx": + presentation = Presentation(str(converted)) + page_count = len(presentation.slides) + if page_count < 1: + raise ValueError("转换后的演示文稿不包含页面") + else: + from pypdf import PdfReader + + page_count = len(PdfReader(str(converted)).pages) + if page_count < 1: + raise ValueError("转换后的 PDF 不包含页面") + staged_output = temp_dir / f"publish{target_suffix}" + shutil.copy2(converted, staged_output) + publish_file(staged_output, destination, overwrite=args.overwrite) + return { + "source": str(source), + "path": str(destination), + "format": target_suffix.lstrip("."), + "page_count": page_count, + "size_bytes": destination.stat().st_size, + "office_stdout": office_output["stdout"], + "office_stderr": office_output["stderr"], + } + + +if __name__ == "__main__": + raise SystemExit(run_cli(main)) diff --git a/skills/pptx/scripts/create_presentation.py b/skills/pptx/scripts/create_presentation.py new file mode 100644 index 0000000..ecdecae --- /dev/null +++ b/skills/pptx/scripts/create_presentation.py @@ -0,0 +1,95 @@ +#!/usr/bin/env python3 + +from __future__ import annotations + +import argparse +import json +import os +import tempfile +from pathlib import Path +from typing import Any + +from _pptx_common import ( + SkillArgumentParser, + find_program, + load_json_argument, + output_file, + publish_file, + run_cli, + run_program, +) + + +def build_parser() -> argparse.ArgumentParser: + parser = SkillArgumentParser(description="按受控 JSON 说明创建 PowerPoint 演示文稿。") + parser.add_argument("--output", required=True) + parser.add_argument("--spec") + parser.add_argument("--spec-file") + parser.add_argument("--timeout", type=int, default=180) + parser.add_argument("--overwrite", action="store_true") + return parser + + +def _validate_spec(spec: dict[str, Any]) -> None: + slides = spec.get("slides") + if not isinstance(slides, list) or not slides: + raise ValueError("spec.slides 必须是非空数组") + if len(slides) > 300: + raise ValueError("单个演示文稿最多支持 300 页") + for index, slide in enumerate(slides): + if not isinstance(slide, dict): + raise ValueError(f"spec.slides[{index}] 必须是对象") + elements = slide.get("elements", []) + if not isinstance(elements, list): + raise ValueError(f"spec.slides[{index}].elements 必须是数组") + if len(elements) > 1000: + raise ValueError(f"spec.slides[{index}].elements 超过 1000 项限制") + + +def main() -> dict[str, Any]: + args = build_parser().parse_args() + destination = output_file(args.output, {".pptx"}, overwrite=args.overwrite) + spec = load_json_argument(args.spec, args.spec_file, label="演示文稿说明") + _validate_spec(spec) + builder = Path(__file__).with_name("_presentation_builder.js").resolve() + if not builder.is_file(): + raise FileNotFoundError(f"内部构建器不存在:{builder}") + + with tempfile.TemporaryDirectory(prefix="pptx-create-") as temp_name: + temp_dir = Path(temp_name) + spec_path = temp_dir / "spec.json" + staged_output = temp_dir / "presentation.pptx" + spec_path.write_text( + json.dumps(spec, ensure_ascii=False), + encoding="utf-8", + ) + completed = run_program( + [ + find_program("node"), + str(builder), + "--spec", + str(spec_path), + "--output", + str(staged_output), + ], + timeout=args.timeout, + cwd=Path.cwd(), + env=os.environ.copy(), + ) + if not staged_output.is_file() or staged_output.stat().st_size <= 0: + raise RuntimeError("PptxGenJS 未生成输出文件") + try: + builder_result = json.loads(completed.stdout.strip()) + except json.JSONDecodeError as exc: + raise RuntimeError("内部构建器返回了无效结果") from exc + publish_file(staged_output, destination, overwrite=args.overwrite) + + return { + "path": str(destination), + "size_bytes": destination.stat().st_size, + **builder_result, + } + + +if __name__ == "__main__": + raise SystemExit(run_cli(main)) diff --git a/skills/pptx/scripts/download_presentation.py b/skills/pptx/scripts/download_presentation.py new file mode 100644 index 0000000..8eec5c0 --- /dev/null +++ b/skills/pptx/scripts/download_presentation.py @@ -0,0 +1,271 @@ +#!/usr/bin/env python3 + +from __future__ import annotations + +import argparse +import os +import socket +import sys +import tempfile +import urllib.error +import urllib.parse +import urllib.request +import zipfile +from pathlib import Path +from typing import Any, NoReturn, Optional + +from _pptx_common import ( + PRESENTATION_INPUT_SUFFIXES, + emit, + failure_message, + inspect_archive, + output_file, + parse_xml_bytes, + publish_file, +) + + +DEFAULT_TIMEOUT_SECONDS = 60 +DEFAULT_MAX_BYTES = 100 * 1024 * 1024 +MAX_ALLOWED_BYTES = 512 * 1024 * 1024 +CHUNK_SIZE = 1024 * 1024 +USER_AGENT = "wechat-robot-pptx-skill/1.0" +OLE_COMPOUND_MAGIC = bytes.fromhex("D0CF11E0A1B11AE1") +CONTENT_TYPES_NS = "http://schemas.openxmlformats.org/package/2006/content-types" +PRESENTATION_CONTENT_TYPES = { + ( + "application/vnd.openxmlformats-officedocument." + "presentationml.presentation.main+xml" + ): ".pptx", + ( + "application/vnd.openxmlformats-officedocument." + "presentationml.template.main+xml" + ): ".potx", + ( + "application/vnd.openxmlformats-officedocument." + "presentationml.slideshow.main+xml" + ): ".ppsx", +} + + +class SkillArgumentParser(argparse.ArgumentParser): + def error(self, message: str) -> NoReturn: + raise ValueError(f"参数错误:{message}") + + +def _validate_https_url(value: str) -> str: + url = value.strip() + if not url: + raise ValueError("演示文稿 URL 不能为空") + if any(character.isspace() or ord(character) < 32 for character in url): + raise ValueError("演示文稿 URL 不能包含空白字符或控制字符") + parsed = urllib.parse.urlsplit(url) + if parsed.scheme.lower() != "https" or not parsed.hostname: + raise ValueError("演示文稿 URL 必须是有效的 HTTPS 地址") + if parsed.username is not None or parsed.password is not None: + raise ValueError("演示文稿 URL 不允许包含用户名或密码") + try: + parsed.port + except ValueError as exc: + raise ValueError("演示文稿 URL 端口格式不正确") from exc + return url + + +class HTTPSOnlyRedirectHandler(urllib.request.HTTPRedirectHandler): + def redirect_request(self, req, fp, code, msg, headers, newurl): + return super().redirect_request( + req, + fp, + code, + msg, + headers, + _validate_https_url(newurl), + ) + + +def _parse_args(argv: list[str]) -> argparse.Namespace: + parser = SkillArgumentParser(description="下载并校验远程 HTTPS 演示文稿") + parser.add_argument("--url", required=True) + parser.add_argument("--output", required=True) + parser.add_argument("--timeout", type=int, default=DEFAULT_TIMEOUT_SECONDS) + parser.add_argument("--max-bytes", type=int, default=DEFAULT_MAX_BYTES) + parser.add_argument("--overwrite", action="store_true") + args = parser.parse_args(argv) + args.url = _validate_https_url(args.url) + if args.timeout < 1 or args.timeout > 600: + raise ValueError("timeout 必须在 1 到 600 秒之间") + if args.max_bytes < 1 or args.max_bytes > MAX_ALLOWED_BYTES: + raise ValueError(f"max-bytes 必须在 1 到 {MAX_ALLOWED_BYTES} 之间") + args.output = output_file( + args.output, + PRESENTATION_INPUT_SUFFIXES, + overwrite=args.overwrite, + ) + return args + + +def _detect_ooxml_suffix(path: Path) -> str: + with zipfile.ZipFile(path) as archive: + root = parse_xml_bytes( + archive.read("[Content_Types].xml"), + label="[Content_Types].xml", + ) + for element in root.iter(f"{{{CONTENT_TYPES_NS}}}Override"): + if element.attrib.get("PartName") != "/ppt/presentation.xml": + continue + detected = PRESENTATION_CONTENT_TYPES.get( + element.attrib.get("ContentType", "") + ) + if detected: + return detected + raise ValueError("下载内容不是受支持的 PowerPoint OOXML 文件") + + +def _validate_ooxml(path: Path, expected_suffix: str) -> dict[str, Any]: + archive = inspect_archive(path) + if archive["missing_required_parts"]: + raise ValueError( + "下载内容不是有效的 PowerPoint OOXML 文件;缺少:" + + "、".join(archive["missing_required_parts"]) + ) + if archive["duplicate_members"]: + raise ValueError( + "PowerPoint 压缩包含重复成员:" + + "、".join(archive["duplicate_members"][:10]) + ) + detected_suffix = _detect_ooxml_suffix(path) + if detected_suffix != expected_suffix: + raise ValueError( + "下载内容的实际格式为 " + f"{detected_suffix},但 output 使用了 {expected_suffix}" + ) + + try: + from pptx import Presentation + except ImportError as exc: + raise RuntimeError("当前 Python 未加载环境预置的 python-pptx 模块") from exc + try: + presentation = Presentation(str(path)) + slide_count = len(presentation.slides) + except Exception as exc: + raise ValueError("下载内容不是可解析的 PowerPoint 演示文稿") from exc + return { + "format": detected_suffix.lstrip("."), + "slide_count": slide_count, + "validation": "ooxml-and-python-pptx", + "archive": archive, + } + + +def _validate_legacy_ppt(path: Path) -> dict[str, Any]: + with path.open("rb") as stream: + magic = stream.read(len(OLE_COMPOUND_MAGIC)) + if magic != OLE_COMPOUND_MAGIC: + raise ValueError("下载内容不是有效的旧版 PowerPoint 复合文件") + return { + "format": "ppt", + "validation": "ole-compound-signature", + } + + +def _validate_presentation(path: Path, suffix: str) -> dict[str, Any]: + if suffix == ".ppt": + return _validate_legacy_ppt(path) + return _validate_ooxml(path, suffix) + + +def _download(args: argparse.Namespace) -> dict[str, Any]: + output: Path = args.output + request = urllib.request.Request( + args.url, + headers={ + "Accept": ( + "application/vnd.openxmlformats-officedocument." + "presentationml.presentation," + "application/vnd.ms-powerpoint," + "application/octet-stream;q=0.9,*/*;q=0.1" + ), + "Accept-Encoding": "identity", + "User-Agent": USER_AGENT, + }, + method="GET", + ) + opener = urllib.request.build_opener(HTTPSOnlyRedirectHandler()) + temp_path: Optional[Path] = None + downloaded_bytes = 0 + try: + with tempfile.NamedTemporaryFile( + mode="wb", + prefix=f".{output.stem}.", + suffix=f".part{output.suffix}", + dir=str(output.parent), + delete=False, + ) as temp_file: + temp_path = Path(temp_file.name) + with opener.open(request, timeout=args.timeout) as response: + _validate_https_url(response.geturl()) + content_length = response.headers.get("Content-Length") + if content_length: + try: + expected_bytes = int(content_length) + except ValueError: + expected_bytes = 0 + if expected_bytes > args.max_bytes: + raise ValueError( + f"远程文件超过大小限制:最多允许 {args.max_bytes} 字节" + ) + while True: + chunk = response.read(CHUNK_SIZE) + if not chunk: + break + downloaded_bytes += len(chunk) + if downloaded_bytes > args.max_bytes: + raise ValueError( + f"远程文件超过大小限制:最多允许 {args.max_bytes} 字节" + ) + temp_file.write(chunk) + temp_file.flush() + os.fsync(temp_file.fileno()) + if downloaded_bytes == 0: + raise ValueError("远程服务器返回了空文件") + details = _validate_presentation(temp_path, output.suffix.lower()) + publish_file(temp_path, output, overwrite=args.overwrite) + temp_path = None + return { + "path": str(output), + "size_bytes": downloaded_bytes, + **details, + } + finally: + if temp_path is not None: + try: + temp_path.unlink(missing_ok=True) + except OSError: + pass + + +def _failure_message(exc: Exception) -> str: + if isinstance(exc, urllib.error.HTTPError): + return f"下载失败:远程服务器返回 HTTP {exc.code}" + if isinstance(exc, (TimeoutError, socket.timeout)): + return "下载失败:连接或读取超时" + if isinstance(exc, urllib.error.URLError): + if isinstance(exc.reason, (TimeoutError, socket.timeout)): + return "下载失败:连接或读取超时" + return "下载失败:无法访问远程服务器" + return failure_message(exc) + + +def main(argv: Optional[list[str]] = None) -> int: + try: + args = _parse_args(sys.argv[1:] if argv is None else argv) + result = _download(args) + except Exception as exc: + emit({"ok": False, "error": _failure_message(exc)}) + return 1 + emit({"ok": True, **result}) + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/skills/pptx/scripts/duplicate_slide.py b/skills/pptx/scripts/duplicate_slide.py new file mode 100644 index 0000000..3797c95 --- /dev/null +++ b/skills/pptx/scripts/duplicate_slide.py @@ -0,0 +1,292 @@ +#!/usr/bin/env python3 + +from __future__ import annotations + +import argparse +import posixpath +import re +import tempfile +import zipfile +from pathlib import Path, PurePosixPath +from typing import Any + +from _pptx_common import ( + P_NS, + R_NS, + REL_NS, + SkillArgumentParser, + input_file, + inspect_archive, + output_file, + publish_file, + run_cli, +) + + +CONTENT_TYPES_NS = "http://schemas.openxmlformats.org/package/2006/content-types" +SLIDE_REL_TYPE_SUFFIX = "/slide" +SLIDE_CONTENT_TYPE = ( + "application/vnd.openxmlformats-officedocument." + "presentationml.slide+xml" +) +DROP_RELATIONSHIP_SUFFIXES = ("/notesSlide", "/comments", "/comment") + + +def build_parser() -> argparse.ArgumentParser: + parser = SkillArgumentParser( + description="安全复制现有 PPTX 页面并更新 OOXML 包关系。" + ) + parser.add_argument("--input", required=True) + parser.add_argument("--output", required=True) + parser.add_argument("--slide", type=int, required=True) + parser.add_argument("--after", type=int) + parser.add_argument("--overwrite", action="store_true") + return parser + + +def _parse_xml(payload: bytes, label: str) -> Any: + from lxml import etree + + parser = etree.XMLParser( + resolve_entities=False, + no_network=True, + recover=False, + huge_tree=False, + remove_blank_text=False, + ) + try: + return etree.fromstring(payload, parser=parser) + except etree.XMLSyntaxError as exc: + raise ValueError(f"{label} 解析失败:{exc}") from exc + + +def _serialize_xml(root: Any) -> bytes: + from lxml import etree + + return etree.tostring( + root, + encoding="UTF-8", + xml_declaration=True, + standalone=True, + ) + + +def _next_relationship_id(existing: set[str]) -> str: + number = 1 + while f"rId{number}" in existing: + number += 1 + return f"rId{number}" + + +def _slide_rels_name(slide_part: str) -> str: + path = PurePosixPath(slide_part) + return (path.parent / "_rels" / f"{path.name}.rels").as_posix() + + +def _resolve_part_target(owner_part: str, target: str) -> str: + if not target or target.startswith("#"): + raise ValueError(f"关系目标无效:{target!r}") + if target.startswith("/"): + normalized = posixpath.normpath(target).lstrip("/") + else: + normalized = posixpath.normpath( + posixpath.join(posixpath.dirname(owner_part), target) + ) + if normalized == ".." or normalized.startswith("../"): + raise ValueError(f"关系目标逃逸 OOXML 包根目录:{target}") + return normalized.lstrip("/") + + +def _duplicate( + source: Path, + staged: Path, + *, + source_slide: int, + insert_after: int, +) -> dict[str, Any]: + with zipfile.ZipFile(source, "r") as incoming: + infos = incoming.infolist() + names = {info.filename for info in infos} + presentation_root = _parse_xml( + incoming.read("ppt/presentation.xml"), + "ppt/presentation.xml", + ) + presentation_rels_root = _parse_xml( + incoming.read("ppt/_rels/presentation.xml.rels"), + "ppt/_rels/presentation.xml.rels", + ) + content_types_root = _parse_xml( + incoming.read("[Content_Types].xml"), + "[Content_Types].xml", + ) + + slide_id_list = presentation_root.find(f"{{{P_NS}}}sldIdLst") + if slide_id_list is None: + raise ValueError("presentation.xml 缺少 sldIdLst") + slide_ids = list(slide_id_list) + if source_slide < 1 or source_slide > len(slide_ids): + raise ValueError(f"slide 超出页面总数 {len(slide_ids)}") + if insert_after < 1 or insert_after > len(slide_ids): + raise ValueError(f"after 超出页面总数 {len(slide_ids)}") + + rel_targets: dict[str, str] = {} + existing_rel_ids: set[str] = set() + slide_relationship_type = None + for relationship in presentation_rels_root: + relationship_id = relationship.attrib.get("Id", "") + existing_rel_ids.add(relationship_id) + target = relationship.attrib.get("Target", "") + if relationship.attrib.get("TargetMode") == "External": + continue + posix_target = _resolve_part_target( + "ppt/presentation.xml", + target, + ) + rel_targets[relationship_id] = posix_target + if relationship.attrib.get("Type", "").endswith(SLIDE_REL_TYPE_SUFFIX): + slide_relationship_type = relationship.attrib.get("Type") + if not slide_relationship_type: + raise ValueError("presentation.xml.rels 不包含 slide 关系类型") + + source_rel_id = slide_ids[source_slide - 1].attrib.get(f"{{{R_NS}}}id") + source_part = rel_targets.get(source_rel_id or "") + if not source_part or source_part not in names: + raise ValueError("无法解析待复制页面的 slide 部件") + + existing_numbers = [ + int(match.group(1)) + for name in names + if ( + match := re.fullmatch(r"ppt/slides/slide(\d+)\.xml", name) + ) + ] + new_part_number = max(existing_numbers, default=0) + 1 + new_part = f"ppt/slides/slide{new_part_number}.xml" + new_rels_part = _slide_rels_name(new_part) + new_rel_id = _next_relationship_id(existing_rel_ids) + numeric_ids = [ + int(item.attrib["id"]) + for item in slide_ids + if item.attrib.get("id", "").isdigit() + ] + new_slide_id = str(max(numeric_ids, default=255) + 1) + + from lxml import etree + + new_relationship = etree.Element( + f"{{{REL_NS}}}Relationship", + Id=new_rel_id, + Type=slide_relationship_type, + Target=f"slides/slide{new_part_number}.xml", + ) + presentation_rels_root.append(new_relationship) + new_slide_id_element = etree.Element( + f"{{{P_NS}}}sldId", + id=new_slide_id, + ) + new_slide_id_element.set(f"{{{R_NS}}}id", new_rel_id) + slide_id_list.insert(insert_after, new_slide_id_element) + + override_exists = any( + node.attrib.get("PartName") == f"/{new_part}" + for node in content_types_root.findall( + f"{{{CONTENT_TYPES_NS}}}Override" + ) + ) + if not override_exists: + content_types_root.append( + etree.Element( + f"{{{CONTENT_TYPES_NS}}}Override", + PartName=f"/{new_part}", + ContentType=SLIDE_CONTENT_TYPE, + ) + ) + + replacements = { + "ppt/presentation.xml": _serialize_xml(presentation_root), + "ppt/_rels/presentation.xml.rels": _serialize_xml( + presentation_rels_root + ), + "[Content_Types].xml": _serialize_xml(content_types_root), + } + additions = {new_part: incoming.read(source_part)} + source_rels_part = _slide_rels_name(source_part) + dropped_relationships: list[str] = [] + shared_relationships: list[dict[str, str]] = [] + if source_rels_part in names: + slide_rels_root = _parse_xml( + incoming.read(source_rels_part), + source_rels_part, + ) + for relationship in list(slide_rels_root): + relationship_type = relationship.attrib.get("Type", "") + if relationship_type.endswith(DROP_RELATIONSHIP_SUFFIXES): + dropped_relationships.append(relationship_type) + slide_rels_root.remove(relationship) + continue + if relationship_type.endswith( + ("/chart", "/diagramData", "/diagramDrawing", "/oleObject") + ): + shared_relationships.append( + { + "type": relationship_type, + "target": relationship.attrib.get("Target", ""), + } + ) + additions[new_rels_part] = _serialize_xml(slide_rels_root) + + with zipfile.ZipFile( + staged, + "w", + compression=zipfile.ZIP_DEFLATED, + compresslevel=6, + ) as outgoing: + for info in infos: + payload = replacements.get(info.filename) + if payload is None: + payload = incoming.read(info.filename) + outgoing.writestr(info, payload) + for name, payload in additions.items(): + outgoing.writestr(name, payload) + + archive = inspect_archive(staged) + return { + "source_slide": source_slide, + "insert_after": insert_after, + "new_slide": insert_after + 1, + "new_slide_part": new_part, + "dropped_relationship_types": dropped_relationships, + "shared_relationships": shared_relationships, + "archive": archive, + } + + +def main() -> dict[str, Any]: + from pptx import Presentation + + args = build_parser().parse_args() + source = input_file(args.input, {".pptx"}) + destination = output_file(args.output, {".pptx"}, overwrite=args.overwrite) + if source == destination: + raise ValueError("不能覆盖输入演示文稿;请使用新的 output 路径") + insert_after = args.slide if args.after is None else args.after + with tempfile.TemporaryDirectory(prefix="pptx-duplicate-") as temp_name: + staged = Path(temp_name) / "duplicated.pptx" + result = _duplicate( + source, + staged, + source_slide=args.slide, + insert_after=insert_after, + ) + presentation = Presentation(str(staged)) + result["slide_count"] = len(presentation.slides) + publish_file(staged, destination, overwrite=args.overwrite) + return { + "source": str(source), + "path": str(destination), + **result, + } + + +if __name__ == "__main__": + raise SystemExit(run_cli(main)) diff --git a/skills/pptx/scripts/edit_presentation.py b/skills/pptx/scripts/edit_presentation.py new file mode 100644 index 0000000..888c6de --- /dev/null +++ b/skills/pptx/scripts/edit_presentation.py @@ -0,0 +1,367 @@ +#!/usr/bin/env python3 + +from __future__ import annotations + +import argparse +import re +import tempfile +import zipfile +from pathlib import Path +from typing import Any, Iterable, Optional + +from _pptx_common import ( + OOXML_PRESENTATION_SUFFIXES, + SkillArgumentParser, + input_file, + load_json_argument, + output_file, + parse_xml_bytes, + publish_file, + run_cli, +) + + +def build_parser() -> argparse.ArgumentParser: + parser = SkillArgumentParser(description="按受控操作编辑 PowerPoint 演示文稿。") + parser.add_argument("--input", required=True) + parser.add_argument("--output", required=True) + parser.add_argument("--spec") + parser.add_argument("--spec-file") + parser.add_argument("--allow-external-links", action="store_true") + parser.add_argument("--overwrite", action="store_true") + return parser + + +def _external_relationship_count(path: Path) -> int: + count = 0 + with zipfile.ZipFile(path) as archive: + for name in archive.namelist(): + if not name.endswith(".rels"): + continue + root = parse_xml_bytes(archive.read(name), label=name) + count += sum( + 1 + for node in root + if node.attrib.get("TargetMode") == "External" + ) + return count + + +def _package_risks(path: Path) -> list[str]: + risks: list[str] = [] + with zipfile.ZipFile(path) as archive: + names = archive.namelist() + if any(name.startswith("ppt/comments/") for name in names): + risks.append("comments") + if any( + name.endswith((".bin", ".vbaProject")) + for name in names + ): + risks.append("binary_embedded_objects_or_macros") + if any(name.startswith("ppt/activeX/") for name in names): + risks.append("activex") + return risks + + +def _iter_shapes(shapes: Any) -> Iterable[Any]: + from pptx.enum.shapes import MSO_SHAPE_TYPE + + for shape in shapes: + yield shape + if shape.shape_type == MSO_SHAPE_TYPE.GROUP: + yield from _iter_shapes(shape.shapes) + + +def _iter_text_frames(slide: Any, *, include_notes: bool) -> Iterable[Any]: + for shape in _iter_shapes(slide.shapes): + if getattr(shape, "has_text_frame", False): + yield shape.text_frame + if getattr(shape, "has_table", False): + for row in shape.table.rows: + for cell in row.cells: + yield cell.text_frame + if include_notes: + try: + text_frame = slide.notes_slide.notes_text_frame + if text_frame is not None: + yield text_frame + except (AttributeError, KeyError, ValueError): + pass + + +def _find_spans( + text: str, + needle: str, + *, + match_case: bool, + whole_word: bool, + limit: Optional[int], +) -> list[tuple[int, int]]: + flags = 0 if match_case else re.IGNORECASE + escaped = re.escape(needle) + if whole_word: + escaped = rf"(? tuple[int, int]: + cursor = 0 + for index, run in enumerate(runs): + next_cursor = cursor + len(run.text) + if offset < next_cursor or (end and offset == next_cursor): + return index, offset - cursor + cursor = next_cursor + if not runs: + raise ValueError("段落没有可编辑的文本 Run") + return len(runs) - 1, len(runs[-1].text) + + +def _replace_in_paragraph( + paragraph: Any, + needle: str, + replacement: str, + *, + match_case: bool, + whole_word: bool, + limit: Optional[int], +) -> int: + runs = list(paragraph.runs) + if not runs: + return 0 + text = "".join(run.text for run in runs) + spans = _find_spans( + text, + needle, + match_case=match_case, + whole_word=whole_word, + limit=limit, + ) + for start, end in reversed(spans): + start_index, start_offset = _run_at_offset(runs, start) + end_index, end_offset = _run_at_offset(runs, end, end=True) + if start_index == end_index: + original = runs[start_index].text + runs[start_index].text = ( + original[:start_offset] + replacement + original[end_offset:] + ) + continue + prefix = runs[start_index].text[:start_offset] + suffix = runs[end_index].text[end_offset:] + runs[start_index].text = prefix + replacement + for index in range(start_index + 1, end_index): + runs[index].text = "" + runs[end_index].text = suffix + return len(spans) + + +def _selected_slides(presentation: Any, indexes: Optional[list[Any]]) -> list[Any]: + if indexes is None: + return list(presentation.slides) + if not isinstance(indexes, list) or not indexes: + raise ValueError("slides 必须是非空页码数组") + selected: list[Any] = [] + for value in indexes: + if not isinstance(value, int) or isinstance(value, bool): + raise ValueError("slides 中的页码必须是整数") + if value < 1 or value > len(presentation.slides): + raise ValueError(f"页码超出范围:{value}") + selected.append(presentation.slides[value - 1]) + return selected + + +def _replace_text(presentation: Any, operation: dict[str, Any]) -> dict[str, Any]: + needle = operation.get("find") + replacement = operation.get("replace") + if not isinstance(needle, str) or not needle: + raise ValueError("replace_text.find 必须是非空字符串") + if not isinstance(replacement, str): + raise ValueError("replace_text.replace 必须是字符串") + match_case = bool(operation.get("match_case", True)) + whole_word = bool(operation.get("whole_word", False)) + required = bool(operation.get("required", True)) + include_notes = bool(operation.get("include_notes", False)) + limit_value = operation.get("count") + if limit_value is not None: + if ( + not isinstance(limit_value, int) + or isinstance(limit_value, bool) + or limit_value < 1 + or limit_value > 10000 + ): + raise ValueError("replace_text.count 必须是 1 到 10000 的整数") + remaining = limit_value + changed = 0 + for slide in _selected_slides(presentation, operation.get("slides")): + for text_frame in _iter_text_frames(slide, include_notes=include_notes): + for paragraph in text_frame.paragraphs: + per_paragraph_limit = remaining + replacements = _replace_in_paragraph( + paragraph, + needle, + replacement, + match_case=match_case, + whole_word=whole_word, + limit=per_paragraph_limit, + ) + changed += replacements + if remaining is not None: + remaining -= replacements + if remaining <= 0: + break + if remaining is not None and remaining <= 0: + break + if remaining is not None and remaining <= 0: + break + if required and changed == 0: + raise ValueError(f"未找到必须替换的文本:{needle}") + return {"type": "replace_text", "replacement_count": changed} + + +def _set_properties(presentation: Any, operation: dict[str, Any]) -> dict[str, Any]: + properties = operation.get("properties") + if not isinstance(properties, dict): + raise ValueError("set_properties.properties 必须是对象") + allowed = { + "title", + "subject", + "author", + "keywords", + "comments", + "category", + "content_status", + "identifier", + "language", + "last_modified_by", + "revision", + "version", + } + changed: list[str] = [] + for key, value in properties.items(): + if key not in allowed: + raise ValueError(f"set_properties 不支持字段:{key}") + if key == "revision": + if ( + not isinstance(value, int) + or isinstance(value, bool) + or value < 1 + ): + raise ValueError("set_properties.revision 必须是正整数") + elif value is not None and not isinstance(value, str): + raise ValueError(f"set_properties.{key} 必须是字符串或 null") + setattr(presentation.core_properties, key, value) + changed.append(key) + return {"type": "set_properties", "fields": changed} + + +def _delete_slides(presentation: Any, operation: dict[str, Any]) -> dict[str, Any]: + indexes = operation.get("slides") + if not isinstance(indexes, list) or not indexes: + raise ValueError("delete_slides.slides 必须是非空页码数组") + normalized: set[int] = set() + for value in indexes: + if not isinstance(value, int) or isinstance(value, bool): + raise ValueError("delete_slides.slides 中的页码必须是整数") + if value < 1 or value > len(presentation.slides): + raise ValueError(f"要删除的页码超出范围:{value}") + normalized.add(value) + if len(normalized) >= len(presentation.slides): + raise ValueError("不能删除演示文稿中的全部页面") + slide_ids = presentation.slides._sldIdLst + for index in sorted(normalized, reverse=True): + slide_id = slide_ids[index - 1] + relationship_id = slide_id.rId + slide_ids.remove(slide_id) + presentation.part.drop_rel(relationship_id) + return {"type": "delete_slides", "deleted": sorted(normalized)} + + +def _reorder_slides(presentation: Any, operation: dict[str, Any]) -> dict[str, Any]: + order = operation.get("order") + expected = list(range(1, len(presentation.slides) + 1)) + if ( + not isinstance(order, list) + or any( + not isinstance(value, int) or isinstance(value, bool) + for value in order + ) + or sorted(order) != expected + ): + raise ValueError( + "reorder_slides.order 必须完整且不重复地列出当前全部页码" + ) + slide_ids = presentation.slides._sldIdLst + original = list(slide_ids) + for slide_id in original: + slide_ids.remove(slide_id) + for number in order: + slide_ids.append(original[number - 1]) + return {"type": "reorder_slides", "order": order} + + +def _apply_operation(presentation: Any, operation: dict[str, Any]) -> dict[str, Any]: + operation_type = operation.get("type") + if operation_type == "replace_text": + return _replace_text(presentation, operation) + if operation_type == "set_properties": + return _set_properties(presentation, operation) + if operation_type == "delete_slides": + return _delete_slides(presentation, operation) + if operation_type == "reorder_slides": + return _reorder_slides(presentation, operation) + raise ValueError(f"不支持的编辑操作:{operation_type}") + + +def main() -> dict[str, Any]: + from pptx import Presentation + + args = build_parser().parse_args() + source = input_file(args.input, OOXML_PRESENTATION_SUFFIXES) + destination = output_file(args.output, {".pptx"}, overwrite=args.overwrite) + spec = load_json_argument(args.spec, args.spec_file, label="编辑说明") + operations = spec.get("operations") + if not isinstance(operations, list) or not operations: + raise ValueError("编辑说明的 operations 必须是非空数组") + if len(operations) > 1000: + raise ValueError("编辑操作不能超过 1000 项") + + external_relationships = _external_relationship_count(source) + if external_relationships and not args.allow_external_links: + raise ValueError( + "输入演示文稿包含外部链接;如用户接受外部链接可能变化的风险," + "请显式传 --allow-external-links" + ) + risks = _package_risks(source) + if risks: + raise ValueError( + "输入演示文稿包含 python-pptx 不能可靠保留的内容:" + + "、".join(risks) + + ";请改用 unpack_presentation.py 做最小化 OOXML 编辑" + ) + + presentation = Presentation(str(source)) + results: list[dict[str, Any]] = [] + for index, operation in enumerate(operations): + if not isinstance(operation, dict): + raise ValueError(f"operations[{index}] 必须是对象") + results.append(_apply_operation(presentation, operation)) + + with tempfile.TemporaryDirectory(prefix="pptx-edit-") as temp_name: + staged = Path(temp_name) / "edited.pptx" + presentation.save(str(staged)) + reopened = Presentation(str(staged)) + slide_count = len(reopened.slides) + publish_file(staged, destination, overwrite=args.overwrite) + return { + "source": str(source), + "path": str(destination), + "slide_count": slide_count, + "external_relationship_count": external_relationships, + "operations": results, + } + + +if __name__ == "__main__": + raise SystemExit(run_cli(main)) diff --git a/skills/pptx/scripts/extract_presentation.py b/skills/pptx/scripts/extract_presentation.py new file mode 100644 index 0000000..abf0086 --- /dev/null +++ b/skills/pptx/scripts/extract_presentation.py @@ -0,0 +1,57 @@ +#!/usr/bin/env python3 + +from __future__ import annotations + +import argparse +import tempfile +from pathlib import Path +from typing import Any + +from _pptx_common import ( + OOXML_PRESENTATION_SUFFIXES, + SkillArgumentParser, + input_file, + output_file, + publish_file, + run_cli, +) + + +def build_parser() -> argparse.ArgumentParser: + parser = SkillArgumentParser(description="把演示文稿正文提取为 Markdown。") + parser.add_argument("--input", required=True) + parser.add_argument("--output", required=True) + parser.add_argument("--max-chars", type=int, default=2_000_000) + parser.add_argument("--overwrite", action="store_true") + return parser + + +def main() -> dict[str, Any]: + from markitdown import MarkItDown + + args = build_parser().parse_args() + if args.max_chars < 1000 or args.max_chars > 10_000_000: + raise ValueError("max-chars 必须在 1000 到 10000000 之间") + source = input_file(args.input, OOXML_PRESENTATION_SUFFIXES) + destination = output_file(args.output, {".md"}, overwrite=args.overwrite) + + result = MarkItDown(enable_plugins=False).convert(str(source)) + markdown = result.text_content or "" + truncated = len(markdown) > args.max_chars + if truncated: + markdown = markdown[: args.max_chars] + markdown += "\n\n\n" + with tempfile.TemporaryDirectory(prefix="pptx-extract-") as temp_name: + staged = Path(temp_name) / "presentation.md" + staged.write_text(markdown, encoding="utf-8") + publish_file(staged, destination, overwrite=args.overwrite) + return { + "source": str(source), + "path": str(destination), + "characters": len(markdown), + "truncated": truncated, + } + + +if __name__ == "__main__": + raise SystemExit(run_cli(main)) diff --git a/skills/pptx/scripts/inspect_presentation.py b/skills/pptx/scripts/inspect_presentation.py new file mode 100644 index 0000000..044b242 --- /dev/null +++ b/skills/pptx/scripts/inspect_presentation.py @@ -0,0 +1,404 @@ +#!/usr/bin/env python3 + +from __future__ import annotations + +import argparse +import zipfile +from pathlib import Path +from typing import Any, Optional + +from _pptx_common import ( + OOXML_PRESENTATION_SUFFIXES, + SkillArgumentParser, + input_file, + inspect_archive, + parse_xml_bytes, + run_cli, +) + + +EMU_PER_INCH = 914400 + + +def build_parser() -> argparse.ArgumentParser: + parser = SkillArgumentParser(description="分段检查演示文稿的结构、文本和媒体。") + parser.add_argument("--input", required=True) + parser.add_argument("--start-slide", type=int, default=1) + parser.add_argument("--max-slides", type=int, default=30) + parser.add_argument("--max-shapes", type=int, default=200) + parser.add_argument("--max-table-cells", type=int, default=500) + parser.add_argument("--max-chars", type=int, default=100000) + parser.add_argument("--include-runs", action="store_true") + return parser + + +def _inches(value: Any) -> Optional[float]: + try: + return round(int(value) / EMU_PER_INCH, 4) + except (TypeError, ValueError): + return None + + +def _iter_shapes(shapes: Any): + from pptx.enum.shapes import MSO_SHAPE_TYPE + + for shape in shapes: + yield shape + if shape.shape_type == MSO_SHAPE_TYPE.GROUP: + yield from _iter_shapes(shape.shapes) + + +def _shape_text_char_count(shape: Any) -> int: + texts: list[str] = [] + if getattr(shape, "has_text_frame", False): + texts.append(shape.text_frame.text) + if getattr(shape, "has_table", False): + for row in shape.table.rows: + texts.extend(cell.text for cell in row.cells) + return sum( + 1 + for text in texts + for character in text + if character.isalnum() + ) + + +def _contains_picture(shape: Any) -> bool: + from pptx.enum.shapes import MSO_SHAPE_TYPE + + if shape.shape_type == MSO_SHAPE_TYPE.PICTURE: + return True + if shape.shape_type == MSO_SHAPE_TYPE.GROUP: + return any(_contains_picture(child) for child in shape.shapes) + return False + + +def _slide_media_profile( + slide: Any, + slide_width: int, + slide_height: int, +) -> dict[str, Any]: + from pptx.enum.shapes import MSO_SHAPE_TYPE + + shapes = list(_iter_shapes(slide.shapes)) + image_count = sum( + 1 for shape in shapes + if shape.shape_type == MSO_SHAPE_TYPE.PICTURE + ) + chart_count = sum( + 1 for shape in shapes + if getattr(shape, "has_chart", False) + ) + native_text_char_count = sum( + _shape_text_char_count(shape) for shape in shapes + ) + image_area = 0.0 + slide_area = slide_width * slide_height + if slide_area > 0: + for shape in slide.shapes: + if not _contains_picture(shape): + continue + try: + raw_left = int(shape.left) + raw_top = int(shape.top) + left = max(0, raw_left) + top = max(0, raw_top) + right = min(slide_width, raw_left + int(shape.width)) + bottom = min(slide_height, raw_top + int(shape.height)) + except (AttributeError, TypeError, ValueError): + continue + if right > left and bottom > top: + image_area += (right - left) * (bottom - top) / slide_area + return { + "image_count": image_count, + "chart_count": chart_count, + "image_area_ratio": round(min(1.0, image_area), 4), + "native_text_char_count": native_text_char_count, + "has_images": image_count > 0, + } + + +def _run_payload(run: Any) -> dict[str, Any]: + font = run.font + hyperlink = None + try: + hyperlink = run.hyperlink.address + except (AttributeError, KeyError, ValueError): + pass + return { + "text": run.text, + "bold": font.bold, + "italic": font.italic, + "underline": font.underline, + "font_name": font.name, + "font_size_pt": round(font.size.pt, 2) if font.size is not None else None, + "hyperlink": hyperlink, + } + + +def _paragraph_payload(paragraph: Any, *, include_runs: bool) -> dict[str, Any]: + payload: dict[str, Any] = { + "text": paragraph.text, + "level": paragraph.level, + "alignment": str(paragraph.alignment) if paragraph.alignment else None, + } + if include_runs: + payload["runs"] = [_run_payload(run) for run in paragraph.runs] + return payload + + +def _text_frame_payload(text_frame: Any, *, include_runs: bool) -> dict[str, Any]: + return { + "text": text_frame.text, + "paragraphs": [ + _paragraph_payload(paragraph, include_runs=include_runs) + for paragraph in text_frame.paragraphs + ], + } + + +def _chart_payload(shape: Any, *, max_points: int = 100) -> dict[str, Any]: + chart = shape.chart + series_payload: list[dict[str, Any]] = [] + for series in list(chart.series)[:50]: + values: list[Any] = [] + try: + values = list(series.values)[:max_points] + except (AttributeError, TypeError, ValueError): + pass + series_payload.append( + { + "name": getattr(series, "name", None), + "point_count": len(values), + "values": values, + } + ) + categories: list[str] = [] + try: + categories = [str(item.label) for item in chart.plots[0].categories][:max_points] + except (AttributeError, IndexError, TypeError, ValueError): + pass + return { + "chart_type": str(chart.chart_type), + "has_title": chart.has_title, + "series": series_payload, + "categories": categories, + } + + +def _shape_payload( + shape: Any, + *, + include_runs: bool, + max_table_cells: int, +) -> dict[str, Any]: + payload: dict[str, Any] = { + "name": shape.name, + "shape_type": str(shape.shape_type), + "x": _inches(shape.left), + "y": _inches(shape.top), + "w": _inches(shape.width), + "h": _inches(shape.height), + } + if getattr(shape, "has_text_frame", False): + payload["text_frame"] = _text_frame_payload( + shape.text_frame, + include_runs=include_runs, + ) + if getattr(shape, "has_table", False): + rows: list[list[str]] = [] + cell_count = 0 + truncated = False + for row in shape.table.rows: + values: list[str] = [] + for cell in row.cells: + if cell_count >= max_table_cells: + truncated = True + break + values.append(cell.text) + cell_count += 1 + if values: + rows.append(values) + if truncated: + break + payload["table"] = { + "row_count": len(shape.table.rows), + "column_count": len(shape.table.columns), + "rows": rows, + "truncated": truncated, + } + if getattr(shape, "has_chart", False): + payload["chart"] = _chart_payload(shape) + if shape.shape_type == 13: + try: + payload["image"] = { + "content_type": shape.image.content_type, + "filename": shape.image.filename, + "size_bytes": len(shape.image.blob), + } + except (AttributeError, KeyError, ValueError): + payload["image"] = {"readable": False} + return payload + + +def _slide_notes(slide: Any) -> str: + try: + text_frame = slide.notes_slide.notes_text_frame + return text_frame.text if text_frame is not None else "" + except (AttributeError, KeyError, ValueError): + return "" + + +def _comment_summary(path: Path) -> dict[str, Any]: + comment_parts: list[str] = [] + authors: set[str] = set() + with zipfile.ZipFile(path) as archive: + for name in archive.namelist(): + if name.startswith("ppt/comments/comment") and name.endswith(".xml"): + comment_parts.append(name) + if name in { + "ppt/commentAuthors.xml", + "ppt/authors.xml", + }: + root = parse_xml_bytes(archive.read(name), label=name) + for node in root.iter(): + author = node.attrib.get("name") + if author: + authors.add(author) + return { + "part_count": len(comment_parts), + "authors": sorted(authors), + } + + +def _external_relationships(path: Path) -> list[dict[str, str]]: + relationships: list[dict[str, str]] = [] + with zipfile.ZipFile(path) as archive: + for name in archive.namelist(): + if not name.endswith(".rels"): + continue + root = parse_xml_bytes(archive.read(name), label=name) + for node in root: + if node.attrib.get("TargetMode") != "External": + continue + relationships.append( + { + "part": name, + "type": node.attrib.get("Type", ""), + "target": node.attrib.get("Target", ""), + } + ) + return relationships[:200] + + +def main() -> dict[str, Any]: + from pptx import Presentation + + args = build_parser().parse_args() + if args.start_slide < 1: + raise ValueError("start-slide 必须大于 0") + if args.max_slides < 1 or args.max_slides > 100: + raise ValueError("max-slides 必须在 1 到 100 之间") + if args.max_shapes < 1 or args.max_shapes > 1000: + raise ValueError("max-shapes 必须在 1 到 1000 之间") + if args.max_table_cells < 1 or args.max_table_cells > 10000: + raise ValueError("max-table-cells 必须在 1 到 10000 之间") + if args.max_chars < 1000 or args.max_chars > 1_000_000: + raise ValueError("max-chars 必须在 1000 到 1000000 之间") + + source = input_file(args.input, OOXML_PRESENTATION_SUFFIXES) + archive = inspect_archive(source) + if archive["missing_required_parts"]: + raise ValueError( + "演示文稿缺少必要部件:" + + "、".join(archive["missing_required_parts"]) + ) + presentation = Presentation(str(source)) + slide_count = len(presentation.slides) + if args.start_slide > slide_count and slide_count > 0: + raise ValueError(f"start-slide 超出页面总数 {slide_count}") + + end_slide = min(slide_count, args.start_slide + args.max_slides - 1) + payload_slides: list[dict[str, Any]] = [] + char_count = 0 + truncated_by_chars = False + for slide_number in range(args.start_slide, end_slide + 1): + slide = presentation.slides[slide_number - 1] + title = slide.shapes.title.text if slide.shapes.title is not None else "" + media_profile = _slide_media_profile( + slide, + int(presentation.slide_width), + int(presentation.slide_height), + ) + shapes: list[dict[str, Any]] = [] + for shape in list(slide.shapes)[: args.max_shapes]: + item = _shape_payload( + shape, + include_runs=args.include_runs, + max_table_cells=args.max_table_cells, + ) + item_chars = len(str(item)) + if char_count + item_chars > args.max_chars: + truncated_by_chars = True + break + shapes.append(item) + char_count += item_chars + notes = _slide_notes(slide) + if char_count + len(notes) > args.max_chars: + notes = notes[: max(0, args.max_chars - char_count)] + truncated_by_chars = True + char_count += len(notes) + payload_slides.append( + { + "number": slide_number, + "title": title, + "layout": getattr(slide.slide_layout, "name", None), + "shape_count": len(slide.shapes), + "shapes_truncated": len(slide.shapes) > args.max_shapes, + "shapes": shapes, + "speaker_notes": notes, + "media": media_profile, + } + ) + if truncated_by_chars: + break + + last_slide = payload_slides[-1]["number"] if payload_slides else args.start_slide - 1 + next_slide = last_slide + 1 if last_slide < slide_count else None + properties = presentation.core_properties + return { + "source": str(source), + "slide_count": slide_count, + "slide_size": { + "width_inches": _inches(presentation.slide_width), + "height_inches": _inches(presentation.slide_height), + }, + "properties": { + "title": properties.title, + "subject": properties.subject, + "author": properties.author, + "keywords": properties.keywords, + "comments": properties.comments, + "last_modified_by": properties.last_modified_by, + }, + "selection": { + "start_slide": args.start_slide, + "end_slide": last_slide, + "has_more": next_slide is not None, + "next_slide": next_slide, + "truncated_by_chars": truncated_by_chars, + }, + "slides": payload_slides, + "image_slides": [ + slide["number"] + for slide in payload_slides + if slide["media"]["has_images"] + ], + "comments": _comment_summary(source), + "external_relationships": _external_relationships(source), + "archive": archive, + } + + +if __name__ == "__main__": + raise SystemExit(run_cli(main)) diff --git a/skills/pptx/scripts/ocr_presentation.py b/skills/pptx/scripts/ocr_presentation.py new file mode 100644 index 0000000..d7bf74b --- /dev/null +++ b/skills/pptx/scripts/ocr_presentation.py @@ -0,0 +1,727 @@ +#!/usr/bin/env python3 + +from __future__ import annotations + +import contextlib +import difflib +import importlib.metadata +import io +import logging +import os +import re +import shutil +import tempfile +import time +import unicodedata +from pathlib import Path +from typing import Any, Iterable + +from _pptx_common import ( + OOXML_PRESENTATION_SUFFIXES, + SkillArgumentParser, + find_program, + input_file, + run_cli, + run_program, + run_soffice_convert, +) + + +DEFAULT_DPI = 260 +DEFAULT_MAX_CHARS = 24000 +DEFAULT_TIMEOUT_SECONDS = 180 +MAX_SLIDES_PER_CALL = 4 +MAX_PIXELS_PER_SLIDE = 20_000_000 +MIN_MEAN_CONFIDENCE = 0.60 +MIN_MEANINGFUL_CHARS = 5 +WHITESPACE_PATTERN = re.compile(r"[ \t]+") + + +for variable, value in ( + ("OMP_NUM_THREADS", "2"), + ("OPENBLAS_NUM_THREADS", "1"), + ("MKL_NUM_THREADS", "1"), + ("NUMEXPR_NUM_THREADS", "1"), +): + os.environ.setdefault(variable, value) + +for logger_name in ("rapidocr", "RapidOCR", "onnxruntime"): + logging.getLogger(logger_name).setLevel(logging.ERROR) + + +def build_parser(): + parser = SkillArgumentParser( + description="渲染指定演示文稿页面并用本地 OCR 提取图片文字。" + ) + parser.add_argument("--input", required=True) + parser.add_argument( + "--slides", + required=True, + help="要识别的页码,例如 2 或 2,5-6;单次最多 4 页", + ) + parser.add_argument( + "--start-offset", + type=int, + default=0, + help="续读单页图片文字时的字符偏移量", + ) + parser.add_argument( + "--max-chars", + type=int, + default=DEFAULT_MAX_CHARS, + help=f"单次最多返回字符数,默认 {DEFAULT_MAX_CHARS}", + ) + parser.add_argument( + "--dpi", + type=int, + default=DEFAULT_DPI, + help=f"OCR 渲染分辨率,默认 {DEFAULT_DPI} DPI", + ) + parser.add_argument( + "--timeout", + type=int, + default=DEFAULT_TIMEOUT_SECONDS, + help=f"转换和单页渲染超时秒数,默认 {DEFAULT_TIMEOUT_SECONDS}", + ) + return parser + + +def _parse_slide_spec(value: str, slide_count: int) -> list[int]: + if not value.strip(): + raise ValueError("slides 不能为空") + slides: set[int] = set() + for raw_part in value.split(","): + part = raw_part.strip() + if not part: + continue + if "-" in part: + pieces = part.split("-", 1) + try: + start = int(pieces[0]) + end = int(pieces[1]) + except ValueError as exc: + raise ValueError(f"页码范围格式错误:{part}") from exc + if start > end: + raise ValueError(f"页码范围起始值不能大于结束值:{part}") + else: + try: + start = end = int(part) + except ValueError as exc: + raise ValueError(f"页码格式错误:{part}") from exc + if start < 1 or end > slide_count: + raise ValueError(f"页码必须在 1 到 {slide_count} 之间:{part}") + slides.update(range(start, end + 1)) + if not slides: + raise ValueError("slides 不能为空") + return sorted(slides) + + +def _clean_text(value: Any) -> str: + text = str(value or "").replace("\x00", "").strip() + return "\n".join( + WHITESPACE_PATTERN.sub(" ", line).strip() + for line in text.replace("\r\n", "\n").replace("\r", "\n").split("\n") + if line.strip() + ) + + +def _comparison_key(value: str) -> str: + normalized = unicodedata.normalize("NFKC", value).casefold() + return "".join(character for character in normalized if character.isalnum()) + + +def _iter_shapes(shapes: Any) -> Iterable[Any]: + from pptx.enum.shapes import MSO_SHAPE_TYPE + + for shape in shapes: + yield shape + if shape.shape_type == MSO_SHAPE_TYPE.GROUP: + yield from _iter_shapes(shape.shapes) + + +def _shape_text_fragments(shape: Any) -> list[str]: + fragments: list[str] = [] + if getattr(shape, "has_text_frame", False): + fragments.extend( + line + for line in _clean_text(shape.text_frame.text).splitlines() + if line + ) + if getattr(shape, "has_table", False): + for row in shape.table.rows: + for cell in row.cells: + fragments.extend( + line + for line in _clean_text(cell.text).splitlines() + if line + ) + return fragments + + +def _normalized_box( + shape: Any, + slide_width: int, + slide_height: int, +) -> tuple[float, float, float, float] | None: + try: + left = float(shape.left) + top = float(shape.top) + right = left + float(shape.width) + bottom = top + float(shape.height) + except (AttributeError, TypeError, ValueError): + return None + if slide_width <= 0 or slide_height <= 0: + return None + x0 = max(0.0, min(1.0, left / slide_width)) + y0 = max(0.0, min(1.0, top / slide_height)) + x1 = max(0.0, min(1.0, right / slide_width)) + y1 = max(0.0, min(1.0, bottom / slide_height)) + if x1 <= x0 or y1 <= y0: + return None + return (x0, y0, x1, y1) + + +def _contains_picture(shape: Any) -> bool: + from pptx.enum.shapes import MSO_SHAPE_TYPE + + if shape.shape_type == MSO_SHAPE_TYPE.PICTURE: + return True + if shape.shape_type == MSO_SHAPE_TYPE.GROUP: + return any(_contains_picture(child) for child in shape.shapes) + return False + + +def _slide_profile( + slide: Any, + slide_width: int, + slide_height: int, +) -> dict[str, Any]: + from pptx.enum.shapes import MSO_SHAPE_TYPE + + all_shapes = list(_iter_shapes(slide.shapes)) + native_fragments: list[str] = [] + picture_count = 0 + chart_count = 0 + for shape in all_shapes: + native_fragments.extend(_shape_text_fragments(shape)) + if shape.shape_type == MSO_SHAPE_TYPE.PICTURE: + picture_count += 1 + if getattr(shape, "has_chart", False): + chart_count += 1 + + unique_fragments = list(dict.fromkeys(native_fragments)) + native_keys = [ + key + for fragment in unique_fragments + if (key := _comparison_key(fragment)) + ] + native_boxes: list[dict[str, Any]] = [] + picture_area = 0.0 + for shape in slide.shapes: + box = _normalized_box(shape, slide_width, slide_height) + shape_fragments = _shape_text_fragments(shape) + if box and shape_fragments: + native_boxes.append( + { + "box": box, + "keys": [ + key + for fragment in shape_fragments + if (key := _comparison_key(fragment)) + ], + } + ) + if box and _contains_picture(shape): + picture_area += (box[2] - box[0]) * (box[3] - box[1]) + + return { + "picture_count": picture_count, + "chart_count": chart_count, + "image_area_ratio": round(min(1.0, picture_area), 4), + "native_text_char_count": sum( + 1 + for fragment in unique_fragments + for character in fragment + if character.isalnum() + ), + "_native_keys": native_keys, + "_native_boxes": native_boxes, + } + + +def _box_points(value: Any) -> list[list[float]] | None: + if value is None: + return None + try: + points = [ + [round(float(point[0]), 2), round(float(point[1]), 2)] + for point in value + ] + except (IndexError, TypeError, ValueError): + return None + return points if len(points) == 4 else None + + +def _ordered_lines(result: Any) -> list[dict[str, Any]]: + texts = list(getattr(result, "txts", None) or ()) + scores = list(getattr(result, "scores", None) or ()) + raw_boxes = getattr(result, "boxes", None) + boxes = list(raw_boxes) if raw_boxes is not None else [] + + lines: list[dict[str, Any]] = [] + for index, raw_text in enumerate(texts): + text = _clean_text(raw_text) + if not text: + continue + try: + confidence = float(scores[index]) + except (IndexError, TypeError, ValueError): + confidence = 0.0 + confidence = max(0.0, min(1.0, confidence)) + box = _box_points(boxes[index] if index < len(boxes) else None) + if box: + left = min(point[0] for point in box) + top = min(point[1] for point in box) + else: + left = float(index) + top = float(index) + lines.append( + { + "text": text, + "confidence": confidence, + "box": box, + "_left": left, + "_top": top, + "_index": index, + } + ) + + lines.sort( + key=lambda line: ( + round(line["_top"] / 10.0), + line["_left"], + line["_index"], + ) + ) + return lines + + +def _similar_to_any( + candidate: str, + references: list[str], + *, + threshold: float, +) -> bool: + if not candidate: + return False + for reference in references: + if not reference: + continue + if candidate == reference: + return True + shorter = min(len(candidate), len(reference)) + longer = max(len(candidate), len(reference)) + if shorter >= 3 and candidate in reference: + return True + if ( + shorter >= 3 + and reference in candidate + and longer <= round(shorter * 1.25) + ): + return True + if shorter >= 3 and difflib.SequenceMatcher( + None, + candidate, + reference, + ).ratio() >= threshold: + return True + return False + + +def _line_center( + box: list[list[float]] | None, + image_width: int, + image_height: int, +) -> tuple[float, float] | None: + if not box or image_width <= 0 or image_height <= 0: + return None + return ( + sum(point[0] for point in box) / len(box) / image_width, + sum(point[1] for point in box) / len(box) / image_height, + ) + + +def _line_is_native( + line: dict[str, Any], + profile: dict[str, Any], + image_width: int, + image_height: int, +) -> bool: + candidate = _comparison_key(line["text"]) + if _similar_to_any( + candidate, + profile["_native_keys"], + threshold=0.82, + ): + return True + + center = _line_center(line["box"], image_width, image_height) + if center is None: + return False + x, y = center + padding = 0.01 + for native_box in profile["_native_boxes"]: + x0, y0, x1, y1 = native_box["box"] + if ( + x0 - padding <= x <= x1 + padding + and y0 - padding <= y <= y1 + padding + and _similar_to_any( + candidate, + native_box["keys"], + threshold=0.68, + ) + ): + return True + return False + + +def _create_ocr_engine(): + try: + from rapidocr import RapidOCR + except ImportError as exc: + raise RuntimeError("环境预置的 rapidocr 模块不可用") from exc + + captured_stdout = io.StringIO() + captured_stderr = io.StringIO() + with ( + contextlib.redirect_stdout(captured_stdout), + contextlib.redirect_stderr(captured_stderr), + ): + return RapidOCR() + + +def _ocr_slide( + engine: Any, + image_path: Path, + profile: dict[str, Any], +) -> dict[str, Any]: + from PIL import Image + + with Image.open(image_path) as image: + image_width, image_height = image.size + + captured_stdout = io.StringIO() + captured_stderr = io.StringIO() + started = time.monotonic() + with ( + contextlib.redirect_stdout(captured_stdout), + contextlib.redirect_stderr(captured_stderr), + ): + result = engine(str(image_path)) + elapsed = time.monotonic() - started + + raw_lines = _ordered_lines(result) + image_lines: list[dict[str, Any]] = [] + seen: set[str] = set() + filtered_native = 0 + filtered_duplicates = 0 + for line in raw_lines: + if _line_is_native(line, profile, image_width, image_height): + filtered_native += 1 + continue + key = _comparison_key(line["text"]) + if key and key in seen: + filtered_duplicates += 1 + continue + if key: + seen.add(key) + image_lines.append(line) + + text = "\n".join(line["text"] for line in image_lines) + weighted_chars = [ + max(1, sum(1 for character in line["text"] if not character.isspace())) + for line in image_lines + ] + total_weight = sum(weighted_chars) + mean_confidence = ( + sum( + line["confidence"] * weight + for line, weight in zip(image_lines, weighted_chars) + ) + / total_weight + if total_weight + else 0.0 + ) + meaningful_chars = sum(1 for character in text if character.isalnum()) + low_confidence_lines = sum( + 1 + for line in image_lines + if line["confidence"] < MIN_MEAN_CONFIDENCE + ) + + reasons: list[str] = [] + if not text: + status = "no_image_text" + reasons.append("未识别到原生文本之外的图片文字") + elif meaningful_chars < MIN_MEANINGFUL_CHARS: + status = "sparse" + reasons.append( + f"图片中的有效文字少于 {MIN_MEANINGFUL_CHARS} 个字符" + ) + elif mean_confidence < MIN_MEAN_CONFIDENCE: + status = "low_confidence" + reasons.append( + "图片文字 OCR 平均置信度低于 " + f"{round(MIN_MEAN_CONFIDENCE * 100)}%" + ) + else: + status = "good" + + return { + "text": text, + "status": status, + "usable_for_summary": status == "good", + "needs_review": status in {"sparse", "low_confidence"}, + "raw_ocr_line_count": len(raw_lines), + "image_line_count": len(image_lines), + "filtered_native_line_count": filtered_native, + "filtered_duplicate_line_count": filtered_duplicates, + "low_confidence_line_count": low_confidence_lines, + "mean_confidence": round(mean_confidence, 4), + "meaningful_chars": meaningful_chars, + "reasons": reasons, + "ocr_seconds": round(elapsed, 3), + } + + +def _pdf_pages(path: Path) -> tuple[int, dict[int, tuple[float, float]]]: + from pypdf import PdfReader + + page_sizes: dict[int, tuple[float, float]] = {} + with path.open("rb") as stream: + reader = PdfReader(stream, strict=False) + if reader.is_encrypted: + raise ValueError("LibreOffice 生成了加密 PDF,无法执行 OCR") + page_count = len(reader.pages) + for page_number, page in enumerate(reader.pages, start=1): + page_sizes[page_number] = ( + abs(float(page.cropbox.width)), + abs(float(page.cropbox.height)), + ) + return page_count, page_sizes + + +def _render_slide( + pdf_path: Path, + slide_number: int, + page_size: tuple[float, float], + dpi: int, + timeout: int, + temp_dir: Path, +) -> tuple[Path, float]: + width_points, height_points = page_size + estimated_pixels = ( + width_points * dpi / 72.0 + * height_points * dpi / 72.0 + ) + if estimated_pixels > MAX_PIXELS_PER_SLIDE: + raise ValueError( + f"第 {slide_number} 页按 {dpi} DPI 渲染预计超过 " + f"{MAX_PIXELS_PER_SLIDE} 像素,请降低 dpi" + ) + + prefix = temp_dir / f"slide-{slide_number:04d}" + output = prefix.with_suffix(".png") + started = time.monotonic() + run_program( + [ + find_program("pdftoppm"), + "-f", + str(slide_number), + "-l", + str(slide_number), + "-singlefile", + "-png", + "-r", + str(dpi), + str(pdf_path), + str(prefix), + ], + timeout=timeout, + ) + elapsed = time.monotonic() - started + if not output.is_file() or output.stat().st_size <= 0: + raise RuntimeError(f"第 {slide_number} 页没有生成有效 PNG") + return output, elapsed + + +def _package_version(name: str) -> str | None: + try: + return importlib.metadata.version(name) + except importlib.metadata.PackageNotFoundError: + return None + + +def main() -> dict[str, Any]: + from pptx import Presentation + + args = build_parser().parse_args() + if args.start_offset < 0: + raise ValueError("start-offset 不能小于 0") + if args.max_chars < 1 or args.max_chars > 60000: + raise ValueError("max-chars 必须在 1 到 60000 之间") + if args.dpi < 150 or args.dpi > 400: + raise ValueError("dpi 必须在 150 到 400 之间") + if args.timeout < 1 or args.timeout > 600: + raise ValueError("timeout 必须在 1 到 600 秒之间") + + source = input_file(args.input, OOXML_PRESENTATION_SUFFIXES) + presentation = Presentation(str(source)) + slide_count = len(presentation.slides) + if slide_count < 1: + raise ValueError("演示文稿没有可执行 OCR 的页面") + requested_slides = _parse_slide_spec(args.slides, slide_count) + if len(requested_slides) > MAX_SLIDES_PER_CALL: + raise ValueError( + f"单次最多 OCR {MAX_SLIDES_PER_CALL} 页,请拆分 slides 后重试" + ) + if args.start_offset > 0 and len(requested_slides) != 1: + raise ValueError("使用 start-offset 时 slides 必须只包含一页") + + slide_width = int(presentation.slide_width) + slide_height = int(presentation.slide_height) + profiles = { + slide_number: _slide_profile( + presentation.slides[slide_number - 1], + slide_width, + slide_height, + ) + for slide_number in requested_slides + } + + engine = _create_ocr_engine() + page_outputs: list[dict[str, Any]] = [] + returned_chars = 0 + next_slide: int | None = None + next_offset = 0 + remaining_slides: list[int] = [] + office_output = {"stdout": "", "stderr": ""} + + with tempfile.TemporaryDirectory(prefix="pptx-ocr-") as temp_name: + temp_dir = Path(temp_name) + staged_input = temp_dir / f"presentation{source.suffix.lower()}" + shutil.copy2(source, staged_input) + pdf_path, office_output = run_soffice_convert( + staged_input, + target_format="pdf", + output_dir=temp_dir / "pdf", + timeout=args.timeout, + ) + pdf_page_count, page_sizes = _pdf_pages(pdf_path) + if pdf_page_count != slide_count: + raise RuntimeError( + f"演示文稿有 {slide_count} 页,但渲染结果有 " + f"{pdf_page_count} 页" + ) + + for index, slide_number in enumerate(requested_slides): + budget = args.max_chars - returned_chars + if budget <= 0: + next_slide = slide_number + remaining_slides = requested_slides[index:] + break + + image_path, render_seconds = _render_slide( + pdf_path, + slide_number, + page_sizes[slide_number], + args.dpi, + args.timeout, + temp_dir, + ) + result = _ocr_slide( + engine, + image_path, + profiles[slide_number], + ) + full_text = result.pop("text") + offset = args.start_offset if index == 0 else 0 + if offset > len(full_text): + raise ValueError( + f"start-offset 超过第 {slide_number} 页图片文字长度 " + f"{len(full_text)}" + ) + + usable = bool(result["usable_for_summary"]) + if not usable: + slide_text = "" + complete = True + else: + remaining_text = full_text[offset:] + slide_text = remaining_text[:budget] + complete = len(slide_text) == len(remaining_text) + + profile = profiles[slide_number] + page_outputs.append( + { + "slide": slide_number, + "text": slide_text, + "char_count": len(full_text), + "offset_start": offset if usable else 0, + "offset_end": offset + len(slide_text) if usable else 0, + "complete": complete, + "render_seconds": round(render_seconds, 3), + "picture_count": profile["picture_count"], + "chart_count": profile["chart_count"], + "image_area_ratio": profile["image_area_ratio"], + "native_text_char_count": profile[ + "native_text_char_count" + ], + **result, + } + ) + returned_chars += len(slide_text) + + if not complete: + next_slide = slide_number + next_offset = offset + len(slide_text) + remaining_slides = requested_slides[index + 1 :] + break + + all_processed = len(page_outputs) == len(requested_slides) + all_complete = all(page["complete"] for page in page_outputs) + all_safe = all( + page["status"] in {"good", "no_image_text"} + for page in page_outputs + ) + return { + "source": str(source), + "slide_count": slide_count, + "engine": "rapidocr", + "engine_version": _package_version("rapidocr"), + "runtime": "onnxruntime", + "runtime_version": _package_version("onnxruntime"), + "offline": True, + "dpi": args.dpi, + "requested_slides": requested_slides, + "processed_slides": [page["slide"] for page in page_outputs], + "returned_chars": returned_chars, + "slides": page_outputs, + "usable_for_summary": any( + page["usable_for_summary"] for page in page_outputs + ), + "complete_ocr_coverage": ( + all_processed and all_complete and all_safe + ), + "needs_review": any(page["needs_review"] for page in page_outputs), + "has_more": next_slide is not None, + "next_slide": next_slide, + "next_offset": next_offset, + "remaining_slides": remaining_slides, + "office_stdout": office_output["stdout"], + "office_stderr": office_output["stderr"], + } + + +if __name__ == "__main__": + raise SystemExit(run_cli(main)) diff --git a/skills/pptx/scripts/pack_presentation.py b/skills/pptx/scripts/pack_presentation.py new file mode 100644 index 0000000..a2f4bb6 --- /dev/null +++ b/skills/pptx/scripts/pack_presentation.py @@ -0,0 +1,55 @@ +#!/usr/bin/env python3 + +from __future__ import annotations + +import argparse +import tempfile +from pathlib import Path +from typing import Any + +from _pptx_common import ( + SkillArgumentParser, + output_file, + pack_presentation_directory, + publish_file, + run_cli, +) + + +def build_parser() -> argparse.ArgumentParser: + parser = SkillArgumentParser(description="把 OOXML 目录安全打包为 PowerPoint 文件。") + parser.add_argument("--input-dir", required=True) + parser.add_argument("--output", required=True) + parser.add_argument("--overwrite", action="store_true") + return parser + + +def main() -> dict[str, Any]: + from pptx import Presentation + + args = build_parser().parse_args() + source_dir = Path(args.input_dir).expanduser().resolve() + destination = output_file( + args.output, + {".pptx", ".potx", ".ppsx"}, + overwrite=args.overwrite, + ) + with tempfile.TemporaryDirectory(prefix="pptx-pack-") as temp_name: + staged = Path(temp_name) / destination.name + archive = pack_presentation_directory(source_dir, staged) + try: + presentation = Presentation(str(staged)) + slide_count = len(presentation.slides) + except Exception as exc: + raise ValueError("打包结果不能被 python-pptx 打开") from exc + publish_file(staged, destination, overwrite=args.overwrite) + return { + "source_dir": str(source_dir), + "path": str(destination), + "slide_count": slide_count, + "archive": archive, + } + + +if __name__ == "__main__": + raise SystemExit(run_cli(main)) diff --git a/skills/pptx/scripts/render_icon.py b/skills/pptx/scripts/render_icon.py new file mode 100644 index 0000000..10781ba --- /dev/null +++ b/skills/pptx/scripts/render_icon.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 + +from __future__ import annotations + +import argparse +import json +import os +import tempfile +from pathlib import Path +from typing import Any + +from _pptx_common import ( + SkillArgumentParser, + find_program, + output_file, + publish_file, + run_cli, + run_program, +) + + +def build_parser() -> argparse.ArgumentParser: + parser = SkillArgumentParser(description="把 React Icons 图标渲染为透明 PNG。") + parser.add_argument("--library", required=True) + parser.add_argument("--name", required=True) + parser.add_argument("--output", required=True) + parser.add_argument("--color", default="111827") + parser.add_argument("--background") + parser.add_argument("--size", type=int, default=256) + parser.add_argument("--title") + parser.add_argument("--timeout", type=int, default=60) + parser.add_argument("--overwrite", action="store_true") + return parser + + +def main() -> dict[str, Any]: + args = build_parser().parse_args() + destination = output_file(args.output, {".png"}, overwrite=args.overwrite) + renderer = Path(__file__).with_name("_icon_renderer.js").resolve() + if not renderer.is_file(): + raise FileNotFoundError(f"内部图标渲染器不存在:{renderer}") + spec = { + "library": args.library, + "name": args.name, + "color": args.color, + "background": args.background, + "size": args.size, + "title": args.title, + } + with tempfile.TemporaryDirectory(prefix="pptx-icon-") as temp_name: + temp_dir = Path(temp_name) + spec_path = temp_dir / "spec.json" + staged = temp_dir / "icon.png" + spec_path.write_text(json.dumps(spec, ensure_ascii=False), encoding="utf-8") + completed = run_program( + [ + find_program("node"), + str(renderer), + "--spec", + str(spec_path), + "--output", + str(staged), + ], + timeout=args.timeout, + cwd=Path.cwd(), + env=os.environ.copy(), + ) + try: + result = json.loads(completed.stdout.strip()) + except json.JSONDecodeError as exc: + raise RuntimeError("内部图标渲染器返回了无效结果") from exc + publish_file(staged, destination, overwrite=args.overwrite) + return { + "path": str(destination), + "size_bytes": destination.stat().st_size, + **result, + } + + +if __name__ == "__main__": + raise SystemExit(run_cli(main)) diff --git a/skills/pptx/scripts/render_presentation.py b/skills/pptx/scripts/render_presentation.py new file mode 100644 index 0000000..f15cb70 --- /dev/null +++ b/skills/pptx/scripts/render_presentation.py @@ -0,0 +1,224 @@ +#!/usr/bin/env python3 + +from __future__ import annotations + +import argparse +import math +import shutil +import tempfile +from pathlib import Path +from typing import Any, Optional + +from _pptx_common import ( + OOXML_PRESENTATION_SUFFIXES, + SkillArgumentParser, + find_program, + input_file, + output_directory, + publish_file, + run_cli, + run_program, + run_soffice_convert, +) + + +def build_parser() -> argparse.ArgumentParser: + parser = SkillArgumentParser( + description="通过 LibreOffice 和 Poppler 把演示文稿渲染为逐页 PNG。" + ) + parser.add_argument("--input", required=True) + parser.add_argument("--output-dir", required=True) + parser.add_argument("--start-slide", type=int, default=1) + parser.add_argument("--end-slide", type=int) + parser.add_argument("--max-slides", type=int, default=30) + parser.add_argument("--dpi", type=int, default=150) + parser.add_argument("--timeout", type=int, default=180) + parser.add_argument("--include-pdf", action="store_true") + parser.add_argument("--contact-sheet", action="store_true") + parser.add_argument("--overwrite", action="store_true") + return parser + + +def _contact_sheet( + images: list[Path], + slide_numbers: list[int], + destination: Path, +) -> None: + from PIL import Image, ImageDraw, ImageFont + + columns = min(4, max(1, len(images))) + rows = math.ceil(len(images) / columns) + thumb_width = 420 + label_height = 34 + gap = 18 + opened: list[Image.Image] = [] + try: + for path in images: + opened.append(Image.open(path).convert("RGB")) + aspect = opened[0].height / opened[0].width + thumb_height = max(1, round(thumb_width * aspect)) + canvas_width = columns * thumb_width + (columns + 1) * gap + canvas_height = rows * (thumb_height + label_height) + (rows + 1) * gap + canvas = Image.new("RGB", (canvas_width, canvas_height), "white") + draw = ImageDraw.Draw(canvas) + try: + font = ImageFont.truetype( + "/usr/share/fonts/truetype/dejavu/DejaVuSans.ttf", + 20, + ) + except OSError: + font = ImageFont.load_default() + for index, (image, slide_number) in enumerate(zip(opened, slide_numbers)): + row, column = divmod(index, columns) + x = gap + column * (thumb_width + gap) + y = gap + row * (thumb_height + label_height) + thumbnail = image.copy() + thumbnail.thumbnail((thumb_width, thumb_height)) + paste_x = x + (thumb_width - thumbnail.width) // 2 + paste_y = y + (thumb_height - thumbnail.height) // 2 + canvas.paste(thumbnail, (paste_x, paste_y)) + label = f"Slide {slide_number}" + label_box = draw.textbbox((0, 0), label, font=font) + label_width = label_box[2] - label_box[0] + draw.text( + (x + (thumb_width - label_width) // 2, y + thumb_height + 6), + label, + fill="black", + font=font, + ) + canvas.save(destination, format="PNG", optimize=True) + finally: + for image in opened: + image.close() + + +def main() -> dict[str, Any]: + from pypdf import PdfReader + + args = build_parser().parse_args() + if args.start_slide < 1: + raise ValueError("start-slide 必须大于 0") + if args.end_slide is not None and args.end_slide < args.start_slide: + raise ValueError("end-slide 不能小于 start-slide") + if args.max_slides < 1 or args.max_slides > 100: + raise ValueError("max-slides 必须在 1 到 100 之间") + if args.dpi < 72 or args.dpi > 300: + raise ValueError("dpi 必须在 72 到 300 之间") + + source = input_file(args.input, OOXML_PRESENTATION_SUFFIXES) + destination_dir = output_directory(args.output_dir) + with tempfile.TemporaryDirectory(prefix="pptx-render-") as temp_name: + temp_dir = Path(temp_name) + staged_input = temp_dir / f"presentation{source.suffix.lower()}" + shutil.copy2(source, staged_input) + pdf_path, office_output = run_soffice_convert( + staged_input, + target_format="pdf", + output_dir=temp_dir / "pdf", + timeout=args.timeout, + ) + slide_count = len(PdfReader(str(pdf_path)).pages) + if slide_count < 1: + raise ValueError("LibreOffice 生成的 PDF 没有页面") + if args.start_slide > slide_count: + raise ValueError(f"start-slide 超出页面总数 {slide_count}") + requested_end = slide_count if args.end_slide is None else args.end_slide + if requested_end > slide_count: + raise ValueError(f"end-slide 超出页面总数 {slide_count}") + actual_end = min( + requested_end, + args.start_slide + args.max_slides - 1, + ) + + raw_prefix = temp_dir / "raw-slide" + run_program( + [ + find_program("pdftoppm"), + "-png", + "-r", + str(args.dpi), + "-f", + str(args.start_slide), + "-l", + str(actual_end), + str(pdf_path), + str(raw_prefix), + ], + timeout=args.timeout, + ) + raw_pages = sorted( + temp_dir.glob("raw-slide-*.png"), + key=lambda path: int(path.stem.rsplit("-", 1)[1]), + ) + expected_count = actual_end - args.start_slide + 1 + if len(raw_pages) != expected_count: + raise RuntimeError( + f"Poppler 应生成 {expected_count} 页,实际生成 {len(raw_pages)} 页" + ) + slide_numbers = list(range(args.start_slide, actual_end + 1)) + destinations = [ + destination_dir / f"slide-{number:04d}.png" + for number in slide_numbers + ] + contact_destination = destination_dir / ( + f"contact-sheet-{args.start_slide:04d}-{actual_end:04d}.png" + ) + if args.contact_sheet: + destinations.append(contact_destination) + if args.include_pdf: + destinations.append(destination_dir / "presentation.pdf") + if not args.overwrite: + existing = [str(path) for path in destinations if path.exists()] + if existing: + raise FileExistsError("以下渲染目标已存在:" + "、".join(existing)) + + output_paths: list[str] = [] + published_pages: list[Path] = [] + for slide_number, raw_page in zip(slide_numbers, raw_pages): + destination = destination_dir / f"slide-{slide_number:04d}.png" + publish_file(raw_page, destination, overwrite=args.overwrite) + output_paths.append(str(destination)) + published_pages.append(destination) + + contact_sheet_path: Optional[str] = None + if args.contact_sheet: + staged_contact = temp_dir / "contact-sheet.png" + _contact_sheet( + published_pages, + slide_numbers, + staged_contact, + ) + publish_file( + staged_contact, + contact_destination, + overwrite=args.overwrite, + ) + contact_sheet_path = str(contact_destination) + + pdf_output: Optional[str] = None + if args.include_pdf: + destination_pdf = destination_dir / "presentation.pdf" + staged_pdf = temp_dir / "publish.pdf" + shutil.copy2(pdf_path, staged_pdf) + publish_file(staged_pdf, destination_pdf, overwrite=args.overwrite) + pdf_output = str(destination_pdf) + + next_slide = actual_end + 1 if actual_end < requested_end else None + return { + "source": str(source), + "slide_count": slide_count, + "start_slide": args.start_slide, + "end_slide": actual_end, + "rendered_slides": output_paths, + "contact_sheet": contact_sheet_path, + "pdf": pdf_output, + "has_more": next_slide is not None, + "next_slide": next_slide, + "dpi": args.dpi, + "office_stdout": office_output["stdout"], + "office_stderr": office_output["stderr"], + } + + +if __name__ == "__main__": + raise SystemExit(run_cli(main)) diff --git a/skills/pptx/scripts/unpack_presentation.py b/skills/pptx/scripts/unpack_presentation.py new file mode 100644 index 0000000..4bbb2ce --- /dev/null +++ b/skills/pptx/scripts/unpack_presentation.py @@ -0,0 +1,40 @@ +#!/usr/bin/env python3 + +from __future__ import annotations + +import argparse +from typing import Any + +from _pptx_common import ( + OOXML_PRESENTATION_SUFFIXES, + SkillArgumentParser, + input_file, + output_directory, + run_cli, + safe_extract_presentation, +) + + +def build_parser() -> argparse.ArgumentParser: + parser = SkillArgumentParser(description="安全解包 PowerPoint OOXML 文件。") + parser.add_argument("--input", required=True) + parser.add_argument("--output-dir", required=True) + return parser + + +def main() -> dict[str, Any]: + args = build_parser().parse_args() + source = input_file(args.input, OOXML_PRESENTATION_SUFFIXES) + destination = output_directory(args.output_dir) + if any(destination.iterdir()): + raise ValueError("output-dir 必须为空目录") + archive = safe_extract_presentation(source, destination) + return { + "source": str(source), + "path": str(destination), + "archive": archive, + } + + +if __name__ == "__main__": + raise SystemExit(run_cli(main)) diff --git a/skills/pptx/scripts/validate_presentation.py b/skills/pptx/scripts/validate_presentation.py new file mode 100644 index 0000000..f8c0c7f --- /dev/null +++ b/skills/pptx/scripts/validate_presentation.py @@ -0,0 +1,478 @@ +#!/usr/bin/env python3 + +from __future__ import annotations + +import argparse +import posixpath +import re +import shutil +import tempfile +import zipfile +from pathlib import Path, PurePosixPath +from typing import Any, Optional + +from _pptx_common import ( + OOXML_PRESENTATION_SUFFIXES, + P_NS, + R_NS, + SkillArgumentParser, + input_file, + inspect_archive, + parse_xml_bytes, + run_cli, + run_soffice_convert, +) + + +REL_NS = "http://schemas.openxmlformats.org/package/2006/relationships" +C_NS = "http://schemas.openxmlformats.org/drawingml/2006/chart" +PLACEHOLDER_RE = re.compile( + r"\b(?:lorem|ipsum|todo|x{3,})\b|\[insert|this\s+(?:page|slide).+layout", + re.IGNORECASE, +) +EMU_PER_INCH = 914400 + + +def build_parser() -> argparse.ArgumentParser: + parser = SkillArgumentParser(description="校验 PowerPoint 的结构、关系和可渲染性。") + parser.add_argument("--input", required=True) + parser.add_argument("--original") + parser.add_argument("--check-render", action="store_true") + parser.add_argument("--timeout", type=int, default=180) + return parser + + +def _issue( + code: str, + message: str, + *, + part: Optional[str] = None, + slide: Optional[int] = None, +) -> dict[str, Any]: + payload: dict[str, Any] = {"code": code, "message": message} + if part is not None: + payload["part"] = part + if slide is not None: + payload["slide"] = slide + return payload + + +def _owner_part_for_rels(name: str) -> str: + if name == "_rels/.rels": + return "" + path = PurePosixPath(name) + if path.parent.name != "_rels" or not path.name.endswith(".rels"): + raise ValueError(f"关系部件路径无效:{name}") + owner_name = path.name[: -len(".rels")] + owner_parent = path.parent.parent + return (owner_parent / owner_name).as_posix() + + +def _resolve_relationship_target(owner_part: str, target: str) -> Optional[str]: + if not target or target.startswith("#"): + return None + if target.startswith("/"): + normalized = posixpath.normpath(target).lstrip("/") + if normalized == ".." or normalized.startswith("../"): + return None + return normalized + base = posixpath.dirname(owner_part) + normalized = posixpath.normpath(posixpath.join(base, target)) + if normalized == ".." or normalized.startswith("../"): + return None + return normalized.lstrip("/") + + +def _validate_relationships( + archive: zipfile.ZipFile, + names: set[str], +) -> tuple[list[dict[str, Any]], list[dict[str, Any]], dict[str, dict[str, str]]]: + errors: list[dict[str, Any]] = [] + externals: list[dict[str, Any]] = [] + maps: dict[str, dict[str, str]] = {} + for name in sorted(item for item in names if item.endswith(".rels")): + try: + root = parse_xml_bytes(archive.read(name), label=name) + except ValueError as exc: + errors.append(_issue("invalid_relationship_xml", str(exc), part=name)) + continue + owner_part = _owner_part_for_rels(name) + relationships: dict[str, str] = {} + for node in root.findall(f"{{{REL_NS}}}Relationship"): + relationship_id = node.attrib.get("Id", "") + target = node.attrib.get("Target", "") + if not relationship_id or relationship_id in relationships: + errors.append( + _issue( + "duplicate_or_missing_relationship_id", + f"关系 Id 缺失或重复:{relationship_id!r}", + part=name, + ) + ) + continue + if node.attrib.get("TargetMode") == "External": + externals.append( + { + "part": name, + "id": relationship_id, + "type": node.attrib.get("Type", ""), + "target": target, + } + ) + relationships[relationship_id] = target + continue + resolved = _resolve_relationship_target(owner_part, target) + if resolved is None: + errors.append( + _issue( + "unsafe_relationship_target", + f"关系目标不安全:{target}", + part=name, + ) + ) + continue + relationships[relationship_id] = resolved + if resolved not in names: + errors.append( + _issue( + "missing_relationship_target", + f"关系目标不存在:{resolved}", + part=name, + ) + ) + maps[owner_part] = relationships + return errors, externals, maps + + +def _validate_slide_order( + archive: zipfile.ZipFile, + relationship_maps: dict[str, dict[str, str]], +) -> tuple[list[dict[str, Any]], list[str]]: + errors: list[dict[str, Any]] = [] + ordered_parts: list[str] = [] + root = parse_xml_bytes( + archive.read("ppt/presentation.xml"), + label="ppt/presentation.xml", + ) + slide_ids: set[str] = set() + relationship_ids: set[str] = set() + presentation_relationships = relationship_maps.get( + "ppt/presentation.xml", + {}, + ) + for node in root.findall(f".//{{{P_NS}}}sldId"): + slide_id = node.attrib.get("id", "") + relationship_id = node.attrib.get(f"{{{R_NS}}}id", "") + if not slide_id or slide_id in slide_ids: + errors.append( + _issue( + "duplicate_or_missing_slide_id", + f"页面 id 缺失或重复:{slide_id!r}", + part="ppt/presentation.xml", + ) + ) + slide_ids.add(slide_id) + if not relationship_id or relationship_id in relationship_ids: + errors.append( + _issue( + "duplicate_or_missing_slide_relationship", + f"页面关系 id 缺失或重复:{relationship_id!r}", + part="ppt/presentation.xml", + ) + ) + relationship_ids.add(relationship_id) + target = presentation_relationships.get(relationship_id) + if not target or not target.startswith("ppt/slides/slide"): + errors.append( + _issue( + "invalid_slide_relationship", + f"页面关系 {relationship_id!r} 未指向有效 slide 部件", + part="ppt/presentation.xml", + ) + ) + else: + ordered_parts.append(target) + if not ordered_parts: + errors.append( + _issue( + "no_slides", + "演示文稿不包含页面", + part="ppt/presentation.xml", + ) + ) + return errors, ordered_parts + + +def _validate_charts( + archive: zipfile.ZipFile, + names: set[str], +) -> list[dict[str, Any]]: + errors: list[dict[str, Any]] = [] + for name in sorted( + item + for item in names + if item.startswith("ppt/charts/chart") and item.endswith(".xml") + ): + root = parse_xml_bytes(archive.read(name), label=name) + declared_axis_ids: set[str] = set() + for axis_name in ("catAx", "valAx", "dateAx", "serAx"): + for axis in root.findall(f".//{{{C_NS}}}{axis_name}"): + node = axis.find(f"{{{C_NS}}}axId") + if node is not None and node.attrib.get("val"): + declared_axis_ids.add(node.attrib["val"]) + referenced_axis_ids = { + node.attrib["val"] + for node in root.findall(f".//{{{C_NS}}}axId") + if node.attrib.get("val") + } + # PptxGenJS writes the standard primary series-axis id on ordinary + # two-dimensional charts even though no serAx element is required. + # Secondary value/category ids must still have real declarations. + tolerated_series_axis_ids = {"2094734556"} + missing_axis_ids = sorted( + referenced_axis_ids + - declared_axis_ids + - tolerated_series_axis_ids + ) + if declared_axis_ids and missing_axis_ids: + errors.append( + _issue( + "undeclared_chart_axis", + "图表引用了未声明的坐标轴:" + + "、".join(missing_axis_ids), + part=name, + ) + ) + for bar_chart in root.findall(f".//{{{C_NS}}}barChart"): + grouping = bar_chart.find(f"{{{C_NS}}}grouping") + grouping_value = grouping.attrib.get("val") if grouping is not None else "" + if grouping_value not in {"stacked", "percentStacked"}: + continue + invalid_label = bar_chart.find( + f".//{{{C_NS}}}dLblPos[@val='outEnd']" + ) + if invalid_label is not None: + errors.append( + _issue( + "invalid_stacked_chart_label_position", + "堆积柱形或条形图不能使用 outEnd 数据标签位置", + part=name, + ) + ) + return errors + + +def _visual_structure(path: Path) -> tuple[list[dict[str, Any]], list[dict[str, Any]], int]: + from pptx import Presentation + + errors: list[dict[str, Any]] = [] + warnings: list[dict[str, Any]] = [] + presentation = Presentation(str(path)) + slide_width = int(presentation.slide_width) + slide_height = int(presentation.slide_height) + tolerance = 2000 + for slide_number, slide in enumerate(presentation.slides, start=1): + if len(slide.shapes) == 0: + warnings.append( + _issue("empty_slide", "页面没有可见形状", slide=slide_number) + ) + for shape in slide.shapes: + left = int(shape.left) + top = int(shape.top) + right = left + int(shape.width) + bottom = top + int(shape.height) + if ( + left < -tolerance + or top < -tolerance + or right > slide_width + tolerance + or bottom > slide_height + tolerance + ): + errors.append( + _issue( + "shape_out_of_bounds", + f"形状 {shape.name!r} 超出页面边界", + slide=slide_number, + ) + ) + text = getattr(shape, "text", "") + if text and PLACEHOLDER_RE.search(text): + warnings.append( + _issue( + "placeholder_text", + f"形状 {shape.name!r} 可能残留占位文本", + slide=slide_number, + ) + ) + return errors, warnings, len(presentation.slides) + + +def _validate_once(path: Path, *, check_render: bool, timeout: int) -> dict[str, Any]: + archive_info = inspect_archive(path) + errors: list[dict[str, Any]] = [] + warnings: list[dict[str, Any]] = [] + if archive_info["missing_required_parts"]: + errors.append( + _issue( + "missing_required_parts", + "缺少必要部件:" + + "、".join(archive_info["missing_required_parts"]), + ) + ) + return { + "errors": errors, + "warnings": warnings, + "archive": archive_info, + "slide_count": 0, + "external_relationships": [], + } + if archive_info["duplicate_members"]: + errors.append( + _issue( + "duplicate_archive_members", + "压缩包存在重复成员:" + + "、".join(archive_info["duplicate_members"][:20]), + ) + ) + + with zipfile.ZipFile(path) as archive: + names = set(archive.namelist()) + for name in sorted( + item for item in names if item.endswith((".xml", ".rels")) + ): + try: + parse_xml_bytes(archive.read(name), label=name) + except ValueError as exc: + errors.append(_issue("invalid_xml", str(exc), part=name)) + relationship_errors, externals, maps = _validate_relationships( + archive, + names, + ) + errors.extend(relationship_errors) + slide_errors, ordered_parts = _validate_slide_order(archive, maps) + errors.extend(slide_errors) + orphaned_slides = sorted( + name + for name in names + if name.startswith("ppt/slides/slide") + and name.endswith(".xml") + and name not in set(ordered_parts) + ) + if orphaned_slides: + warnings.append( + _issue( + "orphaned_slide_parts", + "存在未被 presentation.xml 引用的页面部件:" + + "、".join(orphaned_slides[:20]), + ) + ) + errors.extend(_validate_charts(archive, names)) + + try: + shape_errors, shape_warnings, slide_count = _visual_structure(path) + errors.extend(shape_errors) + warnings.extend(shape_warnings) + except Exception as exc: + errors.append( + _issue( + "python_pptx_open_failed", + f"python-pptx 无法打开演示文稿:{exc}", + ) + ) + slide_count = 0 + + render_result: Optional[dict[str, Any]] = None + if check_render and not any( + item["code"] in {"missing_required_parts", "invalid_xml"} + for item in errors + ): + from pypdf import PdfReader + + with tempfile.TemporaryDirectory(prefix="pptx-validate-render-") as temp_name: + temp_dir = Path(temp_name) + staged = temp_dir / f"presentation{path.suffix.lower()}" + shutil.copy2(path, staged) + pdf, office_output = run_soffice_convert( + staged, + target_format="pdf", + output_dir=temp_dir / "pdf", + timeout=timeout, + ) + pdf_pages = len(PdfReader(str(pdf)).pages) + render_result = { + "pdf_pages": pdf_pages, + "office_stdout": office_output["stdout"], + "office_stderr": office_output["stderr"], + } + if pdf_pages != slide_count: + errors.append( + _issue( + "render_page_count_mismatch", + f"结构页数为 {slide_count},渲染页数为 {pdf_pages}", + ) + ) + return { + "errors": errors, + "warnings": warnings, + "archive": archive_info, + "slide_count": slide_count, + "external_relationships": externals, + "render": render_result, + } + + +def _signature(issue: dict[str, Any]) -> tuple[Any, ...]: + return ( + issue.get("code"), + issue.get("part"), + issue.get("slide"), + issue.get("message"), + ) + + +def main() -> dict[str, Any]: + args = build_parser().parse_args() + source = input_file(args.input, OOXML_PRESENTATION_SUFFIXES) + result = _validate_once( + source, + check_render=args.check_render, + timeout=args.timeout, + ) + baseline_count = 0 + if args.original: + original = input_file(args.original, OOXML_PRESENTATION_SUFFIXES) + baseline = _validate_once( + original, + check_render=False, + timeout=args.timeout, + ) + baseline_signatures = { + _signature(item) + for item in baseline["errors"] + if item["code"] + in { + "invalid_xml", + "shape_out_of_bounds", + "python_pptx_open_failed", + } + } + before = len(result["errors"]) + result["errors"] = [ + item + for item in result["errors"] + if _signature(item) not in baseline_signatures + ] + baseline_count = before - len(result["errors"]) + result["original"] = str(original) + status = "valid" if not result["errors"] else "invalid" + return { + "source": str(source), + "status": status, + "issue_count": len(result["errors"]), + "warning_count": len(result["warnings"]), + "baselined_issue_count": baseline_count, + **result, + } + + +if __name__ == "__main__": + raise SystemExit(run_cli(main))