feat: 优化docx pdf pptx xlsx 技能
This commit is contained in:
parent
9a04c4e4c6
commit
48d446c27e
@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
name: docx
|
name: docx
|
||||||
description: "创建、读取、编辑、转换、批注、接受修订、校验和渲染本地或远程 HTTPS Microsoft Word 文档,并按需识别文档图片、截图和扫描页中的文字。用户提到 Word、文档、报告、备忘录、合同、信函、模板、目录、页眉页脚、页码、表格、图片、图片文字 OCR、批注或修订,或提供 HTTPS Word 地址、.docx、.dotx、.doc 文件时使用;支持安全下载、结构化创建、跨 Run 查找替换、本地 RapidOCR、安全 OOXML 解包/打包、旧格式转换、关系与 XML 校验及逐页视觉检查。若主要交付物是 PDF、电子表格、Google Docs 或普通代码,则不要使用。"
|
description: "创建、读取、编辑、转换、批注、接受修订、校验和渲染本地或远程 HTTPS Microsoft Word 文档,并下载 Word 任务所需且不超过 25 MiB 的图片、音视频、压缩包和其他 HTTPS 附件,按需识别文档图片、截图和扫描页中的文字。用户提到 Word、文档、报告、备忘录、合同、信函、模板、目录、页眉页脚、页码、表格、图片、图片文字 OCR、批注或修订,或提供 HTTPS Word 地址、.docx、.dotx、.doc 文件时使用;支持源文档安全下载、结构化创建、跨 Run 查找替换、本地 RapidOCR、安全 OOXML 解包/打包、旧格式转换、关系与 XML 校验及逐页视觉检查。若主要交付物是 PDF、电子表格、Google Docs 或普通代码,则不要使用。"
|
||||||
---
|
---
|
||||||
|
|
||||||
# Word 文档处理
|
# Word 文档处理
|
||||||
@ -14,7 +14,7 @@ description: "创建、读取、编辑、转换、批注、接受修订、校验
|
|||||||
- LibreOffice、Pandoc、Poppler 和 ZIP 操作只允许由固定 Python 脚本在内部调用。
|
- LibreOffice、Pandoc、Poppler 和 ZIP 操作只允许由固定 Python 脚本在内部调用。
|
||||||
- 每次检查脚本返回 JSON;只有 `ok` 为 `true` 时才继续。`validate_document.py` 还必须返回 `status: valid`。
|
- 每次检查脚本返回 JSON;只有 `ok` 为 `true` 时才继续。`validate_document.py` 还必须返回 `status: valid`。
|
||||||
- 只在需要读取图片、截图或扫描页中的文字时调用 `ocr_document.py`。只使用 `pages[]` 中 `usable_for_summary: true` 的 `text`;低置信度结果不得作为可靠正文。
|
- 只在需要读取图片、截图或扫描页中的文字时调用 `ocr_document.py`。只使用 `pages[]` 中 `usable_for_summary: true` 的 `text`;低置信度结果不得作为可靠正文。
|
||||||
- 远程地址只交给 `download_document.py`;不要在回复、日志摘要或文件名中复述可能含敏感查询参数的完整 URL。
|
- 远程 Word 源文档只交给 `download_document.py`;任务所需的远程图片、视频、音频、压缩包或其他附件只交给 `download_attachment.py`。不要在回复、日志摘要或文件名中复述可能含敏感查询参数的完整 URL。
|
||||||
- 不覆盖用户提供的源文件。Word 最终文件一律写入 `/usr/local/src/word/`,下载缓存、中间文件和渲染结果一律写入 `/usr/local/src/word/tmp/<任务名>/`。始终传绝对路径;固定脚本会自动创建目录并拒绝该根目录之外的输出。
|
- 不覆盖用户提供的源文件。Word 最终文件一律写入 `/usr/local/src/word/`,下载缓存、中间文件和渲染结果一律写入 `/usr/local/src/word/tmp/<任务名>/`。始终传绝对路径;固定脚本会自动创建目录并拒绝该根目录之外的输出。
|
||||||
- 环境已预置依赖,不安装软件包,也不提示用户安装依赖。
|
- 环境已预置依赖,不安装软件包,也不提示用户安装依赖。
|
||||||
|
|
||||||
@ -23,6 +23,7 @@ description: "创建、读取、编辑、转换、批注、接受修订、校验
|
|||||||
| 脚本 | 用途 | 底层能力 |
|
| 脚本 | 用途 | 底层能力 |
|
||||||
| --- | --- | --- |
|
| --- | --- | --- |
|
||||||
| `scripts/download_document.py` | 下载并校验远程 HTTPS Word 文档 | Python `urllib`、安全 OOXML 解析、`python-docx` |
|
| `scripts/download_document.py` | 下载并校验远程 HTTPS Word 文档 | Python `urllib`、安全 OOXML 解析、`python-docx` |
|
||||||
|
| `scripts/download_attachment.py` | 下载图片、音视频、压缩包等通用 HTTPS 附件 | Python `urllib`、HEAD 大小探测、流式硬限制 |
|
||||||
| `scripts/inspect_document.py` | 分段读取正文、表格、样式、批注和修订 | `python-docx`、安全 OOXML 解析 |
|
| `scripts/inspect_document.py` | 分段读取正文、表格、样式、批注和修订 | `python-docx`、安全 OOXML 解析 |
|
||||||
| `scripts/ocr_document.py` | 按页识别图片、截图和扫描页中的文字 | RapidOCR、ONNX Runtime、LibreOffice、Poppler、`pdfplumber` |
|
| `scripts/ocr_document.py` | 按页识别图片、截图和扫描页中的文字 | RapidOCR、ONNX Runtime、LibreOffice、Poppler、`pdfplumber` |
|
||||||
| `scripts/create_document.py` | 按受控 JSON 创建专业 DOCX | `python-docx`、Pillow |
|
| `scripts/create_document.py` | 按受控 JSON 创建专业 DOCX | `python-docx`、Pillow |
|
||||||
@ -61,13 +62,25 @@ description: "创建、读取、编辑、转换、批注、接受修订、校验
|
|||||||
可选参数:
|
可选参数:
|
||||||
|
|
||||||
- `--timeout <1-600>`:连接和读取超时秒数,默认 `60`。
|
- `--timeout <1-600>`:连接和读取超时秒数,默认 `60`。
|
||||||
- `--max-bytes <字节数>`:默认 `104857600`(100 MiB),最高 `536870912`(512 MiB)。
|
- `--max-bytes <字节数>`:默认且最高 `26214400`(25 MiB),只允许设置更小的限制。
|
||||||
- `--overwrite`:只在目标是本次任务生成的旧缓存时使用。
|
- `--overwrite`:只在目标是本次任务生成的旧缓存时使用。
|
||||||
|
|
||||||
`output` 扩展名必须是 `.docx`、`.dotx` 或 `.doc`。脚本阻止 HTTPS 重定向降级到 HTTP,流式限制大小,先写同目录临时文件,再原子发布;DOCX/DOTX 会检查 ZIP 路径、成员大小、必要部件和内容类型,并用 `python-docx` 打开。实际 OOXML 格式与 `output` 扩展名不一致时,根据错误中的实际格式更正缓存扩展名,再调用同一脚本。
|
`output` 扩展名必须是 `.docx`、`.dotx` 或 `.doc`。脚本阻止 HTTPS 重定向降级到 HTTP,流式限制大小,先写同目录临时文件,再原子发布;DOCX/DOTX 会检查 ZIP 路径、成员大小、必要部件和内容类型,并用 `python-docx` 打开。实际 OOXML 格式与 `output` 扩展名不一致时,根据错误中的实际格式更正缓存扩展名,再调用同一脚本。
|
||||||
|
|
||||||
成功结果包含 `path`、`size_bytes`、`format` 和 `validation`;OOXML 还包含段落、表格和章节数量。后续脚本只使用返回的本地 `path`,不再访问原 URL。
|
成功结果包含 `path`、`size_bytes`、`format` 和 `validation`;OOXML 还包含段落、表格和章节数量。后续脚本只使用返回的本地 `path`,不再访问原 URL。
|
||||||
|
|
||||||
|
## 下载通用附件
|
||||||
|
|
||||||
|
需要下载作为 Word 任务素材的图片、视频、音频、压缩包或其他文件时,调用 `scripts/download_attachment.py`:
|
||||||
|
|
||||||
|
```text
|
||||||
|
--url 'https://example.com/asset.bin?signature=...' --output '/usr/local/src/word/tmp/<任务名>/asset.bin'
|
||||||
|
```
|
||||||
|
|
||||||
|
只接受 HTTPS 地址,`output` 可使用任意附件扩展名。可选参数只有 `--timeout <1-600>`(默认 `60`)和 `--overwrite`。附件上限固定为 25 MiB(26214400 字节),不可调高:脚本先用 HEAD 探测远端声明大小,再检查 GET 响应声明,并在流式接收时持续兜底计数;任一阶段发现超限都会返回 `ok: false` 和明确的“已拒绝下载”错误,且不会发布部分文件。
|
||||||
|
|
||||||
|
成功结果包含 `path`、实际 `size_bytes`、`declared_size_bytes`、`size_limit_bytes`、`size_probe` 和 `content_type`。本脚本不校验文件业务格式;远程 Word 源文档仍使用 `download_document.py`。
|
||||||
|
|
||||||
## 检查文档
|
## 检查文档
|
||||||
|
|
||||||
调用 `scripts/inspect_document.py`:
|
调用 `scripts/inspect_document.py`:
|
||||||
|
|||||||
256
skills/docx/scripts/download_attachment.py
Normal file
256
skills/docx/scripts/download_attachment.py
Normal file
@ -0,0 +1,256 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import os
|
||||||
|
import socket
|
||||||
|
import sys
|
||||||
|
import tempfile
|
||||||
|
import urllib.error
|
||||||
|
import urllib.parse
|
||||||
|
import urllib.request
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any, NoReturn, Optional
|
||||||
|
|
||||||
|
from _docx_common import emit, failure_message, output_file, publish_file
|
||||||
|
|
||||||
|
|
||||||
|
DEFAULT_TIMEOUT_SECONDS = 60
|
||||||
|
MAX_ATTACHMENT_BYTES = 25 * 1024 * 1024
|
||||||
|
CHUNK_SIZE = 1024 * 1024
|
||||||
|
USER_AGENT = "wechat-robot-docx-attachment-downloader/1.0"
|
||||||
|
|
||||||
|
|
||||||
|
class SkillArgumentParser(argparse.ArgumentParser):
|
||||||
|
def error(self, message: str) -> NoReturn:
|
||||||
|
raise ValueError(f"参数错误:{message}")
|
||||||
|
|
||||||
|
|
||||||
|
def _validate_https_url(value: str) -> str:
|
||||||
|
url = value.strip()
|
||||||
|
if not url:
|
||||||
|
raise ValueError("附件 URL 不能为空")
|
||||||
|
if any(character.isspace() or ord(character) < 32 for character in url):
|
||||||
|
raise ValueError("附件 URL 不能包含空白字符或控制字符")
|
||||||
|
|
||||||
|
parsed = urllib.parse.urlsplit(url)
|
||||||
|
if parsed.scheme.lower() != "https" or not parsed.hostname:
|
||||||
|
raise ValueError("附件 URL 必须是有效的 HTTPS 地址")
|
||||||
|
if parsed.username is not None or parsed.password is not None:
|
||||||
|
raise ValueError("附件 URL 不允许包含用户名或密码")
|
||||||
|
try:
|
||||||
|
parsed.port
|
||||||
|
except ValueError as exc:
|
||||||
|
raise ValueError("附件 URL 端口格式不正确") from exc
|
||||||
|
return url
|
||||||
|
|
||||||
|
|
||||||
|
class HTTPSOnlyRedirectHandler(urllib.request.HTTPRedirectHandler):
|
||||||
|
def redirect_request(self, req, fp, code, msg, headers, newurl):
|
||||||
|
return super().redirect_request(
|
||||||
|
req,
|
||||||
|
fp,
|
||||||
|
code,
|
||||||
|
msg,
|
||||||
|
headers,
|
||||||
|
_validate_https_url(newurl),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _parse_args(argv: list[str]) -> argparse.Namespace:
|
||||||
|
parser = SkillArgumentParser(
|
||||||
|
description="下载不超过 25 MiB 的远程 HTTPS 通用附件"
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--url",
|
||||||
|
"--attachment-url",
|
||||||
|
"--attachment_url",
|
||||||
|
dest="url",
|
||||||
|
required=True,
|
||||||
|
help="远程 HTTPS 附件地址",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--output",
|
||||||
|
required=True,
|
||||||
|
help="附件本地输出路径;允许图片、音视频、压缩包及其他文件类型",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--timeout",
|
||||||
|
type=int,
|
||||||
|
default=DEFAULT_TIMEOUT_SECONDS,
|
||||||
|
help=f"连接和读取超时秒数,默认 {DEFAULT_TIMEOUT_SECONDS}",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--overwrite",
|
||||||
|
action="store_true",
|
||||||
|
help="允许覆盖本次任务已存在的缓存文件",
|
||||||
|
)
|
||||||
|
args = parser.parse_args(argv)
|
||||||
|
|
||||||
|
args.url = _validate_https_url(args.url)
|
||||||
|
if args.timeout < 1 or args.timeout > 600:
|
||||||
|
raise ValueError("timeout 必须在 1 到 600 秒之间")
|
||||||
|
args.output = output_file(args.output, overwrite=args.overwrite)
|
||||||
|
return args
|
||||||
|
|
||||||
|
|
||||||
|
def _content_length(headers: Any) -> Optional[int]:
|
||||||
|
raw_value = headers.get("Content-Length")
|
||||||
|
if raw_value is None:
|
||||||
|
return None
|
||||||
|
try:
|
||||||
|
size = int(raw_value)
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
return None
|
||||||
|
return size if size >= 0 else None
|
||||||
|
|
||||||
|
|
||||||
|
def _reject_if_too_large(size_bytes: int, *, source: str) -> None:
|
||||||
|
if size_bytes <= MAX_ATTACHMENT_BYTES:
|
||||||
|
return
|
||||||
|
raise ValueError(
|
||||||
|
"附件超过 25 MiB(26214400 字节)限制,已拒绝下载;"
|
||||||
|
f"{source}大小为 {size_bytes} 字节"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _probe_size(
|
||||||
|
opener: urllib.request.OpenerDirector,
|
||||||
|
url: str,
|
||||||
|
timeout: int,
|
||||||
|
) -> Optional[int]:
|
||||||
|
request = urllib.request.Request(
|
||||||
|
url,
|
||||||
|
headers={
|
||||||
|
"Accept": "*/*",
|
||||||
|
"Accept-Encoding": "identity",
|
||||||
|
"User-Agent": USER_AGENT,
|
||||||
|
},
|
||||||
|
method="HEAD",
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
with opener.open(request, timeout=timeout) as response:
|
||||||
|
_validate_https_url(response.geturl())
|
||||||
|
size = _content_length(response.headers)
|
||||||
|
except (
|
||||||
|
urllib.error.HTTPError,
|
||||||
|
urllib.error.URLError,
|
||||||
|
TimeoutError,
|
||||||
|
socket.timeout,
|
||||||
|
):
|
||||||
|
return None
|
||||||
|
|
||||||
|
if size is not None:
|
||||||
|
_reject_if_too_large(size, source="远端声明")
|
||||||
|
return size
|
||||||
|
|
||||||
|
|
||||||
|
def _content_type(headers: Any) -> Optional[str]:
|
||||||
|
raw_value = headers.get("Content-Type")
|
||||||
|
if not raw_value:
|
||||||
|
return None
|
||||||
|
media_type = raw_value.split(";", 1)[0].strip().lower()
|
||||||
|
return media_type or None
|
||||||
|
|
||||||
|
|
||||||
|
def _download(args: argparse.Namespace) -> dict[str, Any]:
|
||||||
|
output: Path = args.output
|
||||||
|
opener = urllib.request.build_opener(HTTPSOnlyRedirectHandler())
|
||||||
|
probed_size = _probe_size(opener, args.url, args.timeout)
|
||||||
|
request = urllib.request.Request(
|
||||||
|
args.url,
|
||||||
|
headers={
|
||||||
|
"Accept": "*/*",
|
||||||
|
"Accept-Encoding": "identity",
|
||||||
|
"User-Agent": USER_AGENT,
|
||||||
|
},
|
||||||
|
method="GET",
|
||||||
|
)
|
||||||
|
|
||||||
|
temp_path: Optional[Path] = None
|
||||||
|
downloaded_bytes = 0
|
||||||
|
response_size: Optional[int] = None
|
||||||
|
response_type: Optional[str] = None
|
||||||
|
try:
|
||||||
|
with tempfile.NamedTemporaryFile(
|
||||||
|
mode="wb",
|
||||||
|
prefix=f".{output.stem}.",
|
||||||
|
suffix=f".part{output.suffix}",
|
||||||
|
dir=str(output.parent),
|
||||||
|
delete=False,
|
||||||
|
) as temp_file:
|
||||||
|
temp_path = Path(temp_file.name)
|
||||||
|
with opener.open(request, timeout=args.timeout) as response:
|
||||||
|
_validate_https_url(response.geturl())
|
||||||
|
response_size = _content_length(response.headers)
|
||||||
|
response_type = _content_type(response.headers)
|
||||||
|
if response_size is not None:
|
||||||
|
_reject_if_too_large(response_size, source="下载响应声明")
|
||||||
|
|
||||||
|
while True:
|
||||||
|
chunk = response.read(CHUNK_SIZE)
|
||||||
|
if not chunk:
|
||||||
|
break
|
||||||
|
downloaded_bytes += len(chunk)
|
||||||
|
_reject_if_too_large(downloaded_bytes, source="已接收")
|
||||||
|
temp_file.write(chunk)
|
||||||
|
|
||||||
|
temp_file.flush()
|
||||||
|
os.fsync(temp_file.fileno())
|
||||||
|
|
||||||
|
if downloaded_bytes == 0:
|
||||||
|
raise ValueError("远程服务器返回了空附件")
|
||||||
|
|
||||||
|
publish_file(temp_path, output, overwrite=args.overwrite)
|
||||||
|
temp_path = None
|
||||||
|
declared_size = (
|
||||||
|
response_size if response_size is not None else probed_size
|
||||||
|
)
|
||||||
|
if response_size is not None:
|
||||||
|
size_probe = "get-content-length"
|
||||||
|
elif probed_size is not None:
|
||||||
|
size_probe = "head-content-length"
|
||||||
|
else:
|
||||||
|
size_probe = "stream"
|
||||||
|
return {
|
||||||
|
"path": str(output),
|
||||||
|
"size_bytes": downloaded_bytes,
|
||||||
|
"declared_size_bytes": declared_size,
|
||||||
|
"size_limit_bytes": MAX_ATTACHMENT_BYTES,
|
||||||
|
"size_probe": size_probe,
|
||||||
|
"content_type": response_type,
|
||||||
|
}
|
||||||
|
finally:
|
||||||
|
if temp_path is not None:
|
||||||
|
try:
|
||||||
|
temp_path.unlink(missing_ok=True)
|
||||||
|
except OSError:
|
||||||
|
pass
|
||||||
|
|
||||||
|
|
||||||
|
def _failure_message(exc: Exception) -> str:
|
||||||
|
if isinstance(exc, urllib.error.HTTPError):
|
||||||
|
return f"附件下载失败:远程服务器返回 HTTP {exc.code}"
|
||||||
|
if isinstance(exc, (TimeoutError, socket.timeout)):
|
||||||
|
return "附件下载失败:连接或读取超时"
|
||||||
|
if isinstance(exc, urllib.error.URLError):
|
||||||
|
if isinstance(exc.reason, (TimeoutError, socket.timeout)):
|
||||||
|
return "附件下载失败:连接或读取超时"
|
||||||
|
return "附件下载失败:无法访问远程服务器"
|
||||||
|
return failure_message(exc)
|
||||||
|
|
||||||
|
|
||||||
|
def main(argv: Optional[list[str]] = None) -> int:
|
||||||
|
try:
|
||||||
|
args = _parse_args(sys.argv[1:] if argv is None else argv)
|
||||||
|
result = _download(args)
|
||||||
|
except Exception as exc:
|
||||||
|
emit({"ok": False, "error": _failure_message(exc)})
|
||||||
|
return 1
|
||||||
|
emit({"ok": True, **result})
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
@ -26,8 +26,8 @@ from _docx_common import (
|
|||||||
|
|
||||||
|
|
||||||
DEFAULT_TIMEOUT_SECONDS = 60
|
DEFAULT_TIMEOUT_SECONDS = 60
|
||||||
DEFAULT_MAX_BYTES = 100 * 1024 * 1024
|
MAX_ALLOWED_BYTES = 25 * 1024 * 1024
|
||||||
MAX_ALLOWED_BYTES = 512 * 1024 * 1024
|
DEFAULT_MAX_BYTES = MAX_ALLOWED_BYTES
|
||||||
CHUNK_SIZE = 1024 * 1024
|
CHUNK_SIZE = 1024 * 1024
|
||||||
USER_AGENT = "wechat-robot-docx-skill/1.0"
|
USER_AGENT = "wechat-robot-docx-skill/1.0"
|
||||||
OLE_COMPOUND_MAGIC = bytes.fromhex("D0CF11E0A1B11AE1")
|
OLE_COMPOUND_MAGIC = bytes.fromhex("D0CF11E0A1B11AE1")
|
||||||
@ -244,7 +244,7 @@ def _download(args: argparse.Namespace) -> dict[str, Any]:
|
|||||||
expected_bytes = 0
|
expected_bytes = 0
|
||||||
if expected_bytes > args.max_bytes:
|
if expected_bytes > args.max_bytes:
|
||||||
raise ValueError(
|
raise ValueError(
|
||||||
"远程文件超过大小限制:"
|
"远程文件超过大小限制,已拒绝下载:"
|
||||||
f"最多允许 {args.max_bytes} 字节"
|
f"最多允许 {args.max_bytes} 字节"
|
||||||
)
|
)
|
||||||
|
|
||||||
@ -256,7 +256,7 @@ def _download(args: argparse.Namespace) -> dict[str, Any]:
|
|||||||
downloaded_bytes += len(chunk)
|
downloaded_bytes += len(chunk)
|
||||||
if downloaded_bytes > args.max_bytes:
|
if downloaded_bytes > args.max_bytes:
|
||||||
raise ValueError(
|
raise ValueError(
|
||||||
"远程文件超过大小限制:"
|
"远程文件超过大小限制,已拒绝下载:"
|
||||||
f"最多允许 {args.max_bytes} 字节"
|
f"最多允许 {args.max_bytes} 字节"
|
||||||
)
|
)
|
||||||
temp_file.write(chunk)
|
temp_file.write(chunk)
|
||||||
|
|||||||
@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
name: pdf
|
name: pdf
|
||||||
description: "处理本地 PDF 文件或远程 HTTPS PDF 链接,包括安全下载、元数据与页面检查、多引擎分段文本提取和质量检测、扫描页本地 OCR、表格提取、按需页面 PNG 渲染、从文本创建 PDF、合并、拆分、旋转及最终质量校验。当用户提供 .pdf 文件或 HTTPS PDF 地址,或要求总结、读取、识别扫描件、生成、编辑、转换或审阅 PDF 时使用。"
|
description: "处理本地 PDF 文件或远程 HTTPS PDF 链接,并下载 PDF 任务所需且不超过 25 MiB 的图片、音视频、压缩包和其他 HTTPS 附件;包括源 PDF 安全下载、元数据与页面检查、多引擎分段文本提取和质量检测、扫描页本地 OCR、表格提取、按需页面 PNG 渲染、从文本创建 PDF、合并、拆分、旋转及最终质量校验。当用户提供 .pdf 文件或 HTTPS PDF 地址,或要求总结、读取、识别扫描件、生成、编辑、转换或审阅 PDF 时使用。"
|
||||||
---
|
---
|
||||||
|
|
||||||
# PDF 处理
|
# PDF 处理
|
||||||
@ -16,6 +16,7 @@ description: "处理本地 PDF 文件或远程 HTTPS PDF 链接,包括安全
|
|||||||
- 每次检查脚本返回的 JSON;只有 `ok` 为 `true` 时才继续。
|
- 每次检查脚本返回的 JSON;只有 `ok` 为 `true` 时才继续。
|
||||||
- 收到 `ok: false` 时,依据 `error` 调整合法参数或向用户说明失败原因,不要把参数改传给其他脚本碰运气。
|
- 收到 `ok: false` 时,依据 `error` 调整合法参数或向用户说明失败原因,不要把参数改传给其他脚本碰运气。
|
||||||
- 阅读或总结时只使用 `pages[]` 中 `usable_for_summary: true` 的文本。`needs_ocr: false` 时不得为了“常规检查”继续 OCR、渲染或调用图片识别。
|
- 阅读或总结时只使用 `pages[]` 中 `usable_for_summary: true` 的文本。`needs_ocr: false` 时不得为了“常规检查”继续 OCR、渲染或调用图片识别。
|
||||||
|
- 远程 PDF 源文件只交给 `download_pdf.py`;任务所需的远程图片、视频、音频、压缩包或其他附件只交给 `download_attachment.py`。不要在回复、日志摘要或文件名中复述可能含敏感查询参数的完整 URL。
|
||||||
- PDF 最终文件一律写入 `/usr/local/src/pdf/`,下载缓存、中间文件和渲染结果一律写入 `/usr/local/src/pdf/tmp/<任务名>/`。始终传绝对路径;固定脚本会自动创建目录并拒绝该根目录之外的输出。
|
- PDF 最终文件一律写入 `/usr/local/src/pdf/`,下载缓存、中间文件和渲染结果一律写入 `/usr/local/src/pdf/tmp/<任务名>/`。始终传绝对路径;固定脚本会自动创建目录并拒绝该根目录之外的输出。
|
||||||
- 不把 PDF 密码作为脚本参数;工具调用参数可能进入运行日志。
|
- 不把 PDF 密码作为脚本参数;工具调用参数可能进入运行日志。
|
||||||
|
|
||||||
@ -24,6 +25,7 @@ description: "处理本地 PDF 文件或远程 HTTPS PDF 链接,包括安全
|
|||||||
| 脚本 | 用途 | 底层能力 |
|
| 脚本 | 用途 | 底层能力 |
|
||||||
| --- | --- | --- |
|
| --- | --- | --- |
|
||||||
| `scripts/download_pdf.py` | 下载并校验远程 HTTPS PDF | `urllib`、`pypdf` |
|
| `scripts/download_pdf.py` | 下载并校验远程 HTTPS PDF | `urllib`、`pypdf` |
|
||||||
|
| `scripts/download_attachment.py` | 下载图片、音视频、压缩包等通用 HTTPS 附件 | `urllib`、HEAD 大小探测、流式硬限制 |
|
||||||
| `scripts/inspect_pdf.py` | 检查页数、加密、元数据、页面尺寸和表单数量 | `pypdf` |
|
| `scripts/inspect_pdf.py` | 检查页数、加密、元数据、页面尺寸和表单数量 | `pypdf` |
|
||||||
| `scripts/extract_text.py` | 多引擎提取、质量检测并分段返回正文 | Poppler `pdftotext`、`pdfplumber`;`pypdf` 校验 |
|
| `scripts/extract_text.py` | 多引擎提取、质量检测并分段返回正文 | Poppler `pdftotext`、`pdfplumber`;`pypdf` 校验 |
|
||||||
| `scripts/ocr_text.py` | 对指定扫描页执行离线 OCR 并返回可靠文字 | RapidOCR、ONNX Runtime、Poppler `pdftoppm` |
|
| `scripts/ocr_text.py` | 对指定扫描页执行离线 OCR 并返回可靠文字 | RapidOCR、ONNX Runtime、Poppler `pdftoppm` |
|
||||||
@ -60,11 +62,23 @@ description: "处理本地 PDF 文件或远程 HTTPS PDF 链接,包括安全
|
|||||||
可选参数:
|
可选参数:
|
||||||
|
|
||||||
- `--timeout <秒>`:默认 `60`。
|
- `--timeout <秒>`:默认 `60`。
|
||||||
- `--max-bytes <字节数>`:默认 `104857600`(100 MiB)。
|
- `--max-bytes <字节数>`:默认且最高 `26214400`(25 MiB),只允许设置更小的限制。
|
||||||
- `--overwrite`:仅在目标是本次任务生成的缓存时使用。
|
- `--overwrite`:仅在目标是本次任务生成的缓存时使用。
|
||||||
|
|
||||||
脚本会创建父目录、流式下载、阻止 HTTPS 重定向降级到 HTTP,并验证 PDF。成功结果包含 `path`、`size_bytes`、`page_count` 和 `encrypted`。
|
脚本会创建父目录、流式下载、阻止 HTTPS 重定向降级到 HTTP,并验证 PDF。成功结果包含 `path`、`size_bytes`、`page_count` 和 `encrypted`。
|
||||||
|
|
||||||
|
## 下载通用附件
|
||||||
|
|
||||||
|
需要下载作为 PDF 任务素材的图片、视频、音频、压缩包或其他文件时,调用 `scripts/download_attachment.py`:
|
||||||
|
|
||||||
|
```text
|
||||||
|
--url 'https://example.com/asset.bin?signature=...' --output '/usr/local/src/pdf/tmp/<任务名>/asset.bin'
|
||||||
|
```
|
||||||
|
|
||||||
|
只接受 HTTPS 地址,`output` 可使用任意附件扩展名。可选参数只有 `--timeout <1-600>`(默认 `60`)和 `--overwrite`。附件上限固定为 25 MiB(26214400 字节),不可调高:脚本先用 HEAD 探测远端声明大小,再检查 GET 响应声明,并在流式接收时持续兜底计数;任一阶段发现超限都会返回 `ok: false` 和明确的“已拒绝下载”错误,且不会发布部分文件。
|
||||||
|
|
||||||
|
成功结果包含 `path`、实际 `size_bytes`、`declared_size_bytes`、`size_limit_bytes`、`size_probe` 和 `content_type`。本脚本不校验文件业务格式;远程 PDF 源文件仍使用 `download_pdf.py`。
|
||||||
|
|
||||||
## 检查 PDF
|
## 检查 PDF
|
||||||
|
|
||||||
调用 `scripts/inspect_pdf.py`:
|
调用 `scripts/inspect_pdf.py`:
|
||||||
|
|||||||
266
skills/pdf/scripts/download_attachment.py
Normal file
266
skills/pdf/scripts/download_attachment.py
Normal file
@ -0,0 +1,266 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import os
|
||||||
|
import socket
|
||||||
|
import sys
|
||||||
|
import tempfile
|
||||||
|
import urllib.error
|
||||||
|
import urllib.parse
|
||||||
|
import urllib.request
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any, NoReturn, Optional
|
||||||
|
|
||||||
|
from _pdf_common import emit, ensure_output_path, failure_message, publish_temp_file
|
||||||
|
|
||||||
|
|
||||||
|
DEFAULT_TIMEOUT_SECONDS = 60
|
||||||
|
MAX_ATTACHMENT_BYTES = 25 * 1024 * 1024
|
||||||
|
CHUNK_SIZE = 1024 * 1024
|
||||||
|
USER_AGENT = "wechat-robot-pdf-attachment-downloader/1.0"
|
||||||
|
|
||||||
|
|
||||||
|
class SkillArgumentParser(argparse.ArgumentParser):
|
||||||
|
def error(self, message: str) -> NoReturn:
|
||||||
|
raise ValueError(f"参数错误:{message}")
|
||||||
|
|
||||||
|
|
||||||
|
def _validate_https_url(value: str) -> str:
|
||||||
|
url = value.strip()
|
||||||
|
if not url:
|
||||||
|
raise ValueError("附件 URL 不能为空")
|
||||||
|
if any(character.isspace() or ord(character) < 32 for character in url):
|
||||||
|
raise ValueError("附件 URL 不能包含空白字符或控制字符")
|
||||||
|
|
||||||
|
parsed = urllib.parse.urlsplit(url)
|
||||||
|
if parsed.scheme.lower() != "https" or not parsed.hostname:
|
||||||
|
raise ValueError("附件 URL 必须是有效的 HTTPS 地址")
|
||||||
|
if parsed.username is not None or parsed.password is not None:
|
||||||
|
raise ValueError("附件 URL 不允许包含用户名或密码")
|
||||||
|
try:
|
||||||
|
parsed.port
|
||||||
|
except ValueError as exc:
|
||||||
|
raise ValueError("附件 URL 端口格式不正确") from exc
|
||||||
|
return url
|
||||||
|
|
||||||
|
|
||||||
|
class HTTPSOnlyRedirectHandler(urllib.request.HTTPRedirectHandler):
|
||||||
|
def redirect_request(self, req, fp, code, msg, headers, newurl):
|
||||||
|
return super().redirect_request(
|
||||||
|
req,
|
||||||
|
fp,
|
||||||
|
code,
|
||||||
|
msg,
|
||||||
|
headers,
|
||||||
|
_validate_https_url(newurl),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _output_file(value: str, *, overwrite: bool) -> Path:
|
||||||
|
path = ensure_output_path(Path(value))
|
||||||
|
if path.exists() and not path.is_file():
|
||||||
|
raise ValueError(f"目标路径不是文件:{path}")
|
||||||
|
if path.exists() and not overwrite:
|
||||||
|
raise FileExistsError(f"目标文件已存在:{path}")
|
||||||
|
path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
return path
|
||||||
|
|
||||||
|
|
||||||
|
def _parse_args(argv: list[str]) -> argparse.Namespace:
|
||||||
|
parser = SkillArgumentParser(
|
||||||
|
description="下载不超过 25 MiB 的远程 HTTPS 通用附件"
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--url",
|
||||||
|
"--attachment-url",
|
||||||
|
"--attachment_url",
|
||||||
|
dest="url",
|
||||||
|
required=True,
|
||||||
|
help="远程 HTTPS 附件地址",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--output",
|
||||||
|
required=True,
|
||||||
|
help="附件本地输出路径;允许图片、音视频、压缩包及其他文件类型",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--timeout",
|
||||||
|
type=int,
|
||||||
|
default=DEFAULT_TIMEOUT_SECONDS,
|
||||||
|
help=f"连接和读取超时秒数,默认 {DEFAULT_TIMEOUT_SECONDS}",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--overwrite",
|
||||||
|
action="store_true",
|
||||||
|
help="允许覆盖本次任务已存在的缓存文件",
|
||||||
|
)
|
||||||
|
args = parser.parse_args(argv)
|
||||||
|
|
||||||
|
args.url = _validate_https_url(args.url)
|
||||||
|
if args.timeout < 1 or args.timeout > 600:
|
||||||
|
raise ValueError("timeout 必须在 1 到 600 秒之间")
|
||||||
|
args.output = _output_file(args.output, overwrite=args.overwrite)
|
||||||
|
return args
|
||||||
|
|
||||||
|
|
||||||
|
def _content_length(headers: Any) -> Optional[int]:
|
||||||
|
raw_value = headers.get("Content-Length")
|
||||||
|
if raw_value is None:
|
||||||
|
return None
|
||||||
|
try:
|
||||||
|
size = int(raw_value)
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
return None
|
||||||
|
return size if size >= 0 else None
|
||||||
|
|
||||||
|
|
||||||
|
def _reject_if_too_large(size_bytes: int, *, source: str) -> None:
|
||||||
|
if size_bytes <= MAX_ATTACHMENT_BYTES:
|
||||||
|
return
|
||||||
|
raise ValueError(
|
||||||
|
"附件超过 25 MiB(26214400 字节)限制,已拒绝下载;"
|
||||||
|
f"{source}大小为 {size_bytes} 字节"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _probe_size(
|
||||||
|
opener: urllib.request.OpenerDirector,
|
||||||
|
url: str,
|
||||||
|
timeout: int,
|
||||||
|
) -> Optional[int]:
|
||||||
|
request = urllib.request.Request(
|
||||||
|
url,
|
||||||
|
headers={
|
||||||
|
"Accept": "*/*",
|
||||||
|
"Accept-Encoding": "identity",
|
||||||
|
"User-Agent": USER_AGENT,
|
||||||
|
},
|
||||||
|
method="HEAD",
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
with opener.open(request, timeout=timeout) as response:
|
||||||
|
_validate_https_url(response.geturl())
|
||||||
|
size = _content_length(response.headers)
|
||||||
|
except (
|
||||||
|
urllib.error.HTTPError,
|
||||||
|
urllib.error.URLError,
|
||||||
|
TimeoutError,
|
||||||
|
socket.timeout,
|
||||||
|
):
|
||||||
|
return None
|
||||||
|
|
||||||
|
if size is not None:
|
||||||
|
_reject_if_too_large(size, source="远端声明")
|
||||||
|
return size
|
||||||
|
|
||||||
|
|
||||||
|
def _content_type(headers: Any) -> Optional[str]:
|
||||||
|
raw_value = headers.get("Content-Type")
|
||||||
|
if not raw_value:
|
||||||
|
return None
|
||||||
|
media_type = raw_value.split(";", 1)[0].strip().lower()
|
||||||
|
return media_type or None
|
||||||
|
|
||||||
|
|
||||||
|
def _download(args: argparse.Namespace) -> dict[str, Any]:
|
||||||
|
output: Path = args.output
|
||||||
|
opener = urllib.request.build_opener(HTTPSOnlyRedirectHandler())
|
||||||
|
probed_size = _probe_size(opener, args.url, args.timeout)
|
||||||
|
request = urllib.request.Request(
|
||||||
|
args.url,
|
||||||
|
headers={
|
||||||
|
"Accept": "*/*",
|
||||||
|
"Accept-Encoding": "identity",
|
||||||
|
"User-Agent": USER_AGENT,
|
||||||
|
},
|
||||||
|
method="GET",
|
||||||
|
)
|
||||||
|
|
||||||
|
temp_path: Optional[Path] = None
|
||||||
|
downloaded_bytes = 0
|
||||||
|
response_size: Optional[int] = None
|
||||||
|
response_type: Optional[str] = None
|
||||||
|
try:
|
||||||
|
with tempfile.NamedTemporaryFile(
|
||||||
|
mode="wb",
|
||||||
|
prefix=f".{output.stem}.",
|
||||||
|
suffix=f".part{output.suffix}",
|
||||||
|
dir=str(output.parent),
|
||||||
|
delete=False,
|
||||||
|
) as temp_file:
|
||||||
|
temp_path = Path(temp_file.name)
|
||||||
|
with opener.open(request, timeout=args.timeout) as response:
|
||||||
|
_validate_https_url(response.geturl())
|
||||||
|
response_size = _content_length(response.headers)
|
||||||
|
response_type = _content_type(response.headers)
|
||||||
|
if response_size is not None:
|
||||||
|
_reject_if_too_large(response_size, source="下载响应声明")
|
||||||
|
|
||||||
|
while True:
|
||||||
|
chunk = response.read(CHUNK_SIZE)
|
||||||
|
if not chunk:
|
||||||
|
break
|
||||||
|
downloaded_bytes += len(chunk)
|
||||||
|
_reject_if_too_large(downloaded_bytes, source="已接收")
|
||||||
|
temp_file.write(chunk)
|
||||||
|
|
||||||
|
temp_file.flush()
|
||||||
|
os.fsync(temp_file.fileno())
|
||||||
|
|
||||||
|
if downloaded_bytes == 0:
|
||||||
|
raise ValueError("远程服务器返回了空附件")
|
||||||
|
|
||||||
|
publish_temp_file(temp_path, output, args.overwrite)
|
||||||
|
temp_path = None
|
||||||
|
declared_size = (
|
||||||
|
response_size if response_size is not None else probed_size
|
||||||
|
)
|
||||||
|
if response_size is not None:
|
||||||
|
size_probe = "get-content-length"
|
||||||
|
elif probed_size is not None:
|
||||||
|
size_probe = "head-content-length"
|
||||||
|
else:
|
||||||
|
size_probe = "stream"
|
||||||
|
return {
|
||||||
|
"path": str(output),
|
||||||
|
"size_bytes": downloaded_bytes,
|
||||||
|
"declared_size_bytes": declared_size,
|
||||||
|
"size_limit_bytes": MAX_ATTACHMENT_BYTES,
|
||||||
|
"size_probe": size_probe,
|
||||||
|
"content_type": response_type,
|
||||||
|
}
|
||||||
|
finally:
|
||||||
|
if temp_path is not None:
|
||||||
|
try:
|
||||||
|
temp_path.unlink(missing_ok=True)
|
||||||
|
except OSError:
|
||||||
|
pass
|
||||||
|
|
||||||
|
|
||||||
|
def _failure_message(exc: Exception) -> str:
|
||||||
|
if isinstance(exc, urllib.error.HTTPError):
|
||||||
|
return f"附件下载失败:远程服务器返回 HTTP {exc.code}"
|
||||||
|
if isinstance(exc, (TimeoutError, socket.timeout)):
|
||||||
|
return "附件下载失败:连接或读取超时"
|
||||||
|
if isinstance(exc, urllib.error.URLError):
|
||||||
|
if isinstance(exc.reason, (TimeoutError, socket.timeout)):
|
||||||
|
return "附件下载失败:连接或读取超时"
|
||||||
|
return "附件下载失败:无法访问远程服务器"
|
||||||
|
return failure_message(exc)
|
||||||
|
|
||||||
|
|
||||||
|
def main(argv: Optional[list[str]] = None) -> int:
|
||||||
|
try:
|
||||||
|
args = _parse_args(sys.argv[1:] if argv is None else argv)
|
||||||
|
result = _download(args)
|
||||||
|
except Exception as exc:
|
||||||
|
emit({"ok": False, "error": _failure_message(exc)})
|
||||||
|
return 1
|
||||||
|
emit({"ok": True, **result})
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
@ -18,7 +18,8 @@ from typing import NoReturn, Optional
|
|||||||
from _pdf_common import ensure_output_path
|
from _pdf_common import ensure_output_path
|
||||||
|
|
||||||
DEFAULT_TIMEOUT_SECONDS = 60
|
DEFAULT_TIMEOUT_SECONDS = 60
|
||||||
DEFAULT_MAX_BYTES = 100 * 1024 * 1024
|
MAX_ALLOWED_BYTES = 25 * 1024 * 1024
|
||||||
|
DEFAULT_MAX_BYTES = MAX_ALLOWED_BYTES
|
||||||
CHUNK_SIZE = 1024 * 1024
|
CHUNK_SIZE = 1024 * 1024
|
||||||
USER_AGENT = "wechat-robot-pdf-skill/1.0"
|
USER_AGENT = "wechat-robot-pdf-skill/1.0"
|
||||||
|
|
||||||
@ -96,8 +97,10 @@ def _parse_args(argv: list[str]) -> argparse.Namespace:
|
|||||||
args.url = _validate_https_url(args.url)
|
args.url = _validate_https_url(args.url)
|
||||||
if args.timeout <= 0:
|
if args.timeout <= 0:
|
||||||
raise ValueError("timeout 必须大于 0")
|
raise ValueError("timeout 必须大于 0")
|
||||||
if args.max_bytes <= 0:
|
if args.max_bytes <= 0 or args.max_bytes > MAX_ALLOWED_BYTES:
|
||||||
raise ValueError("max-bytes 必须大于 0")
|
raise ValueError(
|
||||||
|
f"max-bytes 必须在 1 到 {MAX_ALLOWED_BYTES} 之间"
|
||||||
|
)
|
||||||
|
|
||||||
output = Path(args.output).expanduser()
|
output = Path(args.output).expanduser()
|
||||||
if output.suffix.lower() != ".pdf":
|
if output.suffix.lower() != ".pdf":
|
||||||
@ -169,7 +172,8 @@ def _download(args: argparse.Namespace) -> dict:
|
|||||||
expected_bytes = 0
|
expected_bytes = 0
|
||||||
if expected_bytes > args.max_bytes:
|
if expected_bytes > args.max_bytes:
|
||||||
raise ValueError(
|
raise ValueError(
|
||||||
f"远程文件超过大小限制:最多允许 {args.max_bytes} 字节"
|
"远程文件超过大小限制,已拒绝下载:"
|
||||||
|
f"最多允许 {args.max_bytes} 字节"
|
||||||
)
|
)
|
||||||
|
|
||||||
downloaded_bytes = 0
|
downloaded_bytes = 0
|
||||||
@ -180,7 +184,8 @@ def _download(args: argparse.Namespace) -> dict:
|
|||||||
downloaded_bytes += len(chunk)
|
downloaded_bytes += len(chunk)
|
||||||
if downloaded_bytes > args.max_bytes:
|
if downloaded_bytes > args.max_bytes:
|
||||||
raise ValueError(
|
raise ValueError(
|
||||||
f"远程文件超过大小限制:最多允许 {args.max_bytes} 字节"
|
"远程文件超过大小限制,已拒绝下载:"
|
||||||
|
f"最多允许 {args.max_bytes} 字节"
|
||||||
)
|
)
|
||||||
temp_file.write(chunk)
|
temp_file.write(chunk)
|
||||||
|
|
||||||
|
|||||||
@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
name: pptx
|
name: pptx
|
||||||
description: "创建、读取、编辑、复制页面、转换、校验和渲染本地或远程 HTTPS PowerPoint 演示文稿与模板,并按需识别页面图片、截图和图表中的文字。用户提到 PPT、PPTX、PowerPoint、演示文稿、幻灯片、路演稿、汇报材料、演讲者备注、模板、版式或图片文字 OCR,或提供 HTTPS 演示文稿地址、.pptx、.potx、.ppsx、.ppt 文件时使用;支持安全下载、结构化创建、保留 Run 格式的文本替换、页面删除/重排/复制、Markdown 提取、本地 RapidOCR、旧格式转换、受控 OOXML 解包/打包、关系与图表校验及逐页视觉检查。若主要交付物不是演示文稿且不需要读取或修改 PPT 内容,则不要使用。"
|
description: "创建、读取、编辑、复制页面、转换、校验和渲染本地或远程 HTTPS PowerPoint 演示文稿与模板,并下载 PPT 任务所需且不超过 25 MiB 的图片、音视频、压缩包和其他 HTTPS 附件,按需识别页面图片、截图和图表中的文字。用户提到 PPT、PPTX、PowerPoint、演示文稿、幻灯片、路演稿、汇报材料、演讲者备注、模板、版式或图片文字 OCR,或提供 HTTPS 演示文稿地址、.pptx、.potx、.ppsx、.ppt 文件时使用;支持源演示文稿安全下载、结构化创建、保留 Run 格式的文本替换、页面删除/重排/复制、Markdown 提取、本地 RapidOCR、旧格式转换、受控 OOXML 解包/打包、关系与图表校验及逐页视觉检查。若主要交付物不是演示文稿且不需要读取或修改 PPT 内容,则不要使用。"
|
||||||
---
|
---
|
||||||
|
|
||||||
# PowerPoint 演示文稿处理
|
# PowerPoint 演示文稿处理
|
||||||
@ -15,7 +15,7 @@ description: "创建、读取、编辑、复制页面、转换、校验和渲染
|
|||||||
- 每次检查脚本返回的 JSON;只有 `ok` 为 `true` 时才继续。`validate_presentation.py` 还必须返回 `status: valid`、`issue_count: 0`。
|
- 每次检查脚本返回的 JSON;只有 `ok` 为 `true` 时才继续。`validate_presentation.py` 还必须返回 `status: valid`、`issue_count: 0`。
|
||||||
- 只在需要读取图片、截图或视觉图表中的文字时调用 `ocr_presentation.py`。只使用 `slides[]` 中 `usable_for_summary: true` 的 `text`;低置信度结果不得作为可靠正文。
|
- 只在需要读取图片、截图或视觉图表中的文字时调用 `ocr_presentation.py`。只使用 `slides[]` 中 `usable_for_summary: true` 的 `text`;低置信度结果不得作为可靠正文。
|
||||||
- 不覆盖用户提供的源文件。PPT 最终文件一律写入 `/usr/local/src/ppt/`,下载缓存、中间文件和渲染结果一律写入 `/usr/local/src/ppt/tmp/<任务名>/`。始终传绝对路径;固定脚本会自动创建目录并拒绝该根目录之外的输出。
|
- 不覆盖用户提供的源文件。PPT 最终文件一律写入 `/usr/local/src/ppt/`,下载缓存、中间文件和渲染结果一律写入 `/usr/local/src/ppt/tmp/<任务名>/`。始终传绝对路径;固定脚本会自动创建目录并拒绝该根目录之外的输出。
|
||||||
- 远程地址只交给 `download_presentation.py`;不要在回复、日志摘要或文件名中复述可能含敏感查询参数的完整 URL。
|
- 远程 PowerPoint 源文件只交给 `download_presentation.py`;任务所需的远程图片、视频、音频、压缩包或其他附件只交给 `download_attachment.py`。不要在回复、日志摘要或文件名中复述可能含敏感查询参数的完整 URL。
|
||||||
- 环境已预置全部依赖,不安装软件包,也不提示用户安装依赖。
|
- 环境已预置全部依赖,不安装软件包,也不提示用户安装依赖。
|
||||||
|
|
||||||
## 脚本清单
|
## 脚本清单
|
||||||
@ -23,6 +23,7 @@ description: "创建、读取、编辑、复制页面、转换、校验和渲染
|
|||||||
| 脚本 | 用途 | 底层能力 |
|
| 脚本 | 用途 | 底层能力 |
|
||||||
| --- | --- | --- |
|
| --- | --- | --- |
|
||||||
| `scripts/download_presentation.py` | 下载并校验远程 HTTPS 演示文稿 | Python `urllib`、安全 OOXML 解析、`python-pptx` |
|
| `scripts/download_presentation.py` | 下载并校验远程 HTTPS 演示文稿 | Python `urllib`、安全 OOXML 解析、`python-pptx` |
|
||||||
|
| `scripts/download_attachment.py` | 下载图片、音视频、压缩包等通用 HTTPS 附件 | Python `urllib`、HEAD 大小探测、流式硬限制 |
|
||||||
| `scripts/inspect_presentation.py` | 分段读取页面、文本、表格、图表、图片和备注 | `python-pptx`、安全 OOXML 解析 |
|
| `scripts/inspect_presentation.py` | 分段读取页面、文本、表格、图表、图片和备注 | `python-pptx`、安全 OOXML 解析 |
|
||||||
| `scripts/ocr_presentation.py` | 按页识别图片、截图和视觉图表中的文字 | RapidOCR、ONNX Runtime、LibreOffice、Poppler |
|
| `scripts/ocr_presentation.py` | 按页识别图片、截图和视觉图表中的文字 | RapidOCR、ONNX Runtime、LibreOffice、Poppler |
|
||||||
| `scripts/extract_presentation.py` | 提取整份演示文稿为 Markdown | `markitdown[pptx]` |
|
| `scripts/extract_presentation.py` | 提取整份演示文稿为 Markdown | `markitdown[pptx]` |
|
||||||
@ -60,11 +61,23 @@ description: "创建、读取、编辑、复制页面、转换、校验和渲染
|
|||||||
可选参数:
|
可选参数:
|
||||||
|
|
||||||
- `--timeout <1-600>`:连接和读取超时秒数,默认 `60`。
|
- `--timeout <1-600>`:连接和读取超时秒数,默认 `60`。
|
||||||
- `--max-bytes <字节数>`:默认 100 MiB,最高 512 MiB。
|
- `--max-bytes <字节数>`:默认且最高 25 MiB(26214400 字节),只允许设置更小的限制。
|
||||||
- `--overwrite`:只覆盖本次任务生成的旧缓存。
|
- `--overwrite`:只覆盖本次任务生成的旧缓存。
|
||||||
|
|
||||||
`output` 扩展名必须与远程内容的真实格式一致,支持 `.pptx/.potx/.ppsx/.ppt`。脚本限制重定向只能继续使用 HTTPS,流式限制大小,先写临时文件,再原子发布。
|
`output` 扩展名必须与远程内容的真实格式一致,支持 `.pptx/.potx/.ppsx/.ppt`。脚本限制重定向只能继续使用 HTTPS,流式限制大小,先写临时文件,再原子发布。
|
||||||
|
|
||||||
|
## 下载通用附件
|
||||||
|
|
||||||
|
需要下载作为 PowerPoint 任务素材的图片、视频、音频、压缩包或其他文件时,调用 `scripts/download_attachment.py`:
|
||||||
|
|
||||||
|
```text
|
||||||
|
--url 'https://example.com/asset.bin?signature=...' --output '/usr/local/src/ppt/tmp/<任务名>/asset.bin'
|
||||||
|
```
|
||||||
|
|
||||||
|
只接受 HTTPS 地址,`output` 可使用任意附件扩展名。可选参数只有 `--timeout <1-600>`(默认 `60`)和 `--overwrite`。附件上限固定为 25 MiB(26214400 字节),不可调高:脚本先用 HEAD 探测远端声明大小,再检查 GET 响应声明,并在流式接收时持续兜底计数;任一阶段发现超限都会返回 `ok: false` 和明确的“已拒绝下载”错误,且不会发布部分文件。
|
||||||
|
|
||||||
|
成功结果包含 `path`、实际 `size_bytes`、`declared_size_bytes`、`size_limit_bytes`、`size_probe` 和 `content_type`。本脚本不校验文件业务格式;远程 PowerPoint 源文件仍使用 `download_presentation.py`。
|
||||||
|
|
||||||
## 检查和提取内容
|
## 检查和提取内容
|
||||||
|
|
||||||
调用结构检查:
|
调用结构检查:
|
||||||
|
|||||||
256
skills/pptx/scripts/download_attachment.py
Normal file
256
skills/pptx/scripts/download_attachment.py
Normal file
@ -0,0 +1,256 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import os
|
||||||
|
import socket
|
||||||
|
import sys
|
||||||
|
import tempfile
|
||||||
|
import urllib.error
|
||||||
|
import urllib.parse
|
||||||
|
import urllib.request
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any, NoReturn, Optional
|
||||||
|
|
||||||
|
from _pptx_common import emit, failure_message, output_file, publish_file
|
||||||
|
|
||||||
|
|
||||||
|
DEFAULT_TIMEOUT_SECONDS = 60
|
||||||
|
MAX_ATTACHMENT_BYTES = 25 * 1024 * 1024
|
||||||
|
CHUNK_SIZE = 1024 * 1024
|
||||||
|
USER_AGENT = "wechat-robot-pptx-attachment-downloader/1.0"
|
||||||
|
|
||||||
|
|
||||||
|
class SkillArgumentParser(argparse.ArgumentParser):
|
||||||
|
def error(self, message: str) -> NoReturn:
|
||||||
|
raise ValueError(f"参数错误:{message}")
|
||||||
|
|
||||||
|
|
||||||
|
def _validate_https_url(value: str) -> str:
|
||||||
|
url = value.strip()
|
||||||
|
if not url:
|
||||||
|
raise ValueError("附件 URL 不能为空")
|
||||||
|
if any(character.isspace() or ord(character) < 32 for character in url):
|
||||||
|
raise ValueError("附件 URL 不能包含空白字符或控制字符")
|
||||||
|
|
||||||
|
parsed = urllib.parse.urlsplit(url)
|
||||||
|
if parsed.scheme.lower() != "https" or not parsed.hostname:
|
||||||
|
raise ValueError("附件 URL 必须是有效的 HTTPS 地址")
|
||||||
|
if parsed.username is not None or parsed.password is not None:
|
||||||
|
raise ValueError("附件 URL 不允许包含用户名或密码")
|
||||||
|
try:
|
||||||
|
parsed.port
|
||||||
|
except ValueError as exc:
|
||||||
|
raise ValueError("附件 URL 端口格式不正确") from exc
|
||||||
|
return url
|
||||||
|
|
||||||
|
|
||||||
|
class HTTPSOnlyRedirectHandler(urllib.request.HTTPRedirectHandler):
|
||||||
|
def redirect_request(self, req, fp, code, msg, headers, newurl):
|
||||||
|
return super().redirect_request(
|
||||||
|
req,
|
||||||
|
fp,
|
||||||
|
code,
|
||||||
|
msg,
|
||||||
|
headers,
|
||||||
|
_validate_https_url(newurl),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _parse_args(argv: list[str]) -> argparse.Namespace:
|
||||||
|
parser = SkillArgumentParser(
|
||||||
|
description="下载不超过 25 MiB 的远程 HTTPS 通用附件"
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--url",
|
||||||
|
"--attachment-url",
|
||||||
|
"--attachment_url",
|
||||||
|
dest="url",
|
||||||
|
required=True,
|
||||||
|
help="远程 HTTPS 附件地址",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--output",
|
||||||
|
required=True,
|
||||||
|
help="附件本地输出路径;允许图片、音视频、压缩包及其他文件类型",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--timeout",
|
||||||
|
type=int,
|
||||||
|
default=DEFAULT_TIMEOUT_SECONDS,
|
||||||
|
help=f"连接和读取超时秒数,默认 {DEFAULT_TIMEOUT_SECONDS}",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--overwrite",
|
||||||
|
action="store_true",
|
||||||
|
help="允许覆盖本次任务已存在的缓存文件",
|
||||||
|
)
|
||||||
|
args = parser.parse_args(argv)
|
||||||
|
|
||||||
|
args.url = _validate_https_url(args.url)
|
||||||
|
if args.timeout < 1 or args.timeout > 600:
|
||||||
|
raise ValueError("timeout 必须在 1 到 600 秒之间")
|
||||||
|
args.output = output_file(args.output, overwrite=args.overwrite)
|
||||||
|
return args
|
||||||
|
|
||||||
|
|
||||||
|
def _content_length(headers: Any) -> Optional[int]:
|
||||||
|
raw_value = headers.get("Content-Length")
|
||||||
|
if raw_value is None:
|
||||||
|
return None
|
||||||
|
try:
|
||||||
|
size = int(raw_value)
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
return None
|
||||||
|
return size if size >= 0 else None
|
||||||
|
|
||||||
|
|
||||||
|
def _reject_if_too_large(size_bytes: int, *, source: str) -> None:
|
||||||
|
if size_bytes <= MAX_ATTACHMENT_BYTES:
|
||||||
|
return
|
||||||
|
raise ValueError(
|
||||||
|
"附件超过 25 MiB(26214400 字节)限制,已拒绝下载;"
|
||||||
|
f"{source}大小为 {size_bytes} 字节"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _probe_size(
|
||||||
|
opener: urllib.request.OpenerDirector,
|
||||||
|
url: str,
|
||||||
|
timeout: int,
|
||||||
|
) -> Optional[int]:
|
||||||
|
request = urllib.request.Request(
|
||||||
|
url,
|
||||||
|
headers={
|
||||||
|
"Accept": "*/*",
|
||||||
|
"Accept-Encoding": "identity",
|
||||||
|
"User-Agent": USER_AGENT,
|
||||||
|
},
|
||||||
|
method="HEAD",
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
with opener.open(request, timeout=timeout) as response:
|
||||||
|
_validate_https_url(response.geturl())
|
||||||
|
size = _content_length(response.headers)
|
||||||
|
except (
|
||||||
|
urllib.error.HTTPError,
|
||||||
|
urllib.error.URLError,
|
||||||
|
TimeoutError,
|
||||||
|
socket.timeout,
|
||||||
|
):
|
||||||
|
return None
|
||||||
|
|
||||||
|
if size is not None:
|
||||||
|
_reject_if_too_large(size, source="远端声明")
|
||||||
|
return size
|
||||||
|
|
||||||
|
|
||||||
|
def _content_type(headers: Any) -> Optional[str]:
|
||||||
|
raw_value = headers.get("Content-Type")
|
||||||
|
if not raw_value:
|
||||||
|
return None
|
||||||
|
media_type = raw_value.split(";", 1)[0].strip().lower()
|
||||||
|
return media_type or None
|
||||||
|
|
||||||
|
|
||||||
|
def _download(args: argparse.Namespace) -> dict[str, Any]:
|
||||||
|
output: Path = args.output
|
||||||
|
opener = urllib.request.build_opener(HTTPSOnlyRedirectHandler())
|
||||||
|
probed_size = _probe_size(opener, args.url, args.timeout)
|
||||||
|
request = urllib.request.Request(
|
||||||
|
args.url,
|
||||||
|
headers={
|
||||||
|
"Accept": "*/*",
|
||||||
|
"Accept-Encoding": "identity",
|
||||||
|
"User-Agent": USER_AGENT,
|
||||||
|
},
|
||||||
|
method="GET",
|
||||||
|
)
|
||||||
|
|
||||||
|
temp_path: Optional[Path] = None
|
||||||
|
downloaded_bytes = 0
|
||||||
|
response_size: Optional[int] = None
|
||||||
|
response_type: Optional[str] = None
|
||||||
|
try:
|
||||||
|
with tempfile.NamedTemporaryFile(
|
||||||
|
mode="wb",
|
||||||
|
prefix=f".{output.stem}.",
|
||||||
|
suffix=f".part{output.suffix}",
|
||||||
|
dir=str(output.parent),
|
||||||
|
delete=False,
|
||||||
|
) as temp_file:
|
||||||
|
temp_path = Path(temp_file.name)
|
||||||
|
with opener.open(request, timeout=args.timeout) as response:
|
||||||
|
_validate_https_url(response.geturl())
|
||||||
|
response_size = _content_length(response.headers)
|
||||||
|
response_type = _content_type(response.headers)
|
||||||
|
if response_size is not None:
|
||||||
|
_reject_if_too_large(response_size, source="下载响应声明")
|
||||||
|
|
||||||
|
while True:
|
||||||
|
chunk = response.read(CHUNK_SIZE)
|
||||||
|
if not chunk:
|
||||||
|
break
|
||||||
|
downloaded_bytes += len(chunk)
|
||||||
|
_reject_if_too_large(downloaded_bytes, source="已接收")
|
||||||
|
temp_file.write(chunk)
|
||||||
|
|
||||||
|
temp_file.flush()
|
||||||
|
os.fsync(temp_file.fileno())
|
||||||
|
|
||||||
|
if downloaded_bytes == 0:
|
||||||
|
raise ValueError("远程服务器返回了空附件")
|
||||||
|
|
||||||
|
publish_file(temp_path, output, overwrite=args.overwrite)
|
||||||
|
temp_path = None
|
||||||
|
declared_size = (
|
||||||
|
response_size if response_size is not None else probed_size
|
||||||
|
)
|
||||||
|
if response_size is not None:
|
||||||
|
size_probe = "get-content-length"
|
||||||
|
elif probed_size is not None:
|
||||||
|
size_probe = "head-content-length"
|
||||||
|
else:
|
||||||
|
size_probe = "stream"
|
||||||
|
return {
|
||||||
|
"path": str(output),
|
||||||
|
"size_bytes": downloaded_bytes,
|
||||||
|
"declared_size_bytes": declared_size,
|
||||||
|
"size_limit_bytes": MAX_ATTACHMENT_BYTES,
|
||||||
|
"size_probe": size_probe,
|
||||||
|
"content_type": response_type,
|
||||||
|
}
|
||||||
|
finally:
|
||||||
|
if temp_path is not None:
|
||||||
|
try:
|
||||||
|
temp_path.unlink(missing_ok=True)
|
||||||
|
except OSError:
|
||||||
|
pass
|
||||||
|
|
||||||
|
|
||||||
|
def _failure_message(exc: Exception) -> str:
|
||||||
|
if isinstance(exc, urllib.error.HTTPError):
|
||||||
|
return f"附件下载失败:远程服务器返回 HTTP {exc.code}"
|
||||||
|
if isinstance(exc, (TimeoutError, socket.timeout)):
|
||||||
|
return "附件下载失败:连接或读取超时"
|
||||||
|
if isinstance(exc, urllib.error.URLError):
|
||||||
|
if isinstance(exc.reason, (TimeoutError, socket.timeout)):
|
||||||
|
return "附件下载失败:连接或读取超时"
|
||||||
|
return "附件下载失败:无法访问远程服务器"
|
||||||
|
return failure_message(exc)
|
||||||
|
|
||||||
|
|
||||||
|
def main(argv: Optional[list[str]] = None) -> int:
|
||||||
|
try:
|
||||||
|
args = _parse_args(sys.argv[1:] if argv is None else argv)
|
||||||
|
result = _download(args)
|
||||||
|
except Exception as exc:
|
||||||
|
emit({"ok": False, "error": _failure_message(exc)})
|
||||||
|
return 1
|
||||||
|
emit({"ok": True, **result})
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
@ -26,8 +26,8 @@ from _pptx_common import (
|
|||||||
|
|
||||||
|
|
||||||
DEFAULT_TIMEOUT_SECONDS = 60
|
DEFAULT_TIMEOUT_SECONDS = 60
|
||||||
DEFAULT_MAX_BYTES = 100 * 1024 * 1024
|
MAX_ALLOWED_BYTES = 25 * 1024 * 1024
|
||||||
MAX_ALLOWED_BYTES = 512 * 1024 * 1024
|
DEFAULT_MAX_BYTES = MAX_ALLOWED_BYTES
|
||||||
CHUNK_SIZE = 1024 * 1024
|
CHUNK_SIZE = 1024 * 1024
|
||||||
USER_AGENT = "wechat-robot-pptx-skill/1.0"
|
USER_AGENT = "wechat-robot-pptx-skill/1.0"
|
||||||
OLE_COMPOUND_MAGIC = bytes.fromhex("D0CF11E0A1B11AE1")
|
OLE_COMPOUND_MAGIC = bytes.fromhex("D0CF11E0A1B11AE1")
|
||||||
@ -212,7 +212,8 @@ def _download(args: argparse.Namespace) -> dict[str, Any]:
|
|||||||
expected_bytes = 0
|
expected_bytes = 0
|
||||||
if expected_bytes > args.max_bytes:
|
if expected_bytes > args.max_bytes:
|
||||||
raise ValueError(
|
raise ValueError(
|
||||||
f"远程文件超过大小限制:最多允许 {args.max_bytes} 字节"
|
"远程文件超过大小限制,已拒绝下载:"
|
||||||
|
f"最多允许 {args.max_bytes} 字节"
|
||||||
)
|
)
|
||||||
while True:
|
while True:
|
||||||
chunk = response.read(CHUNK_SIZE)
|
chunk = response.read(CHUNK_SIZE)
|
||||||
@ -221,7 +222,8 @@ def _download(args: argparse.Namespace) -> dict[str, Any]:
|
|||||||
downloaded_bytes += len(chunk)
|
downloaded_bytes += len(chunk)
|
||||||
if downloaded_bytes > args.max_bytes:
|
if downloaded_bytes > args.max_bytes:
|
||||||
raise ValueError(
|
raise ValueError(
|
||||||
f"远程文件超过大小限制:最多允许 {args.max_bytes} 字节"
|
"远程文件超过大小限制,已拒绝下载:"
|
||||||
|
f"最多允许 {args.max_bytes} 字节"
|
||||||
)
|
)
|
||||||
temp_file.write(chunk)
|
temp_file.write(chunk)
|
||||||
temp_file.flush()
|
temp_file.flush()
|
||||||
|
|||||||
@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
name: xlsx
|
name: xlsx
|
||||||
description: "创建、读取、编辑、修复、转换、重算、校验和渲染本地或远程 HTTPS Excel 工作簿及表格数据。用户提到 Excel、电子表格、工作簿、工作表、单元格、公式、图表、数据清洗,或提供 HTTPS Excel/CSV/TSV 地址、.xlsx、.xlsm、.xltx、.xls、.csv、.tsv 文件时使用;支持安全下载、保留现有样式、批量写入、公式与缓存值检查、表格/图表/图片/数据验证/条件格式、旧格式转换和逐页视觉检查。最终交付物必须是电子表格文件;若主要交付物是 Word、PDF、HTML、数据库程序或在线 Google Sheets,则不要使用。"
|
description: "创建、读取、编辑、修复、转换、重算、校验和渲染本地或远程 HTTPS Excel 工作簿及表格数据,并下载 Excel 任务所需且不超过 25 MiB 的图片、音视频、压缩包和其他 HTTPS 附件。用户提到 Excel、电子表格、工作簿、工作表、单元格、公式、图表、数据清洗,或提供 HTTPS Excel/CSV/TSV 地址、.xlsx、.xlsm、.xltx、.xls、.csv、.tsv 文件时使用;支持源文件安全下载、保留现有样式、批量写入、公式与缓存值检查、表格/图表/图片/数据验证/条件格式、旧格式转换和逐页视觉检查。最终交付物必须是电子表格文件;若主要交付物是 Word、PDF、HTML、数据库程序或在线 Google Sheets,则不要使用。"
|
||||||
---
|
---
|
||||||
|
|
||||||
# Excel 工作簿处理
|
# Excel 工作簿处理
|
||||||
@ -13,7 +13,7 @@ description: "创建、读取、编辑、修复、转换、重算、校验和渲
|
|||||||
- 不把 `python3`、`soffice`、`libreoffice`、`pdftoppm`、`zip`、`unzip`、`rm` 或其他系统命令作为脚本参数。
|
- 不把 `python3`、`soffice`、`libreoffice`、`pdftoppm`、`zip`、`unzip`、`rm` 或其他系统命令作为脚本参数。
|
||||||
- 外部程序只允许由固定脚本在内部以无 shell 参数数组方式调用。
|
- 外部程序只允许由固定脚本在内部以无 shell 参数数组方式调用。
|
||||||
- 每次检查脚本返回的 JSON;只有 `ok` 为 `true` 时才继续。`status: errors_found` 虽然表示脚本成功运行,但工作簿不合格,必须修复。
|
- 每次检查脚本返回的 JSON;只有 `ok` 为 `true` 时才继续。`status: errors_found` 虽然表示脚本成功运行,但工作簿不合格,必须修复。
|
||||||
- 远程地址只交给 `download_workbook.py`;不要在回复、日志摘要或文件名中复述可能含敏感查询参数的完整 URL。
|
- 远程 Excel/CSV/TSV 源文件只交给 `download_workbook.py`;任务所需的远程图片、视频、音频、压缩包或其他附件只交给 `download_attachment.py`。不要在回复、日志摘要或文件名中复述可能含敏感查询参数的完整 URL。
|
||||||
- 不覆盖用户提供的源文件。Excel 最终文件一律写入 `/usr/local/src/excel/`,下载缓存、中间文件和渲染结果一律写入 `/usr/local/src/excel/tmp/<任务名>/`。始终传绝对路径;固定脚本会自动创建目录并拒绝该根目录之外的输出。
|
- 不覆盖用户提供的源文件。Excel 最终文件一律写入 `/usr/local/src/excel/`,下载缓存、中间文件和渲染结果一律写入 `/usr/local/src/excel/tmp/<任务名>/`。始终传绝对路径;固定脚本会自动创建目录并拒绝该根目录之外的输出。
|
||||||
- 环境已预置依赖,不安装软件包,也不提示用户安装依赖。
|
- 环境已预置依赖,不安装软件包,也不提示用户安装依赖。
|
||||||
|
|
||||||
@ -22,6 +22,7 @@ description: "创建、读取、编辑、修复、转换、重算、校验和渲
|
|||||||
| 脚本 | 用途 | 底层能力 |
|
| 脚本 | 用途 | 底层能力 |
|
||||||
| --- | --- | --- |
|
| --- | --- | --- |
|
||||||
| `scripts/download_workbook.py` | 下载并校验远程 HTTPS Excel/CSV/TSV | Python `urllib`、安全 OOXML 解析、`openpyxl` |
|
| `scripts/download_workbook.py` | 下载并校验远程 HTTPS Excel/CSV/TSV | Python `urllib`、安全 OOXML 解析、`openpyxl` |
|
||||||
|
| `scripts/download_attachment.py` | 下载图片、音视频、压缩包等通用 HTTPS 附件 | Python `urllib`、HEAD 大小探测、流式硬限制 |
|
||||||
| `scripts/inspect_workbook.py` | 分段读取结构、公式、缓存值和样式 | `openpyxl`、Python `csv` |
|
| `scripts/inspect_workbook.py` | 分段读取结构、公式、缓存值和样式 | `openpyxl`、Python `csv` |
|
||||||
| `scripts/apply_workbook.py` | 按受控 JSON 创建或编辑工作簿 | `openpyxl`、Pillow |
|
| `scripts/apply_workbook.py` | 按受控 JSON 创建或编辑工作簿 | `openpyxl`、Pillow |
|
||||||
| `scripts/convert_workbook.py` | 转换 `.xls/.csv/.tsv/.xlsx/.xlsm/.xltx` | `openpyxl`、LibreOffice |
|
| `scripts/convert_workbook.py` | 转换 `.xls/.csv/.tsv/.xlsx/.xlsm/.xltx` | `openpyxl`、LibreOffice |
|
||||||
@ -52,13 +53,25 @@ description: "创建、读取、编辑、修复、转换、重算、校验和渲
|
|||||||
可选参数:
|
可选参数:
|
||||||
|
|
||||||
- `--timeout <1-600>`:连接和读取超时秒数,默认 `60`。
|
- `--timeout <1-600>`:连接和读取超时秒数,默认 `60`。
|
||||||
- `--max-bytes <字节数>`:默认 `104857600`(100 MiB),最高 `536870912`(512 MiB)。
|
- `--max-bytes <字节数>`:默认且最高 `26214400`(25 MiB),只允许设置更小的限制。
|
||||||
- `--overwrite`:只在目标是本次任务生成的旧缓存时使用。
|
- `--overwrite`:只在目标是本次任务生成的旧缓存时使用。
|
||||||
|
|
||||||
`output` 扩展名必须是 `.xlsx`、`.xlsm`、`.xltx`、`.xltm`、`.xls`、`.csv` 或 `.tsv`。脚本阻止 HTTPS 重定向降级到 HTTP,流式限制大小,先写同目录临时文件,再原子发布;OOXML 会检查 ZIP 路径、成员大小、内容类型并用 `openpyxl` 打开,CSV/TSV 会拒绝二进制或网页响应。实际 OOXML 格式与 `output` 扩展名不一致时,根据错误中的实际格式更正缓存扩展名,再调用同一脚本。
|
`output` 扩展名必须是 `.xlsx`、`.xlsm`、`.xltx`、`.xltm`、`.xls`、`.csv` 或 `.tsv`。脚本阻止 HTTPS 重定向降级到 HTTP,流式限制大小,先写同目录临时文件,再原子发布;OOXML 会检查 ZIP 路径、成员大小、内容类型并用 `openpyxl` 打开,CSV/TSV 会拒绝二进制或网页响应。实际 OOXML 格式与 `output` 扩展名不一致时,根据错误中的实际格式更正缓存扩展名,再调用同一脚本。
|
||||||
|
|
||||||
成功结果包含 `path`、`size_bytes`、`format` 和 `validation`;OOXML 还包含 `sheet_count`。后续脚本只使用返回的本地 `path`,不再访问原 URL。
|
成功结果包含 `path`、`size_bytes`、`format` 和 `validation`;OOXML 还包含 `sheet_count`。后续脚本只使用返回的本地 `path`,不再访问原 URL。
|
||||||
|
|
||||||
|
## 下载通用附件
|
||||||
|
|
||||||
|
需要下载作为 Excel 任务素材的图片、视频、音频、压缩包或其他文件时,调用 `scripts/download_attachment.py`:
|
||||||
|
|
||||||
|
```text
|
||||||
|
--url 'https://example.com/asset.bin?signature=...' --output '/usr/local/src/excel/tmp/<任务名>/asset.bin'
|
||||||
|
```
|
||||||
|
|
||||||
|
只接受 HTTPS 地址,`output` 可使用任意附件扩展名。可选参数只有 `--timeout <1-600>`(默认 `60`)和 `--overwrite`。附件上限固定为 25 MiB(26214400 字节),不可调高:脚本先用 HEAD 探测远端声明大小,再检查 GET 响应声明,并在流式接收时持续兜底计数;任一阶段发现超限都会返回 `ok: false` 和明确的“已拒绝下载”错误,且不会发布部分文件。
|
||||||
|
|
||||||
|
成功结果包含 `path`、实际 `size_bytes`、`declared_size_bytes`、`size_limit_bytes`、`size_probe` 和 `content_type`。本脚本不校验文件业务格式;远程 Excel/CSV/TSV 源文件仍使用 `download_workbook.py`。
|
||||||
|
|
||||||
## 检查工作簿
|
## 检查工作簿
|
||||||
|
|
||||||
调用 `scripts/inspect_workbook.py`:
|
调用 `scripts/inspect_workbook.py`:
|
||||||
|
|||||||
256
skills/xlsx/scripts/download_attachment.py
Normal file
256
skills/xlsx/scripts/download_attachment.py
Normal file
@ -0,0 +1,256 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import os
|
||||||
|
import socket
|
||||||
|
import sys
|
||||||
|
import tempfile
|
||||||
|
import urllib.error
|
||||||
|
import urllib.parse
|
||||||
|
import urllib.request
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any, NoReturn, Optional
|
||||||
|
|
||||||
|
from _xlsx_common import emit, failure_message, output_file, publish_file
|
||||||
|
|
||||||
|
|
||||||
|
DEFAULT_TIMEOUT_SECONDS = 60
|
||||||
|
MAX_ATTACHMENT_BYTES = 25 * 1024 * 1024
|
||||||
|
CHUNK_SIZE = 1024 * 1024
|
||||||
|
USER_AGENT = "wechat-robot-xlsx-attachment-downloader/1.0"
|
||||||
|
|
||||||
|
|
||||||
|
class SkillArgumentParser(argparse.ArgumentParser):
|
||||||
|
def error(self, message: str) -> NoReturn:
|
||||||
|
raise ValueError(f"参数错误:{message}")
|
||||||
|
|
||||||
|
|
||||||
|
def _validate_https_url(value: str) -> str:
|
||||||
|
url = value.strip()
|
||||||
|
if not url:
|
||||||
|
raise ValueError("附件 URL 不能为空")
|
||||||
|
if any(character.isspace() or ord(character) < 32 for character in url):
|
||||||
|
raise ValueError("附件 URL 不能包含空白字符或控制字符")
|
||||||
|
|
||||||
|
parsed = urllib.parse.urlsplit(url)
|
||||||
|
if parsed.scheme.lower() != "https" or not parsed.hostname:
|
||||||
|
raise ValueError("附件 URL 必须是有效的 HTTPS 地址")
|
||||||
|
if parsed.username is not None or parsed.password is not None:
|
||||||
|
raise ValueError("附件 URL 不允许包含用户名或密码")
|
||||||
|
try:
|
||||||
|
parsed.port
|
||||||
|
except ValueError as exc:
|
||||||
|
raise ValueError("附件 URL 端口格式不正确") from exc
|
||||||
|
return url
|
||||||
|
|
||||||
|
|
||||||
|
class HTTPSOnlyRedirectHandler(urllib.request.HTTPRedirectHandler):
|
||||||
|
def redirect_request(self, req, fp, code, msg, headers, newurl):
|
||||||
|
return super().redirect_request(
|
||||||
|
req,
|
||||||
|
fp,
|
||||||
|
code,
|
||||||
|
msg,
|
||||||
|
headers,
|
||||||
|
_validate_https_url(newurl),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _parse_args(argv: list[str]) -> argparse.Namespace:
|
||||||
|
parser = SkillArgumentParser(
|
||||||
|
description="下载不超过 25 MiB 的远程 HTTPS 通用附件"
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--url",
|
||||||
|
"--attachment-url",
|
||||||
|
"--attachment_url",
|
||||||
|
dest="url",
|
||||||
|
required=True,
|
||||||
|
help="远程 HTTPS 附件地址",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--output",
|
||||||
|
required=True,
|
||||||
|
help="附件本地输出路径;允许图片、音视频、压缩包及其他文件类型",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--timeout",
|
||||||
|
type=int,
|
||||||
|
default=DEFAULT_TIMEOUT_SECONDS,
|
||||||
|
help=f"连接和读取超时秒数,默认 {DEFAULT_TIMEOUT_SECONDS}",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--overwrite",
|
||||||
|
action="store_true",
|
||||||
|
help="允许覆盖本次任务已存在的缓存文件",
|
||||||
|
)
|
||||||
|
args = parser.parse_args(argv)
|
||||||
|
|
||||||
|
args.url = _validate_https_url(args.url)
|
||||||
|
if args.timeout < 1 or args.timeout > 600:
|
||||||
|
raise ValueError("timeout 必须在 1 到 600 秒之间")
|
||||||
|
args.output = output_file(args.output, overwrite=args.overwrite)
|
||||||
|
return args
|
||||||
|
|
||||||
|
|
||||||
|
def _content_length(headers: Any) -> Optional[int]:
|
||||||
|
raw_value = headers.get("Content-Length")
|
||||||
|
if raw_value is None:
|
||||||
|
return None
|
||||||
|
try:
|
||||||
|
size = int(raw_value)
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
return None
|
||||||
|
return size if size >= 0 else None
|
||||||
|
|
||||||
|
|
||||||
|
def _reject_if_too_large(size_bytes: int, *, source: str) -> None:
|
||||||
|
if size_bytes <= MAX_ATTACHMENT_BYTES:
|
||||||
|
return
|
||||||
|
raise ValueError(
|
||||||
|
"附件超过 25 MiB(26214400 字节)限制,已拒绝下载;"
|
||||||
|
f"{source}大小为 {size_bytes} 字节"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _probe_size(
|
||||||
|
opener: urllib.request.OpenerDirector,
|
||||||
|
url: str,
|
||||||
|
timeout: int,
|
||||||
|
) -> Optional[int]:
|
||||||
|
request = urllib.request.Request(
|
||||||
|
url,
|
||||||
|
headers={
|
||||||
|
"Accept": "*/*",
|
||||||
|
"Accept-Encoding": "identity",
|
||||||
|
"User-Agent": USER_AGENT,
|
||||||
|
},
|
||||||
|
method="HEAD",
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
with opener.open(request, timeout=timeout) as response:
|
||||||
|
_validate_https_url(response.geturl())
|
||||||
|
size = _content_length(response.headers)
|
||||||
|
except (
|
||||||
|
urllib.error.HTTPError,
|
||||||
|
urllib.error.URLError,
|
||||||
|
TimeoutError,
|
||||||
|
socket.timeout,
|
||||||
|
):
|
||||||
|
return None
|
||||||
|
|
||||||
|
if size is not None:
|
||||||
|
_reject_if_too_large(size, source="远端声明")
|
||||||
|
return size
|
||||||
|
|
||||||
|
|
||||||
|
def _content_type(headers: Any) -> Optional[str]:
|
||||||
|
raw_value = headers.get("Content-Type")
|
||||||
|
if not raw_value:
|
||||||
|
return None
|
||||||
|
media_type = raw_value.split(";", 1)[0].strip().lower()
|
||||||
|
return media_type or None
|
||||||
|
|
||||||
|
|
||||||
|
def _download(args: argparse.Namespace) -> dict[str, Any]:
|
||||||
|
output: Path = args.output
|
||||||
|
opener = urllib.request.build_opener(HTTPSOnlyRedirectHandler())
|
||||||
|
probed_size = _probe_size(opener, args.url, args.timeout)
|
||||||
|
request = urllib.request.Request(
|
||||||
|
args.url,
|
||||||
|
headers={
|
||||||
|
"Accept": "*/*",
|
||||||
|
"Accept-Encoding": "identity",
|
||||||
|
"User-Agent": USER_AGENT,
|
||||||
|
},
|
||||||
|
method="GET",
|
||||||
|
)
|
||||||
|
|
||||||
|
temp_path: Optional[Path] = None
|
||||||
|
downloaded_bytes = 0
|
||||||
|
response_size: Optional[int] = None
|
||||||
|
response_type: Optional[str] = None
|
||||||
|
try:
|
||||||
|
with tempfile.NamedTemporaryFile(
|
||||||
|
mode="wb",
|
||||||
|
prefix=f".{output.stem}.",
|
||||||
|
suffix=f".part{output.suffix}",
|
||||||
|
dir=str(output.parent),
|
||||||
|
delete=False,
|
||||||
|
) as temp_file:
|
||||||
|
temp_path = Path(temp_file.name)
|
||||||
|
with opener.open(request, timeout=args.timeout) as response:
|
||||||
|
_validate_https_url(response.geturl())
|
||||||
|
response_size = _content_length(response.headers)
|
||||||
|
response_type = _content_type(response.headers)
|
||||||
|
if response_size is not None:
|
||||||
|
_reject_if_too_large(response_size, source="下载响应声明")
|
||||||
|
|
||||||
|
while True:
|
||||||
|
chunk = response.read(CHUNK_SIZE)
|
||||||
|
if not chunk:
|
||||||
|
break
|
||||||
|
downloaded_bytes += len(chunk)
|
||||||
|
_reject_if_too_large(downloaded_bytes, source="已接收")
|
||||||
|
temp_file.write(chunk)
|
||||||
|
|
||||||
|
temp_file.flush()
|
||||||
|
os.fsync(temp_file.fileno())
|
||||||
|
|
||||||
|
if downloaded_bytes == 0:
|
||||||
|
raise ValueError("远程服务器返回了空附件")
|
||||||
|
|
||||||
|
publish_file(temp_path, output, overwrite=args.overwrite)
|
||||||
|
temp_path = None
|
||||||
|
declared_size = (
|
||||||
|
response_size if response_size is not None else probed_size
|
||||||
|
)
|
||||||
|
if response_size is not None:
|
||||||
|
size_probe = "get-content-length"
|
||||||
|
elif probed_size is not None:
|
||||||
|
size_probe = "head-content-length"
|
||||||
|
else:
|
||||||
|
size_probe = "stream"
|
||||||
|
return {
|
||||||
|
"path": str(output),
|
||||||
|
"size_bytes": downloaded_bytes,
|
||||||
|
"declared_size_bytes": declared_size,
|
||||||
|
"size_limit_bytes": MAX_ATTACHMENT_BYTES,
|
||||||
|
"size_probe": size_probe,
|
||||||
|
"content_type": response_type,
|
||||||
|
}
|
||||||
|
finally:
|
||||||
|
if temp_path is not None:
|
||||||
|
try:
|
||||||
|
temp_path.unlink(missing_ok=True)
|
||||||
|
except OSError:
|
||||||
|
pass
|
||||||
|
|
||||||
|
|
||||||
|
def _failure_message(exc: Exception) -> str:
|
||||||
|
if isinstance(exc, urllib.error.HTTPError):
|
||||||
|
return f"附件下载失败:远程服务器返回 HTTP {exc.code}"
|
||||||
|
if isinstance(exc, (TimeoutError, socket.timeout)):
|
||||||
|
return "附件下载失败:连接或读取超时"
|
||||||
|
if isinstance(exc, urllib.error.URLError):
|
||||||
|
if isinstance(exc.reason, (TimeoutError, socket.timeout)):
|
||||||
|
return "附件下载失败:连接或读取超时"
|
||||||
|
return "附件下载失败:无法访问远程服务器"
|
||||||
|
return failure_message(exc)
|
||||||
|
|
||||||
|
|
||||||
|
def main(argv: Optional[list[str]] = None) -> int:
|
||||||
|
try:
|
||||||
|
args = _parse_args(sys.argv[1:] if argv is None else argv)
|
||||||
|
result = _download(args)
|
||||||
|
except Exception as exc:
|
||||||
|
emit({"ok": False, "error": _failure_message(exc)})
|
||||||
|
return 1
|
||||||
|
emit({"ok": True, **result})
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
@ -25,8 +25,8 @@ from _xlsx_common import (
|
|||||||
|
|
||||||
|
|
||||||
DEFAULT_TIMEOUT_SECONDS = 60
|
DEFAULT_TIMEOUT_SECONDS = 60
|
||||||
DEFAULT_MAX_BYTES = 100 * 1024 * 1024
|
MAX_ALLOWED_BYTES = 25 * 1024 * 1024
|
||||||
MAX_ALLOWED_BYTES = 512 * 1024 * 1024
|
DEFAULT_MAX_BYTES = MAX_ALLOWED_BYTES
|
||||||
CHUNK_SIZE = 1024 * 1024
|
CHUNK_SIZE = 1024 * 1024
|
||||||
MAX_ARCHIVE_MEMBERS = 20_000
|
MAX_ARCHIVE_MEMBERS = 20_000
|
||||||
MAX_ARCHIVE_UNCOMPRESSED_BYTES = 512 * 1024 * 1024
|
MAX_ARCHIVE_UNCOMPRESSED_BYTES = 512 * 1024 * 1024
|
||||||
@ -357,7 +357,7 @@ def _download(args: argparse.Namespace) -> dict[str, Any]:
|
|||||||
expected_bytes = 0
|
expected_bytes = 0
|
||||||
if expected_bytes > args.max_bytes:
|
if expected_bytes > args.max_bytes:
|
||||||
raise ValueError(
|
raise ValueError(
|
||||||
"远程文件超过大小限制:"
|
"远程文件超过大小限制,已拒绝下载:"
|
||||||
f"最多允许 {args.max_bytes} 字节"
|
f"最多允许 {args.max_bytes} 字节"
|
||||||
)
|
)
|
||||||
|
|
||||||
@ -369,7 +369,7 @@ def _download(args: argparse.Namespace) -> dict[str, Any]:
|
|||||||
downloaded_bytes += len(chunk)
|
downloaded_bytes += len(chunk)
|
||||||
if downloaded_bytes > args.max_bytes:
|
if downloaded_bytes > args.max_bytes:
|
||||||
raise ValueError(
|
raise ValueError(
|
||||||
"远程文件超过大小限制:"
|
"远程文件超过大小限制,已拒绝下载:"
|
||||||
f"最多允许 {args.max_bytes} 字节"
|
f"最多允许 {args.max_bytes} 字节"
|
||||||
)
|
)
|
||||||
temp_file.write(chunk)
|
temp_file.write(chunk)
|
||||||
|
|||||||
1
tests/__init__.py
Normal file
1
tests/__init__.py
Normal file
@ -0,0 +1 @@
|
|||||||
|
"""Tests for bundled skill scripts."""
|
||||||
306
tests/test_download_attachments.py
Normal file
306
tests/test_download_attachments.py
Normal file
@ -0,0 +1,306 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import contextlib
|
||||||
|
import importlib.util
|
||||||
|
import io
|
||||||
|
import json
|
||||||
|
import sys
|
||||||
|
import tempfile
|
||||||
|
import unittest
|
||||||
|
from pathlib import Path
|
||||||
|
from unittest import mock
|
||||||
|
|
||||||
|
|
||||||
|
REPOSITORY_ROOT = Path(__file__).resolve().parents[1]
|
||||||
|
SKILLS = {
|
||||||
|
"pdf": {
|
||||||
|
"common_module": "_pdf_common",
|
||||||
|
"root_name": "PDF_OUTPUT_ROOT",
|
||||||
|
"source_downloader": "download_pdf.py",
|
||||||
|
"source_suffix": ".pdf",
|
||||||
|
},
|
||||||
|
"docx": {
|
||||||
|
"common_module": "_docx_common",
|
||||||
|
"root_name": "WORD_OUTPUT_ROOT",
|
||||||
|
"source_downloader": "download_document.py",
|
||||||
|
"source_suffix": ".docx",
|
||||||
|
},
|
||||||
|
"xlsx": {
|
||||||
|
"common_module": "_xlsx_common",
|
||||||
|
"root_name": "EXCEL_OUTPUT_ROOT",
|
||||||
|
"source_downloader": "download_workbook.py",
|
||||||
|
"source_suffix": ".xlsx",
|
||||||
|
},
|
||||||
|
"pptx": {
|
||||||
|
"common_module": "_pptx_common",
|
||||||
|
"root_name": "PPT_OUTPUT_ROOT",
|
||||||
|
"source_downloader": "download_presentation.py",
|
||||||
|
"source_suffix": ".pptx",
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def load_script(skill: str, filename: str):
|
||||||
|
script_path = REPOSITORY_ROOT / "skills" / skill / "scripts" / filename
|
||||||
|
scripts_directory = str(script_path.parent)
|
||||||
|
sys.path.insert(0, scripts_directory)
|
||||||
|
try:
|
||||||
|
module_name = f"_test_{skill}_{script_path.stem}"
|
||||||
|
spec = importlib.util.spec_from_file_location(module_name, script_path)
|
||||||
|
if spec is None or spec.loader is None:
|
||||||
|
raise RuntimeError(f"无法加载测试脚本:{script_path}")
|
||||||
|
module = importlib.util.module_from_spec(spec)
|
||||||
|
sys.modules[module_name] = module
|
||||||
|
spec.loader.exec_module(module)
|
||||||
|
return module
|
||||||
|
finally:
|
||||||
|
sys.path.remove(scripts_directory)
|
||||||
|
|
||||||
|
|
||||||
|
class FakeResponse:
|
||||||
|
def __init__(
|
||||||
|
self,
|
||||||
|
*,
|
||||||
|
headers: dict[str, str],
|
||||||
|
payload: bytes = b"",
|
||||||
|
generated_size: int = 0,
|
||||||
|
) -> None:
|
||||||
|
self.headers = headers
|
||||||
|
self._payload = payload
|
||||||
|
self._payload_offset = 0
|
||||||
|
self._remaining = generated_size
|
||||||
|
|
||||||
|
def __enter__(self):
|
||||||
|
return self
|
||||||
|
|
||||||
|
def __exit__(self, exc_type, exc, traceback):
|
||||||
|
return False
|
||||||
|
|
||||||
|
def geturl(self) -> str:
|
||||||
|
return "https://example.test/attachment"
|
||||||
|
|
||||||
|
def read(self, size: int) -> bytes:
|
||||||
|
if self._remaining:
|
||||||
|
chunk_size = min(size, self._remaining)
|
||||||
|
self._remaining -= chunk_size
|
||||||
|
return b"x" * chunk_size
|
||||||
|
if self._payload_offset >= len(self._payload):
|
||||||
|
return b""
|
||||||
|
end = min(self._payload_offset + size, len(self._payload))
|
||||||
|
chunk = self._payload[self._payload_offset : end]
|
||||||
|
self._payload_offset = end
|
||||||
|
return chunk
|
||||||
|
|
||||||
|
|
||||||
|
class FakeOpener:
|
||||||
|
def __init__(
|
||||||
|
self,
|
||||||
|
*,
|
||||||
|
head_headers: dict[str, str],
|
||||||
|
get_headers: dict[str, str],
|
||||||
|
payload: bytes = b"",
|
||||||
|
generated_size: int = 0,
|
||||||
|
) -> None:
|
||||||
|
self.head_headers = head_headers
|
||||||
|
self.get_headers = get_headers
|
||||||
|
self.payload = payload
|
||||||
|
self.generated_size = generated_size
|
||||||
|
self.methods: list[str] = []
|
||||||
|
|
||||||
|
def open(self, request, timeout: int):
|
||||||
|
del timeout
|
||||||
|
method = request.get_method()
|
||||||
|
self.methods.append(method)
|
||||||
|
if method == "HEAD":
|
||||||
|
return FakeResponse(headers=self.head_headers)
|
||||||
|
return FakeResponse(
|
||||||
|
headers=self.get_headers,
|
||||||
|
payload=self.payload,
|
||||||
|
generated_size=self.generated_size,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
class DownloadAttachmentTests(unittest.TestCase):
|
||||||
|
@classmethod
|
||||||
|
def setUpClass(cls) -> None:
|
||||||
|
cls.modules = {
|
||||||
|
skill: load_script(skill, "download_attachment.py")
|
||||||
|
for skill in SKILLS
|
||||||
|
}
|
||||||
|
cls.source_modules = {
|
||||||
|
skill: load_script(skill, details["source_downloader"])
|
||||||
|
for skill, details in SKILLS.items()
|
||||||
|
}
|
||||||
|
|
||||||
|
def test_each_skill_downloads_an_arbitrary_attachment_type(self) -> None:
|
||||||
|
payload = b"small video payload"
|
||||||
|
for skill, module in self.modules.items():
|
||||||
|
with self.subTest(skill=skill), tempfile.TemporaryDirectory() as tmp:
|
||||||
|
output = Path(tmp) / "material.mp4"
|
||||||
|
common = sys.modules[SKILLS[skill]["common_module"]]
|
||||||
|
opener = FakeOpener(
|
||||||
|
head_headers={"Content-Length": str(len(payload))},
|
||||||
|
get_headers={
|
||||||
|
"Content-Length": str(len(payload)),
|
||||||
|
"Content-Type": "video/mp4; charset=binary",
|
||||||
|
},
|
||||||
|
payload=payload,
|
||||||
|
)
|
||||||
|
with mock.patch.object(
|
||||||
|
common,
|
||||||
|
SKILLS[skill]["root_name"],
|
||||||
|
Path(tmp).resolve(),
|
||||||
|
):
|
||||||
|
with mock.patch.object(
|
||||||
|
module.urllib.request,
|
||||||
|
"build_opener",
|
||||||
|
return_value=opener,
|
||||||
|
):
|
||||||
|
args = module._parse_args(
|
||||||
|
[
|
||||||
|
"--url",
|
||||||
|
"https://example.test/material.mp4",
|
||||||
|
"--output",
|
||||||
|
str(output),
|
||||||
|
]
|
||||||
|
)
|
||||||
|
result = module._download(args)
|
||||||
|
|
||||||
|
self.assertEqual(output.read_bytes(), payload)
|
||||||
|
self.assertEqual(opener.methods, ["HEAD", "GET"])
|
||||||
|
self.assertEqual(result["size_bytes"], len(payload))
|
||||||
|
self.assertEqual(
|
||||||
|
result["size_limit_bytes"],
|
||||||
|
25 * 1024 * 1024,
|
||||||
|
)
|
||||||
|
self.assertEqual(result["content_type"], "video/mp4")
|
||||||
|
|
||||||
|
def test_head_probe_rejects_oversize_without_get(self) -> None:
|
||||||
|
for skill, module in self.modules.items():
|
||||||
|
with self.subTest(skill=skill), tempfile.TemporaryDirectory() as tmp:
|
||||||
|
output = Path(tmp) / "too-large.zip"
|
||||||
|
common = sys.modules[SKILLS[skill]["common_module"]]
|
||||||
|
opener = FakeOpener(
|
||||||
|
head_headers={
|
||||||
|
"Content-Length": str(module.MAX_ATTACHMENT_BYTES + 1)
|
||||||
|
},
|
||||||
|
get_headers={},
|
||||||
|
)
|
||||||
|
stdout = io.StringIO()
|
||||||
|
with mock.patch.object(
|
||||||
|
common,
|
||||||
|
SKILLS[skill]["root_name"],
|
||||||
|
Path(tmp).resolve(),
|
||||||
|
):
|
||||||
|
with mock.patch.object(
|
||||||
|
module.urllib.request,
|
||||||
|
"build_opener",
|
||||||
|
return_value=opener,
|
||||||
|
):
|
||||||
|
with contextlib.redirect_stdout(stdout):
|
||||||
|
return_code = module.main(
|
||||||
|
[
|
||||||
|
"--url",
|
||||||
|
"https://example.test/too-large.zip",
|
||||||
|
"--output",
|
||||||
|
str(output),
|
||||||
|
]
|
||||||
|
)
|
||||||
|
|
||||||
|
response = json.loads(stdout.getvalue())
|
||||||
|
self.assertEqual(return_code, 1)
|
||||||
|
self.assertFalse(response["ok"])
|
||||||
|
self.assertRegex(
|
||||||
|
response["error"],
|
||||||
|
"超过 25 MiB.*已拒绝下载",
|
||||||
|
)
|
||||||
|
self.assertEqual(opener.methods, ["HEAD"])
|
||||||
|
self.assertFalse(output.exists())
|
||||||
|
|
||||||
|
def test_stream_limit_rejects_and_removes_partial_file(self) -> None:
|
||||||
|
for skill, module in self.modules.items():
|
||||||
|
with self.subTest(skill=skill), tempfile.TemporaryDirectory() as tmp:
|
||||||
|
output = Path(tmp) / "unknown-size.bin"
|
||||||
|
opener = FakeOpener(
|
||||||
|
head_headers={},
|
||||||
|
get_headers={"Content-Type": "application/octet-stream"},
|
||||||
|
generated_size=module.MAX_ATTACHMENT_BYTES + 1,
|
||||||
|
)
|
||||||
|
args = argparse.Namespace(
|
||||||
|
url="https://example.test/unknown-size.bin",
|
||||||
|
output=output,
|
||||||
|
timeout=60,
|
||||||
|
overwrite=False,
|
||||||
|
)
|
||||||
|
with mock.patch.object(
|
||||||
|
module.urllib.request,
|
||||||
|
"build_opener",
|
||||||
|
return_value=opener,
|
||||||
|
):
|
||||||
|
with self.assertRaisesRegex(
|
||||||
|
ValueError,
|
||||||
|
"超过 25 MiB.*已拒绝下载",
|
||||||
|
):
|
||||||
|
module._download(args)
|
||||||
|
|
||||||
|
self.assertEqual(opener.methods, ["HEAD", "GET"])
|
||||||
|
self.assertFalse(output.exists())
|
||||||
|
self.assertEqual(list(Path(tmp).iterdir()), [])
|
||||||
|
|
||||||
|
def test_get_content_length_rejects_when_head_has_no_size(self) -> None:
|
||||||
|
for skill, module in self.modules.items():
|
||||||
|
with self.subTest(skill=skill), tempfile.TemporaryDirectory() as tmp:
|
||||||
|
output = Path(tmp) / "get-declared-large.mov"
|
||||||
|
opener = FakeOpener(
|
||||||
|
head_headers={},
|
||||||
|
get_headers={
|
||||||
|
"Content-Length": str(module.MAX_ATTACHMENT_BYTES + 1)
|
||||||
|
},
|
||||||
|
)
|
||||||
|
args = argparse.Namespace(
|
||||||
|
url="https://example.test/get-declared-large.mov",
|
||||||
|
output=output,
|
||||||
|
timeout=60,
|
||||||
|
overwrite=False,
|
||||||
|
)
|
||||||
|
with mock.patch.object(
|
||||||
|
module.urllib.request,
|
||||||
|
"build_opener",
|
||||||
|
return_value=opener,
|
||||||
|
):
|
||||||
|
with self.assertRaisesRegex(
|
||||||
|
ValueError,
|
||||||
|
"超过 25 MiB.*已拒绝下载",
|
||||||
|
):
|
||||||
|
module._download(args)
|
||||||
|
|
||||||
|
self.assertEqual(opener.methods, ["HEAD", "GET"])
|
||||||
|
self.assertFalse(output.exists())
|
||||||
|
self.assertEqual(list(Path(tmp).iterdir()), [])
|
||||||
|
|
||||||
|
def test_source_downloaders_cannot_raise_the_25_mib_limit(self) -> None:
|
||||||
|
expected_limit = 25 * 1024 * 1024
|
||||||
|
for skill, module in self.source_modules.items():
|
||||||
|
with self.subTest(skill=skill):
|
||||||
|
self.assertEqual(module.DEFAULT_MAX_BYTES, expected_limit)
|
||||||
|
self.assertEqual(module.MAX_ALLOWED_BYTES, expected_limit)
|
||||||
|
with self.assertRaisesRegex(
|
||||||
|
ValueError,
|
||||||
|
f"1 到 {expected_limit}",
|
||||||
|
):
|
||||||
|
module._parse_args(
|
||||||
|
[
|
||||||
|
"--url",
|
||||||
|
"https://example.test/source",
|
||||||
|
"--output",
|
||||||
|
"/outside/source"
|
||||||
|
+ SKILLS[skill]["source_suffix"],
|
||||||
|
"--max-bytes",
|
||||||
|
str(expected_limit + 1),
|
||||||
|
]
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
unittest.main()
|
||||||
Loading…
Reference in New Issue
Block a user