Sync third-party and MCP marketplace plugins

Constraint: Public skills are published only by explicit administrator action unless they are tracked third-party market sources.
Confidence: high
Scope-risk: narrow
Directive: Keep private/internal skills out of the public marketplace and preserve normal incremental market Git history.
Tested: Marketplace validation passed.
This commit is contained in:
KeyInfo Bot
2026-08-11 00:02:04 +08:00
parent 271c8ee14b
commit 572930a9e0
60 changed files with 2291 additions and 503 deletions
+8 -8
View File
@@ -24,8 +24,8 @@
"repo": "https://github.com/Yeachan-Heo/oh-my-codex.git",
"ref": "main",
"adapter": "codex-plugin",
"commit": "a62d5bd77bef6d2bc7df467dcae68082b8616239",
"syncedAt": "2026-08-01T16:00:00Z"
"commit": "b30127a0979c96046c8cd6312a8cd922c3516cad",
"syncedAt": "2026-08-10T15:59:59Z"
},
{
"id": "ui-ux-pro-max",
@@ -60,8 +60,8 @@
"repo": "https://github.com/shadcn-ui/ui.git",
"ref": "main",
"adapter": "claude-skill",
"commit": "6261bd89f72d794aea491482cc2acfd8dc3d63e2",
"syncedAt": "2026-08-08T16:00:01Z"
"commit": "deda4df80fb350230b2fce2b575e769a90cae076",
"syncedAt": "2026-08-10T15:59:59Z"
},
{
"id": "frontend-slides",
@@ -96,8 +96,8 @@
"repo": "https://github.com/hugohe3/ppt-master.git",
"ref": "main",
"adapter": "claude-skill",
"commit": "108b4b82db1b6076284bd68f603422044899eaf2",
"syncedAt": "2026-08-09T16:00:00Z"
"commit": "182c6b8229a44990cdc5b394545f90992be377d6",
"syncedAt": "2026-08-10T15:59:59Z"
},
{
"id": "next-skills",
@@ -105,8 +105,8 @@
"repo": "https://github.com/vercel/next.js.git",
"ref": "canary",
"adapter": "skill-collection",
"commit": "62d9b8e03160727191bdcecd96c2284575621501",
"syncedAt": "2026-08-09T16:00:00Z"
"commit": "1722e45c3957d1e6ce2aaff482a9b9ece28864a1",
"syncedAt": "2026-08-10T15:59:59Z"
}
]
}
@@ -3,5 +3,5 @@
"name": "playwright浏览器自动化操作",
"version": "20260605",
"keySource": "none",
"syncedAt": "2026-08-09T16:02:07Z"
"syncedAt": "2026-08-10T16:02:03Z"
}
@@ -2,8 +2,8 @@
"sourceId": "next-skills",
"repo": "https://github.com/vercel/next.js.git",
"ref": "canary",
"commit": "62d9b8e03160727191bdcecd96c2284575621501",
"commit": "1722e45c3957d1e6ce2aaff482a9b9ece28864a1",
"adapter": "skill-collection",
"sourcePath": "skills",
"syncedAt": "2026-08-09T16:00:00Z"
"syncedAt": "2026-08-10T15:59:59Z"
}
@@ -1,6 +1,6 @@
{
"name": "oh-my-codex",
"version": "0.20.4",
"version": "0.20.5",
"description": "oh-my-codex 是 Codex CLI 的多 Agent 编排、结构化工作流、插件级 hooks、MCP 和 HUD 扩展插件。",
"author": {
"name": "Yeachan Heo",
@@ -2,8 +2,8 @@
"sourceId": "oh-my-codex",
"repo": "https://github.com/Yeachan-Heo/oh-my-codex.git",
"ref": "main",
"commit": "a62d5bd77bef6d2bc7df467dcae68082b8616239",
"commit": "b30127a0979c96046c8cd6312a8cd922c3516cad",
"adapter": "codex-plugin",
"sourcePath": "plugins/oh-my-codex",
"syncedAt": "2026-08-01T16:00:00Z"
"syncedAt": "2026-08-10T15:59:59Z"
}
@@ -34,7 +34,7 @@ Complex tasks often fail silently: partial implementations get declared "done",
- Use `run_in_background: true` for long operations (installs, builds, test suites)
- The documented-leader preflight is not a general native-session gate. Run `omx ralplan preflight --json` only when native role routing reports `role_routing_unavailable` and Ralph attempts adapted Ralplan Planner, Architect, or Critic authority, adapted role-intent, or adapted consensus authority. On `unsupported_documented_leader_proof`, stop before that adapted authority and use a Codex surface with documented root proof or a reviewed alternative workflow. Do not infer root authority from `session_id`, undocumented `thread_id`, session/pointer/transcript/cwd state, absent child data, or prompt labels. Ordinary native planning, lifecycle, state, status, health, HUD, runtime, setup, install, sync, and unrelated delegation remain outside this preflight boundary and under their existing controls.
- When the native surface exposes `agent_type` role routing, set `agent_type` to an installed OMX role and never omit it for OMX work; use `reasoning_effort` for per-dispatch intensity when needed.
- **OMX adapted role-pass protocol:** when native routing is `role_routing_unavailable`, do not fabricate `agent_type`. On documented Codex 0.144.5, only an attempted adapted Ralplan Planner, Architect, Critic, role-intent, or consensus authority path requires `omx ralplan preflight --json` and fails closed on `unsupported_documented_leader_proof`; do not use prompt labels, task-name carriers, pending intents, markers, or `omx ralplan role-intent write` as substitutes.
- **OMX adapted role-pass protocol:** when native routing is `role_routing_unavailable`, do not fabricate `agent_type`. On the exact reviewed Codex releases 0.144.5, 0.145.0, 0.146.1, and 0.148.0-alpha.5, only an attempted adapted Ralplan Planner, Architect, Critic, role-intent, or consensus authority path requires `omx ralplan preflight --json` and fails closed on `unsupported_documented_leader_proof`; every other version remains unknown and fails closed. Do not use prompt labels, task-name carriers, pending intents, markers, or `omx ralplan role-intent write` as substitutes.
- Preserve legacy Ralph tier intent through native reasoning effort: LOW -> `low`, STANDARD -> `medium`, THOROUGH -> `xhigh`
- Deliver the full implementation: no scope reduction, no partial completion, no deleting tests to make them pass
- Apply the shared workflow guidance pattern: outcome-first framing, concise visible updates for multi-step execution, local overrides for the active workflow branch, validation proportional to risk, explicit stop rules, and automatic continuation for safe reversible steps. Ask only for material, destructive, credentialed, external-production, or preference-dependent branches.
@@ -80,7 +80,7 @@ Complex tasks often fail silently: partial implementations get declared "done",
- Standard changes: `task(agent_type="architect", reasoning_effort="medium", prompt="...")`
- >20 files or security/architectural changes: `task(agent_type="architect", reasoning_effort="xhigh", prompt="...")`
- Ralph floor: always run an explicit `architect` native subagent, even for small changes
- On `role_routing_unavailable`, do not invoke `omx ralplan role-intent write` or manufacture an Architect identity. Run the documented-leader preflight only if Architect verification would attempt adapted Ralplan Architect authority; on documented Codex 0.144.5 that adapted path is unavailable, so surface the leader-proof diagnostic/remediation and stop before it. Ordinary Ralph delegation remains under its existing controls. On a future or other surface, use an adapted route only after its documented positive root proof has been reviewed and implemented.
- On `role_routing_unavailable`, do not invoke `omx ralplan role-intent write` or manufacture an Architect identity. Run the documented-leader preflight only if Architect verification would attempt adapted Ralplan Architect authority; on the exact reviewed Codex releases 0.144.5, 0.145.0, 0.146.1, and 0.148.0-alpha.5 that adapted path is unavailable, while every other version remains unknown and fails closed. Surface the leader-proof diagnostic and the supported-surface recovery guidance, then stop before that authority. Ordinary Ralph delegation remains under its existing controls. Use an adapted route only after its documented positive root proof has been reviewed and implemented.
7.5 **Mandatory Deslop Pass**:
- After Step 7 passes, run `oh-my-codex:ai-slop-cleaner` on **all files changed during the Ralph session**.
- Scope the cleaner to **changed files only**; do not widen the pass beyond Ralph-owned edits.
@@ -204,7 +204,7 @@ Why bad: These are independent tasks that should run in parallel, not sequential
- [ ] Fresh test run output shows all tests pass
- [ ] Fresh build output shows success
- [ ] lsp_diagnostics shows 0 errors on affected files
- [ ] Architect verification passed: on a routing-capable surface via explicit `task(agent_type="architect", reasoning_effort="medium"...)` minimum. On documented Codex 0.144.5 role-routing-unavailable surfaces, no adapted Ralplan Architect pass is valid; when that adapted authority is attempted, preflight must have stopped with the leader-proof diagnostic. Ordinary Ralph work remains subject to its existing controls.
- [ ] Architect verification passed: on a routing-capable surface via explicit `task(agent_type="architect", reasoning_effort="medium"...)` minimum. On the exact reviewed Codex releases 0.144.5, 0.145.0, 0.146.1, and 0.148.0-alpha.5 when role routing is unavailable, no adapted Ralplan Architect pass is valid; every other version remains unknown and fails closed. When adapted authority is attempted, preflight must have stopped with the leader-proof diagnostic. Ordinary Ralph work remains subject to its existing controls.
- [ ] Codex goal-mode completion audit passed, and `update_goal({status: "complete"})` was called when an active goal exists
- [ ] ai-slop-cleaner pass completed on changed files (or --no-deslop specified)
- [ ] Post-deslop regression tests pass
@@ -51,7 +51,7 @@ The consensus workflow:
2. **User feedback** *(--interactive only)*: If `--interactive` is set, use the structured question UI (`omx question` in attached tmux; native structured input outside tmux when available) to present the draft plan **plus the Principles / Drivers / Options summary** before review (Proceed to review / Request changes / Skip review). Otherwise, automatically proceed to review.
**Native role-routing preflight:** Keyword routing may already have selected Ralplan, but it is not authority. Run `omx ralplan preflight --json` only when the native task surface reports `role_routing_unavailable` and this workflow attempts adapted Ralplan Planner, Architect, or Critic authority, adapted role-intent, or adapted consensus authority. On `unsupported_documented_leader_proof`, stop before that adapted authority and use a Codex surface with documented root proof or a reviewed alternative workflow. Do not infer root identity from `session_id`, undocumented `thread_id`, session/pointer/transcript/cwd state, absence of child data, or a prompt label. Ordinary native planning, lifecycle, state, status, health, HUD, runtime, setup, install, sync, and unrelated delegation are outside this preflight boundary and remain governed by existing controls.
**Native role-routing rule:** When the native surface exposes `agent_type` role routing, set `agent_type` to an installed OMX role and never omit it for OMX work. When it does not (`role_routing_unavailable`), do not fabricate `agent_type`. On documented Codex 0.144.5, adapted Ralplan Planner, Architect, Critic, role-intent, and consensus authority are unavailable because they lack documented root proof; do not silently weaken routing with a prompt role label or inferred carrier. Use a Codex surface with documented root proof or a reviewed alternative workflow for that authority. A direct `omx ralplan role-intent write` attempt is denied with machine reason `unsupported_documented_leader_proof`.
**Native role-routing rule:** When the native surface exposes `agent_type` role routing, set `agent_type` to an installed OMX role and never omit it for OMX work. When it does not (`role_routing_unavailable`), do not fabricate `agent_type`. On the exact reviewed Codex releases 0.144.5, 0.145.0, 0.146.1, and 0.148.0-alpha.5, adapted Ralplan Planner, Architect, Critic, role-intent, and consensus authority are unavailable because they lack documented root proof; every other version remains unknown and fails closed. Do not silently weaken routing with a prompt role label or inferred carrier. Use a Codex surface with documented root proof or a reviewed alternative workflow for that authority. A direct `omx ralplan role-intent write` attempt is denied with machine reason `unsupported_documented_leader_proof`.
3. **Architect** reviews for architectural soundness and must provide the strongest steelman antithesis, at least one real tradeoff tension, and (when possible) synthesis — **await completion before step 4**. Launch this as a subsequent role-specific `Architect` subagent and pass the full task statement, context snapshot, PRD/test-spec paths, and relevant prior findings; do not substitute an unvalidated reviewer identity or a short improvised reviewer prompt. In deliberate mode, Architect should explicitly flag principle violations.
4. **Critic** evaluates against quality criteria — run only after step 3 completes. Launch this as a subsequent role-specific `Critic` subagent with the full task statement, context snapshot, PRD/test-spec paths, and the completed Architect review; do not ask the Architect subagent to perform the Critic gate and do not substitute an unvalidated reviewer identity or a short improvised reviewer prompt. Critic must enforce principle-option consistency, fair alternatives, risk mitigation clarity, testable acceptance criteria, and concrete verification steps. In deliberate mode, Critic must reject missing/weak pre-mortem or expanded test plan.
@@ -2,8 +2,8 @@
"sourceId": "ppt-master",
"repo": "https://github.com/hugohe3/ppt-master.git",
"ref": "main",
"commit": "108b4b82db1b6076284bd68f603422044899eaf2",
"commit": "182c6b8229a44990cdc5b394545f90992be377d6",
"adapter": "claude-skill",
"sourcePath": "skills/ppt-master",
"syncedAt": "2026-08-09T16:00:00Z"
"syncedAt": "2026-08-10T15:59:59Z"
}
@@ -29,6 +29,7 @@ PPT Master is a routed presentation workflow. This entry owns global execution d
| Selected route / profile | Runtime authority |
|---|---|
| Generate PPTX — Image to PPTX | [`workflows/profiles/image-to-pptx.md`](workflows/profiles/image-to-pptx.md); Codex-supported, always Quick |
| Generate PPTX — Beautify | [`workflows/profiles/beautify-pptx.md`](workflows/profiles/beautify-pptx.md); explicit Quick intent selects Quick, otherwise Default |
| Generate PPTX — ordinary Default | [`workflows/generate-pptx.md`](workflows/generate-pptx.md) |
| Generate PPTX — ordinary explicit Quick | [`workflows/profiles/quick-generate.md`](workflows/profiles/quick-generate.md) |
@@ -37,10 +38,10 @@ PPT Master is a routed presentation workflow. This entry owns global execution d
| Enhance Native PPTX | [`workflows/native-enhance-pptx.md`](workflows/native-enhance-pptx.md) |
**Hard rule — selected authority only**: Do not load another top-level route's
procedure after routing. Beautify selects exactly one Generate runtime from the
explicit Quick signal; never load both Default and Quick. Profiles, stages,
governance files, and child workflows refine one selected route; they never
compete with it.
procedure after routing. Image to PPTX and Beautify are mutually exclusive;
Image to PPTX activates Quick, while Beautify selects from explicit Quick
intent. Never load both runtimes. Supporting documents refine one route; they
never compete with it.
---
@@ -27,10 +27,13 @@ history, or resumable planning state. Context loss restarts the Quick run.
| `sources/*.facts.json` | Fact provenance contract | Stable external `fact_id` → claim/source mapping created by topic research | Default Strategist cites IDs in §IX and Executor resolves them for attribution; Quick's current agent carries the same IDs into visible attribution. Scenario data never enters this file. |
| `sources/` converted-source originals | Source archive | Imported source files that have a converted content contract (`.pdf` / `.pptx` / `.docx` / `.xlsx` / `.html` / `.epub` / `.tex` / `.rst` / `.ipynb` / `.typ`, etc.) and source-adjacent extracted assets | Read via the converted `<stem>.md` in the main pipeline; direct-PPTX workflows read the `.pptx` by route |
| `sources/*.conversion_profile.json`, `sources/*_files/image_manifest.json` | Pipeline sidecar | Conversion audit record / asset index | NOT read as slide content; open only to audit a conversion or resolve assets |
| `sources/` Image to PPTX source containers | Visible-surface source contract | One or more immutable raster inputs that may each contain one page or several visibly bounded page frames | The Codex-supported Quick profile normalizes them into the ordered canonical frame roster without overwriting originals; frame count, not input-file count, owns slide count. |
| `analysis/source_profile.json` | Machine fact index | Compact PPTX intake digest | Default Strategist or Quick's current agent reads it as factual context and recommendation candidates |
| `analysis/<stem>.identity.json` | Native deck identity facts | Canvas, theme palette/fonts, observed usage | Read selectively when detailed identity facts are needed |
| `analysis/<stem>.slide_library.json` | Native PPTX structure facts | Text slots, geometry, native tables, native chart caches, SmartArt nodes/connections | Direct PPTX workflows use as native fill/structure contract |
| `analysis/beautify_inventory.json` | Beautify frozen validation ledger | Complete self-contained per-slide ledger of frozen content/data and validation exceptions, deterministically inlined from the source extracts | Keep the complete file as the sole Beautify ledger. Models normally consume derived read-only stdout projections from `beautify_inventory.py --summary` or `--page`; add `--with-geometry` only when raw source geometry is required. Never persist a projection or treat it as a second authority. |
| `analysis/image_analysis.csv` | Regenerated image fact view | Measured facts about the current `images/` folder | Re-run `analyze_images.py` before reading image facts after changes |
| `analysis/reconstruction_inventory.json` | Image to PPTX source-evidence ledger | Source-container/frame mapping, hashes, page order/canvas, visible-region bboxes, observed text/graphic/image family, transcription confidence, overlap/z-order observations, and unresolved evidence | The active Codex Quick agent writes it before layer decisions. It contains no final realization, prompts, output filenames, or SVG bindings; those belong to Quick active context plus required operational image evidence. |
| `design_spec.md` | Strategist design authority | Human-readable design intent, page brief, rationale, resources, and production mechanics | Consume final confirmation once, write/audit here, and apply enabled refinement to this same artifact. §I records effective Speaker Notes, Custom Animations, and Narration Audio outcomes plus provenance; a newer explicit user instruction updates only its owning outcome without reopening Confirm UI. After Gate 1 plus conditional approval, later roles read this file instead of `result.json`; §IX owns Executor page content. |
| `spec_lock.md` | Execution anchor and routing contract | Machine-readable stable color/type roles, icons, images, page rhythm, charts, `template_reuse_scope`, and the route's PowerPoint structure mode; mirror/layout template routes additionally own input prototypes, the Master roster, and the complete page-to-Master/Layout mapping | Strategist authors the route-specific anchors from the audited Design Spec plus current project/page/template context. Executor retains the complete lock once per valid execution context; local uncertainty consults that retained copy before the owning Design Spec fragment. Sparse page-local color/font garnish needs no lock row; a recurring semantic role or new adaptive Layout identity requires Strategist repair before reuse. |
| `project_manager.py page-context` stdout | Derived on-demand page context | Read-only model-facing anchor set + current-page delta + fingerprints for large references | Use only for explicit diagnostics/telemetry or an unresolved page/template/chart path-SHA projection. Never edit or persist it as a replacement source of truth, and never run it as a routine pre-page gate. `global` is a bounded anchor set, not a whitelist. `reference_set` carries path/SHA/load policy but never appends reference payloads. |
@@ -68,9 +71,9 @@ history, or resumable planning state. Context loss restarts the Quick run.
| Invariant | Rule |
|---|---|
| Content authority | Content-type files in `sources/` own the factual/text origin for content, tables, chart values, and SmartArt wording. Default Strategist resolves them into §IX and Executor realizes that contract without drafting a second outline. Quick's current agent resolves them once in active context before SVG authoring. `slide_library.json` does not own content values. |
| Content authority | Content-type files in `sources/` own the factual/text origin for content, tables, chart values, and SmartArt wording. For Beautify, `sources/<stem>.md` remains that origin; `analysis/beautify_inventory.json` only deterministically inlines the frozen values needed by its self-contained validation contract and does not become a second content authority. Default Strategist resolves source content into §IX and Executor realizes that contract without drafting a second outline. Quick's current agent resolves it once in active context before SVG authoring. `slide_library.json` does not own content values. |
| Sources read policy | In `sources/`, read content-type files (`.md` / `.markdown` / `.txt` / `.csv` / `.tsv` / `.json` / `.jsonl` / `.yaml` / `.yml`) and judge by content — a `.json` / `.csv` may be core content or just data. Exclude known sidecars: `*.conversion_profile.json` and `*_files/image_manifest.json`. `analysis/` facts (`source_profile.json`, `<stem>.slide_library.json`) are read per Step 4 / direct-PPTX workflow, not in the `sources/` content scan. |
| PPTX structure | `slide_library.json` owns native geometry, slot facts, and SmartArt layout/relationships for direct PPTX workflows. |
| PPTX structure | `analysis/<stem>.slide_library.json` owns native geometry, slot facts, and SmartArt layout/relationships for direct PPTX workflows. The Beautify inventory only deterministically inlines the subset required by its validation contract; neither the full ledger nor its stdout projections replace `slide_library.json` as the native-structure fact source. |
| Design contract | Final confirmation once → audited `design_spec.md` → optional same-file refinement/approval → context-authored lock. Never maintain a parallel draft/lock. Executor may apply `Template Application` prose but never replace identity. Repair divergence from the approved Design Spec/context unless it fails active-decision fidelity. |
| Flat packaging authority | Free-design, brand-only, Style-only, and every plan with `template_reuse_scope: style` declare `pptx_structure.mode: flat` and omit `pptx_masters`, `pptx_layouts`, `page_pptx_layouts`, and `page_layouts`. A Style installed alongside Layout/Deck changes only Direction / method and does not force the non-Style structure plan to flat. `svg_output/` owns the complete Slide-local visual design without root Master/Layout identity, fixed-layer ownership, or placeholder metadata. Export materializes one clean project-owned Master plus one Blank Layout, applies the locked theme defaults, removes stock content placeholders/Layout inventory, and retains only the standard date/footer/slide-number capability hooks. |
| Template structure authority | `template_reuse_scope: mirror|layout` uses `page_layouts` for each page's authoring-input prototype. `pptx_masters` / `pptx_layouts` own the unique reusable output definitions, while `page_pptx_layouts` owns page assignment. Strict keeps the prototype contract; adaptive may use a current or new Layout already declared by Strategist. A construction-discovered structural change returns upstream for definition and assignment repair before authoring resumes. Mirror additionally preserves literal visuals/text topology; layout allows project-controlled reflow/re-skinning. Unused definitions may register without a published Slide. Templates validate provenance but never add missing visible page objects during export. |
@@ -78,6 +81,7 @@ history, or resumable planning state. Context loss restarts the Quick run.
| Imported-template authoring | Editable SVGs under `authoring-svg/` own create-template edits, `authoring_summary.json` owns model-facing orientation, and `authoring_manifest.json` owns tool-only source-object identity. Lossless `svg/` owns immutable native payload and fallback evidence; optional `svg-flat/` owns only complete-page verification. Materialized `templates/*.svg` own the validated deliverable contract and contain no IR-only source refs. |
| Legacy template input | Old unmapped/distilled/preserve structured projects and incomplete template packages are not migrated in place. [`create-template`](../workflows/create-template.md) authors a new current workspace: original PPTX Type A may preserve existing native topology in mirror; legacy SVG-only Type B is visual reference for `standard` / `fidelity`. Intentional free-design, Brand-only, and Style-only `flat` projects are already current. The exporter does not migrate or visually cluster legacy structure. |
| Image facts | `images/` is live state; `analysis/image_analysis.csv` is a regenerated view, not a durable cache. |
| Image to PPTX | Normalized page frames own visible-surface truth, not literal reuse of their pixel bytes; `analysis/reconstruction_inventory.json` owns only source evidence. Native text restores visible wording. Identity graphics use an exact vector, deterministic redraw, sufficient source asset, or Codex reference-based high-resolution reconstruction; data graphics stay native-and-verified or exact and are never generatively recreated. Scene imagery uses registered clean-base/midground/subject/foreground layers; padded-bbox-disjoint objects may share one generated plate before independent crop/slice realization. Never infer hidden semantics, substitute a merely similar graphic, or use a full-page screenshot skin as the editable result. |
| SVG source | `svg_output/` is the only author source for generated pages. |
| Page-design closure | On SVG-authoring routes, every visible exported-slide object exists in the corresponding page SVG or an explicitly referenced visual asset. |
| Package-behavior separation | Speaker notes, animations, transitions, narration, and direct native-PPTX workflows keep their owning artifacts; do not force them into SVG metadata. |
@@ -42,36 +42,60 @@ Content purpose?
└── Print → 1240x1754
```
## Layout Principles
## Platform Keep-clear
### Landscape (16:9, 4:3, 2.35:1)
- Visual flow: Z-pattern, left to right
- Margins: 40-80px
- Layouts: multi-column, left-right split, grid
- Card dimensions (16:9): single-row 530-600px, double-row 265-295px
Canvas dimensions do not imply a title band, content topology, or recurring
chrome. Reserve space only for a real output obstruction. For `story`, keep
meaning-bearing text, identity, and calls to action within `y=120..1740` by
default because common mobile story controls occupy the top and bottom; images,
backgrounds, and nonessential texture may remain full bleed. An exact target-
platform overlay guide or installed template overrides this advisory band.
### Portrait (3:4, 9:16)
- Visual flow: top to bottom
- Margins: 60-120px
- Layouts: single-column, top-bottom split, card stacking
- Card dimensions (3:4): height 400-600px, gap 40-60px
## Typography Scale Start
### Square (1:1)
- Visual flow: center-radiating
- Margins: 60-100px
- Core area: ~800x800px
**Hard rule — normative owner**: This section owns the initial body-size anchor
and sanity band for every registered or custom canvas. Strategist and Quick
consume it directly. Confirm UI maintains an exact executable mirror and must
not infer alternate canvas classes or values. All values are unitless SVG px.
## Format-specific Design
**PPT reading modes**:
| Format | Title Area | Content Area | Special Notes |
|--------|-----------|--------------|---------------|
| PPT | 80-100px | Full width utilization | Page number bottom-right |
| Xiaohongshu (RED) | 180-240px (bold) | Generous top/bottom whitespace | Brand area at bottom 120-160px |
| WeChat Moments | 200-280px | Center 500-600px | QR code area at bottom 150-200px |
| Story | — | Middle 1500px | Top safe zone 120px, bottom 180px |
| WeChat Article Header | Center/left-aligned 48-72px | — | Image on right or as background |
| Canvas | Reading mode | Advisory body band | Initial body |
|---|---|---:|---:|
| `ppt169` / `ppt43` | `text` | 1821 | 20 |
| `ppt169` / `ppt43` | `balanced` | 2225 | 24 |
| `ppt169` / `ppt43` | `presentation` | 2832 | 32 |
> **Body font baseline scales with canvas and reading mode** — a PPT 16:9 baseline confirmed for read-close / business / projection cannot be carried onto tall canvases (Xiaohongshu / Story / A4). Pick the baseline from the confirmed canvas, not the recommended one; see the per-canvas px anchors in [`strategist.md`](strategist.md) §g "Typography Plan Confirmation" (the system is px-only — all sizes are unitless px on every canvas).
**Non-PPT registered and custom canvases**: derive one effective canvas span
from the canonical or custom `W x H`, then calculate the advisory band and
initial body anchor:
```text
short = min(W, H)
long = max(W, H)
span = min(long, 3 * short)
low = round(span * 0.025)
start = round(span * 0.029)
high = round(span * 0.033)
```
| Canvas | Effective span | Advisory body band | Initial body |
|---|---:|---:|---:|
| `wechat` | 900 | 2330 | 26 |
| `moments` | 1080 | 2736 | 31 |
| `xiaohongshu` | 1660 | 4255 | 48 |
| `story` | 1920 | 4863 | 56 |
| `banner` | 1920 | 4863 | 56 |
| `a4` | 1754 | 4458 | 51 |
**Default — starting anchor, not a floor (may override when confirmed identity,
source fidelity, or target viewing conditions require it)**: Start from the
table or formula, then resolve the complete role ramp and page density from the
active content and delivery context. The advisory band only surfaces unusual
values; falling outside it is not a validation failure. Apply the
viewing-distance baseline in
[`shared-standards-core.md`](./shared-standards-core.md) instead of silently
shrinking a recurring role to make content fit.
## ViewBox Examples
@@ -54,7 +54,7 @@ Before the first SVG page, output a confirmation listing: the compact communicat
### 2.1 Execution context validity (Mandatory)
- **Valid**: if the exact complete Design Spec and lock remain in the unchanged, uncompacted active context, reuse both for every page. Do not reread or poll them.
- **Valid**: if the exact complete Design Spec and lock remain in the unchanged, uncompacted active context, reuse both for every page. Do not reread or poll them. The one scheduled exception is the five-page lightweight lock re-read defined below.
- **Invalid**: fresh/resumed/restarted execution, compaction/summary-only recovery, or an external/unknown change requires one complete read of `design_spec.md`, then `spec_lock.md`, plus triggered references/template inputs. Mid-deck recovery also reads the latest completed SVG and, when images are used, current image metadata.
- **Uncertain**: consult the retained lock first, then only the owning Design Spec fragment; use sources only for facts. Design Spec remains upstream on conflict.
@@ -119,6 +119,17 @@ Apply the content-vs-expression contract above within the selected reading mode.
Return upstream before any derived/accent identity becomes recurring or structural, or when an undeclared display size reaches its third occurrence, then update the retained context under §2.1. Local garnish, same-role `±2`px adjustments, and at most two sparse display-size occurrences need no lock row. Never expand the lock to silence a comparison. New icon acquisition, images, structural fonts, role anchors, and resources keep their preparation/role rules.
**Five-page lightweight lock re-read (Default Generate only)**: after
completing P05, P10, P15, … and only when another page follows, read
`spec_lock.md` in full once before starting the next page. This is a pure
context re-anchor for the locked palette, typography, icon style, and
`page_rhythm` values under long context. No checker runs, no per-page
self-check output, no pause, and no mid-deck repair loop. If the re-read
reveals an external/unknown change, follow the Invalid recovery branch above;
otherwise continue directly. Compaction, resume, or restart still follows the
full recovery reads above. Quick Generate does not read a project-root
`spec_lock.md` and does not use this rule.
**Per-page layout rhythm — `page_rhythm` section**:
Before drawing each page, look up its entry in `page_rhythm` (key format `P<NN>` matching the page index in §IX of `design_spec.md`) and apply the corresponding layout discipline:
@@ -129,7 +140,7 @@ Before drawing each page, look up its entry in `page_rhythm` (key format `P<NN>`
| `dense` | Information-heavy. Card grids, multi-column layouts, KPI dashboards, tables, and charts are all permitted. This is the baseline behavior. |
| `breathing` | Low-density impact page. Avoid **multi-card grid layouts** — do not organize content as multiple parallel rounded containers (3-card row, 4-card KPI grid, 2×2 matrix rendered as cards). Use naked text blocks, dividers, whitespace, or full-bleed imagery as the content structure. Single rounded visual elements (hero image corners, callouts, tags, one emphasis block) are fine — the rule is about grid structure, not about the `rx` attribute. Proportions follow information weight (not a preset ratio). Typical forms: hero quote, single large number with one-line interpretation, full-bleed image with floating caption, section transition. |
> Without rhythm variation, every page defaults to card grids (the "AI-generated" look). Context recovery follows §2.1.
> Mechanical repetition comes from reusing the same carrier and topology without a page job—not from semantic cards themselves. Cards remain appropriate when they express real peer grouping, comparison, hierarchy, or capacity; vary rhythm when the content relationship changes. Context recovery follows §2.1.
**Missing or empty `page_rhythm` section — fixed compatibility default** → emit `warning: spec_lock.md missing/empty page_rhythm — defaulting all pages to dense` once, fall back to `dense` for all pages.
@@ -148,11 +159,11 @@ Before drawing each page, look up its entry in `page_rhythm` (key format `P<NN>`
- **Main-agent ownership**: SVG generation must run in the main agent (not sub-agents) — pages share upstream context for cross-page visual continuity
- **Generation rhythm**: P01 → first-page gate → uninterrupted remaining pages → final gate, in one context without batches or mid-run checker calls.
- **Fact provenance**: when a §IX page lists `Fact IDs`, resolve each ID from `sources/*.facts.json` and keep the claim/value unchanged. Render a compact source footnote using the source name and a short URL/domain when space permits; when speaker notes are enabled, state the attribution naturally there too. When §IX says `Data class: scenario`, place a visible localized `Scenario data` / `情景数据` label adjacent to the affected KPI/chart and, when notes are enabled, state naturally there that the number is illustrative. Never attach an external fact ID to scenario data or let an unlabeled invented KPI look factual.
- **Default — stage each page with the style's composition geometry (may override when the content genuinely calls for a plain grid)**: an SVG page is a canvas, not a DOM. Before defaulting to stacked rounded-rect cards or uniform equal columns, pick one page-scale move from the locked visual style's §1 `Composition geometry` (a bleed shape, diagonal split, oversized numeral, orbit rings, …) to stage the page's primary zone. Card grids are one option among many, not the house layout.
- **Default — stage each page with the style's composition geometry (may override when the content genuinely calls for a plain grid)**: an SVG page is a canvas, not a DOM. Resolve the page-scale move from `spec_lock.md`: a preset uses that selected style's §1 `Composition geometry`; `custom` executes `visual_style_behavior` first, then uses §1 geometry only from exact `visual_style_references` that the behavior assigns a shape or composition job. Other bases contribute only their assigned job, and an unreferenced novel custom follows its behavior alone. Before defaulting to stacked rounded-rect cards or uniform equal columns, use that resolved geometry to stage the page's primary zone. Card grids are one option among many, not the house layout.
- **Default — vary a planned deck motif instead of cloning it (may omit where it has no page job)**: when §III `Theme` names a cross-page motif, use the current §IX `Layout` to preserve its recognizable contour, direction, material, or relationship while varying scale, crop, density, position, and content interaction by page role. Apply it only where it supports hierarchy or continuity; do not paste identical ornament or invent a second recurring identity.
- **Inherited containers**: preserve meaningful template frames; restyle radius, fill, stroke, and depth from the active Design Spec and `spec_lock.md`. Selected Chart/Table reference adaptation is owned by [`executor-visualization.md`](./executor-visualization.md); preview effects never override project styling or structural roles.
- **Reference — prefer semantic geometry over preset stacks**: for relationships such as ascending, converging, breaking through, or stacking, first seek a basic primitive, one exact preset, or a clear Boolean result. Only when none can faithfully express the relationship should one page-specific polygon/path replace a stack of generic arrows.
- **Reference — create depth with restraint**: use rhythm, spacing, typography, accent bars, and subtle tints before shadows. Reserve lift for a few genuinely floating elements; keep peer grids, dividers, and body containers flat.
- **Reference — create depth with restraint**: use rhythm, spacing, typography, accent bars, and subtle tints before shadows. Reserve lift for a few genuinely floating elements; keep peer grids, dividers, and ordinary body containers flat. When material layering itself is part of the resolved visual style, follow that style's hierarchy instead of flattening its body planes.
- **Phased generation** (recommended):
1. **Visual Construction Phase**: generate all SVG pages sequentially for visual consistency. Apply every triggered information-model branch while drawing. **MUST embed one object-scoped plot-area marker** per §IX-named or Quick-promoted value-driven chart object under [`executor-chart.md`](./executor-chart.md) §2; coordinate calibration is a post-generation step (see [`verify-charts`](../workflows/stages/verify-charts.md)). Write every `<object-key>=yes` native marker plus JSON metadata atomically under [`native-data-interface.md`](./native-data-interface.md) §2. **Reach for native presets** per §3.0 as you draw each page: a block arrow, chevron, banner/ribbon, callout, standard flowchart node, or star is authored through `preset_shape_svg.py` at draw time — decided by the object's intent as you create it, never by scanning finished paths, and never committed to a bare `<path>`/`<polygon>` when a preset expresses it (a gradient fill/stroke or a pattern fill is the one paint exception — keep those ordinary SVG). **First-page gate (Mandatory)**: after completing the first page, run `python3 ${SKILL_DIR}/scripts/svg_quality_checker.py <project_path> --stage first-page --json` directly without output filtering. Review the whole P01 issue set, make one consolidated edit pass for every error and any selected warnings, then perform one verification rerun. If it still fails, treat that complete output as the next batch; never check between individual fixes. After it passes, draw P02 through the last page without checker calls.
2. **Quality Check Gate**: only after every planned SVG exists, run `python3 ${SKILL_DIR}/scripts/svg_quality_checker.py <project_path> --stage final --json` directly on `svg_output/` without `tail` / `head` / `grep` filtering. One run already reports all pages. Review its complete issue set, fix every `error` plus any selected advisory warnings in one consolidated edit pass, then perform one verification rerun. If it still fails, its complete output begins the next batch cycle; never use checker calls to discover or fix one next issue at a time. Every `warning` is advisory: it never sends the page back for required modification, never authorizes automatic rewriting of compatible user syntax, and needs no acknowledgement/disposition line. Recommendation warnings describe the generated-SVG default; fidelity/quality warnings may be surfaced when material, while the existing input remains releasable. Prototype-identical diagnostics are recorded as `inherited`, source conversion losses as `source-import`, changed/new advisories as `introduced`, and release failures as `blocking` in `validation/svg_quality_report.json`. If release truly depends on a condition, it belongs in `errors`. On success, use the exit status and terminal summary; do not open or `cat` the complete JSON into model context. If terminal output is truncated on failure, read only the relevant issue arrays from the report written by that same run. Do NOT defer error handling to after `finalize_svg.py` — finalize rewrites SVG and masks some violations.
@@ -28,6 +28,20 @@ Construct the chart in this order:
4. Add data labels, axis/category labels, annotations, units, source notes, and visible exceptions from the active page contract.
5. Apply project typography, palette, effects, and container treatment without changing the encoding.
**Perceptual reading**: choose the least ambiguous presentation of the same
authoritative data. Preserve source or semantic order when it carries meaning;
otherwise sort categories for the page's comparison task. Prefer direct series
labels when they remain legible, and keep legends, grid lines, ticks, and other
decoding aids only when they materially reduce lookup or comparison effort.
Comparable panels and small multiples use the same domain, scale, and category
order unless a visibly disclosed difference is itself the message. Bars and
columns whose length compares magnitude start from zero; schedule spans and
other true interval marks retain their authoritative domain. Any non-zero
baseline or axis break must be explicit and must not exaggerate the comparison.
A dual-axis chart is valid only when both series share the exact time/category
domain and the units and visual identities stay unambiguous; otherwise separate
the views.
**Per-object completeness**: preserve every authoritative series, category, point, label, unit, qualifier, source, and scale cue needed to read the chart. When the source cannot determine a required scale or derived value, return the ambiguity upstream in Default or resolve it from explicit source facts in Quick; never fabricate it at draw time.
**Hard rule — schedule geometry**: A schedule is a Gantt chart when dates or
@@ -2,7 +2,7 @@
# Executor Speaker-notes Branch
Conditional late-stage authority for generating the complete speaker-notes document.
Conditional late-stage authority for generating or validating the complete speaker-notes document.
**Trigger**: Default Generate loads this after the final quality check when the
effective Speaker Notes outcome in `design_spec.md §I` is enabled. Quick
@@ -15,6 +15,13 @@ create `notes/total.md`.
Write the complete deck to `notes/total.md` in one batch for coherent transitions. Use `# <number>_<page_title>` per page and `---` between pages; only the heading is stripped before TTS.
**Pre-SVG narration branch**: when `notes/total.md` already exists because the
user supplied a final/literal script or Quick is directly delivering narrated
video, validate it instead of regenerating it. Retain every word and segment of
a final/literal script. Agent-authored Quick narration may be repaired only for
final-SVG inconsistency and before audio generation. A `# Slide <number>`
heading remains valid until Generate Step 7.1 resolves the authored roster.
**Pure spoken narration**: `notes_to_audio.py` reads the body verbatim. Write prose only; never add Markdown list/bullet markup, stage markers, key-point labels, duration lines, or other metadata.
**Length follows content**: size natural sentences to semantic burden. Two to five is typical, not a cap; anchor pages may use less and dense pages more. Honor the active Design Spec or Quick context plus source rules. Duration is pacing guidance only: never pad, repeat, compress, or omit meaning to hit it.
@@ -25,6 +32,13 @@ Write the complete deck to `notes/total.md` in one batch for coherent transition
Before drafting, internally inventory the visible title/subtitle and every information-bearing direct-root `<g id>`; structured placeholder content still counts. Coverage requires its unique claim, evidence, example, relationship, qualifier, or implication—not merely its label—to enter the narration.
For a pre-SVG narration branch, apply the same inventory in reverse: every
independent visible claim or relationship must be supported by its script
segment. Repair the visual page or return to planning for final/literal input;
for agent-authored Quick narration, repair the narration before audio without
inventing unsupported claims. Every spoken idea that requires orientation must
likewise have a visible state or an explicit speech-only role in the active plan.
- Text blocks, comparisons, and processes retain every independent fact or relationship; combine related short groups causally or comparatively.
- Charts, tables, and KPIs state the takeaway, decisive values or trend, comparison basis, implication, and material uncertainty—not every axis, row, or cell.
- Quotes retain the decisive clause, material attribution, and relevance. Explain semantic images or text-free diagrams only from the SVG plus locked plan/source; never infer facts from appearance.
@@ -306,43 +306,97 @@ python3 scripts/slice_images.py <project>/images/illus_sheet.png --grid 2x3 \
---
### 4.4 Registered subject/base pairs
### 4.4 Registered reconstruction groups and shared plates
Use this preparation when a person, product, creature, or other foreground
subject must cross native titles, panels, frames, cards, or shapes while the
original scene remains behind them. The final pair is one compositional asset:
Use this preparation when a person, product, creature, effect, or other scene
element must cross native titles, panels, frames, cards, or shapes while the
original scene remains behind it. A clean base plus one subject/foreground
output is the minimum group; add layers only when overlap or independent
editing requires them:
| Output | Required content |
|---|---|
| Clean base | Full original canvas with the foreground subject removed and the hidden background reconstructed |
| Subject cutout | The same full canvas and subject position, with only the foreground subject visible on RGBA transparency |
| Clean base | Full original canvas with every planned removable scene element removed and the hidden background reconstructed |
| Optional midground | Full canvas with only the scene content that must sit between the base and primary subjects |
| Subject / foreground | Full canvas with one subject or one z-order-compatible set visible on RGBA transparency |
| Shared layer plate | Several mutually non-overlapping objects isolated together in one full-canvas or regular-cell output |
**Mandatory — preserve registration**: Start both derivatives from the same
source. Preserve canvas dimensions, subject pose, scale, and position; do not
trim or independently crop either final output. Record the shared source and
pair relationship in the owning §VIII rows or Quick active-context resources.
**Mandatory — preserve registration**: derive every full-canvas member
independently from the same canonical source. Preserve canvas dimensions,
subject pose, scale, position, lighting, and visible style; do not trim or
independently crop registered final outputs. Record the shared source and group
relationship in the owning §VIII rows or Quick active-context resources.
**Image to PPTX override — Codex required**: when
[`image-to-pptx.md`](../workflows/profiles/image-to-pptx.md) is active, follow
its §3 per-region decision. A complete, separable, final-resolution-sufficient
region may remain source-derived. Otherwise use Codex's native reference-image
capability for required editing or reconstruction. Inspect every prepared
member plus the final recomposition. Do not adapt `image_gen.py`, its manifest,
or provider backends for this profile. Other hosts are unsupported. The
ordinary Path A / Path B procedure below applies outside this profile.
**Preparation procedure**:
1. Use the active image-editing path to remove the subject and inpaint the clean base.
2. Use the same source to isolate the subject. Prefer a direct RGBA result.
3. When the active path cannot return transparency, place the isolated subject on one exact flat key color, then run `slice_images.py` as a `1x1` sheet with `--alpha` and without `--trim` so the full-canvas coordinates remain unchanged.
4. Save both final files under `<project>/images/` and mark them `no-crop` in the active resource authority.
1. From the canonical reference, remove every planned separate subject,
foreground object, source/data graphic, and editable text, then inpaint one
clean base without redesigning visible background content.
2. From that same reference, prepare the subject/foreground content as an
exact source-derived layer or a reference reconstruction according to the
selected profile's source-sufficiency decision. Never derive a layer from
the generated base or another generated layer.
3. Prefer one shared plate when several objects do not overlap and use the same
isolation treatment. Their padded bboxes, including visible shadows and
effects, must be pairwise disjoint. One object does not imply one generation
call.
4. For a registered plate, retain the original full-canvas positions and use
one nested-SVG picture crop per recorded bbox under
[`svg-effects.md`](./svg-effects.md) §6.5. For a rearranged regular-cell
plate, follow §4.3 and run
`slice_images.py --grid ... --names ... --trim --alpha`; place each
resulting asset at its recorded source bbox.
5. Prefer direct RGBA. When transparency is unavailable, use one exact flat
key color for the whole layer/plate, then run `slice_images.py` once as a
`1x1` sheet with `--alpha` and without `--trim` so full-canvas coordinates
remain unchanged. Do not generate one keyed image per object.
6. Save final files under `<project>/images/`. Mark registered full-canvas
members `no-crop`; ordinary trimmed cell slices retain their own measured
dimensions.
Path A may use the existing single-image edit mode for each derivative; Path B
may perform the same edits with the host-native image tool:
Objects that overlap one another or require different z-order use separate
plates/layers. A shared output is valid only when every required final object
still becomes an independent SVG/PPT picture object.
**Shared registered-plate prompt core**:
> Using the supplied canonical page as the only visual reference, isolate the
> following foreground objects together on one full-canvas extraction plate:
> {stable object ids/descriptions}. Preserve each object's visible identity,
> silhouette, pose, scale, rotation, lighting, shadow, and exact original canvas
> position. Keep the original aspect ratio and canvas registration. Retain only
> those listed objects; remove the background and every unlisted element. Do not
> rearrange, resize, merge, duplicate, or let the listed objects touch one
> another. Retain an explicitly listed source graphic or wordmark exactly when
> it is one of the requested objects; remove editable slide text and every
> unlisted logo/source graphic. Return RGBA transparency if supported;
> otherwise use one uniform exact {key HEX} matte with no gradient, texture,
> spill, or extra marks.
Outside Image to PPTX, Path A may use the existing single-image edit mode for
each registered derivative; Path B may perform the same edits with the
host-native image tool:
```bash
python3 scripts/image_gen.py "Remove the foreground subject and reconstruct the hidden background; preserve the exact canvas" \
--reference-image <project>/images/<source>.png -o <project>/images -f <pair>_base
python3 scripts/image_gen.py "Isolate the same foreground subject at the exact original pose, scale, and position on a flat #00FF00 background" \
--reference-image <project>/images/<source>.png -o <project>/images -f <pair>_subject_key
python3 scripts/slice_images.py <project>/images/<pair>_subject_key.png --grid 1x1 \
--names <pair>_subject --alpha --bg "#00FF00"
python3 scripts/image_gen.py "Remove the planned foreground subjects and reconstruct the hidden background; preserve the exact canvas" \
--reference-image <project>/images/<source>.png -o <project>/images -f <group>_base
python3 scripts/image_gen.py "Isolate the planned non-overlapping foreground objects at their exact original positions on one flat #00FF00 plate" \
--reference-image <project>/images/<source>.png -o <project>/images -f <group>_plate_key
python3 scripts/slice_images.py <project>/images/<group>_plate_key.png --grid 1x1 \
--names <group>_plate --alpha --bg "#00FF00"
```
The positional edit commands are allowed here as the declared derivation step
for already-planned pair rows; keep the final pair in the ordinary resource
These positional edit commands remain the declared derivation exception for
already-planned group rows. Keep every final member in the ordinary resource
authority and operational sidecar. SVG realization follows
[`image-layout-patterns.md`](./image-layout-patterns.md) `#A2-03`.
@@ -46,7 +46,7 @@ Compact composition vocabulary for prepared images, illustrations, and rendered
| Several visuals should read as one system | `#P3-05` grid, `#P3-14` mosaic with text cell, `#P3-20` tessellation, `#P3-21` split tiling, `#P3-22` curve array, `#P3-23` depth row |
| A foreground needs an opening or reveal | `#M1-06` true hole, `#M1-07` cut scrim, `#M1-08` background-registered fill, `#M1-05` text subtraction |
| Text needs contrast without discarding the visual | `#M2-01` directional scrim, `#M2-05` spotlight, `#A3-02` prepared frosted panel, `#M2-09` grid scrim |
| A subject should cross or re-layer around native content | `#A2-02` frame breakout or `#A2-03` registered subject/base pair |
| A subject should cross or re-layer around native content | `#A2-02` frame breakout or `#A2-03` registered reconstruction group |
| A cover, divider, or promotional page needs image-led structure | `#P1-01`, `#P1-04`, `#P1-13`, or `#P3-15``#P3-19` |
| Consecutive pages should share one visual world | `#C1-01` persistent state, `#C2-01` pan, `#C2-02` push/pull, or `#C3-01` matched framing |
@@ -178,7 +178,7 @@ The following three patterns are topologically different and are not interchange
- **#A2-01 · Transparent sticker or cutout** — use a prepared RGBA asset and preserve its open silhouette.
- **#A2-02 · Subject breaking out of a container** — register a prepared foreground subject across its frame boundary.
- **#A2-03 · Registered subject/base pair** — align a clean base with its prepared transparent subject cutout in one coordinate system. Draw the base first, the native title/panel/shape second, and the cutout last. Give both images the same `x`, `y`, `width`, `height`, and aspect mapping; never trim or independently crop either layer.
- **#A2-03 · Registered reconstruction group** — align a clean base with one or more prepared transparent midground/subject/foreground layers in one coordinate system. Draw each member at its required z-order. Give every full-canvas member the same `x`, `y`, `width`, `height`, and aspect mapping; never trim or independently crop it. Several padded-bbox-disjoint objects may share one prepared plate while remaining separate nested-SVG picture crops.
### 5.3 A3 · Registered Derivatives
@@ -71,7 +71,7 @@ Every coordinated Stage-2 direction carries one complete `rendering: custom` can
**Hard rule**: three complete rendering candidates are mandatory in every fresh Stage-2 direction set; AI source recommendation remains independent. See [`strategist-image.md`](../strategist-image.md) for the Stage-2 carrier and downstream lock behavior.
Write `image_rendering_references` only when the confirmed custom direction actually uses catalog material. Keep the list exact: a blend of `screen-print` and `watercolor` lists both ids. A genuinely new rendering with no catalog source omits the field and proceeds from its standalone behavior; never invent a reference merely to legitimize `custom`.
Write `image_rendering_references` only when the confirmed custom direction actually uses catalog material. A custom may use zero, one, or many renderings: keep one when it owns the whole specialized treatment, or include every rendering that contributes a distinct executable job across line, texture, depth, material, or mood. Reference count has no fixed cap; count is an outcome, not a target. A four-basis direction may assign `vector-illustration` to silhouette clarity, `minimalist-swiss` to negative-space composition, `screen-print` to restrained halftone texture, and `warm-scene` to light and mood; list all four ids. Omit every rendering whose contribution cannot be stated and never add a second merely to imply synthesis. A genuinely new rendering with no catalog source omits the field and proceeds from its standalone behavior; never invent a reference merely to legitimize `custom`.
---
@@ -68,7 +68,7 @@ Each mode keeps its own authoritative file with: narrative skeleton, page-struct
**Quick custom**: do not display a candidate set. Use `custom` only when a project-specific specialization or fusion serves the deck better than one preset; retain the behavior and exact bases in active context and persist nothing.
**Mandatory — select before detail reading**: Use this index to freeze every catalog source actually used, then read only those exact files before writing the behavior. A `pyramid` + `narrative` fusion therefore reads those two files and writes `mode_references: pyramid, narrative` beside `mode_behavior`; Quick retains those bases only in active context. Do not open candidates for comparison after this gate or add loosely related references after the fact. A genuinely new cadence names and reads no catalog source.
**Mandatory — select before detail reading**: Use this index to freeze every catalog source actually used, then read only those exact files before writing the behavior. A custom may use zero, one, or many sources: keep one when it owns the whole specialized cadence, or include every mode that owns a distinct executable act, posture, title voice, rhythm, or register. Reference count has no fixed cap; count is an outcome, not a target. A three-basis direction may use `pyramid` for a conclusion-first opening, `narrative` for the risk-tension act, and `instructional` for the closing action sequence; it reads those three files and writes all three ids beside `mode_behavior`. Quick retains its bases only in active context. Omit every source whose contribution cannot be stated, never add a second merely to imply synthesis, and do not open candidates for comparison after this gate. A genuinely new cadence names and reads no catalog source.
> **One value per deck — fusion is *one* `custom`, not several modes.** A deck always resolves a single `mode`. A multi-mode blend is expressed as **one** custom behavior whose paragraph describes the acts — never as several simultaneous modes.
>
@@ -16,11 +16,15 @@ Mandatory reference for every route that authors or regenerates slide visuals th
|---|---|
| Text-block rhythm | Use §4.2 leading. Make the baseline step into a new paragraph visibly larger than the intra-paragraph line step; keep the extra gap between list items smaller than paragraph separation but large enough to scan each item. Repeated peer blocks share one rhythm unless their hierarchy differs. |
| Typography roles | Use the fewest semantic text roles that preserve hierarchy, and make their differences legible at slide-thumbnail scale. Consolidate near-neighbor sizes that serve the same role; otherwise distinguish roles through a deliberate combination of size, weight, color, position, and surrounding space. |
| Viewing-distance legibility | Resolve delivery context and viewing distance before fixing density and type scale. Preserve necessary text at a readable scale by applying only actions the active route's content and page invariants permit: restructure, shorten, split, or reflow. If none is permitted, surface the unresolved fit instead of silently miniaturizing it. Do not turn this into one universal font-size floor: captions and metadata may be smaller when their role and context remain legible. |
| Contrast and semantic encoding | Within the active profile's fidelity boundary, keep meaning-bearing text distinguishable from its actual background. For newly authored distinctions, combine luminance, weight, scale, shape, position, or explicit labeling; color may reinforce meaning but never carry a required distinction alone. When fidelity requires preserving source-only color encoding, reproduce it rather than inventing a cue. Reserve lower-contrast treatment for genuinely secondary metadata that remains legible. |
| Natural wrapping | Break at semantic phrase or punctuation boundaries where possible. Reflow the text frame or adjust neighboring geometry before using any permitted local size reduction. Let the final line run naturally shorter; avoid mechanically equal lines or a stranded single-character / single-word line when an earlier natural break preserves meaning. |
| Content field | Establish the usable body frame before placing modules. Divide it into a small set of unequal-weight macro-regions from information weight and reading order, then give each region its own local axes / micro-grid while retaining only the cross-region anchors the composition needs. On a dense page, let the planned content system organize that frame; create breathing room through gutters, module spacing, and intentional voids between semantic clusters. Unorganized residual space that leaves content stranded in one part of the frame is leftover blank, not negative space. |
| Content field | Establish the usable body frame before placing modules. Divide it into one or a small set of macro-regions from information weight and reading order: use unequal weight when the information differs, while true peers may share equal weight. Give each region its own local axes / micro-grid while retaining only the cross-region anchors the composition needs. On a dense page, let the planned content system organize that frame; create breathing room through gutters, module spacing, and intentional voids between semantic clusters. Unorganized residual space that leaves content stranded in one part of the frame is leftover blank, not negative space. |
| Alignment and proximity | Establish shared axes from the current composition. Align related titles, copy, labels, images, and diagram nodes to those edges, centers, or baselines; group related elements more tightly than unrelated groups so spacing carries hierarchy. Break an axis only when the offset performs hierarchy, direction, or tension. |
| Visual weight | Judge weight from area, darkness, saturation, density, stroke, image detail, and elevation together. Distribute it to support the focal path; symmetry is optional, and deliberate imbalance may create direction. |
| Containers | Use a card or panel when it expresses grouping, hierarchy, boundary, capacity, or a distinct material plane. Otherwise prefer spacing, rules, or direct text / geometry; peer containers share treatment unless a semantic difference justifies contrast. |
| Boundary strength | Match the relationship with the lightest sufficient boundary from this expressive ladder: spacing / alignment → rule / bracket → tint field → outline → filled panel → true floating layer. Peer relationships use comparable strength while focus, hierarchy, or material difference may move to a stronger treatment. The ladder is not a required sequence or per-page quota. |
| Containers | Use a card or panel when it expresses grouping, hierarchy, boundary, capacity, or a distinct material plane. Otherwise prefer spacing, rules, or direct text / geometry; peer containers share treatment unless a semantic difference justifies contrast. An unplanned repeated web-card grid is a carrier / topology problem, not a reason to suppress meaningful borders, shapes, or containers. |
| Titles and page chrome | Treat the semantic page title as part of the current composition rather than an automatic fixed header band; its position, scale, and relationship may change with page role while preserving the active route's content invariants. Add or retain running headers, footers, and page numbers only when they carry navigation, identity, attribution, or another explicit page job. Fidelity profiles preserve required source chrome. |
**Default — active effects vocabulary (may resolve to no added technique when no visual job is diagnosed)**: Default and Quick Generate run the already-loaded [`svg-effects.md`](./svg-effects.md) §6.1 Visual Job Router before completing each page; apply a compatible technique only for a diagnosed visual job.
@@ -150,9 +154,10 @@ Unknown or unmapped declarations fail Checker preflight and native export.
> **PPT preset patterns and native chart/table/template metadata are
> conditional** — see [`native-data-interface.md`](./native-data-interface.md) and [`pptx-structure-interface.md`](./pptx-structure-interface.md).
DrawingML has no arbitrary per-pixel alpha-compositing path. Effects that rely
on one, including text-knockout image fills and arbitrary alpha composites,
must be baked into a raster asset before SVG export.
DrawingML has no arbitrary per-pixel alpha-compositing path. A registered
single-image text picture/texture fill follows [`svg-effects.md`](./svg-effects.md)
§6.3; arbitrary text-knockout composites, multi-layer image text, and arbitrary
alpha composites remain bake-required before SVG export.
---
@@ -18,13 +18,13 @@ For illustration, apply this precedence: confirmed `none` → explicit user inte
**Default — one coherent sheet for compatible same-family spots (may override when aspect, detail, quality, or semantic needs differ)**: prefer one Illustration Sheet when several AI-generated spots can share a useful cell shape and production treatment; generate them independently when forcing one sheet would weaken a planned element. When a sheet is chosen, plan one unplaced `ai` Illustration Sheet row plus one placed `slice` row per used element; only slice rows enter `spec_lock.md images`. State the intended placement shape family in the sheet reference and use separate sheets for incompatible shapes. [`image-generator.md`](./image-generator.md) §4.3 owns grid, ratio, slicing, and execution details. Final Stage 2 chooses the AI execution path under `image-generator.md` §7; do not pre-empt or re-pick it here.
**Mandatory — subject-layer capability scan**: When a supplied reference or the intended composition shows a subject crossing a native title, panel, frame, or shape, plan a registered subject/base pair before writing §VIII. Add one clean full-canvas base row and one same-canvas RGBA subject row, give both `Crop Policy: no-crop`, name their shared source and registration relationship in `Reference`, and suggest `#A2-03` for the subject row. Use `Acquire Via: user` only when both final assets are already supplied; otherwise use `ai` for the prepared derivatives. [`image-generator.md`](./image-generator.md) §4.4 owns preparation. A simple floating cutout that does not re-layer over its source may use `#A2-01` instead.
**Mandatory — subject-layer capability scan**: When a supplied reference or the intended composition shows a subject crossing a native title, panel, frame, or shape, plan a registered reconstruction group before writing §VIII. Add one clean full-canvas base row and the minimum same-canvas RGBA midground/subject/foreground rows, give full-canvas members `Crop Policy: no-crop`, name their shared source and registration relationship in `Reference`, and suggest `#A2-03`. Padded-bbox-disjoint objects may share one generated plate if the final SVG gives each one an independent picture crop. Use `Acquire Via: user` only when every final asset is already supplied; otherwise use `ai` for the prepared derivatives. [`image-generator.md`](./image-generator.md) §4.4 owns preparation. A simple floating cutout that does not re-layer over its source may use `#A2-01` instead.
## 2. AI Image Strategy — always propose three; lock only for confirmed `ai`
Before any rendering detail, read only [`image-renderings/_index.md`](./image-renderings/_index.md). First author exactly three complete, project-fit solution intents; use the index to freeze each intent's exact rendering bases, then read once only the deduplicated referenced sibling files. Project one complete `image_strategy` into each direction regardless of `recommend.image_usage`. Every candidate carries localized `name`, `rendering: custom`, `visual`, `mood`, and non-empty localized `behavior`. Mood includes a recognizable real-world analogy. All three must credibly serve their owning whole solution, but they need not use different bases or span artificial safe / shifted / bold extremes. Image colors always inherit that direction's deck HEX roles; never add an image palette or alter deck colors to rescue a rendering.
Every direction is a `custom` rendering. If it specializes, combines, or borrows existing renderings, name every exact index-selected id in its visible behavior and read only those files after selection; if it is genuinely novel, name no basis and read none. Under a template it obeys inherited identity and application. Only a confirmed custom locks its edited behavior as `image_rendering_behavior`; when catalog material is actually used, also project the exact ids as `image_rendering_references`, otherwise omit that field. Unselected candidates remain recommendation-only. Do not write a separate fourth `custom_candidates.image_strategy`; ignore legacy `image_palette`.
Every direction is a `custom` rendering. It may use zero, one, or many index-selected bases: one may own a specialized treatment, while several must each own a distinct line, texture, depth, material, or mood contribution. Name every actual id in the visible behavior and read only those files after selection; reference count has no fixed cap, and a second basis is never required. If the direction is genuinely novel, name no basis and read none. Under a template it obeys inherited identity and application. Only a confirmed custom locks its edited behavior as `image_rendering_behavior`; when catalog material is actually used, also project the exact ids as `image_rendering_references`, otherwise omit that field. Unselected candidates remain recommendation-only. Do not write a separate fourth `custom_candidates.image_strategy`; ignore legacy `image_palette`.
The UI hides these candidates while AI is not selected. If the user adds AI, it reveals the already-authored three without another backend recommendation; source selection never creates or rewrites rendering candidates. After confirmation, Image_Generator reads only the selected preset or exact custom references and must not blend unselected candidate identities.
@@ -36,7 +36,7 @@ solution + production gate:
| **1 — communication contract + template choice** | `primary_language` · `c` audience · open-ended communication intent · audience outcome · core message / delivery context (primary + optional secondary) / artifact afterlife · `content_divergence` (all prose fields may be blank) · `a` canvas · explicit `free_design` or `templates` choice and selected roots | confirmed together; candidate workspaces do not influence the communication recommendation |
| **2 — final solution + production** (authored once from the user's *actual* Stage 1) | reading mode (`delivery_purpose`, PPT only) · `d` mode + visual style · `b` page count · `e` color · `f` icon · `g` typography · `h` image source + generated-image rendering · conditional natural-language template application · formula policy · conditional AI-image acquisition path · generation mode · refine-spec toggle · proactive speaker notes / custom animations / narration audio | derived as one coherent plan from the confirmed contract; internal template exporter modes remain hidden |
Do not force communication intent into one catalog label; Stage 1 records composite intent in prose. Editable prose fields are recommendation drafts, not required inputs: confirmation preserves current text and blanks; never repopulate a cleared field. Stage 2 confirms narrative spine, reading density, page budget, visual system, image direction, production mechanics, and how any installed template should be used. It never chooses or installs a template. Inspect only project-local template spec/prototypes, present one editable application plan, and keep exporter reuse/adherence internal. First author exactly three complete, project-fit solution directions from the confirmed contract and source; only then project each direction into mode, visual style, color, type, icons, and generated-image rendering for lower-level adjustment. Every direction projects a project-specific `custom` mode, `custom` visual style, and `custom` generated-image rendering; the fixed catalogs remain conservative lower-level single-select alternatives. All three must be viable and distinguishable as whole solutions, but do not force safe / shifted / bold archetypes, different catalog bases, or artificial extremes. Every direction carries a complete generated-image rendering candidate even when AI imagery is not recommended; `recommend.image_usage` independently decides whether AI is proposed. Generated images inherit deck colors—there is no second image palette. Proactive defaults are speaker notes `true`, custom animations `false`, and narration audio `false`; a prior explicit user instruction overrides the matching recommendation, and effective narration audio requires effective speaker notes. Author each stage once; same-stage edits update only visible browser state through documented deterministic dependencies, without another AI/backend recommendation. Launch/derive/wait mechanics live in [`generate-pptx.md`](../workflows/generate-pptx.md) Step 4; item specs keep `a``h`.
Do not force communication intent into one catalog label; Stage 1 records composite intent in prose. Editable prose fields are recommendation drafts, not required inputs: confirmation preserves current text and blanks; never repopulate a cleared field. Stage 2 confirms narrative spine, reading density, page budget, visual system, image direction, production mechanics, and how any installed template should be used. It never chooses or installs a template. Inspect only project-local template spec/prototypes, present one editable application plan, and keep exporter reuse/adherence internal. First author exactly three complete, project-fit solution directions from the confirmed contract and source; only then project each direction into mode, visual style, color, type, icons, and generated-image rendering for lower-level adjustment. Every direction projects a project-specific `custom` mode, `custom` visual style, and `custom` generated-image rendering; the fixed catalogs remain conservative lower-level single-select alternatives. All three must be viable and distinguishable as whole solutions, but do not force safe / shifted / bold archetypes, different catalog bases, or artificial extremes. After all three bundles are complete, compare them against the confirmed contract and source, choose the strongest overall fit, and write its actual zero-based index as `design_directions.selected` (`0`, `1`, or `2`); array order never determines preference. Every direction carries a complete generated-image rendering candidate even when AI imagery is not recommended; `recommend.image_usage` independently decides whether AI is proposed. Generated images inherit deck colors—there is no second image palette. Proactive defaults are speaker notes `true`, custom animations `false`, and narration audio `false`; a prior explicit user instruction overrides the matching recommendation, and effective narration audio requires effective speaker notes. Author each stage once; same-stage edits update only visible browser state through documented deterministic dependencies, without another AI/backend recommendation. Launch/derive/wait mechanics live in [`generate-pptx.md`](../workflows/generate-pptx.md) Step 4; item specs keep `a``h`.
**Hard rule — Stage-1 source boundary**: Build the communication recommendation only from the current user request, source facts, conversation constraints, and project-initialization state. Author it before loading index summaries for a chat listing, and do not read any candidate spec, prototype, asset, or template-owned canvas. The same Stage-1 surface may display template controls, but their values are confirmation state, not recommendation evidence. Do not load or apply [`strategist-template.md`](./strategist-template.md) until Stage 1 is confirmed and the selected workspace has been installed for Stage 2.
@@ -44,7 +44,7 @@ Do not force communication intent into one catalog label; Stage 1 records compos
>
> **One opt-in exception**: present the refinement line with the split-mode note ([`generate-pptx.md`](../workflows/generate-pptx.md) Step 4). Only explicit opt-in runs [`refine-spec`](../workflows/stages/refine-spec.md): write the Design Spec once, pass Gate 1, then stop before the lock for unrestricted chat revision. Never enter it unprompted.
> **Default presentation surface — Confirm UI.** Before the first actual confirmation phase, apply [`confirm_ui.md`](../scripts/docs/confirm_ui.md)'s sticky per-run surface decision; its explicit chat branch skips every UI command, and a chat selection after UI launch follows its in-run switch procedure. Chat-question tools alone do not select a branch. In the UI branch, `template_options.json` and `recommendations.stage1.json` open one Stage-1 page; its single submission writes the pure Strategist `result.json` plus the sidecar `template_selection.json`. After installation/free-design closure, `template_handoff.json` gates `.stage2.json`. The chat/delegated branch keeps equivalent state without fabricating those receipts. Replace only the active unconfirmed stage and print the URL plus combined Stage-1 summary/fallback without treating that handoff as confirmation. Stage 1 writes canonical BCP-47 `primary_language` apart from UI `lang`; Strategist projects it through Design Spec §I to lock communication. Stage 2 carries exactly three immutable `design_directions`, each with a unique stable id, `custom` mode, `custom` visual style, six-role HEX palette, primary-language heading/body typography plus an English companion only for non-English decks, icons, and `custom` generated-image rendering. A direction card applies its whole bundle; lower controls may then diverge through the three projected custom candidates or the fixed single-select catalogs, and clicking that card again restores the authored bundle. The final result stores only the current component values, never the direction id as execution authority. Step 4 retains final confirmation from the selected channel for Design Spec authoring. `confirm_ui.md` owns selection-surface and staged-confirmation lifecycle.
> **Default presentation surface — Confirm UI.** Before the first actual confirmation phase, apply [`confirm_ui.md`](../scripts/docs/confirm_ui.md)'s sticky per-run surface decision; its explicit chat branch skips every UI command, and a chat selection after UI launch follows its in-run switch procedure. Chat-question tools alone do not select a branch. In the UI branch, `template_options.json` and `recommendations.stage1.json` open one Stage-1 page; its single submission writes the pure Strategist `result.json` plus the sidecar `template_selection.json`. After installation/free-design closure, `template_handoff.json` gates `.stage2.json`. The chat/delegated branch keeps equivalent state without fabricating those receipts. Replace only the active unconfirmed stage and print the URL plus combined Stage-1 summary/fallback without treating that handoff as confirmation. Stage 1 writes canonical BCP-47 `primary_language` apart from UI `lang`; Strategist projects it through Design Spec §I to lock communication. Stage 2 carries exactly three immutable `design_directions`, each with a unique stable id, `custom` mode, `custom` visual style, six-role HEX palette, primary-language heading/body typography plus an English companion only for non-English decks, icons, and `custom` generated-image rendering. Its `selected` index marks the Strategist's post-comparison preference and initializes the whole bundle. An inactive direction card applies its bundle; lower controls may then diverge through the three projected custom candidates or the fixed single-select catalogs, and the adjusted active card exposes an explicit restore action. The final result stores only the current component values, never the direction id as execution authority. Step 4 retains final confirmation from the selected channel for Design Spec authoring. `confirm_ui.md` owns selection-surface and staged-confirmation lifecycle.
**Confirmed-value semantics**: confirmation preserves both the value and the owning field's semantic type. Apply the type to the affected property, not automatically to the whole object:
@@ -123,7 +123,7 @@ When authoring §IX, translate every purpose named in `communication_intent` int
Two independent layers, each locks one preset or `custom`. Output: `d. Mode: <mode> + Visual style: <visual_style>`.
> **Top-down custom direction construction.** Author three complete solution intents from the confirmed project contract and source before selecting any catalog basis; do not assemble three apparent solutions from independent mode/style/rendering picks. At that point only the three indexes may be in context. Use their summaries to freeze the exact reference ids for each direction, then read once only the deduplicated union of those referenced detail files and author the final behaviors. Every direction MUST serialize `mode: custom`, `visual_style: custom`, and `image_strategy.rendering: custom`, each with visible, non-empty behavior prose. A catalog-based custom names only its actual bases; a genuinely novel custom names none and reads no detail file. These three projections are the project-specific choices above the conservative fixed catalogs, not a fourth Custom proposal. Never glob a catalog, read an unselected sibling, or write bespoke prose as an enum value.
> **Top-down custom direction construction.** Author three complete solution intents from the confirmed project contract and source before selecting any catalog basis; do not assemble three apparent solutions from independent mode/style/rendering picks. At that point only the three indexes may be in context. Use their summaries to freeze the exact reference ids for each direction, then read once only the deduplicated union of those referenced detail files and author the final behaviors. Every direction MUST serialize `mode: custom`, `visual_style: custom`, and `image_strategy.rendering: custom`, each with visible, non-empty behavior prose. A custom may use zero, one, or many catalog bases: one may specialize a strong dominant basis, several may divide distinct executable jobs, and a genuinely novel custom uses none. Reference count has no fixed cap and is an outcome, not a target; omit every basis whose contribution the behavior cannot state, and never add a second basis merely to make the result look synthesized. These three projections are the project-specific choices above the conservative fixed catalogs, not a fourth Custom proposal. Never glob a catalog, read an unselected sibling, or write bespoke prose as an enum value.
#### Layer 1 — Communication mode
@@ -137,7 +137,7 @@ The deck's **narrative + persuasion skeleton** — how the argument is organized
- Each direction crystallizes one project-specific cadence / posture in `mode_behavior`: it may specialize one preset, fuse several modes into a multi-act sequence, or be genuinely novel. A catalog-based direction retains only the exact ids it actually uses; a novel direction invents no basis. One deck locks one `custom` value, never several simultaneous modes.
- No user structure or cadence → derive each whole solution from the confirmed `communication_intent`, `audience_outcome`, source texture, and delivery context, then project its custom mode. The three directions may share a catalog basis when that is honestly best; distinguish them through project-specific behavior or other fields instead of forcing different bases.
Record the confirmed mode and rationale in `design_spec.md` first, including the exact catalog basis when a selected custom uses one. Then project `- mode:` to `spec_lock.md`; for `custom`, also project `- mode_behavior:` and, only when catalog material is actually used, `- mode_references: <id>, <id>`. Executor reads only those exact references; an unreferenced novel custom follows the behavior directly.
Record the confirmed mode and rationale in `design_spec.md` first, including every exact catalog basis when a selected custom uses any. Then project `- mode:` to `spec_lock.md`; for `custom`, also project `- mode_behavior:` and, only when catalog material is actually used, `- mode_references: <id>[, <id> ...]`. Executor reads only those exact references; an unreferenced novel custom follows the behavior directly.
#### Layer 2 — Visual style
@@ -147,13 +147,13 @@ The deck's **visual aesthetic** — shape language, decoration density, whitespa
**Source**:
- User named a style (chat / template / beautify) → it is truth: retain it as the required basis or inherited anchor in every custom behavior. Keep the three whole solutions meaningful by varying only fields the user left open; when visual variation is forbidden, the three style behaviors may be identical.
- No user description → author three project-fit whole solutions first, then project one complete custom aesthetic for each. Directions may share catalog bases when their overall systems still differ meaningfully. Do not force different bases, a safe-to-bold ladder, or one deliberately extreme option merely to manufacture variety. Give each direction a localized name and concise real-world outcome note; the Confirm UI exposes the three project-specific styles above all 18 fixed manual alternatives.
- No user description → author three project-fit whole solutions first, then project one complete custom aesthetic for each. Directions may share catalog bases when their overall systems still differ meaningfully. Do not force different bases, a safe-to-bold ladder, or one deliberately extreme option merely to manufacture variety. Give each direction a localized name and use its localized note as a compact, user-facing style summary. The note may reuse localized display labels from Confirm UI's `visual_styles` catalog (for example, `瑞士极简`, `柔和圆角`, or `编辑出版`) when they concisely describe the result, but these labels are optional vocabulary, not a selection constraint or required mapping. Use concise natural language wherever the catalog wording does not fit, and never force the nearest label. Keep the summary to one or two short sentences without exposing catalog ids or reference mechanics. The Confirm UI exposes these three project-specific styles above all 18 fixed manual alternatives.
**Forbidden — a non-catalog name as `visual_style`**: every direction recommendation uses literal `custom` for `visual_style`; bespoke prose belongs only in `visual_style_behavior`, while optional `visual_style_references` contain only first-column catalog ids. A name from the `_index` "Paired rendering" column (`flat`, `vector-illustration`, `digital-dashboard`, `3d-isometric`, `corporate-photo`, …) is an image-rendering id, not a style reference. Generic words such as flat / modern / clean / simple / minimal are also insufficient behavior: use the index to choose an exact basis when applicable, then state the project-specific shape language, composition, density, whitespace, typography, and texture.
**Carries no color.** A visual style governs how the deck's HEX (locked at `e`) is *used* — never which colors, same discipline as [`image-renderings`](./image-renderings/_index.md). When the deck has AI images, prefer the style's paired rendering so layout and illustration share one aesthetic.
Record the confirmed visual style and rationale in `design_spec.md` first, including the exact catalog basis when a selected custom uses one. Then project `- visual_style:` to `spec_lock.md`; for `custom`, also project `- visual_style_behavior:` and, only when catalog material is actually used, `- visual_style_references: <id>, <id>`. Executor reads only those exact references; an unreferenced novel custom follows the behavior directly.
Record the confirmed visual style and rationale in `design_spec.md` first, including every exact catalog basis when a selected custom uses any. Then project `- visual_style:` to `spec_lock.md`; for `custom`, also project `- visual_style_behavior:` and, only when catalog material is actually used, `- visual_style_references: <id>[, <id> ...]`. Executor reads only those exact references; an unreferenced novel custom follows the behavior directly.
**Conditional template workspace**: When the Stage-1 template choice has been installed into `<project_path>/templates/`, read [`strategist-template.md`](./strategist-template.md) before completing Stage 2. Read the installed project-local spec and prototypes only; never reopen the library/external source root. The module owns the editable natural-language application plan, confirmed-value consumption, AI-authored prototype selection, internal reuse/adherence derivation, inherited design precedence, and structured-lock planning. This plan decides how to use the installed template, never which template to select. Bare names, style words, and free-design projects do not trigger it.
@@ -229,13 +229,13 @@ See [`../templates/icons/README.md`](../templates/icons/README.md) for the curre
**Size anchors — px only**: Every authoring layer carries bare px numbers. PowerPoint's displayed pt is an export result (`px × 0.75`), never an input or confirmation value.
| Reading mode on PPT | Initial body | Information posture |
|---|---:|---|
| `text` | 20 | read-close / dense |
| `balanced` | 24 | mixed reading + presentation |
| `presentation` | 32 | projected / sparse |
Other canvases use the body baseline in [`canvas-formats.md`](canvas-formats.md). The confirmed role-anchor values always win: take Confirm UI `body_size` / `sizes` verbatim as anchors; a manually edited anchor remains pinned, and changing canvas does not secretly rescale it.
**Mandatory — canvas-owned body start**: Read
[`canvas-formats.md`](canvas-formats.md) § "Typography Scale Start" before
authoring size candidates. It owns the initial body anchor and sanity band for
PPT reading modes plus registered/custom non-PPT canvases; do not reproduce or
rederive them here. The confirmed role-anchor values always win: take Confirm UI
`body_size` / `sizes` verbatim as anchors; a manually edited anchor remains
pinned, and changing canvas does not secretly rescale it.
| Recurring role | Ratio to body |
|---|---:|
@@ -384,6 +384,13 @@ When enabled, match SVG names where possible (`01_cover.svg` →
`notes/01_cover.md`); `notes/slide01.md` remains compatible. Split files contain
no `#` heading lines; `notes/total.md` uses `#` headings.
**Prepared final narration**: when the user explicitly marks a script as
final/literal and intends it for notes or generated audio, preserve its wording
and order. Segment it by semantic scene while resolving §IX, record its source
and verbatim policy in §X `Content`, and let Generate write the frozen
`notes/total.md` only after the final roster and lock pass their gates. Do not
copy the full script into on-slide `Content` or rewrite it as visible body text.
---
## 2. Mode & Visual-Style Catalogs (Reference for Confirmation Item d)
@@ -491,6 +498,7 @@ includes transitions.
| Natural-language template application | §I records it and the relevant layout/prototype choices realize it without silently dropping a requested use or exclusion |
| Formula policy, AI-image acquisition path, generation mode, refine-spec toggle | §I records them as production mechanics; their owning Generate stage consumes the Design Spec, and formula policy also shapes §VIII when formula-worthy content exists |
| Proactive speaker notes, custom animations, and narration audio | §I records the three resolved effective outcomes with provenance, while §X records enabled note requirements or `Generation: disabled`; they remain outside `spec_lock.md`. §IX Motion suggestions remain optional advice regardless of the animation outcome |
| Explicit final/literal narration script | §IX segments the argument by semantic scene and gives each segment a supporting visible state; §X records the source plus verbatim policy, and Generate freezes the actual segments in `notes/total.md` after Gate 2 |
**GATE 1 — active-decision fidelity.** Do not create `spec_lock.md` until the initial Design Spec passes the comparison above and any enabled refinement is explicitly approved. Before Gate 2, every requested revision must be present and every unaffected decision intact. Missing/substituted values, unapplied revisions, or silently changed semantic types block despite schema validity; bounded Reference adaptation and unused Permission remain valid.
@@ -57,7 +57,7 @@ different jobs.
| Picture/card/overlay elevation or boundary is unclear | Object or picture/carrier shadow, restrained glow, or hairline | §6.4; equal peers stay flat; one light direction |
| Native copy and image do not integrate | Scrim, fade, wash, vignette, off-center spotlight, or faux glass | §6.5 and the Image-Treatment Implementation Map; verify contrast; no backdrop blur |
| Relationship state, direction, continuity, or boundary is unclear | Draft/optional/future → dash; direction → marker; undirected → solid; continuous flow → gradient stroke; repeated boundary → frame/contour/crop edge; exact grid → multi-subpath | §6.6 / §6.3; every line needs a job |
| Short display text needs notation or silhouette | Removed/former → strike; eyebrow distinction → tracking; display silhouette → outline/gradient; luminous metric → glow; semantic list → native bullet | §6.7 / §6.4; no decorative body-copy treatment |
| Short display text needs notation, silhouette, or material/image emphasis | Removed/former → strike; eyebrow distinction → tracking; display silhouette → outline/gradient; justified material/image emphasis → native picture/texture fill; luminous metric → glow; semantic list → native bullet | §6.7 / §6.3 / §6.4; no decorative body-copy treatment |
| Tilt, repetition, or reversible asset direction helps composition | Rotate, translate/mirror, or local `<use>` | §6.8; never mirror text, logos, or directional evidence |
| Resolved style needs hand, print, pixel, facets, layers, ribbon, or line-plus-area | Matching constructed recipe | §6.11; no generic decorative freeform |
| Meaning needs an unmatched silhouette, radial hierarchy, gauge, or custom route | Freeform, explicit arc/sector, or calculated arrowhead | §6.9 / §6.10; prefer an equal stock shape/marker |
@@ -212,6 +212,40 @@ has both dimensions. Checker and exporter reject the degenerate stroke form.
stroke="url(#flow)" stroke-width="12"/>
```
**Native text picture/texture fill**:
| Concern | Contract |
|---|---|
| Target | Direct `fill="url(#id)"` on `<text>` or a non-positional `<tspan>`; the text remains editable |
| Definition | Direct `<pattern>` child of `<defs>` with unique `id` and exact `data-pptx-text-image-fill="stretch"` or `"tile"` |
| Image | Exactly one direct SVG-namespace `<image>` child; project-local or data-URI source; explicit positive `width` / `height` |
| Native result | `stretch` → run-level `a:blipFill/a:stretch`; `tile` → run-level `a:blipFill/a:tile` |
| Alpha | Text `fill-opacity` multiplies the native picture-fill alpha |
| Forbidden | Preset-pattern attributes; `patternTransform`; additional pattern children; image style/alpha/clip/filter/mask/transform; use outside text; unannotated custom image patterns; multi-image/layer knockout composites |
Use this registered pattern when the design calls for a photograph, material,
or texture inside editable glyphs. The pattern is an authoring carrier for a
PowerPoint run picture fill, not a general SVG pattern promise. PowerPoint owns
the final run bounding box: `stretch` is `Native-normalized`, while `tile` may
normalize tile scale or phase and needs visual review. Forward SVG→PPTX export
is native; PPTX→SVG does not reconstruct run-level picture fills yet.
```xml
<defs>
<pattern id="titleTexture" data-pptx-text-image-fill="stretch"
patternUnits="objectBoundingBox"
patternContentUnits="objectBoundingBox" width="1" height="1">
<image href="../images/cloud-texture.png"
x="0" y="0" width="1" height="1"
preserveAspectRatio="none"/>
</pattern>
</defs>
<text x="96" y="220" fill="url(#titleTexture)" fill-opacity="0.85"
font-family="Microsoft YaHei" font-size="72" font-weight="700">
国风之美
</text>
```
Preset patterns are a separate PPT interface in [`native-data-interface.md`](./native-data-interface.md).
---
@@ -0,0 +1,214 @@
# Video-delivery Design Reference Manual
Conditional design guidance for presentations whose intended use is a recorded,
self-running, or video delivery.
**Trigger**: load this reference when the effective delivery purpose is video,
recorded narration, or unattended playback. Also load its script rules for an
explicit final/literal narration input. Speaker notes, animation, or audio
requested for an otherwise ordinary deck do not activate it alone. Explicit
video/MP4 delivery does; Quick additionally activates §3's direct-delivery
contract.
**Ownership**: this is a conditional Generate reference, not a profile or a new
artifact route. Default keeps its Strategist and confirmation flow; Quick keeps
its one-pass active-context flow. Existing notes, animation, audio, and native
PowerPoint export stages retain their schemas and commands.
When Beautify is active, its wording/page/order invariants still bind; apply
this reference only inside the design and motion freedom that profile permits.
---
## 1. Intake and Script State
Classify supplied spoken material before planning:
| Material | Treatment |
|---|---|
| Ordinary source or rough transcript | Use as source material; edit, condense, and reorganize under the selected route's normal content-divergence contract |
| Explicit final/literal narration script | Preserve every spoken word and its order; segment only at semantic scene boundaries |
| SRT used to generate new TTS | Preserve cue text when it is explicitly final; use source timecodes only as pacing evidence because the new synthesis timing becomes authoritative |
| SRT bound to an existing recording | Preserve its text/audio timing authority; do not regenerate TTS or pretend that one long recording was split automatically |
| Already page-separated final script | Preserve the supplied page boundaries unless the user explicitly permits restructuring |
| Target platform, canvas, or duration | Use the existing canvas registry; resolve scene granularity, page count, and notes length together |
**Hard rule — final means explicit**: freeze wording only when the user identifies
the script as final, literal, or verbatim. Never promote ASR output, subtitles,
or a draft transcript into a literal contract by inference.
**Default — semantic segmentation (may override for a user-authored page
plan)**: one scene represents one coherent visual state or mental-map step, not
one sentence, subtitle cue, or effect. Several cues may share a scene; one scene
may contain several ordered reveals.
**Final-script production input**: after the page roster is final but before SVG
authoring, write the resolved per-slide script once to `notes/total.md`. Use
`# Slide <number>` headings and `---` separators so the file can exist before SVG
filenames do; preserve each body segment verbatim. It is a production input, not
a storyboard or substitute Design Spec. Run `total_md_split.py` only after the
SVG roster exists.
---
## 2. Scene and Page Planning
Judge quality by reduced audience understanding cost, not spectacle. Resolve the
spoken argument, visible evidence, mental map, and motion job together before
choosing effects.
| Narrative relationship | Page treatment |
|---|---|
| Several lines explain one idea | Keep one page/scene and reveal only the semantic units needed for that explanation |
| One process or system persists | Retain its main roles and relative positions as a stable visual anchor |
| New evidence expands one part of a known map | Keep the map, enlarge or emphasize the active region, and de-emphasize context only as needed |
| The same object changes position, scale, containment, or state | Author compatible adjacent endpoints and use the Morph contract when motion is active |
| The audience must adopt a genuinely new mental map | Start a new composition and make the transition explicit |
**Default — stable visual anchors (may override when the subject resets)**:
within one explanation, keep recurring roles, relative positions, and governing
structure stable. Prefer moving, enlarging, revealing, or dimming within that
map over replacing the entire layout at every beat.
**Default — one semantic focus change per beat (may override for one inseparable
idea)**: several elements may change together only when they communicate one
unit. Do not change unrelated title, diagram, annotation, and footer regions at
the same moment merely to make the frame busier.
**Default — scene chrome earns its place (may override for navigation, identity,
attribution, or fidelity)**: for newly authored recorded, self-running, or video
scenes, do not carry a report-style fixed header, footer, or page number merely
by deck convention. Let the semantic title participate in the scene composition,
and omit nonessential running chrome especially on cover, ending, and breathing
scenes. Retain source or template chrome when the active profile's fidelity
boundary requires it, and retain new chrome when it genuinely orients the
audience or carries required identity or attribution.
**Default — screen for orientation, notes for speech (may override for literal
on-screen copy)**: place keywords, structure, evidence, and relationships on the
slide; keep full explanation in notes. Do not duplicate the narration script as
body copy.
**Page-count rule**: derive page count from semantic scenes, visual-state
endpoints, and target duration. Never derive it from subtitle-cue or sentence
count.
---
## 3. Default and Quick Planning Handoff
**Default**: Stage 1 confirms the existing open-text `delivery_context`; it does
not ask a separate video question. When the confirmed value identifies
recorded/self-running/video delivery, load this reference before authoring the
three Stage-2 whole solutions. Apply its scene grammar to every direction; it
does not add a style catalog or confirmation field. Record delivery context and
afterlife in §I, visible states and optional motion jobs in §IX, and script/notes
policy plus target duration in §X. When the final-script branch is active,
create the frozen `notes/total.md` after the approved roster/lock is final and
before Step 5 or split-mode handoff.
**Default — reading mode (may override for durable close reading)**: recorded
explanation leans `presentation`; choose `balanced` when close-reading afterlife
materially outweighs video delivery.
**Quick**: there is no Stage 1 or separate video-purpose confirmation. Explicit
video/recorded/self-running intent activates this reference after source
sufficiency is known and before the one-pass roster, resource, and motion
decisions; absent that intent, keep ordinary Quick behavior. Load the script
rules alone when an explicit final/literal narration will become notes/audio.
Keep the applicable scene grammar and final-script handling in active context.
A pre-SVG `notes/total.md` is an enabled production artifact, not a forbidden
planning checkpoint; Quick still creates no root Design Spec, lock,
confirmation payload, or storyboard.
**Mandatory — Quick direct video input**: when Quick must deliver a narrated
video or MP4 rather than only a deck for later recording, enable Speaker Notes,
Narration Audio, and video export; write the complete per-scene narration to
`notes/total.md` before P01 and use it as page-design input. After the SVG
roster, only agent-authored wording may be finalized; final/literal input remains
verbatim. Before audio, choose static/page-transition-only,
narration-independent deck-wide exporter motion, or page/object-specific Custom
Animations; for the last, decide whether narration governs any group timing.
**Production outcomes**:
| Need | Decision |
|---|---|
| Spoken delivery or a supplied final script | Enable Speaker Notes |
| User asks the workflow to synthesize narration | Enable Narration Audio; Speaker Notes is its dependency |
| Progressive reveal, continuing geometry, or timed emphasis materially aids explanation | Enable/load the appropriate animation capability |
| Quick directly delivers a narrated video or MP4 | Enable Speaker Notes, Narration Audio, and video export; resolve motion before audio, requiring timestamped page-local SRT only for narration-cue sync or subtitle delivery |
| Video is manually narrated or static playback is sufficient | Keep Narration Audio and/or object animation off as applicable |
**Capability boundary**: a deck intended for later video use does not force
object animation or generated audio. Direct Quick narrated-video delivery uses
the mandatory row above but may intentionally remain static or
page-transition-only. Explicit user instructions remain authoritative.
---
## 4. SVG, Notes, and Motion Realization
When §3 created `notes/total.md` before SVG, read it once before the first SVG
and design each page around its corresponding spoken segment. Give every
independently narrated or timed semantic unit a descriptive direct-root `<g
id>`; keep inseparable units grouped. Preserve a final/literal script exactly;
agent-authored direct-video narration may change only during its final-SVG
validation before audio.
**Hard rule — script/design consistency**: a final script is literal content.
If the finished visual page introduces an independent claim or relationship the
script does not explain, repair the page or return to planning; never rewrite or
pad the final script during the late notes pass. Conversely, every spoken idea
that requires visual orientation must have a visible state or deliberate
speech-only treatment.
**Motion readiness**: load `animations.md` before SVG authoring whenever the
plan needs compatible Morph endpoints or page/object-specific motion. Author
every required start/end state and real semantic group before the final checker;
post-processing cannot invent missing visual endpoints or target IDs.
**Motion restraint**: use transitions, reveals, emphasis, and Morph only for a
named communication job. There is no motion-coverage quota, and `effect: none`
remains valid. Auto-running narration uses `after-previous` / `with-previous`,
never `on-click`.
**Mandatory when narration governs object motion**: before SVG authoring, load
`animations.md` and preserve real semantic groups; before audio, create and
validate canonical `animations.json`. After the base PPTX/report and timestamped
page audio/SRT, map timed groups in `narration_timing.json`, derive
`narration_animations.json`, and export the narrated PPTX/MP4. Only derived
triggers/delays wait for SRT; object identity, effect, and order do not. `-a
auto` or inherited fixed stagger is not semantic synchronization. For
static/page-transition-only or narration-independent deck-wide motion, omit
these sidecars and the object-sync claim.
**Sound effects**: add them only on explicit request and only from prepared,
project-local assets. Do not introduce sound search, trimming, or normalization
as an implicit video step.
**Production sequence**: after the final SVG check, validate any pre-SVG
narration against the visible pages; ordinary draft-source runs instead use the
final-SVG-grounded notes generation. Split notes, execute the resolved motion
path, and export the editable PPTX. Direct Quick video continues
through audio and, when required, timestamped SRT. Custom Animations use the
narrated-sidecar flow when narration governs group timing;
narration-independent custom motion exports its canonical timing without an
object-sync claim before the narrated PPTX and MP4.
---
## 5. Delivery Boundary
**Canonical artifact**: the editable PPTX remains canonical. `generate-audio` owns provider/voice/rate
selection, page audio/SRT generation, semantic narration timing, narrated PPTX
export, and optional native PowerPoint video export.
**Conditional MP4**: run `powerpoint_video.py --check` only when MP4 delivery is
selected. If native Windows PowerPoint export is unavailable, keep the narrated
PPTX as the successful upstream artifact; do not substitute screenshots, HTML,
or a third-party renderer and call it equivalent.
**Current boundary**: importing and automatically splitting one long finished
recording is unsupported. Require page-level audio or an explicit page/time map;
otherwise deliver the designed deck and frozen notes without claiming audio
integration.
@@ -21,7 +21,7 @@ If the static checker has not been run or has failed, the subagent must abort wi
Each review subagent processes a **batch** of pages (see §6.1 for batch sizing). The inputs are:
1. **Page batch** — a list of `(svg_path, png_path, page_role)` tuples, one per assigned page. `svg_path` resolves under `<project>/svg_output/<page>.svg`, `png_path` under `<project>/.preview/<page>.png`. `page_role` is one of `cover` / `chapter` / `tldr` / `content` / `data` / `closing` / `breathing`, parsed from `design_spec.md §IX` by the orchestrator — subagents do **not** guess.
1. **Page batch** — a list of `(svg_path, png_path, page_role, canvas)` records, one per assigned page. `svg_path` resolves under `<project>/svg_output/<page>.svg`, `png_path` under `<project>/.preview/<page>.png`. `page_role` is one of `cover` / `chapter` / `tldr` / `content` / `data` / `closing` / `breathing`, parsed from `design_spec.md §IX` by the orchestrator — subagents do **not** guess. `canvas` is copied verbatim from that page's successful renderer record and contains the root-SVG `view_box`, exact `width` / `height`, and raster `png_width` / `png_height`; subagents never assume a fixed canvas or recompute it from the PNG.
2. **Path to this rubric file**
3. **`<project>/design_spec.md`** (read-only) — §IX outline is the source of truth for "what should this page deliver"
4. **`<project>/spec_lock.md`** (read-only) — brand-locked values
@@ -36,13 +36,13 @@ Style Review Focus is supplemental acceptance context, not a second rubric. It c
| # | Category | Trigger | Permitted fix |
|---|----------|---------|---------------|
| H1 | Out-of-bounds | element bbox falls outside `0,0,1280,720` | shrink or reposition into canvas |
| H1 | Out-of-bounds | element bbox falls outside the bounds declared by `canvas.view_box` | shrink or reposition into canvas |
| H2 | Text overflow | text bbox extends past its visual container | reduce font-size or line-break |
| H3 | Text overlap | two `<text>` elements' bboxes intersect (tspans within one text excluded) | reposition or resize |
| H4 | Readability | contrast < 4.5 (small text) / < 3.0 (font-size ≥ 24px); OR text directly atop a complex image with no scrim | if **neither** the foreground nor the background color is a brand token: position-only escape — add a `<rect>` scrim under the text, or raise the offending text's font-size to ≥ 24px so the 3.0 threshold applies. If **either** color is a brand token: do not edit the SVG → goto §1.1 escalation. |
| ~~H5~~ | Font-ramp drift | *covered by `svg_quality_checker.py` — see §0 prerequisites* | n/a (do not re-check) |
| H6 | Element collision | rect/circle/path bboxes overlap with z-order violating semantics | open spacing |
| H7 | Anchored element displaced | page number / header / footer covered, missing, or out of canvas | restore to anchor position |
| H7 | Declared page chrome displaced | page number / header / footer is explicitly declared by `design_spec §IX`, `spec_lock.md`, or the installed template with a concrete anchor, but is covered, missing, or outside `canvas.view_box` | restore only that declared chrome to its declared anchor; never invent undeclared chrome |
| H8 | Image rendering broken | `<image>` empty / broken-image / severe distortion | fix `href`; for `adaptive`, choose `meet` or a safer crop; a new complete-display requirement returns to §VIII `Crop Policy` and lock projection |
| H9 | Missing key element | element required by `design_spec §IX` outline is absent from rendered slide | recreate from spec |
@@ -73,11 +73,11 @@ Subagents must apply the **明显** ("clearly bad") threshold — when in doubt,
|---|----------|---------|---------------|
| S1 | Vertical rhythm tight | Within the **same logical text block**, consecutive baselines have gap < 1.05× larger font-size | open to 1.151.3× |
| S2 | Vertical rhythm hollow | Within one logical block, > 150 px non-decorative whitespace; `breathing` pages exempt | tighten |
| S3 | Visual centroid off | hero/title block centroid offset from canvas center exceeds threshold by `page_role`: `cover` > 35%, `chapter` > 25%, `tldr`/`closing`/`breathing` > 25%, `content`/`data` > 20% | shift toward intended anchor |
| S3 | Intended anchor missed | hero/title block is clearly displaced from a concrete anchor explicitly declared by `design_spec §IX`, `spec_lock.md`, or the installed template. Canvas center counts only when the plan explicitly calls for centered placement; intentional asymmetry, negative space, and image-focal placement are exempt. | restore toward the declared anchor |
| S4 | Alignment drift | same-column elements differ in `x` by > 4 px (or same-row baselines by > 4 px) **and** are semantically meant to be on the same grid line | snap to grid |
| S5 | Grid non-uniform | N-card row: neighbor `x`-spacing differs by > 5% of the average | re-distribute |
| S6 | CJK letter-spacing | CJK characters with `letter-spacing / font-size > 5%` | reduce to ≤ 2% |
| S7 | Accent overload | > 2 accent colors across ≥ 3 distinct elements | collapse to 1 primary + 1 secondary |
| S7 | Decorative accents compete | multiple unlocked colors with no semantic role visibly compete with one another and with the intended primary emphasis. Brand colors, natural-media colors, data-series/category colors, status encoding, and other semantic colors are exempt. | consolidate only the competing decorative colors into existing page/deck accents; never recolor locked or semantic elements |
| S8 | Emphasis mismatch | most visually prominent element ≠ the element `design_spec §IX` declares as the page's primary | rescale to match intent |
| S9 | Image-text relationship | caption > 60 px from its image; text on busy image without scrim; image clearly purposeless | tighten / add scrim / remove |
| S10 | Breathing violation | only when `page_role = breathing`: ≥3 rounded card grid | replace with naked text / single hero |
@@ -101,7 +101,7 @@ If a "violation" requires reinterpreting `design_spec.md` to fix → mark `needs
Run before applying any rule:
- PNG file exists and is non-zero bytes
- PNG dimensions = 1280 × 720
- PNG dimensions = that page record's `canvas.png_width` × `canvas.png_height`
- PNG is **not** all-background (a histogram check: count of background-color pixels < 99% of total) — guards against blank/white-out renders only, **does not** filter sparse dark layouts
Any check fails → status = `render_failed`, abort without scanning rules.
@@ -145,6 +145,13 @@ Each subagent writes exactly one file to `<project>/.review/<page>.json`:
{
"page": "02_three_steps.svg",
"page_role": "content",
"canvas": {
"view_box": [0, 0, 1242, 1660],
"width": 1242,
"height": 1660,
"png_width": 1242,
"png_height": 1660
},
"status": "ok" | "fixed" | "needs_human" | "render_failed" | "prereq_failed",
"iterations_run": 1,
"screenshot_paths": [
@@ -198,7 +205,7 @@ This rubric is consumed by subagents spawned via the `visual-review` stage. Mand
The orchestrator partitions the N pages into `ceil(N/K)` batches of ≤ K pages each (default **K = 5**; configurable per run via the orchestrator prompt) and spawns one subagent per batch.
- Spawn all batch subagents in **one assistant message** (parallel `Agent` calls). Sequential dispatch breaks pipelining.
- Each subagent prompt is **self-contained** — no prior conversation context. Inline the absolute paths for §0.1 inputs 15 explicitly, plus the full `(svg_path, png_path, page_role)` list for that batch. Do not assume the subagent knows the project root.
- Each subagent prompt is **self-contained** — no prior conversation context. Inline the absolute paths for §0.1 inputs 15 explicitly, plus the full `(svg_path, png_path, page_role, canvas)` records for that batch. Do not assume the subagent knows the project root.
- `subagent_type: general-purpose`. Tool restrictions: Read, Edit, Bash (for `cp` backups), Write (for JSON output). MCP playwright is **not** required by subagents — orchestrator pre-renders PNGs.
- `name` / `team_name` parameters may be unavailable from nested teammate context. Dispatch must remain functional with anonymous subagents — do not require named addressing.
@@ -232,7 +239,8 @@ Larger K is **not** always better: subagent context fills with prior pages' SVG
`visual_review.py <project> [pages...]` must guarantee:
- Output PNG matches what the user would see in the live-preview browser (inlined `<use data-icon>`, resolved `<image href>`)
- Output dimensions = 1280 × 720
- The root SVG `viewBox` is the canvas source of truth; each successful page record includes its exact `view_box`, `width` / `height`, and raster `png_width` / `png_height`
- Output dimensions = that page record's `png_width` × `png_height`
- File-lock serialization at `<project>/.preview/.render.lock`
- Clean exit codes:
- `0` — all requested pages rendered
@@ -96,7 +96,7 @@ Each coordinated Default Stage-2 direction authors one visible, non-empty `custo
Quick does not display a candidate spectrum. It reads this index, resolves one preset or custom behavior, then reads only the selected detail files and persists nothing.
**Mandatory — select before detail reading**: Freeze every catalog source actually used from this index, then read only those exact files before writing the behavior. Default persists those ids as `visual_style_references`; Quick retains them only in active context. Do not open candidates for comparison after this gate or attach references merely because they are adjacent. A genuinely new aesthetic names and reads no catalog source.
**Mandatory — select before detail reading**: Freeze every catalog source actually used from this index, then read only those exact files before writing the behavior. A custom may use zero, one, or many sources: keep one when it owns the whole specialized aesthetic, or include every style that contributes a distinct executable job across shape language, composition, decoration, whitespace, typography, or texture. Reference count has no fixed cap; count is an outcome, not a target. A coherent three-basis direction may assign `swiss-minimal` to grid and whitespace, `soft-rounded` to selective surface contours and elevation, and `editorial` to evidence hierarchy and rules. Default persists every actual id as `visual_style_references`; Quick retains them only in active context. Omit every source whose contribution cannot be stated, never add a second merely to imply synthesis, and do not open candidates for comparison after this gate. A genuinely new aesthetic names and reads no catalog source.
---
@@ -11,10 +11,14 @@ to fill with judgment (hidden shapes, combo charts, overcrowded pages, ...).
Usage:
python3 scripts/beautify_inventory.py <slide_library.json> [--images <image_manifest.json>] [-o inventory.json]
python3 scripts/beautify_inventory.py <inventory.json> --summary
python3 scripts/beautify_inventory.py <inventory.json> --page N [--with-geometry]
Examples:
python3 scripts/beautify_inventory.py projects/x/analysis/<stem>.slide_library.json \
--images projects/x/images/image_manifest.json -o projects/x/analysis/beautify_inventory.json
python3 scripts/beautify_inventory.py projects/x/analysis/beautify_inventory.json --summary
python3 scripts/beautify_inventory.py projects/x/analysis/beautify_inventory.json --page 7
Dependencies:
None (standard library only).
@@ -28,13 +32,23 @@ import argparse
import json
import sys
from pathlib import Path
from typing import Optional
from typing import Any, Optional
from console_encoding import configure_utf8_stdio
configure_utf8_stdio()
_GEOMETRY_KEYS = {
"geometry",
"display_ratio",
"display_left_emu",
"display_top_emu",
"display_width_emu",
"display_height_emu",
}
def _images_by_slide(manifest: list) -> dict[int, list[dict]]:
"""Map slide_index -> [image entries on that slide], from ppt_to_md occurrences."""
by_slide: dict[int, list[dict]] = {}
@@ -138,35 +152,172 @@ def build_inventory(slide_library: dict, images_by_slide: dict[int, list[dict]])
}
def _view_payload(inventory: dict, view: str, slides: list[dict]) -> dict:
"""Wrap projected slides with the canonical deck-level facts."""
return {
"schema": "beautify_inventory.view.v1",
"inventory_schema": inventory.get("schema", "beautify_inventory.v1"),
"view": view,
"source": inventory.get("source"),
"slide_count": inventory.get("slide_count", len(inventory.get("slides", []))),
"canvas_px": inventory.get("canvas_px"),
"slides": slides,
}
def _summary_view(inventory: dict) -> dict:
"""Project the whole roster into compact per-slide counts and review flags."""
slides = []
for slide in inventory.get("slides", []):
slides.append({
"slide_index": slide.get("slide_index"),
"page_type": slide.get("page_type"),
"text_block_count": len(slide.get("text_blocks", [])),
"table_count": len(slide.get("tables", [])),
"chart_count": len(slide.get("charts", [])),
"diagram_count": len(slide.get("diagrams", [])),
"image_count": len(slide.get("images", [])),
"ignored": slide.get("ignored", []),
"needs_confirmation": slide.get("needs_confirmation", []),
})
return _view_payload(inventory, "summary", slides)
def _without_geometry(value: Any) -> Any:
"""Remove explicit source-layout geometry while preserving content and data."""
if isinstance(value, dict):
return {
key: _without_geometry(item)
for key, item in value.items()
if key not in _GEOMETRY_KEYS
}
if isinstance(value, list):
return [_without_geometry(item) for item in value]
return value
def _page_view(inventory: dict, slide_index: int, with_geometry: bool) -> Optional[dict]:
"""Project one slide by its canonical slide_index."""
for slide in inventory.get("slides", []):
if slide.get("slide_index") != slide_index:
continue
projected = slide if with_geometry else _without_geometry(slide)
return _view_payload(inventory, "page", [projected])
return None
def _positive_int(value: str) -> int:
"""Parse a positive slide index for argparse."""
parsed = int(value)
if parsed < 1:
raise argparse.ArgumentTypeError("page must be a positive integer")
return parsed
def build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(
description="Merge slide_library.json + image_manifest.json into a per-slide beautify inventory.",
description=(
"Build a beautify inventory from a slide library, or print a read-only "
"model view from a canonical beautify inventory."
),
formatter_class=argparse.RawDescriptionHelpFormatter,
)
parser.add_argument("slide_library", help="slide_library.json from `template_fill_pptx.py analyze`")
parser.add_argument("--images", help="image_manifest.json from `ppt_to_md.py` (optional)")
parser.add_argument("-o", "--output", help="Write JSON here (default: stdout)")
parser.add_argument(
"input_json",
help="slide_library.json in builder mode, or beautify_inventory.v1 JSON in a view mode",
)
parser.add_argument(
"--images",
help="image_manifest.json for a slide_library input (optional)",
)
parser.add_argument(
"-o",
"--output",
help="Write the full built inventory here (builder mode only; default: stdout)",
)
view_group = parser.add_mutually_exclusive_group()
view_group.add_argument(
"--summary",
action="store_true",
help="Print deck facts, per-slide object counts, and review flags to stdout",
)
view_group.add_argument(
"--page",
type=_positive_int,
metavar="N",
help="Print one slide's frozen content/data by slide_index to stdout",
)
parser.add_argument(
"--with-geometry",
action="store_true",
help="Retain source-layout geometry in a --page view",
)
return parser
def main(argv: Optional[list[str]] = None) -> int:
args = build_parser().parse_args(argv)
parser = build_parser()
args = parser.parse_args(argv)
lib_path = Path(args.slide_library)
if not lib_path.is_file():
print(f"[ERROR] slide_library not found: {lib_path}", file=sys.stderr)
view_requested = args.summary or args.page is not None
if view_requested and args.output:
parser.error("--summary/--page write to stdout and cannot be combined with --output")
if args.with_geometry and args.page is None:
parser.error("--with-geometry requires --page")
input_path = Path(args.input_json)
if not input_path.is_file():
print(f"[ERROR] input JSON not found: {input_path}", file=sys.stderr)
return 1
slide_library = json.loads(lib_path.read_text(encoding="utf-8"))
input_data = json.loads(input_path.read_text(encoding="utf-8"))
input_schema = input_data.get("schema")
if input_schema == "beautify_inventory.view.v1":
parser.error(
"beautify_inventory.view.v1 is a read-only stdout projection and "
"cannot be used as input"
)
is_inventory = input_schema == "beautify_inventory.v1"
if is_inventory and not view_requested:
parser.error(
"beautify_inventory.v1 input requires --summary or --page; "
"builder mode requires slide_library.json"
)
if is_inventory and args.images:
parser.error("--images cannot be combined with a beautify_inventory.v1 input")
if view_requested and not is_inventory:
parser.error("--summary/--page require a canonical beautify_inventory.v1 input")
manifest: list = []
if args.images:
if args.images and not is_inventory:
img_path = Path(args.images)
if not img_path.is_file():
print(f"[ERROR] image manifest not found: {img_path}", file=sys.stderr)
return 1
manifest = json.loads(img_path.read_text(encoding="utf-8"))
inventory = build_inventory(slide_library, _images_by_slide(manifest))
inventory = input_data if is_inventory else build_inventory(
input_data,
_images_by_slide(manifest),
)
if args.summary:
print(json.dumps(_summary_view(inventory), ensure_ascii=False, indent=2))
return 0
if args.page is not None:
page_view = _page_view(inventory, args.page, args.with_geometry)
if page_view is None:
available = ", ".join(
str(slide.get("slide_index"))
for slide in inventory.get("slides", [])
)
print(
f"[ERROR] slide_index {args.page} not found; available: {available or 'none'}",
file=sys.stderr,
)
return 1
print(json.dumps(page_view, ensure_ascii=False, indent=2))
return 0
payload = json.dumps(inventory, ensure_ascii=False, indent=2)
if args.output:
@@ -1520,13 +1520,16 @@ def _stage2_design_directions_error(
*,
main_language: object = '',
) -> Optional[str]:
"""Require exactly three complete top-down custom Stage 2 design systems."""
"""Require three complete custom systems and a valid preferred direction."""
main_language = main_language or _recommendation_language(recommendations)
directions = recommendations.get('design_directions')
if isinstance(directions, dict):
candidates = _candidate_list(directions)
if len(candidates) != 3:
return 'Stage 2 design_directions must include exactly 3 candidates'
selected = directions.get('selected', 0)
if type(selected) is not int or not 0 <= selected < len(candidates):
return 'Stage 2 design_directions.selected must be an integer from 0 to 2'
typography_candidates = []
direction_ids = set()
for index, candidate in enumerate(candidates, start=1):
@@ -64,10 +64,11 @@
sec_mode: "Generation mode",
sec_refine: "Review the Design Spec first",
sec_design_directions: "Coherent design directions",
design_directions_hint: "Choose a complete direction first, then fine-tune the projected fields below. Clicking a direction again restores its authored bundle.",
design_directions_hint: "The recommended complete direction is applied first. Choose another or fine-tune the projected fields below; use Restore to return an adjusted direction to its authored bundle.",
direction_active: "Applied",
direction_adjusted: "Adjusted",
direction_reset_hint: "Click to apply or restore this complete direction.",
direction_apply_hint: "Click to apply this complete direction.",
direction_restore: "Restore authored direction",
scheme_component_options: "Project-specific custom choices · select a card to edit",
sec_template_application: "Template application",
template_application_hint: "The AI recommends how to apply the installed template to this deck. Revise the plan directly in natural language.",
@@ -147,7 +148,7 @@
body_size_unit_relation: "SVG px to PPT pt: 1px = 0.75pt.",
body_size_pt_hint: "Approximately {pt} pt (1px = 0.75pt; saved as px).",
role_size_pt_hint: "≈ {pt} pt",
body_size_hint_canvas: "This canvas suggests ~{lo}{hi}px (scales with canvas height).",
body_size_hint_canvas: "This canvas suggests ~{lo}{hi}px (from its effective canvas span).",
body_size_hint_purpose: "This reading mode recommends {def}px — one fixed size, not a range.",
body_size_hint_oor: "(Current value is outside the usual range for this canvas — check the unit is right and that it fits.)",
delivery_purpose: "Reading mode",
@@ -246,10 +247,11 @@
sec_mode: "生成モード",
sec_refine: "先に設計仕様を確認",
sec_design_directions: "統合デザイン方針",
design_directions_hint: "まず全体案を選び、その下に反映された各項目を微調整ます。同じ全体案をもう一度押すと、元の組み合わせに戻ります。",
design_directions_hint: "おすすめの全体案が最初に適用されています。別案を選ぶか、下の各項目を微調整できます。調整後は「元の案に戻す」で最初の組み合わせを復元できます。",
direction_active: "適用中",
direction_adjusted: "調整済み",
direction_reset_hint: "クリックすると、この全体案を適用または元の状態に戻します。",
direction_apply_hint: "クリックすると、この全体案を適用します。",
direction_restore: "元の案に戻す",
scheme_component_options: "プロジェクト専用カスタム案 · カードを選んで編集",
sec_template_application: "テンプレートの適用方法",
template_application_hint: "AIが現在の内容に合わせたテンプレートの使い方を提案します。自然言語で直接修正できます。",
@@ -329,7 +331,7 @@
body_size_unit_relation: "SVG px と PPT pt の換算:1px = 0.75pt。",
body_size_pt_hint: "約 {pt} pt1px = 0.75pt 換算、保存は px)。",
role_size_pt_hint: "約 {pt} pt",
body_size_hint_canvas: "このキャンバスの目安は約{lo}–{hi}px(キャンバスの高さに応じて変化)。",
body_size_hint_canvas: "このキャンバスの目安は約{lo}–{hi}px(有効キャンバス尺度から算出)。",
body_size_hint_purpose: "この閲覧モードの推奨は{def}px — 範囲ではなく固定値です。",
body_size_hint_oor: "(現在の値はこのキャンバスの通常範囲外です — 単位とサイズ感を確認してください。)",
delivery_purpose: "閲覧モード",
@@ -428,10 +430,11 @@
sec_mode: "生成模式",
sec_refine: "先审核设计规范",
sec_design_directions: "成套设计方向",
design_directions_hint: "先选择一套完整方案,在下方微调各项;再次点击同一方案会恢复它原本的整套预设。",
design_directions_hint: "AI 最倾向的成套方案已默认应用;你可以改选其他方案,在下方微调各项。调整后可用“恢复原方案”还原整套预设。",
direction_active: "已应用",
direction_adjusted: "已调整",
direction_reset_hint: "点击应用或恢复这套完整方案。",
direction_apply_hint: "点击应用这套完整方案。",
direction_restore: "恢复原方案",
scheme_component_options: "项目专属自定义方案 · 选中卡片后可编辑",
sec_template_application: "模板应用方式",
template_application_hint: "AI 会根据当前内容推荐如何使用已安装模板;你可以直接用自然语言修改。",
@@ -511,7 +514,7 @@
body_size_unit_relation: "SVG px 与 PPT pt 的换算:1px = 0.75pt。",
body_size_pt_hint: "约 {pt} pt(按 1px = 0.75pt 换算;提交仍保存 px)。",
role_size_pt_hint: "约 {pt} pt",
body_size_hint_canvas: "当前画布建议 ~{lo}{hi}px随画布高度缩放)。",
body_size_hint_canvas: "当前画布建议 ~{lo}{hi}px按有效画布跨度计算)。",
body_size_hint_purpose: "该阅读模式推荐 {def}px(单一固定值,非区间)。",
body_size_hint_oor: "(当前数值超出该画布的常用范围——请确认单位无误、是否合适。)",
delivery_purpose: "阅读模式",
@@ -724,6 +727,16 @@
return node;
}
function fitTextareaToContent(input) {
if (!input) return;
window.requestAnimationFrame(function () {
if (!input.isConnected || input.offsetParent === null) return;
input.style.height = "auto";
var borderHeight = input.offsetHeight - input.clientHeight;
input.style.height = (input.scrollHeight + borderHeight) + "px";
});
}
function previewNode(kind, id) {
var node = el("div", "option-preview option-preview-" + kind);
node.setAttribute("aria-hidden", "true");
@@ -1625,6 +1638,13 @@
return spec.candidates || spec.options || [];
}
function selectedDesignDirectionIndex() {
var candidates = designDirectionCandidates();
var selected = Number(designDirectionSpec().selected);
if (!isFinite(selected) || selected < 0) selected = 0;
return Math.min(Math.floor(selected), Math.max(candidates.length - 1, 0));
}
function designDirectionId(candidate, index) {
var value = candidate && candidate.id;
return String(value || ("direction-" + (Number(index) + 1)));
@@ -1632,9 +1652,7 @@
function selectedDesignDirection() {
var candidates = designDirectionCandidates();
var selected = Number(designDirectionSpec().selected || 0);
if (!isFinite(selected) || selected < 0) selected = 0;
return candidates[Math.min(selected, Math.max(candidates.length - 1, 0))] || {};
return candidates[selectedDesignDirectionIndex()] || {};
}
function directionField(field) {
@@ -1811,6 +1829,19 @@
};
}
function directionCardSummary(candidate) {
candidate = candidate || {};
var strategy = normalizedImageStrategy(candidate.image_strategy || {});
return String(
localized(candidate, "note") ||
directionBehavior(candidate, "visual_style") ||
directionBehavior(candidate, "mode") ||
strategy.visual ||
strategy.behavior ||
""
);
}
function usesCustomImagePlanValue(value) {
var ids = (CAT.image_usage || []).map(function (item) { return item.id; });
if (Array.isArray(value)) return false;
@@ -1849,8 +1880,12 @@
function imageStrategySelectedIndex() {
var spec = imageStrategySpec();
var direct = spec.candidates || spec.options || [];
var idx = direct.length ? (spec.selected || 0) : (designDirectionSpec().selected || 0);
return Math.min(idx, Math.max(imageStrategyRecommendationCandidates().length - 1, 0));
var idx = direct.length ? Number(spec.selected || 0) : selectedDesignDirectionIndex();
if (!isFinite(idx) || idx < 0) idx = 0;
return Math.min(
Math.floor(idx),
Math.max(imageStrategyRecommendationCandidates().length - 1, 0)
);
}
// ---- section renderers ----------------------------------------------
@@ -1994,7 +2029,7 @@
STATE.typography.body_size = typography.body_size ||
defaultBodySizeForCanvas(STATE.canvas, STATE.delivery_purpose);
STATE.typography.sizes = Object.assign({}, typography.sizes || {});
syncUnpinnedTypographySizes(true);
syncUnpinnedTypographySizes(false);
}
if (candidate.icons) STATE.icons = normalizeRecId("icons", candidate.icons);
if (candidate.image_strategy) {
@@ -2012,28 +2047,51 @@
var sec = section("B", "sec_design_directions", t("design_directions_hint"));
var grid = el("div", "font-grid design-direction-grid");
var cardStates = [];
var recommendedIndex = selectedDesignDirectionIndex();
candidates.forEach(function (candidate, idx) {
var card = el("div", "font-card design-direction-card");
card.title = t("direction_reset_hint");
card.title = t("direction_apply_hint");
var head = el("div", "font-card-head");
head.appendChild(el("span", "font-card-name",
localized(candidate, "name") || (t("option_prefix") + " " + (idx + 1))));
if (idx === recommendedIndex) {
head.appendChild(el("span", "rec-badge", "★ " + t("recommended")));
}
var status = el("span", "rec-badge direction-status");
status.style.display = "none";
head.appendChild(status);
card.appendChild(head);
var customVisual = candidate.visual_style === "custom";
if (candidate.visual_style) {
var preview = el("div", "design-direction-preview");
appendVisualStyleImage(preview, candidate.visual_style);
if (customVisual) {
preview.classList.add("design-direction-custom-preview");
preview.appendChild(
el("div", "design-direction-custom-label", t("custom"))
);
preview.appendChild(el(
"div",
"design-direction-custom-copy",
directionCardSummary(candidate) || t("custom")
));
} else {
appendVisualStyleImage(preview, candidate.visual_style);
}
card.appendChild(preview);
}
var meta = [];
if (candidate.mode) meta.push(directionComponentValueLabel(candidate, "mode"));
if (candidate.visual_style) {
if (candidate.mode && candidate.mode !== "custom") {
meta.push(directionComponentValueLabel(candidate, "mode"));
}
if (candidate.visual_style && !customVisual) {
meta.push(directionComponentValueLabel(candidate, "visual_style"));
}
var typographyName = candidate.typography &&
(localized(candidate.typography, "name") || candidate.typography.name);
if (typographyName) meta.push(typographyName);
if (candidate.icons) meta.push(humanizeId(candidate.icons));
if (candidate.image_strategy && candidate.image_strategy.rendering) {
if (candidate.image_strategy && candidate.image_strategy.rendering &&
candidate.image_strategy.rendering !== "custom") {
meta.push(comparisonValueLabel("rendering", candidate.image_strategy.rendering));
}
if (meta.length) card.appendChild(el("div", "font-card-meta", meta.join(" · ")));
@@ -2049,9 +2107,26 @@
});
if (swatches.childElementCount) card.appendChild(swatches);
var note = localized(candidate, "note");
if (note) card.appendChild(el("div", "color-note", note));
card.addEventListener("click", function () { applyDesignDirection(candidate, idx); });
cardStates.push({ candidate: candidate, index: idx, card: card, status: status });
if (note && !customVisual) card.appendChild(el("div", "color-note", note));
var restore = el("button", "direction-reset-button", t("direction_restore"));
restore.type = "button";
restore.hidden = true;
restore.addEventListener("click", function (event) {
event.stopPropagation();
applyDesignDirection(candidate, idx);
});
card.appendChild(restore);
card.addEventListener("click", function () {
if (designDirectionId(candidate, idx) === ACTIVE_DIRECTION_ID) return;
applyDesignDirection(candidate, idx);
});
cardStates.push({
candidate: candidate,
index: idx,
card: card,
status: status,
restore: restore
});
grid.appendChild(card);
});
refreshDesignDirectionState = function () {
@@ -2062,6 +2137,8 @@
entry.card.classList.toggle("adjusted", adjusted);
entry.status.style.display = active ? "inline-block" : "none";
entry.status.textContent = adjusted ? t("direction_adjusted") : t("direction_active");
entry.restore.hidden = !(active && adjusted);
entry.card.title = active ? "" : t("direction_apply_hint");
});
};
refreshDesignDirectionState();
@@ -2235,6 +2312,7 @@
((STATE.image_strategy || {}).behavior || ""));
entry.editor.value = current;
}
if (selected) fitTextareaToContent(entry.editor);
}
if (entry.note) entry.note.style.display = selected ? "none" : "block";
});
@@ -2374,7 +2452,7 @@
// edits call it so the conditional AI path stays synchronized on the page.
var refreshImageProduction = function () {};
// Replaced when the typography section mounts; the canvas section calls it so
// the body-size hint tracks the chosen canvas height.
// the body-size hint tracks the chosen canvas dimensions.
var refreshBodySizeHint = function () {};
// Replaced when the typography section mounts; body-size / reading-mode
// changes call it so unpinned per-role values update locally.
@@ -2407,42 +2485,25 @@
return Math.round(raw);
}
// Canvas height (viewBox user units) parsed from a catalog `dim` like
// Canvas dimensions (viewBox user units) parsed from a catalog `dim` like
// "1242×1660" or from a custom canvas string containing WxH; null if unknown.
function canvasHeight(canvasVal) {
function canvasDimensions(canvasVal) {
var dim = null;
(CAT.canvas || []).forEach(function (o) { if (o.id === canvasVal) dim = o.dim; });
var m = String(dim || canvasVal || "").match(/(\d{2,5})\s*[×xX*]\s*(\d{2,5})/);
return m ? parseInt(m[2], 10) : null;
}
function bodySizeRatioBand(canvasVal) {
var dim = null;
(CAT.canvas || []).forEach(function (o) { if (o.id === canvasVal) dim = o.dim; });
var raw = String(dim || canvasVal || "");
var id = String(canvasVal || "").toLowerCase();
var isPpt = id === "ppt169" || id === "ppt43" ||
/1280\s*[×xX*]\s*720/.test(raw) ||
/1024\s*[×xX*]\s*768/.test(raw);
return isPpt ? { lo: 0.031, hi: 0.047 } : { lo: 0.025, hi: 0.033 };
return m ? { width: parseInt(m[1], 10), height: parseInt(m[2], 10) } : null;
}
// PPT canvases (16:9 / 4:3) take the fixed per-reading-mode body px;
// social / print canvases scale the body px by canvas height instead.
// other canvases use the canvas-owned effective-span rule below.
function isPptCanvas(canvasVal) {
var dim = null;
(CAT.canvas || []).forEach(function (o) { if (o.id === canvasVal) dim = o.dim; });
var raw = String(dim || canvasVal || "");
var id = String(canvasVal || "").toLowerCase();
return id === "ppt169" || id === "ppt43" ||
/1280\s*[×xX*]\s*720/.test(raw) ||
/1024\s*[×xX*]\s*768/.test(raw);
return id === "ppt169" || id === "ppt43";
}
// Body baseline in **px** per reading mode (legacy key:
// delivery_purpose; see strategist.md §g). The
// system is px-only — these are the SVG/execution px values, recalibrated for
// the 1280×720 PPT canvas. No pt layer, no conversion. `def` is the fixed
// delivery_purpose). The system is px-only — these mirror the registered-PPT
// values in references/canvas-formats.md. No pt layer, no conversion. `def` is the fixed
// recommendation; lo/hi are a sanity envelope for the out-of-range flag only.
function deliveryBodyPx(purposeId) {
if (purposeId === "text") return { lo: 18, hi: 21, def: 20 };
@@ -2450,12 +2511,25 @@
return { lo: 22, hi: 25, def: 24 }; // balanced — the default
}
// Mirrors references/canvas-formats.md § "Typography Scale Start". The
// 3× short-side cap keeps extreme aspect ratios from inflating the scale;
// lo/hi remain advisory and def is the initial anchor, never a hard floor.
function bodySizeBandForCanvas(canvasVal, purposeId) {
if (isPptCanvas(canvasVal)) return deliveryBodyPx(purposeId);
var dims = canvasDimensions(canvasVal);
if (!dims) return { lo: 32, hi: 48, def: 40 }; // legacy invalid-custom fallback
var shortSide = Math.min(dims.width, dims.height);
var longSide = Math.max(dims.width, dims.height);
var span = Math.min(longSide, 3 * shortSide);
return {
lo: Math.round(span * 0.025),
hi: Math.round(span * 0.033),
def: Math.round(span * 0.029)
};
}
function defaultBodySizeForCanvas(canvasVal, purposeId) {
if (isPptCanvas(canvasVal)) return deliveryBodyPx(purposeId).def;
var h = canvasHeight(canvasVal);
if (!h) return 40;
var band = bodySizeRatioBand(canvasVal);
return Math.round(h * (band.lo + band.hi) / 2);
return bodySizeBandForCanvas(canvasVal, purposeId).def;
}
// Resolve the only deterministic same-stage size dependency locally. The
@@ -2922,10 +2996,12 @@
var sizeInput = el("input", "num-input font-size-input");
sizeInput.type = "number";
sizeInput.min = "8";
sizeInput.max = "96";
sizeInput.max = "256";
sizeInput.step = "1";
sizeInput.value = (STATE.typography && STATE.typography.body_size) || "";
sizeInput.placeholder = isPptCanvas(STATE.canvas) ? "20 / 24 / 32" : "40 / 48";
sizeInput.placeholder = isPptCanvas(STATE.canvas)
? "20 / 24 / 32"
: String(bodySizeBandForCanvas(STATE.canvas, STATE.delivery_purpose).def);
sizeInput.addEventListener("input", function () {
if (!STATE.typography) STATE.typography = { name: "", heading: {}, body: {} };
STATE.typography.body_size = sizeInput.value;
@@ -2939,7 +3015,7 @@
var sizePtHint = el("div", "toggle-desc body-size-pt");
var sizeHint = el("div", "toggle-desc body-size-hint");
// PPT body is one fixed px value per reading mode (not a range); non-PPT
// canvases scale px to canvas height. A manually pinned value is never
// canvases use the canvas-owned effective span. A manually pinned value is never
// overwritten by later reading-mode changes.
// Everything is px — lo/hi are only a sanity envelope for the OOR flag.
refreshBodySizeHint = function () {
@@ -2950,13 +3026,10 @@
lo = pb.lo; hi = pb.hi;
txt += " " + t("body_size_hint_purpose").replace("{def}", pb.def);
} else {
var h = canvasHeight(STATE.canvas);
var band = bodySizeRatioBand(STATE.canvas);
if (h) {
lo = Math.round(h * band.lo); hi = Math.round(h * band.hi);
txt += " " + t("body_size_hint_canvas")
.replace("{lo}", lo).replace("{hi}", hi);
}
var band = bodySizeBandForCanvas(STATE.canvas, STATE.delivery_purpose);
lo = band.lo; hi = band.hi;
txt += " " + t("body_size_hint_canvas")
.replace("{lo}", lo).replace("{hi}", hi);
}
// Flag (hint only) a value far outside the
// canvas's usual px range, so an accidental extreme value is visible
@@ -3000,7 +3073,7 @@
wrap.appendChild(el("div", "hex-cell-label", t("size_role_" + role)));
var inputLine = el("div", "role-size-line");
var inp = document.createElement("input");
inp.type = "number"; inp.min = "6"; inp.max = "200"; inp.step = "1";
inp.type = "number"; inp.min = "6"; inp.max = "512"; inp.step = "1";
inp.addEventListener("input", function () {
if (!STATE.typography) STATE.typography = { name: "", heading: {}, body: {} };
if (!STATE.typography.sizes) STATE.typography.sizes = {};
@@ -3027,7 +3100,7 @@
if (!STATE.typography.sizes) STATE.typography.sizes = {};
sizeInput.value = STATE.typography.body_size || "";
var bodyVal = parseFloat(STATE.typography.body_size) ||
(isPptCanvas(STATE.canvas) ? deliveryBodyPx(STATE.delivery_purpose).def : 40);
defaultBodySizeForCanvas(STATE.canvas, STATE.delivery_purpose);
SIZE_ROLES.forEach(function (role) {
var cur = STATE.typography.sizes[role];
var hasVal = cur !== undefined && cur !== null && cur !== "";
@@ -3129,7 +3202,8 @@
var sacc = hexOr(pal.secondary_accent, acc);
var txt = hexOr(pal.body_text, "#1d2430");
// body_size is px everywhere — preview it directly, no conversion.
var rawSize = parseFloat(typ.body_size) || (isPptCanvas(STATE.canvas) ? 24 : 18);
var rawSize = parseFloat(typ.body_size) ||
defaultBodySizeForCanvas(STATE.canvas, STATE.delivery_purpose);
var bodyPx = Math.max(12, Math.min(34, rawSize));
var headPrimaryStack = previewFontStack(head.primary, head.css);
var headEnglishStack = previewFontStack(head.english, head.css);
@@ -3746,7 +3820,7 @@
initCreativeSelection("visual_style", CAT.visual_styles, "visual_style_behavior");
var cc = colorRecommendationCandidates();
var csel = (REC.color && REC.color.selected != null) ? REC.color.selected :
(designDirectionSpec().selected || 0);
selectedDesignDirectionIndex();
var c0 = cc[Math.min(csel, Math.max(cc.length - 1, 0))] || {};
STATE.color = {
name: localized(c0, "name") || c0.name || "",
@@ -3757,7 +3831,7 @@
var tc = typographyRecommendationCandidates();
var tsel = (REC.typography && REC.typography.selected != null) ? REC.typography.selected :
(designDirectionSpec().selected || 0);
selectedDesignDirectionIndex();
var t0 = normTypography(tc[Math.min(tsel, Math.max(tc.length - 1, 0))] || {});
STATE.typography = {
name: localized(t0, "name") || t0.name || "",
@@ -3769,14 +3843,14 @@
if (t0.custom) STATE.typography.custom = t0.custom;
// Guarantee a body baseline even when a candidate omitted body_size, on
// any canvas (PPT → px default by purpose, non-PPT → px from canvas height),
// any canvas (PPT → px default by purpose, non-PPT → px from effective span),
// so role sizes never derive from an empty anchor.
if (STATE.typography && !STATE.typography.body_size) {
STATE.typography.body_size = defaultBodySizeForCanvas(STATE.canvas, STATE.delivery_purpose);
}
// A freshly authored Stage 2 starts from one deterministic reading-mode
// baseline.
if (stageNumber(REC) === 2) syncUnpinnedTypographySizes(true);
// Preserve an authored positive candidate anchor; derive only the roles
// that remain unpinned. A user reading-mode change may reset it later.
if (stageNumber(REC) === 2) syncUnpinnedTypographySizes(false);
var rawImageUsage = recValue("image_usage");
STATE.image_usage = selectedImageUsageIds(rawImageUsage);
if (!STATE.image_usage.length) {
@@ -3798,9 +3872,7 @@
}
var directions = designDirectionCandidates();
if (directions.length) {
var selected = Number(designDirectionSpec().selected || 0);
if (!isFinite(selected) || selected < 0) selected = 0;
selected = Math.min(selected, directions.length - 1);
var selected = selectedDesignDirectionIndex();
applyDesignDirection(directions[selected], selected, false);
}
}
@@ -482,7 +482,25 @@ textarea.text-input { resize: vertical; line-height: 1.5; }
border-color: #d97706;
background: #fffbeb;
}
.design-direction-card.selected,
.design-direction-card.adjusted { cursor: default; }
.direction-status { margin-left: auto; }
.direction-reset-button {
margin-top: 10px;
border: 1px solid #d97706;
border-radius: 7px;
background: #fff;
color: #92400e;
padding: 6px 10px;
font-size: 12px;
font-weight: 650;
cursor: pointer;
}
.direction-reset-button:hover { background: #fef3c7; }
.direction-reset-button:focus-visible {
outline: 2px solid #d97706;
outline-offset: 2px;
}
.scheme-component-options { margin-bottom: 14px; }
.scheme-component-grid {
flex-direction: row;
@@ -502,7 +520,11 @@ textarea.text-input { resize: vertical; line-height: 1.5; }
.scheme-component-editor {
margin-top: 8px;
min-height: 84px;
resize: vertical;
max-height: 380px;
overflow-y: auto;
font-size: 12px;
line-height: 1.5;
resize: none;
cursor: text;
}
.scheme-component-card .design-direction-preview { margin-top: 8px; }
@@ -515,6 +537,29 @@ textarea.text-input { resize: vertical; line-height: 1.5; }
overflow: hidden;
}
.design-direction-preview img { width: 100%; height: 100%; object-fit: cover; }
.design-direction-custom-preview {
display: flex;
flex-direction: column;
justify-content: center;
gap: 8px;
padding: 14px;
background: var(--accent-soft);
}
.design-direction-custom-label {
color: var(--accent);
font-size: 11px;
font-weight: 700;
letter-spacing: .04em;
}
.design-direction-custom-copy {
display: -webkit-box;
overflow: hidden;
color: var(--ink);
font-size: 12.5px;
line-height: 1.55;
-webkit-box-orient: vertical;
-webkit-line-clamp: 5;
}
.design-direction-swatches { display: flex; gap: 5px; margin-top: 8px; }
.design-direction-swatches .swatch { width: 24px; height: 24px; }
.locked-summary-chip { cursor: default; }
@@ -307,15 +307,15 @@ its file time is newer than the handoff.
The following fields belong to the Strategist stages, not to the
template-selection receipt.
- **Enumerable + custom** — canvas / icons retain blank manual inputs. Mode and visual style first show three project-specific `custom` values projected from the complete directions, then the full fixed base catalog as conservative single-select alternatives. Selecting a projected card expands its behavior editor in place; edits change only the current value, while clicking the owning whole-direction card restores the authored text.
- **Visual examples for hard-to-name choices** — the full-screen confirmation page loads real SVG page samples from `static/style_previews/` for `visual_style`, and renders real sample SVGs from `templates/icons` for `icons`. These thumbnails make style and icon-library choices visually comparable before the user locks them. Preview copy is fixed role text (big title / section title / body / points), not project content from recommendation files, so users compare visual treatment rather than copywriting. These previews are a confirmation aid only: they do not add fields to recommendation stage files or `result.json`, and they do not replace the later Step 6 live preview.
- **Enumerable + custom** — canvas / icons retain blank manual inputs. Mode and visual style first show three project-specific `custom` values projected from the complete directions, then the full fixed base catalog as conservative single-select alternatives. Selecting a projected card expands its behavior editor in place; edits change only the current value, while the adjusted active whole-direction card exposes an explicit restore action for the authored text.
- **Visual examples for hard-to-name choices** — the full-screen confirmation page loads real SVG page samples from `static/style_previews/` for fixed `visual_style` catalog choices, and renders real sample SVGs from `templates/icons` for `icons`. Project-specific `custom` direction cards show their authored summary instead of requesting a nonexistent preset asset. These thumbnails and summaries make style and icon-library choices visually comparable before the user locks them. Preview copy is fixed role text (big title / section title / body / points), not project content from recommendation files, so users compare visual treatment rather than copywriting. These previews are a confirmation aid only: they do not add fields to recommendation stage files or `result.json`, and they do not replace the later Step 6 live preview.
- **Image usage multi-select** — image sources are selected as one or more catalog ids: `ai` = AI-generated, `web` = Web-sourced, `provided` = User-provided, `placeholder` = Placeholder, `none` = No images. `none` is exclusive. A confirmed non-`none` set is the allowed acquisition-source boundary, not a requirement to use every selected source; only explicit `image_notes` wording can require a source, asset, or page role. Recommendation and result values may be a legacy single string, but new files should use an array. When several sources are recommended, write the source ids to `recommend.image_usage` and write the actual usage strategy to `image_notes`, not a custom prose value.
- **Closed enumerable** — PPT reading mode (`delivery_purpose` compatibility key), formula policy / generation mode / refine spec, plus AI source only when image usage includes `ai`. These have no Custom box; out-of-catalog values snap back to the recommended option.
- **Proactive execution booleans** — Final Stage 2 carries top-level `proactive_speaker_notes`, `proactive_custom_animations`, and `proactive_narration_audio` values. Defaults are `true`, `false`, and `false`, respectively. They control what the Agent does proactively only when the user has not explicitly instructed otherwise; the latest explicit user instruction always wins. These three values are raw confirmation evidence: the UI and server neither couple nor rewrite them, and every boolean combination is valid. When narration audio is enabled, Strategist later resolves the effective Speaker Notes outcome to enabled and records `Narration Audio dependency` as its Design Spec provenance. Disabling proactive custom animation does not suppress the Strategist's advisory motion recommendations.
- **Open prose**`audience`, `communication_intent`, `audience_outcome`, `core_message`, `delivery_context`, `artifact_afterlife`, `content_divergence`, and `page_count`. `communication_intent` may carry several purposes plus priority / sequence; common paths appear only as help text. `delivery_context` states one primary presenter-led / reader-led / hybrid / recorded-self-running context plus optional secondary use; a hybrid recommendation names which context leads. `content_divergence` is the source-treatment axis. `page_count` may be a range here; Strategist resolves the exact §IX roster, leaving Executor no pagination latitude.
- **Coordinated generative directions**`design_directions` carries exactly three complete candidates authored top-down from the project contract. Each has a unique stable id and bundles `custom` mode, `custom` visual style, color, typography, icon id, and `custom` generated-image rendering regardless of recommended image source. All three must be viable as whole solutions; they do not need different catalog bases or forced safe / shifted / bold archetypes. The page can still render legacy top-level `color`, `typography`, and `image_strategy` candidates, but new staged recommendations use the coordinated bundle.
- **Coordinated generative directions**`design_directions` carries exactly three complete candidates authored top-down from the project contract. Each has a unique stable id and bundles `custom` mode, `custom` visual style, color, typography, icon id, and `custom` generated-image rendering regardless of recommended image source. Its localized note is a compact, user-facing style summary. It may reuse localized display labels from `catalogs.visual_styles` when they describe the result concisely, but those labels are optional vocabulary rather than a selection constraint or required mapping. Otherwise it uses concise natural language and never forces the nearest label. The summary stays within one or two short sentences and does not expose catalog ids or reference mechanics. All three must be viable as whole solutions; they do not need different catalog bases or forced safe / shifted / bold archetypes. After completing all three bundles, Strategist compares them against the confirmed contract and source, then writes the strongest overall fit's zero-based index to `selected`; array position does not determine preference. That bundle becomes the initial default and applies its three custom projections coherently. The page can still render legacy top-level `color`, `typography`, and `image_strategy` candidates, but new staged recommendations use the coordinated bundle.
Direction-local custom projections apply to mode, visual style, and generated-image rendering; all three are editable after selection and a selected custom value cannot be blank. The original recommendation remains immutable so a whole-direction click can reset every edited component. Legacy standalone `custom_candidates` remain readable but are optional and are not authored in new files. Color / typography keep their existing manual Custom cards. Image usage uses source ids plus `image_notes`; closed sets have no Custom path.
Direction-local custom projections apply to mode, visual style, and generated-image rendering; all three are editable after selection and a selected custom value cannot be blank. The original recommendation remains immutable so the active whole-direction card can explicitly restore every edited component without making an ordinary card click destructive. Legacy standalone `custom_candidates` remain readable but are optional and are not authored in new files. Color / typography keep their existing manual Custom cards. Image usage uses source ids plus `image_notes`; closed sets have no Custom path.
**Stage-2 catalog read gate.** Before choosing component bases, Strategist reads only `modes/_index.md`, `visual-styles/_index.md`, and `image-renderings/_index.md`. It authors the three whole solution intents first, freezes every basis id from those indexes, and only then reads the deduplicated selected detail files before completing the custom behaviors. Unselected sibling files never enter context; a novel custom reads none.
@@ -463,16 +463,16 @@ corresponding direct values (`formula_policy`, `generation_mode`, boolean
"proactive_narration_audio": { "value": false },
"refine_spec": { "value": false },
"design_directions": {
"selected": 0,
"selected": 1,
"candidates": [
{
"id": "executive-clarity",
"name_zh": "稳妥专业",
"note_zh": "像成熟咨询简报",
"note_zh": "以瑞士极简为主,融合柔和圆角与编辑出版风格。",
"mode": "custom",
"mode_behavior_zh": "以 pyramid 的结论先行为主轴,在风险段引入 narrative 的局势—张力—解法转折;标题保持判断句,每章以可执行结论收束。",
"mode_behavior_zh": "以 pyramid 作为唯一目录基底,为当前风险决策材料定制两次结论闸门;标题保持判断句,每章先给判断,再用证据展开并以可执行结论收束。",
"visual_style": "custom",
"visual_style_behavior_zh": " swiss-minimal 精确栅格和大留白为基底,融入 editorial 细规则、边注与证据层级;标题锐利,正文中性,装饰只用于标记推理关系。",
"visual_style_behavior_zh": " swiss-minimal 负责精确栅格和大留白soft-rounded 负责少量关键容器的轮廓与轻微抬升,editorial 负责细规则、边注与证据层级;标题锐利,正文中性,装饰只标记推理关系。",
"icons": "tabler-outline",
"color": { "name_zh": "冷静专业", "palette": {
"background": "#FFFFFF", "secondary_bg": "#F4F6F8",
@@ -490,7 +490,7 @@ corresponding direct values (`formula_policy`, `generation_mode`, boolean
"rendering": "custom",
"visual_zh": "简化矢量主体配合编辑式注释与局部材质对比",
"mood_zh": "审慎、可信,像调查报道中的证据插图",
"behavior_zh": "融合 vector-illustration 清晰轮廓与 editorial 的证据编排;使用平面色块、细规则和局部纸张纹理,避免写实景深与装饰性渐变,颜色继承当前演示文稿角色。"
"behavior_zh": " vector-illustration 负责清晰轮廓minimalist-swiss 负责留白构图,screen-print 负责克制的半调纹理,warm-scene 负责暖光与可信氛围;四者服从同一平面主体和当前演示文稿颜色角色,避免写实景深与装饰性渐变。"
}
}
]
@@ -498,11 +498,11 @@ corresponding direct values (`formula_policy`, `generation_mode`, boolean
}
```
The example shows one candidate's complete shape; the actual array repeats that shape for exactly three candidates. Every candidate requires a unique stable `id`, literal `custom` mode/style/rendering with localized behavior, an icon value, complete six-role palette, complete heading/body stack, and complete image strategy even when `recommend.image_usage` excludes AI. The server rejects any fixed mode/style/rendering inside a new direction. Legacy grids remain readable only with three complete palettes and complete typography.
The example shows one candidate's complete shape; the actual array repeats that shape for exactly three candidates. Every candidate requires a unique stable `id`, literal `custom` mode/style/rendering with localized behavior, an icon value, complete six-role palette, complete heading/body stack, and complete image strategy even when `recommend.image_usage` excludes AI. The server rejects any fixed mode/style/rendering inside a new direction and any explicit `design_directions.selected` outside integer `0` through `2`; omission remains readable as a legacy index-`0` fallback. Legacy grids remain readable only with three complete palettes and complete typography.
- `design_directions.selected` owns the initial complete bundle. New direction mode/style/rendering values are always literal `custom`; prose stays in the required behavior sibling. `recommend.*` remains a compatibility hint and may mirror `custom`, but it never replaces the selected direction. Legacy aliases remain accepted; new files write canonical ids.
- `design_directions.selected` owns the initial complete bundle and MUST be the actual zero-based index (`0`, `1`, or `2`) that Strategist chooses only after all three bundles are complete. The chosen card is marked Recommended and initially applies its matching mode, visual-style, and generated-image custom candidates together; array order carries no recommendation meaning. New direction mode/style/rendering values are always literal `custom`; prose stays in the required behavior sibling. `recommend.*` remains a compatibility hint and may mirror `custom`, but it never replaces the selected direction. Legacy aliases remain accepted; new files write canonical ids.
- The three proactive-execution fields are top-level boolean `{ "value": ... }` objects, not catalog ids. Omitted fields use `true / false / false` for notes / custom animation / narration audio. These are absence-of-instruction defaults, not permission to override the user's latest explicit request. Preserve all three raw values independently through `result.json`; do not couple or rewrite them. Strategist derives effective Speaker Notes as enabled when audio is `true` and records `Narration Audio dependency` as provenance in the Design Spec. `proactive_custom_animations: false` leaves Strategist animation suggestions unchanged; it only prevents unrequested custom-animation execution.
- `custom_candidates` is an optional legacy recommendation-only shape. New files place all three project-specific variants inside `design_directions`; each remains separately selectable and becomes editable in place. A custom behavior names exact catalog ids only when it actually uses them as synthesis bases. The UI rejects a selected blank and submits only the edited current value. Template-backed variants obey inherited identity, prototype capacity, and `template_application`.
- `custom_candidates` is an optional legacy recommendation-only shape. New files place all three project-specific variants inside `design_directions`; each remains separately selectable and becomes editable in place. A custom behavior may use zero, one, or many exact catalog bases. Reference count has no fixed cap: every named id must contribute a distinct executable job, the behavior omits any id whose contribution it cannot state, and one basis never requires a decorative second. The UI rejects a selected blank and submits only the edited current value. Template-backed variants obey inherited identity, prototype capacity, and `template_application`.
- Seed `audience`, `communication_intent`, `audience_outcome`, and `delivery_context` when evidence supports them; users need not supply them, and every Stage-1 prose field may end blank. The contract and `primary_language` stay in `result.json` and `design_spec.md`; `spec_lock.md communication` receives `primary_language`, compact `audience` / `objective` / `core_message`, and reading mode. `communication_intent` may preserve multiple purposes and priority/sequence; never add a `primary_job` enum.
- Do not write `recommend.template_reuse_scope` or `recommend.template_adherence`. Strategist records those internal exporter values later in `spec_lock.md` after inspecting the actual template and current content.
- For a confirmed templates-mode handoff, write one editable prose field as top-level `template_application.value`. It summarizes **how to use** the already selected project-local template: actual page/prototype use and preservation/reorganization decisions. It never chooses, changes, or reinstalls a workspace. Omit it for free design. The UI returns the current string through final Stage 2; Strategist then persists the final effective plan as `- **Template Application**: ...` in `design_spec.md §I`, which Executor reads from the retained Design Spec. Never replace it with internal reuse/adherence ids or a fixed option menu.
@@ -522,12 +522,12 @@ Template-mode-only Stage-2 fragment:
- Final Stage 2 shows and submits `recommend.image_ai_path` as one of `auto` / `api` / `host-native` / `manual` only while its current `image_usage` includes `ai`; changing sources refreshes that production control on the same page.
- **Color candidates carry the user-facing core `palette`**: `background`, `secondary_bg`, `primary`, `accent`, `secondary_accent`, and `body_text`. The page renders every role as a labelled swatch with its HEX value visible, and offers per-role override inputs for precise single-role edits, plus a **Custom color card with a free-text box** — the user can describe the palette in words or paste HEX values instead of filling each role; this writes `color: { "name": "custom", "custom": "<text>" }` to `result.json` for the AI to interpret. Legacy `text` is accepted as an alias for `body_text`, but new files should write `body_text`. Strategist derives secondary text, borders, state colors, and visual-style neutral tiers while writing `design_spec.md`, then projects the machine values to `spec_lock.md`; those are not user-facing confirmation choices.
- **Candidate display text may be multilingual**: color / typography candidates can provide `name_zh` / `name_en` / `name_ja` and `note_zh` / `note_en` / `note_ja`; the page falls back to legacy `name` / `note`. Labels resolve in the page language first, then fall back across the others (a `ja` page: ja → en → zh; zh/en pages keep their zh↔en fallback and try `_ja` last), so when `lang` is `ja` always include the `_ja` variants — otherwise the candidate labels render in English.
- **Typography candidates** use concrete heading/body `primary`; non-English decks also use `english`, while English-primary decks omit it. `cjk` / `latin` remain legacy aliases. Localized `name` labels the pair and `css` only previews. Bundles differ overall; font pairs may repeat without blocking. Fixed pairs require `fixed: true`. Catalog `fonts` supplies language-filtered dropdowns plus Other without limiting recommendations; edits mark Custom and refresh the preview. Include topic samples. PPT baselines are `text` 20 · `balanced` 24 · `presentation` 32 px; cards preserve sizes and submit px.
- **Per-role size override** (parallel to color's per-role HEX override): besides `body_size`, the page exposes editable inputs for `title` / `subtitle` / `annotation`. The browser applies one documented deterministic dependency chain: `reading mode → body baseline → unpinned role sizes` (role ramp: `body ×` the §g ratios). Changing reading mode updates the body and all unpinned roles locally; changing body updates unpinned roles locally. Editing body or a role pins that value, so later reading-mode changes do not overwrite it. A font-only selection preserves current sizes; applying a complete direction restores that direction's typography baseline and derived unpinned sizes, because the top card is an explicit reset action. This is a browser-only state update: it performs no fetch and asks the backend to author no new recommendations. Each role input is labelled as px and shows an approximate pt equivalent (`1px = 0.75pt`) for orientation. The final values are written to `result.json` as `typography.sizes: { "title", "subtitle", "annotation" }` in **px** — every canvas, no pt and no `sizes_pt` provenance. These confirmed values are Strategist input anchors: the completed page plan may add recurring roles, and downstream execution owns bounded per-occurrence treatment. Candidate `sizes` remain accepted for compatibility, but the fresh Stage-2 baseline is normalized through the same local ramp before first render.
- **Typography candidates** use concrete heading/body `primary`; non-English decks also use `english`, while English-primary decks omit it. `cjk` / `latin` remain legacy aliases. Localized `name` labels the pair and `css` only previews. Bundles differ overall; font pairs may repeat without blocking. Fixed pairs require `fixed: true`. Catalog `fonts` supplies language-filtered dropdowns plus Other without limiting recommendations; edits mark Custom and refresh the preview. Include topic samples. [`canvas-formats.md`](../../references/canvas-formats.md) § "Typography Scale Start" is the single owner of initial body anchors and sanity bands; the browser mirrors that rule, and submitted values remain px.
- **Per-role size override** (parallel to color's per-role HEX override): besides `body_size`, the page exposes editable inputs for `title` / `subtitle` / `annotation`. The browser applies one documented deterministic dependency chain: PPT uses `reading mode → body baseline`, non-PPT uses `canvas → body baseline`, then every canvas uses `body baseline → unpinned role sizes` (role ramp: `body ×` the §g ratios). Changing reading mode updates a PPT body and all unpinned roles locally; changing body updates unpinned roles locally. Editing body or a role pins that value, so later reading-mode changes do not overwrite it. A font-only selection preserves current sizes; applying a different complete direction, or using the active card's explicit restore action, restores that direction's typography baseline and derived unpinned sizes. This is a browser-only state update: it performs no fetch and asks the backend to author no new recommendations. Each role input is labelled as px and shows an approximate pt equivalent (`1px = 0.75pt`) for orientation. The final values are written to `result.json` as `typography.sizes: { "title", "subtitle", "annotation" }` in **px** — every canvas, no pt and no `sizes_pt` provenance. These confirmed values are Strategist input anchors: the completed page plan may add recurring roles, and downstream execution owns bounded per-occurrence treatment. Candidate `sizes` remain accepted for compatibility; fresh Stage 2 preserves a candidate `body_size` as its baseline and derives only missing or unpinned role sizes from the same local ramp before first render.
- **`delivery_purpose` compatibility key / Reading mode** (enumerable, PPT only) decides where meaning is carried, not merely how large type is: `text` makes pages self-contained with complete sentences, short prose, captions, tables, and necessary detail; `balanced` shares explanation between page and presenter; `presentation` uses one idea, concise claims, and visual evidence while speech / notes carry the detail. It therefore governs page grammar, granularity, density / rhythm, and note burden. Reading-mode cards intentionally show **no px value**; the typography section owns the separately visible body / role sizes and applies any local default. It is surfaced in Stage 2 beside the visual system, separate from communication intent. `recommend.delivery_purpose` pre-selects one; `result.json` retains the key, while `spec_lock.md` uses canonical `consumption_mode`. Non-PPT canvases omit it.
- **Combined style preview** — a compact live "overall impression" strip sits just above the color section and is **sticky**: it pins under the topbar so it stays visible while the user scrolls through the color / icon / typography sections, keeping the picking controls and their combined effect on screen together. It applies the currently selected color palette **and** typography (heading sample in `primary` over `background`, body sample in `body_text`, an `accent` bar, a `secondary_bg` chip) and repaints on every color / HEX-override / font / `body_size` change. It does not replace the per-candidate swatches or font samples (those stay for picking); it is deliberately an abstract style chip, **not** a slide-layout preview — page layout preview remains the live-preview server's job (Step 6). No schema field; it derives entirely from the existing color + typography selections.
- **Generated-image direction** appears only for current `image_usage: ai`, but all three custom project candidates already exist in `design_directions` before that toggle. Turning AI on reveals those candidates immediately without a backend rerun, followed by the 20 fixed system styles. Selecting a project candidate expands its behavior editor in that card; a fixed preset submits its id, while a project candidate submits `rendering: "custom"` + edited non-empty `behavior`. Turning AI off omits `image_strategy` from the final result without deleting the authored recommendation candidates. Catalog-based custom behavior names exact ids for optional `image_rendering_references`; a novel behavior has none. The left preview follows selection. No image palette is written; deck colors remain authoritative, and legacy `image_strategy.palette` is ignored.
- **`design_directions`** is the canonical Stage-2 starting set: exactly three top-down, project-fit bundles with stable ids, localized copy, custom mode/style/rendering, icons, complete language-aware typography, and HEX `background`, `secondary_bg`, `primary`, `accent`, `secondary_accent`, `body_text`. Clicking a card applies every field it owns; projected custom fields can then be edited in place and all lower controls may diverge. The card shows an adjusted state, while clicking the same or another whole direction reapplies its immutable authored bundle. `result.json` stores the edited current components, never a direction id.
- **`design_directions`** is the canonical Stage-2 starting set: exactly three top-down, project-fit bundles with stable ids, localized copy, custom mode/style/rendering, icons, complete language-aware typography, and HEX `background`, `secondary_bg`, `primary`, `accent`, `secondary_accent`, `body_text`. The `selected` card carries the persistent Recommended marker and is applied first. A custom direction card uses its localized style-summary note—or the required behavior fallback—instead of requesting a preset-style preview. Newly authored notes may borrow localized catalog display labels where useful or use concise natural language freely; they never force an approximate label or expose internal catalog ids. Clicking an inactive card applies every field it owns; projected custom fields can then be edited in place and all lower controls may diverge. The active card shows an adjusted state and exposes an explicit restore action for its immutable authored bundle. `result.json` stores the edited current components, never a direction id.
- `recommend.generation_mode` and `refine_spec` mirror [`generate-pptx`](../../workflows/generate-pptx.md) Step 4. `split` / `true` are explicit opt-ins. Refinement adds no UI stage: after Gate 1 it stops before the lock for unrestricted chat revision until approval.
- `content_divergence` is a **free-text** Stage-1 source-treatment field. Blank means a balanced default; facts stay sourced at every level. Strategist consumes it while authoring §IX and records it in `design_spec.md §I`; it is not written to `spec_lock.md`. Beautify sends `{ "value": "keep source wording and page structure verbatim", "locked": true }`, so the UI displays it read-only and the server restores it on every staged submit. Template-fill does not use this confirmation flow and does not surface it.
- `lang` is the soft UI-language default (`zh` / `en` / `ja`); the persisted user choice wins. It never sets `primary_language`.
@@ -582,7 +582,9 @@ custom` + `visual_style_behavior`, or `image_strategy.rendering: custom` +
`behavior`. During Design Spec and lock authoring, Strategist projects optional
`mode_references`, `visual_style_references`, or
`image_rendering_references` only when that confirmed behavior actually uses
named catalog sources; genuinely novel custom behavior has no reference list.
named catalog sources. These lists have no fixed item limit; each item must own
a distinct executable contribution, while genuinely novel custom behavior has
no reference list.
The Stage-1 intermediate write retains the communication contract for Stage 2.
**Final-result consumption contract.** A final result is the user-confirmed input contract for the Strategist's Design Spec, not another recommendation input. After the final wait, Generate Step 4 reads the complete final object exactly once and retains it while Strategist writes and audits `design_spec.md` against every explicitly present field. Normal lock authoring and downstream execution do not reopen `result.json`; the completed Design Spec is the durable authority. Only after that audit passes does Strategist author `spec_lock.md` from the Design Spec plus current execution context, selecting stable anchors and routing rather than copying every field or enumerating every legal color/font. Every value must be consumed at the semantic type owned by [`strategist.md`](../../references/strategist.md) §1 and its field owner: do not omit or substitute it, and do not silently strengthen or weaken its type. If a confirmed requirement cannot be honored, the owning workflow reports or pauses under failure recovery; it never deletes the requirement to keep the pipeline moving.
@@ -43,7 +43,7 @@ Advance mode:
| click | Click advance only |
| after | Timed advance only |
| both | Click or timed advance, whichever occurs first |
| narration | Timed advance from audio duration plus padding; click disabled |
| narration | Timed advance from narration lead-in, audio duration, and page-tail padding; click disabled |
**Hard rule**: enter=none may coexist with a timed advance. The valid result is
a timing-only p:transition with no visual-effect child.
@@ -647,7 +647,7 @@ Behavior:
- `notes_to_audio.py` uses `edge-tts` by default, or a configured cloud TTS provider (`elevenlabs`, `minimax`, `qwen`, `cosyvoice`), and generates one audio file per slide into `audio/`
- Narration text is read strictly from the matching `notes/*.md` file; the script only skips Markdown heading lines (`# ...`) and does not summarize, rewrite, or filter delivery notes
- `--recorded-narration audio` prepares PowerPoint's "recorded timings and narrations": every slide must have matching `m4a` / `mp3` / `wav` audio, `ffprobe` must read every duration, and `--animation-trigger on-click` is rejected
- `--recorded-narration audio` keeps speaker notes, embeds each matching audio file, and writes slide auto-advance timings from audio duration
- `--recorded-narration audio` keeps speaker notes, embeds each matching audio file, and writes slide auto-advance timings from page-start lead-in + audio duration + page-tail padding. `--narration-start-floor` and `--narration-padding` are independent optional seconds; their defaults are `0.8` and `0.5`, and the post-transition lead-in is `max(0, start floor - transition duration)`
- When either animation sidecar exists, narrated export defaults to `<project>/narration_animations.json`; a canonical `animations.json` without that derived file remains a blocking synchronization error
- Without animation sidecars, Generate narration reads base-report deck motion via `--inherit-motion-from`; direct low-level omission keeps legacy `fade` / no object builds. Use `--animation-config animations.json` for canonical animation, or `--no-animations` to remove object/page motion while retaining narration timings
- Non-narrated export keeps the existing optional `<project>/animations.json` default
@@ -57,7 +57,12 @@ from pptx_animations import ( # noqa: E402
normalize_animation_effect,
normalize_animation_trigger,
)
from pptx_transitions import read_slide_transition_xml # noqa: E402
from pptx_transitions import ( # noqa: E402
DEFAULT_TRANSITION_DURATION,
normalize_transition_effect_request,
read_slide_transition_xml,
validate_seconds,
)
from svg_to_pptx.animation_config import ( # noqa: E402
animation_group_effect_entries,
scan_project_targets,
@@ -66,8 +71,11 @@ from svg_to_pptx.animation_config import ( # noqa: E402
validate_transition_config,
)
from svg_to_pptx.pptx_package.narration import ( # noqa: E402
DEFAULT_NARRATION_START_FLOOR,
NARRATION_EXTENSIONS,
narration_lead_in_seconds,
probe_audio_duration,
read_narration_start_delay_xml,
)
configure_utf8_stdio()
@@ -157,6 +165,15 @@ class SubtitleMergeResult:
minimum_video_correlation: float | None = None
@dataclass(frozen=True)
class PowerPointTiming:
"""One slide's transition, narration, and advance timing in milliseconds."""
transition_ms: int
narration_delay_ms: int
advance_ms: int
def _timestamp_to_ms(value: str) -> int:
hours_text, minutes_text, remainder = value.split(":")
seconds_text, milliseconds_text = remainder.split(",")
@@ -301,9 +318,20 @@ def _subtitle_fingerprint(slide_names: list[str], subtitle_dir: Path) -> str:
return digest.hexdigest()
def _finite_non_negative_seconds(value: object, field: str) -> float:
if (
isinstance(value, bool)
or not isinstance(value, (int, float))
or not math.isfinite(float(value))
or value < 0
):
raise ValueError(f'{field} must be a finite non-negative number')
return float(value)
def _load_timing_plan(
path: Path,
) -> tuple[str, float, dict[str, list[TimingPlanEntry]]]:
) -> tuple[str, float, float | None, dict[str, list[TimingPlanEntry]]]:
raw = json.loads(path.read_text(encoding="utf-8"))
if not isinstance(raw, dict):
raise ValueError(f"Narration timing plan must be a JSON object: {path}")
@@ -311,6 +339,7 @@ def _load_timing_plan(
"version",
"srt_sha256",
"narration_padding",
"narration_start_floor",
"slides",
}
if unknown_top:
@@ -331,17 +360,18 @@ def _load_timing_plan(
'Narration timing plan field "srt_sha256" must be a lowercase '
"SHA-256 digest of the ordered page-local SRT files"
)
narration_padding = raw.get("narration_padding")
if (
isinstance(narration_padding, bool)
or not isinstance(narration_padding, (int, float))
or not math.isfinite(float(narration_padding))
or narration_padding < 0
):
raise ValueError(
'Narration timing plan field "narration_padding" must be a '
"finite non-negative number"
narration_padding = _finite_non_negative_seconds(
raw.get("narration_padding"),
'Narration timing plan field "narration_padding"',
)
narration_start_floor = (
_finite_non_negative_seconds(
raw["narration_start_floor"],
'Narration timing plan field "narration_start_floor"',
)
if "narration_start_floor" in raw
else None
)
slides = raw.get("slides")
if not isinstance(slides, dict):
raise ValueError('Narration timing plan field "slides" must be an object')
@@ -397,7 +427,7 @@ def _load_timing_plan(
entries.append(TimingPlanEntry(group_id, cue_number))
seen_groups.add(group_id)
result[slide_name] = entries
return srt_sha256, float(narration_padding), result
return srt_sha256, narration_padding, narration_start_floor, result
def _load_canonical_animation_config(path: Path) -> dict[str, Any]:
@@ -497,6 +527,58 @@ def _animation_scope(
return value
def _transition_scope(
scope: dict[str, Any],
*,
label: str,
) -> dict[str, Any]:
value = scope.get("transition", {})
if not isinstance(value, dict):
raise ValueError(f'{label} field "transition" must be an object')
return value
def _effective_transition_duration_ms(
config: dict[str, Any],
slide_cfg: dict[str, Any],
) -> int:
"""Resolve the destination slide's effective transition duration."""
defaults = config.get("defaults", {})
if not isinstance(defaults, dict):
raise ValueError('Canonical animations.json field "defaults" must be an object')
default_transition = _transition_scope(
defaults,
label="Canonical animations.json defaults",
)
default_effect, _default_options = normalize_transition_effect_request(
default_transition.get("effect", "fade"),
default_transition.get("effect_options"),
)
default_duration = validate_seconds(
default_transition.get("duration", DEFAULT_TRANSITION_DURATION),
"canonical transition duration",
allow_zero=default_effect is None,
)
slide_transition = _transition_scope(
slide_cfg,
label="Canonical animations.json slide",
)
if "effect" in slide_transition:
effect, _effect_options = normalize_transition_effect_request(
slide_transition["effect"],
slide_transition.get("effect_options"),
)
else:
effect = default_effect
duration = validate_seconds(
slide_transition.get("duration", default_duration),
"canonical slide transition duration",
allow_zero=effect is None,
)
return 0 if effect is None else round(duration * 1000)
def _effective_slide_animation(
config: dict[str, Any],
slide_cfg: dict[str, Any],
@@ -963,10 +1045,17 @@ def rebuild_animations(
output_path: Path,
narration_padding: float,
force: bool,
narration_start_floor: float = DEFAULT_NARRATION_START_FLOOR,
) -> AnimationBuildResult:
"""Derive narration timing without modifying the canonical animation file."""
if not math.isfinite(narration_padding) or narration_padding < 0:
raise ValueError("Narration padding must be finite and non-negative")
narration_padding = _finite_non_negative_seconds(
narration_padding,
"Narration padding",
)
narration_start_floor = _finite_non_negative_seconds(
narration_start_floor,
"Narration start floor",
)
_reject_output_alias(
output_path,
[canonical_path, plan_path],
@@ -983,9 +1072,12 @@ def rebuild_animations(
timing_plan: dict[str, list[TimingPlanEntry]] | None = None
if plan_path.is_file():
expected_srt_sha256, planned_padding, loaded_plan = _load_timing_plan(
plan_path
)
(
expected_srt_sha256,
planned_padding,
planned_start_floor,
loaded_plan,
) = _load_timing_plan(plan_path)
if not math.isclose(
narration_padding,
planned_padding,
@@ -996,6 +1088,16 @@ def rebuild_animations(
"Narration padding differs from the timing plan: "
f"plan={planned_padding}, command={narration_padding}"
)
if planned_start_floor is not None and not math.isclose(
narration_start_floor,
planned_start_floor,
rel_tol=0,
abs_tol=1e-9,
):
raise ValueError(
"Narration start floor differs from the timing plan: "
f"plan={planned_start_floor}, command={narration_start_floor}"
)
current_srt_sha256 = _subtitle_fingerprint(slide_names, subtitle_dir)
if current_srt_sha256 != expected_srt_sha256:
raise ValueError(
@@ -1049,6 +1151,17 @@ def rebuild_animations(
f'Derived animation slide "{slide_name}" must be an object'
)
settings = _effective_slide_animation(canonical, canonical_slide)
transition_duration_ms = _effective_transition_duration_ms(
canonical,
canonical_slide,
)
narration_lead_in_ms = round(
narration_lead_in_seconds(
transition_duration_ms / 1000,
start_floor=narration_start_floor,
)
* 1000
)
plan_entries = timing_plan.get(slide_name) if timing_plan else None
states, used_svg = _resolve_animation_groups(
project_path,
@@ -1122,14 +1235,18 @@ def rebuild_animations(
else:
anchored_count += 1
referenced_cues.add(cue_number)
desired_start_ms = cues[cue_number - 1].start_ms
cue_start_ms = cues[cue_number - 1].start_ms
desired_start_ms = (
narration_lead_in_ms + cue_start_ms
)
actual_start_ms = max(desired_start_ms, previous_end_ms)
delay_ms = actual_start_ms - previous_end_ms
drift_ms = actual_start_ms - desired_start_ms
if drift_ms > 500:
drift_warnings.append(
f"{slide_name}/{state.group_id}: cue {cue_number} "
f"starts at {_seconds_from_ms(desired_start_ms):.3f}s, "
f"starts at {_seconds_from_ms(cue_start_ms):.3f}s after "
f"a {_seconds_from_ms(narration_lead_in_ms):.3f}s lead-in; "
f"animation starts at {_seconds_from_ms(actual_start_ms):.3f}s "
f"(after-previous drift {_seconds_from_ms(drift_ms):.3f}s)"
)
@@ -1159,7 +1276,14 @@ def rebuild_animations(
ignored_cue_count += len(cues) - len(referenced_cues)
advance_ms = int((audio_duration + narration_padding) * 1000)
advance_ms = round(
(
audio_duration
+ narration_padding
+ narration_lead_in_ms / 1000
)
* 1000
)
if previous_end_ms > advance_ms:
raise ValueError(
f'Animations on slide "{slide_name}" end at '
@@ -1270,8 +1394,11 @@ def _presentation_slide_members(package: zipfile.ZipFile) -> list[str]:
return members
def _read_powerpoint_timings(pptx_path: Path, slide_count: int) -> list[tuple[int, int]]:
timings: list[tuple[int, int]] = []
def _read_powerpoint_timings(
pptx_path: Path,
slide_count: int,
) -> list[PowerPointTiming]:
timings: list[PowerPointTiming] = []
with zipfile.ZipFile(pptx_path) as package:
slide_members = _presentation_slide_members(package)
if len(slide_members) != slide_count:
@@ -1280,7 +1407,8 @@ def _read_powerpoint_timings(pptx_path: Path, slide_count: int) -> list[tuple[in
f"but the project has {slide_count}"
)
for slide_index, member in enumerate(slide_members, 1):
summary = read_slide_transition_xml(package.read(member))
slide_xml = package.read(member)
summary = read_slide_transition_xml(slide_xml)
if summary.logical_count != 1:
raise ValueError(
f"Narrated PPTX slide {slide_index} has "
@@ -1292,24 +1420,45 @@ def _read_powerpoint_timings(pptx_path: Path, slide_count: int) -> list[tuple[in
f"Narrated PPTX slide {slide_index} has no recorded advance time"
)
transition_ms = summary.duration_ms or 0
if advance_ms <= 0 or transition_ms < 0:
narration_delay_ms = read_narration_start_delay_xml(
slide_xml.decode("utf-8")
)
if (
advance_ms <= 0
or transition_ms < 0
or narration_delay_ms < 0
or narration_delay_ms >= advance_ms
):
raise ValueError(
f"Narrated PPTX slide {slide_index} has invalid timing values"
)
timings.append((transition_ms, advance_ms))
timings.append(
PowerPointTiming(
transition_ms=transition_ms,
narration_delay_ms=narration_delay_ms,
advance_ms=advance_ms,
)
)
return timings
def _powerpoint_audio_starts(
timings: list[tuple[int, int]],
timings: list[PowerPointTiming],
) -> tuple[list[int], int]:
"""Return theoretical narration starts and the complete PPTX timeline."""
audio_starts: list[int] = []
timeline_ms = 0
for transition_ms, advance_ms in timings:
audio_start_ms = timeline_ms + transition_ms
for timing in timings:
slide_start_ms = timeline_ms
audio_start_ms = (
slide_start_ms
+ timing.transition_ms
+ timing.narration_delay_ms
)
audio_starts.append(audio_start_ms)
timeline_ms = audio_start_ms + advance_ms
timeline_ms = (
slide_start_ms + timing.transition_ms + timing.advance_ms
)
return audio_starts, timeline_ms
@@ -1610,17 +1759,19 @@ def _merge_subtitles_result(
merged_cues: list[SubtitleCue] = []
for slide_name, audio_start_ms, (_transition_ms, advance_ms) in zip(
for slide_name, audio_start_ms, timing in zip(
slide_names,
audio_starts,
timings,
):
slide_cues = local_cues[slide_name]
if slide_cues[-1].end_ms > advance_ms:
narration_window_ms = timing.advance_ms - timing.narration_delay_ms
if slide_cues[-1].end_ms > narration_window_ms:
raise ValueError(
f"{slide_name}.srt ends at "
f"{_seconds_from_ms(slide_cues[-1].end_ms):.3f}s, after the "
f"PowerPoint slide advance at {_seconds_from_ms(advance_ms):.3f}s"
"available narration window before PowerPoint advances at "
f"{_seconds_from_ms(narration_window_ms):.3f}s"
)
for cue in slide_cues:
merged_cue = SubtitleCue(
@@ -1757,6 +1908,16 @@ def build_parser() -> argparse.ArgumentParser:
default=0.5,
help="Seconds added after each narration before slide advance (default: 0.5)",
)
animations.add_argument(
"--narration-start-floor",
type=float,
default=DEFAULT_NARRATION_START_FLOOR,
help=(
"Minimum seconds from transition start to narration start; "
"0 waits only for transition completion "
f"(default: {DEFAULT_NARRATION_START_FLOOR:g})"
),
)
animations.add_argument(
"--force",
action="store_true",
@@ -1865,6 +2026,7 @@ def main(argv: list[str] | None = None) -> int:
audio_dir=audio_dir,
output_path=output_path,
narration_padding=args.narration_padding,
narration_start_floor=args.narration_start_floor,
force=args.force,
)
print(f"Narration animation config written: {output_path}")
@@ -59,6 +59,7 @@ from pptx_transitions import ( # noqa: E402
NATIVE_TRANSITION_KEYS,
apply_slide_motion_xml,
normalize_transition_effect_request,
read_slide_transition_xml,
set_directory_use_timings,
validate_pptx_transition_package,
validate_seconds,
@@ -72,10 +73,12 @@ from svg_to_pptx.pptx_package.narration import ( # noqa: E402
AUDIO_CONTENT_TYPES,
AUDIO_MARKER_PNG_BYTES,
AUDIO_REL_TYPE,
DEFAULT_NARRATION_START_FLOOR,
IMAGE_REL_TYPE,
MEDIA_REL_TYPE,
NARRATION_EXTENSIONS,
inject_narration,
narration_lead_in_seconds,
next_shape_id,
probe_audio_duration,
)
@@ -902,6 +905,7 @@ def _apply_audio(
enter: EnterUpdate,
timings_enabled: bool,
narration_padding: float,
narration_start_floor: float,
audio_duration: float | None = None,
) -> bool:
media_dir = extract_dir / "ppt" / "media"
@@ -925,6 +929,24 @@ def _apply_audio(
slide_xml_path = extract_dir / slide.part_name
slide_xml = slide_xml_path.read_text(encoding="utf-8")
source_animation_fingerprint = object_animation_fingerprint(slide_xml)
source_transition = read_slide_transition_xml(slide_xml)
if enter.policy == "replace":
transition_duration = enter.duration
elif enter.policy == "none":
transition_duration = 0.0
else:
# Legacy spd-only transitions expose no exact milliseconds. Treat
# those as unknown and keep the full configured floor after the
# preserved transition rather than guessing an application duration.
transition_duration = (
source_transition.duration_ms / 1000
if source_transition.duration_ms is not None
else 0.0
)
narration_lead_in = narration_lead_in_seconds(
transition_duration,
start_floor=narration_start_floor,
)
shape_id = next_shape_id(slide_xml)
slide_xml = inject_narration(
slide_xml,
@@ -933,18 +955,38 @@ def _apply_audio(
audio_rid=audio_rid,
media_rid=media_rid,
poster_rid=poster_rid,
start_delay=narration_lead_in,
)
advance = AdvanceUpdate(mode="preserve")
if timings_enabled:
duration = audio_duration
duration = audio_duration
if duration is None and (
timings_enabled or source_transition.advance_after_ms is not None
):
duration = probe_audio_duration(audio_path)
if (
not timings_enabled
and source_transition.advance_after_ms is not None
):
if duration is None:
duration = probe_audio_duration(audio_path)
raise RuntimeError(
f"Unable to validate narration against slide {slide.index} "
f"auto-advance with ffprobe: {audio_path}"
)
required_playback_ms = round((narration_lead_in + duration) * 1000)
if source_transition.advance_after_ms < required_playback_ms:
raise RuntimeError(
f"Slide {slide.index} advances after "
f"{source_transition.advance_after_ms} ms, before delayed "
f"narration can finish at {required_playback_ms} ms; enable "
"timings or lengthen the source auto-advance"
)
if timings_enabled:
if duration is None:
raise RuntimeError(f"Unable to read narration duration with ffprobe: {audio_path}")
advance = AdvanceUpdate(
mode="narration",
after=duration + narration_padding,
after=narration_lead_in + duration + narration_padding,
)
wrote_advance = False
@@ -1223,6 +1265,7 @@ def _build_enhancement_plan(
transition: str | None,
transition_duration: float | None,
narration_padding: float | None,
narration_start_floor: float | None,
apply_transition_without_audio: bool | None,
existing_plan: dict | None = None,
) -> dict:
@@ -1241,6 +1284,19 @@ def _build_enhancement_plan(
"narration padding",
allow_zero=True,
)
raw_start_floor: object
if narration_start_floor is not None:
raw_start_floor = narration_start_floor
else:
raw_start_floor = previous_timings.get(
"narration_start_floor",
DEFAULT_NARRATION_START_FLOOR,
)
resolved_start_floor = validate_seconds(
raw_start_floor,
"narration start floor",
allow_zero=True,
)
transitions_enabled, transition_config = _resolved_draft_transition_config(
project,
previous,
@@ -1297,6 +1353,7 @@ def _build_enhancement_plan(
),
"source": "audio_duration",
"narration_padding": resolved_padding,
"narration_start_floor": resolved_start_floor,
},
"transitions": {
"enabled": transitions_enabled,
@@ -1729,6 +1786,7 @@ def init_project(args: argparse.Namespace) -> int:
transition=args.transition,
transition_duration=args.transition_duration,
narration_padding=args.narration_padding,
narration_start_floor=args.narration_start_floor,
apply_transition_without_audio=args.apply_transition_without_audio,
)
_write_json(_plan_path(project_path), plan)
@@ -1813,6 +1871,7 @@ def plan_project(args: argparse.Namespace) -> int:
transition=args.transition,
transition_duration=args.transition_duration,
narration_padding=args.narration_padding,
narration_start_floor=args.narration_start_floor,
apply_transition_without_audio=args.apply_transition_without_audio,
existing_plan=existing_plan,
)
@@ -1915,6 +1974,25 @@ def apply_project(args: argparse.Namespace) -> int:
except ValueError as exc:
return fail_preflight([str(exc)])
if args.narration_start_floor is not None:
raw_narration_start_floor = args.narration_start_floor
elif "narration_start_floor" in timings_cfg:
raw_narration_start_floor = timings_cfg["narration_start_floor"]
else:
raw_narration_start_floor = DEFAULT_NARRATION_START_FLOOR
try:
if "audio" in modules:
narration_start_floor = validate_seconds(
raw_narration_start_floor,
"narration start floor",
allow_zero=True,
)
else:
narration_start_floor = DEFAULT_NARRATION_START_FLOOR
except ValueError as exc:
return fail_preflight([str(exc)])
output_path = (
Path(args.output).expanduser().resolve()
if args.output
@@ -2034,6 +2112,7 @@ def apply_project(args: argparse.Namespace) -> int:
enter=enter_update,
timings_enabled="timings" in modules,
narration_padding=narration_padding,
narration_start_floor=narration_start_floor,
audio_duration=readiness.audio_durations.get(slide.index),
) or wrote_auto_advance
audio_exts.add(audio.suffix.lower())
@@ -2330,6 +2409,12 @@ def build_parser() -> argparse.ArgumentParser:
)
init.add_argument("--transition-duration", type=_positive_seconds_arg, default=0.5)
init.add_argument("--narration-padding", type=_non_negative_seconds_arg, default=0.4)
init.add_argument(
"--narration-start-floor",
type=_non_negative_seconds_arg,
default=DEFAULT_NARRATION_START_FLOOR,
help="minimum seconds from transition start to narration start",
)
init.add_argument(
"--apply-transition-without-audio",
action="store_true",
@@ -2361,6 +2446,12 @@ def build_parser() -> argparse.ArgumentParser:
type=_non_negative_seconds_arg,
default=None,
)
plan.add_argument(
"--narration-start-floor",
type=_non_negative_seconds_arg,
default=None,
help="replace the saved narration start floor; omitted values preserve it",
)
plan.add_argument(
"--apply-transition-without-audio",
action="store_true",
@@ -2387,6 +2478,12 @@ def build_parser() -> argparse.ArgumentParser:
)
apply.add_argument("--transition-duration", type=_positive_seconds_arg, default=None)
apply.add_argument("--narration-padding", type=_non_negative_seconds_arg, default=None)
apply.add_argument(
"--narration-start-floor",
type=_non_negative_seconds_arg,
default=None,
help="override the confirmed narration start floor for this export",
)
apply.add_argument("--force", action="store_true", help="apply without a confirmed enhancement plan")
apply.add_argument(
"--apply-transition-without-audio",
@@ -37,7 +37,7 @@
"skills/ppt-master/references/strategist-template.md": 3000,
"skills/ppt-master/templates/design_spec_reference.md": 3750,
"skills/ppt-master/templates/spec_lock_reference.md": 2750,
"skills/ppt-master/workflows/generate-pptx.md": 16000,
"skills/ppt-master/workflows/generate-pptx.md": 18000,
"skills/ppt-master/workflows/stages/apply-template-workspace.md": 3500
},
"load_sets": {
@@ -49,7 +49,7 @@
"skills/ppt-master/SKILL.md",
"skills/ppt-master/workflows/routing.md"
],
"max_tokens": 6750
"max_tokens": 7750
},
"support.sponsor-recommendation.en": {
"description": "English sponsor context for explicit user requests for model, AI image model, API/provider, or hosted-service recommendations.",
@@ -140,7 +140,7 @@
"files": [
"skills/ppt-master/workflows/native-enhance-pptx.md"
],
"max_tokens": 11000
"max_tokens": 13000
},
"route.fill-native-pptx": {
"description": "Raw-PPTX native fill route.",
@@ -258,6 +258,32 @@
"files": [],
"max_tokens": 100000
},
"route.generate.quick-generate.image-to-pptx": {
"description": "Codex-supported Quick-only Image to PPTX with registered reference-image layers and shared plates.",
"scope": "cumulative",
"include": [
"route.generate.quick-generate.in-hand-image",
"profile.generate.image-to-pptx"
],
"files": [
"skills/ppt-master/references/image-base.md",
"skills/ppt-master/references/image-generator.md"
],
"max_tokens": 130000
},
"route.generate.quick-generate.video-design": {
"description": "Quick Generate with conditional video-delivery design, frozen-script notes, custom animation, and narration/audio delivery.",
"scope": "cumulative",
"include": [
"route.generate.quick-generate",
"stage.generate.video-design",
"stage.generate.executor.notes",
"stage.generate.customize-animations",
"stage.shared.generate-audio"
],
"files": [],
"max_tokens": 140000
},
"route.generate.quick-generate.template-fusion": {
"description": "Quick Generate with the maximum Brand, Style, Layout, and Deck exact-root fusion loaded for flat authoring.",
"scope": "cumulative",
@@ -558,7 +584,7 @@
"stage.shared.generate-audio"
],
"files": [],
"max_tokens": 17000
"max_tokens": 19000
},
"route.generate.beautify-flat-no-image": {
"description": "Generate-PPTX 1:1 beautify profile on the flat no-image path.",
@@ -570,6 +596,18 @@
"files": [],
"max_tokens": 175000
},
"route.generate.flat-no-image.video-design": {
"description": "Default flat Generate-PPTX path with conditional video-delivery design, custom animation, and narration/audio delivery.",
"scope": "cumulative",
"include": [
"route.generate.flat-no-image",
"stage.generate.video-design",
"stage.generate.customize-animations",
"stage.shared.generate-audio"
],
"files": [],
"max_tokens": 215000
},
"route.generate.flat-no-image-visualization": {
"description": "Default flat Generate-PPTX path with visualization-reference consumption, all three information-model branches, optional native-data replacement, and chart verification.",
"scope": "cumulative",
@@ -773,6 +811,22 @@
],
"max_tokens": 7750
},
"profile.generate.image-to-pptx": {
"description": "Codex-supported, Quick-only profile for reconstructing raster page frames as layered editable PPTX.",
"scope": "incremental",
"files": [
"skills/ppt-master/workflows/profiles/image-to-pptx.md"
],
"max_tokens": 4000
},
"stage.generate.video-design": {
"description": "Conditional recorded/self-running delivery design and prepared final-script handling shared by Default and Quick.",
"scope": "incremental",
"files": [
"skills/ppt-master/references/video-design.md"
],
"max_tokens": 2500
},
"stage.generate.refine-spec": {
"description": "Explicit post-confirmation spec refinement runbook.",
"scope": "incremental",
@@ -850,7 +904,7 @@
"files": [
"skills/ppt-master/workflows/stages/generate-audio.md"
],
"max_tokens": 6000
"max_tokens": 7000
},
"governance.failure-recovery": {
"description": "Global stop, retry, and resume policy.",
@@ -86,6 +86,7 @@ try:
from svg_to_pptx.drawingml.utils import (
IDENTITY_MATRIX as _IDENTITY_MATRIX,
PROJECT_PAINT_PROPERTIES as _PAINT_PROPERTIES,
PROJECT_TEXT_IMAGE_FILL_ATTR as _TEXT_IMAGE_FILL_ATTR,
detect_text_lang as _detect_text_lang,
matrix_multiply as _matrix_multiply,
parse_inline_style as _parse_inline_style,
@@ -103,6 +104,7 @@ try:
except ImportError:
_IDENTITY_MATRIX = None
_PAINT_PROPERTIES = None
_TEXT_IMAGE_FILL_ATTR = 'data-pptx-text-image-fill'
_detect_text_lang = None
_matrix_multiply = None
_parse_inline_style = None
@@ -3659,7 +3661,9 @@ class SVGQualityChecker:
svg_to_pptx maps <pattern fill> to native <a:pattFill prst="...">. The
preset name comes from `data-pptx-pattern` (e.g. `lgGrid` / `smGrid` /
`dkUpDiag`). Two failure modes worth catching pre-export:
`dkUpDiag`). Patterns marked with `data-pptx-text-image-fill` instead
map to run-level <a:blipFill> and are validated by converter preflight.
Two preset-pattern failure modes are worth catching pre-export:
1. Missing annotation the converter compatibility fallback chooses
`ltUpDiag` (diagonal stripes), which is not an authoring contract.
@@ -3691,6 +3695,8 @@ class SVGQualityChecker:
):
pat_id = pattern.get('id', '<unnamed>')
prst = pattern.get('data-pptx-pattern')
if pattern.get(_TEXT_IMAGE_FILL_ATTR) is not None:
continue
if pat_id in referenced_patterns and not prst:
result['warnings'].append(
f"Fidelity warning: <pattern id=\"{pat_id}\"> has no "
@@ -127,6 +127,8 @@ class ConvertContext:
# Canonical BCP-47 content language from spec_lock.md. ``None`` preserves
# the legacy per-run script heuristic for older projects and lockless quick generation.
primary_language: str | None = None
# Reuse media relationships when the same image fills multiple text runs.
text_image_fill_cache: dict[tuple[str, str], str] = field(default_factory=dict)
def next_id(self) -> int:
"""Allocate the next shape ID."""
@@ -277,6 +279,7 @@ class ConvertContext:
theme_font_spec=self.theme_font_spec,
theme_color_spec=self.theme_color_spec,
primary_language=self.primary_language,
text_image_fill_cache=self.text_image_fill_cache,
)
def sync_from_child(self, child_ctx: ConvertContext) -> None:
@@ -56,7 +56,7 @@ from .utils import (
ctx_x, ctx_y, ctx_w, ctx_h,
rect_to_dml_xfrm,
combine_opacity, parse_hex_color, parse_svg_color,
resolve_url_id, get_effective_filter_id,
resolve_project_text_image_fill, resolve_url_id, get_effective_filter_id,
parse_inline_style, parse_font_family, is_cjk_char,
detect_text_lang, estimate_text_cluster_widths, font_px_to_hpt,
resolve_text_run_fonts, split_project_text_clusters,
@@ -2463,14 +2463,32 @@ def _build_text_fill_xml(
if fill_raw.strip().lower() in ('none', 'transparent'):
return '<a:noFill/>'
grad_id = resolve_url_id(fill_raw)
if grad_id and ctx and grad_id in ctx.defs:
return build_gradient_fill(
ctx.defs[grad_id],
opacity,
ctx.theme_color_spec,
"text",
)
paint_id = resolve_url_id(fill_raw)
if paint_id and ctx and paint_id in ctx.defs:
paint = ctx.defs[paint_id]
paint_tag = paint.tag.rsplit('}', 1)[-1]
if paint_tag in {'linearGradient', 'radialGradient'}:
return build_gradient_fill(
paint,
opacity,
ctx.theme_color_spec,
"text",
)
if paint_tag == 'pattern':
mode, image = resolve_project_text_image_fill(paint)
source = load_project_image_source(image, ctx.svg_dir)
r_id = _register_image_media(
ctx,
source.img_format,
source.img_data,
reuse_text_fill=True,
)
blip_xml = _build_image_blip_xml(r_id, opacity)
if mode == 'stretch':
fill_mode_xml = '<a:stretch><a:fillRect/></a:stretch>'
else:
fill_mode_xml = '<a:tile/>'
return f'<a:blipFill>{blip_xml}{fill_mode_xml}</a:blipFill>'
parsed_color, color_alpha = parse_svg_color(fill_raw)
fill = parsed_color or fill
@@ -4267,6 +4285,34 @@ def _build_image_blip_xml(r_id: str, opacity: float | None) -> str:
)
def _register_image_media(
ctx: ConvertContext,
img_format: str,
img_data: bytes,
*,
reuse_text_fill: bool = False,
) -> str:
"""Register image bytes and return a DrawingML relationship id."""
cache_key: tuple[str, str] | None = None
if reuse_text_fill:
cache_key = (img_format, hashlib.sha256(img_data).hexdigest())
if cache_key in ctx.text_image_fill_cache:
return ctx.text_image_fill_cache[cache_key]
img_idx = len(ctx.media_files) + 1
img_filename = f's{ctx.slide_num}_img{img_idx}.{img_format}'
ctx.media_files[img_filename] = img_data
r_id = ctx.next_rel_id()
ctx.rel_entries.append({
'id': r_id,
'type': 'http://schemas.openxmlformats.org/officeDocument/2006/relationships/image',
'target': f'../media/{img_filename}',
})
if cache_key is not None:
ctx.text_image_fill_cache[cache_key] = r_id
return r_id
def convert_image(elem: ET.Element, ctx: ConvertContext) -> ShapeResult | None:
"""Convert SVG <image> to DrawingML picture element.
@@ -4342,16 +4388,7 @@ def convert_image(elem: ET.Element, ctx: ConvertContext) -> ShapeResult | None:
visible_source_fraction=visible_source_fraction,
)
img_idx = len(ctx.media_files) + 1
img_filename = f's{ctx.slide_num}_img{img_idx}.{img_format}'
ctx.media_files[img_filename] = img_data
r_id = ctx.next_rel_id()
ctx.rel_entries.append({
'id': r_id,
'type': 'http://schemas.openxmlformats.org/officeDocument/2006/relationships/image',
'target': f'../media/{img_filename}',
})
r_id = _register_image_media(ctx, img_format, img_data)
# Resolve clip-path → DrawingML geometry
clip_geom = _resolve_clip_geometry(elem, ctx, raw_x, raw_y, raw_w, raw_h)
@@ -4965,16 +5002,7 @@ def convert_nested_svg(elem: ET.Element, ctx: ConvertContext) -> ShapeResult:
),
)
img_idx = len(ctx.media_files) + 1
img_filename = f's{ctx.slide_num}_img{img_idx}.{img_format}'
ctx.media_files[img_filename] = img_data
r_id = ctx.next_rel_id()
ctx.rel_entries.append({
'id': r_id,
'type': 'http://schemas.openxmlformats.org/officeDocument/2006/relationships/image',
'target': f'../media/{img_filename}',
})
r_id = _register_image_media(ctx, img_format, img_data)
shape_id = _claim_element_shape_id(elem, ctx)
xfrm_attr, off_x, off_y, ext_cx, ext_cy, bounds_emu = _picture_xfrm_from_svg_rect(
@@ -227,6 +227,8 @@ PROJECT_DEFINITION_TAGS = frozenset({
'radialGradient',
})
PROJECT_GRADIENT_TAGS = frozenset({'linearGradient', 'radialGradient'})
PROJECT_TEXT_IMAGE_FILL_ATTR = 'data-pptx-text-image-fill'
PROJECT_TEXT_IMAGE_FILL_MODES = frozenset({'stretch', 'tile'})
# PPTX angle projection can overshoot a unit box by at most ~0.1036.
PROJECT_LINEAR_GRADIENT_COORDINATE_MIN = -0.105
PROJECT_LINEAR_GRADIENT_COORDINATE_MAX = 1.105
@@ -2157,6 +2159,51 @@ def project_marker_errors(root: ET.Element) -> list[str]:
return sorted(errors)
def resolve_project_text_image_fill(pattern: ET.Element) -> tuple[str, ET.Element]:
"""Resolve the controlled one-image pattern used for native text picture fills."""
if _svg_element_tag(pattern) != 'pattern':
raise ValueError('definition must be an SVG <pattern>')
mode = pattern.get(PROJECT_TEXT_IMAGE_FILL_ATTR, '')
if mode not in PROJECT_TEXT_IMAGE_FILL_MODES:
supported = ', '.join(sorted(PROJECT_TEXT_IMAGE_FILL_MODES))
raise ValueError(f'{PROJECT_TEXT_IMAGE_FILL_ATTR} must be one of: {supported}')
preset_attributes = [
name
for name in ('data-pptx-pattern', 'data-pptx-fg', 'data-pptx-bg')
if pattern.get(name) is not None
]
if preset_attributes:
raise ValueError(
'text image fill must not combine preset-pattern attributes: '
f'{", ".join(preset_attributes)}'
)
if pattern.get('patternTransform') is not None:
raise ValueError('pattern must not use patternTransform')
children = list(pattern)
if len(children) != 1 or children[0].tag != f'{{{SVG_NS}}}image':
raise ValueError('pattern must contain exactly one direct SVG <image> child')
image = children[0]
unsupported = [
name
for name in (
'clip-path',
'filter',
'fill-opacity',
'mask',
'opacity',
'style',
'transform',
)
if image.get(name) is not None
]
if unsupported:
raise ValueError(f"pattern image must not use {', '.join(unsupported)}")
return mode, image
def project_paint_reference_errors(root: ET.Element) -> list[str]:
"""Validate local paint-server references and their native contexts."""
definitions, _duplicates = project_definition_index(root)
@@ -2215,7 +2262,11 @@ def project_paint_reference_errors(root: ET.Element) -> list[str]:
elif property_name == 'stroke' and elem_tag_lower in stroke_shape_tags:
allowed_tags = ('lineargradient', 'radialgradient')
elif property_name == 'fill' and elem_tag_lower in {'text', 'tspan'}:
allowed_tags = ('lineargradient', 'radialgradient')
target_tag = (_svg_element_tag(target) or str(target.tag)).lower()
if target_tag == 'pattern' and target.get(PROJECT_TEXT_IMAGE_FILL_ATTR) is not None:
allowed_tags = ('lineargradient', 'radialgradient', 'pattern')
else:
allowed_tags = ('lineargradient', 'radialgradient')
elif property_name == 'fill' and elem_tag_lower == 'g':
allowed_tags = (
('lineargradient', 'radialgradient')
@@ -2239,6 +2290,19 @@ def project_paint_reference_errors(root: ET.Element) -> list[str]:
continue
target_tag = (_svg_element_tag(target) or str(target.tag)).lower()
is_text_image_fill = (
target_tag == 'pattern'
and target.get(PROJECT_TEXT_IMAGE_FILL_ATTR) is not None
)
if is_text_image_fill and not (
property_name == 'fill'
and elem_tag_lower in {'text', 'tspan'}
):
errors.add(
f'<{elem_tag}> {property_name}=url(#{reference_id}) uses a '
'text image fill pattern outside <text>/<tspan>'
)
continue
if target_tag not in allowed_tags:
tag_labels = {
'lineargradient': 'linearGradient',
@@ -2251,6 +2315,16 @@ def project_paint_reference_errors(root: ET.Element) -> list[str]:
f'to <{_svg_element_tag(target) or target.tag}>; expected '
f'{expected}'
)
continue
if property_name == 'fill' and elem_tag_lower in {'text', 'tspan'} and target_tag == 'pattern':
try:
resolve_project_text_image_fill(target)
except ValueError as exc:
errors.add(
f'<{elem_tag}> fill=url(#{reference_id}) has an invalid '
f'text image fill: {exc}'
)
return sorted(errors)
@@ -99,10 +99,12 @@ from .narration import (
AUDIO_CONTENT_TYPES,
AUDIO_REL_TYPE,
AUDIO_MARKER_PNG_BYTES,
DEFAULT_NARRATION_START_FLOOR,
IMAGE_REL_TYPE,
MEDIA_REL_TYPE,
apply_recorded_timing,
inject_narration,
narration_lead_in_seconds,
next_shape_id,
probe_audio_duration,
)
@@ -4683,6 +4685,7 @@ def create_pptx_with_native_svg(
transition_effect_options: dict[str, object] | None = None,
text_flow: str | None = None,
primary_language: str | None = None,
narration_start_floor: float = DEFAULT_NARRATION_START_FLOOR,
) -> bool:
"""Create a PPTX file with native DrawingML shapes.
@@ -4723,6 +4726,8 @@ def create_pptx_with_native_svg(
narration_audio: Optional dict mapping SVG stem to narration audio file.
use_narration_timings: Whether to set slide auto-advance from audio duration.
narration_padding: Extra seconds added after each narration before advancing.
narration_start_floor: Minimum seconds from transition start to narration
start. Any remainder after the transition becomes silent lead-in.
merge_paragraphs: Legacy compatibility option. True selects reflow;
False selects split. Do not combine with ``text_flow``.
text_flow: Positional-tspan policy: preserve authored line breaks in
@@ -5586,6 +5591,15 @@ def create_pptx_with_native_svg(
slide_xml = slide_xml_path.read_text(encoding='utf-8')
narration_shape_id = next_shape_id(slide_xml)
narration_transition_duration = (
slide_transition_duration
if slide_transition is not None
else 0.0
)
narration_lead_in = narration_lead_in_seconds(
narration_transition_duration,
start_floor=narration_start_floor,
)
slide_xml = inject_narration(
slide_xml,
shape_id=narration_shape_id,
@@ -5593,6 +5607,7 @@ def create_pptx_with_native_svg(
audio_rid=audio_rid,
media_rid=media_rid,
poster_rid=poster_rid,
start_delay=narration_lead_in,
)
if use_narration_timings:
@@ -5601,13 +5616,16 @@ def create_pptx_with_native_svg(
raise RuntimeError(
f"Unable to read narration duration with ffprobe: {audio_path}"
)
narration_advance_after = (
narration_lead_in + duration + narration_padding
)
slide_xml = apply_recorded_timing(
slide_xml,
advance_after=duration + narration_padding,
advance_after=narration_advance_after,
transition_duration=slide_transition_duration,
transition_effect=slide_transition,
)
resolved_advance_after = duration + narration_padding
resolved_advance_after = narration_advance_after
resolved_advance_on_click = False
package_uses_timings = True
slide_xml_path.write_text(slide_xml, encoding='utf-8')
@@ -70,7 +70,12 @@ from ..drawingml.theme_fonts import (
load_theme_font_spec,
)
from ..drawingml.utils import unsafe_exported_font_faces
from .narration import NARRATION_EXTENSIONS, find_narration_files, probe_audio_duration
from .narration import (
DEFAULT_NARRATION_START_FLOOR,
NARRATION_EXTENSIONS,
find_narration_files,
probe_audio_duration,
)
from .template_structure import (
TemplateStructureError,
load_pptx_structure_lock,
@@ -1096,6 +1101,16 @@ Recorded narration:
'(<project>_<ts>_narrated.pptx) to tell them apart from silent exports.')
parser.add_argument('--narration-padding', type=non_negative_float, default=0.5,
help='Seconds to add after each narration before auto-advance (default: 0.5)')
parser.add_argument(
'--narration-start-floor',
type=non_negative_float,
default=DEFAULT_NARRATION_START_FLOOR,
help=(
'Minimum seconds from slide-transition start to narration start; '
'set 0 to start immediately after the transition '
f'(default: {DEFAULT_NARRATION_START_FLOOR})'
),
)
parser.add_argument(
'--inherit-motion-from',
type=str,
@@ -1899,6 +1914,7 @@ Recorded narration:
narration_audio=narration_audio,
use_narration_timings=use_narration_timings,
narration_padding=args.narration_padding,
narration_start_floor=args.narration_start_floor,
text_flow=args.text_flow,
image_optimize=not args.no_image_optimize,
image_max_dimension=args.image_max_dimension,
@@ -20,6 +20,7 @@ from pptx_transitions import (
parse_source_xml,
read_slide_transition_xml,
serialize_source_xml,
validate_seconds,
)
@@ -40,6 +41,7 @@ AUDIO_CONTENT_TYPES = {
}
NARRATION_EXTENSIONS = tuple(AUDIO_CONTENT_TYPES.keys())
DEFAULT_NARRATION_START_FLOOR = 0.8
AUDIO_MARKER_SIZE_EMU = 457200 # 48 SVG px
AUDIO_MARKER_OFF_CANVAS_EMU = -AUDIO_MARKER_SIZE_EMU
@@ -300,7 +302,30 @@ def _numeric_ids(
return ids
def _create_audio_timing_element(shape_id: int, ctn_id: int) -> ET.Element:
def narration_lead_in_seconds(
transition_duration: float,
*,
start_floor: float = DEFAULT_NARRATION_START_FLOOR,
) -> float:
"""Return silence after a slide transition before narration starts."""
transition_seconds = validate_seconds(
transition_duration,
"narration transition duration",
allow_zero=True,
)
floor_seconds = validate_seconds(
start_floor,
"narration start floor",
allow_zero=True,
)
return max(0.0, floor_seconds - transition_seconds)
def _create_audio_timing_element(
shape_id: int,
ctn_id: int,
start_delay_ms: int,
) -> ET.Element:
audio = ET.Element(_qn(PML_NS, "audio"))
media_node = ET.SubElement(
audio,
@@ -313,7 +338,11 @@ def _create_audio_timing_element(shape_id: int, ctn_id: int) -> ET.Element:
{"id": str(ctn_id), "fill": "hold", "display": "0"},
)
start_conditions = ET.SubElement(time_node, _qn(PML_NS, "stCondLst"))
ET.SubElement(start_conditions, _qn(PML_NS, "cond"), {"delay": "0"})
ET.SubElement(
start_conditions,
_qn(PML_NS, "cond"),
{"delay": str(start_delay_ms)},
)
target = ET.SubElement(media_node, _qn(PML_NS, "tgtEl"))
ET.SubElement(target, _qn(PML_NS, "spTgt"), {"spid": str(shape_id)})
return audio
@@ -466,8 +495,9 @@ def inject_narration(
audio_rid: str,
media_rid: str,
poster_rid: str,
start_delay: float = 0.0,
) -> str:
"""Inject a hidden narration media shape and slide-entry autoplay timing."""
"""Inject a hidden narration shape and delayed slide-entry autoplay timing."""
if isinstance(shape_id, bool) or not isinstance(shape_id, int) or shape_id <= 0:
raise ValueError("narration shape_id must be a positive integer")
if shape_id > MAX_OOXML_UNSIGNED_INT:
@@ -475,6 +505,17 @@ def inject_narration(
"narration shape_id exceeds the OOXML unsigned-integer limit: "
f"{shape_id}"
)
start_delay_seconds = validate_seconds(
start_delay,
"narration start delay",
allow_zero=True,
)
start_delay_ms = round(start_delay_seconds * 1000)
if start_delay_ms > MAX_OOXML_UNSIGNED_INT:
raise ValueError(
"narration start delay exceeds the OOXML unsigned-integer limit: "
f"{start_delay_ms} ms"
)
root = parse_source_xml(slide_xml)
if root.tag != _qn(PML_NS, "sld"):
@@ -533,15 +574,101 @@ def inject_narration(
"tmRoot/p:childTnLst",
)
child_nodes.append(
_create_audio_timing_element(shape_id, next_timing_id)
_create_audio_timing_element(
shape_id,
next_timing_id,
start_delay_ms,
)
)
else:
audio_timing = _create_audio_timing_element(shape_id, next_timing_id + 1)
audio_timing = _create_audio_timing_element(
shape_id,
next_timing_id + 1,
start_delay_ms,
)
_insert_root_timing(root, _new_timing(audio_timing, next_timing_id))
return serialize_source_xml(root, slide_xml).decode("utf-8")
def read_narration_start_delay_xml(slide_xml: str) -> int:
"""Return the latest embedded narration picture's autoplay delay in ms."""
root = parse_source_xml(slide_xml)
if root.tag != _qn(PML_NS, "sld"):
raise ValueError("narration source XML root must be p:sld")
audio_shape_properties: list[ET.Element] = []
for picture in root.iter(_qn(PML_NS, "pic")):
if not any(
element.tag == _qn(DRAWINGML_NS, "audioFile")
for element in picture.iter()
):
continue
properties = list(picture.iter(_qn(PML_NS, "cNvPr")))
if len(properties) != 1:
raise ValueError(
"narration audio picture must contain exactly one p:cNvPr"
)
audio_shape_properties.extend(properties)
if not audio_shape_properties:
raise ValueError("narration source has no embedded audio picture")
narration_shape_id = max(
_numeric_ids(
audio_shape_properties,
"narration audio shape",
minimum=1,
)
)
delays: list[int] = []
for audio in root.iter(_qn(PML_NS, "audio")):
media_node = audio.find(_qn(PML_NS, "cMediaNode"))
if media_node is None:
continue
target = media_node.find(
f"{_qn(PML_NS, 'tgtEl')}/{_qn(PML_NS, 'spTgt')}"
)
if target is None or target.get("spid") != str(narration_shape_id):
continue
time_node = _direct_child(
media_node,
_qn(PML_NS, "cTn"),
"p:audio/p:cMediaNode/p:cTn",
)
start_conditions = _direct_child(
time_node,
_qn(PML_NS, "stCondLst"),
"p:cTn/p:stCondLst",
)
condition = _direct_child(
start_conditions,
_qn(PML_NS, "cond"),
"p:stCondLst/p:cond",
)
raw_delay = condition.get("delay")
try:
delay = int(raw_delay)
except (TypeError, ValueError) as exc:
raise ValueError(
f"narration timing has invalid start delay: {raw_delay!r}"
) from exc
if delay < 0 or delay > MAX_OOXML_UNSIGNED_INT:
raise ValueError(
"narration timing start delay is outside the OOXML "
f"unsigned-integer range: {delay}"
)
delays.append(delay)
if not delays:
raise ValueError("narration source has no autoplay timing for its audio picture")
if len(set(delays)) != 1:
raise ValueError(
"narration source timing branches disagree on autoplay delay: "
f"{sorted(set(delays))}"
)
return delays[0]
def apply_recorded_timing(
slide_xml: str,
*,
@@ -2,10 +2,10 @@
"""
PPT Master - Visual Review Renderer
Renders project SVGs to 1280x720 PNGs that match the live-preview browser view
(inlined <use data-icon>, resolved <image href>, full font fallback including CJK).
The pure renderer for the visual-review stage does not edit SVGs, does not
interpret the rubric.
Renders project SVGs at their root viewBox dimensions to PNGs that match the
live-preview browser view (inlined <use data-icon>, resolved <image href>, full
font fallback including CJK). The pure renderer for the visual-review stage
does not edit SVGs, does not interpret the rubric.
Backend: Playwright (Chromium). The cairosvg backend was evaluated and rejected
because cairo's text API has no font-fallback chain — CJK characters render as
@@ -30,6 +30,7 @@ from __future__ import annotations
import argparse
import io
import json
import math
import os
import sys
import time
@@ -37,11 +38,14 @@ import urllib.error
import urllib.parse
import urllib.request
from contextlib import contextmanager
from decimal import Decimal
from pathlib import Path
from xml.etree import ElementTree as ET
from console_encoding import configure_utf8_stdio
from server_common import lock_pid, process_alive, read_lock
from slide_roster import discover_slide_svgs
from svg_to_pptx.canvas_contract import parse_project_svg_root
configure_utf8_stdio()
@@ -101,8 +105,8 @@ def is_all_background(png_bytes: bytes) -> bool:
return False
img = Image.open(io.BytesIO(png_bytes)).convert('RGB')
pixels = list(img.getdata())
total = len(pixels)
pixels = img.getdata()
total = img.width * img.height
if total == 0:
return True
counts: dict[tuple[int, int, int], int] = {}
@@ -113,17 +117,42 @@ def is_all_background(png_bytes: bytes) -> bool:
return dominant / total >= ALL_BG_THRESHOLD
def fetch_slide_text(server_url: str, page_name: str, timeout: float = 5.0) -> int:
"""Probe that the server can return the slide. Returns content length.
Used only for failure detection the actual fetch happens inside the
browser via fetch() so the response is parsed by JS, not Python."""
def fetch_slide_content(server_url: str, page_name: str, timeout: float = 5.0) -> str:
"""Return the live-preview server's inlined SVG content for one slide."""
url = f"{server_url.rstrip('/')}/api/slide/{urllib.parse.quote(page_name)}"
req = urllib.request.Request(url, headers={'Accept': 'application/json'})
with urllib.request.urlopen(req, timeout=timeout) as resp:
payload = json.loads(resp.read().decode('utf-8'))
if 'content' not in payload:
content = payload.get('content') if isinstance(payload, dict) else None
if not isinstance(content, str):
raise RuntimeError(f'unexpected response shape from {url}: {payload!r}')
return len(payload['content'])
return content
def _json_number(value: Decimal) -> int | float:
"""Keep integral canvas values compact while preserving fractional input."""
if value == value.to_integral_value():
return int(value)
return float(value)
def parse_slide_canvas(svg_content: str, page_name: str) -> dict:
"""Read the authoritative canvas from one inlined SVG root viewBox."""
try:
root = ET.fromstring(svg_content)
except ET.ParseError as exc:
raise ValueError(f'{page_name}: unable to parse root SVG: {exc}') from exc
viewbox = parse_project_svg_root(root, context=page_name)
width = _json_number(viewbox.width)
height = _json_number(viewbox.height)
return {
'view_box': [_json_number(value) for value in viewbox.values],
'width': width,
'height': height,
'png_width': math.ceil(float(viewbox.width)),
'png_height': math.ceil(float(viewbox.height)),
}
def render_pages(server_url: str, pages: list[str], preview_dir: Path) -> list[dict]:
@@ -140,26 +169,25 @@ def render_pages(server_url: str, pages: list[str], preview_dir: Path) -> list[d
records: list[dict] = []
inject_js = """
async (pageName) => {
const res = await fetch('/api/slide/' + encodeURIComponent(pageName) + '?_=' + Date.now());
if (!res.ok) throw new Error('fetch /api/slide/' + pageName + ' returned ' + res.status);
const data = await res.json();
({svgContent, width, height}) => {
document.documentElement.innerHTML =
'<head><style>html,body{margin:0;padding:0;background:#0E1116;overflow:hidden}'
+ ' svg{display:block;width:1280px;height:720px}</style></head>'
+ '<body>' + data.content + '</body>';
return { len: data.content.length };
+ ' svg{display:block;width:' + width + 'px;height:' + height + 'px}</style></head>'
+ '<body>' + svgContent + '</body>';
return { len: svgContent.length };
}
"""
with sync_playwright() as p:
browser = p.chromium.launch()
try:
context = browser.new_context(viewport={'width': 1280, 'height': 720})
context = browser.new_context()
for page_name in pages:
rec: dict = {'page': page_name, 'ok': False}
try:
fetch_slide_text(server_url, page_name)
svg_content = fetch_slide_content(server_url, page_name)
canvas = parse_slide_canvas(svg_content, page_name)
rec['canvas'] = canvas
except urllib.error.URLError as e:
rec['error'] = f'server_unreachable: {e!r}'
records.append(rec)
@@ -172,22 +200,36 @@ async (pageName) => {
stem = page_name[:-4] if page_name.endswith('.svg') else page_name
out_path = preview_dir / f'{stem}.png'
pg = None
try:
pg = context.new_page()
pg.set_viewport_size({
'width': canvas['png_width'],
'height': canvas['png_height'],
})
pg.goto(server_url, wait_until='domcontentloaded')
pg.evaluate(inject_js, page_name)
pg.evaluate(inject_js, {
'svgContent': svg_content,
'width': canvas['width'],
'height': canvas['height'],
})
# Wait one frame so font/text shaping settles before capture.
pg.wait_for_timeout(100)
png_bytes = pg.screenshot(type='png', full_page=False)
pg.close()
out_path.write_bytes(png_bytes)
rec['ok'] = True
rec['path'] = str(out_path)
rec['bytes'] = len(png_bytes)
rec['all_background'] = is_all_background(png_bytes)
rec['ok'] = True
except Exception as e: # noqa: BLE001 — best-effort per-page
rec['error'] = f'{type(e).__name__}: {e}'
finally:
if pg is not None:
try:
pg.close()
except Exception: # noqa: BLE001 — cleanup is best-effort
pass
records.append(rec)
finally:
browser.close()
@@ -351,6 +351,12 @@ grammar 和构形思维,而不是页面示例清单。先还原顺序、层级
两条路径都不依赖 `structure/<key>` 或固定 SVG,不能因 Quick 跳过 §VII/lock 就跳过
Structure 载体判断。
**已批准的图表容量迁移**
| Canonical key | 已批准边界 | 兼容处理 |
|---|---|---|
| `gauge_chart` | 2026-08-10:中性预览从三个并列 Gauge 重组为一个有界域、明确目标或阈值的 KPI;多个同级 KPI 改用 `bullet_chart``progress_bar_chart` | canonical key 保持不变;旧三指标示例不再作为可读取容量合同,无需 alias |
**Canonical Table set**:
| Canonical key | 核心信息关系 |
@@ -388,7 +394,7 @@ Table;日期或持续时间决定 `x`/`width` 的排期是 `chart/gantt_chart`
**Hard rule**: 修改一个仍属 canonical 的模板时先冻结可见文本、数据和结构层级,
再简化确认无语义的效果、补齐直属根 bounds,并完成文本差异、独立渲染与双路线
验证。经本节明确登记的 catalog 合并/退役不要求保留旧示例文案;除此之外,未经
验证。经本节明确登记的 catalog 合并/重组/退役不要求保留旧示例文案;除此之外,未经
说明的文本删除、改写或结构边界丢失都会阻断变更。
---
@@ -8,7 +8,7 @@
"filePattern": "{key}.svg",
"libraryPositioning": "Data-driven visualization templates whose source values determine geometry or visual encoding.",
"summaryGrammar": "Each chart summary is a selection rule, not a description. Format: 'Pick for <content shape + scale>. Skip if <reason → alternative>'.",
"updated": "2026-08-08"
"updated": "2026-08-10"
},
"aliases": {
"project_schedule_table": "gantt_chart"
@@ -51,7 +51,7 @@
"summary": "Pick for project schedule with 6-12 tasks whose dates or durations determine bar position and length, with optional dependencies. Skip for milestones without duration or qualitative lane-stage handoffs (build a Structure from the page relationships)."
},
"gauge_chart": {
"summary": "Pick for one hero metric or single KPI's goal achievement rate. Skip for multiple metrics (use bullet_chart for target+actual or progress_bar_chart for completion %)."
"summary": "Pick for one bounded metric against an explicit target or threshold on a known domain. Skip for an unbounded hero number (build a KPI-led Structure) or multiple metrics (use bullet_chart for target+actual or progress_bar_chart for completion %)."
},
"grouped_bar_chart": {
"summary": "Pick for 2-4 series side-by-side across the same categories (e.g. YoY/QoQ). Skip if showing composition within each category (use stacked_bar_chart)."
@@ -75,7 +75,7 @@
"summary": "Pick for showing 80/20 contribution with descending bars and a cumulative line. Skip if the cumulative line is not the message (use column_chart)."
},
"pie_chart": {
"summary": "Pick for simple 3-6 part proportions of one whole. Skip for >=7 parts (use donut_chart or treemap_chart) or hierarchical composition (use treemap_chart)."
"summary": "Pick for simple 3-6 part proportions of one whole. Skip for >=7 flat parts (use horizontal_bar_chart) or hierarchical composition (use treemap_chart)."
},
"pie_of_pie_chart": {
"summary": "Pick for pie composition where small slices need a secondary pie detail. Skip if the secondary detail is easier to compare as bars (use bar_of_pie_chart) or hierarchy matters (use sunburst_chart)."
@@ -3,140 +3,59 @@
font-size="12">
<!--
Gauge Chart Template
Usage: Single metric attainment display
Usage: One bounded metric against one explicit goal or threshold
Scenarios: KPI completion, health index, performance score
Supports: 3 gauges side by side
Skip: Multiple peer metrics; use bullet_chart or progress_bar_chart
-->
<!-- Background -->
<rect width="1280" height="720" fill="#F8FAFC"/>
<!-- Title -->
<g id="header" data-pptx-bounds="40 15 1200 125">
<text x="640" y="60"
font-size="32" font-weight="700" fill="#0F172A" text-anchor="middle">Core Business KPI Dashboard</text>
font-size="32" font-weight="700" fill="#0F172A" text-anchor="middle">Sales Target Attainment</text>
<text x="640" y="95"
font-size="18" fill="#64748B" text-anchor="middle">Q4 Performance Real-time Data</text>
font-size="18" fill="#64748B" text-anchor="middle">Q4 2025 · Actual 85% against an 80% target</text>
</g>
<!-- Gauge 1: Sales Completion 85% -->
<g id="gauge1" data-pptx-bounds="10 191 427 351" transform="translate(220, 380)">
<!-- Background Arc -->
<path d="M -150,0 A 150,150 0 0,1 150,0" fill="none" stroke="#E2E8F0" stroke-width="30" stroke-linecap="round"/>
<!-- One hero gauge: Sales target attainment 85% -->
<g id="heroGauge" data-pptx-bounds="395 125 490 440" transform="translate(640, 380)">
<g id="gaugeScale" transform="scale(1.35)">
<!-- Scales (0-60% Red, 60-80% Amber, 80-100% Emerald) -->
<path d="M -150,0 A 150,150 0 0,1 46.35,-142.66" fill="none" stroke="#FECACA" stroke-width="30" stroke-linecap="round"/>
<path d="M 46.35,-142.66 A 150,150 0 0,1 121.35,-88.17" fill="none" stroke="#FDE68A" stroke-width="30" stroke-linecap="round"/>
<path d="M 121.35,-88.17 A 150,150 0 0,1 150,0" fill="none" stroke="#A7F3D0" stroke-width="30" stroke-linecap="round"/>
<!-- Value Arc 85% -->
<path d="M -150,0 A 150,150 0 0,1 133.65,-68.10" fill="none" stroke="#059669" stroke-width="30" stroke-linecap="round"/>
<!-- Tick Marks -->
<line x1="-135" y1="0" x2="-120" y2="0" stroke="#64748B" stroke-width="2"/>
<line x1="-95.5" y1="-95.5" x2="-84.85" y2="-84.85" stroke="#64748B" stroke-width="2"/>
<line x1="0" y1="-135" x2="0" y2="-120" stroke="#64748B" stroke-width="2"/>
<line x1="95.5" y1="-95.5" x2="84.85" y2="-84.85" stroke="#64748B" stroke-width="2"/>
<line x1="135" y1="0" x2="120" y2="0" stroke="#64748B" stroke-width="2"/>
</g>
<!-- Tick Labels -->
<text x="-155" y="25"
<text x="-215" y="30"
font-size="14" fill="#64748B" text-anchor="middle">0%</text>
<text x="0" y="-155"
<text x="0" y="-230"
font-size="14" fill="#64748B" text-anchor="middle">50%</text>
<text x="155" y="25"
<text x="215" y="30"
font-size="14" fill="#64748B" text-anchor="middle">100%</text>
<!-- Needle (85% = 153°, start from -90°) -->
<g transform="rotate(63)">
<path d="M 0,-100 L -8,0 L 0,15 L 8,0 Z" fill="#0F172A"/>
<g id="gaugeNeedle" transform="scale(1.35)">
<g transform="rotate(63)">
<path d="M 0,-100 L -8,0 L 0,15 L 8,0 Z" fill="#0F172A"/>
</g>
<circle cx="0" cy="0" r="15" fill="#0F172A"/>
<circle cx="0" cy="0" r="8" fill="#FFFFFF"/>
</g>
<circle cx="0" cy="0" r="15" fill="#0F172A"/>
<circle cx="0" cy="0" r="8" fill="#FFFFFF"/>
<!-- Center Value -->
<text x="0" y="60"
font-size="48" font-weight="700" fill="#10B981" text-anchor="middle">85%</text>
<!-- Title -->
<text x="0" y="100"
font-size="18" font-weight="600" fill="#0F172A" text-anchor="middle">Sales Target Attainment</text>
<text x="0" y="80"
font-size="64" font-weight="700" fill="#10B981" text-anchor="middle">85%</text>
<text x="0" y="125"
font-size="24" font-weight="600" fill="#0F172A" text-anchor="middle">Current achievement</text>
<!-- Status Tag -->
<rect x="-40" y="115" width="80" height="26" rx="13" fill="#D1FAE5"/>
<text x="0" y="133"
font-size="13" font-weight="600" fill="#10B981" text-anchor="middle">Excellent</text>
<rect x="-68" y="142" width="136" height="35" rx="18" fill="#D1FAE5"/>
<text x="0" y="166"
font-size="17" font-weight="600" fill="#047857" text-anchor="middle">Above target</text>
</g>
<!-- Gauge 2: Customer Satisfaction 72% -->
<g id="gauge2" data-pptx-bounds="430 191 427 351" transform="translate(640, 380)">
<!-- Background Arc -->
<path d="M -150,0 A 150,150 0 0,1 150,0" fill="none" stroke="#E2E8F0" stroke-width="30" stroke-linecap="round"/>
<!-- Scales -->
<path d="M -150,0 A 150,150 0 0,1 46.35,-142.66" fill="none" stroke="#FECACA" stroke-width="30" stroke-linecap="round"/>
<path d="M 46.35,-142.66 A 150,150 0 0,1 121.35,-88.17" fill="none" stroke="#FDE68A" stroke-width="30" stroke-linecap="round"/>
<path d="M 121.35,-88.17 A 150,150 0 0,1 150,0" fill="none" stroke="#A7F3D0" stroke-width="30" stroke-linecap="round"/>
<!-- Value Arc 72% -->
<path d="M -150,0 A 150,150 0 0,1 95.61,-115.58" fill="none" stroke="#2563EB" stroke-width="30" stroke-linecap="round"/>
<!-- Tick Marks -->
<line x1="-135" y1="0" x2="-120" y2="0" stroke="#64748B" stroke-width="2"/>
<line x1="-95.5" y1="-95.5" x2="-84.85" y2="-84.85" stroke="#64748B" stroke-width="2"/>
<line x1="0" y1="-135" x2="0" y2="-120" stroke="#64748B" stroke-width="2"/>
<line x1="95.5" y1="-95.5" x2="84.85" y2="-84.85" stroke="#64748B" stroke-width="2"/>
<line x1="135" y1="0" x2="120" y2="0" stroke="#64748B" stroke-width="2"/>
<!-- Tick Labels -->
<text x="-155" y="25"
font-size="14" fill="#64748B" text-anchor="middle">0%</text>
<text x="0" y="-155"
font-size="14" fill="#64748B" text-anchor="middle">50%</text>
<text x="155" y="25"
font-size="14" fill="#64748B" text-anchor="middle">100%</text>
<!-- Needle -->
<g transform="rotate(39.6)">
<path d="M 0,-100 L -8,0 L 0,15 L 8,0 Z" fill="#0F172A"/>
</g>
<circle cx="0" cy="0" r="15" fill="#0F172A"/>
<circle cx="0" cy="0" r="8" fill="#FFFFFF"/>
<!-- Center Value -->
<text x="0" y="60"
font-size="48" font-weight="700" fill="#3B82F6" text-anchor="middle">72%</text>
<!-- Title -->
<text x="0" y="100"
font-size="18" font-weight="600" fill="#0F172A" text-anchor="middle">Customer Satisfaction</text>
<!-- Status Tag -->
<rect x="-35" y="115" width="70" height="26" rx="13" fill="#DBEAFE"/>
<text x="0" y="133"
font-size="13" font-weight="600" fill="#3B82F6" text-anchor="middle">Good</text>
</g>
<!-- Gauge 3: Inventory Turnover 58% -->
<g id="gauge3" data-pptx-bounds="850 191 427 351" transform="translate(1060, 380)">
<!-- Background Arc -->
<path d="M -150,0 A 150,150 0 0,1 150,0" fill="none" stroke="#E2E8F0" stroke-width="30" stroke-linecap="round"/>
<!-- Scales -->
<path d="M -150,0 A 150,150 0 0,1 46.35,-142.66" fill="none" stroke="#FECACA" stroke-width="30" stroke-linecap="round"/>
<path d="M 46.35,-142.66 A 150,150 0 0,1 121.35,-88.17" fill="none" stroke="#FDE68A" stroke-width="30" stroke-linecap="round"/>
<path d="M 121.35,-88.17 A 150,150 0 0,1 150,0" fill="none" stroke="#A7F3D0" stroke-width="30" stroke-linecap="round"/>
<!-- Value Arc 58% -->
<path d="M -150,0 A 150,150 0 0,1 37.30,-145.29" fill="none" stroke="#D97706" stroke-width="30" stroke-linecap="round"/>
<!-- Tick Marks -->
<line x1="-135" y1="0" x2="-120" y2="0" stroke="#64748B" stroke-width="2"/>
<line x1="-95.5" y1="-95.5" x2="-84.85" y2="-84.85" stroke="#64748B" stroke-width="2"/>
<line x1="0" y1="-135" x2="0" y2="-120" stroke="#64748B" stroke-width="2"/>
<line x1="95.5" y1="-95.5" x2="84.85" y2="-84.85" stroke="#64748B" stroke-width="2"/>
<line x1="135" y1="0" x2="120" y2="0" stroke="#64748B" stroke-width="2"/>
<!-- Tick Labels -->
<text x="-155" y="25"
font-size="14" fill="#64748B" text-anchor="middle">0%</text>
<text x="0" y="-155"
font-size="14" fill="#64748B" text-anchor="middle">50%</text>
<text x="155" y="25"
font-size="14" fill="#64748B" text-anchor="middle">100%</text>
<!-- Needle -->
<g transform="rotate(14.4)">
<path d="M 0,-100 L -8,0 L 0,15 L 8,0 Z" fill="#0F172A"/>
</g>
<circle cx="0" cy="0" r="15" fill="#0F172A"/>
<circle cx="0" cy="0" r="8" fill="#FFFFFF"/>
<!-- Center Value -->
<text x="0" y="60"
font-size="48" font-weight="700" fill="#F59E0B" text-anchor="middle">58%</text>
<!-- Title -->
<text x="0" y="100"
font-size="18" font-weight="600" fill="#0F172A" text-anchor="middle">Inventory Turnover</text>
<!-- Status Tag -->
<rect x="-40" y="115" width="80" height="26" rx="13" fill="#FEF3C7"/>
<text x="0" y="133"
font-size="13" font-weight="600" fill="#F59E0B" text-anchor="middle">Warning</text>
</g>
<!-- Legend -->
<!-- Threshold legend -->
<g id="legendBar" data-pptx-bounds="265 593 750 54" transform="translate(340, 600)">
<rect x="0" y="0" width="600" height="40" rx="8" fill="#FFFFFF" stroke="#E2E8F0"/>
<rect x="30" y="12" width="16" height="16" rx="3" fill="#FECACA"/>
@@ -154,6 +73,6 @@
<!-- Data Source -->
<g id="source" data-pptx-bounds="40 670 1200 40">
<text x="640" y="695"
font-size="14" fill="#94A3B8" text-anchor="middle">Source: Business Intelligence Platform · Live Data</text>
font-size="14" fill="#94A3B8" text-anchor="middle">Source: Business Intelligence Platform · Q4 2025</text>
</g>
</svg>

Before

Width:  |  Height:  |  Size: 9.2 KiB

After

Width:  |  Height:  |  Size: 4.3 KiB

@@ -216,6 +216,10 @@ When Speaker Notes is disabled, keep §X with only
`- **Generation**: disabled`; do not write filename, duration, style, or purpose
placeholders. An explicit notes-off/audio-on conflict blocks before authoring.
When an explicit final/literal narration script will become notes or generated
audio, make §X `Content` name that source and say `preserve verbatim`; keep the
full segmented script in `notes/total.md`, not in §IX or this Design Spec.
Append either or both optional lines only when the capability earns a place;
never write an empty or `none` placeholder:
@@ -129,8 +129,8 @@ Catalog-based custom example:
```markdown
## mode
- mode: custom
- mode_references: pyramid, narrative
- mode_behavior: Lead each act with the decision-first clarity of pyramid, then develop it through a narrative tension-and-resolution arc.
- mode_references: pyramid, narrative, instructional
- mode_behavior: Open conclusion-first with pyramid, develop the risk through a narrative tension-and-resolution act, then close with an instructional action sequence.
```
---
@@ -18,6 +18,12 @@ Create one reusable template workspace under either the **global template librar
> **Boundary against template-fill and in-place structure edits**: Create Template does not fill content into a PPTX, add Master/Layout structure to an existing PPTX/SVG, or directly output the user's final generated deck. It authors a separate reusable workspace; an optional PPTX is review evidence only. To generate a deck, return the workspace root as an exact candidate to [`generate-pptx`](./generate-pptx.md) Step 3, confirm it with Stage 1, then author new SVG pages from the installed state. A project-scoped workspace selected for its own project is consumed in place after that confirmation.
> **Boundary against page-image reconstruction**: screenshots and page visuals
> in this route are evidence for reusable rules and prototypes. When the user
> instead wants each supplied page image reconstructed into one final editable
> slide, use the Codex-supported [`image-to-pptx.md`](./profiles/image-to-pptx.md);
> do not turn a final-deck request into a template workspace.
## Child Workflow Dispatch
Create Template is the fixed user-facing entry and common contract. It selects one child workflow, then that child owns the kind-specific lifecycle. Do not execute two children for one workspace or blend their schemas.
@@ -254,6 +254,7 @@ Then load only the extra role modules triggered by the current plan:
| Deterministic trigger | Additional Strategist reference |
|---|---|
| Stage 1 is confirmed and its template choice installed a selected Brand/Style/Layout/Deck workspace into this project | `references/strategist-template.md` before Stage 2 |
| The confirmed Stage-1 `delivery_context` identifies recorded/self-running/video delivery, or input is an explicit final/literal narration script | `references/video-design.md` before the three Stage-2 whole solutions and page roster |
| The confirmed Stage 2 `image_usage` contains a source other than `none`, the user supplied an explicit non-`none` image constraint, or formula-worthy content activates formula planning | `references/image-layout-spec.md` + `references/image-layout-patterns.md` before production detail, formula resources, or §VIII |
After Stage 1 and template handoff, load `strategist-image.md` plus only the
@@ -448,6 +449,7 @@ python3 ${SKILL_DIR}/scripts/analyze_images.py <project_path>/images
**Output**:
- `<project_path>/design_spec.md` — complete human-readable design narrative and durable confirmed production state
- `<project_path>/spec_lock.md` — machine-readable stable execution anchors/routing, authored after conditional review approval
- `<project_path>/notes/total.md` — only when the prepared final narration branch is active; frozen verbatim production input
For a new project, use the reference-first whole-document sequence:
@@ -459,6 +461,13 @@ For a new project, use the reference-first whole-document sequence:
Final state → initial Design Spec mismatch, approved Design Spec/context → lock mismatch, or an unapplied revision blocks despite schema validity. `validate` does not prove fidelity. Repair from retained confirmation before refinement; during it, preserve unaffected values and apply explicit revisions. After approval, derive the lock from that Design Spec/context. Resume/refine edits existing files, never scaffolds. Fresh recovery alone may reread persisted final evidence once.
**Prepared final narration branch**: follow `video-design.md` §1 and §3 when an
explicit final/literal script will become notes or generated audio. Segment it
by semantic scene during Stage 2; §IX gives each segment a supporting visible
state and §X records its source/verbatim policy. After Gate 2, before Step 5 or
split handoff, write the exact segments once to `notes/total.md`; split them only
in Step 7.1. This is frozen production input, not a third planning artifact.
**✅ Internal checkpoint — Phase deliverables complete**: facts read; confirmation consumed once; final Stage-2 production fields resolved (formula policy, generation mode, refine-spec, proactive choices, and conditional AI path); Design Spec passed Gate 1; enabled refinement approved; lock derived from it; split handling resolved; communication and every §IX `Audience move` validated. Do not print this checklist; auto-proceed.
---
@@ -488,7 +497,7 @@ Then **lazy-load the path-specific reference** for each row that actually needs
A deck with only `ai` rows never loads `image-searcher.md`; a deck with only `web` rows never loads `image-generator.md`. A mixed deck loads both, processes each row through its own path, and writes both `image_prompts.json` and `image_sources.json`.
> ⚠️ **In-pipeline ai rows MUST use the manifest contract** — even when only 1 ai row exists. Always write `images/image_prompts.json` first and render `image_prompts.md` with `image_gen.py --render-md`. Then execute the confirmed path from `image-generator.md §7`: `image_gen.py --manifest` is **Path A only**; `host-native` is **Path B** and MUST skip `--manifest`; `manual` writes the prompts and stops for external generation. The positional form (`image_gen.py "prompt" ...`) is reserved for **out-of-pipeline one-off testing / single-image fixups**, except for the already-planned registered subject/base derivation in `image-generator.md` §4.4. That narrow exception keeps both final rows in the resource authority and operational sidecar; it does not authorize unrelated in-pipeline generation outside the manifest contract.
> ⚠️ **In-pipeline ai rows MUST use the manifest contract** — even when only 1 ai row exists. Always write `images/image_prompts.json` first and render `image_prompts.md` with `image_gen.py --render-md`. Then execute the confirmed path from `image-generator.md §7`: `image_gen.py --manifest` is **Path A only**; `host-native` is **Path B** and MUST skip `--manifest`; `manual` writes the prompts and stops for external generation. The positional form (`image_gen.py "prompt" ...`) is reserved for **out-of-pipeline one-off testing / single-image fixups**, except for the already-planned registered reconstruction-group derivation in `image-generator.md` §4.4. That narrow exception keeps every final member in the resource authority and operational sidecar; it does not authorize unrelated in-pipeline generation outside the manifest contract.
> ⚠️ **web path — batch multiple rows**: when ≥2 rows are `Acquire Via: web`, write all queries into `images/image_queries.json` and run `image_search.py --batch` once (concurrent acquisition, status written back), instead of one CLI call per row. A single web row may use the positional single-query form. See [image-searcher.md](../references/image-searcher.md) §5.
@@ -529,8 +538,14 @@ Workflow:
**Page content**: §IX is preferred wording and semantic authority. Use it when it works; adapt it when presentation benefits while preserving intent, facts, and explicit literal requirements. Read sources only to verify requested evidence; return incomplete blocks to Step 4 instead of enriching them during execution.
**Prepared final narration**: when §X records a literal script, read the frozen
`notes/total.md` once before P01 and design each visible state/semantic group
around its exact segment; never edit or pad it.
**Planning context**: follow [`executor-base.md`](../references/executor-base.md) §2.1. Reuse the complete Design Spec and lock in an unchanged, uncompacted context. Fresh/resumed/restarted, compacted/summary-only, or externally/unknown changed execution reads both once and reloads triggered inputs. For a local question, consult the retained lock first, then only the owning Design Spec fragment; do not poll files merely to prove validity.
**Scheduled lock re-read (Default Generate only)**: when another page follows, re-read `spec_lock.md` once after P05/P10/P15/… per [`executor-base.md`](../references/executor-base.md) §2.1.
**Artifact ownership**: `svg_output/` is the author source, `svg_final/` is derived, and image facts come from the regenerated `analysis/image_analysis.csv`; see [`references/artifact-ownership.md`](../references/artifact-ownership.md).
Read the execution references for this deck's locked `mode` + `visual_style`
@@ -561,6 +576,7 @@ Keep the core's shared visual-quality defaults and `svg-effects.md` §6.1 Visual
| Used preset pattern fill, or independent Chart/Table with §IX `<object-key>=yes` | `native-data-interface.md` before that object |
| `spec_lock.md images` / §VIII has an image/formula row, or the template has bundled images | `executor-image.md` + `image-layout-spec.md` + `image-layout-patterns.md` + `svg-image-embedding.md` |
| At least one placed image has `Status: Sourced` | `executor-web-image.md` after the image branch |
| §I records recorded/self-running/video delivery, or §X records a final/literal narration script | `video-design.md` before the first SVG; retain it through notes/motion handling |
| All SVG pages and SVG quality gates are complete, and the effective Speaker Notes outcome in `design_spec.md §I` is enabled | `executor-notes.md` before generating speaker notes |
No branch is loaded by analogy. For each page, after §IX content/communication
@@ -661,10 +677,13 @@ python3 ${SKILL_DIR}/scripts/svg_quality_checker.py <project_path> --stage final
**Logic Construction Phase (conditional)**: after the SVG quality gate passes,
when the effective Speaker Notes outcome in `design_spec.md §I` is enabled, load
[`executor-notes.md`](../references/executor-notes.md), ground each page's
narration in all information-bearing content in its final SVG, and generate
speaker notes → `<project_path>/notes/total.md`. When the outcome is `disabled`,
do not load the notes branch and do not require or create `notes/total.md`.
[`executor-notes.md`](../references/executor-notes.md). When the prepared final
narration branch already created `notes/total.md`, validate its exact segments
against every information-bearing final SVG group and repair the visual page or
upstream plan on mismatch; never rewrite the script. Otherwise ground each
page's narration in its final SVG and generate complete speaker notes →
`<project_path>/notes/total.md`. When the outcome is `disabled`, do not load the
notes branch and do not require or create `notes/total.md`.
**✅ Internal checkpoint — execution complete**: verify live preview timing,
the P01 method gate, uninterrupted remaining-page generation, consolidated
@@ -19,6 +19,7 @@ Maintainer-only inventory for adding, moving, or removing workflow documents. Ru
| ID | Class | Path | Parent / lifecycle slot |
|---|---|---|---|
| `image-to-pptx` | Generation profile | [`profiles/image-to-pptx.md`](./profiles/image-to-pptx.md) | Codex-supported, Quick-only page-frame normalization plus layered reconstruction |
| `beautify-pptx` | Generation profile | [`profiles/beautify-pptx.md`](./profiles/beautify-pptx.md) | Generate PPTX |
| `quick-generate` | Generation profile | [`profiles/quick-generate.md`](./profiles/quick-generate.md) | Generate PPTX direct SVG-to-PPTX short circuit |
| `apply-template-workspace` | Template-input stage | [`stages/apply-template-workspace.md`](./stages/apply-template-workspace.md) | Non-free Default Stage-1 selection after confirmation, or Quick direct exact-root input |
@@ -50,7 +50,7 @@ source.pptx
|---|---:|---|
| `narration.notes` | Enabled | Add or replace speaker notes generated from slide content |
| `narration.audio` | Enabled | Embed one audio file per slide |
| `narration.timings` | Enabled | Set narrated slides to auto-advance by audio duration |
| `narration.timings` | Enabled | Set narrated slides to auto-advance from page-start lead-in, audio duration, and page-tail padding |
| `narration.transitions` | Enabled | Add page-level transitions for narrated/selected slides |
| `delivery.check` | Enabled | Read-only package/font/media/hidden-slide/file-size and existing-motion audit |
| `media` | Planned | Background music, video, media compression |
@@ -137,7 +137,7 @@ Present the plan to the user before generating notes or audio:
|---|---|---|
| `notes` | Enabled; required whenever audio is enabled | Add/replace speaker notes generated from slide content? |
| `audio` | Enabled when user wants narration/video/autoplay | After notes are complete, generate one narration audio file per slide? |
| `timings` | Enabled with audio | Set slide auto-advance from audio duration? |
| `timings` | Enabled with audio; 0.8s page-start floor and 0.4s page-tail padding | Set slide auto-advance from narration timing? |
| `transitions` | Enabled, `fade` 0.5s | Add page transitions? Which canonical native effect, Effect Options, and duration? |
| `delivery.check` | Always on, read-only | No confirmation required; review errors and advisories |
@@ -156,8 +156,20 @@ the notes artifact.
| Transitions enabled with an effect | Replace with that exact effect and duration | Preserve unless timings is enabled |
| Transitions disabled with a non-`none` configured effect | Preserve the source effect, including unknown `AlternateContent` | Preserve unless timings is enabled |
| Explicit `none` | Remove the visual effect | Preserve, or write timing-only advance when timings is enabled |
| Timings enabled with audio | Keep the resolved enter policy | Use audio duration plus narration padding; click disabled |
| Timings disabled | Apply the confirmed enter policy only | Audio readiness may probe decodability; do not use duration or add/change `advTm` or `useTimings` |
| Timings enabled with audio | Keep the resolved enter policy | Use page-start lead-in plus audio duration plus page-tail padding; click disabled |
| Timings disabled | Apply the confirmed enter policy only | Use duration only to reject source auto-advance that would truncate narration; do not derive, add, or change `advTm` / `useTimings` |
The timing module's `narration_start_floor` and `narration_padding` are
independent optional values. When omitted, use `0.8` and `0.4` seconds
respectively. For a destination-page transition of `T` seconds, the
post-transition lead-in is
`max(0, narration_start_floor - T)`: narration does not begin during the
transition, and the transition duration itself remains unchanged. A start
floor of `0` means start immediately after the transition completes. A
preserved legacy transition that exposes only `spd` has no exact millisecond
duration; keep the full configured floor after that transition instead of
guessing a PowerPoint-specific duration. When timings are disabled, preserve
source `advTm`, but fail if it would advance before the delayed narration ends.
The confirmed `modules.transitions` object may include `effect_options` beside
an explicit canonical `effect`. Use
@@ -314,6 +326,7 @@ Optional:
python3 skills/ppt-master/scripts/native_enhance_pptx.py apply "<project>" \
--transition fade \
--transition-duration 0.5 \
--narration-start-floor 0.8 \
--narration-padding 0.4 \
--apply-transition-without-audio \
--overwrite
@@ -376,7 +389,7 @@ Check:
| Visible content | No intentional changes |
| Notes | Present on intended slides |
| Audio media | Present under `ppt/media/` when generated |
| Auto-play | Narrated slides advance by audio duration |
| Auto-play | Narrated slides wait for the resolved page-start lead-in, then advance after audio duration plus page-tail padding |
| Transition | Requested effect remains exact; preserved `AlternateContent` keeps its primary and fallback branches |
| Timings disabled | Source `advTm` and package `useTimings` are not changed |
| Delivery check | No newly introduced structural errors; source baseline and font/media/hidden-slide advisories reviewed |
@@ -32,6 +32,13 @@ Beautify constraints in this file apply in either runtime.
**Distinct from mirror templates**: `replication_mode: mirror` ([`executor-structured.md`](../../references/executor-structured.md) §1.1) keeps layout + visuals verbatim and edits text. Beautify is the inverse — content verbatim, layout redone, identity inherited.
**Distinct from page-image reconstruction**: when the authoritative input is
an ordered raster page roster and the user wants its visible layout preserved,
activate the Codex-supported, Quick-only
[`image-to-pptx.md`](./image-to-pptx.md) instead.
Beautify requires a semantic source PPTX and deliberately redesigns layout; the
two fidelity profiles never compose.
**When this profile is wrong — re-architecture belongs to ordinary Generate**: this profile preserves the source's page count and page order 1:1. It is for "keep this deck, just lay it out better". When the user instead wants the original page breakdown reconsidered — merge / split / reorder pages, re-outline the structure, build a *better deck* from the same content rather than a prettier version of the same pages — do not activate this profile. This includes re-pagination for fit: "keep every word but split a crowded page so it reads better" changes page count. Convert the deck with [`ppt_to_md`](../../scripts/source_to_md/ppt_to_md.py) and use ordinary Quick when Quick was explicit, otherwise the Default main pipeline. The deciding question: is the source's page split information to preserve, or just the previous author's structure to improve? Preserve → activate this profile; improve → ordinary Generate in the selected runtime.
---
@@ -97,13 +104,6 @@ python3 ${SKILL_DIR}/scripts/pptx_intake.py <project_path>/sources/<source.pptx>
> Note: `theme` is what the deck declares; `observed` is a frequency sample of run-level overrides (not a complete style resolution — it misses `schemeClr` and master/layout inheritance, and counts chart/gradient fills). A hand-edited deck can diverge from `theme` — Step 5 recommends which to inherit and the user confirms.
**Chart + table data (for regeneration)**: read `<project_path>/analysis/<stem>.slide_library.json`. It contains the source chart and table *data* so they can be redrawn natively in the inherited style:
| `<stem>.slide_library.json` field | Use |
|---|---|
| `slides[].charts[]` (`chart_type` / `categories` / `series[].values`) | regenerate as a native SVG chart; use the §VII catalog key only when recall selects a real reference, otherwise plan the custom chart in §IX |
| `slides[].tables[]` (`row_count` / `column_count` / cell text) | regenerate as a native SVG table |
**Hard rule — regenerate visuals, do not carry them over**: charts / tables / images are rebuilt from their data in the inherited style, never spliced in byte-for-byte. This keeps the deck style-consistent and natively editable. **Data values are frozen** (categories / series / cell text / numbers unchanged); only their rendering is the deck's own. Pictures (`ppt_to_md`-extracted files) are reused but re-laid-out — position / crop / size follow the new layout, not the source slot. A user who wants an original element verbatim copies it across themselves.
**Optional source-SVG visual reference**: when the source deck has complex vector decoration, distinctive page chrome, or a visual language that cannot be captured by `<stem>.identity.json` colors/fonts alone, create a read-only SVG reference package under `analysis/`. This is for understanding style only; it is not a carry-over asset path.
@@ -120,7 +120,13 @@ Use the cleaned `analysis/source_svg_import/svg-flat/slide_*.svg` files plus `an
Default: do **not** copy these candidates into the project `icons/`, do **not** list them as reusable output assets, and do **not** preserve original vector decorations byte-for-byte in the beautified deck. The Executor still regenerates fresh native shapes from the confirmed plan.
Optional reuse gate: if a candidate is a non-text brand/logo/motif/decorative asset that should survive the beautification, list it in the Step 5 plan with source slide, candidate filename, intended reuse, and dependency notes from the inventory. Wait for user confirmation. Only confirmed candidates may be promoted into `<project_path>/icons/imported/` and referenced from generated SVGs with `<use data-icon="imported/<name>"/>`; `finalize_svg.py` then re-inlines them as native shapes. Never promote text-bearing groups, charts/tables, source page layouts, or dense slide composites as reusable assets.
**Optional reuse gate**: retain source slide, filename, use, and dependencies
for a non-text brand/logo/motif/decorative candidate. Default lists it in Step 5
and waits; only confirmed candidates are promoted. Quick's current main agent
decides directly and stops only when frozen facts lack a lossless preservation
path. Promote to `<project_path>/icons/imported/` and reference with
`<use data-icon="imported/<name>"/>`; Quick never runs `finalize_svg.py`. Never
promote text-bearing groups, charts/tables, page layouts, or dense composites.
**Assemble the inventory** — the deterministic join into one per-slide ledger, `analysis/beautify_inventory.json`, the contract Step 5 confirms and Step 7 verifies against:
@@ -136,6 +142,21 @@ If `images/image_manifest.json` does not exist because the source deck has no ex
| `ignored` | hidden slides / shapes, master-only text, image crop / opacity / rotation / mask (not captured upstream) |
| `needs_confirmation` | unreadable SmartArt data; combo / dual-axis / waterfall charts; merged-cell or multi-header tables; density-outlier pages — **either** overcrowded **or** near-empty / title-only |
**Mandatory — bounded inventory reads**: the complete inventory is the Step 7
validation ledger, not the default authoring prompt. Read its compact roster,
then the current page; add geometry only for structural ambiguity:
```bash
python3 ${SKILL_DIR}/scripts/beautify_inventory.py \
<project_path>/analysis/beautify_inventory.json --summary
python3 ${SKILL_DIR}/scripts/beautify_inventory.py \
<project_path>/analysis/beautify_inventory.json --page <N>
python3 ${SKILL_DIR}/scripts/beautify_inventory.py \
<project_path>/analysis/beautify_inventory.json --page <N> --with-geometry
```
During authoring, do not bulk-read either complete file.
**SmartArt output boundary**: Preserve its extracted wording and semantic relationships, then redraw it through SVG as ordinary editable PowerPoint shapes. Do not attempt to regenerate a native SmartArt object or reuse persisted-drawing text as a second content source.
```markdown
@@ -162,8 +183,22 @@ active context. Explicit user requirements remain authoritative; otherwise use
the source identity as the default. Resolve `ignored` and `needs_confirmation`
without creating a confirmation payload, Design Spec, lock, or substitute
plan. If a flagged complex object cannot be regenerated without losing frozen
facts, stop as a hard prerequisite instead of simplifying it. Then continue to
the Quick branch in §6.
facts, stop as a hard prerequisite instead of simplifying it.
**Mandatory — close the transient Quick state before authoring**: before
entering §6 and [`quick-generate.md`](./quick-generate.md) §3, resolve every
row below in the active context:
| Transient state | Required closure |
|---|---|
| Roster and message | Exact source-order roster and one core message per page |
| Identity and type | Source identity, palette, fonts, body size, and type-role anchors |
| Page geometry | Per-page density, body frame, primary zone, and composition direction |
| Meaning and rhythm | Frozen relationships, reading path, neighbor/section rhythm, and ending |
| Resources and capabilities | Required local resources are usable; triggered notes, motion, audio, image, icon, formula, Chart/Table, and verification outcomes are decided |
Keep it transient: create no page/resource plan, Design Spec, lock,
confirmation payload, or substitute artifact. Then continue to §6 Quick.
### Default branch — Recommend & Confirm
@@ -265,6 +300,12 @@ its source order, hand-author every page, run the lockless Quick final checker,
and export with `--quick-generate`. Do not run Confirm UI, write a Design Spec
or lock, run the Default first-page gate, or call `finalize_svg.py`.
**Quick — lightweight long-deck review cadence (may adapt for a short deck or
semantic boundary)**: after about five pages or at a section
boundary, reread only the inventory summary/current-page views and cross-page
anchors. Do not run a checker; this is neither a gate nor an approval stop. Send
one `authored/total` status after each batch.
**Default**: run the standard pipeline as follows.
Run the standard pipeline ([`generate-pptx`](../generate-pptx.md) Steps 67). The Executor re-lays-out each page — hierarchy, spacing, alignment, page rhythm — using the semantic anchors in `spec_lock.md` plus current page/source/template context; valid page-local colors, gradients, effects, and export-safe display faces need not be added to the lock. It regenerates charts / tables as native SVG from the extracted data and re-lays-out the source pictures.
@@ -0,0 +1,337 @@
---
description: Quick-only Generate profile for reconstructing one or more source images into layered, editable PPTX slides.
---
# Image to PPTX Profile
> Quick-only Generate profile, not a top-level route. Normalize one or more
> supplied images into the represented page roster, then rebuild each page as
> native text, identity-faithful source graphics, and independently placeable
> image layers.
Reconstructs approved mockups, rendered slide images, contact sheets, or
flattened page visuals into a new PPTX. It does not redesign the deck and does
not run Strategist or confirmation. The visible result is the reference truth,
but the output is not a screenshot skin: image content is rebuilt into the
smallest useful background / foreground / subject stack, visible text is
restored as native text, and source graphics remain visually exact.
**Trigger**: the user supplies one or more raster visuals and explicitly asks
to restore the represented pages as a PPTX. Ordinary photos, illustrations, and
moodboards used only as resources for a new story do not activate this profile.
**Support boundary — Codex required**: this profile is currently documented
and validated for Codex because it depends on Codex's native reference-image
generation/editing capability and direct inspection of every derived layer.
Other hosts are not adapted or supported by this workflow; they may happen to
work, but the repository makes no compatibility claim and defines no alternate
host or generic image-backend fallback for this profile.
**Hard rule — Quick only**: this profile always loads
[`quick-generate.md`](./quick-generate.md) and never
[`generate-pptx.md`](../generate-pptx.md). The user does not need to say
"Quick" separately. Skip Strategist, Confirm UI, template selection,
`design_spec.md`, `spec_lock.md`, and the Default first-page gate. The current
main agent decides the reconstruction directly, prepares all resources, hand-
authors SVG pages, runs the lockless final checker, and exports.
**Hard rule — source surface, not template application**: stay in Quick free
design and do not install or apply a Brand/Style/Layout/Deck workspace for this
profile. A template would compete with the canonical page geometry and visual
identity. A reusable-system request routes to Create Template instead.
---
## 1. Routing and Output Boundary
| Request shape | Behavior |
|---|---|
| One or more image files represent pages that must become a final editable PPTX | Activate this profile and Quick Generate |
| One file contains several clearly separated slide frames | Split it into the ordered page roster first, then reconstruct each frame |
| Images are ordinary content assets or inspiration for a new deck | Use ordinary Generate |
| Images should define a reusable Brand/Style/Layout/Deck workspace | Use [`create-template.md`](../create-template.md) |
| A semantic source PPTX exists and only its layout should improve | Use [`beautify-pptx.md`](./beautify-pptx.md) |
**Hard rule — mutually exclusive fidelity profiles**: Image to PPTX and
Beautify never compose. Beautify preserves semantic PPTX content while
redesigning layout; Image to PPTX preserves a rendered page surface while
rebuilding its object and image-layer boundaries.
**Hard rule — flat final deck**: output stays `pptx_structure.mode: flat`. Page
pixels do not prove Master/Layout identity, placeholders, theme ancestry,
hidden objects, notes, animations, chart data sources, or authoring history.
Do not infer them. A reusable-system request routes to Create Template instead.
---
## 2. Normalize Source Images into Pages
🚧 **GATE**: an ordered canonical page-image roster exists before image-layer
decisions or SVG authoring.
| Source form | Normalization |
|---|---|
| One file containing one complete page | Keep it as one canonical page image |
| Several files containing one page each | Preserve the explicit or filename-natural order |
| One or more regular contact sheets | Split row-major with `slice_images.py --grid`, without alpha removal |
| A file containing several non-grid page frames | Record each visibly bounded page bbox and crop it into a separate lossless canonical page image |
| Boundaries or order are genuinely ambiguous | Mark the roster blocked instead of silently merging, dropping, or reordering pages |
Archive original files under `sources/`; keep normalized page images under
`images/source-pages/` or another project-local source-page folder. Never
overwrite an original file.
**Hard rule — one normalized frame, one slide**: frame count, not input-file
count, owns slide count. Every normalized page frame maps to one output slide
in the same order. Preserve the frame's aspect ratio. Mixed aspect ratios are
blocked until the current agent resolves one explicit whole-deck treatment.
**Mandatory — inspect every canonical page**: ordinary image-resource
inspection limits do not apply to this page roster. Inspect each normalized
page once to identify text, source graphics, scene-image regions, overlap,
region-level source sufficiency, boundary completeness, occlusion, and the
minimum useful layer stack. Reopen only the current page or a specifically
unresolved region afterward.
---
## 3. Reconstruct by Content Family
Classify visible regions by what they are, not by how easy they are to crop.
| Content family | Default realization | Non-negotiable boundary |
|---|---|---|
| Editable text | `native_text` | Restore exact visible wording, line grouping, alignment, emphasis, and approximate font metrics; do not bake ordinary slide text into generated images |
| Source graphic | `source_graphic` | Logos, icons, badges, vector-like ornaments, and decorative marks preserve visible identity. Use an exact vector, deterministic redraw, sufficient source pixels, or Codex reference reconstruction according to the quality ladder below; never substitute a merely similar graphic |
| Data graphic | `native_chart`, `native_table`, or exact `source_graphic` | Preserve every visible value, label, relationship, and geometry. Rebuild natively only when the source is legible enough to verify; otherwise use an exact crop/vector or mark `manual_required`. Never ask a generative model to recreate chart/table/data content |
| Simple exact geometry | `native_shape` | Use a native shape only when fill, stroke, geometry, and layering can be matched faithfully; otherwise prepare an identity-faithful source-graphic asset |
| Scene image | `image_layer` | Photos, people, characters, products, environmental backgrounds, textures, and complex illustrations may be reference-edited or regenerated as registered layers |
| Unreadable or unsafe region | `manual_required` | Block rather than invent wording, identity, values, or a visually different replacement |
**Hard rule — separate layer need from realization**: source clarity never
decides whether a required editable, movable, or overlapping object becomes a
separate layer; it decides only how that layer is prepared.
**Mandatory — assess source sufficiency per region**: judge each region at final
display size without a page-wide score or threshold. Inspect detail,
contamination/occlusion, and whether identity, geometry, lettering, or data
remain verifiable.
| Source evidence for a required independent image layer | Realization |
|---|---|
| Complete, cleanly separable, and sufficient at final display size | Prepare a source-derived crop or RGBA layer at the recorded geometry |
| Contaminated, occluded, incomplete, or too low-resolution; identity and geometry remain verifiable | Reference-edit or reconstruct the layer and exposed background from that evidence |
| Required identity, wording, values, or geometry cannot be verified | Mark `manual_required`; do not invent authoritative content |
**Graphic identity is authoritative; source pixel bytes are not**: use an exact
known vector when available. Deterministically redraw a simple, fully legible
graphic as SVG/native geometry. Reuse source pixels only when they are complete
and sufficient at final display size. When a complex logo, icon, badge,
ornament, or wordmark is visibly identifiable but too low-resolution, use its
source crop as the Codex reference and reconstruct a higher-resolution asset
that preserves the same silhouette, proportions, colors, lettering, bbox, and
z-order. Never merely interpolate low-resolution pixels, redesign the brand,
replace it with a similar library icon, or invent unreadable identity. If those
properties cannot be verified, mark the graphic `manual_required`.
**Visible-surface authority**: preserve every legible string, number, label,
relative position, crop, z-order, color relationship, and emphasis visible in
the source. Do not improve the layout, rewrite copy, correct claims through
research, reveal invented semantics, or silently replace branded graphics.
---
## 4. Build the Minimum Useful Layer Stack
For each page, decide the smallest stack that makes the intended objects
independent. Do not split a page merely to maximize layer count.
Typical bottom-to-top order:
1. `base` — clean full-canvas background with all planned removable subjects,
foreground objects, and editable text removed; hidden background pixels are
reconstructed where necessary.
2. `midground-*` — optional scene layers that must sit between the base and
primary subject.
3. `subject-*` — people, characters, products, props, or other independently
movable cutouts.
4. `foreground-*` — effects, foliage, particles, framing objects, or other
scene elements that cross the subject or native slide objects.
5. `source-graphic-*` — exact or identity-faithfully reconstructed logos,
icons, badges, and ornamental marks, plus exact/native data graphics at
their visible z-order.
6. `native-text-*` and exact native shapes.
**Registered-group rule**: every base/midground/subject/foreground layer in a
group stays registered to the same canonical page or scene bbox. A
source-derived member retains recorded geometry; every Codex-derived member
starts from that canonical source. Preserve canvas, position, scale,
pose, lighting, and style. Do not trim registered full-canvas layers;
transparent pixels retain alignment.
When one or more scene layers require reference editing or reconstruction, use
[`image-generator.md`](../../references/image-generator.md) §4.4's registered
reconstruction group as the primitive:
- create one clean base by removing **all** scene subjects/foreground objects,
source/data graphics, and editable text planned for separate realization,
then reconstruct only the newly exposed background;
- create at least one independent subject/foreground output from the same
canonical source whenever the page contains scene content that must be
independently editable; the base plus that output are the minimum two
independently prepared image layers, while §3 decides whether each layer
retains sufficient source pixels or requires reference reconstruction;
- derive every additional layer independently from the canonical source, never
from the base or another generated layer;
- preserve the original pose, scale, and coordinates on RGBA transparency;
- repeat only for additional layers that genuinely need independent movement,
overlap, or animation.
**Batch non-overlapping objects**: one object does not imply one generation.
When several subjects, props, effects, or source-graphic reconstructions have
pairwise-disjoint padded bboxes—including visible shadows and effects—and can
share one isolation treatment, ask Codex for one `layer-plate` containing all
of them with clear separation. Use either:
- a full-canvas registered plate that keeps the source positions, then create
one nested-SVG picture crop per recorded bbox; or
- a regular isolated-cell sheet when source coordinates are unnecessary, then
use `slice_images.py --grid ... --names ... --trim --alpha` and place the
resulting assets at their recorded source bboxes.
Both paths yield independent PPT picture objects from one generated output.
If transparency is unavailable, use one exact flat key color for the whole
plate and remove it once; never regenerate a separate green-background image
for each object. Objects that overlap one another or require different z-order
must use separate plates/layers.
The reference-image CLI does not inherit source dimensions automatically. Pass
an explicit matching aspect ratio/size, then verify that every member of one
registration group has the same final pixel canvas. In SVG, place the base and
all full-canvas layers at identical `x`, `y`, `width`, and `height` with
`no-crop` behavior.
**Reconstruct for final resolution**: apply the §3 source-sufficiency decision
per region. Retaining complete source pixels is valid only when they remain
sharp enough at final display size; a clear source may still require reference
reconstruction when separation needs hidden or uncontaminated pixels. When
detail is insufficient, use Codex reference reconstruction; interpolation alone
does not recover detail.
**Reference-edit, not reinterpretation**: reconstruction prompts name the
canonical source page/region and ask to preserve the visible composition and
style. They may inpaint hidden scene pixels or complete an occluded subject,
but must not redesign the scene, change a character/person, introduce text,
substitute or alter a logo, or invent extra decorative graphics.
When Codex cannot return transparency, generate the isolated layer or shared
plate on one exact flat key color and use `slice_images.py` as a `1x1` sheet
with `--alpha` and **without** `--trim`, preserving full-canvas registration.
Several plate members share this single keyed output.
---
## 5. Source Evidence without a Quick Plan
Before deciding layers, write source evidence to:
```text
<project_path>/analysis/reconstruction_inventory.json
```
The inventory records what is visibly present, not a resumable implementation
plan. Keep it limited to:
- original file and normalized page path;
- page order, source-frame bbox, SHA-256, and pixel dimensions;
- visible regions with stable ids, source bboxes, observed family
(`text`, `graphic`, `image`, or `unknown`), verbatim text when applicable,
and confidence;
- observed source sufficiency, boundary completeness, occlusion/contamination,
and identity/data verifiability at final placement;
- overlap/z-order observations and unresolved evidence.
Do **not** put final layer choices, generation prompts, output filenames, or SVG
bindings into this inventory. The current main agent keeps those decisions in
active context and writes only required operational image manifests/evidence.
Context loss restarts the Quick run; the inventory is not a resume artifact.
Low-confidence visible text, an uncertain page boundary, or an unidentified
branded/data graphic is unresolved evidence and blocks successful delivery.
---
## 6. Image Preparation
When any `image_layer` or low-resolution `source_graphic` requires reference
editing or generation, load
[`image-base.md`](../../references/image-base.md) and
[`image-generator.md`](../../references/image-generator.md). The current Codex
main agent resolves the layer stack directly, uses Codex's native
reference-image capability, and finishes every required layer before SVG
authoring. Do not adapt `image_gen.py`, its generic manifest, or provider
backends for this profile.
- Use `text_policy: none` for scene reconstruction layers. Use `embedded` only
when an exact visible wordmark/letterform is integral to a reconstructed
source graphic; ordinary slide text always remains native.
- Exhaust the available Codex image path automatically; block before export if a
required layer remains `Needs-Manual`.
- Preserve each prepared image layer's source page/region, source hash,
realization method, operation, output path/hash, registration group, and
z-order in the applicable operational evidence; include prompt and
backend/model when the layer was reference-edited or reconstructed.
- Re-run `analyze_images.py` after assets change.
- A generated candidate is not usable until its expected file exists, it has
been inspected once, and its registration group or plate has been checked
against the canonical page.
- Inspect the recomposed page once after all generated layers, plate crops,
source graphics, native shapes, and native text are in place. This narrow
readback is mandatory fidelity validation, not resource reselection.
---
## 7. SVG Authoring and Release Gate
Follow Quick Generate after source normalization and resource preparation.
Hand-author pages serially from the prepared base, registered scene layers,
identity-faithful source graphics, native shapes, and native text. Give independently
movable layers stable direct-root group ids so later animation can target them.
**Forbidden — screenshot skin**: do not use the complete source page as the
sole full-slide picture and add token editable text above it. The source page
is a comparison reference, not a hidden backing layer in the delivered slide.
Verify each page against its canonical image:
| Final check | Required evidence |
|---|---|
| Page roster | Every normalized frame becomes one slide in the same order and canvas treatment |
| Native text | Every legible string/number is present verbatim and remains editable |
| Source graphics | Logos, icons, and decorative graphics preserve the original identity and geometry at adequate final resolution; no similar substitute or unverified redesign appears |
| Data graphics | Every chart/table/data value and relationship is native-and-verified or retained from an exact source asset; none is generatively recreated |
| Layer registration | Base, subject, foreground, and other generated layers share the expected canvas/placement and show no jumps, seams, halos, or independent-crop drift |
| Visible image fidelity | The recomposed scene preserves the source's visible subject identity, pose, crop, lighting, color relationships, and z-order |
| Honest reconstruction | AI-recovered hidden pixels are identified as reconstruction, not claimed as original source detail |
| Independent objects | Every layer requested for editing or animation is a distinct SVG/PPT picture object; non-overlapping members may originate from one shared generated plate |
| Reference exclusion | Canonical full-page source images remain comparison evidence and are not referenced or packaged as delivered slide media |
| Package quality | Quick's lockless final SVG checker and PPTX postflight pass |
If a generated layer drifts, retry from the canonical reference with a narrower
edit instruction. Do not compensate by changing native text/graphics or by
flattening the full page. If the Codex image path is exhausted,
mark the affected layer `Needs-Manual` and block successful export.
```markdown
## ✅ Image to PPTX Complete
- [x] Source files were normalized into the complete ordered page roster
- [x] Visible text is native and verbatim
- [x] Source graphics preserve identity and are sharp enough at final size
- [x] Required background / foreground / subject layers are independent and registered
- [x] Shared plates were split/cropped into the required independent objects
- [x] Recombined pages match the supplied visual references
- [x] Canonical full-page source images are absent from delivered slide media
- [x] Quick's SVG quality gate and PPTX postflight pass
- [ ] **Next**: Report the PPTX and identify native, exact-source, and AI-reconstructed objects
```
@@ -69,6 +69,14 @@ defaults when no row supplies a concrete communication job. When several
signals apply, perform every required action and use the earliest required load
point; a before-authoring signal always overrides a before-export-only timing.
**Direct-video application**: when [`video-design.md`](../../references/video-design.md)
is active for Quick direct delivery, use this gate as the required pre-audio
motion choice, not an animation quota. The page/object-specific row activates
Custom Animations; when narration should govern those semantic reveals, it also
activates narration-cue synchronization. The deck-wide row remains independent
of narration and must not be reported as synchronized; static or
page-transition-only motion remains valid.
---
## 2. Source and Resource Preparation
@@ -78,6 +86,7 @@ Prepare source facts before initialization:
| Input | Action |
|---|---|
| Topic or requirements without supporting facts | Run [`topic-research`](../stages/topic-research.md) immediately and retain its Markdown supplement plus fact-provenance JSON for import |
| One or more PNG / JPEG / WebP files representing page frames under Image to PPTX | Do not call `source_to_md.py`; normalize single-page files and multi-frame contact sheets into the canonical ordered frame roster through that profile, then import the originals below |
| PDF / DOCX / Office document / XLSX / XLSM / PPTX / EPUB / HTML / LaTeX / RST / web URL | Run `python3 ${SKILL_DIR}/scripts/source_to_md.py <file_or_URL_or_dir> [<file_or_URL_or_dir> ...]` |
| CSV / TSV | Read directly as a plain-text table source |
| Markdown or direct conversation text | Read directly |
@@ -97,6 +106,7 @@ After reading every direct and converted source, assess factual sufficiency:
| Material state | Action |
|---|---|
| Image to PPTX page surface | Treat as a closed visible corpus; unreadable/occluded regions become `manual_required`, never external research |
| The requested outcome is supported | Continue |
| A required externally verifiable claim remains unsupported | Run [`topic-research`](../stages/topic-research.md) for those gaps only |
| Closed corpus / source-only / no external enrichment | Stay within the supplied material |
@@ -106,8 +116,20 @@ require inventing, omitting, or leaving unsupported an externally verifiable
claim. File presence or length does not establish sufficiency. Research gathers
facts only; image acquisition remains part of the resource preparation below.
**Conditional video-delivery context**: when the intended use is recorded,
self-running, or video-directed—or an explicit final/literal narration script
will become notes/audio—read
[`video-design.md`](../../references/video-design.md) now and retain it through
roster, SVG, notes, and motion decisions. This changes neither the Quick profile
nor its artifacts.
Before initialization, resolve exactly one template branch:
When [`image-to-pptx.md`](./image-to-pptx.md) is active, its canonical page
surface owns the design: select **Free design** directly and do not inspect,
install, or apply a supplied template workspace. The branches below apply to
ordinary Quick and other compatible profiles.
- **Direct template application**: one or more exact current workspace roots
were supplied in the request, or Create Template returned an exact validated
root in the current conversation. Accept at most one root per declared kind.
@@ -159,6 +181,10 @@ target project; every external path is copied and remains untouched. Use
wrote Markdown beside the original source, pass that source path or directory
once; when `-o` wrote it elsewhere, pass both locations. Direct supported bitmap
inputs are archived under `sources/` and copied collision-safely into `images/`.
When [`image-to-pptx.md`](./image-to-pptx.md) is active, its
normalized frame roster is canonical page-surface input and the current main
agent writes the source-evidence-only `analysis/reconstruction_inventory.json`
before deciding the layer stack in active context.
For each imported PPTX, `import-sources` automatically writes
`analysis/<stem>.identity.json`, `analysis/<stem>.slide_library.json`, and the
@@ -203,9 +229,11 @@ use; otherwise decide which prototypes to use, skip, repeat, reorder, or adapt
while authoring. Persist no separate template-application artifact. If no
template was installed, make the same design choices freely.
Before resolving the one-pass design, read only these three indexes:
Before resolving the one-pass design, read the canvas authority and only these
three choice indexes:
```
Read references/canvas-formats.md
Read references/modes/_index.md
Read references/visual-styles/_index.md
Read references/image-renderings/_index.md
@@ -227,7 +255,9 @@ checkpoint, or persist a page/resource plan.
Before writing P01, resolve in active context:
- the exact slide roster and one compact core message for every page, used to choose its composition and hierarchy;
- the canvas, visual direction, palette, wording, and one concrete typography plan using installed font families, with stable size anchors for title, body, annotation, and every other recurring role the roster uses; explicit user, template, or resolved-style requirements may call for a deliberate exception;
- the effective Speaker Notes, Custom Animations, and Narration Audio outcomes; generated narration requires notes. A deck intended only for later recording forces neither audio nor object animation. Quick direct video follows [`video-design.md`](../../references/video-design.md): enable notes/narration/video delivery, decide motion before audio, and enable Custom Animations only for narration-cued or other page/object-specific motion;
- the canvas, visual direction, wording, intended viewing distance, and effective reading mode: choose `presentation` for distance-first projected or recorded viewing, `balanced` for mixed viewing, or `text` for close content-heavy reading. Take the initial body anchor and sanity band from [`canvas-formats.md`](../../references/canvas-formats.md) § "Typography Scale Start" for the resolved canvas—PPT remains reading-mode-driven, while registered/custom non-PPT canvases use their canvas-derived start—then resolve one concrete typography plan using installed font families, with stable size anchors for title, body, annotation, and every other recurring role the roster uses. When content does not fit, preserve its core message and apply only fitting actions the source/profile invariants permit—restructure, shorten, or split; if none is permitted, surface the unresolved fit instead of shrinking a recurring role. Explicit user, template, fidelity-profile, or resolved-style requirements may call for a deliberate exception;
- the semantic color roles actually needed by the roster, each with a concrete active-context color anchor, including background/surface, primary/secondary text, dominant/accent, and status roles as applicable. Honor explicit user, installed template/brand, fidelity-profile source-identity, and resolved-style color semantics before deriving only the missing roles that the active profile permits; decide which roles dominate, support, or remain rare, and preserve sufficient contrast for meaning-bearing text. Pair newly authored color-coded states, categories, or relationships with a label, symbol, line, or geometry cue; when fidelity forbids adding one, preserve the source encoding;
- an ordinary body-content frame and a density judgment for every page, adapted to the canvas and any user / template / style geometry; use `anchor`, `dense`, `breathing`, or an equivalent active-context distinction instead of one uniform fill level;
- for each page not bound to literal supplied geometry, a primary visual zone and page-scale composition direction tied to its core message; use cards or equal grids when the content relationship calls for them, not as the automatic page grammar;
- for each page, preserve its semantic units, source-stated qualitative relationships, intended entry, and outcome so §3 can make the sole Structure decision before geometry;
@@ -242,6 +272,15 @@ Before writing P01, resolve in active context:
otherwise choose the registered automatic/default path without another
interaction.
**Prepared final narration**: when the user explicitly marks a script as
final/literal and intends it for notes or generated audio, segment it by semantic
scene while resolving the roster and preserve every spoken word. Before writing
P01, write the ordered segments once to `notes/total.md` with
`# Slide <number>` headings and `---` separators. Keep that file as exact
production input for page design; it is not a planning checkpoint. Do not split
it until the SVG roster exists. Draft narration instead remains source material
and uses the ordinary post-SVG notes branch when notes are enabled.
**Mandatory — subject-layer capability scan**: Before resource preparation,
inspect the intended layer relationship rather than treating every reference
visual as one flat image. When a subject crosses a native title, panel, frame,
@@ -317,19 +356,24 @@ Prepare only the resource paths needed by the decided pages:
| Resource | Required preparation |
|---|---|
| Supplied/extracted image | Copy the selected file into `images/`; preserve its factual/provenance context and use the measured file rather than an invented substitute |
| Image-to-PPTX reconstruction asset | In Codex, preserve identity graphics through an exact vector, deterministic redraw, sufficient source asset, or reference-based high-resolution reconstruction; keep data graphics native-and-verified or exact. For scene imagery, build the minimum registered clean-base/midground/subject/foreground group; batch padded-bbox-disjoint objects into one shared plate, then split them with grid slicing or independent nested-SVG bbox crops |
| Bundled/custom icon | Follow the [icon library contract](../../templates/icons/README.md), choose one coherent primary library, sync a useful project pool covering recurring semantics and likely page-local needs without assigning icons to pages, and choose from that prepared pool during SVG authoring |
| Formula | Follow the [`latex_render.py` contract](../../scripts/docs/image.md), write `images/formula_manifest.json`, run the renderer, and keep the rendered PNG under `images/` |
| AI image | Follow `image-base.md` + `image-generator.md`; apply only the chosen rendering preset or exact custom bases, never blend unselected catalog identities, and keep `image_prompts.json` plus its human-readable sidecar |
| Web image | Follow `image-base.md` + `image-searcher.md`; keep query/status data and `image_sources.json`, including any required on-slide attribution |
| Illustration slice | Generate or obtain the parent sheet, run `slice_images.py`, and place only the resulting element files |
| Registered subject/base pair | Follow `image-generator.md` §4.4; keep the clean base and transparent subject on the same full canvas, then place both with `crop=no-crop` |
| Registered reconstruction group | Follow `image-generator.md` §4.4; keep full-canvas members registered with `crop=no-crop`, and materialize every required shared-plate member as an independent picture object |
| Visualization | Keep Chart values, Table cell topology, and chosen treatment in active context; load the applicable Chart/Table authority in §3 and write native replacement metadata only for an independently selected native-ready object |
**Image inspection boundary**: acquisition-time suitability review follows the
owning AI/web/slice reference. Once resources reach terminal status, SVG
authoring follows `executor-image.md`'s narrow placement inspection: inspect only
one specifically ambiguous `Existing`/`Sourced` asset and never routinely reopen
`Generated` outputs.
`Generated` outputs. Image to PPTX is the narrow fidelity exception: inspect
every normalized page once for its inventory, inspect every generated
reconstruction layer or shared plate once, and inspect the final recomposition
against the canonical frame. Reopen only the current page or one unresolved
region after that required comparison.
After image resources change, run `analyze_images.py` so
`analysis/image_analysis.csv` reflects the files that SVG authoring will use.
@@ -369,9 +413,7 @@ page and reuse throughout the valid execution context:
[`image-layout-patterns.md`](../../references/image-layout-patterns.md), and
[`svg-image-embedding.md`](../../references/svg-image-embedding.md); add
[`executor-web-image.md`](../../references/executor-web-image.md) for a sourced
web image. Reread only after a known file change or context invalidation. Load
[`canvas-formats.md`](../../references/canvas-formats.md) only for a non-default
canvas.
web image. Reread only after a known file change or context invalidation.
`executor-structure.md` is loaded once before all SVG authoring so Quick cannot
omit shape-composition reasoning. Reuse it throughout the valid execution
@@ -389,6 +431,13 @@ relationship better. Formula-only pages use
[`image-layout-spec.md`](../../references/image-layout-spec.md) without forcing a
multi-image system.
Image to PPTX replaces this open composition decision for its canonical page
frame: preserve the source geometry, restore text natively, preserve
source-graphic identity through the prepared exact or reconstructed asset, and
use the active-context registered layer/plate stack for scene imagery. Run the
ordinary decision only for an additional non-source image whose placement is
not already fixed by that surface.
**Mandatory — per-page Structure decision**: after the current page's content
and communication move are determined, but before choosing any geometry or
shape, decide whether geometry must carry qualitative `order`, `link`, `parent`,
@@ -416,7 +465,12 @@ native-ready or replaces the per-page Structure decision.
Keep the core's shared visual-quality / leading defaults and `svg-effects.md` §6.1 Visual Job Router active while authoring. Explicit user/template requirements and the resolved style override compatible aesthetic defaults, never technical Required / Forbidden boundaries.
**Per-page execution anchors**: apply the transient core-message, typography-role, body-frame, density, and composition anchors resolved in §2 while authoring; they guide the current run without creating a persisted planning artifact.
**Per-page execution anchors**: apply the transient core-message, typography-role, semantic-color, body-frame, density, and composition anchors resolved in §2 while authoring; they guide the current run without creating a persisted planning artifact.
When `notes/total.md` was frozen from a final script, retain its corresponding
segment while authoring each page. The visible state and real direct-root
semantic groups must support that spoken segment without duplicating the full
script as body copy or changing its wording.
Use one zero-padded filename width sized for the resolved roster, such as
`01_cover.svg` through `12_end.svg` or `001_cover.svg` through `120_end.svg`.
@@ -453,6 +507,16 @@ active context; create no planning artifact or approval stop. After every page
exists, run the one final checker below. Apply other supporting tools and
stages only when their capability is actually needed.
**Hard rule — direct page authoring stays with the current main agent**: write
every page SVG directly in the active context. Do not delegate page generation
to another agent, and do not run a Python, Node, shell, or other generator that
writes slide files into `svg_output/`. Documented fragment-only helpers remain
allowed after the current main agent chooses the object's role, operands,
paint, and z-order and integrates the fragment itself. This boundary does not
restrict resource preparation, inspection, checker, verification,
post-processing, or export tools; a run fails this profile only when a delegated
agent or generator authors a page SVG on the main agent's behalf.
This is not a resume protocol. If the active context is lost before delivery,
start a clean Quick run rather than inferring an unfinished plan from the files
already present.
@@ -473,15 +537,47 @@ python3 ${SKILL_DIR}/scripts/svg_quality_checker.py <project_path> \
Fix every blocking error and rerun the same command. Then export:
When Speaker Notes is enabled, load
[`executor-notes.md`](../../references/executor-notes.md) after the passing final
check. Validate an already frozen final script or direct-video pre-SVG narration
without regenerating it; otherwise generate `notes/total.md` from the final SVG
roster. Then run:
```bash
python3 ${SKILL_DIR}/scripts/svg_to_pptx.py <project_path> --quick-generate
python3 ${SKILL_DIR}/scripts/total_md_split.py <project_path>
```
Run [`customize-animations`](../stages/customize-animations.md) after that notes
pass when the active-context outcome or an existing sidecar triggers it. Resolve
deck-wide-only motion through the selected exporter flags instead.
For Quick direct video, preserve the motion choice made under
[`video-design.md`](../../references/video-design.md). A static,
page-transition-only, or narration-independent deck-wide choice may export
without `animations.json` and does not claim object-level audio sync. Any
page/object-specific choice completes the custom stage before base export;
`generate-audio` derives cue timing only for narration-governed groups and
otherwise exports canonical timing without an object-sync claim.
Choose exactly one notes mode for the base export:
```bash
# Speaker Notes enabled
python3 ${SKILL_DIR}/scripts/svg_to_pptx.py <project_path> \
--quick-generate --with-notes
# Speaker Notes disabled
python3 ${SKILL_DIR}/scripts/svg_to_pptx.py <project_path> \
--quick-generate --no-notes
```
`--quick-generate` reads `svg_output/` as the page source and resolves the
project-local assets referenced by those SVGs. It infers one consistent canvas,
uses a lockless flat PowerPoint package, and does not force-disable ordinary
export options. Notes, custom object animation, and narration remain off unless
selected by the agent. Do not run `finalize_svg.py`.
selected by the agent. Do not run `finalize_svg.py`. After the validated base
export, run [`generate-audio`](../stages/generate-audio.md) when Narration Audio
is enabled; it owns page audio/SRT, narrated PPTX, and optional native MP4.
The exporter requires a passing `final` report whose SVG fingerprint matches
the current `svg_output/`; missing, blocking, non-final, or stale reports stop
@@ -502,7 +598,8 @@ or lock.
- [x] Every role declared by an installed template spec is locatable in the finished pages, or its non-use is deliberate — checked per installed spec, not from memory
- [x] Every triggered capability-specific preparation and pre-checker verification completed
- [x] The lockless final SVG quality report passes and matches the current SVGs
- [x] Enabled notes were validated/generated and split; enabled custom motion ran through its owning stage
- [x] One native PPTX exists under `exports/` or the explicit output path
- [x] No Strategist, confirmation, root project Design Spec, or lock artifact was created
- [ ] **Next**: Report the PPTX path
- [ ] **Next**: Report the base PPTX and any enabled narrated PPTX/MP4 outputs
```
@@ -30,7 +30,7 @@ route selection. After selection, the active runtime authority owns execution.
| Route | Request shape | Authority | Preconditions | Mutation model | Output contract |
|---|---|---|---|---|---|
| Generate PPTX | Create a new presentation; regenerate an existing deck visually; use source material or a topic; optionally select a registered library template or supply an explicit workspace | Beautify: [`beautify-pptx`](./profiles/beautify-pptx.md), which selects Default or Quick; ordinary Default: [`generate-pptx`](./generate-pptx.md); ordinary explicit Quick: [`quick-generate`](./profiles/quick-generate.md) | Source facts exist or research can gather them; explicit quick intent activates its runtime | Author new SVG pages and export a new PPTX | Default: spec, lock, SVG, validation, and PPTX; Quick: optional source/resource artifacts, no spec/lock, SVG, and one PPTX |
| Generate PPTX | Create, reconstruct, or visually regenerate a presentation/video from sources or a topic; templates remain optional | Image to PPTX: [`image-to-pptx`](./profiles/image-to-pptx.md), always Quick; Beautify: [`beautify-pptx`](./profiles/beautify-pptx.md), Default or Quick; ordinary [`generate-pptx`](./generate-pptx.md) / [`quick-generate`](./profiles/quick-generate.md) | Facts exist or research can gather them; Image to PPTX also requires Codex and an ordered page-frame roster | Author SVG pages and export a new PPTX | Default: spec/lock/SVG/PPTX; Quick: optional source/resource artifacts, no spec/lock, SVG/PPTX; either may derive narrated PPTX/MP4 |
| Create Template | Create a reusable brand/style/layout/deck template from one or more PPTX/SVG files, images/PDFs, direct or file-based text, documents/websites, brand assets, or a mixed reference bundle | [`create-template`](./create-template.md) | A reusable-template request exists; reference material is optional, and project scope additionally requires an initialized target project | Author a new portable workspace; never modify any reference file in place | Workspace with required `templates/`, optional `images/` / `icons/`, and optional review `exports/` |
| Fill Native PPTX | Use a raw PPTX's native slide shells and replace/fill content | [`template-fill-pptx`](./template-fill-pptx.md) | Source PPTX plus new material/topic | Clone and patch PPTX through OOXML; no SVG pipeline | New filled PPTX in project `exports/` |
| Enhance Native PPTX | Keep a finished PPTX's visible slides stable while adding notes, audio, timings, or transitions | [`native-enhance-pptx`](./native-enhance-pptx.md) | Finished source PPTX exists | Append/update scoped OOXML parts; no slide regeneration | New enhanced PPTX in project `exports/` |
@@ -41,26 +41,29 @@ route selection. After selection, the active runtime authority owns execution.
| Request condition | Generate-route behavior |
|---|---|
| One or more raster files represent page frames that must be reconstructed into a layered editable PPTX | Activate the Codex-supported [`image-to-pptx`](./profiles/image-to-pptx.md); normalize the represented frame roster and activate `quick-generate` directly without requiring a separate Quick request |
| Existing PPTX must preserve wording, page count, and page order 1:1 | Activate [`beautify-pptx`](./profiles/beautify-pptx.md); it selects `quick-generate` when that profile's explicit trigger also matches, otherwise `generate-pptx` |
| Explicit quick/fast, skip-strategy, or direct SVG-to-PPTX intent without Beautify | Load [`quick-generate`](./profiles/quick-generate.md) directly without loading `generate-pptx.md`: prepare sources/resources as needed, let the current agent decide without interaction, directly apply at most one exact workspace root per kind supplied for this run, otherwise use free design, omit Strategist/Confirm UI/spec/lock, hand-author SVG, run the lockless final checker, and export the final PPTX |
| The effective delivery purpose is recorded, self-running, or video-directed | Inside the already selected Default or explicit Quick runtime, load [`video-design`](../references/video-design.md) before whole-solution/page planning. This is a conditional design reference, not a profile or fifth route; notes, animation, audio, and optional native MP4 remain owned by their existing stages |
| Explicit quick/fast, skip-strategy, or direct SVG-to-PPTX intent without an active fidelity profile | Load [`quick-generate`](./profiles/quick-generate.md) directly without loading `generate-pptx.md`: prepare sources/resources as needed, let the current agent decide without interaction, directly apply at most one exact workspace root per kind supplied for this run, otherwise use free design, omit Strategist/Confirm UI/spec/lock, hand-author SVG, run the lockless final checker, and export the final PPTX |
| Topic only, or supplied sources leave planning-critical factual gaps | Run [`topic-research`](./stages/topic-research.md) inside the selected Generate profile's source preparation: immediately for topic-only input, or after conversion and reading for source-backed input; research only the identified gaps |
| Existing PPTX may be split, merged, dropped, reordered, or re-outlined | Treat the PPTX as source content through the selected Generate authority's source intake; continue the default pipeline unless explicit Quick Generate intent selected that profile |
| Existing PPTX may be split, merged, dropped, reordered, or re-outlined | Treat the PPTX as source content through the selected Generate authority's source intake; continue Default unless explicit Quick intent selected that runtime |
| Default Generate reaches planning | Step 3 prepares template candidates without interaction. Stage 1 then confirms the communication contract and free-design/template choice together; only a confirmed non-free choice runs [`apply-template-workspace`](./stages/apply-template-workspace.md) before Stage 2 |
| Explicit current brand/style/layout/deck workspace root | Default Generate preserves the exact path as a Stage-1 template candidate; Quick Generate validates and installs it directly without Steps 34 or Confirm UI. Classify it as `library` only when its normalized root exactly matches a registered index entry; otherwise retain `explicit`. Consume the workspace root, never only its inner `templates/` directory |
| Explicit current brand/style/layout/deck workspace root outside Image to PPTX | Default Generate preserves the exact path as a Stage-1 template candidate; Quick Generate validates and installs it directly without Steps 34 or Confirm UI. Classify it as `library` only when its normalized root exactly matches a registered index entry; otherwise retain `explicit`. Consume the workspace root, never only its inner `templates/` directory |
| Split-mode project resumes in a fresh chat | Run [`resume-execute`](./stages/resume-execute.md) inside the active Generate route |
| Existing generated project needs a deck-wide `colors.*` or universal `typography.font_family` substitution | Stay in Generate; load [`update_spec.py`](../scripts/docs/update_spec.md), honor its supported-key boundary, then rerun the final quality gate and Step 7 export |
| User explicitly requests spec refinement | Run [`refine-spec`](./stages/refine-spec.md) after Design Spec Gate 1 and before lock Gate 2 |
| Data charts exist | Run [`verify-charts`](./stages/verify-charts.md) before export |
| User explicitly requests visual review | Run [`visual-review`](./stages/visual-review.md) before post-processing |
| User requests preview, selection, or annotation application | Use the default Generate pipeline and run [`live-preview`](./stages/live-preview.md) at the stage defined there; explicit Quick + preview intent falls back to default rather than dropping preview |
| User requests preview, selection, or annotation application outside Image to PPTX | Use the default Generate pipeline and run [`live-preview`](./stages/live-preview.md) at the stage defined there; explicit Quick + preview intent falls back to default rather than dropping preview. Image to PPTX remains Quick-only and uses its mandatory canonical-frame recomposition comparison instead of this interactive stage |
| User requests page transitions, auto-advance, or deck-wide animation settings without page-specific motion planning or an existing `animations.json` | Load [`animations`](../references/animations.md) and apply its export-level contract |
| `<project_path>/animations.json` already exists, the user explicitly requests per-slide/object-level animation control, or the effective Custom Animations outcome in `design_spec.md §I` is enabled | Run [`customize-animations`](./stages/customize-animations.md) after the final SVG quality gate and any enabled speaker-note pass, before Generate Step 7. A §IX `Motion suggestion` informs an active pass but never triggers it alone |
| Generate PPTX receives an explicit narration request or has effective Narration Audio enabled in `design_spec.md §I`; Enhance Native PPTX has a confirmed `audio.enabled: true` module | Run [`generate-audio`](./stages/generate-audio.md) after the owning route's notes/export readiness; Generate audio implies effective Speaker Notes enabled |
**Hard rule — profile, not fifth route**: Beautify changes content/page
invariants, then selects one existing Generate runtime: explicit Quick intent
uses Quick; otherwise it uses Default. It does not define a separate artifact
lifecycle or load both runtimes.
**Hard rule — fidelity profiles, not fifth routes**: Image to PPTX and Beautify
change different source/page invariants and are mutually exclusive. Image to
PPTX always activates Quick; Beautify uses Quick only on explicit Quick intent
and otherwise uses Default. Neither defines a separate artifact lifecycle or
loads both runtimes.
**Hard rule — direct-generation profile, not a fifth route**: `quick-generate`
stays inside Generate PPTX but owns an explicit SVG → PPTX short circuit. Page
@@ -70,7 +73,8 @@ or agent-selected. Quick may consume exact Brand/Style/Layout/Deck workspaces as
flat authoring inputs; compiling reusable Master/Layout/placeholder structure
still requires the default lock-backed Generate pipeline. Once selected, Quick
is the complete runtime procedure and never loads `generate-pptx.md`; Default
never loads `quick-generate.md`. Beautify may select either one, but not both.
never loads `quick-generate.md`. Image to PPTX is the narrow profile-owned
Quick activation; Beautify may select either runtime, but never both.
---
@@ -89,6 +93,7 @@ When a PPTX already contains native Master/Layout parts, `create-template` mirro
| Input | Route behavior |
|---|---|
| One or more images containing page frames + explicit final-deck reconstruction intent | Generate PPTX with the Codex-supported, Quick-only [`image-to-pptx`](./profiles/image-to-pptx.md); normalize page frames first and do not infer reusable native structure from pixels |
| Raw PPTX called a template + new content | Fill Native PPTX unless the user explicitly asks for a reusable template workspace |
| Any supported reference bundle or direct-text brief + reusable template request | Create Template |
| Current template workspace root + content | [`generate-pptx`](./generate-pptx.md) Stage-1 template choice |
@@ -27,6 +27,8 @@ this stage.
- Per-page narration files exist at `notes/*.md`. In Generate PPTX, split `notes/total.md` during Step 7.1. In Enhance Native PPTX, the notes module writes numeric files such as `001.md`.
- Default mode: `edge-tts` is installed (`python3 -m pip install edge-tts`).
- The stage is page-level only: one note becomes `audio/<stem>.<audio-ext>` plus `audio/<stem>.srt` on provider-timed paths, or one audio file with Qwen / explicit CosyVoice audio-only mode. Never substitute one long track or automatic splitting.
- Final/literal script notes are synthesized verbatim. Source SRT timecodes are pacing evidence only; new provider timing owns the generated audio/SRT set.
- SRT bound to an authoritative existing recording does not enter TTS. Recorded narration requires page-level audio or an explicit page/time map; automatic long-track splitting is unsupported.
- A fully successful run writes a compact `audio/manifest.json` with only provider/model, audio/subtitle format, relevant voice settings, and a SHA-256 fingerprint instead of the raw cloud voice ID. It has no per-slide inventory, artifact hashes, or API keys and is not a normal generation input. The flat `audio/` directory is the single active narration set; do not create provider subdirectories unless the user explicitly asks to preserve multiple variants.
- PPT narration assets must be PowerPoint-reliable audio: `m4a` (AAC), `mp3`, or `wav`. The built-in TTS path defaults to `mp3`; provider formats such as `pcm`, `opus`, or `flac` must be transcoded before embedding.
- PowerPoint recorded narration export requires `ffprobe` so slide timings can be written from actual audio duration.
@@ -105,11 +107,20 @@ For each candidate, write a **one-line Chinese description** covering: 性别 ·
---
## Step 3: One-shot user interaction (mandatory)
## Step 3: Resolve generation settings
Send a single message to the user that resolves all five configuration decisions at once and provides a recommended value for each. Before offering automatic video export, run `python3 skills/ppt-master/scripts/powerpoint_video.py --check`; do not present an unavailable local capability as executable. Do NOT split into multiple rounds.
**Quick exception**: do not pause. Apply explicit user values, then resolve
unspecified provider, voice, rate, and embed choices from the recommended-value
rules below. Keep video off unless the caller selected direct video; then embed
the narrated PPTX and continue to native video only when
`powerpoint_video.py --check` succeeds. Require a timestamp-capable provider
only when narration-cue sync or subtitle delivery needs page-local SRT.
**Cloned-voice fast path**: if the user mentioned a cloned voice / 克隆音色 / 复刻音色 / "my own voice" along with a `voice_id`, skip the voice-recommendation list — set the provider to whichever the user named (`elevenlabs` / `minimax` / `qwen` / `cosyvoice`), pin the `voice_id` they gave you, and only confirm rate + embed + video.
**Default / Enhance Native — one-shot interaction (mandatory)**:
For Default or Enhance Native, send one message that resolves all five configuration decisions and recommends each value. Before offering automatic video export, run `python3 skills/ppt-master/scripts/powerpoint_video.py --check`; do not present an unavailable local capability as executable. Do NOT split into multiple rounds.
**Cloned-voice fast path**: if the user mentioned a cloned voice / 克隆音色 / 复刻音色 / "my own voice" along with a `voice_id`, skip the voice-recommendation list — set the named provider (`elevenlabs` / `minimax` / `qwen` / `cosyvoice`) and pin that `voice_id`. Quick applies its exception above; Default and Enhance Native confirm only rate + embed + video.
**Message template** (Chinese; translate to user's chat language if different). “Embed” means caller-specific integration: SVG re-export for Generate PPTX, or native OOXML application for Enhance Native PPTX.
@@ -179,29 +190,33 @@ python3 skills/ppt-master/scripts/notes_to_audio.py <project_path> \
--provider cosyvoice --voice-id <chosen-voice> \
--cosyvoice-model cosyvoice-v3-flash
# 2A. Only when page-local SRT exists and animations.json is active, author or
# refresh narration_timing.json
# 2A. Only when narration-cue sync is selected and page SRT + animations.json
# exist, author or refresh narration_timing.json
# by matching SVG group semantics to SRT topics, then derive the narrated
# sidecar. Reuse current SVG semantics when complete; otherwise read only
# the missing or stale svg_output pages.
python3 skills/ppt-master/scripts/narration_sync.py animations <project_path> \
--narration-padding 0.5 --force
--narration-start-floor 0.8 --narration-padding 0.5 --force
# 2B. Re-export with audio embedded
# Use the base export's [REPORT] path to preserve source-bound deck motion.
# Quick Generate adds --quick-generate --with-notes to every re-export below.
python3 skills/ppt-master/scripts/svg_to_pptx.py <project_path> \
--recorded-narration audio --narration-padding 0.5 \
--recorded-narration audio \
--narration-start-floor 0.8 --narration-padding 0.5 \
--inherit-motion-from "<base_postflight_report>"
# Optional: use the canonical presentation animation instead
python3 skills/ppt-master/scripts/svg_to_pptx.py <project_path> \
--recorded-narration audio --narration-padding 0.5 \
--recorded-narration audio \
--narration-start-floor 0.8 --narration-padding 0.5 \
--animation-config animations.json \
--inherit-motion-from "<base_postflight_report>"
# Optional: export narration with no object or page-transition animation
python3 skills/ppt-master/scripts/svg_to_pptx.py <project_path> \
--recorded-narration audio --narration-padding 0.5 \
--recorded-narration audio \
--narration-start-floor 0.8 --narration-padding 0.5 \
--no-animations
# 2C. Only when page-local SRT exists, merge it against timing values read
@@ -234,11 +249,17 @@ Provider-timed paths share punctuation-first, `--subtitle-max-chars`-bounded reg
Before generation starts, `notes_to_audio.py` removes stale `audio/manifest.json` and `audio/total.srt`; an incomplete run therefore cannot claim the previous set's provenance or merged timeline. A successful audio-only provider run also removes same-stem stale SRT files. The new manifest is published atomically only after the complete page roster succeeds.
**Mandatory when `animations.json` is consumed — semantic animation context**: Before writing or refreshing `<project_path>/narration_timing.json`, determine whether the active context already contains the current top-level SVG group IDs and visible group-content semantics for every affected page. Reuse that context without rereading SVG when it is complete and still matches the current `svg_output/`. If any page is missing, stale, or represented only by group IDs/order without content meaning, read only that page's SVG as a read-only source and extract the missing group semantics. Always combine those semantics with the page SRT topics/timestamps and `animations.json`; group order alone is not a semantic narration mapping.
**Mandatory when narration-cue sync is selected — semantic animation context**: Before writing or refreshing `<project_path>/narration_timing.json`, determine whether the active context already contains the current top-level SVG group IDs and visible group-content semantics for every affected page. Reuse that context without rereading SVG when it is complete and still matches the current `svg_output/`. If any page is missing, stale, or represented only by group IDs/order without content meaning, read only that page's SVG as a read-only source and extract the missing group semantics. Always combine those semantics with the page SRT topics/timestamps and `animations.json`; group order alone is not a semantic narration mapping.
> Active `animations.json` requires `narration_timing.json`; explicit `--no-animations` bypasses both. Without a sidecar, `narration_sync.py animations` maps groups **positionally** (group N → cue N) and warns when later objects may reveal during an earlier topic. Treat that warning as required repair: author the semantic plan and re-derive.
> Narration-cue sync with `animations.json` requires `narration_timing.json`.
> Narration-independent custom motion instead passes `--animation-config animations.json`
> and makes no object-sync claim. Explicit `--no-animations`
> bypasses both. Without a timing sidecar, `narration_sync.py animations` maps
> groups **positionally** (group N → cue N) and warns when later objects may
> reveal during an earlier topic. Treat that warning as required repair: author
> the semantic plan and re-derive.
**Narration animation ownership**: When `animations.json` is consumed, it remains read-only. The audio stage deep-copies it to `narration_animations.json`, preserves transitions, effects, durations, order, and explicit `effect: none`, then changes only the derived trigger/delay values needed for click-free narration playback. The authored `narration_timing.json` maps each animated content group—not each effect row—to the SRT cue that speaks about that content. For `effects[]`, the cue anchors the group's first active row; later rows keep global order and their relative delay. The command may still read an affected SVG page to resolve structural group order when a sparse sidecar cannot identify every effective group; this structural fallback does not replace the semantic-context step and never edits SVG, notes, or `animations.json`. Unmatched groups keep their canonical relative delay.
**Narration animation ownership**: When narration-cue sync is selected, `animations.json` remains read-only. The audio stage deep-copies it to `narration_animations.json`, preserves transitions, effects, durations, order, and explicit `effect: none`, then changes only the derived trigger/delay values needed for click-free narration playback. The authored `narration_timing.json` maps each animated content group—not each effect row—to the SRT cue that speaks about that content. For `effects[]`, the cue anchors the group's first active row; later rows keep global order and their relative delay. The command may still read an affected SVG page to resolve structural group order when a sparse sidecar cannot identify every effective group; this structural fallback does not replace the semantic-context step and never edits SVG, notes, or `animations.json`. Unmatched groups keep their canonical relative delay.
**Title timing handoff when canonical animation exists**: preserve the title reveal decision already made by the custom-animation pass. Assign a title group to an SRT cue only when the user's request or the active motion plan explicitly chose `narration-cued`; otherwise leave its `cue` omitted in `narration_timing.json` so it keeps the canonical relative delay from `animations.json`. Do not infer `narration-cued` merely because speaker notes mention the title.
@@ -246,14 +267,27 @@ Before generation starts, `notes_to_audio.py` removes stale `audio/manifest.json
| Sidecar state | Behavior |
|---|---|
| `narration_animations.json` exists | Use it by default |
| Only canonical `animations.json` exists | Block until narration synchronization creates the derived sidecar |
| `narration_animations.json` exists and narration-cue sync is selected | Use it |
| Only canonical `animations.json` exists and narration-cue sync is selected | Block until narration synchronization creates the derived sidecar |
| Canonical `animations.json` exists and motion is narration-independent, whether or not a derived sidecar also exists | Pass `--animation-config animations.json`; do not claim object sync |
| Both are absent | Create no sidecar; inherit the base report's deck motion |
Generate passes the base report through `--inherit-motion-from`: inherited
`-a none` preserves explicit objects-off, while final Stage-2 `false` does not.
Only explicit all-motion-off uses `--no-animations`. Invalid reports block;
audio duration plus padding owns final advance.
page-start lead-in, audio duration, and page-tail padding own final advance.
**Narration pacing controls**: page-front and page-tail timing are independent,
optional parameters. Unless the user supplies values, use
`narration_start_floor=0.8` seconds and `narration_padding=0.5` seconds without
adding a confirmation question. For a destination-page transition of `T`
seconds, the post-transition lead-in is
`max(0, narration_start_floor - T)`: narration never begins during the
transition, while a longer transition is not stretched. Apply the same
lead-in to embedded narration, cue-bound object animation, subtitle offsets,
and slide advance. Uncued title or decorative animation keeps its canonical
relative timing. Setting the start floor to `0` means narration begins as soon
as the transition completes; it does not bypass the transition.
When canonical custom animation is synchronized,
`<project_path>/narration_timing.json` is the explicit semantic mapping for
@@ -273,6 +307,7 @@ python3 skills/ppt-master/scripts/narration_sync.py fingerprint <project_path>
{
"version": 1,
"srt_sha256": "<sha256 of the ordered page-local SRT set>",
"narration_start_floor": 0.8,
"narration_padding": 0.5,
"slides": {
"01_title": {
@@ -301,12 +336,12 @@ This stage keeps subtitles as external SRT files. It does not burn subtitles int
| Caller | After audio generation |
|---|---|
| Generate PPTX | With page-local SRT from Edge, ElevenLabs, MiniMax, or timestamp-capable CosyVoice and an existing `animations.json`, derive `narration_animations.json`; with no sidecar, inherit the base report's resolved motion, while explicit all-motion-off uses `--no-animations`. Export with `--recorded-narration audio`, optionally continue through `powerpoint_video.py`, then generate the delivery SRT from the finished video. |
| Generate PPTX | When narration-cue sync is selected, combine page-local SRT with `animations.json` and derive `narration_animations.json`; narration-independent custom motion passes `--animation-config animations.json`; with no sidecar, inherit the base report's resolved motion, while explicit all-motion-off uses `--no-animations`. Export with `--recorded-narration audio`; Quick also passes `--quick-generate --with-notes`. Optionally continue through `powerpoint_video.py`, then generate the delivery SRT from the finished video. |
| Enhance Native PPTX | Return to [`native-enhance-pptx`](../native-enhance-pptx.md) Step 9; its `apply` command owns audio relationships, timings, transitions, and the enhanced export. If video was selected, pass that final PPTX to `powerpoint_video.py`. |
For Qwen or explicit CosyVoice audio-only mode, embed/export the audio normally but skip `narration_timing.json`, `narration_sync.py animations`, SRT merge, and final-video subtitle alignment. Never present those missing subtitle artifacts as generated.
For Qwen or explicit CosyVoice audio-only mode, embed/export the audio normally but skip `narration_timing.json`, `narration_sync.py animations`, SRT merge, and final-video subtitle alignment. Pass canonical narration-independent custom motion explicitly when present. Never present those missing subtitle artifacts or object sync as generated.
For Generate PPTX, `--recorded-narration audio` prepares PowerPoint's recorded timings and narrations: every slide must have a matching supported audio file, every duration must be readable by `ffprobe`, and object animations must not use `--animation-trigger on-click`. Use `after-previous` or `with-previous` for narrated/video export. Narration changes the slide-advance layer only: the resolved page-transition effect remains unchanged, `-t none` remains visually transition-free, and narration advance disables click while using audio duration plus padding. The re-export is saved as `exports/<project_name>_<timestamp>_narrated.pptx`, telling it apart from silent exports.
For Generate PPTX, `--recorded-narration audio` prepares PowerPoint's recorded timings and narrations: every slide must have a matching supported audio file, every duration must be readable by `ffprobe`, and object animations must not use `--animation-trigger on-click`. Use `after-previous` or `with-previous` for narrated/video export. Narration changes the slide-advance layer only: the resolved page-transition effect remains unchanged, `-t none` remains visually transition-free, and narration advance disables click while using page-start lead-in plus audio duration plus page-tail padding. The re-export is saved as `exports/<project_name>_<timestamp>_narrated.pptx`, telling it apart from silent exports.
**Narrated SVG export**: use the default text-flow mode. It keeps authored line breaks in one editable, no-wrap text frame; narration does not require per-line text frames.
@@ -320,9 +355,11 @@ Output one summary block listing:
- For provider-timed subtitles, number of matching page-local SRT files and their location (`<project_path>/audio/*`); for Qwen or explicit CosyVoice audio-only mode, report that no page-local SRT was generated.
- Narration provider/model plus the `<project_path>/audio/manifest.json` provenance path.
- For narrated object animation, whether current SVG semantics were reused or which missing/stale pages were reread, plus semantic mapping coverage and fallback count.
- For Generate PPTX with page-local SRT and canonical custom animation, derived narration animation group count and `narration_animations.json` path; otherwise report inherited base motion or explicit all-motion-off.
- For Generate PPTX, report derived narration animation coverage/path when cue sync ran, the canonical config path for narration-independent custom motion, or inherited/all-motion-off state.
- When video export was selected, the final MP4 path and native PowerPoint export status.
- When a finished video exists, the final aligned sidecar SRT path.
- When page-local SRT was merged, the PPTX-timeline `audio/total.srt` path.
- When final-video subtitle alignment ran, the aligned delivery SRT path;
otherwise do not claim a video-aligned subtitle.
- The provider, voice, and rate/settings actually used.
- The caller-owned integration result: narrated SVG export path, enhanced native PPTX path, or “audio only”.
- For Generate PPTX when embedding was skipped, one-line hint: `python3 skills/ppt-master/scripts/svg_to_pptx.py <project_path> --recorded-narration audio`.
@@ -37,6 +37,7 @@ Verify the project's planning-session artifacts before doing anything else:
|---|---|---|
| `<project_path>/spec_lock.md` | Always | Strategist's execution anchors and routing contract; read it completely once in this fresh execution context |
| `<project_path>/design_spec.md` | Always | Complete approved design narrative and Section IX page outline; read it completely once in this fresh execution context |
| `<project_path>/notes/total.md` | Design Spec §X records a supplied final/literal narration script | Frozen verbatim narration input; read it once before SVG authoring and never reconstruct it from the planning chat |
| `<project_path>/images/` plus files whose row status requires existence | `spec_lock images` references any image | `Existing` / `Generated` / `Sourced` / `Rendered` files must exist; an absent `Needs-Manual` file remains allowed until the Step 7 readiness gate |
| `<project_path>/templates/` | `spec_lock page_layouts` references any | Layout / mirror prototypes required by execution |
| Resolver-returned Chart/Table SVG | `spec_lock page_visualizations` or legacy `page_charts` references a live Chart/Table key | Shared page-local SVG selected through the two live catalogs |
@@ -62,6 +63,7 @@ recover its relationship from §IX, or return to Step 4 when §IX is insufficien
If any required artifact is missing, report it and stop this stage. Do not enter Step 6 or invent a replacement artifact. Recover by artifact owner:
- Missing `design_spec.md` / `spec_lock.md` → use [`failure-recovery.md`](../governance/failure-recovery.md) §3.
- Missing frozen `notes/total.md` when §X declares a final/literal script → return to Generate Step 4's prepared final narration branch; never rewrite the script from memory.
- Missing `images/`, or a file whose status requires existence → recover by provenance: an `Acquire Via: user` / `Status: Existing` file is a required manual artifact, so use `failure-recovery.md` §2 and wait for the user to restore that exact file; a template-bundled bitmap returns to [`generate-pptx`](../generate-pptx.md) Step 3 to restore the selected workspace; an AI, web, formula, or slice output uses its matching row in `failure-recovery.md` §1 to reacquire, rerender, or derive it. An absent `Needs-Manual` file is not a Step 1 failure.
- Missing `templates/` inputs → restore the selected workspace through [`generate-pptx`](../generate-pptx.md) Step 3 and [`apply-template-workspace`](apply-template-workspace.md). If the workspace is unavailable or invalid, run Create Template again rather than reconstructing a template inside this stage.
@@ -76,6 +78,7 @@ Read skills/ppt-master/workflows/generate-pptx.md
Then jump to `### Step 6: Executor Phase` and run the documented pipeline:
- Read the complete project Design Spec, then the complete `spec_lock.md`, once to establish the fresh execution context
- When §X records a final/literal narration script, read the frozen `notes/total.md` once and retain its page segments through SVG authoring and the late notes validation
- Resolve the effective Speaker Notes, Custom Animations, and Narration Audio
outcomes from `design_spec.md §I`. Missing outcomes use the workflow defaults
`enabled` / `disabled` / `disabled`; these production decisions never come
@@ -58,7 +58,7 @@ The renderer (`visual_review.py`) does **not** auto-start the live-preview serve
python3 skills/ppt-master/scripts/visual_review.py <project_path>
```
This writes one PNG per page to `<project_path>/.preview/<page>.png` at 1280×720, with `<use data-icon>` inlined and `<image href>` resolved exactly as the live-preview browser sees them. Renders are serialized via a project-local file lock — safe to invoke concurrently.
This writes one PNG per page to `<project_path>/.preview/<page>.png`, sized from that SVG root's `viewBox`, with `<use data-icon>` inlined and `<image href>` resolved exactly as the live-preview browser sees them. Each successful page record in the JSON summary includes the exact canvas plus its raster dimensions. Renders are serialized via a project-local file lock — safe to invoke concurrently.
Exit codes:
@@ -69,6 +69,15 @@ Exit codes:
If any page comes back with `"all_background": true` in the JSON summary, that page rendered to a blank surface — investigate before continuing (broken `<use>` reference, missing image asset, etc.).
**Mandatory — normalize partial renders before dispatch**: parse the renderer
summary before Step 2. Dispatch only records with `"ok": true`,
`"all_background": false`, and a complete `canvas` object. For every other page,
the main agent adds a `render_failed` row directly to the aggregate with the
renderer error or blank-surface reason; no per-page `.review/<page>.json` is
expected until that page renders successfully. Exit `2` or `3` stops dispatch
entirely. Exit `4` may still review the successful subset, but the stage cannot
finish cleanly until every failed page is retried or handed off per Step 4.
---
## Step 2 — Spawn the review team
@@ -90,7 +99,7 @@ Agent(
The orchestrator prompt must be self-contained and is the **single** place where dispatch shape, batch size, and forbid lists are stated — the rubric (`references/visual-review.md`) defines the contract those prompts must satisfy. Required fields (all absolute paths):
- `<project_path>` — project root
- Full page list with `page_role` per page (parse `<project>/design_spec.md` §IX outline; **fixed compatibility default**: if an existing `design_spec.md` lacks §IX, use `content` for every page and flag this in the final report; if `design_spec.md` itself is missing, restore it through [`failure-recovery.md`](../governance/failure-recovery.md) §3 before dispatch)
- Full page list with `page_role` and the successful renderer record's `canvas` per page (parse `<project>/design_spec.md` §IX outline; **fixed compatibility default**: if an existing `design_spec.md` lacks §IX, use `content` for every page and flag this in the final report; if `design_spec.md` itself is missing, restore it through [`failure-recovery.md`](../governance/failure-recovery.md) §3 before dispatch). Pass `canvas` through verbatim; do not assume a fixed slide size.
- Batch size `K` (default 5; raise to 10 for token-sensitive runs on large decks, lower to 3 for high-fidelity short decks — see rubric §6.1)
- Iteration budget per page (default 1; 2 only for high-stakes / final-cut runs — see [Appendix: Iteration loop](#appendix-iteration-loop-opt-in))
- Path to the rubric: `skills/ppt-master/references/visual-review.md`
@@ -2,8 +2,8 @@
"sourceId": "shadcn",
"repo": "https://github.com/shadcn-ui/ui.git",
"ref": "main",
"commit": "6261bd89f72d794aea491482cc2acfd8dc3d63e2",
"commit": "deda4df80fb350230b2fce2b575e769a90cae076",
"adapter": "claude-skill",
"sourcePath": "skills/shadcn",
"syncedAt": "2026-08-08T16:00:01Z"
"syncedAt": "2026-08-10T15:59:59Z"
}