Skip to content
Back to skills

taiwan-exam-generator

ASecurity

Create original Taiwan GSAT/CAP (學測/會考) exams for 國綜、國寫、英文、數A、數B、社會、自然, with separate question/solution PDFs and official scope, difficulty, originality, answer, visual and layout QA.

  • 15 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 29, 2026
ai-agentsjavascriptpythonrustjavanodeexpressrailsgitapisecurity

Works with

  • terminal
  • cli
  • api

Security analysis

A100/100

Pro scans all 20 files and shows the line behind each finding

Scanned September 29, 2026

npx -y skills add niansia/taiwan-exam --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of taiwan-exam-generator?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for taiwan-exam-generator
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/niansia-taiwan-exam-generator/badge)](https://www.skillsdirectory.com/skills/niansia-taiwan-exam-generator)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: taiwan-exam-generator
description: Create original Taiwan GSAT/CAP (學測/會考) exams for 國綜、國寫、英文、數A、數B、社會、自然, with separate question/solution PDFs and official scope, difficulty, originality, answer, visual and layout QA.
---

# Taiwan Exam Generator

**Hosted web generation:** start with
[references/hosted-execution.md](references/hosted-execution.md), the single
hosted execution route. Extract once with `read_web_knowledge.py <knowledge.md>
--subject <科目> --output-dir <versioned-refs> --reading-plan` and read
`reading/preflight.md`. Read the subject authoring, review and layout packets
when needed. Canonical inputs remain intact for validators. The long local
execution/maintenance procedure below is not an additional hosted prerequisite;
the hosted route retains all applicable subject and quality requirements.
Do not print this entire multi-subject manual, executable sources or full data
maps merely to begin one hosted paper. Use the original template attachments,
the maintained body renderer and the single hosted render/review pipeline.

Use the repository as the source of truth. The model writes new questions; the Exam Pack, metadata, learned blueprint, and validators decide what is in scope and how the paper is composed.

This skill is an editorial constraint and validation system, not a rewriting template library. For every run, the LLM must newly invent the disciplinary object, information mechanism, solution graph, representation, context, and distractor logic. Historical material supplies aggregate boundaries only.

Do not use or package any hard-coded paper builder as a content source, including a newly written one. A renderer or validator may be reused; new questions belong in this run's structured exam data, not in a reusable Python/JavaScript question bank. Never implement five forms by rotating stems, swapping constants, reusing passage frames, cycling difficulty labels, or fixing answer patterns. Writing a new batch generator is not an alternative way to follow this Skill.

## Mandatory execution contract

For **every complete paper**, including internal tests, smoke tests and stress tests, first read [references/exam-pack-execution-contract.md](references/exam-pack-execution-contract.md). Resolve the actual `exam_packs/<exam>/subjects/<subject>/` records and reference pages before writing. A profile id, a file inventory or a `verified` string is not evidence that its contents were inspected. Keep the reference/measurement pass separate from the original item-writing pass.

Source papers are protected execution dependencies, not disposable build output. A
Git ignore rule controls versioning only; it never means `歷屆試題/`, `模擬考/`,
raw intake bundles or their registries may be deleted during cleanup, packaging or
updating. Before a local full-paper run, use
`python scripts/bootstrap_exam_sources.py --subject <科目> --verify-only`; if the
selected source pack is incomplete, run the same command without `--verify-only`
and verify it again. Never overwrite a conflicting local source or claim a
source-verified paper from profiles alone. Read [references/data-ingestion.md](references/data-ingestion.md)
when acquiring, restoring, indexing, moving or publishing source material.

Complete-paper tests must use the same content and subject-layout checks as ordinary generation. Missing empirical pilot statistics may be disclosed; missing correct answers, curriculum boundaries, sources, official response modes or a usable layout may not be excused by calling the result a preview. Never silently change a full-paper request into generic practice. If a prerequisite is missing, produce a precise gap report and repair the evidence/template first; do not fabricate evidence or output a substitute batch.

When the user requests fast or timed generation, also read [references/fast-full-paper-workflow.md](references/fast-full-paper-workflow.md). Treat an under-20-minute paper as a measured warm-run performance target, not as permission to skip candidate competition, independent solving, source/rights review, subject balance, rendering, or all-page inspection. Report the actual clock boundaries and cache state. If the target is missed, finish the valid paper and report the bottleneck honestly; never relabel a partial or unchecked artifact as a complete paper.

For hosted body layout, use [references/hosted-body-workflow.md](references/hosted-body-workflow.md): reuse measured section/item components, never placeholder questions or diagram topology. Project body specs from the saved exam instead of retyping items, review each authored batch's item crops early, reserve time for final QA, and prepare both booklets' page/item review together. Retain a prior actual visual review only within the same paper: an item crop needs an unchanged authored item record and a pixel-identical rendering or the same printed glyphs, rules and images within 0.02 pt; a page needs identical pixels and item content. Changed parts require new inspection. A body renderer never replaces the original fixed PDF layers.

Before full-booklet layout, finish content/answer/difficulty review and run
`run_hosted_workflow.py lock-content --state <latest-state>`; `build` refuses to
run before the lock. Run `check-figures` once the figures exist, read each
checkpoint's `evidence_attention`, and after a content correction run
`refresh-evidence` and complete its drafts from an actual review. Check
pagination with `plan` (at most three, using `--compare`), build at most twice
and review once; every result reports its `iteration_budget`. Repair pagination through layout hints; a necessary content
correction requires renewed dependent reviews and an explicit re-lock with
`--reason`. Continue every repair from the
state returned by the last build/review, and inspect its pending `review_batches`.
Before yielding a turn, use `clock --state <latest-state> --operation pause`;
on continuation use `resume`, and record `touch` during long thinking/review
periods. These are workflow commands, not automatic chat-platform hooks. Report
inclusive wall time separately from estimated activity and recorded tool time.

Use `templates/hosted-subject-layouts.json` to select the requested subject's own question/solution layout pair. Seven subject examples are available; load only that pair, never treat the common math-oriented gallery as every subject's paper. Preview PDFs are optional; their placeholders, partial item coverage and page density cannot be used as a full exam.

In a complete local checkout, use subject renderers for body/internal proofs, compose formal booklets from the original fixed PDFs, and use the single `scripts/validate_exam_release.py` content/delivery gate. `scripts/validate_exam_pack_contract.py` is only its renderer handoff adapter, not a second validation policy. Set `metadata.run_contract` to the external run-contract.json path relative to exam.json; see the execution contract for the maintained evidence format. On a hosted web surface without the repository executable tree, follow the hosted-equivalent gate in `references/web-platform-use.md`; the absence of a local command is not a release blocker, but every observable content, answer, template and all-page inspection check remains required. A generic renderer cannot render a complete paper merely because metadata claims a verified layout. Renderer output is a **review proof**, not an accepted exam; PDF page inspection and final acceptance are separate. No helper script or export/marking test can self-certify educational quality.

Distinguish a **new paper** from a **user-requested correction of an existing paper**. In correction mode, inspect the supplied exam and audit findings, preserve unaffected items, replace confirmed failed mechanisms, and retain revision provenance. This is not a new-original-paper claim. The new-build inheritance prohibition below applies to new papers, not maintenance. After corrections, rerun whole-paper checks; metadata marking, packaging, or fixing one defect never clears unrelated failed acceptance checks.

## Route the request

For first use in a fresh installation or a new chat environment, read
[references/first-use.md](references/first-use.md). Handle ordinary-language
requests without requiring the user to operate repository commands. Installing
this folder does not install runtimes or certify reference data; inspect actual
capabilities and repair feasible prerequisites without weakening the full-paper
gates or claiming that a file upload is persistent skill installation.
For one-time installation requests, follow [INSTALL.md](INSTALL.md). On later
exam requests use the existing enabled Skill; do not reinstall or redownload it
as a routine prerequisite. Reuse accessible, still-valid reference evidence;
missing subject calibration is not a reason to reinstall the Skill.

When the Skill is used in ChatGPT, Claude.ai, Gemini or another hosted web
surface, also read [references/web-platform-use.md](references/web-platform-use.md).
For PDF production, also read [references/hosted-pdf-production.md](references/hosted-pdf-production.md).
Before expensive drafting, prove exact-template acquisition and composition in
the actual file runtime. Use the embedded fetch/composition/inspection helpers;
an optional uploaded data-only resource PDF handles runtimes without networking.
If necessary, request that one resource file before writing the paper. A full
paper request never authorizes a generic-layout draft: obtain explicit consent
before producing any downgrade; a disclaimer is not consent. Mechanical reports
and correct scores do not certify difficulty, source literacy or visual quality.
Use the platform's persistent Skill/Gem mechanism when available, not a claim
that one ordinary chat attachment installs anything globally. A complete-paper
delivery must contain two separate downloadable files: a student question PDF
and an answer-with-full-solutions PDF. Generate and visually inspect both with
the surface's file/code tools. If that surface cannot create files, execute the
required checks, or inspect every PDF page, state the exact limitation and do
not label text-only output or an unchecked PDF as the completed formal paper.
Use the same download naming convention for **all subjects and both exam types**:
`{考試}_{科目}_{paper_id}_題本.pdf` and `{考試}_{科目}_{paper_id}_詳解.pdf`.
Use the actual exam (`學測` or `會考`) and canonical subject name (學測:
`國綜/國寫/英文/數學A/數學B/社會/自然`; 會考: its subject record).
Choose one short, unique `paper_id` at run creation, such as `20260920-01`,
and preserve it across both booklets and resumed work. A new paper gets a new
ID. For example: `學測_數學A_20260920-01_題本.pdf` and
`學測_數學A_20260920-01_詳解.pdf`. Deliver the actual files using these names,
not just renamed link labels. Hosted `finalize` supplies these paths in
`delivery`; hand over those copies, not internal `question.pdf`/`solution.pdf`
proofs. Browser-added duplicate suffixes such as `(1)` are outside the Skill's
control. Honour an explicitly requested filename instead of this default.
For a current-form full paper on a hosted surface, the release-calibration and
live-spot-check procedure in that reference is mandatory. Use the embedded
verified Paper/Layout/difficulty profiles and subject references as the
hash-bound 111–115 aggregate, and use
`exam_packs/學測/metadata/official-current-web-sources.json` for time-boxed live
checks. Do not claim a timed-out PDF was opened, but do not refuse or downgrade
solely because an immutable CEEC URL times out when the compatible embedded
release calibration has no relevant unresolved fields.
Read the hosted reference's bounded-loading/continuation procedure at generation
time: load only applicable sections, fetch the requested subject's components
concurrently with verified cache reuse, and resume the same paper from saved
phase evidence instead of repeatedly reinstalling or restarting. Its embedded
source map includes each year's actual Paper Profile; preserve `needs_review`
states and reconcile specific gaps. A ready aggregate or verified layout does
not promote a pending structure to verified.

For every hosted full paper, use [references/hosted-run-evidence.md](references/hosted-run-evidence.md)
from the first item onward: save small recoverable checkpoints, review difficulty
and originality during writing, and reserve time for both PDFs. Execute the extracted `check_hosted_run.py` before formal delivery. It now
rechecks both actual PDFs against canonical fixed assets and executes the embedded
four-band/math design validators. A narrative equivalent or an author-written
pass JSON is not execution. If the helper cannot run, preserve a pending checkpoint;
do not use the missing executable as permission to skip its checks. Zero mechanically blocking pages
does not clear unresolved layout review flags or missing editorial checks.
Measure phases from preflight, review shortest routes using the available review mode
before final composition, and reserve answer rails below measured content and
figures. Final delivery requires readable item crops as well as every page;
follow [references/hosted-quality-gates.md](references/hosted-quality-gates.md).
The time benchmark never waives QA. Never manufacture missing timing or reviews.
For ALL seven GSAT subjects, prefer a real separate difficulty reviewer when the
platform provides one. Without that capability, automatically use the documented
single-context second-pass review; do not stop ordinary full-paper generation,
ask users to open a second chat, or invent a reviewer identity. Same-context review
is not independent or blind, even when its input artifact omits answers/labels.
Elsewhere, independently solving/checking an answer means a distinct reasoning
or calculation pass; it does not by itself require another model context.
Preserve an explicit user requirement for independent review; that requirement
does need a real separate reviewer. Supply actual ordered questions, continuations,
visuals and verified calibration. In single-context mode solve from the answer-free
packet before comparing saved answers, then audit shortcuts and difficulty.
Disclose the actual review mode at delivery; all substantive quality gates still
apply and unresolved errors still block. Before drafting, run
`prepare_hosted_run.py` to check embedded subject calibration, template bytes and
small question/answer layout proofs. The default offline basis is the canonical
aggregate profile plus expert review under the selected mode; original official PDFs remain
an optional, stronger item-to-item comparison when actually available. Never
claim an unseen official page was read or promote aggregate targets to achieved
difficulty. Neither difficulty nor density QA may start a new original-PDF
download dependency at final delivery. See hosted-quality-gates.md for the two
explicit evidence formats. Adopt the reviewed bands and
rebalance before PDF production. In Math A/B, assess shortcuts using options and
earlier subquestions: routine arithmetic steps and supplied intermediate results
are not new decisions. Never make a hard label fit by lengthening the solution.
Only after content review, compose both booklets and run
`verify_fixed_template_pdf.py` on their final saved bytes (also repeated inside the
delivery gate). Questions require the original subject cover and alternating
inner furniture, plus the original Math A/B final formula body. Answers require
the same subject's inner furniture on every page. “Nonofficial mock”, “116”, or
“completed in 13 minutes” does not waive any fixed-template or difficulty gate.
For Math A/B, follow [references/math-current-events-and-sourcing.md](references/math-current-events-and-sourcing.md):
use 2–4 genuinely model-dependent recent contexts in a default full paper,
verify events/results within the previous year, retain internal provenance, and
keep source-note rows/URLs off the math question and ordinary solution booklets.

For full papers, multi-form tests and Skill quality/acceptance work, first read
[references/pack-and-release-verification.md](references/pack-and-release-verification.md).
Verify actual `exam_packs` source PDFs, scored-slot structure and separate layout
evidence before trusting any `verified` flag. Preserve the requested product in
an external run contract; never substitute a generic smoke suite for independent
exam generation. `validate_exam_release.py` content AND delivery gates are
mandatory for formal full-paper claims in a complete local checkout; hosted
surfaces must execute the equivalent observable checks defined in the hosted
reference and record which mode ran. Reconcile/demote unsupported profiles;
do not repair them with confident labels. Diagnostic PDFs and provenance checks
are not exam acceptance.

For adapting, rebranding or packaging the Skill itself, read `NOTICE`,
`ORIGIN.json` and [references/attribution-and-forks.md](references/attribution-and-forks.md).
Keep current branding distinct from upstream attribution. These public guidance
and packaging checks do not add an account/approval step to ordinary exam use.
Before distributing a Skill update, also follow [references/software-release-security.md](references/software-release-security.md).
Publishing tools are source-checkout-only; their absence from an installed Skill
is intentional and does not require reinstalling it to generate exams.
Antivirus checks apply to software publication, not every exam request. Never
ask users to disable protection, restore quarantined scripts, or add exclusions.

1. Map the requested exam and subject to an existing folder under `exam_packs/`.
2. For generation, read that pack's `manifest.json`, the subject's `metadata/papers.jsonl`, `blueprints/writer-blueprint.json`, and compatible aggregate difficulty/layout profiles when they exist. Never read `metadata/questions.jsonl`, review queues, or `blueprints/learned-blueprint.json` during the writing pass. Those source-level files belong only to ingestion and analysis.
3. If the user is adding or analyzing source material, read [references/data-ingestion.md](references/data-ingestion.md).
4. For GSAT analysis or generation, read [references/gsat-subject-patterns.md](references/gsat-subject-patterns.md).
   Resolve the subject's controlling CEEC examination specification through [references/official-gsat-specifications.md](references/official-gsat-specifications.md). A third-party summary or overview image is never sufficient for detailed scope.
5. If the user asks about difficulty, questions, or a full mock exam, read [references/difficulty-calibration.md](references/difficulty-calibration.md) and [references/generation-protocol.md](references/generation-protocol.md). For current GSAT Math A/B generation, also read [references/math-difficulty-design.md](references/math-difficulty-design.md); its anti-collapse and discrimination-design gates are release-blocking. The difficulty settings are the GSAT defaults and each gate accepts a range (數學A/B 中偏難+難 70 and 難 30 points ±3, 80–92 minutes ±5; 國綜 mean 答對率 at most 0.62, 社會/自然 choice items at most 0.65, each +0.03; 英文 cross-sentence, reading and higher-vocabulary counts one item below 12/8/5). A paper inside the range is finished; do not run further review rounds to reach the target. **Only when the user asks for a different difficulty** (「數A 難題 40%」「自然平均答對率 0.5」), record `metadata.user_difficulty_request` = {`request`: the user's own words, and the numbers they gave: `hard_percent` (數學: 難 points of 100; 國綜/社會/自然: share of reviewed items with 答對率 < 0.30), `challenge_percent` and `easy_percent` (數學), `mean_p` (國綜/社會/自然)}. The review and plan gates then check those numbers (±3 points for 數學, ±0.03 答對率 and ±5 percentage points of 難 items otherwise) in place of the defaults. Never add the record, or lower difficulty, on your own initiative; for 英文 and 國寫 a request guides the design but no numeric gate changes.
   Every subject's full paper also needs a declared 簡單/中/中偏難/難 count-and-score distribution and `scripts/validate_paper_difficulty_balance.py` review. The levels must describe actual item demands, never an all-中 default, hidden column, or quota-only relabelling. Use each subject's own 111–115 evidence; transfer Math A's design-and-shortcut-review method, not its percentage mix.
6. Before generating from any official, mock, screenshot, or publisher corpus, read [references/originality-firewall.md](references/originality-firewall.md) and [references/llm-original-item-generation.md](references/llm-original-item-generation.md). Blind source separation, LLM candidate competition, and the structural skin-swap audit are release-blocking requirements.
7. For current-form GSAT mathematics, also read [references/current-gsat-math-form.md](references/current-gsat-math-form.md) and [references/current-gsat-math-scope.md](references/current-gsat-math-scope.md); their typography, formula, stem-rhetoric, visual-placement, option-geometry, machine-marking, curriculum-code, and boundary rules are hard constraints.
   When present, also read `exam_packs/學測/shared-data/current-math-form-writer-profile.json`; it is an aggregate-only form envelope. Treat common-range mock statistics as shared-unit/form evidence, never as a Math A or Math B full-paper distribution.
   For a complete Math B paper, accessible and medium-labelled items still require at least three genuine decisions; a definition lookup or one exposed formula substitution is not an acceptable entry item. The first three fill-in items must each combine a representation choice with a constraint, case, comparison, or consistency check. A latitude/longitude item may use coordinate conversion only as an intermediate step: its assessed mechanism must be spherical two-point distance, route comparison, or a navigation constraint. Every Math B item must record and pass the non-routine innovation audit described in [references/math-difficulty-design.md](references/math-difficulty-design.md); changing names, numbers, or scenery is not innovation.
   For current-form GSAT English, read [references/current-gsat-english-form.md](references/current-gsat-english-form.md). Its section sequence, vocabulary envelope, option-competition rules, passage lengths, source transformation, response formats, and page geometry are hard constraints, and its printed-form catalogue (the ten official headings with scores, `(A)`–`(D)` labels, one-row vocabulary options, stacked reading options, Chinese 說明/提示 composition prompt, `1.` `2.` translation numbers, 180–470-word passages) is enforced by `scripts/validate_english_layout_contract.py` on every surface through `scripts/hosted_subject_gates.py`. The vocabulary section must be written from contextual distinctions among plausible alternatives, not from one obvious word surrounded by wrong-part-of-speech fillers.
   English difficulty must come from discourse inference, evidence integration, semantic precision, and competitive distractors while remaining inside the CEEC reference-vocabulary boundary. Do not raise difficulty with off-list words, rare trivia, opaque syntax, or longer padding. In the ten vocabulary items, leave at least two locally plausible wrong options in most items and vary near-synonym, collocation, polysemy, argument-structure, register, semantic-prosody, word-family/form, and discourse-relation competition across the paper. Keep all alternatives grammatically usable in the slot; a visibly wrong suffix or part of speech is not useful difficulty. Follow the higher-demand floor in the English reference instead of assigning `中偏難` or `難` labels to direct one-clue sentences. In a complete 115-form paper, cloze/completion/structure must include cross-sentence decisions, reading groups must contain cross-paragraph or text–visual inference, and mixed items must require transformation rather than copying. Run `scripts/validate_english_difficulty_design.py` in addition to the lexical and layout checks.
   English novelty is also an item-level release gate, not a property of the passage topic alone. Every scored vocabulary, cloze, completion, discourse, reading, mixed, translation, and composition task must carry the `item_spec.subject_innovation_audit` defined in [references/llm-original-item-generation.md](references/llm-original-item-generation.md) and realize an English-specific semantic, discourse, evidence-transformation, translation-constraint, or writing-decision mechanism. A fresh article followed by a recoverable stock question template still fails. Apply the section-specific tests in [references/current-gsat-english-form.md](references/current-gsat-english-form.md), require the full-paper `metadata.subject_innovation_review`, and let `scripts/validate_english_difficulty_design.py` reject missing or merely declarative records.
   Every English vocabulary answer explanation must bind the answer label to the exact printed surface form. It may state the lemma or a related form only after explicitly naming the selected word; silently explaining a different derivative or near-synonym is a release failure. Student-facing English-composition directions must be in Traditional Chinese and state the minimum 120-word requirement. Reject a prompt that uses a photo as a pretext for an unrelated abstract lesson or requires facts not visible in the material; record and pass the prompt-coherence contract in the English reference.
   For the verified 115 English profile, render cloze, text completion, and discourse gaps inline inside the passage; print their option blocks once in the official order, never as repeated worksheet-style blank rows. Part I is 62 points. Section-start underlines, heading placement, and page transitions require a section-by-section visual comparison before release.
   For current-form GSAT Social Studies, read [references/current-gsat-social-form.md](references/current-gsat-social-form.md). Its discipline balance, source ecology, evidence operations, cross-disciplinary grouping, constructed-response contract, and page geometry are hard constraints. Current events may supply evidence and a real decision problem, but may not replace history, geography, or civics reasoning.
   Group the objective section's standalone items into contiguous disciplinary blocks, following the selected official year's block order and counts. Do not alternate history, geography and civics item by item. This restriction does not apply to shared-stimulus objective groups or the mixed/constructed section, including its single-choice subparts. A mixed group may integrate all three disciplines through one coherent evidence problem; prefer genuine, curriculum-bounded cross-disciplinary inference when the material supports it, without forcing every group or every subpart to involve all three. Follow the subject reference's integration and ordering checks.
   For Social Studies, at least ten scored items must rest on verified events or substantive updates within the year before the editorial lock, four of them within 180 days, spread over five materials and at least four in each part (maintainer decision 2026-09-24); print the year and month in the material. Keep the rest of the paper on the curriculum rather than turning it into a news quiz. Apply the reference's party-stance-free rules to prompts, options, images and explanations; assess evidence and syllabus concepts, never allegiance to a party or policy position. Plan answer-bearing visuals as 111–115 print them: 2–4 real photographs or archival images, 7–13 charts, maps and tables, and 18 or more items that cite a 圖/表/照片. Search the web for fitting real photographs first and use them from your own environment, recording where each came from; when the runtime cannot fetch images, use `scripts/photo_library.py` (59 traceable photographs). A traceable source is enough; no license verdict is required. User examples are candidates, not mandatory recurring topics.
   Social Studies novelty applies to every scored standalone and grouped item, not only to cross-disciplinary mixed groups or current-event items. Each item must record `item_spec.subject_innovation_audit` and introduce a new evidence configuration, source tension, spatial/temporal comparison, institutional constraint, quantitative relation, or cross-domain inference that changes the reasoning path. A different place, year, policy name, person, photograph, or dataset attached to the same textbook-definition question is a failed skin swap. Follow the per-domain audit in [references/current-gsat-social-form.md](references/current-gsat-social-form.md), require `metadata.subject_innovation_review`, and enforce it through `scripts/validate_social_item_design.py`.
   For current-form GSAT 國綜 or 自然, read [references/current-gsat-chinese-natural-form.md](references/current-gsat-chinese-natural-form.md). Its ROC 111–115 item-length and page-density envelopes, source-novelty rule, and no-unattributed-passage rule are hard constraints. 國綜 does not permit model-authored literary, classical, expository, or practical-text passages. 自然 may define a school-level model or ask students to transform source data, but every printed empirical datum and real-world claim must be traceable to a frozen source or a transparent calculation from it.
   國綜 and 自然 also require subject-level novelty beyond source novelty or visual novelty. Every scored item must carry `item_spec.subject_innovation_audit` and pass the applicable section in [references/current-gsat-chinese-natural-form.md](references/current-gsat-chinese-natural-form.md). 國綜 must create a new language/interpretive problem and evidence relation; merely selecting a previously unused author or excerpt is insufficient. 自然 must create a new model–evidence, experiment, constraint, uncertainty, multi-representation, or cross-disciplinary reasoning architecture; a new mission, organism, apparatus, photograph, graph skin, or numeric tuple around the same routine is insufficient. Require `metadata.subject_innovation_review` and enforce both subjects through `scripts/validate_chinese_natural_scope.py`.
   The printed 國綜 form is fixed across 111–115 and enforced by `scripts/validate_chinese_layout_contract.py` on every surface: the four headings with scores (第壹部分、選擇題(占76分)/一、單選題(占48分)/二、多選題(占28分)/第貳部分、混合題或非選擇題(占24分)), item 1 字音 and item 2 字形, standalone items 1–5, seven standalone multiple-choice items 25–31 with at least two language-knowledge items and no (應選n項), one mixed group 32–36, `(A)`–`(E)` labels with every option on its own line.
   國綜 has two additional editorial constraints: each independently answered short-response subpart is at most 40 Chinese characters and at most 4 points (a full short explanation is designed for 4 points); core classical selections must account for 20–25% of the whole paper's score. Apply the counting, rubric and source-dependency rules in the 國綜 section of that reference. These are not 國寫 limits, not a quota for all classical-language material, and not permission to alter historical official profiles.
   A complete Natural Science paper must place Questions 1–36 in one first part worth exactly 72 points: the printed heading is `第壹部分、選擇題(占72分)` and its boxed direction states `說明:第1題至第36題,含單選題及多選題,每題2分。` (spacing may follow the measured font, but no fact may be omitted). These 36 items are **not** all single-choice and must not be recorded as an unresolved generic choice block after the controlling paper has been reviewed. Reproduce a measured official mix: the official 111–115 booklets print 18, 15, 19, 18 and 12 multiple-choice items among Questions 1–36 (single-choice is the remainder), so a paper must carry 12–19 multiple-choice items there and record its actual counts in `metadata.natural_choice_form_contract`; every multiple-choice item prints 應選2項 or 應選3項. The mixed part numbers 37 through 56–60 as exactly six groups of 3–6 items, each with at least one constructed-response item (official bands: 3–9 single, 5–10 multiple, 8–9 constructed). Every Natural Science selected-response item uses five options labelled `A`–`E`. Every multiple-choice item anywhere in the paper—including a selected-response subpart inside the mixed section—must print `(應選 n 項)`, where `n` is derived from and checked against the independently verified answer key. The cover must print both the single-choice and partial-credit multiple-choice scoring rules; a generic `依題本說明` sentence is not an acceptable substitute. Questions 1–36 must also form four uninterrupted nine-item discipline blocks, one each for physics, chemistry, biology, and earth science; record the actual block order and do not interleave disciplines. This block rule stops at Question 36. Do **not** force the mixed/constructed part into four isolated mini-papers: one coherent shared stimulus may and often should integrate two or more of physics, chemistry, biology, and earth science. Give every subpart one primary scored domain, add valid codes for every discipline genuinely required by its solution, and record the non-ornamental evidence bridge in `metadata.natural_mixed_group_designs`. Merely mentioning a second discipline is not integration. Apply the same anti-surface rule used for Social Studies: no pure definition, named-law recall, one-step formula substitution, or decorative experiment/data material. Every item must assess a core or high-frequency curriculum anchor through at least two linked operations; medium and harder items must normally require three. Difficulty may be raised through experimental design, competing models, variable control, multi-representation evidence, uncertainty, or constraint reconciliation, never through peripheral content or calculation bulk.
   Every full Natural Science paper is current-affairs-aware by default, not only on request: read [references/current-form-topicality.md](references/current-form-topicality.md), set an `as_of_date` and editorial lock date, record `metadata.natural_source_ecology_plan` and `metadata.current_context_plan`, and meet the measured floor enforced by `scripts/validate_current_context.py` (at least five verified sources within the year carrying eight scored items in both parts, two of them within 180 days, plus Taiwan-hazard, climate/energy and Taiwan-place contexts). Choose sources freely among fresh, checkable ones — a journal paper, a conference result, a Nature/Science news item, an agency data release, a monitoring series such as ENSO or CO2, a space mission, a hazard report, local Taiwan data; at most one Nobel prize, at most two sources from one publisher, at least three source families, and one event feeds one group (maintainer decision 2026-09-24). Official 111–115 papers carry 0–5 such contexts a year with a typhoon in four of five years; a paper with none is outside the form. Give the preceding 12 months a larger share of the genuinely time-sensitive source groups, but never let a numerical news quota distort the four-discipline, difficulty, form, or solving-time balance. Sample the remaining dated sources across multiple earlier years and source families rather than clustering on one convenient year; keep evergreen school models in a separate bucket instead of pretending they are dated events. Every recent item must depend on a source-specific measurement, comparison, image feature, method, or constraint—not merely a fashionable name. Favor additional graphs, tables, maps, photographs, and observation images when they carry answer evidence; the visual floor is a minimum, not a target or permission to add decoration.
   For current-form GSAT writing, read [references/current-gsat-writing-form.md](references/current-gsat-writing-form.md) and [references/gsat-writing-source-ecology.md](references/gsat-writing-source-ecology.md). During the writing pass, do not read past question text, year-by-year topic summaries, or prior generated writing prompts. The LLM must discover new articles independently, then design a new material sequence, rhetorical tension, task decision, and title from those sources. Every printable event, case, datum, attributed viewpoint, and concrete anecdote must map to an identified source; the LLM may paraphrase, translate, juxtapose, and ask a new question, but it may not invent a supposedly real or generic case to complete the material. **Do not use United Daily News, any other newspaper, publisher, magazine, platform, archive, or institution as a default, preferred, or de facto exclusive source.** 國寫 source eligibility is publisher-neutral: any traceable and rights-safe source may compete when its transformed material fits the selected official length envelope and rhetorical role. The full-paper candidate pool must still span at least four publishers and four unrelated domains, and one easy-to-search publisher must not occupy more than half the pool. Length compliance is necessary but never substitutes for traceability, rights safety, a concrete carrier, a semantic hinge, material dependence, or role fit. For the second task, search contemporary Chinese essays, new poetry, picture-book prose and other Chinese literary writing first. A foreign work may advance only through a traceable published Chinese translation; do not expose a raw English web bibliography as the normal student-facing source line.
   In every subject the printed answer key must look like an official key, that is unpatterned: `scripts/answer_key_patterns.py` (run by the release gate, the hosted final checker and each saved batch) rejects four identical positions in a row, a period-2/3/4 cycle that continues past two repeats, an option bank keyed in label order, five answers stepping through the labels, two item groups with the same answer sequence, and label counts differing by more than one. Write the item, shuffle the options, then derive the key; a key such as 1-4-3-2 repeated or A–J in order is a release failure even when every answer is correct.
   For every full 英文, 國綜 and 國寫 paper, also apply [references/current-form-topicality.md](references/current-form-topicality.md): 英文 needs two verified recent passages carrying six items and a composition prompt tied to a verified current social trend; 國綜 needs two recent groups carrying four items and Taiwan-anchored passages; 國寫 needs one task tied to a verified current trend; 社會 needs ten items within the year, four of them within 180 days, and ten answer-bearing visuals. Record `metadata.current_context_plan` and each item's `item_spec.current_context`; `scripts/validate_current_context.py` runs in the release gate and the hosted final checker, and `append_items.py` reports progress toward the floor after every batch.
8. For every competence-oriented item or group stimulus, read [references/stimulus-generation.md](references/stimulus-generation.md).
   For corpus rechecks, current-example/literacy complaints, source-note cleanup, or release review, also read [references/evidence-backed-editorial-audit.md](references/evidence-backed-editorial-audit.md). Report coverage gaps; author-declared pass flags never substitute for a comparison. Keep full provenance internal and print only notes justified by the official form, answerability, or rights.
   If the stimulus is drawn from a dated article, event, dataset, research release, or technical update, also read [references/current-source-transformation.md](references/current-source-transformation.md). A recognizable topic name is not evidence of literacy or originality. When users request real/current-event literacy, anonymous hypothetical cases do not fulfill that request: source actual dated evidence first, then require its specific relations to enter the curriculum reasoning. Keep event, publication and page-update dates distinct; never backfill invented observations under a real institution's name.
9. If the user wants a printable paper, HTML, or PDF, also read [references/layout-fidelity.md](references/layout-fidelity.md) and [references/rendering.md](references/rendering.md).
10. If any source or generated item uses a figure, graph, map, table, photograph, or visual evidence, read [references/visual-generation.md](references/visual-generation.md).
11. For PDF export or watermark requests, read [references/pdf-provenance.md](references/pdf-provenance.md). Use non-visible metadata provenance without changing page pixels. Disclose its presence and limits; never hide instructions or promise unremovable marks. Provenance verification is separate from exam acceptance.

Official subject folders are:

- 學測: 國文、英文、數學A、數學B、社會、自然. 國綜 and 國寫 share the 國文 data folder but are separate examinations/booklets with separate timing, numbering, scoring and layout contracts. Do not merge them into one 180-minute/150-point paper. If both are requested, produce two separately validated booklets. English translation/composition, by contrast, remain inside the English booklet.
- 會考: 國文、英語、數學、社會、自然、寫作測驗. Treat reading and listening as sections of 英語.

The auxiliary `數學(舊制)` folder is historical-only. Do not merge pre-111 GSAT mathematics into Math A or Math B. The auxiliary `數學(共同範圍模考)` folder holds current-regime mock exams that cover only the common books and are not labeled A or B; use their item patterns only for shared units, never as the full-paper distribution of Math A or Math B. When a blueprint contains more than one curriculum, require an explicit curriculum and never sample across regimes.

For a current GSAT paper, use ROC years 111–115 as the primary form corpus. Official 111–115 papers and same-period publisher mocks determine the current item-writing mode: stem and option length, stimulus density, representation changes, literacy construction, reasoning depth, distractor competition, local difficulty curve, item-block height, option-row arrangement, and page density. ROC years 100–110 may expand the curriculum-content and historical item-archetype pool, but they must not set current wording, layout, literacy, or difficulty distributions. This priority is mandatory even when older files are more extractable.

When `exam_packs/學測/shared-data/historical-content-envelope.json` exists, the writing pass may read it only to broaden curriculum coverage and abstract archetype diversity. Its forbidden-use list is binding. Do not read bundle-specific reports or source-level records to obtain the same information.

Do not silently substitute a neighboring subject, curriculum regime, or exam system.

## Calibration gate

Run `python scripts/exam_data.py status` before claiming that output is calibrated. A subject and curriculum are data-calibrated only when metadata validates and `calibration_by_curriculum` is `ready` in the learned blueprint. For GSAT this requires, separately, official semantic anchors, official objective-item P/D curves, full difficulty vectors, competence/source annotations, constructed-response calibration, and a verified formal Layout Profile.

- If calibrated data exists, use its observed joint patterns for unit, item type, score, position, and difficulty.
- If only an official structure or objective-item curve exists, it may support an analysis or content plan, but not a complete-paper generation claim. Label partial diagnostics explicitly.
- If neither exists, do not invent an "official-like" distribution. Explain what data is missing, or proceed only if the user explicitly accepts an exploratory uncalibrated draft.
- Never treat a sample file or a single year as a ten-year trend.

## Full-paper structure gate

For a complete mock exam, select one compatible Paper Profile from `metadata/papers.jsonl` before planning individual items. It must specify total numbered/scored items and every major section's count, type mix, scoring rule, and order. Use a `verified` official profile for a formal GSAT claim. `auto_parsed`, `needs_review`, or `pending_ocr` records are evidence queues, not formal generation templates.

When compatible official and publisher-mock profiles both pass the gate, use `official_past_exam` for full-paper structure. Use mock exams as supplementary evidence for explanation style and item variation; never let a mock override a verified official structure for the same regime and year.

Do not average incompatible paper structures or infer a section recipe from loose per-question frequencies. If the user explicitly requests fidelity to a specific historical official administration, use that year's verified profile. A request for a 116 or later mock instead defaults to the existing compatible current-regime profile and 115 measured assets; the corpus endpoint is not an expiry date. Follow the academic-year/regime/reference-year policy in [references/official-gsat-specifications.md](references/official-gsat-specifications.md#academic-year-regime-and-reference-year). Do not require a future-year original or new template just to change the printed mock year, and do not relabel historical evidence as future-year evidence. If only an unverified reference architecture exists, stop at analysis or planning; never claim the observed counts are guaranteed for every future official administration.

## Formal layout gate

A Paper Profile is not a Layout Profile. A formal paper must also select a subject/regime-compatible Layout Profile whose `fidelity_status` and instruction transcription are both `verified`. It must reproduce the official cover hierarchy, full作答注意事項, score explanations, section labels, page geometry, running headers/footers, typeface roles, apparent type size, line pitch, page-density behavior, answer-sheet references, and subject-specific answer spaces. A generic readable renderer may be called a preview only. The observed page count is reference evidence, not a target that outranks typography or substantive content: never shrink type, narrow margins, compress line spacing, enlarge figures, add blank answer lines, or truncate material merely to force the same number of physical pages.

For a current-regime GSAT booklet using the maintained 115 reference templates, first load [references/gsat-115-template-assets.md](references/gsat-115-template-assets.md) and `exam_packs/學測/templates/115/template-pack.json`. Use its subject-specific deterministic template for the cover, signature banner, full answer instructions, scoring rules, alternating running header/footer, and (for Mathematics A/B) the correct reference-formula variant. The LLM may supply only the named dynamic fields: academic year, test name, actual current page, and actual total inner pages. It must not paraphrase, shorten, expand, or regenerate the locked cover text. Mathematics A and B are separate formula assets; Math B must not inherit Math A's angle-addition block. Render body content first, obtain the real inner-page total, and only then fill page furniture. Never force body text into the reference year's page count, and never treat `blank-template.pdf` as a fixed-page exam skeleton. These assets are layout-only and must not become a reusable question generator, question bank, or batch-content source.

On a hosted web surface, also load `exam_packs/學測/templates/115/hosted-web-template-assets.json`. Every persistent or packaged projection must retain all 30 distinct per-file `download_url` records for all seven subjects together with each file's SHA-256, byte count and page count. Installation stores rules and URLs but downloads no template PDF binaries. At generation time, acquire only the subject's three production components, or four for Mathematics. The embedded `scripts/fetch_hosted_template_assets.py` supports raw download, GitHub Contents API/base64, and an uploaded data-only template resource PDF. Verify bytes inside the composing runtime; a base64 response in another tool alone is not successful handoff. Use `scripts/compose_hosted_pdf.py` with the mapped overlay geometry and transparent body-only pages. Keep fixed cover text, fractions, grids, rules, signature and formula body as original PDF layers. No OCR, retyping, HTML/Word conversion, screenshots or visual imitation of locked material. Never download or return `github-pages.zip`, a repository archive or an all-template ZIP. Persisting PDF bytes is optional, not an installation gate; apply a natively saved Skill immediately in its creation chat and later selectable chats. If bounded network transport fails, use or request the single offline resource PDF according to `references/hosted-pdf-production.md`. If required capability is still absent, report it before drafting. Do not generate or deliver any generic-layout substitute unless the user explicitly accepts that downgrade first; labelling a draft does not authorize it. Run the mechanical inspection helper on both final saved PDFs, then actually review every page and every content gate. None of the helper reports certifies formal acceptance.

## Regression gates learned from full-paper review

These checks apply to ordinary generation, previews, smoke tests, timed tests, and regenerated papers. A test paper is not allowed to omit them merely because its purpose is to test speed or a renderer.

- **Final answer positions:** this rule applies to every full paper containing single-choice items—國綜、英文、數學 A、數學 B、社會、自然;國寫 has no such population, and multiple-selection items are excluded from this count. Audit the final printed order, not the drafting order. For each population with one stable option-label set and at least twice as many items as labels, build the position plan from a near-even multiset and then shuffle it: when the item count is divisible by the label count the totals must be exactly equal; otherwise the largest and smallest totals may differ by at most one, and every label must appear. Use a fresh per-paper shuffle rather than a fixed A-B-C-D rotation, then reject four identical answers in succession and any period-2 to period-4 cycle repeated three times. “Random” here means balanced and pattern-screened, not unconstrained randomness that can create a visible cluster. For a small section that cannot meet the whole-paper arithmetic exactly, avoid an omitted label or conspicuous concentration, while the whole-paper gate still controls. Multiple-selection answers require a separate inclusion-frequency and set-size audit so one option position is not systematically absent or selected. Never change the truth of an item to fill a quota: author and solve the item first, permute already valid options, then remap and re-solve every dependent record.
- **Printed-option binding:** after any option move, update the key, independent `derived_answer`, exact option verdicts, explanation labels/surface forms, lexical or distractor records, question and answer hashes, and every displayed answer table. Rerender both the student and explanation PDFs. A balanced count in stale metadata is a failure.
- **Formal page furniture:** the selected Layout Profile alone controls the alternating running header, actual current/total inner-page count, signature strip, subject/year wording, footer number, section heading, instruction box, question-number hanging column, option indentation, score placement, and answer-space form. Do not invent a house header, hard-code the reference paper's page total, print internal audit/provenance text in the student booklet, or add response lines because space remains.
- **Template-first rendering:** hosted full papers use the existing verified PDFs and `scripts/compose_hosted_pdf.py`, not template-source reconstruction. The same original-PDF rule applies in a complete local checkout. HTML template modules are internal component proofs, never a substitute for fixed PDF layers in a delivered booklet. `scripts/render_gsat_template_assets.py --build-all` is a maintainer asset-build/preview command, not a hosted substitute for unavailable binaries. A production paper uses its selected measured Layout Profile and fills year/test/page fields after actual pagination. Any change to locked instructions, signature wording, formula membership, field positions, or type roles requires a new template version, regression tests, regenerated assets, and fresh all-page visual review. A previously rendered PDF is stale after such a change.
- **Mathematics fraction geometry:** every inline fraction, fill-format fraction, fixed denominator, scoring fraction, and reference-sheet fraction must be one nonbreaking semantic/geometry unit. At final PDF size, verify the numerator is centred above exactly one fraction bar, the denominator is centred below it, neither level collides with neighbouring prose, and no numerator, bar, denominator, sign, slot circle, or answer-row label is clipped or split across lines. A fixed denominator must not draw a second underline beneath itself. Text extraction that happens to contain the right digits is not a visual pass.
- **Response-format integrity:** if the same-role official page uses a bordered inference or completion table, encode its caption, heading, row meanings, slots and dimensions as structured student-facing content and bind it into the content hash. A semantic table may not be replaced by generic horizontal lines, and a table may not be invented solely to occupy white space.
- **Substantive density:** compare complete printed passages, item blocks, figures, and occupied page roles at the verified typeface role, apparent size, line pitch, and margins. In hosted mode, a metric-compatible Traditional-Chinese fallback permitted by `web-platform-use.md` may satisfy the role even when its internal family name differs. Passing page count or container height does not excuse short materials, oversized gaps, a sparse terminal page, repeated pseudo-tables, or decorative figures. Repair content and page assignment before release; never stretch, shrink, or pad.
- **Difficulty and scope:** reject definition lookup, one-clue recognition, one-step substitution, topical-name decoration, peripheral syllabus trivia, and options where only one is remotely plausible. Every stated difficulty label must be supported by the actual shortest solution route. Social Studies and Natural Science additionally require core/high-frequency curriculum anchors and at least two linked evidence operations for every scored item, with medium/hard items normally requiring three. English difficulty must come from in-scope semantic/discourse competition, not rare vocabulary. For 國綜、國寫、英文、社會 and 自然, apply the construction floors in [references/current-form-literacy-load.md](references/current-form-literacy-load.md) before drafting; they are the non-mathematics equivalent of the mathematics difficulty-design gate and are measured from official ROC 111–115 papers.
- **Reading load and shared stimulus:** 國綜、國寫、英文、社會 and 自然 are 素養 papers whose items mostly sit in 題組 that share one substantial printed stimulus, **including inside 第壹部分**. Assembling a paper from standalone one- or two-sentence scenarios is a structural failure, not a style choice: it is the direct cause of both the short-paper and easy-paper defects. Meet the per-subject group counts, group-stimulus lengths and whole-paper substantive-text floors in `references/current-form-literacy-load.md`, measured against `exam_packs/學測/shared-data/current-form-literacy-envelope.json`. A 自然 paper must additionally place at least two of its six 第貳部分 題組 genuinely across 物理/化學/生物/地科 (CEEC declares exactly two 合科 groups in every official 111–115 paper), recorded in `metadata.natural_mixed_group_designs`; a second discipline that can be deleted without changing the solution does not count. These are floors taken below the weakest official year, never targets, and never a reason to pad prose, answer space or figures.
- **Post-render evidence:** run all content gates again, rasterize every student page, and compare the cover, first content page, every section transition, every page containing a large figure/table, and the final two content pages at readable scale against the matching official page roles. Any renderer change invalidates the earlier visual pass until both student and explanation PDFs are regenerated and rechecked.

## Required generation sequence

1. Resolve exam, subject, curriculum, year/regime, full-paper or custom-practice mode, difficulty target, simulated exam date, exclusions, answer profile, and output format. Use pack defaults only when those defaults are present in data.
2. In full-paper mode, lock a verified Paper Profile and Layout Profile. Copy the official section counts, order, scores, duration, instructions, and page contract into the plan. Then create a new full-paper distribution inside the multi-year aggregate envelope; do not inherit a previous generated paper or copy one historical year's unit sequence. A visual item must include a Visual Spec; `requires_diagram: true` without one is invalid. Source analysis must end in an abstraction artifact; the item writer must not receive source stems, numeric tuples, equations in source order, source-question ids, or source-figure topology.
   For current GSAT generation, every Item Spec must cite a 111–115 aggregate form-pattern cluster built from multiple compatible administrations. A pre-111 record may influence only aggregate curriculum coverage and historical archetype diversity; it is never an item seed.
   For a 20-item Mathematics B paper, allocate slots through the five-strand blueprint and narrow-family caps in `references/current-gsat-math-scope.md` before choosing surface contexts. Never call unconstrained model sampling a balanced distribution. Maintain the rolling three-form Math B rotation ledger: one-point perspective and latitude/longitude or sphere-coordinate conversion must each appear at least once in every three successive forms, while each single form still preserves broad coverage instead of mechanically repeating both topics.
3. Attach the official difficulty target for every compatible objective slot and the separate official/pilot rubric calibration for every constructed-response slot. Preserve section resets, position curves, annual exceptions, and the distinction between target and achieved P. For current mathematics, add a `difficulty_design` record to every scored item before writing prose. It must name the irreducible linked decisions, representation changes or constraint checks, misconception paths, discrimination move, and the result of a shortcut-collapse audit. Counting algebra lines or naming several syllabus concepts is not evidence of difficulty.
   When timed reviewer feedback exists for an earlier generated paper, preserve it as a separate feedback-calibration record. Use it to adjust the next paper's actual reasoning architecture, not its official target labels. A complete Math A paper must include a non-routine `D-10-3` item meeting the counting-depth gate in `references/math-difficulty-design.md`; direct factorial multiplication or a single symmetry halving step cannot be the paper's sole combinatorics coverage.
4. Build a dated inspiration pool for competence items from sources available before the simulated editorial lock. Match source family and lead-time distributions; never use a current event merely because it feels topical. Validate the pool with `scripts/validate_inspiration_pool.py`. Record corpus coverage separately: a file being hashed, indexed, or visually scanned does not mean its questions were semantically read. When OCR or item annotation is incomplete, disclose the exact gap and do not claim the model has learned every supplied question.
   For 國綜 and 自然, screen every proposed source title, author-work pair, distinctive quoted phrase, dataset name, and canonical URL against the full supplied official/mock corpus before drafting. Run `scripts/audit_source_novelty.py source-registry.json output-report.json`. A zero lexical hit is only a triage pass: also review aliases, translations, alternate titles, excerpt boundaries, and whether the same source object appeared through another representation. Store the result in the source registry; do not expose source-paper wording to the writing pass.
5. Write original questions that satisfy each item spec. Test the intended concepts, not superficial number changes. Require a new information mechanism, solution graph, and representation semantics before drafting the surface context. Generate at least three mutually dissimilar mechanism candidates for every scored item; do not bind a curriculum unit to a user example or recurring domain. Render answer-bearing diagrams from new semantic data; use image generation only for non-exact visual context and follow the visual-generation reference.
   `Original` describes the assessment design, not the authorship of the reading material. Do not invent a passage, poem, diary, archival record, interview, public notice, research result, or allegedly real dataset and then print it as source material. Select a traceable published source that has passed the corpus-novelty screen; quote only when rights permit, otherwise make an accurate attributed adaptation with an audit map back to source propositions. Do not print labels such as `自擬`, `本卷自擬`, or `數值為自擬` in a student paper. If a task genuinely requires simulated data, it may be used only when the user explicitly allows it and the paper is labelled as a custom exercise rather than a source-grounded full mock.
   For competence-oriented items, the stimulus must be required evidence rather than decoration. Apply the stimulus-removal test from `references/stimulus-generation.md`; if the item remains answerable with the same reasoning after removing the stimulus, rewrite it or label it a basic item.
   Also apply the source-relation test: anonymizing names is allowed, but removing the source-derived relation, representation, constraint, or decision must change the reasoning. A topical name alone is ornamental. See `references/evidence-backed-editorial-audit.md`; current material must supply a mathematical or rhetorical affordance, not prestige vocabulary.
   Apply bounded novelty: unfamiliar professional, technological, daily-life, or university-adjacent mechanisms are allowed only when every external rule is defined in the item, no outside domain knowledge is required, and the complete solution reduces to named concepts in the selected curriculum. Novel context must not inflate difficulty beyond the slot target.
   Do not hard-code a topic-to-unit association. Spaceflight, language models, telescopes, public bicycles, energy systems, or any user example are candidates only. For each slot, compare at least three mutually dissimilar information mechanisms, including a non-topical alternative, and select by evidence necessity, curriculum fit, solution quality, and target difficulty.
   Apply the originality firewall to every scored item without exception: text-only, single-choice, multiple-selection, fill-in, constructed response, and every mixed-group subpart. For fill-ins, only the official marking rail may be reused. For mixed groups, the shared object and the dependency among subparts must also be newly invented.
   A photograph must carry an auditable provenance record: original, public-domain, licensed, user-supplied, or found on the web with a traceable source (`web_sourced`; no license verdict is required, maintainer decision 2026-09-24). When it will print in grayscale, freeze the exact crop and tonal conversion before item review. Every answer-bearing feature must remain distinguishable at final print size without hue. A color-dependent prompt is rejected unless the same distinction is redundantly encoded by labels, patterns, shapes, positions, or printed values. Never ask students to identify a color from a grayscale reproduction.
   Before solving, perform a substantive-surface audit against the selected current official booklet. Compare unique stimulus volume, complete item-block volume, source/representation mix, and the amount of genuinely occupied page area. Short standalone scenarios are allowed only where the matching current-form role is also short; a full paper may not be assembled from one- or two-sentence mini-scenarios plus generic options. A curriculum code, a declared two-step operation, or a filled metadata field does not compensate for missing evidence in the printed material.
6. Solve every question in a separate reasoning pass. For high-risk mathematics, science, ambiguous reading, or constructed response, use an independent second route or deterministic calculation where practical. The printable answer material must preserve each item's actual calculation, evidence comparison, model boundary, or scoring points. Boilerplate such as “符合題示資料與模型”, “排除超出資料支持範圍的敘述”, or the same generic two-line rationale repeated across items is not a solution and blocks release. For selected response, state the decisive evidence and at least the main distractor distinction; for constructed response, show the computation or auditable rubric elements actually used to award points.
7. Validate scope twice: first map every mathematical operation to one or more official learning-content codes, then audit the actual wording and solution path for hidden out-of-scope terminology, theorems, or procedures. Run `scripts/validate_math_curriculum.py` for Math A/B and require human review of every `defined-bridge` item. Then validate answerability, unique answer where applicable, distractors, units, diagrams, answer distribution, duplicated concepts, total score, difficulty-vector match, stimulus necessity, and blueprint fit. The answer-distribution audit runs on the **final printed option order**: reject a conspicuous omitted label, two-label concentration, mechanical pattern, or unexplained long run. Reordering options requires remapping the key, independent solve, every option verdict and explanation label, followed by fresh content/answer hashes and a complete rerender. A near-even count is an editorial default, never permission to alter which statement is true. Run lexical triage plus the mandatory structural skin-swap audit from `references/originality-firewall.md` against official, mock, and already-generated items. Also validate the paper-level diversity matrix from `references/llm-original-item-generation.md`. Replace only failed items and recheck the whole paper. Any build that inherits legacy generated question content is rejected in full.
   For 國綜 and 自然, validate every item against the controlling CEEC examination specification at the learning-performance/content-code level. The natural-science paper must visibly balance physics, chemistry, biology, earth science, and inquiry/practice across both major parts; naming a discipline in metadata is not a scope audit. Run `scripts/validate_chinese_natural_scope.py generated-exam.json --report output/scope-report.json`, then run `scripts/validate_source_grounding.py` with the frozen source registry and passing novelty report before rendering. A made-up, misspelled, or wrong-subject curriculum code is release-blocking even when the prose seems on topic.
   For source-bearing items, verify that printed facts agree with the frozen source snapshot, invented values are explicitly labelled as simplified or simulated, URLs and dates remain in the audit registry rather than cluttering the formal paper, and removing the source-derived relation changes the solution. For 國寫, verify that each supplied passage has a distinct rhetorical job, every material paragraph has a `material_source_map`, every printed case comes from an identified source, and no source record has invented modelling values. The second task must include an authored Chinese literary source, or a traceable published Chinese translation, suitable for affective expression rather than defaulting to institutional explainers or newly found English-language web essays. On the student page, source attribution stays in full-width parentheses at the end of the material paragraph; it is not a separate source block. Do not title the supplied passages `材料一:` or `材料二:`; use the official-style `甲`/`乙` markers only when multiple texts need labels. Run `scripts/validate_writing_source_grounding.py` and treat any failure as release-blocking.
   For 國綜、國寫、英文、社會 and 自然, run `scripts/validate_literacy_load.py generated-exam.json --subject <科目> --report output/literacy-load.json` on the authored exam **before** layout. It rejects a paper that prints materially less than an official booklet, that carries too few shared-stimulus 題組, whose group stimuli are too short, or whose 自然 mixed groups never leave one discipline. Repair a failure by restoring genuine source material and complete item blocks; never by enlarging type, inflating figures, widening answer space or appending filler. A structural pass is not an editorial pass.
   For mathematics, run `scripts/validate_math_difficulty_design.py` and reject any item whose planned difficulty collapses to direct substitution, one familiar formula, one routine linear-system solve, or repeated execution of the same operation. Treat model-estimated discrimination only as a design label until representative pilot data exist. Then validate explicit `option_layout` per item (`row-5`, `row-4`, `grid-3-2`, `grid-2`, or `stack`). Mathematics 111-115 print options five abreast or one per line; the renderer prints 數學A/B options that overrun the chosen tab one per line rather than wrapping them inside a cell, so choose `stack` for sentence-length options. Never infer the printed arrangement only from character count. Validate the rendered option baselines and inter-option whitespace against recent reference pages.
   Every mathematics fill-in item must declare a machine-marking `answer_format`: integer/decimal/fraction/sign form, exact numerator and denominator slot counts, and the sequential answer-row ids printed in the booklet. The renderer must show the same circle/rail/fraction geometry that the student will mark; a generic blank line is a hard failure.
   A current mathematics paper must also meet its learned answer-bearing visual quota across more than one section. Until item-level 111–115 visual annotation is complete, the internal review floor is four required visuals or structured graphical representations across at least three sections, including at least one item before the mixed-response section. Decorative pictures do not count.
   The same stability requirement applies to full current Mathematics B, Natural Science, Social Studies, and English papers; it is not a Mathematics A-only feature. Until each subject/year has complete item-level annotation, use these conservative internal release floors: Mathematics A/B `4 visuals / 3 sections / 2 kinds`; Natural Science `16 / 2 / 4`, spanning physics, chemistry, biology, and earth science, including at least **1 traceable real photograph or observation image** (official 111–115 print 16–28 labelled figures a year, almost all drawn graphs, apparatus and tables, and 0–2 photographs), every figure and table captioned 「圖N/表N」 and cited in its stem; Social Studies `10 / 2 / 4`, spanning history, geography, and civics, including at least **2 traceable real photographs or archival images** and **18 items that cite a 圖/表/照片** (official 111–115: 2–4 photographs or archival images and 7–13 charts, maps and tables a year, 18–45 figure-citing items); English `3 / 2 / 2`, including at least **1 traceable real photograph** and the noncontinuous mixed material or selected composition form. The photo numbers are lower bounds only: Natural Science and Social Studies have **no photo-count upper bound**. Do not stop at two merely because the validator floor has passed; retain every additional sourced photo that introduces a distinct, answer-bearing observation and improves section, source-family, or discipline spread without harming rights safety, grayscale survival, page density, or solving time. Equally, never add a decorative image merely to increase the count. Generated photorealism, screenshots, decorative pictures, and photographs whose decisive evidence disappears in grayscale do not satisfy these real-image floors. These are product floors, not claimed CEEC frequencies. A selected annotated profile may raise the minimum, never silently lower it or impose an arbitrary maximum.
   Label every figure and table in Chinese outside 英文 (樣品、硫酸根、電解液、反應進程、時間), as the official booklets do; keep only symbols, units, formulas and acronyms (x, t (s), mol, NaCl, DNA, NOAA) in Latin letters. `validate_visual_item_contract.py` rejects English words in `semantic_data` labels, SVG text and PDF figures; it cannot read a raster, so check a PNG's labels when reviewing the page.
   Every counted visual must be evidence or required for solution, set `item_spec.requires_diagram: true`, include a schema-complete `visual_asset.visual_spec`, enumerate answer-bearing features, and fail the visual-removal test. Run `scripts/validate_visual_item_contract.py generated-exam.json`. If a full paper misses its subject envelope, replace the failed or text-only item with a newly designed visual item and re-solve it; never attach a decorative image to preserve an old stem or enlarge a figure to fill the page.
   The declared visual kind must match the **rendered scientific topology**, not merely its metadata label. A coordinate graph needs axes, scales and plotted marks; a profile or cross-section needs spatial layers/paths; a spectrum needs a wavelength axis and spectral lines; an apparatus or circuit needs connected components; a flowchart needs meaningful nodes and directed links. A one-column box that restates prompt values is not a graph, map, profile, spectrum, apparatus, process diagram, or evidence matrix. Do not route heterogeneous visual kinds through one generic label-and-row panel. A genuine data table must have an explicit row/column comparison structure used by the solution; a vertical list of already printed facts is not a data table. If removing the figure leaves every number and relation needed for the answer in the prose, remove the redundant figure or rewrite the item so the figure actually carries evidence. Before release, record `representation_audit` with the topology family, rendered primitive types, semantic channels, prompt-redundancy result, and visual-removal result; reject repeated near-identical panel topology masquerading as representation diversity.
   A sourced photograph must preserve the downloaded original beside the placed derivative and hash-bind both files. Record the stable source page, creator or agency, the license when one is stated, original-file path and hash, fixed crop, processing steps, target print width, and effective raster resolution. Convert the placed asset to a fixed grayscale or bilevel file before pagination; do not rely on printer conversion. The final-size review must locate every answer-bearing boundary, object, count, texture, label, or relative tone in the monochrome page without consulting the color original.
   Full Social Studies papers must keep history, geography, and civics genuinely balanced in both item count and score, and full Natural Science papers must do the same for physics, chemistry, biology, and earth science. Until an annotated selected-year profile supplies tighter values, the internal fail-closed limits are a largest-to-smallest item-count gap of at most `3` and a score-share gap of at most `8` percentage points. `validate_social_item_design.py` and `validate_chinese_natural_scope.py` enforce these limits; merely naming every discipline once is not balance.
   For English, run `scripts/validate_english_vocabulary_scope.py generated-exam.json <CEEC-reference-vocabulary.pdf> --report output/english-vocabulary-scope.json`, `scripts/validate_english_difficulty_design.py generated-exam.json`, and `scripts/validate_english_layout_contract.py generated-exam.json`. Treat an out-of-envelope non-reading word, a vocabulary target above level 5, a missing exact surface-form binding between option/answer/explanation, fewer than two plausible distractors for most vocabulary items, a narrow distractor-family mix, more than one simple vocabulary anchor, fewer than five medium-hard/hard vocabulary items, a one-cue or suffix-only shortcut disguised by a hard label, a non-Chinese composition direction, an incoherent forced writing prompt, or a failed section-display contract as release-blocking. Use near-synonym, collocation, polysemy, argument-structure, register, semantic-prosody, discourse-relation, and controlled word-form/word-family competition across the section; a form-based distractor counts only when it remains syntactically plausible and is defeated by full-sentence evidence. Level 6 or off-list words in authentic reading material require local support or a documented reading-only exception; rarity may not be the intended source of difficulty.
   For Social Studies, run `scripts/validate_social_item_design.py generated-exam.json --report output/social-item-design.json`. `curriculum_codes` must contain exact Grade 10–11 required **learning-content** codes from `references/social-required-content-codes.json`; keep learning-performance codes in a separate field and CEEC `H/G/C/S` assessment targets in `ceec_assessment_targets`. Treat a performance code masquerading as content, an elective/invented/wrong-domain code, a missing curriculum-alignment record, unsupported current-event claim, ornamental proper noun, non-self-contained domain rule, evidence-free competence label, bare definition recall, or peripheral low-priority curriculum target as release-blocking. Every scored item—including basic anchors—must record a core or high-frequency curriculum anchor, forbid recall-only solution paths, and require at least two linked operations grounded in evidence, relations, causes, constraints, scale, or procedure. Preserve identifiable history, geography, and civics coverage alongside cross-disciplinary groups instead of making the whole paper a collection of topical news passages.
   For a full current Social Studies paper, meet the within-one-year floors in `references/current-gsat-social-form.md` (10 items, 4 within 180 days, 5 materials, 4 in each part) and its three-subject 題組 floor (at least one in 第壹部分 and two in 第貳部分 whose items lead with 歷史, 地理 and 公民 in turn); an event name, date, or fashionable technology used only as decoration does not count. Run `scripts/validate_social_layout_contract.py` too: it enforces the printed form measured on 111–115 (equal-length options, the key the longest option in at most four items, 「26-27 為題組」, 「(3 分,35 字內)」, the two part headings, no printed authoring notes or blank answer boxes).
   For every answer-bearing photograph or raster figure, require `grayscale_evidence_survival` and `color_independence` in its Visual Spec and inspect the final rasterized page at print scale. Schema compliance alone is not a visual pass. Newness means a new information mechanism, solution graph, semantic data, and visual topology—not an unusual topic, grayscale filter, crop, rotation, relabeling, or artistic restyle. Precise diagrams, axes, boundaries, scales, labels, state regions, maps, and measured values must be deterministic; image models may supply only non-exact context beneath a deterministic answer-bearing overlay.
   Keep photo rights and credit in the internal provenance ledger. Do not print optional photo-source lines in the student booklet unless the selected official Layout Profile explicitly includes them; textual material/source lines remain subject to their own official-form rule. If a license requires visible attribution incompatible with the form, replace the source or obtain suitable permission; internal metadata never substitutes for required visible credit.
8. Keep student paper and answer material separate. Produce structured `exam.json`; for fixed-page proofs, render HTML and pass `scripts/validate_fixed_page_html.py` before PDF export. The gate must cover horizontal and vertical overflow plus containment inside bordered instruction boxes, tables, response examples, and the printable page frame. Then rasterize the PDF and inspect every page at readable scale before a formal claim. Page count, extracted-text density, font inventory, or a contact sheet alone cannot establish that nothing is clipped. A minimum-height box is not printed content: compare actual text/figure endings and distributed working space with the selected reference. Verify function powers and log bases by their actual raised/lowered positions, not merely smaller digits or valid-looking tags. After shared renderer changes, regenerate affected PDFs and withdraw stale visual passes until the new pages are reviewed. For current Math A/B mixed-response items, do not print workbook-style answer lines in the question booklet unless the controlling official Layout Profile explicitly contains them.
   For 國綜 and 自然, also compare every rendered content page with the ROC 111–115 density envelope by running `scripts/validate_current_form_density.py`. In this project page density is one fixed rule (`scripts/hosted_density.py`, shared by plan, inspector and final check): a body page leaves at most 32% of the printable body blank (英文 42%), the last body page at most 60%, and nothing waives a page over the limit. The candidate's complete-paper extractable text volume must also reach at least 80% of the controlling official paper (or the lowest recent official total when no year controls). This is an anti-padding floor, not a target or permission to copy. CSS minimum height, enlarged spacing, blank answer rules, decorative labels, and figure bounding boxes do not count as substantive material. Do not repair under-filled pages by shrinking the page count, enlarging type, or adding decorative filler; rebalance substantive source excerpts, item blocks, reference-supported answer space, exact figures, and page transitions while preserving readability. Do not auto-append `來源補充`, add answer-revealing commentary, inflate figures, or over-allocate response lines to pass this metric. Apply the role-specific typography and SVG text/stroke checks in `references/current-gsat-chinese-natural-form.md`; preserve failed full-form status when a corrected proof exposes inadequate substantive length. A matching total page count with half-empty pages is not layout fidelity.
   For every current-form subject, run the applicable PDF density comparison after export: `validate_reference_page_density.py` for a selected official booklet and `validate_current_form_density.py` for the maintained 國綜/自然 multi-year envelope. Compare matching page roles and require at least 80% of the official reference's substantive text volume unless a reviewed answer-bearing visual replaces prose. A page that passes overflow checks but ends materially earlier than its official counterpart, or a paper whose containers reach the bottom while its text volume is short, is still a release failure. Repair it by restoring authentic source passages, necessary context, complete item blocks, noncontinuous evidence, or natural page flow. Page-count difference alone is not a release failure; record it and review the resulting page roles. For 國寫, explicitly compare each printed source packet's prose length and rhetorical function with the corresponding official task; a short prompt plus oversized response area is not a full-form writing task. Never create one large terminal void, and never disguise one by shrinking or inflating type, changing verified margins/line pitch, enlarging answer space or figures, or adding irrelevant prose. Small distributed spacing between complete items is acceptable only after the substantive content-length contract is satisfied.

The output must say which pack, subject, curriculum/regime, blueprint fingerprint, and calibration level it used.

## Originality and private data

Historical files in `歷屆試題/` and `模擬考/` are persistent calibration data.
They may be omitted from the ordinary Git tree or Skill archive for size and
distribution reasons, but must remain recoverable through the checked source-pack
manifest and bootstrap workflow. Never delete, move or replace them merely because
they are ignored, absent from a public clone, or excluded by a packager. A cleanup
request aimed at generated output does not include source papers, source registries,
raw intake bundles, templates or calibration evidence. Destructive changes to those
exact resources require an explicit request naming them.

- Learn abstractions: unit, concept, skill, stimulus class, item type, difficulty vector, score, position, visual function, and common misconception. Keep those abstractions in a separate artifact from source wording and figure topology.
- Do not copy, lightly paraphrase, translate, or merely swap numbers in a source item.
- Do not publish source passages, images, answer keys, or full historical questions unless the user confirms the necessary rights.
- If similarity is uncertain, change the information mechanism, solution graph, representation semantics, and surface context, then validate again. Surface changes alone never cure a structural match.

## Repository tools

Use the repository helpers when local execution is available (some require PDF/browser dependencies):

```text
python scripts/audit_exam_pack.py --output output/pack-audit.json
python scripts/validate_exam_release.py exam.json --contract run-contract.json --stage content --output output/content-gate.json
python scripts/validate_exam_release.py exam.json --contract run-contract.json --stage delivery --output output/delivery-gate.json
python scripts/audit_generated_suite.py form1.json form2.json --output output/suite-audit.json
python scripts/exam_data.py index-sources
python scripts/build_paper_profiles.py
python scripts/analyze_pdf_visuals.py
python scripts/build_visual_queue.py
python scripts/build_question_candidates.py --promote
python scripts/download_ceec_gsat_statistics.py --min-roc-year 100
python scripts/import_ceec_gsat_difficulty.py
python scripts/build_gsat_difficulty_profiles.py
python scripts/analyze_gsat_official_patterns.py
python scripts/analyze_mock_exam_dataset.py --intake <本次資料夾>
python scripts/analyze_mock_bundle.py --bundle "<正式套卷名稱>"
python scripts/analyze_historical_content.py
python scripts/build_layout_review_queue.py
python scripts/analyze_stimulus_ecology.py --exam 學測 --subject 社會 --output output/social-stimulus-ecology.json
python scripts/analyze_writing_source_corpus.py <資料夾...> --output output/writing-source-corpus.json
python scripts/validate_inspiration_pool.py output/current-inspiration-pool.json
python scripts/audit_item_originality.py generated-exam.json output/originality-audit.json
python scripts/validate_llm_originality_contract.py generated-exam.json
python scripts/validate_math_curriculum.py generated-exam.json
python scripts/validate_math_difficulty_design.py generated-exam.json
python scripts/validate_english_vocabulary_scope.py generated-exam.json <CEEC-reference-vocabulary.pdf> --report output/english-vocabulary-scope.json
python scripts/validate_social_item_design.py generated-exam.json --report output/social-item-design.json
python scripts/validate_chinese_natural_scope.py generated-exam.json --report output/chinese-natural-scope.json
python scripts/validate_source_grounding.py generated-exam.json source-registry.json --novelty-report source-novelty.json --report output/source-grounding.json
python scripts/validate_current_form_density.py generated-student.pdf --subject 國綜 --report output/form-density.json
python scripts/exam_data.py validate
python scripts/exam_data.py build-blueprints
python scripts/exam_data.py plan --exam 學測 --subject 數學A --count 10
python scripts/exam_data.py plan --exam 學測 --subject 數學A --full-paper --paper-year 2025
python scripts/render_visual.py templates/visual-spec.json output/example-graph.svg
python scripts/render_exam.py examples/synthetic-exam.json output/exam.html
python scripts/render_pdf.py examples/synthetic-exam.json output/exam.pdf
python scripts/validate_fixed_page_html.py output/paper.html
```

`render_pdf.py` uses a locally installed Chrome, Chromium, or Edge and preserves the CSS A4 page size. If no supported browser or local execution is available, deliver the print-ready HTML or follow the schemas manually. Never claim a validator ran when it did not.

Use `templates/llm-originality-record.json` for every scored item and `templates/paper-originality-matrix.json` for the paper-level and mixed-group records. These are audit forms only; their placeholder values are not evidence of a pass.

Files in this skill

  • AGENTS.md1.8 KB
  • CONTRIBUTING.md1.6 KB
  • INSTALL.md14.7 KB
  • NOTICE1 KB
  • ORIGIN.json423 B
  • SKILL.md88.1 KB
  • SOFTWARE_RELEASE_STATUS.json1.8 KB
  • agents/openai.yaml441 B
  • core/taxonomy.json696 B
  • core/visual-taxonomy.json1.3 KB
  • docs/.nojekyll1 B
  • docs/brand-assets.md4.8 KB
  • docs/brand-research.md2.1 KB
  • docs/chinese-form-audit-2026-09-22.md4.8 KB
  • docs/download-web-knowledge.html9.2 KB
  • docs/english-form-audit-2026-09-22.md5.3 KB
  • docs/fixed-template-web-regression-2026-09-13.md5.3 KB
  • docs/github-collaboration-plan.md1.9 KB
  • docs/hosted-claude-run-2026-09-19.md4.5 KB
  • docs/hosted-efficiency-2026-09-17.md9.8 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…