STAGE 0.7.0
Status: verified local release candidate; owner acceptance and external release operations pending
Prepared: 2026-09-19
Plugin: 0.7.0 · Method: 0.7
Reviewable Result
STAGE now uses Understand, Change, Verify, Deliver for ordinary authorized work. The agent makes a task-only local commit after relevant mechanical checks pass, while human product acceptance and external operations remain separate. Five conditional references cover gameplay, visual work, quality, continuity, and the optional Director attention board. A persistent work item has five fields.
Correctness repairs precede that simplification: read-only Git observations no longer refresh indexes or fabricate success, revision inventories have explicit first-parent/root semantics, output writers reject linked paths, and both publication workflows use the same current-main quality eligibility guard.
The agent prepares the complete release proposal, including both package lockfiles, current references, changelog, compatibility notes, and this evidence. Release Please remains draft-only. This is a local candidate, not permission to push, merge, tag, publish, or accept product judgment. The 1.0 review leaves that future decision explicitly undecided.
Compatibility And Support
- Preserve S1–S8, optional-contract semantics, plugin identity, and four skill names. Delivery is implicit; orientation, verification, and mapping are explicit.
- Preserve exact project-map schemas
0.1and0.3and case-study schemas0.1,0.3, and0.4; installing 0.7 does not create or migrate game artifacts. - Preserve command names, arguments, and successful fields. Failed orientation
now returns nonzero with
conflictand a diagnostic. Failed verification inventory returns nonzero withinconclusive, without a fabricated inventory. - A committed revision compares with its first parent; an initial commit lists every introduced path. Merge reports name the first-parent comparison.
- External refresh rejects linked output paths/ancestors and target overlap, preserves earlier external output on publication failure, and reports source changes without restoring concurrent work. Bootstrap rejects linked manifests and linked project-relative ancestors before mutation.
- Templates, instructional wording, reference layout, and optional board labels are guidance. Existing project boards can keep their labels. Optional mapping, Unity bootstrap, and research tools remain on correctness, compatibility, and maintenance support; new capability requires a named present consumer.
See ADR 0065 and artifact compatibility for the decision boundaries.
Mechanical Evidence
Evidence is dated here rather than copied as current counts into living guides.
The audit baseline was 25ce5e378a2994c61567ebaa4e1b0dc08127c550 (plugin 0.6.11).
| Packet | Evidence already completed |
|---|---|
| Preservation and observation | 90 focused tests, the change gate, the 168-test Solo Unity suite, and changed-skill/plugin validators passed |
| Publication eligibility | 18 focused tests and the change gate passed; a read-only live guard resolved the existing successful quality run 30717465161 for then-current main |
| Maintenance contracts | Structural/dependency/advisory tests and change gate passed; fresh Python environment installation and pip check passed |
| Workflow and guides | Change gate, changed-skill/plugin validators, and documentation build passed at 856e810 |
| Final candidate check | Observed result on 2026-09-19 |
|---|---|
| Change gate | Passed |
| Full release gate | Passed: 432 tests across 38 modules, supported schemas/maps/studies, Solo Unity contracts, package closure, links, documentation discovery, and diff hygiene |
| Clean Python dependencies | Fresh virtual environment installed every pinned requirement; pip check passed; full release gate ran from this environment |
| Clean Node dependencies | Root and website npm ci --ignore-scripts passed in separate empty directories; both lockfiles retain the 0.7.0 package mirrors |
| Documentation | Production build and public-artifact screen passed; 275 generated files inspected |
| Codex validators | All four skills and the plugin passed direct validators |
| Installed candidate | All 31 installed files match source; fresh implicit delivery probe passed |
| Unfinished placeholders | No unfinished implementation markers in changed active sources; command examples retain intentional caller-supplied parameters |
The local execution used Python 3.11, Node 26.9.0, and the recorded Codex CLI. Hosted CI remains configured for Node 24; this turn did not dispatch or claim a new hosted run. A final commit-bound verification log accompanies the local proposal; these results do not authorize external writes.
An isolated maintenance probe accepted an equivalent skill heading/prose rewrite
and a real resolved @commitlint/cli update from 21.2.1 to 21.2.2 without changing
the release's selected dependency. Negative probes rejected mutable action pins,
invalid permissions, broken package references, version drift, and unsupported
schema 0.2. The initial schema probe lacked a Git fixture and was inconclusive;
the corrected fixture returned the explicit unsupported-version diagnostic.
Both attempts are retained. The first rewritten-doc build exposed a removed
anchor; preserving that heading fixed the build. No failure was hidden by a
policy relaxation or retry-until-green.
| Instruction surface | 0.6.11 words | 0.7 candidate words |
|---|---|---|
| Repository agent guidance | 1,098 | 490 |
| Delivery entry | 1,199 | 862 |
| Delivery plus every conditional reference | 5,270 | 2,341 |
These are whitespace-separated editorial measurements, not acceptance quotas or performance evidence. Conditional reference files went from seven to five; the work-item starter went from 45 fields to five. S1–S8 are byte-identical to the audit baseline and no supported schema file changed.
Behavioral Evaluation
Seven paired fresh-context probes ran against baseline 0.6.11 and candidate
guidance at 856e810, using Codex CLI 0.155.1, GPT-6 Astra, and high reasoning
effort. Both conditions used the same task text and fixture per case, with
different read-only packaged guidance paths. User configuration and plugins
were disabled and host skill discovery was skipped. Each run used a new
ephemeral context and could change only its disposable fixture. The installed
probe separately used normal plugin discovery without an explicit invocation
or source-skill path in the task request.
| Case | 0.6.11 observation | Candidate observation |
|---|---|---|
| Ordinary fix with unrelated staged/dirty work | Fixed, checked, and committed; protected state retained | Fixed, checked, and committed; protected state retained |
| Diagnosis only | Supported diagnosis; no source change or commit | Supported diagnosis; no source change or commit |
| Narrow authorized repair beside uncertain authored state | Exact font repair; surrounding bytes retained; no commit | Exact font repair; surrounding bytes retained; local commit |
| Explicitly rejected implementation route | Repaired existing route; no rejected substitute or commit | Repaired existing route; no rejected substitute; local commit |
| New gameplay premise awaiting direction/play | Proposed an unapproved hypothesis; asked a product question; no implementation | Proposed an unapproved hypothesis; asked a product question; no implementation |
| Required engine check unavailable | Bounded source fix; truthful unavailable result; no commit | Bounded source fix; truthful unavailable result; no commit |
| Stale handoff versus current repository | Preserved approved multiplier and newer work; fixed zero floor; no commit | Preserved approved multiplier and newer work; fixed zero floor; local commit |
Protected file contents and relevant index entries remained intact in all 14 runs. Both conditions made safe progress and respected the tested stopping boundaries. The candidate made local checkpoints in all four eligible repair cases; the baseline did so in one. The gameplay question was necessary. No avoidable approval interruption was observed in the other cases. These are source/trace review findings for one sample per condition per case, not a statistical performance estimate. Paths, Git identities, model sampling, and real project complexity limit generalization. No paired run was repeated.
The installed 0.7.0 task selected stage-deliver-change implicitly, read its
versioned cache, fixed the production calculation, passed five tests and both
CLI checks, created a task-only commit, and preserved unrelated staged, dirty,
and untracked work. This confirms one ordinary installed invocation; it does
not establish universal routing or game-engine behavior.
All prompts, guidance snapshots, fixtures, traces, outcomes, and review judgments
for the 14 paired runs and one installed run are retained locally in
stage-0.7-behavior-probes.tar.gz (SHA-256
9236f726b407c6604811de56405f6e1d3299c4c80e3438a67507946efcaa62ff). Raw traces are
private local evidence and are not included in the generated public site.
Instructions remain model-neutral. These probes establish neither model
superiority, product quality, nor longitudinal productivity gains.
Dependency Advisory Review
The complete website dependency-graph report obtained during preparation contains 32 findings: seven high, 24 moderate, and one low. This replaces the empty production-only audit claim; development dependencies include code bundled into the browser as well as build tools. Counts alone do not establish exposure. A valid advisory report remains maintainer input under existing policy; failure to obtain a valid report fails the check. No gate, suppression, or baseline was weakened to hide findings. The final valid report and exact lockfiles are retained with candidate evidence.
The seven high-severity package findings are brace-expansion, fast-uri,
image-size, js-yaml, nanoid, serialize-javascript, and svgo. Review
updates through the existing dependency workflow, beginning with the open
documentation dependency proposal
and its affected build/browser paths. An npm fix suggestion is not by itself
an accepted upgrade or downgrade. React, React DOM, MDX, Prism, and the
Docusaurus client are browser-facing dependencies despite their placement in
devDependencies; the site also depends on build-only tooling.
Human Decisions And Follow-Through
Review the complete diff and evidence, accept or revise 0.7, then authorize any remaining push, merge, tag, public publication, and release refresh operations. The local installed-candidate probe is verification, not a public release. Use the existing maintainer process and Release Please rather than another release automation system. No open Release Please proposal existed at the read-only check; this branch supplies the complete local proposal for review. Existing dependency pull requests were inspected only as context.
For the next 8–12 real game changes, record one brief observation in the existing evaluation context: human active time, time to a fair inspection/play surface, acceptance or useful rejection, rework, escaped defects, and STAGE maintenance. This period does not block ordinary delivery or the 0.7 release. Compatibility, owner operational readiness, and transferability remain separate evidence dimensions.