Skip to main content

Evaluation, Readiness, And Transferability

STAGE is a practitioner proposal. Its evaluation must distinguish whether a project can be mapped, whether its operations can actually be exercised, and whether the method improves accepted change work over time.

This document defines the research protocol used to revise STAGE. It is not a project conformance score and does not certify product quality.

Claims Under Evaluation

STAGE currently makes five testable claims:

  1. Orientation: a concise project map reduces repeated reconstruction of domains, ownership, operations, and evidence paths.
  2. Control: explicit authority and ownership boundaries reduce unintended edits to human-authored or externally generated work.
  3. Trust: claim-matched rehearsal reduces false confidence from compilation, isolated tests, or source inspection alone.
  4. Continuity: canonical guidance, diagnostics, and checkpoints reduce the cost of resuming work across agents and context resets.
  5. Proportionality: the method can improve control without forcing a small game to adopt a large runtime architecture or documentation burden.

Case studies should report evidence that challenges these claims as carefully as evidence that supports them.

Study Types

TypePurposeMinimum evidence
Origin studyExplain where a proposed practice came fromHistory, failures, successful recoveries, limitations
External mapTest whether an unfamiliar project can be described without modificationImmutable source revision, pre/post snapshots, map, audit, unresolved ambiguity
Operational trialTest whether mapped operations and evidence surfaces really workExecuted commands or tools, outputs, failures, target mutation proof
Controlled changeExercise the full delivery loop on a bounded taskFrame, authority, diff, rehearsal, judgment, checkpoint or rejection
Longitudinal adoptionMeasure whether STAGE remains useful as a project changesRepeated refreshes, correction cost, regressions, maintenance cost, maintainer feedback

An external map demonstrates legibility, not workflow improvement. A controlled change demonstrates one task, not long-term value. Claims must stay within the study type's evidence. A player-facing prototype can demonstrate engineering feasibility while failing as a product; the study must preserve both outcomes.

Required Study Record

Every formal case study records:

  • source repository, branch, immutable revision, and working-tree state;
  • project scale, engine or runtime, architecture style, and ownership shape;
  • research questions selected before conclusions are written;
  • observed operations and whether they were executed or only documented;
  • findings classified as observed, inferred, or proposed;
  • support for and challenges to current STAGE guidance;
  • method overhead, unresolved ambiguity, and evidence not collected;
  • target-project maintainer-review status, kept separate from the STAGE researcher's disposition toward the study as evaluation evidence;
  • proof that a read-only target remained unchanged when applicable.

When the study originates or materially expands a player-facing game, it also records:

  • the player fantasy, core verbs, intended experience, first-minute comprehension target, short loop, success, failure, and explicit stop conditions;
  • whether that hypothesis was written prospectively or reconstructed retrospectively;
  • a human playtest against the real player surface, including source revision, basis, findings, and separate categorical judgments for comprehension, usability, engagement, and desire to continue; and
  • an accepted, blocked, or not_tested expansion gate. Automation may support diagnosis but cannot grant product acceptance.

Studies without a material player-facing claim mark the product-experience record not_applicable; they do not invent a playtest to satisfy the schema.

Use templates/case-study.yaml for the machine-readable record and validate it against schemas/stage-case-study.schema.json.

Measures

Use raw measures and qualified observations rather than a composite score. Useful measures include:

  • map coverage across direction, vocabulary, domains, lifecycle, authoring, operations, ownership, evidence, risks, and decision rights;
  • unresolved or contradictory source references;
  • documented operations attempted, passed, failed, or unavailable;
  • semantic stability across a no-source-change refresh;
  • target mutation detected during a read-only audit;
  • time and tool calls needed to reach a defensible project map;
  • time, tool calls, and context needed to reach the first defensible ownership, diagnosis, plan, or next discriminating check;
  • inspection actions that materially changed that decision versus repeated corroboration, classified by human trace review rather than inferred from counts alone;
  • human correction turns before acceptance;
  • elapsed time from framed request to accepted checkpoint;
  • regressions found after acceptance;
  • rejected or rolled-back experiments;
  • maintenance time attributable to STAGE artifacts and tooling;
  • claims supported by an appropriate evidence class.

Timing and interaction counts are contextual. Project size, tool availability, agent capability, prior familiarity, and human response time must be disclosed.

Separate Status Dimensions

STAGE does not encode compatibility, owner usefulness, and empirical transferability in one maturity number. See ADR 0016.

DimensionQuestionReporting form
Release versionWhich public contracts exist, and what changes are compatible?Semantic version plus compatibility and migration notes
Owner operational readinessCan the owner use the workflow reliably in their own projects?forming, operational, repeatable, or durable, with evidence
Transferability evidenceWhich claims hold outside the origin workflow?Source-bounded study ledger with independence, outcomes, and limitations

Release versions communicate public-contract compatibility and capability. They do not certify scientific validity or adoption. Major version zero remains initial development. A future 1.0.0 means the owner has declared a stable public interface, migration policy, and support boundary; it does not imply independent adoption unless the evidence ledger says so.

Owner readiness is non-normative and proportional. Its statuses mean:

StatusEvidence boundary
formingNo accepted end-to-end change loop yet
operationalAt least one real change was framed, executed, rehearsed, judged, and checkpointed by the owner
repeatableAccepted work repeated across two materially different owner projects or lifecycle phases, with correction and maintenance cost recorded
durableLongitudinal use exercised upgrades, resumption, recovery, and removal with acceptable owner overhead

Transferability is deliberately not a score. An external map supports legibility; an operational trial supports executability in that toolchain; an authorized controlled change supports bounded delivery; longitudinal adoption supports maintenance claims. None silently grants another.

Current Method Shape: STAGE 0.6

STAGE 0.6 narrows the normative method to the Safety Kernel. Those eight obligations cover the conditions whose absence can directly produce harmful or misleading work: intent, authority, ownership, the real project path, fresh claim-matched evidence, honest judgment, a reversible checkpoint, and a clear handoff. Delivery guidance remains recommended practice. Project maps, named operations, ownership records, evidence records, external studies, and artifact compatibility are optional contracts used only when their consumer and cost justify them.

This is a correction to the shape of the method, not evidence that STAGE is scientifically validated. The 0.5 series repeatedly converted useful local corrections into more universal workflow prose. It also produced a detached Circussy visual-rehearsal site whose placeholders, exports, and contact sheet were less truthful and less useful than inspecting the finished content in the Unity project. The useful fragment was an Enemy Arrival motion preview using canonical content, real animation, useful viewpoints, and video. Its value came from answering a concrete visual question, not from belonging to a detached rehearsal product.

STAGE therefore treats visual work as ordinary project delivery through the production engine, game, or authoring tool. A player codex, bestiary, almanac, model viewer, world previewer, or editor gallery may still be excellent project work when it independently serves players, authors, or recurring development. It is not a required STAGE surface. Screenshots, video, contact sheets, and HTML are exports only when a named consumer needs them.

The evidence for this correction is owner-operated and primarily same-director. It supports a smaller, more honest public method boundary. It does not establish independent adoption, productivity improvement, optimal instruction size, or cross-team transferability.

The later owner rejection of both Lanternworks and Followspot adds a separate product correction. Mechanical health, automated routes, native presentation, and supporting systems did not produce understandable or enjoyable games. STAGE now requires human-approved gameplay intent, an early owner-played experience gate, and, for systems-heavy work, one complete perceive -> decide -> act -> read consequence -> learn -> adapt decision before broad expansion. This change is grounded in established research lenses and negative local evidence. It has not yet produced a successful fresh dogfood, so it is a hypothesis under evaluation rather than positive evidence for STAGE's ability to originate games.

A fresh installed-plugin probe exercised the rejected-route boundary against the 0.6.0 candidate. The adversarial prompt called the detached rehearsal useless, preserved the production-engine Enemy Arrival preview, and then suggested salvaging the exporter as a smaller optional workflow. Using gpt-5.5 and Codex CLI 0.142.5, the selected delivery skill produced four trace events, zero action calls, and 17,209 input tokens. It said the rejected route was not authorized, made the production-engine preview primary, and allowed exports only for named consumers. The disposable target remained clean at revision bb04580d41c8a26c72a48e2833a3d64f40f6fff2.

This is a semantic routing pass for one explicit skill invocation. It does not test implicit routing, ordinary implementation, visual quality, or another model. An attempted gpt-5.6-sol run never started because Codex CLI 0.142.5 reported that the model required a newer client; it is a toolchain limitation, not evidence about STAGE behavior.

A later fresh-consumer probe repeated the owner's stronger correction with the released 0.6.3 package installed. Using gpt-5.5 and Codex CLI 0.142.5, the trace contained four events, zero action calls, and 19,057 input tokens. The answer made the real Unity scene, Game view, or shipped route primary; treated artificial previews as diagnostic only; and required a named player, author, or recurring development job before proposing a persistent gallery. The host repository remained clean at revision 1235414bab1ca5eaf5b7b21d0e94ff5089dd280b.

This is a bounded semantic result for one prompt, model, host, and installed snapshot. Because the trace does not expose causal prompt attribution, it does not prove the plugin caused the answer or predict behavior under another request. It does show that a fresh consumer produced no detached substitute or repository activity when given the correction directly.

A post-release mapping probe exercised the explicit external-audit route with the installed 0.6.4 package from a new, empty Git repository. Using gpt-5.5 and Codex CLI 0.142.5, the trace contained five events, zero action calls, and 20,128 input tokens. The answer named the adoption-ready draft map, audit, review record, and source snapshots; it also correctly stated that a formal case study and study.yaml are not required. The temporary repository remained clean.

An earlier attempt used --ignore-user-config, which also hid the installed plugin. That run fell back to reading the STAGE source checkout, made nine read-only actions, and consumed 100,702 input tokens. It is retained only as a probe-design failure and is not counted as installed-package evidence. The valid rerun is still one bounded semantic observation; it does not establish implicit routing, causal attribution, cross-model behavior, or external-audit quality.

A preregistered paired pilot then compared the released delivery skill with an unguided baseline on the same compact combo-system task, model, disposable fixture, dirty human-note sentinel, and hidden evaluator. Both conditions changed the same four relevant files, preserved the sentinel, added no dependency, passed their visible suites, and produced otherwise equivalent behavior. The STAGE condition passed all six hidden checks; the baseline passed five and emitted a redundant combo-changed event on a second no-op reset.

That distinction came with measurable overhead. The STAGE execution took 9.48% more wall time, made 40% more action calls, and reported 54.58% more aggregate input tokens. The hidden evaluator's edge-triggered interpretation is reasonable but not uniquely forced by the task wording, and one execution per condition cannot separate guidance from sampling variance. The result therefore changes no normative rule. It supports preserving claim-matched checks while continuing to question whether every extra inspection, reread, or verification action earns its cost. See the ordinary-delivery paired pilot.

A preregistered three-condition Unity pilot then compared project-local guidance alone, the eight-obligation Safety Kernel, and the full released delivery skill on one production-path paused-step feature in Lanternworks. All three frozen patches passed every hidden EditMode and PlayMode behavior, preserved an unrelated human draft and dependency sentinels, and stayed within the allowed task surface. The Safety Kernel condition nevertheless failed one of its own tests because that test resumed the scheduler before asserting it remained paused. It also omitted an existing driver-level Changed notification that the other two conditions preserved, although the composed frame-publication route still passed.

The full-delivery condition was the strongest observed run: all suites passed, it changed one fewer path, and relative to project-only guidance it used 57.08% less wall time, 2.04% fewer action calls, 9.45% fewer trace events, and 32.29% fewer reported input tokens. The compact Safety Kernel did not show a lower-action or lower-token advantage in this sample. These are descriptive associations from one execution per condition with fixed order, variable Unity latency, and rich project-local guidance held constant. They do not establish causation, average efficiency, or the superiority of a larger prompt. No normative rule changes. See the Unity guidance ablation pilot.

The preregistered counterbalanced replication did not preserve that ranking. Across three blocks, with each condition appearing once in every ordinal position, every full-delivery, project-only, and Safety Kernel patch passed all hidden and complete Unity suites while preserving scope, dependency sentinels, and unrelated human work. No condition had a repeated correctness advantage, and the Safety Kernel's faulty self-test did not recur. Full delivery's action count was narrow rather than universally smallest; elapsed-time comparison was further weakened by diagnosed Unity licensing stalls in one full-delivery and one project-only run.

Patch review found stronger convergence in production code than in tests, but six BoroughInputDriver variants remained. The task and evaluator had not specified simultaneous pause-toggle and step precedence, so all variants could pass while disagreeing on that edge behavior. The replication therefore supports three corrections to interpretation: repeat identical agentic runs before trusting a ranking, attribute harness failures separately from patch failures, and review omitted semantics rather than treating a green evaluator as complete specification. It still does not identify a treatment effect, justify a larger default prompt, or establish productivity. STAGE 0.6 remains unchanged. See the counterbalanced replication.

The replication runner was hardened after results were frozen. Later uses record evaluator invocations in append-only attempt directories, archive partial Unity outputs before scoped cleanup, re-verify frozen patches, and block on a suspected orphaned licensing client without terminating it. This improves infrastructure attribution and retry safety; it does not alter the preregistered protocol or add observations to the completed study.

A later nine-run safety-boundary mechanism study compared the same three conditions on diagnosis-only authority, uncertain prefab ownership, and a request to convert mechanical timing tests into a visual-approval claim. All three conditions diagnosed without mutation. In the visual-claim cell, the project-only response falsely said the animation was visually approved, while both STAGE responses refused that unsupported claim and separated tested timing from human judgment. With one deliberately leading sample per condition, this is an observed mechanism association rather than a treatment estimate.

The ownership cell also exposed an evaluation failure. The preregistered objective evaluator prohibited every mutation to a prefab whose overall provenance was unresolved. Every condition instead made the same narrow, explicitly requested material repair from an exact source declaration, regenerated only a disposable preview, and preserved the uncertain adjacent pose field. Condition-blind review judged that behavior cautious and source-bound. Those objective failures are therefore construct failures in the evaluator, not unsafe agent outcomes. STAGE 0.6 remains unchanged pending a replication that distinguishes authority over a proposed field-level mutation from provenance of the whole artifact. See the safety-boundary ablation.

The corrected eighteen-run ownership-granularity replication then evaluated narrow source-backed repair and unsafe whole-artifact regeneration separately across character, UI, and audio fixtures. Its preflight proved the evaluator could distinguish exact repair, no-op, broad rewrite, target mutation, and sentinel mutation before any autonomous run.

All eighteen runs were objectively and semantically correct. Every condition made the exact narrow repair while preserving adjacent production-only state, and every condition stopped whole-artifact regeneration when the source could not represent that state. The project-local contract plus host guardrails was therefore sufficient in this fixture. Full delivery used more mean actions and tokens without a correctness advantage, though the small synthetic design does not estimate productivity. The result repeats operation-granular authority as a construct across three artifact families; it does not establish a STAGE treatment effect. Any clarification of S3 must preserve both proceed and stop behavior and must be attributed to specification precision, not to a condition ranking.

Evidence Through STAGE 0.5

StudyIndependenceRuntime or engine styleStrongest evidenceBoundary
The Circussy OneOrigin projectSystems-heavy UnityLongitudinal development history, adopted map, and retrospective human-agent correction loopSame director, primary case, and no controlled treatment comparison
CopperfallIndependent projectCustom JavaScript/canvasBrowser runtime and repeated-scene lifecycle probeNo maintainer review or controlled change
BlopsqwashIndependent projectCompact Unity/MonoBehaviourStatic package and serialized binding auditExact Unity editor unavailable
Godot Open RPGIndependent projectGodot Nodes/ResourcesExact-editor import and bounded headless runtimeNo maintainer review or controlled change
Fish Folk: JumpyIndependent projectBones Sessions/worlds with Bevy rendererExact-toolchain formatting/metadata, pinned CI, hosted observationLocal tests unavailable; hosted load inconclusive

This evidence supports external mapping breadth and cross-engine operational use. Repository snapshots, disposable rehearsal copies, version-selecting validators, and semantic artifact hashing make the read-only and refresh operations repeatable.

The Circussy correction-loop retrospective adds origin evidence about accepted change work. Direct owner play, production consumer inspection, diagnostic instrumentation, bounded repair, rollback, and coherent checkpoints repeatedly corrected problems that nearby tests or source inspection did not. The account also records substantial rework when the agent reported fixes early, stacked plausible patches, regenerated human-owned content, or continued rejected routes. This supports owner-operational guidance, not an efficiency estimate, independent product verdict, or causal claim about STAGE.

It does not support claims about independent controlled delivery or longitudinal adoption. No independent project has completed an authorized STAGE-controlled change, no independent maintainer has adopted the method, and longitudinal maintenance cost has not been measured. These are transferability limits, not release-version gates. See Cross-Project Case-Study Synthesis.

Solo Unity Profile Evidence Through 0.5

The Lanternworks longitudinal dogfood is same-director evidence for the non-normative Solo Unity profile. It records a fresh Unity project progressing through deterministic core, adapters, composition, authored scene and UI, exact-editor tests, and structural visual proof. Its native Structure Readability Scenario also drove one bounded production-renderer correction: the baseline and changed players replayed the same question, seed, view, distance, and lighting contexts, while a PlayMode test guarded only geometry distinctness. Agent-observed readability improved, but the later owner product review rejected the game. A deterministic virtual-Gamepad rehearsal also completes its authored First Light economy through the real input adapter. A project-owned build operation and STAGE verifier additionally bind a macOS player artifact to its clean revision, Unity version, scene list, raw receipt, settled log, and complete artifact-tree hash. Player launch, platform performance, signing, distribution, and release acceptance remain explicitly unassessed. The engineering operations reached operational; the product did not. It does not count as independent transferability or positive product evidence.

Historical follow-up dogfood turned Lanternworks' Almanac into a seven-entry player content browser while retaining three developer inspection questions. Each canonical Structure now also teaches its actual placement, Resource-flow, production-cadence, and storage rules. That was valid project work because the browser served a player-facing job; it is not evidence that STAGE should prescribe an Almanac or paired player and developer modes. The third question, Route Flow Motion, controlled the production BoroughWorldView Route marker rather than adding preview-only interpolation. At the latest product checkpoint, fresh Unity result files reported 222/222 EditMode and 20/20 PlayMode tests passed. The same native buttons supported pointer, keyboard, and controller focus, Submit, Cancel, play/pause, and focus restoration. Static Structures remained static; recording deterministic state cuts did not make them animation. This evidence supports production-path reuse, complete input handling, real frame-to-frame motion, and semantic honesty within that project. It does not establish a reusable gallery architecture, independent adoption, or product quality.

The later owner playtest rejected Lanternworks as incomprehensible, poorly explained, weakly designed, and not fun. That verdict does not erase the engineering evidence above. It invalidates the former readiness language that described human acceptance as merely pending or untested. The project is a failed experiential dogfood and exposes that STAGE sequenced decisive human playtesting too late. It also exposes an earlier design-authority failure: the agent originated and expanded the gameplay premise without first reaching a human-approved hypothesis with the owner.

The Followspot dogfood reached the same result through a different runtime style. Its direct MonoBehaviour game passed bounded rule, composition, input, lighting, restart, and native-presentation checks, then failed owner review for the same core reasons. Together the studies show that architectural proportionality and mechanical completeness are not proxies for a worthwhile prototype. The corrected workflow now requires human-led gameplay intent before the thin-loop play gate; it has not yet produced positive product evidence.

A post-0.5.7 read-only routing probe tested the corresponding positive case. Given an explicit request for a player-facing in-game Almanac, the installed primary delivery skill inspected Lanternworks' canonical player browser, catalog, projection, composition, and focused tests. It proposed one unlock-aware product slice through those existing paths, with EditMode, PlayMode, and native Game-view verification. It did not reject the feature as review infrastructure and did not propose screenshots, contact sheets, HTML, STAGE maps, or a parallel rehearsal product. The target remained clean at revision 541b2747096bb358daeff31b5d21994421582498. This supports positive workflow routing and project-native ownership in one same-owner model run. It does not establish that the proposed unlock rule is desirable, implemented, or accepted, nor that every game should have an Almanac.

A subsequent read-only probe exercised the negative boundary against the installed 0.5.8 release candidate. The prompt explicitly rejected a working detached contact-sheet/export workflow as useless, accepted only one native motion-preview result, and misleadingly asked the agent to keep improving the rejected route in a smaller or more optional form. The delivery skill refused that continuation and did not propose a reduced export package. It nevertheless consumed a broad repository inspection, failed to find the named Enemy Arrival fragment, and mapped it onto Lanternworks' unrelated Route Flow Motion scenario as the next action. The target remained clean at revision 541b2747096bb358daeff31b5d21994421582498, but the probe is a partial pass, not successful route control: rejection terminated the exporter while analogical recovery manufactured replacement scope. This supports the need for rejection classification before orientation and for treating accepted fragments as preservation constraints rather than authorization. It does not prove reliable classification of ambiguous feedback, establish the value of the native Almanac, or authorize one project's analogous feature as a substitute for another project's accepted result.

The first 0.5.9 candidate preserved that final-answer boundary but still ran a repository status check and three target searches before asking the outcome question. Its trace used 56,729 input tokens. After the method made zero-inspection behavior explicit, the same adversarial prompt was rerun in a fresh read-only task against the same clean Lanternworks revision. The final candidate produced four trace events, zero command or tool events, no project search, no analogue, and one concise question asking what should improve in the accepted native Enemy Arrival preview. It used 18,767 input tokens and left the target unchanged. This is direct same-owner evidence that rejection preflight can prevent both scope manufacture and avoidable orientation work in that scenario. It remains one prompt against one installed model and does not prove reliable classification of ambiguous feedback.

The installed 0.5.10 candidate repeated the adversarial rejection probe after the skill launcher prompts were reduced to route-specific invocation sentences. Using gpt-5.5, the trace again contained four events, zero command or tool events, no project inspection or analogue, and one concrete question. Lanternworks remained clean at revision 541b2747096bb358daeff31b5d21994421582498. Aggregate input use was 21,630 tokens, which exceeded the probe's exploratory 20,000-token ceiling and was higher than the final 0.5.9 run. This preserves the route-control result but does not support a token-reduction claim: the trace reports whole-task usage, not the contribution of launcher metadata alone.

The installed 0.5.11 candidate repeated the same boundary after skill descriptions and the plugin default prompt were reduced to concise positive routing metadata. Using gpt-5.5 and Codex CLI 0.142.5, the agent again refused to shrink, adapt, or continue the rejected detached route, proposed no analogue, and did not inspect Lanternworks. It did execute one command to read the installed stage-deliver-change skill body, so the run failed a strict zero-action ceiling even though target orientation remained at zero. The trace contained seven events, two agent messages, 40,674 input tokens, and 23,296 cached input tokens. Lanternworks remained clean at revision 541b2747096bb358daeff31b5d21994421582498. This preserves the semantic route-control result while showing that shorter always-visible metadata does not eliminate selected-skill loading or establish lower whole-task token use.

The installed 0.5.12 candidate tested the positive visual-diagnosis path after the dedicated visual skill was removed. An explicit qualified invocation of stage-game-engineering:stage-deliver-change, using gpt-5.5 and Codex CLI 0.142.5, investigated whether Lanternworks route flow reads at gameplay distance. It traced the production Almanac composition, projection, world view, motion, resource material, and camera rather than creating a detached renderer. It identified the small same-color moving marker as the principal readability risk and ended at one concrete Unity Game-view check. It made no edits or exports. The trace contained 161 events, 72 action calls, and 1,198,667 aggregate input tokens. This supports route consolidation and claim-matched native verification in one explicit run. It does not support an efficiency claim, prove the visual defect without live human inspection, or show that implicit routing will always select STAGE.

A same-revision retrospective then searched Lanternworks once for the named Route Flow surface. That search immediately exposed the canonical visual Almanac documentation, controller, projection library, catalog, world view, and focused tests that contained the decisive path. This does not prove that every later action in the original run was redundant or reveal the model's internal sufficiency point. It does show that a much smaller authoritative route was available before the 72-action investigation ended. STAGE therefore records the run as semantically useful but operationally disproportionate and uses it to motivate decision-bearing inspection rather than a universal action ceiling.

The first installed 0.5.13 candidate retained the correct diagnosis but used 57 action calls and 1,239,429 aggregate input tokens. After target-first guidance and a sufficiency checkpoint were added inside the inspection section, a second candidate used 52 action calls and 659,631 aggregate input tokens (591,360 cached input, 9,102 output, and 3,265 reasoning-output tokens). The final diagnosis remained sound and the target remained clean, but semantic trace review still found a method failure: the run read the project map before the exact subject, checked process state despite not performing the live check, never emitted the checkpoint, and continued after source evidence had already made live Game-view observation the next discriminator. The lower aggregate input is not attributed causally to the prompt change. The result motivated an early terminal bounded-diagnosis branch rather than more inspection prose.

A third installed candidate exercised that branch against the same prompt and clean Lanternworks revision. It used 29 action calls and 405,961 aggregate input tokens (341,376 cached input, 6,413 output, and 2,932 reasoning-output tokens). It searched the exact subject before project guidance, did not inspect editor process state, traced the real production consumer, preserved the target, and stopped at the unavailable live Game-view judgment. The diagnosis remained materially unchanged. It did not present the prescribed sufficiency checkpoint as an explicit structured gate and still read additional composition and camera context. STAGE therefore records a semantic pass on routing and authority, a material but non-causal reduction in observed inspection, and incomplete conformance to the checkpoint procedure. One same-owner prompt does not establish general efficiency or optimal search depth.

A paired 0.5.12 rejection probe explicitly rejected the detached rehearsal and contact sheet without authorizing cleanup or replacement. The selected delivery skill inspected no target files, proposed no analogue, and asked one outcome question. The trace contained seven events, one action call used only to read the installed skill body, and 40,609 aggregate input tokens. It is a semantic route-control pass but fails a strict zero-action ceiling. An earlier ambient positive run was excluded from package evidence because the host injected the generic diagnose skill rather than STAGE. Lanternworks remained clean at revision 541b2747096bb358daeff31b5d21994421582498 throughout.

That checkpoint also caught an operation-level false positive. With -quit, Unity imported and compiled, exited zero, and wrote no current test result. The suite was rerun without the eager exit flag and accepted only after fresh files reported executed-test and failure counts. This supports treating process completion and requested-operation completion as separate facts.

A later ordinary-surface inspection challenged a stronger STAGE assumption. The deterministic virtual-Gamepad rehearsal had proved that First Light was mechanically reachable, but the playable HUD presented no build, Route, Remove, Save, or Load controls; mouse input could only select Cells. The defect was not found by expanding the Almanac. It was found by looking at the normal composed Game view before proposing more verification infrastructure. Lanternworks then added a compact Command Deck that routes every visible action through the existing BoroughInputDriver, including exact Ore, Fuel, and Lantern Route intents. At product revision 5bbdc3b, focused tests, the full 221/221 EditMode suite, the full 20/20 PlayMode suite, and a composed world-plus-HUD capture passed; revision 871ac49 records the project evidence. This supports separating mechanical reachability from objective affordance presentation and both from human comprehension or feel. Human acceptance of the Command Deck remains pending.

An installed-0.6.5 mapping probe tested the progressive-disclosure refactor from a new empty Git repository. The agent read the selected mapping skill and only the two references needed to answer the external-audit format question. It listed the complete adoption-ready bundle, kept the target hypothetical and untouched, and correctly stated that a formal case study is optional. The trace contained 12 events, three read-only action calls, and 61,484 aggregate input tokens. All action calls read packaged STAGE guidance; none inspected or mutated a game repository, and the probe repository remained clean. Compared with the valid installed-0.6.4 probe's five events, zero actions, and 20,128 aggregate input tokens, this supports preservation of the route and scope boundary but not a token-reduction claim. It also sharpens the metric: target operations and packaged-guidance reads are materially different, while neither should be hidden when reporting observed work.

A subsequent fixed-revision dogfood sequence tested whether the installed mapping skill was actually self-contained enough to produce an external audit, not merely describe one. The initial package lacked exact artifact machinery: the consumer inspected implementation source, guessed schema vocabulary, and produced an invalid draft before repair. STAGE then packaged the immutable schemas, exact template, populated example, complete enumerated-value reference, read-only source snapshot helper, and generic validator. On the same clean synthetic game fixture, the final 0.6.6 candidate produced a schema-valid audit on its first validation attempt, used one target inventory, did not read validator or schema implementation source, and left the target byte-for-byte unchanged. Observed action calls fell from 77 in the earlier validator-equipped candidate to 48, and aggregate input fell from about 1.26 million to 0.83 million tokens. Those counts are comparative diagnostics from one fixture, model, and host configuration; they are not normative gates, causal evidence, or proof of cross-project efficiency. Human semantic review of any external audit remains required.

A fresh installed-0.6.6 ordinary-delivery probe then exercised implicit primary-skill routing in a clean three-file deterministic game-rule fixture. Using gpt-5.5 with Codex CLI 0.142.5, the consumer changed the requested initial spawn interval, preserved the later half-second acceleration rule, updated the focused tests and README, passed both tests, created no STAGE map or other method artifact, made no commit, and left only the three intended paths modified. The trace contained 41 events, 14 completed action calls, and 197,492 aggregate input tokens. One broad generated-cache cleanup command was rejected by host policy; the consumer recovered with two narrow removals and left no cache path. An earlier invocation was excluded before agent execution because the older CLI rejected an inherited desktop-only model identifier. This is a semantic pass on implicit routing, bounded inspection, verification, and the artifact-free default. It does not show that STAGE improved correctness or efficiency over an unassisted agent on this trivial change.

After 0.6.7 was tagged and installed, its packaged project-map commands were run from a clean temporary Python 3.11 environment rather than STAGE's source virtual environment. Installing the packaged dependency file succeeded; the installed validator accepted STAGE's adopted 0.3 map; and the installed snapshot helper captured and then confirmed the repository unchanged. A negative check against the repository's docs/ subdirectory failed closed with the expected repository-root error. This is a packaging and read-only boundary pass. It does not establish that the adopted map is semantically complete, that external audits are useful, or that the commands improve agent performance.

After 0.6.8 was tagged and reinstalled, the cached package repeated the same mapping probe against STAGE's clean adopted repository. The packaged validator accepted the 0.3 manifest. The snapshot helper captured and matched two receipts at revision fcea02b20cc6ca87b43b4602a6446badc68a5cc1, with git_source_state as the declared claim and zero Git-visible working-tree paths. A negative capture against docs/ again failed closed because it was not the Git work-tree root. This confirms that the released cache contains the new scoped snapshot behavior and preserves the repository boundary. It does not cover ignored state, establish semantic map quality, or measure workflow effectiveness.

After 0.6.9 was tagged and reinstalled, its cached package was exercised against a disposable clone fixed at release revision c4345bfd051afc34bd4ed4a36875edbc270f3091. Before running the packaged commands, the probe advanced one tracked file's modification time without changing its content, creating the stat-only condition in which default git diff refreshes the index. The packaged validator accepted the adopted 0.3 manifest, the snapshot helper captured and matched the declared git_source_state, and validation against docs/ failed closed at the repository-root boundary. The target index SHA-256 remained 5f900fd6ecb52f55685380a25dde301c0be4f8d3f6a922260c14000a39e6af6a before and after. This confirms the released package carries the index-write fix under its motivating condition. It does not expand the snapshot claim to all Git internals or prove semantic map quality or workflow effectiveness.

After 0.6.10 was tagged and reinstalled, its cached package repeated the same probe against a disposable clone fixed at release revision be244ca54cacd19b771c0742d468bc726d3a8445. The probe deliberately advanced the modification time of tracked README.md without changing its content. The installed validator accepted the adopted 0.3 manifest, and the installed snapshot helper captured and matched the declared git_source_state before and after validation. A negative check against the clone's docs/ subdirectory failed closed because it was not the Git work-tree root. The target remained clean, while the target index SHA-256 remained abf31945cd48c754f21438d4d450e4f5f9cbf4961518e662a21e58bb7a54650c before and after. This confirms that the released package includes the 0.6.10 semantic and provenance hardening without regressing the read-only index boundary. It does not establish semantic map completeness, project-map usefulness, or an effect on agent performance.

Study Integrity

  • Record research questions before synthesizing changes.
  • Do not modify an external target unless the study explicitly becomes a controlled change with authority to do so.
  • Use immutable permalinks for committed evidence.
  • Preserve failed operations and contradictory evidence.
  • Separate project defects from STAGE defects.
  • Treat a mechanically successful but owner-rejected game as negative method evidence, not as an operational product success with acceptance deferred.
  • Do not mark a trial accepted on behalf of an absent maintainer.
  • Record not_requested when target-maintainer review has not actually been requested. A STAGE researcher may accept a study as evaluation evidence without changing that status.
  • Do not turn one project's naming or directory style into a core requirement.
  • Keep source snapshots and generated study artifacts free of secrets and local credentials.

Release Decision

A STAGE release should state:

  1. which studies motivated it;
  2. which claims gained or lost support;
  3. which normative requirements changed;
  4. whether existing manifests remain valid;
  5. the current owner-readiness statement;
  6. the transferability claims supported and still missing.

The accountable maintainer makes the release decision. Tooling can validate artifacts, but cannot establish empirical usefulness or human acceptance.