# Performance contract and evidence

Last updated: 2026-07-17 JST

## Evidence reading rule

The Tokyo v2 source, logic/deep matrices, five-attempt local Chrome performance
campaign, 25-trial browser reliability campaign, and counterbalanced stationary
A/B baseline have been measured. The performance campaign retained four passes
and one heap-plateau failure; the contract's ten-attempt target and v2 public
Pages deployment/header verification remain pending at this checkpoint. Every
older browser number retained under
**Legacy v1 evidence archive** is historical and must not be presented as v2,
even where the frozen prose used words such as “current” or “final.”

Evidence states used here:

- **measured**: produced by a retained machine-readable counter or runner;
- **source-enforced**: a fixed capacity or threshold visible in code/tests;
- **reported**: copied from the pinned Fable repository without independent
  reproduction; and
- **unavailable/pending**: absent at the required granularity or not rerun.

## Tokyo v2 current source contract

### Fixed resource budgets

| Resource | Steady bound | Transition/peak bound | Status |
|---|---:|---:|---|
| Streamed descriptors | 600 | 1,200 | source-enforced, 5×5×24 per layer |
| GPU world capacity | 625 | 1,250 | source-enforced, 5×125 per layer |
| World archetype draws | ≤5 | ≤10 | source-enforced instanced batches |
| Attached trophy slots | 60 | 60 | source-enforced reservoir |
| Pickup particle slots | 72 | 72 | source-enforced typed arrays |
| Consumed item IDs | 4,096 | 4,096 | source-enforced insertion-order set |
| Home scenery | 1 draw | 1 draw | one merged geometry; tier zero only |
| Tokyo city scenery | 1 draw | 1 draw | one merged lifetime geometry |
| Tokyo Skytree | 1 draw | 1 draw | one merged geometry/material/mesh |
| BGM voices | 3 | 3 | persistent Web Audio oscillators |
| Concurrent SFX voices | ≤18 | ≤18 | hard admission cap |
| DPR | ≤1.5 | ≤1.5 | adaptive quality cap |

The ordinary ball diameter still covers 0.40 m to a nominal 1.6384 km across
six 4× tiers without increasing the 600-descriptor steady window. Tokyo v2 adds
only fixed-lifetime scenery and goal resources. It does not add buildings as
individual meshes while scale or travel distance increases.

The browser harness enforces 24 total draws during tier overlap and 17 in the
final streaming soak. Tokyo v2 passed those gates: the measured transition
maximum was 19 draws and the final soak maximum was 15. Failed attempts would
remain part of the report rather than silently raising or omitting a budget.

### Current measured logic and telemetry checkpoint

`bun test` on 2026-07-17 produced:

```text
Bun 1.3.11
12 files
58 pass, 0 fail
13,102 expect() calls
253.00 ms runner-reported duration
```

The new coverage includes:

- exact 634 m Skytree visual height and one-mesh lifecycle;
- fixed Skytree coordinate and floating-origin placement;
- zero-height home floor, 12 cm genkan drop, deterministic road/curb/terrace
  heights, and bounded relief across every tier;
- one-draw home and one-draw city scenery visibility/lifecycle;
- tutorial wall envelope, locked front threshold, and genkan-only exit;
- 0.8 s dash, four-second recharge, 2.2× cap, edge-latch semantics, and
  simulation-time determinism;
- foreground wall-clock adjudication that retains visible stalls while bounding
  fixed-step physics work;
- 1.5 s combo boundary, 3× cap, 5,000 rare bonus, 20,000 goal bonus, linear
  time bonus, and inclusive S/A/B/C rank boundaries;
- encoded X intent, canonical production URL, and removal of debug/preview
  source URLs; and
- telemetry schema, token reconciliation, expected-agent fail-closed behavior,
  event/artifact tamper detection, and pinned-import privacy.

The retained telemetry run is
`evidence/v2/goal-checkpoint-20260717/manifest.json`:

| Field | Value | Provenance/limit |
|---|---:|---|
| Goal start counter | 400,444 tokens / 3,644 s | exact goal snapshot |
| Goal checkpoint counter | 556,032 tokens / 4,099 s | exact goal snapshot |
| Delta | **155,588 tokens / 455 s** | exact subtraction, aggregate only |
| Input/cache/output/reasoning split | unavailable | checkpoint did not expose it |
| Per-agent usage | unavailable | aggregate cannot be attributed to root/subagents |
| Run wall-clock boundaries | unavailable | start/end timestamps were not captured |

The orchestration ledger separately observes one root task and eight bounded
subtask streams, with peak concurrency four. That is a task/execution count,
not token attribution or a durable-identity count. It is not directly comparable
to Fable's approximately 148 reported task-agent executions.

The manifest correctly leaves `usage.byAgent` empty. Its zero-valued per-agent
total is not a claim that agents consumed zero tokens; it means no reconciled
per-agent records were imported. Direct token-efficiency comparison with
Fable's reported ultracode totals remains unavailable for this historical
checkpoint.

The final documentation build checkpoint also passed strict TypeScript with
zero diagnostics and Vite 8.1.2:

| Artifact | Raw | gzip |
|---|---:|---:|
| HTML | 9.78 kB | 3.20 kB |
| CSS | 18.02 kB | 4.90 kB |
| JavaScript | 626.53 kB | 161.59 kB |
| JavaScript source map | 3,050.88 kB | — |

Vite transformed 31 modules. The JavaScript chunk advisory remains a
transfer/code-splitting opportunity, not a failed runtime pool gate.

### Deterministic matrix evidence

Two complete core deep-validation runs each produced:

| Metric | Per run |
|---|---:|
| Semantic cases | 7,917,462 |
| Atomic checks | 31,769,395 |
| Failures | 0 |
| Reproducibility digest | '0e6b2224' |

The Tokyo v2 feature matrix separately produced:

| Domain | Semantic cases | Atomic checks |
|---|---:|---:|
| Deterministic terrain/home/genkan | 1,835,529 | 9,177,621 |
| World and rare distribution | 1,825,200 | 19,420,131 |
| Run score/rank/freeze | 1,624,001 | 8,872,004 |
| Dash traces | 500,000 | 5,000,000 |
| X share | 50,000 | 500,000 |
| **Total** | **5,834,730** | **42,969,756** |

Both feature-matrix passes had zero failures and matched digest
'5ee48646fe403f13'. The second pass proves repeatability and is excluded from the
case/check totals. The report records 72 world seeds, 1,752,192 regenerated
descriptors, zero duplicate IDs, 19,448 rare items across all thirteen IDs,
24,000 run traces, 500,000 dash steps, and 50,000 share inputs. The machine
source is 'comparison/data/v2-deep-validation.json'.

Semantic cases and atomic invariant evaluations remain separate units. Neither
number is described as hand-authored tests or added to the fast-suite assertion
count.

### Local Chrome real-frame performance campaign

Chrome 150.0.7871.125 ran five complete production-build attempts at 1440×900,
DPR 1, Ultra, adaptive quality off. The reported WebGL renderer was ANGLE
SwiftShader, so absolute frame times characterize this software-rendered
environment; fixed-resource behavior and the retained distribution are the
primary scale evidence.

| Metric across all 5 attempts | Min | Mean | Median | p95 | Max |
|---|---:|---:|---:|---:|---:|
| Tier 0 steady p95 | 41.1 ms | 44.98 ms | 46.6 ms | 46.8 ms | 47.2 ms |
| Tier 5 steady p95 | 30.9 ms | 40.24 ms | 43.8 ms | 44.4 ms | 48.3 ms |
| Streaming p95 | 34.1 ms | 37.36 ms | 36.6 ms | 38.8 ms | 41.9 ms |
| Streaming distance | 4.723 km | 5.130 km | 4.775 km | 5.544 km | 5.883 km |
| Streaming heap delta | 90,567 B | 1,373,416 B | 721,247 B | 955,247 B | 4,608,227 B |

Four attempts passed all 35 gates. Attempt 4 passed 34/35 and failed only
`streaming.heap-plateau`: heap rose from 6,703,757 to 11,311,984 bytes, a
4,608,227-byte delta above the configured `max(60%, 4 MiB)` allowance. Across
all five attempts, every frame-time, pool, draw, geometry, texture, minimum-
distance, and browser-error gate passed. Maximum tier-5 occupancy remained
1,190 active / 1,200 allocated / 19 draws; the streaming maxima remained
593 / 600 / 15. No failed attempt is removed from the 4/5 denominator.

The compact campaign with every attempt outcome and distribution is
`comparison/data/v2-browser-performance-campaign.json`; its raw run artifacts
are retained under `/tmp/sol-v2-perf-campaign-final`. The latest retained passed
attempt is also copied to `comparison/data/v2-browser-performance.json`.

### Counterbalanced stationary Codex/Fable A/B

The v2 baseline used one Chrome environment, five trials and 600 recorded frames
per site, cache disabled, and counterbalanced site order. All ten trials passed.

| Stationary page metric | Codex Tokyo v2 | Fable |
|---|---:|---:|
| Trials / frames | 5 / 3,000 | 5 / 3,000 |
| Frame p95 | 44.1 ms | 45.1 ms |
| Draw calls/frame | 14 | 44 |
| Mean runtime heap | 4.53 MB | 8.55 MB |
| Mean resource transfer | 645,541 B | 738,351 B |
| Captured browser errors | 0 | 0 |

Both pages were stationary after start. Fable renders a richer OSM/curated Tokyo
world; Codex renders a bounded procedural/fixed Tokyo composition. The table
does not compare gameplay throughput, equivalent visual quality, clear time,
score, or debugging productivity. Machine source:
'comparison/data/v2-cross-browser.json'.

### 25-trial browser reliability campaign

The local v2 build passed **25/25** independent Chrome trials. Each trial opened
and closed a fresh CDP target with cache disabled, yielding a new document, JS
realm, WebGL context, AudioContext, and game instance.

| Metric | Result |
|---|---:|
| Contracts per trial | 62 |
| Exact assertions | 1,550 / 1,550 |
| Unique semantic contracts | 62 |
| Repeated reliability observations | 1,488 |
| Browser errors | 0 |
| Mean / p95 trial duration | 6,000.30 / 6,220.17 ms |

The 1,488 repeats are not added to the 62 unique semantic contracts. Debug
mutation and fixed-step hooks accelerate deterministic tier/goal flow, so this
campaign is a functional reliability gate, not another real-time 600-frame
performance sample. The contracts cover a clear initial camera, the locked home
envelope and unlocked genkan exit, active wall-clock behavior, BGM, dash, rare
bonus/SFX, five transitions, pool/draw bounds, locked/ready fixed-coordinate
Skytree, absorption midpoint, result/rank/X share, resources, and empty browser
errors.
Machine source: 'comparison/data/v2-browser-campaign.json'.

### v2 rerun matrix

| Gate | Required scenario | Status at this checkpoint |
|---|---|---|
| Build/typecheck | strict TypeScript + production Vite output | pass; 31 modules, JS 626.53/161.59 kB raw/gzip, 3,050.88 kB map |
| Functional browser | home, BGM, rare, dash, six tiers, locked/ready Skytree, result, X URL, resources/errors | 25/25 trials pass; 1,550/1,550 assertions |
| Real-frame performance | 170 rAF/tier, transition and steady percentiles, pools, draws, geometry/textures, heap, final stream crossing | five attempts complete; 4 pass / 1 heap-plateau fail, all outcomes retained |
| Repeated Codex campaign | all attempts retained, including failures | functional 25/25; performance 4/5, ten-attempt target pending |
| Direct Codex/Fable runtime | same Chrome, viewport, cache/order, warm-up, frames, and stationary scenario | pass; 5×600 frames/site, all trials pass |
| Public verification | immutable Pages URL, stable alias, assets/docs/security/cache headers, smoke and performance | pending deployment |

The exact procedures and valid comparison units are frozen in
[`V2-BENCHMARK-CONTRACT.md`](V2-BENCHMARK-CONTRACT.md).

## Legacy v1 evidence archive — frozen 2026-07-10 through 2026-07-13

Everything from this heading to the end of the file is historical v1 evidence.
It is preserved so prior measurements and failures remain auditable; names,
coordinates, goal radius, draw plateaus, browser results, campaign counts, and
completion checkboxes below do not describe Tokyo v2.

### Structural budgets

| Metric | Steady budget | Transition budget | Enforcement |
|---|---:|---:|---|
| Streamed world slots | 600 | 1,200 | fixed 5×5×24 descriptor pool |
| GPU world capacity | 625 | 1,250 | five balanced 125-slot archetype buffers |
| Attached trophy slots | 60 | 60 | fixed reservoir |
| Particle slots | 72 | 72 | fixed typed arrays |
| World archetype draws | ≤5 | ≤10 | five instanced batches/layer |
| Total target draws | ≤17 | ≤24 | includes the single-draw final landmark |
| DPR | ≤1.5 | ≤1.5 | adaptive quality cap |
| Collision candidates | target <200/tick | <200/tick | spatial hash telemetry |
| Consumed history | ≤4,096 IDs | ≤4,096 IDs | bounded insertion-order set |
| SOL CITADEL | ≤1 merged draw | ≤1 merged draw | one geometry/material/mesh |

The ordinary tier domain increases from a 0.40 m starting diameter to a
1.6384 km nominal final-tier exit diameter, 4,096×, while the steady world
instance budget remains 600 on every tier. The final objective is a single-draw
landmark with a 512 m equivalent radius. It unlocks at a 711.111… m player
radius and completes only after physical contact and absorption, not on size
alone.

## Live metrics

`window.__SOL_KATAMARI_METRICS__` includes:

- FPS and average/p95/p99 frame time over a 600-frame window
- physical/local radius and view radius
- tier and transition progress
- objective phase/title, target bearing/distance, and SOL CITADEL
  visibility/readiness/absorption state
- active/allocated/attached instance counts
- collision candidate count
- actual total draw calls and triangles from `renderer.info`
- geometry and texture counts
- quality/DPR tier, seed, and deterministic state hash

## Frame-time regression contract

The 33.5 ms value is a comparison floor, not an unconditional absolute ceiling.
The final-tier steady gate is:

```text
final steady p95 <= max(33.5 ms, initial steady p95 × 1.35 + 2 ms jitter)
```

The streaming gate applies the same form using the measured final-tier steady
p95. The 1.35 ratio prevents physical scale from increasing runtime cost, while
the 2 ms allowance absorbs scheduler/display jitter between finite browser
samples. The 33.5 ms floor prevents an unusually fast first sample from making
the relative threshold unrealistically strict.

## Latest local evidence — 2026-07-12

The current mission-first SOL CITADEL revision was measured locally in Chrome
150.0.7871.115, WebGL2, 1440×900 CSS pixels, Ultra quality, with adaptive
quality disabled. All currently enforced local performance gates passed.

| Sample | p95 ms | Max active | Max allocated | Max draws | Geometry | Textures |
|---|---:|---:|---:|---:|---:|---:|
| Tier 0 steady | 33.3 | 600 | 600 | 15 | 16 | 1 |
| Tier 5 steady | 34.1 | 587 | 600 | 17 | 16 | 1 |
| Transition maximum | 34.9 | 1,200 | 1,200 | 22 | 16 | 1 |

The first/final steady p95 delta was 0.8 ms (approximately 2.4%). The 1.024×
first-to-final ratio remained below the configured comparison-floor/relative/
jitter gate. The final-tier
landmark raised the steady budget by one bounded geometry/draw only; geometry
and texture counts remained exactly 16/1 across the measured tiers.

The final-tier streaming soak travelled 6,321.66 m, exceeding the 3,840 m
minimum. Its p95 was 34.0 ms; active/allocated/draw maxima were 587/600/17.
Reported JS heap decreased from 8,904,857 to 6,200,786 bytes. The instance,
draw, resource-plateau, frame-time, distance, and heap gates all passed. The
machine-readable source is `/tmp/sol-katamari-check/performance.json`.

## Current public verification — 2026-07-12 final P1 build

The final P1 deployment completed at
`https://d596cf0b.sol-katamari.pages.dev/`. The stable production alias is:

`https://sol-katamari.pages.dev/`

The final public E2E flow exited zero. In addition to the complete SOL CITADEL
mission and input flow, it independently verified the navigation bearing and
distance equations, the objective panel's target/distance/eight-direction
accessible label, and that absorption is still incomplete at the 0.5 second
midpoint of its 0.95 second animation.

Chrome 150/WebGL2 public performance results were:

| Sample | p95 ms | Max active | Max allocated | Max draws | Geometry | Textures |
|---|---:|---:|---:|---:|---:|---:|
| Tier 0 steady | 33.4 | 600 | 600 | 15 | 16 | 1 |
| Tier 5 steady | 34.5 | 587 | 600 | 17 | 16 | 1 |
| Transition maximum | 34.6 | 1,200 | 1,200 | 22 | 16 | 1 |

The final/initial steady ratio was approximately 1.033×, inside the comparison
floor/relative/jitter regression gate. Transition occupancy remained within the
1,200/1,200 descriptor bounds and 24-draw limit. Geometry/textures returned to
the exact 16/1 plateau.

The public final-tier streaming soak travelled 6,367.58 m at 34.1 ms p95. Its
active/allocated/draw maxima were 587/600/17. Reported heap decreased from
9,094,160 to 6,816,381 bytes. Every public frame-time, instance, draw, resource,
distance, and heap gate passed.

The unique deployment, stable alias, Worklog, and Performance document all
returned HTTP 200. The production HTML returned the configured security
headers, and hashed JavaScript/CSS assets returned
`Cache-Control: public, max-age=31536000, immutable`.

## 2026-07-10 baseline evidence

Verification environment: Chrome 150.0.7871.115, WebGL2, 1440×900 CSS pixels,
Ultra quality, adaptive quality disabled for cross-tier comparability.

Build and structural command:

```bash
bun run check
```

Result at 2026-07-10 14:40 JST:

```text
20 pass, 0 fail, 12,795 expect() calls
strict TypeScript: 0 diagnostics
Vite: 23 modules in 111 ms
JS: 587.33 kB / 149.66 kB gzip
CSS: 12.44 kB / 3.70 kB gzip
```

Coverage includes:

- monotonic 4× tiers and exact physical-boundary continuity
- volume growth, eligibility, score, and all six tier crossings
- deterministic chunk descriptors and unique IDs
- fixed instance caps for every tier and quality density
- descriptor identity reuse across chunk streaming refills
- spatial-hash query isolation
- a 60,000-frame structural soak across all six tiers
- exact projected-ball continuity across boundaries
- implemented camera-damping transition stays within 8% projected-size change
- logarithmic scale interpolation finiteness/monotonicity
- adaptive-quality hysteresis
- custom fog shader/Three uniform contract
- pending-layer-only shader precompile lifecycle

## Browser smoke evidence

`bun run test:e2e` completed with zero console, runtime, network, or shader
errors. It verified:

- real keyboard movement: 2.436 m
- real CDP pointer camera drag: 0.720 radians
- real CDP touch-joystick movement: 0.340 m
- sound toggle and restore
- collision-path pickup: count 1→2, radius 0.2052→0.2105 m, score 90→179
- oversized target remained active after collision and position was repelled
- all five 1.8 s crossfades and all six settled tiers
- finale, restart, pause, and resume
- five mid-transition and five settled-tier screenshots

## Real-frame performance evidence

`bun run test:perf` measured 170 actual `requestAnimationFrame` callbacks per
tier, separating the transition window from the final 35 steady frames.

| Tier | Avg ms | transition p95 | steady p95 | max active | max slots | max draws |
|---:|---:|---:|---:|---:|---:|---:|
| 0 | 16.66 | 17.4 | 17.3 | 600 | 600 | 15 |
| 1 | 17.07 | 17.4 | 17.3 | 1,197 | 1,200 | 21 |
| 2 | 16.67 | 17.4 | 17.3 | 1,199 | 1,200 | 21 |
| 3 | 16.67 | 17.5 | 17.5 | 1,196 | 1,200 | 21 |
| 4 | 16.67 | 17.5 | 17.4 | 1,199 | 1,200 | 21 |
| 5 | 16.67 | 17.6 | 17.5 | 1,194 | 1,200 | 21 |

Final-tier steady p95 is only 0.2 ms (1.2%) above the first tier, far inside
the configured comparison-floor/relative/jitter threshold. Geometry count was 15 at both
ends and texture count stayed exactly 1. Reported JS heap fluctuated between
5.1 and 9.7 MB rather than growing with tier scale.

The final-tier 300-frame autoplay/streaming soak then travelled 4,345 m,
absorbed seven more objects, and ended at 593 active / 600 allocated instances,
with 16 draws, 15 geometries, and 1 texture. The full machine-readable report
is produced at `/tmp/sol-katamari-check/performance.json` by default.

The regression gate requires exact first/last geometry and texture equality,
caps final heap growth at the greater of 60% or 4 MiB, and requires the
streaming soak to cross at least one 3,840 m final-tier chunk. During that soak
it separately asserts p95 frame time, 600/600 active/allocated maxima, 16 draws,
exact resource equality, and the same heap-growth ceiling.

## Public Pages verification — 2026-07-10 build

This section is retained as historical evidence for the earlier deployed
build. Current 2026-07-12 SOL CITADEL public evidence is recorded above.

The same harnesses passed against
`https://sol-katamari.pages.dev/` on 2026-07-10 16:48 JST. Public smoke booted
WebGL2, exercised keyboard, pointer, touch, sound, real absorption, exact-target
oversized rejection, five crossfades, finale, restart, pause, and resume with
zero captured browser/runtime/network errors.

The final strengthened public 170-rAF-per-tier performance run measured 17.5 ms
steady p95 at tier 0 and 17.6 ms at tier 5. Transition p95 never exceeded
17.5 ms; active
instances peaked at 1,199, allocated slots at 1,200, and total draws at 21.
Geometry and texture counts again plateaued at 15 and 1. The final 300-frame
streaming soak travelled 4,288 m with 17.5 ms p95 and ended at 593 active / 600
allocated instances with 16 draws. Its measured heap moved from 5.55 to 6.92
MB, below both the ratio and absolute growth ceilings.

HTTP verification returned 200 for the production alias, the immutable JS/CSS
assets, Worklog, and Performance document. HTML included CSP, COOP,
Permissions-Policy, Referrer-Policy, nosniff, and frame-denial headers. Hashed
JS/CSS returned `Cache-Control: public, max-age=31536000, immutable`; the two
documents returned `text/markdown`.

## Deep deterministic validation — 2026-07-13

`bun run test:deep` exercises the production procedural, growth, navigation,
and quality functions over a deterministic matrix. Set
`DEEP_VALIDATION_OUTPUT` to retain its machine-readable JSON; the comparison
artifact is `comparison/data/deep-validation.json`.

| Domain | Semantic cases | Atomic checks | Principal measured breadth |
|---|---:|---:|---|
| World generation | 1,161,978 | 14,636,922 | 64 seeds × 6 tiers × 121 chunks × 24 objects; every object regenerated |
| Streaming/pool | 4,286,304 | 8,490,528 | 9,504 worlds; 4,276,800 visits; 4,147,200 identity checks |
| Growth/goal | 369,971 | 2,709,785 | 4,096 complete trajectories; 349,868 pickup attempts; 16,007 boundaries |
| Navigation/landmark | 626,176 | 1,882,624 | 622,080 camera vectors; 4,096 landmark placements |
| Adaptive quality | 1,492,520 | 1,968,080 | 2,048 traces plus exact thresholds repartitioned 1–1,000 ways |
| **Total** | **7,936,949** | **29,687,939** | **0 failures; 2.653 s recorded run** |

The sampled world and 9,504 active windows had zero ID collisions; maximum live
occupancy was exactly 600. All 4,096 growth paths reached goal eligibility in
62–96 accepted pickups and visited every tier. Maximum navigation distance and
bearing errors were `3.50e-10 m` and `7.60e-8 rad`. All 4,095 observed quality
transitions matched the reference state model. A second complete run passed in
2.632 seconds with the same `c58f3b92` evidence digest.

The report separates generated semantic cases from atomic invariant
evaluations. The 29.7 million value is not claimed as 29.7 million hand-authored
tests and is not directly interchangeable with another test runner's assertion
count. The reproducible input dimensions and category counts are the useful
comparison unit.

The initial run exposed two numerical boundary defects: exact CITADEL
eligibility could fail after `radius³ → cbrt` rounding, and the four/twelve-
second quality timers could miss when elapsed time was split into floating-
point partitions. The implementation now uses a scale-relative four-ulp pickup
tolerance and a 1 ns timer tolerance. Focused regressions preserve both fixes;
the fast suite is now 26 pass, 0 fail, and 12,826 `expect()` calls.

## Same-Chrome cross-game stationary baseline — 2026-07-13

The current public Codex and Fable deployments were measured by one CDP script
in Chrome 150 at 1440×900 and DPR 1. Each page received five new-tab trials,
disabled cache, counterbalanced order, 60 warm-up frames, and 600 sampled
real-rAF frames. No input was sent after start, making this a stationary render
baseline rather than a gameplay throughput test.

| Aggregate over 3,000 frames/site | Codex / SOL ROLLER | Fable 5 / Tokyo | Codex ÷ Fable |
|---|---:|---:|---:|
| Mean frame interval | 17.58 ms | 29.41 ms | 0.598× |
| p50 / p95 / p99 interval | 16.70 / 32.80 / 33.50 ms | 33.30 / 34.00 / 34.30 ms | 0.502× / 0.965× / 0.977× |
| Maximum / frames above 50 ms | 50.20 ms / 3 | 83.90 ms / 4 | 0.598× / 0.750× |
| Draw calls/frame, p50 and p95 | 15 / 15 | 44 / 44 | 0.341× |
| Mean post-sample runtime heap | 5.64 MB | 9.31 MB | 0.606× |
| Mean resource transfer | 159,950 B | 738,409 B | 0.217× |
| Console/runtime/log/network errors | 0 | 0 | — |

The p50 difference is much larger than the p95 difference, so this browser run
must not be reduced to one universal FPS multiplier. More importantly, the
workloads are not functionally equivalent. Fable presents an OSM-derived Tokyo
world with different assets, UI, content, and engine behavior; SOL ROLLER uses
a bounded procedural world. The table supports claims only about the two
current pages while stationary under this procedure. It does not compare
movement, collection, streaming, growth, transitions, visual richness,
development time, token efficiency, or debugging capability.

All 10 trials and 6,000 frame/draw samples are retained in
`comparison/data/cross-game-benchmark.json`; the summary records SHA-256
`c336b3d7e79c650af679613c95501c982c7e79a0ff01f7e8a4e1732ecd566436`.
Fable's approximately 3,000 assertions across five versions, 25 browser passes,
27 screenshots, 2,376 v4 OSM assertions, and 58% test/review allocation are
article-reported figures, not results rerun here. They demonstrate real-data and
human-review breadth but do not share definitions with Codex semantic/atomic
counts. Source:
`https://qiita.com/otani_ai_memo/items/3c04185ef80b6a9d97c2`.

## Final public repeated-validation campaign — 2026-07-13

The stable public deployment was exercised repeatedly after the deterministic
matrix. The primary campaign ran 25 full browser smoke attempts and 10 complete
performance attempts sequentially through one Chrome endpoint. A second run
used a fresh profile/separate endpoint for five smoke and five performance
attempts while the host was under concurrent Chrome load.

| Campaign | Total pass | Smoke pass | Performance pass | Elapsed | PNG |
|---|---:|---:|---:|---:|---:|
| Primary | 32/35 (91.4%) | 24/25 (96.0%) | 8/10 (80.0%) | 572,976.85 ms | 312 |
| Fresh-profile concurrent-load stress | 5/10 (50.0%) | 5/5 (100%) | 0/5 (0%) | 254,221.80 ms | 65 |
| **Combined observations** | **37/45 (82.2%)** | **29/30 (96.7%)** | **8/15 (53.3%)** | **827,198.65 ms** | **377** |

Failures are part of the result. Primary smoke attempt 23 stopped receiving rAF
callbacks at its movement probe and returned movement 0; 24 other primary and
all five additional smoke attempts completed. Primary performance attempts 2
and 4, plus all five additional stress attempts, exceeded the configured frame-
time gate. The failed timing attempts still remained inside the instance, draw,
and heap structural limits, so they show frame-scheduling/timing instability
rather than unbounded scene growth. The eight passing primary performance runs
completed both timing and structural gates.

Across those eight successful performance trials, transition allocation stayed
at or below 1,200 descriptors and 22 draws; final streaming stayed at 587 active,
600 allocated, and 17 draws. The eight streaming soaks travelled 5,502.90–
5,850.16 m with p95 values 33.70–34.40 ms. Heap delta ranged from -3,386,764 to
+2,881,812 bytes, within the configured bounds.

The phrase “fresh profile” describes browser state only. The second Chrome used
a separate profile and CDP port, but ran on the same machine while another
Chrome process remained active. It is a concurrent-load stress run and is not a
true unloaded/isolated measurement. Its 0/5 timing result must not be used as an
isolated hardware baseline.

The raw and compact machine-readable reports are:

- `comparison/data/validation-campaign.json`
- `comparison/data/validation-campaign-summary.json`
- `comparison/data/validation-recovery-campaign.json`
- `comparison/data/validation-recovery-summary.json`

This campaign does not directly measure Fable. Direct cross-game claims use
only the counterbalanced five-trial-per-site, 6,000-frame stationary benchmark
documented above. Fable development/test figures remain article-reported values.

## Completion gates

- [x] TypeScript strict check passes with installed dependencies.
- [x] Production Vite bundle contains headers, docs, and source maps.
- [x] Local browser initializes WebGL2 without console errors.
- [x] Keyboard, pointer, touch, pause, sound, and restart paths are exercised.
- [x] Real pickup increases radius/count/score and oversized collision blocks.
- [x] All five tier transitions and six settled tiers are captured.
- [x] Actual steady/transition draw and instance counts meet the table.
- [x] Browser p95 frame time is recorded at first and last tier.
- [x] Geometry/texture counts return to the same plateau.
- [x] Local E2E proves size alone does not clear the mission, locked contact is
  rejected, ready contact absorbs SOL CITADEL, and only then shows the finale.
- [x] The 2026-07-12 local performance and streaming gates pass with the
  single-draw landmark active.
- [x] Structural tests pass: 26 pass, 0 fail, 12,826 assertions, including
  persistent outer-extent corridor clearing after density/stream rebuilds.
- [x] Deep validation passes 7,936,949 semantic cases and 29,687,939 atomic
  checks with a stable `c58f3b92` evidence digest.
- [x] Exact goal-eligibility and partitioned quality-timer boundary regressions
  are reproduced, corrected, and retained as focused tests.
- [x] E2E independently verifies navigation distance/bearing, accessible
  eight-direction output, and incomplete absorption at the 0.5 second midpoint.
- [x] The 2026-07-10 public build returned 200 with expected security/cache
  headers and passed its public smoke/performance gates.
- [x] The final P1 revision is deployed at
  `https://d596cf0b.sol-katamari.pages.dev/` and the stable Pages alias.
- [x] Public smoke, performance, asset/header, and published-document checks
  pass against that revision.
- [x] The 45-attempt repeated campaign is retained with all eight failures and
  377 screenshots rather than reporting only successful runs.
- [ ] Repeated performance timing is not uniformly green: 8/15 attempts passed;
  seven frame-time-gate failures occurred under the recorded host conditions.

The original local/public release gates and every structural resource bound are
measured and satisfied. The later repeated campaign additionally exposes
frame-timing reliability as an open environmental/performance investigation;
its failures are not overwritten by the earlier green single-run evidence.
