A themed, evergreen reference of notable tradeoffs and reversals — organized by topic, not by task number or date. Each section states the current rationale in prose, with a brief history note only where a real reversal happened, and cites notes task numbers in brackets for the full problem/attempt/verification detail rather than restating it here.
Maintenance rule: when a new decision is made, fold it into an existing theme below, or add a new H2 only if it's a genuinely new axis of decision-making. Do not append one entry per completed task — that's what notes is for.
Anything that affects simulation outcome (enemy AI timing, loot rolls, weapon spread, elite-loot coinflips) draws from a shared seeded mulberry32 PRNG (src/prng.ts) instead of Math.random(), so a recorded seed plus a recorded per-frame input digest reproduces a byte-identical run. Purely cosmetic randomness (blood particles, SFX pitch, BGM shuffle, console hints) deliberately stays on Math.random() — seeding it would add complexity for a divergence no one would notice. See Architecture for the mechanism.
Replay scope started at a single level (whichever one a run ended on) and was later expanded to cover a full multi-level campaign, once it became clear a single-level replay was too narrow to be useful for a genuinely multi-level run. The recorded payload format was versioned rather than migrated in place — old single-level recordings are recognized and hidden from playback rather than trusted blindly, since their shape can't be verified from the type alone once read back out of localStorage. [notes: Task 35, 36, 37]
Every simulation input must be in the replay payload, and only a playback check can prove it is. The format carried gameplaySeed, difficulty, gore, carryover, astHash and balanceHash but not the bot's rotation-speed multiplier, which the engine read straight off ?botRotSpeedMul — so the shipped board recorded at 2.0/3.5/5.0 and played back at 1.0, and every replay diverged within seconds while CI stayed green. The rule that follows is the one already stated for the styleset (a pure function of (seedFrom(parsed), bonusLevel), so it needs no field): anything not derivable from what's recorded has to be recorded. The corollary is the part that was missing — a hash guard cannot enforce it. astHash answers "same source file" and balanceHash answers "same game"; neither can answer "does this still play back the way it was recorded", because a dropped input changes neither fingerprint. npm run verify:replay answers it directly by re-simulating a shipped entry and comparing frame counts and the carryover score ladder, and is the reason that class of bug now fails loudly instead of silently. A value the engine reads from ambient state (a URL param, window.location, a preference) rather than from its own parameters is the shape to watch for: the same multiplier had already caused an identical, separately-diagnosed measurement bug on the multiplayer axis (multiplayerSessionBootstrap.mjs).
Both ends of a record/playback pair must hash the map in the same state. computeBalanceHash reads maxHp, and the RaycasterEngine constructor rescales enemy hp/maxHp in place by the difficulty multipliers — so when the fingerprint is taken is part of its definition, not an implementation detail. Recording hashes before constructing; playback did it after, which refused every easy (0.7x) and hard (1.5x) replay as balance drift while leaving normal (1x) working. Normal-only test coverage plus a normal-only bundled board is why it went unseen. Playback now snapshots the roster as generated, and the ordering is commented at both call sites rather than left resting on "an async function's prefix runs synchronously".
Determinism constrains what a perf optimization is allowed to change, not whether it can change enemy behavior at all: replacing every chasing enemy's own per-frame windowed BFS flood with one shared full-map path field rooted at the player's tile (pathField.ts) is sim-affecting on purpose — the old flood was clipped to a window and could miss routes the new unclipped field finds — so it's still deterministic (same seed, same rng draw order, byte-identical replay of itself), just not backward-replay-compatible with recordings made before the change. The already-shipped default highscores/replays were regenerated in the same batch rather than left to silently desync. [notes: Task 85]
The original entrypoint heuristic (a filename-convention check, then a fallback scan for any C-family file containing a literal main()) failed in practice on a real shared-library project with no main.c at all — it picked an arbitrary small helper file purely because of alphabetical sort order, and every hand-rolled unit test file in that same repo also had its own main() and was more complex than the real source, ruling out "just pick whichever main()-having file scores highest" as a fix on its own. The replacement is layered: files are split into a real-source bucket and a test/spec-fixture bucket first (by path segment, case-insensitive); within each bucket, a filename-convention check runs before a language-agnostic scan-and-score pass (any language's main/Main entity, catching capitalized conventions too); the test bucket is only consulted if the real-source bucket found nothing at either stage. The whole pass is time-capped, degrading to "first file in tree order" if it fires. [notes: Task 45]
Complexity scales enemy count (one extra enemy per fixed complexity increment) rather than piling more HP onto a single enemy — this was chosen so a big function reads as "several problems," not one implausibly tanky one. Above a threshold placed exactly where a pack would otherwise hit its largest practical size, the room's HP budget is doubled and the pack becomes an Elite pack: a boss-tier encounter (higher damage, guaranteed drop, a distinct look) rather than just a bigger version of the same one. A genuinely different code shape gets a genuinely different encounter shape.
It used to flip to a single Elite, and that was wrong — measurement, not taste (2026-08-08). Because the multiplier stacked on un-split HP, complexity 39 produced a 4-enemy / 976 HP pack and complexity 40 produced one 2,000 HP enemy, with nothing ever generated in between. Across seven repositories of playtest telemetry 1,332 Elites spawned and 2 died — both ~2,000 HP on normal, none at all on hard — against a 21-26% kill rate for everything up to 499 HP. All 514 cleared runs on an Elite-bearing level left the Elite alive. An Elite was not a hard fight, it was terrain to walk past, and routing around it was never the intent.
So the count logic no longer flips. An Elite room's budget is capped (ELITE_MAX_MEMBERS × ELITE_MEMBER_HP_CAP) and split like any other pack, which keeps every enemy inside the band the player demonstrably kills and makes the fight's incoming damage fall as members die. Only the anchor carries Enemy.elite: the damage multiplier is applied per enemy, so flagging the whole pack would multiply the room's DPS by its size — trading an unwinnable fight for an unsurvivable one. See Game Design for the player-facing intent.
Coop deliberately breaks that ceiling, and that is the point (2026-08-12). ELITE_MEMBER_HP_CAP bounds what the generator emits — 350 base, 525 after Hard's 1.5x, the top of the band the kill data actually covers. It does not bound what a player faces, because RaycasterEngine's constructor then applies eliteScalingFor(playerCount) on top (+50% HP and +25% damage per extra player, Elites only), so a four-player Hard session meets an Elite anchor at roughly 1,313 runtime HP — 2.5x the single-player ceiling and outside any measured band. That is intentional, for now: an Elite sized for one player is trivial for four, and a coop encounter that a single competent player can solo is not a coop encounter. The scaling exists to make the pack a fight the team has to cooperate on rather than one anybody clears alone, and overshooting the solo-calibrated ceiling is the mechanism, not a bug in it.
Two things follow, and neither is optional. Do not "fix" this by clamping the product against ELITE_MEMBER_HP_CAP — a doc audit flagged exactly that in 2026-08 on the reasonable-looking grounds that enemies.ts' comment calls the cap absolute; it is absolute over the generator, and the engine pass is a separate axis layered above it. And the constants themselves are still unvalidated: multiplayerScaling.ts says so out loud, and the ceiling argument that justifies 350 in single-player has no coop equivalent, because no coop telemetry campaign has run at the scale that produced the 112,311-kill single-player figure. Treat 1 + 0.5n as a placeholder with a clear purpose, not as a measured value — retuning it is a balance question waiting on data, not a correctness one. Note also that enemy maxHp is lockstep simulation state: changing this pass desyncs mixed-build sessions and invalidates every recorded multiplayer replay carrying an Elite, so it moves with a build-version bump or not at all.
Archetypes have weapons, and their mean DPS is held identical on purpose (2026-08-19). ENEMY_WEAPONS (combatConstants.ts) gives each archetype its own bolt — speed, damage, cooldown window, inherent scatter, colour — where one shared bolt with a damage multiplier used to serve all three. The rule to preserve is not the numbers but the constraint they satisfy: Elite doubles damage and its cooldown window; Edge Case halves both, so each archetype's damage / meanCooldown is exactly what the pre-table model produced.
That is deliberate and load-bearing. The measured difficulty driver is exposure-time x enemy count, and three separate A/Bs redistributed damage without moving the total (122/127/126). A weapon table that also raised DPS would be indistinguishable from that history — the change would be unattributable the moment anyone measured it. Holding the mean fixed makes the shape the only variable: a telegraphed Elite shell you can step out of, an Edge Case spray you cannot. enemyWeapons.test.ts pins the invariant and derives the legacy side from the raw constants, so it is a real second opinion; a retune that intends to move a DPS has to change that test with a number rather than delete it.
radius is the field to be careful with. It is per-weapon but currently identical across all three, because it feeds the hit test — a wider bolt connects more often, moving effective DPS without touching either term the invariant compares. A test pins the uniformity. It lives in the table anyway so that spending that budget is a visible one-line decision instead of a constant buried elsewhere.
Two seams square the archetype ladder if confused, and both have tests: damageMultiplier() in enemyAi.ts is melee-only now, because weapon.damage already carries the ladder; and multiplayer's player-count Elite scaling still multiplies on top, Elite-only.
The cluster fast-path exists because the scorer cannot see AoE, and that is legitimate (2026-08-20, measured). resolveShot resolves a target per pellet, so a shotgun blast really does hit more than one enemy — 23.8% of its shots hit two, against 0.0% for the pistol and gdb. scoreRangedWeapon computes a single-target time-to-kill and structurally cannot express that, so the fast-path is covering a real blind spot rather than overriding the economics arbitrarily. Do not "simplify" it away into the scorer without first giving the scorer a multi-target term.
Friday Hotfix is not a spread weapon in practice, whatever its six pellets suggest. Multi-hit rate by firing distance: shotgun 8.2% / 17.5% / 27.3% / 34.0% / 27.7% across 0-1.5, 1.5-2.5, 2.5-3.5, 3.5-5 and 5-8 tiles; Friday 9.0% / 8.3% / 0.3% / 0.1% / 0.0%. Its 45px spread lands inside one enemy's silhouette at exactly the close range the fast-path gives it. So its close-cluster priority is defensible on DPS — it is the fastest killer in the game — but not on hitting several enemies, and in its own band the shotgun spreads 2.1x better.
Weapon selection for the shotgun and ghidra is structural, not economic (2026-08-20, measured). pickRangedWeapon opens with a cluster fast-path that never reaches scoreRangedWeapon: on two or more clustered enemies it hands out Friday Hotfix inside 2.5 tiles, ghidra beyond 5 if the threat is not an Edge Case, and then falls through to the shotgun with no distance, HP or archetype test at all. The observed shot distances sit on those constants — Friday p50 2.8, shotgun p50 3.0, ghidra p10 5.1 — so essentially all of both weapons' use comes through that one branch.
Consequence: retuning scoreRangedWeapon will not move either weapon's usage. The scorer picks the shotgun in exactly one cell of an HP x distance sweep (HP ~150, distance <=1.5) and only 2.5% of real shotgun shots are in it. Two ~55-minute ghidra A/Bs came back null for this reason, having been designed off a synthetic probe of the scorer. If you want these weapons used differently, change the fast-path's shape, not the economics.
The reload term is what decides ghidra specifically. reloadsNeeded charges the reloads a kill consumes, so a 1-round magazine pays a full 1.6s on every single-rocket kill — for a reload that happens after the target is dead. Deliberate, and documented as such (it exists to stop the scorer preferring shotgun and ghidra "by roughly a reload each"), but it is worth knowing the size: charging only the reloads needed before the killing shot takes ghidra from 2.95s to 1.35s against the pistol's 1.54s, i.e. from never-chosen to chosen. The shotgun still loses either way.
A capture without a null-control arm is not readable (2026-08-20, measured). The density A/B ran a third arm byte-identical to the treatment, purely to measure this harness's own noise. It earned its cost immediately: at the pre-registered endpoint the treatment effect was -3.3pp against a +3.3pp null gap — indistinguishable — while at the endpoint where the effect was real it was -23.6pp against -5.8pp. Without the control both numbers are unscaled, and the temptation is to read the first as a small real effect. Budget the control arm in from the start; it is the arm that looks droppable when time is short and is the one that makes the other two mean anything.
Condition on resolution, not on attempts. The same capture's headline rates are per-attempt, and they hide the effect: an attempt that died at L9 is neither a success nor a failure at L12. Conditioning on reaching L10 turned a 6.7pp per-attempt difference into a 23.6pp conditional one, on the same data.
Read a staged-repo capture as a statement about that staging. stage-campaign selects files spanning the repo's difficulty range, so its levels are far more complex than a repo's average. The same density change was +15.7% enemies on the corpus average and +42.6% on staged curl. Any effect size from a staged capture needs that multiplier stated next to it, or it will be quoted as if it described normal play.
Enemy density is the attributable difficulty lever; archetype mixing is not (2026-08-20). COMPLEXITY_PER_EXTRA_ENEMY went 10 -> 5 because at 10 it was very nearly inert — real complexityScore is p50=1, so 95.5% of entity rooms held exactly one enemy and "packs scale with complexity" was mostly notional. Two facts to keep before touching it again: a cap is a no-op (there is no per-room tail — the 522-enemy laravel level is hundreds of one-enemy rooms plus corridor Edge Cases), and a floor of 2 is +90.4%. Judge any change here on damage taken, never on HP placed.
Inside an entity room, every archetype move is a DPS cut. Only normal (4.21/s) and edgeCase (1.68/s) are available: flagging a member elite would break "one Elite per Elite room", which stage-campaign.mjs and every archetype report rely on, and would drag in 2x melee, a larger sprite and a guaranteed drop. There is no "tougher" archetype to reach for, so per-room DPS neutrality is arithmetically impossible for any swap-based rule — neutrality can only be an aggregate property, held by firing rarely.
That is why "a private method reads as guarded" is expressed on the HP axis instead: a lockable room splits its pack [2s, s, s, …] at unchanged total, which moves instantaneous DPS by exactly zero while moving integrated damage-taken ±12% depending on which member the player kills first. HP shape is orthogonal to ENEMY_WEAPONS; archetype is not. Reach for HP shape when you want texture without a difficulty change.
Per-entity AST facts are the wrong place to look for archetype variety, and that is now measured rather than suspected (2026-08-20). The four proposed additions were checked before building any of them: recursion 10.2%, async 1.5%, deprecation 0.2% of 7,101 entities across 8 languages — the last two worse than the worst fact already available. Together they reach 12.3% of the entities that currently carry no fact at all.
And a cheap fan-in proxy does not exist. Both the raw identifier count and a call-site-only count measure name commonality: correlation with name length -0.370 and -0.300, mean name length ~8 chars in the high group against ~14 in the low, and the top "hubs" are args and path. Any future proposal to approximate call-graph fan-in by counting occurrences should be checked against name length first — it is the tell, and it is cheap to compute.
What to use instead. The only per-entity quantities with full coverage and no language skew are complexityScore and the line span, both already extracted. Key new gameplay off those, or off a room's role and position in the level, rather than off new parsing — it is cheaper and it works on every language.
A fact-driven archetype rule has a low ceiling, and the ceiling is language, not code. Measured over 8,143 corpus entities: private/protected 33.8% overall but 10% (django) to 53% (ripgrep) — C and Rust mark most functions private, Python and PHP barely at all; nestingDepth ≥ 2 14.9% and already triple-used by terrain (labyrinth gate, room footprint, pillar exclusion); switchBranches > 0 4.7%; allocation-dense 2.0%; and 54.7% of entities carry no fact at all. Before proposing another rule over the existing three archetypes, check what fraction of rooms it can reach — the answer has twice been "almost none". Archetype variety needs a fourth archetype.
Density has a second-order reach that offline tests catch and reasoning does not. Raising it broke verify:campaign's Bug-outcome coverage — not by changing any hash, but because a TODO encounter needs a free floor tile near its terminal and denser rooms crowded it out. Coverage that rests on a single instance will break on any change that shifts the generator; the fix is more instances, not a relaxed check.
Simulation constants are hashed as tables, not as picked scalars (2026-08-19). SIMULATION_BALANCE folds in COMBAT_BALANCE and TRAP_BALANCE whole, the same way it already took WEAPONS, rather than naming individual constants. Naming them is what went wrong before: the comment listing five uncovered constants had itself gone stale, locating two of them in modules they had since left, while reading as authoritative. simulationBalanceCoverage.test.ts enumerates each module's value-carrying exports — tables as well as numbers — and fails naming anything missing, so "wholesale" is enforced rather than claimed. Adding an enemy constant now requires no edit in engine.ts at all.
Both changes move the hash and invalidate every shipped replay, which is why the sequencing matters and is not negotiable: batch the balance work, close the hash last, regenerate defaultHighscore.ts once, afterwards (~33 min). Regenerating before the design settles costs a second one, which has happened.
Recurring idea, measured and declined (2026-08-20): extract recursion, call-graph fan-in, async and deprecation per entity, and use them to drive a wider set of enemy archetypes. A prevalence check over 7,101 callable entities across 8 repos in 8 languages, run before building anything, says the route does not exist.
- Prevalence: recursion 10.2%, async 1.5%, deprecation 0.2% —
against the facts already extracted, which this was meant to improve on:
private/protected 33.8%,
nestingDepth >= 214.9%,switchBranches4.7%. Two of the four are rarer than the worst fact already available. - The one that looked transformative is invalid. The proposed cheap fan-in
proxy — a workspace-wide identifier-occurrence count — reads 50.2% at >=10
occurrences, but it measures name commonality, not call-graph position:
mean name length is 7.8 characters in the high group against 14.8 in the low.
Restricting to call sites (
name() does not rescue it; correlation stays -0.300 and the per-language spread runs 12.5% to 72.3%, the same language-marker problem private/protected already has. Real cross-file resolution — imports, namespaces, overloads, dynamic dispatch — is not a small change. - The number that decides it: of the 51.2% of entities carrying no existing fact, all three viable facts together reach 12.3% of them. The factless share would go from 51.2% to ~44.9%. That does not unblock a fourth archetype.
- Caveat, so the verdict stays re-testable: these were text probes over each entity's source span, not real AST extraction. They match the proposed implementations closely (the idea itself specifies same-name matching for recursion and an identifier count for fan-in), and deprecation at 0.2% would survive any measurement error.
So the conclusion is about the approach rather than these four facts: archetype variety cannot be driven by per-entity AST facts, because every candidate is rare, language-bound, or not cheaply resolvable. The facts with full coverage and no language skew are the ones already extracted.
Terrain is a separate question and this does not answer it. Async and deprecation at 1.5%/0.2% would be flavour on a handful of rooms, which may be worth it on its own merits — it is just not an archetype lever.
This is the cost that is invisible from the idea's wording, and it applies to
any reason for extending the parsed AST, not only this one. astHash is
hashRun(JSON.stringify(parsed), campaignName) — so a new field on CodeEntity
changes every AST hash, which makes every shipped replay in
defaultHighscore.ts and every stored highscore entry stale, because playback
compares the hash and refuses on a mismatch. Batch AST-shape changes together,
and expect to regenerate the default board (npm run generate:default-highscore)
as part of the same change.
One design question would also have to be settled before extracting anything:
"async" does not mean the same thing across 15 languages (async fn, suspend,
a goroutine, Task<>, an IO type). Either the game wants one uniform "this
entity does something concurrent", or it wants per-language flavour — and that
is a game-design choice, not an extraction detail.
Every sound in audio.ts is a bounded one-shot that schedules its own stop() and is never referenced again — bgm.ts states the invariant outright, and it is what keeps the module free of lifecycle. Toolchain is the one weapon that breaks it, because it is the only weapon that is held.
Firing a 0.16s buzz once per 0.35s bite left ~0.19s of silence between every buzz, which reads as a tool that keeps stalling rather than a motor that is running. The same gap existed in the viewmodel, for the same structural reason: meleeRecoil decayed under the 0.02 overlay threshold in ~0.29s at 60fps, inside the 0.35s bite interval, so for the last few frames of every cycle the renderer drew the equipped ranged weapon instead and the chainsaw visibly blinked out. A playtest reported it as the chainsaw "stabbing". Both halves were one bug wearing two coats: a continuous thing implemented as a repeated one-shot, with a hole in the middle.
Two rules came out of it.
- Anything the player holds is modelled as a state with a start and a stop, never as a fast enough repeat. The motor is
startChainsaw/revChainsaw/stopChainsawwith retained nodes; the bite cadence now only revs it. The viewmodel equivalent ismeleeActiveplus a non-zero rest value (MELEE_HELD_REST) for the recoil to settle to, rather than a decay that happens to outlast the interval. TuningfireIntervalSeccan no longer silently reintroduce either gap. - A sustained node must be torn down explicitly.
stopChainsawdisconnects every node it retained, theWaveShaperNodeincluded. This is the same failuredistortionNodealready documents — a shaper outliving its voice accumulated edges until the tab died, within seconds of holding this exact trigger. A test asserts the disconnect rather than trusting the comment.
This bends, deliberately, the rule two entries above — that a weapon's audio cue belongs in audio.ts and not the engine, because WebAudio scheduling off ctx.currentTime is sample-accurate while an engine-driven trigger jitters on frame boundaries. That reasoning holds for events, and every one-shot still follows it. It cannot hold for a sustained sound, whose start and stop are by definition engine state (is the key down, is the player alive, is the game paused) rather than a moment in a schedule. The engine therefore drives start/stop, and audio.ts still owns the entire recipe.
A procedural sound is diagnosable from its recipe, and this one was shipped without doing that. The first motor here was two harmonically-related sawtooths (55Hz and its octave) under a smooth sine tremolo, with no noise at all — which is a synth bass through a gate, and got reported as "sounds like scifi techno". Each of those three properties independently pushes a sound toward "electronic": harmonically-related oscillators spell a chord the ear hears as a note, a sine LFO is a tremolo pedal, and a mechanical sound without broadband noise has nothing mechanical left in it. A real chainsaw is mostly noise — the chain, the exhaust, the chips — with the engine supplying only the low body and the pulse rate. The rebuilt version leads with a looped noise bandpass, uses one oscillator off a musical pitch, and chugs on a sawtooth LFO whose asymmetric kick-and-decay is the combustion. Those three properties are asserted in audio.test.ts, because they are the ones that silently drift a machine back into a synthesizer.
Note what no automated run can check here. isSilenced() mutes all SFX under navigator.webdriver, so every Playwright, verify and balancing run is deaf by design. The tests can prove the graph is built, revved and disconnected; only a human can say whether it sounds like a chainsaw.
Every pickup in the game is deliberately the same object — a dark square with a brighter square inside it — so that colour alone separates health from ammo from a keycard (see Colours and Pickups). That rule is fine within the "walk into this" category. It broke the moment a proximity mine was drawn with the same primitive: the one entity that deals 32 damage was wearing the collectible uniform, and the only thing left distinguishing it from a red keycard was about 41 points of red channel. A playtest report closed the loop — "red key and mine look almost the same, i actively avoid it by reflex" — which is the worst possible outcome, since the player was now avoiding the pickup rather than the hazard.
The rule this establishes: category is carried by silhouette, colour only separates items inside a category. A mine is a dome on a base plate with prongs; a keycard is a flat square. They differ in outline, which is the cue that survives the two conditions where hue does not — distance, where the old discriminators (the key floats, the mine pulses) both collapse toward the horizon, and red-green colour blindness, under which the old pair were literally the same object.
Two things were deliberately not changed. The key keeps its red, because it has to match the door it opens. And the mine keeps its red and its pulse, because both are load-bearing from an earlier fix for "mines are too easy to miss" — the point was to buy identity without spending conspicuity. The dome's body tone is fractionally lighter than the rectangle it replaced, which is compensation for the much smaller dark mass a dome covers, not a restyle.
Why no test caught it, and why the bot could not. Nothing pinned the mine's world colours or its shape; the sprite test asserted only that fillRect had been called. And the playtest bot is structurally blind to this class of defect: it reads mine.visible off an engine snapshot, never off pixels, so it can prove the mechanic is survivable while saying nothing about whether a human can identify one. Any future "these two must not look alike" bug needs a shape assertion — sprites.test.ts now pins that a mine draws discs and a keycard never does.
Every generation-time hazard/encounter system keeps a minimum distance from the player's spawn point, on the principle that the player should never be able to take damage before they've had a chance to react — this started as an explicit rule only for placeTraps' corridor scatter (spike traps/mines checked against an avoid list containing spawn) and placePillars/placeAmmoPickups (which skip the spawn room outright), but the TODO/FIXME "technical debt" encounter (see Game Design) was missed — it could resolve to the spawn room itself with no spacing check, occasionally placing a mine or spike trap immediately next to the player. Generalized rather than special-cased for mines: the encounter now reuses the same spacing constant/distance check as placeTraps, and is skipped entirely (not rerolled) if every candidate tile is too close, matching that function's existing "never a hard failure" contract. [notes: Task 51]
The same rule constrained two of the four AST-driven features added later. A Switchboard spur's encounter reuses the identical spacing check before placing a trap or mine (and simply places nothing damaging if every candidate tile is too close, matching the same never-a-hard-failure contract). An Exception Handling Zone is stronger still: its try gauntlet is unavoidable acid plus an invisible mine, so rather than spacing it, the spawn room is excluded from being an anchor at all — there is no distance at which "the corridor out of your spawn room is a damage gauntlet" is acceptable. Vendor Depots, which attach to the spawn room by design, therefore carry no enemies, traps or hazards whatsoever; their stock positions are also fed into the trap-placement avoid list, since a depot mouth is a one-tile choke point right next to spawn and is otherwise exactly what the corridor trap scatter looks for.
Worth stating what this rule does not cover, since it was briefly over-applied: Switchboard spurs are eligible on the spawn room. The rule is about unavoidable damage, and a spur carves none — its worst content is a single weak Edge Case enemy behind a closed door the player has to push open, and the trap/mine outcome is already spacing-checked against spawn. Excluding the spawn room there was cargo-culted from the two features the rule genuinely does bind (exception zones, acid overflows), and it silently cost every level whose first entity holds the switch — i.e. exactly the levels where a player is most likely to meet the feature at all. demo-campaign/stage02_bootstrap.sh was one: a case statement in its opening function, producing nothing.
Until the "Acid Overflow" room event, the engine had exactly two runtime grid mutations — opening a key-locked door (3→0) and revealing a secret wall (6→0) — and both were terminal and zero-valued. That was not an accident of implementation; it is what lets a multiplayer guest apply the host's grid corrections outside its exact-tick gate. applyGridReconciliation's own doc comment states the invariant directly: because every emitted TileMutation.value is 0, applying a stale or future snapshot's grid can at worst re-open an already-open tile, never set one wrong. That decoupling was itself a deliberate fix (a snapshot discarded by the tick gate used to lose its grid corrections permanently, since the host drains its shared delta each interval).
An acid mutation is 0→HAZARD_TILE — non-zero — and the delta is additive-only, so a guest that mis-predicted room entry could never take a phantom flooded tile back. Routing acid through gridDelta would therefore have forced one of two regressions: weaken the invariant, or move grid application back behind the tick gate and undo the earlier fix. Neither was acceptable for a cosmetic-adjacent room event.
The shipped design keeps acid off gridDelta entirely. Each overflow room ships a precomputed, ordered candidate tile list built at the very end of generation from the final grid, so it provably contains no door, teleporter pad, spike, secret wall, lore terminal, pre-existing hazard, or tile a mine or key is standing on. At runtime the grid is then derived from two scalars per room (when the flood started, and where it froze), which ride the reconciliation snapshot like the mine state does — behind the tick gate, alongside the other PRNG-coupled state. The count of tiles actually written stays local and off the wire, and every tick reconciles the grid to whatever those two scalars imply, in both directions. That bidirectional reconcile is the whole point: it can retract a tile, which an additive delta structurally cannot, and it is only sound because the candidate list is disjoint from every other tile-claiming system by construction rather than by convention.
Two consequences worth not "tidying up" later. The planning pass draws no RNG at all (its expansion order is a plain breadth-first walk from the room centre), which is what lets it be appended dead last in generate() without perturbing a single earlier draw — a file with no allocation-dense function generates a byte-identical map either way. And gridVersion is deliberately not bumped for acid: it changes neither walkability nor anything the minimap wall layer or the shared path field cache on, so bumping it would invalidate two caches for nothing and desync from the guest's own gridVersion assignment. The corner minimap's acid markers are recomputed per frame instead, the same shape as the spike-trap markers — the static hazard list is painted with no grid re-check and would have to become retractable to grow at runtime.
Every level used to read as a string of beads: a few rooms scattered across a mostly-empty grid, joined by enormous one-tile corridors, with a metronomic chain of identical little rooms strung along them. Two lines caused it. connectRooms chains rooms in parse order, and tryPlaceRoom placed each room at an independent uniform-random position — so the corridor between two consecutive entities was however far apart two unrelated draws happened to land. Measured over the demo campaign: a mean of 59 tiles between consecutive room centres, peaking at 191, against rooms only 4-18 tiles wide. Everything else followed from that. The corridor-breakup pass had to inject 277 near-identical rooms across 17 levels (21% of all floor area) to keep the sightlines down, and the map had to be sized generously enough for 200 random draws to keep finding gaps, leaving 86-93% of every grid as solid rock. The previous version of this section named proximity-based chaining as the fix and deferred it as out of scope; this is that rework.
tryGrowRoom grows each room off an already-placed one — anchors ordered by distance to the room connectRooms will actually carve to, so the predecessor is first and, when it's boxed in, the next candidate is its nearest neighbour rather than merely the next-most-recent room. growRoomCandidate guarantees at least 2 tiles of span overlap with its anchor, which is what makes the connecting corridor a mostly-straight run instead of a long reaching L. The old random scan survives as a fallback, but with a nearestTo argument: it evaluates its whole attempt budget and returns the closest fit rather than the first, because taking the first random fit put a map-spanning leg straight back into the chain for exactly the one room that needed the fallback. Squeezing the grid without that made it worse, not better — legMax went back up to 113 on the monolith before both orderings were fixed.
The gap between rooms is load-bearing in a way that isn't obvious. ROOM_GAP_MIN/ROOM_PACK_MARGIN are not spacing aesthetics: placeSecretRooms, placeVendorDepots, placeSwitchboards and placeExceptionZones all claim untouched rock plus a one-tile margin via sideCandidateFits, so packing rooms flush would starve all four at once and silently delete four features from the game. npm run verify:campaign is the gate that catches it, since it fails when any feature stops appearing anywhere.
connectLoops then adds a few extra corridors on top of the parse-order chain, between rooms that ended up close in space but far apart along the chain. The chain stays the spine — progression still follows the source file's reading order, and connectRooms alone is still the reachability guarantee — so these only ever add connectivity. A shortcut must never cost the player a key, though, and that constraint is sharper than it looks: placeDoors turns every corridor mouth of a private/protected method room into a door and placeKeys bills one key per doorway, so a shortcut into a locked room lengthens the route by a key-fetch while still reading as a shortcut on the map. Locked rooms are excluded as endpoints and each carve is then verified rather than predicted — doorway counts taken before and after, and the corridor undone tile by tile if any moved. Verifying is what makes it cheap: a corridor's path depends on jittered waypoints, so a geometric test would have to be conservative enough to reject most of the candidates that are actually fine (it was tried; it rejected almost everything).
Map size follows the rooms rather than the file. mapSize sums what roomDimensions will hand out and multiplies the square root by a calibrated ROOM_SPREAD (the packed cluster's bounding side measured 1.56-2.72x that across the campaign), plus a ROCK_RESERVE border for the carvers above. Lines of code was a reasonable proxy only while rooms scattered.
Measured across the 17 campaign levels, before → after: mean corridor leg 59 → 22, longest 191 → 48, longest straight run 34 → 11, floor density within a level's own bounding box 22% → 37%, and 20 loops where the tree topology had 5 (all of those accidental corridor crossings). Every entity still gets a room on every level — report:level-maps prints rooms/entities precisely so a silent drop is visible. [notes: map layout rework]
The pass that keeps straight corridor runs under MAX_CORRIDOR_STRAIGHT_LENGTH was deliberately designed around a first working version's failure mode: repeatedly rescanning the whole grid and bisecting whatever's still too long converges correctly on a small/typical map, but on a dense one it snowballs into far more injected rooms than the level needs without ever fully converging, since each rescan's remaining free space keeps shrinking. The shipped design instead splits every run found by a single initial scan into evenly-spaced segments in one shot (cheap, proportional to length, well-distributed), with only a small, bounded number of wide-search safety-net rescans afterward to mop up genuine edge cases (two unrelated legs landing collinear, a couple of adjacent target points both failing in the same locally-blocked spot). That part still stands, and it is still explicitly best-effort, matching every other placement system in mapGenerator.ts (traps, pillars, TODO encounters) — an unusually large or dense map can still end up with a run over the nominal limit, and that's an accepted trade-off rather than a bug to keep chasing.
What changed is what it places, and why. The original vocabulary had two entries — a small room with a baffle wall, or a forced jog — which was enough while corridors were enormous and a level got dozens of interruptions that blurred together. It was also the single most repeated structure in the game: 277 instances of the box across 17 levels, in 9 footprints that were all the same construction, strung at even spacing down long straight hallways. dressCorridors now draws from six kinds (baffleRoom, pillarHall, gateway, plaza, alcovePair, chicane) over a shared widening primitive, refuses to place the same kind twice running on one corridor, and jitters the spacing. baffleRoom is deliberately no longer the default weight.
It also gained a second pass, and that one is about the game rather than the geometry. Once rooms were packed close together, almost no corridor exceeded the length limit any more — the length pass fired 6 times across the whole campaign. That would have quietly deleted the corridor encounter from the game, because spawnEdgeCaseEnemies populates these rooms and nothing else; the geometric half of the fix would have removed the gameplay half described in Game Design. So an ornament pass spends a room-count-scaled budget on corridors that are already short enough, purely so corridors have character and encounters have somewhere to live. Campaign totals: 84 features against the old 277, every level with at least one, and Edge Case enemies down from 573 to 177 — a large enough drop that stored balancing telemetry taken before this is not comparable.
Every treatment that turns cells back into wall (a baffle, a colonnade, a gateway neck, a plaza's corners) records which cells of its footprint were already floor and refuses to seal one. That rule is not theoretical: sealing a cell an unrelated corridor was routing through once isolated 3 rooms from spawn on a real campaign level, with no route existing at all afterward.
A dependency key belongs to a gate — one locked room — carries a colour, and opens every door of that gate, permanently. It is not consumed.
The engine used to charge one key per doorway, which made a six-mouth room six
key hunts for one space and charged again for a door you had already opened from
the other side. MAX_GATES bounded that by pricing rooms in doorways and
preferring one-mouth rooms; per-room keys remove the reason for that penalty, so
the budget is now MAX_GATE_ROOMS (rooms) and ranking is worth alone.
Three constraints hold this together, and none is optional:
- The generator and the engine must change together. Issuing one key per room
while
openDoorAheadstill spent per doorway was tried on 2026-08-08;assertAllRoomsReachablereported rooms unreachable, correctly. - Every gate has exactly one key.
placeKeyswidens its search through tiers of convenience exclusions before it will give up, and un-gates the room rather than shipping a door with no key. Exception-zone tiles are never relaxed: that exclusion is a measured gameplay decision, not a convenience. - Gate identity is a side table on
GameMap, never newTilevalues.Tileis a closed union mirrored by five hand-maintained solid-tile sets across the engine and scripts; a value per colour would multiply all of them.
In coop the key is granted team-wide on pickup. With one key per gate, a per-player grant strands every teammate but the one who walked over it, and the bots plan routes with no model of a teammate's inventory — so a session wedges at the door rather than slowing down. That is also why there is no key loot drop.
Every room being reachable from spawn was assumed to follow automatically from connectRooms chaining rooms 0..n-1, but two gaps meant it didn't always hold: placeRooms could end a level with a single room (an empty file, or one entity that's the only one that fits), and connectRooms only carves a corridor once a second room exists, so that lone room silently got zero corridors; separately, carveLabyrinth's recursive-division maze (for any entity nested ≥2 deep) kept exactly one connecting gap per split, but a child split's own later wall could land on the exact cell a parent split's gap depended on to reach it, sealing off part of a room despite every individual split "keeping a gap" — confirmed empirically at roughly 9% of seeds across the room shapes/nesting depths generation actually produces. Fixed with a room-count floor (placeFillerRoom, topping a level up to at least 2 rooms; the synthetic filler/fallback entity uses kind: "class" rather than "function" so it can never spawn an enemy whose nameplate would show a placeholder name) and a post-carve labyrinth connectivity repair (flood-fill the room's floor into components, bridge the closest pair via a shortest wall-crossing BFS, repeat until one component remains) — both deterministic, no extra RNG draws, so neither perturbs existing replay determinism. A permanent BFS check (assertAllRoomsReachable, end of every generate() call) now also verifies this invariant and logs loudly if it's ever violated again rather than shipping a silently broken level. It is deliberately not gated behind import.meta.env.DEV — a shipped build is exactly where nothing else would catch a violation, since an unreachable room is unwinnable and silently so; the cost is one BFS per door on a pass that already does far more work than that, paid once per level load rather than per frame. (This entry described it as "dev-time" until 2026-08-26, which the source has contradicted since the check was written. Note also that it console.errors and returns rather than throwing, so "assertion" is the wrong word for what it does.) This is a deliberately different guarantee level from Corridor Dressing: A Vocabulary, Best-Effort above: a corridor's length staying under the nominal cap is still explicitly best-effort, but a corridor existing at all between every room and the exit is now a hard invariant, not a probabilistic one. connectLoops' shortcuts are held to the same line — they only ever add edges, so they cannot weaken the invariant, and assertAllRoomsReachable still runs after them. [notes: sometimes rooms without exit are generated item]
localStorage quota is treated as a real, expected failure mode, not an edge case — a replay payload spanning a multi-level campaign can reach multiple megabytes. The mitigation is layered and ordered by how much it costs: compress the highscore board before ever writing; if a write still fails on quota, retry without the new entry's replay; if that alone still doesn't fit, retry without any stored entry's replay. A run's score is never allowed to be dropped — only the optional replay attached to it. A version prefix distinguishes each storage generation, so old data keeps loading unchanged — readBoard (highscores.ts) accepts all three: today's "bin1:", the older "gz1:", and bare JSON from before either existed. [notes: Task 44]
Why the board is binary-packed and not just gzipped (replayCodec.ts). Gzip alone left far more on the table than it looked. JSON.stringify writes every field of InputSnapshot on every frame — 18 keys, overwhelmingly false/0/null, about 334 bytes per frame — and gzip folds that repetition away nicely. What it cannot fold is dt: measured on the shipped board, 213,740 frames carried 14,425 distinct dt values, all rAF jitter in the twelfth decimal (0.01666666666662786 and neighbours). Each serializes to its own ~20-character string, so the compressor's dictionary never gets a repeat, and that one field ended up dominating the compressed payload. Packing a frame as a held-key bitfield + a flags byte + the raw dt bits (plus an "extras" group needed on under 1% of frames) sidesteps both problems: 1.48 MB of quota → 0.42 MB, measured end to end.
Worth knowing when reasoning about quota at all: localStorage is billed in UTF-16 code units, so a base64 character costs two bytes of quota, not one. The old board's 756 KB of base64 was really ~1.48 MB against the limit.
The codec is exactly lossless, and that is load-bearing rather than tidy. dt is stored as raw IEEE-754 float64 bits — never quantized, never narrowed to float32, which cannot represent any of the 213,740 recorded values exactly (checked, not assumed). A replay is a deterministic input stream, so a rounded dt would drift playback away from the score the run actually recorded, and balanceHash would not catch it: the balance fingerprint would still match, because nothing about the balance changed. Two cheaper-looking encodings were measured and rejected — XOR-delta-encoding dt against the previous frame (Gorilla-style) came out larger after gzip than the raw bits, and base64-ing the packed frames into the existing JSON instead of writing raw bytes into the gzip stream costs 1.31x, since it hands gzip a 4/3-inflated alphabet it only partly wins back. Quantizing dt to whole microseconds is ~2x smaller again and is the one worthwhile lever left, but only if the live game loop quantizes too — otherwise playback and the recorded score disagree by construction.
Separately, a real scoring bug was traced to two causes at once: health lost was being penalized twice (once via a health-fraction bonus, once via a separate malus scaled the same way), and EngineCarryover — rebuilt fresh every level — never forwarded prior score, kill points, or bonuses, so score collapsed back toward zero at every level transition instead of accumulating. Both were fixed together: the duplicate malus was removed, and carryover gained a priorScore field threaded through every level-transition path. [notes: Task 47]
Loot rolls are filtered to exclude outcomes that would be dead weight given current player/run state, rather than rolling them anyway and wasting the result: rocket ammo (and, since gdb and Friday Hotfix each got their own pool, smg/gas ammo) is excluded from the loot table entirely until the matching weapon is owned (its share redistributed to other kinds). The design goal is that a drop the player can't use at all reads as worse than no drop — see Game Design. gdb's ammo split from a shared bullets pool into its own "smg" pool once a full-auto weapon draining a pistol-sized pickup in under a second stopped reading as a distinct resource identity — same reasoning, applied a second time as gdb's role diverged further from the pistol/shotgun's. Friday Hotfix (the flamethrower, "gas" ammo) followed the same pattern from the outset — a fourth full-auto weapon sharing any existing pool would have hit the exact same problem immediately. The shotgun was split off last, into "shells", for a different reason than the other two: not that it drained the pool too fast, but that it and the pistol are both starting weapons, so one shared pool made them rivals from the first minute of every run. Its ammo is never gated by ownership (unlike smg/rockets/gas) precisely because it is a starting weapon. [notes: Task 46, 50, 57, 67]
A regular (non-elite) kill also has a very small (NORMAL_KILL_WEAPON_DROP_CHANCE, 1%, loot.ts's rollBonusWeaponDrop) independent chance to additionally drop a still-locked weapon, gated the same way as the Elite bonus drop (only rolled when a weapon is actually missing) and stacked on top of the kill's normal loot roll rather than replacing it — deliberately rare, so Elites and secret rooms stay the reliable unlock paths and this only ever reads as a lucky bonus. Elites themselves later gained the same independent roll at a much higher ELITE_BONUS_WEAPON_DROP_CHANCE (60%) — their prior behavior was an either/or choice between a guaranteed weapon and a guaranteed health/ammo drop; the health/ammo drop is now always rolled, with the weapon as a genuinely separate, stacked bonus on top. [notes: Task 74, 75]
Toolchain (the unlockable chainsaw) rides this same Elite/secret-room drop machinery but is deliberately not a member of UNLOCKABLE_WEAPONS — that constant is what both the regular-kill 1% roll and the secret-room/Elite candidate lists are built from, so keeping Toolchain out of it means it can't surface via the ordinary per-kill roll or the secret-room/Elite candidate lists automatically; both fold it back in separately (main.ts's computeMissingWeaponIndices, dropEliteLoot), gated by TOOLCHAIN_MIN_LEVEL (campaign level 4). It also has no FORCED_UNLOCK_LEVELS entry — every other unlockable weapon guarantees itself by a fixed level as a safety net; Toolchain is the first deliberately-missable one. [notes: Task 78]
A 450-run, 9-combo balance-telemetry campaign (2026-07-15) found the loot economy needed real rework, not just tuning, once real play data existed to check assumptions against: every regular kill guaranteed a drop of something (a randomly-weighted kind, never nothing), and dynamic (kill-drop) supply so overwhelmingly dominated static level placement (93-100% of everything actually consumed, every resource type, every combo) that ammo never felt meaningfully scarce regardless of level design or difficulty's ammoDropRate. Fixed with two changes: every regular-kill (non-Elite) *_DROP_AMOUNT cut ~30%, and a new REGULAR_KILL_NO_DROP_CHANCE (20%) — a kill's ammo/swap roll can now come up empty, something that structurally couldn't happen before. Elite drop amounts were deliberately left untouched (a bigger reward for a harder kill, not part of the "every regular kill floods you" problem).
Health was cut identically at first, then had to be pulled back out into its own always-on, unconditional grant (not part of the weighted roll or the miss chance at all) after a live verification batch showed the combined cut collapsing Hard difficulty's qualifying rate to 4% (from a ~48% baseline) — health is the one resource that directly ends a run when it runs short, unlike ammo, which stays survivable via the knife/Toolchain melee fallback, so it needed a different, more forgiving rule than the rest of the economy rather than the same cut applied uniformly. rollLoot gained a healthHandledSeparately param so its own weighted table excludes "health" once the caller handles it separately, rather than risking a double-drop.
Toolchain gained a third acquisition path the same day: a regular kill's ammo/swap roll coming up empty now carries its own small independent chance (MISS_CHANCE_TOOLCHAIN_DROP_CHANCE, 5%) to grant Toolchain instead of nothing, on top of (not instead of) the pre-existing secret-room/Elite paths — added after the same campaign found those two paths close to unreachable in practice (a shortest-path bot never explores for secret rooms — see Balancing telemetry — and an early clean Elite kill is rare), a finding that plausibly generalizes to any player who beelines rather than explores.
Everything enemies deal went up 1.5x on 2026-08-23 (ATTACK_DAMAGE 10 -> 15, PROJECTILE_DAMAGE 8 -> 12, and every archetype with them, since ENEMY_WEAPONS derives from those two). The full derivation lives in combatConstants.ts; what belongs here is why it took that shape.
The problem was not distribution, it was scale. Three preceding changes all pushed player-side income up — enemies at 35 HP per complexity point instead of 25, Edge Cases at 25-35 instead of 10-15, and a kill's heal and ammo drop both made proportional to what died — so fights got longer and paid out more, while incoming damage stood still. The result was a health bar that mostly sat full.
Flat, deliberately, because redistribution is a lever class already shown not to move this. Three separate A/Bs redistributed damage between melee and ranged and left total damage taken invariant at 122/127/126. A fourth attempt at the same shape would have been indistinguishable from that history. Judge any future change here on damage taken, never on where the damage is placed.
It compounds with difficulty rather than being folded into it, which is the point: a bite is 12.75 on Easy and 22.5 on Hard, so Hard moved 2.25x relative to the pre-change Normal. That is intended, and it is the caveat to check first if Hard's clear rate falls off.
The raise moves balanceHash, so it invalidated every stored replay and the bundled highscore board was regenerated. That cost is paid once per release regardless of how many simulation constants move — see combatConstants.ts's header for the batching rule.
A kill's ammo drop and its guaranteed health top-up were both flat, so income was a function of kill count rather than of how hard the level hit. Both are now proportional to the dead enemy's maxHp (AMMO_SCALE_REFERENCE_HP = 88, the corpus mean; HEALTH_SCALE_REFERENCE_HP = 100), applied at the drop site rather than in the roll — so the loot table stays uniform across archetypes and only the payout scales.
Three details worth not re-deriving. swap stays flat because it was priced that way and is not an ammo pool. Both are floored at 1, so a rockets drop — base 1 — cannot round away to nothing on a small enemy. And the health grant stays unconditional: it is scaled, not made a roll, because running out of health is the one thing that actually ends a run.
The pre-change asymmetry was the finding this fixes: an Edge Case yielded the same expected loot as an enemy twenty times its size, which balancing-telemetry.md §5.1 had flagged and quantified.
Carryover was unbounded, so a long campaign trivialised itself — measured across 23 repositories, a late level handed the player nearly eighteen times the surplus an early one did, against opposition that does not grow with the reserve. Each pool is now clamped on entry to CARRYOVER_CAP_MULTIPLE = 3 times what that level would give a fresh player (ammo.ts, applied in the RaycasterEngine constructor).
A ceiling, not a floor — level 1 has nothing to clamp and the early game never reaches it, so the change is invisible until deep into a large repository. Capped per pool rather than in total, so a pool the level cannot supply fresh (rockets before ghidra is owned) is capped against its own flat reserve rather than against zero.
K was chosen at 3 from an offline sweep: below 3 the returns are small and the cost is not. Note that CARRYOVER_CAP_MULTIPLE is deliberately not in SIMULATION_BALANCE — it changes what the engine is handed, not how the engine simulates, so it moves no replay and forced no highscore regeneration.
Every overlay — status bar, weapon viewmodel, crosshair, both maps, toasts, floating labels and the blocking screens — is drawn in a fixed 640x400 design space and mapped onto the real canvas by a single ctx.scale in withOverlayScale (overlayScale.ts). Before this, each of them measured in device pixels, so the Sharp render preset (1280x800) halved the apparent size of all of them.
Two invariants hold the design up, and both are easy to break by accident:
withOverlayScaleis the only thing that ever scales the target. A second scale anywhere composes with it silently.- Nothing in shot resolution may import this module.
spreadPxandmaxConeDeviationPxare simulation screen-pixels resolved against thezBuffer; a scale factor leaking into aim would make where a shot lands depend on a render setting. This is why the two live inweapons.tsand not here.
The practical rule for anyone drawing something new: write the numbers as though the canvas were 640x400, and read the w/h the wrapper hands you rather than ctx.canvas.width — the latter is still the backing store and lands at twice the intended coordinate at Sharp. Click hit-testing needs the inverse of both conversions (client -> backing store -> design); stopping after the CSS scale puts every button at twice its intended position, which reads as "the buttons don't work" rather than as a coordinate bug.
TEST_HOOKS_BUILD_ENABLED = import.meta.env.DEV (engine.ts) is a build-time constant rather than a runtime check, specifically so a production bundler can fold the entire branch away instead of shipping dead code behind a flag.
A constant alone is not a control, so scripts/check-bundle-hygiene.mjs runs as build's postbuild step and fails the build on either failure direction: a test-hook global or the mutating debug-handle wiring reaching dist/, or an opt-in diagnostic that is supposed to ship (__codeensteinPerfStats, ?perfDebug, ?ablate) no longer reaching it. The second direction is not symmetry for its own sake — the previous version of this check was vacuous for a month after the Vite 8 bump, and only a both-ways assertion notices that.
Two things its header records that are worth keeping: assert on the hook globals, not on a substring of the gate expression, which changes shape with every refactor; and treat "no bundle found" as a failure rather than a skip.
Every ranged weapon is rate-capped by its own fireIntervalSec. The mechanism already existed — updateFiring honoured it for automatic weapons and for ghidra — but the pistol and shotgun simply defined no interval, so their fire rate was whatever the input device could produce and a shotgun could be click-spammed into an automatic weapon. This was a data gap, not a missing feature; the fix added two numbers and no engine logic.
Four decisions worth not rediscovering:
- The cooldown is per-player, not per-weapon.
weaponCooldownlives onPlayerStateand survives a weapon switch, so the shotgun's pump can't be switch-cancelled by tapping1and back. Quick-melee (Space) bypasses the whole path, which is the intended escape hatch — a player mid-cycle is never left with nothing to do. - A dry fire still starts the cooldown.
fire()early-returns without ammo butupdateFiringsets the timer either way, which also rate-limits the out-of-ammo toast instead of re-triggering it every frame. - The audio cue lives in
audio.ts'splayShotgunBlast, not the engine.viewKindis a weapon's sound identity in that module, and the pump is part of the shotgun's voice rather than a separate game event; WebAudio scheduling offctx.currentTimeis also sample-accurate, whereas an engine-driven trigger would land on frame boundaries and jitter at low FPS. WEAPONSjoinedSIMULATION_BALANCE, wideningcomputeBalanceHash's values fromnumbertounknown. Tuning a weapon changes how a recorded run plays back while leavingastHashintact — precisely the silent-divergence class that guard exists for. Folding the whole table in (rather than naming the two constants that moved) means future weapon edits are covered with no list to maintain. The cost is that it invalidates every replay recorded before it, including the bundled highscore board, which had to be regenerated.
The cadence cap and per-pellet damage are coupled and must be tuned together — see Game Design for the DPS arithmetic that forced the shotgun's 18 → 25 damage bump in the same change.
Damage is the lever for a weapon that reads as weak; cadence is not. gdb's 12 → 16 bump came from a playtest one-liner ("smg dmg too low") that turned out to name a real ordering problem rather than a feel: at 12 it produced 133 DPS against the starting pistol's 147, so a campaign-level-4 unlock killed slower than the level-1 weapon while also carrying the longest reload of the three bullet weapons and the worst damage per round. The enemy-health pass (HP_PER_COMPLEXITY 25 → 35) widened the gap rather than causing it. Two levers could have closed it and only one was right:
- Cadence was rejected.
fireIntervalSec0.09 is what makes the 45-round magazine "about four seconds of held trigger" — the weapon's stated identity. Shortening it would have bought DPS by spending exactly the property the number exists to produce, and drained the pool faster on top. - The ammo economy was deliberately left alone.
SMG_DROP_AMOUNT(21),STARTING_SMG_AMMO(40),VENDOR_SMG_AMOUNT(30) and the smg loot weights are all unchanged, so a round being worth a third more is a straight buff rather than a redistribution. Cutting supply in the same change would have made the result unattributable — the recurring failure mode where a lever moves damage around without reducing it.
The ordering the bump establishes is now asserted rather than commented: weapons.test.ts pins dps(gdb) between the pistol's and the shotgun's, and gdb's damage-per-ammo below both. This was measured offline only — the inversion is arithmetic, and gdb is ~64% of every shot in the telemetry archive, so a bot A/B would have moved every aggregate at once without being able to refute it.
Easy/Normal/Hard originally scaled three axes in lockstep: enemy HP, enemy-dealt damage, and ammoDropRate (Hard tougher/scarcer, Easy the reverse). A 450-run balance-telemetry campaign (2026-07-15) found this produced two real problems, both fixed the same day:
- Easy's damage no longer mirrors its HP reduction. The original symmetric curve (
hp: 0.7, damage: 0.7), combined with a cautious/health-seeking playstyle, could produce near-total immunity for an entire 17-level campaign (100% survival to level 15 in one profile, where the other two were already down to single digits).DIFFICULTY_MULTIPLIERS.easy.damageis now0.85— real stakes on Easy needed the damage floor raised independently of the HP floor. Hard's pair still mirrors (hp/damageboth1.5) — this asymmetry is Easy-specific. - Enemy aim precision is a new fourth axis (
enemyAimSpreadDeg: 10°/4°/0° easy/normal/hard). Enemy ranged bolts previously had zero aim deviation at all — always fired exactly at the player's live position — so the observed ~70-77% hit-rate band was purely a function of the player dodging, never anything difficulty touched; difficulty's whole bite was tougher/scarcer, never actually "smarter."spawnProjectile(projectiles.ts) now rotates a bolt's aim vector by a random angle up to the per-difficulty cap before firing. Hard's0°is deliberately the unchanged prior behavior (still dodgeable, since a bolt has real travel time and isn't hitscan/homing) — only Easy/Normal got worse aim, rather than redesigning Hard's ceiling too.
Both fixes were verified against live telemetry, not assumed from the constants alone — see Balancing telemetry for the accuracy numbers, including the Casual/Easy survival-curve and Gamer/Hard qualifying-rate verification.
Death used to end a run outright: onGameOver recorded the highscore, cleared the campaign save, and showed a one-button Kernel Panic whose only exit was the file tree. A run that died on level 14 of 40 lost everything, and the only way back in was a manual file pick that started over at score 0. Rollbacks (2 Easy / 1 Normal / 0 Hard, single-player only) are a budget of retries against that. The choices worth knowing before touching this:
-
Hard gets none, and that is the point. It is the one tier where death stays final and the Kernel Panic screen keeps its original single button; the badge never draws there either, since it hides at 0. A difficulty whose whole identity is a tighter margin should not also hand the run back. Two knock-on effects worth knowing: a zero grant is a legitimate state now, so nothing may treat it as "the feature is off", and
grantedRollbacks() === 0stopped being proof that the harness opt-out took effect — which is whygetRollbackStatereportsdisabledas its own field andassertRollbacksDisabledchecks that instead. Without it a harness running the hard combo would have readgranted: 0and called the opt-out confirmed whether or not it ever took. -
A rollback restores the level's entry state, not full health. The state is the
EngineCarryoverthe level was launched with, so the failed attempt's score and kills are discarded along with the damage it took. Restoring to full health instead would make a rollback worth more the worse the run was going, and turn "walk into the level hurt" into a reason to die on purpose. The map is identical (AST-seeded) butgameplaySeedis re-rolled per launch, so enemy timing, aim, loot rolls and spread all differ — deliberately, so a rollback is another attempt rather than a rehearsed one, but it does mean "restores the state you entered with" is true of the player and false of the world. Say so in player docs; it reads as a bug otherwise. -
The count is stored, never re-derived from the current difficulty. Difficulty is a standing preference living outside
CampaignSaveon purpose, so re-deriving would silently re-grant the full budget to anyone who switched to Easy mid-campaign.ROLLBACKS_BY_DIFFICULTYis read in exactly two places: a fresh workspace load, and backfilling a save written before the feature existed. -
A separate table, not a field on
DifficultyMultipliers. Every member of that interface is a multiplier applied to a simulation quantity, resolved once at engine construction and multiplied into hp/damage/loot/aim. A retry budget is none of those things and the simulation never reads it. -
A manual file-tree pick re-grants the budget, which diverges from
campaignLevelIndexandcheatsUsed, both of which leak across one. That divergence is deliberate rather than an oversight: a manual pick already resets health, ammo, weapons and score, so it is a fresh run in every respect that matters. The two that leak are defensible for their own reasons (a cheat penalty should stick; the level index is invisible) and are not a rule worth generalising — a player who picked a new file and silently had no rollbacks for the whole run, with nothing on screen saying why, is a different kind of surprise. -
The death screen must never state a bare remaining count. This follows from the point below and was reported as an off-by-one before it was understood as a wording bug: because the rollback is charged at the death,
remainingon that screen is the count after this one is spent — so on the last rollback the original wording read "0 rollbacks left" directly above a live "Roll back" button. The arithmetic was right and the sentence was still false, which is the worse of the two failures. Every line on that screen now describes this choice and what follows it ("This is your last rollback", "1 more rollback after this one"), and a test asserts no wording that offers a rollback may contain a bare zero. The in-play badge and the Continue Run log are a different case and correctly do show a plain balance — nothing is being offered at that moment. -
The rollback is spent at the moment of death, not when the button is pressed. If the decrement waited for the click, closing the tab at the Kernel Panic screen and resuming later would be a free retry — strictly better than pressing either button.
beforeunloadtherefore persists the already-charged entry snapshot. The honest cost of this is that an unrelated crash or OS update also costs a rollback; there is no third option, becausebeforeunloadcannot awaitrecordRunHighscoreto end the run properly either. -
Fixing that exposed a pre-existing bug worth recording.
beforeunloadfirespersistProgress(lastStats)wheneveractiveEngine && lastStats, andactiveEngineis only nulled byresetToFileTree— which has not run while a Kernel Panic overlay is up.lastStatsthere is the death frame. So closing the tab at a death screen wrotehealth: 0as the resume point, afterclearCampaignSave()had run, re-arming a Continue Run that died on its first frame. It had nothing to do with rollbacks; rollbacks just made the death screen a place you might sit for a while. -
A rolled-back attempt is dropped from the replay, not appended.
startLevelwould otherwise push a second segment for the same file path, and playback walks segments in order — the replay viewer'sonGameOversetsendedand onlyonWincallsadvanceLevel(), so playback would stop at the failed attempt and silently never show anything the player did afterwards.CampaignReplayRecorder.dropCurrentLevel()pops it so the recording is exactly the surviving path, which also matches the rule that the failed attempt's score and kills are discarded. -
The count reaches the HUD through a constructor argument, not
EngineCarryover. The siblingcheatsUsedbadge flag rides the carryover, so that was the obvious home. It does not work: level 1 of every campaign launches withcarryover === undefined, which would hide the badge on the level a player is most likely to die on, and synthesizing a partial carryover to carry it zeroes health (createPlayerStatereads health/swap from an unconditionalif (carryover)). The engine never reads the value for anything but the badge. -
A continued run still records, and is marked — and the board now carries the difficulty too. Not disqualified the way a cheated run is: rollbacks are a sanctioned part of the difficulty setting, and disqualifying them would make the feature useless to anyone playing for score, which is most of the reason to keep playing.
HighscoreEntry.rollbacksUsedis omitted when zero so a clean run and one recorded before the feature existed stay the same case.Recording the difficulty forced a second change: the picker now locks for the lifetime of a run. It used to be freely changeable at any moment, applying from the next level — the same "standing preference" shape gore and render quality have, and harmless for those two. It is not harmless here, because score is entirely difficulty-blind (there is no difficulty term anywhere in
scoring.ts): the tier is the only thing separating an Easy 8,000 from a Hard one. Left open, a player could clear fifteen levels on Easy, switch to Hard for the last, and post a Hard-labelled entry whose own replay disagreed with it level by level. Recording the difficulty did not create that hole, it made it visible.setDifficultyLockedcloses it — locked inlaunchLevel, released inresetToFileTreeand on a fresh workspace load — and thetitleon the disabled control explains itself, because a control that is simply dead reads as a bug. Gore and render quality are deliberately left alone: neither touches the simulation or the score.This shipped first as a marker with a stated weakness: the board had never carried a difficulty, so an
RBcolumn marked a row without making two rows comparable — a rank-1 Easy run with two rollbacks outranked a clean Hard run and the table said nothing about it.HighscoreEntry.difficultycloses that, and the two columns are only useful together. Worth knowing about the implementation: entries written before the field exists are not shown as unknown, because the information is already in their replay — everyReplayLevelSegmentrecords the difficulty its level was played at, sodifficultyOf()reads it back and the shipped default board displays real values without being regenerated (pinned by a test against the actual shipped bytes). The replay is only trusted when every segment agrees: a run whose difficulty changed partway has no single honest label, and quietly reporting the first level's would turn "switched to Easy at level 9" into a Hard entry on the board. -
The harness opt-out is not about the balancing bot.
playRunreturns the instant the engine reportsoverand never dismisses the overlay, so the bot cannot spend a rollback and per-level telemetry is unaffected. It is installed there as a guard against a future change to how a run ends, and it is genuinely required byverify-campaign-playthrough.mjsandgenerate-default-highscore.mjs, which both assert on the death path itself. Every call site asserts the resolved state (disabled === true) rather than trusting the install call, because a silently-ignored preset looks identical to one that worked — and note it must bedisabled, notgranted === 0orremaining === 0, neither of which can distinguish the opt-out from Hard's own zero grant or from a budget already spent.
Room decorations (server racks, plants, desks, code-block props as billboards) were implemented once, then disabled behind a flag after playtest feedback found them more annoying than atmospheric — the code was kept rather than deleted, since the underlying goal (rooms feeling less empty) is still considered worth pursuing with a different approach. [notes: Task 21, and the still-open decorations item]
The level-end/run-end player-facing stats screen (PLAYER_STATS_ENABLED, src/engine/playerStats.ts) was cited here as an instance of that pattern — shipped, then disabled by default after measurement. Corrected 2026-08-23: there was no measurement. The 2026-07 commit asserted that recording telemetry on every playthrough slowed gameplay and recorded no number anywhere. Two A/Bs since have failed to reproduce a cost — 2026-08-02 (on a build already dropping ~29% of frames, so barely sensitive) and 2026-08-23 (on a build holding the vsync edge: +0.025ms busy against a 0.16ms calibrated floor, level-end derivation 0.00027ms). The flag now ships on. The pattern this paragraph describes is still real — DECORATIONS_ENABLED remains a genuine instance — but this was not one of them, and it stood as a documented precedent for over a month on the strength of a claim nobody had checked.
The exit compass went through two disable/re-enable cycles: first disabled because its needle read as inverted on at least one axis and it didn't fit visually next to the minimap; root-caused to the needle's rest orientation pointing "east" instead of "up" (a target dead ahead should read as straight up, not sideways); fixed and re-enabled with a redesigned placement (a small badge in the minimap panel's corner rather than a separate bar), then further shrunk after feedback that even the corrected version felt too visually invasive, and shrunk again (the panel's maxPixels halved, 140→70) on the same "still too visually invasive" grounds. [notes: Task 33, 34, 35, 39, 62]
Billboard draw order was originally fixed-by-category (enemies drawn before items, before the exit marker, etc.), which could let a nearer item visually disappear behind a farther category. Replaced with one combined furthest-to-nearest depth-sorted pass across every billboard type against a shared z-buffer — see Architecture. [notes: Task 25]
The weapon viewmodel was drawn 2% right of true screen-center, while the bullet tracer/flame-stream (and the crosshair) have always used a fixed screen-bottom-center/dead-center point — so the gun visibly didn't line up with its own shots. First attempt fixed this backwards: made the tracer/flame-stream chase wherever the gun sprite happened to be drawn (including its head-bob/recoil jitter). Reverted in favor of the opposite direction — the tracer origin is the simulation-adjacent, shared reference point every weapon and the crosshair already agree on; the viewmodel is a purely cosmetic overlay, so it's the one that should conform. Fixed with a one-line change (WEAPON_CENTER_BIAS 0.52 → 0.5 in viewmodel.ts) instead of touching effects.ts/engine.ts at all. [notes: enhance weapon sprites item]
Several docs used to say the bot and the engine could not share code because "src/ cannot import from scripts/". That is true and it is the wrong direction — scripts/ imports src/, and since 2026-08-18 it does.
What makes it work, and the exact constraint. Node has stripped TypeScript types natively since 22.18, and this repo's floor is ^22.22.2 with every CI job on 24, so a plain .mjs can import a .ts module with no build step. But Node ESM does no extensionless resolution, so a module reached this way may not contain a value import written bundler-style. tsconfig.json already sets allowImportingTsExtensions, so writing from "../map/types.ts" satisfies Vite, Vitest and tsc equally. An import type is erased before resolution and needs nothing. Verified by running it: combatConstants.ts loads in bare Node today; player.ts fails with ERR_MODULE_NOT_FOUND on its extensionless ../map/types.
So the rule for any module meant to be shared is: no value import without an explicit .ts extension, and no DOM — the same discipline combatConstants.ts already imposed on itself for the offline solver's sake.
Why the guard is a subprocess. constantMirrors.test.mjs had been importing src/*.ts and passing for a long time, but only because Vitest resolves through Vite. The balancing harness runs under plain node. A test that imports the shared module the same way Vitest imports everything else cannot fail on the one thing that breaks the harness, so the guard spawns a real node --input-type=module -e ... instead. Confirmed it goes red by dropping the extension.
This does not weaken the "keep combatPolicy.mjs liftable into src/engine/" rule — it shortens the lift. The rule exists so the eventual copy is mechanical. Every predicate or constant taken by import is one less thing to copy and one less thing that can silently drift. The bot's own history is the argument: a mirrored ROCKET_TRAVEL_SPEED modelled the player's rocket at the enemy bolt's speed, wrong by 3.6x with nothing in the build able to notice, and the demo-campaign L6 wedge — eleven days, seven wrong root causes — was a mirrored hasLineOfSight built on the route planner's passability rule instead of the engine's.
What deliberately stays divergent. hasLineOfSight takes a sealedCorner flag rather than being unified: false for the engine, because that call decides enemy aggro and widening it is a gameplay change wanting its own A/B; true for the bot, because it is asking whether a bullet can reach, and a corner sealed by two solid tiles cannot be shot through even though a point-sampled ray slips between them. The route-planner tile set (isRouteBlocking, doors passable) is the other deliberate difference. Both are now one argument apart in one file instead of two implementations in two languages.
And two splits that are deliberate for speed, measured not assumed. SOLID_TILES (a Set, for callers holding a tile value) and isWall (a comparison chain, for grid lookups) are two representations of one set, because Set.has costs 3.0x the chain over 2.6M lookups and isWall runs per enemy per AI tick and per sight sample. A test pins them against every tile value the map layer can emit.
hasLineOfSight likewise keeps two loops behind one flag rather than one loop with the flag inside it. The corner rule needs the previous cell, so a shared loop pays two writes per sample to maintain px/py for a caller that never reads them — measured at 1.09x the original over 800k rays. Splitting gives the engine's path its original instruction stream back (1.026x, and the residual is a cross-module call frame, negligible against a few thousand aggro rays a second). The lesson generalises: sharing the predicate and the sampling density is what stops the two copies drifting; sharing four lines of loop body is not, and it costs. Do not "simplify" it back into one loop without re-running the benchmark.
The balancing bot's movement and combat logic is unusually easy to "obviously improve" and unusually hard to improve measurably — the numbers it reports are produced by its own behaviour, so a change that looks better can quietly make the measurement worse. Every item below was implemented, run through the matched-scale A/B (3 profile×difficulty combos, both sides against the same dev server), and reverted on the numbers. They are recorded here so the next reader does not re-attempt them; full detail and per-combo figures are in history.md.
-
Leading the aim ungated (
BOT_AIM_LEADwithBOT_AIM_LEAD_MIN_DIST = 0; the gated lead ships ON since 2026-08-14 — see the correction at the end of this entry). The lag is real and structural —simulate()runsupdateEnemyAibeforeupdateFiringwhile the bot samples positions only after a pump, so every shot resolves against positions one engine frame newer than the ones it was aimed at. Correcting it by differencing each enemy's position across consecutive decisions and extrapolating one frame forward cost Edge Cases 8.9pp of single-pellet hit rate at 0-2 tiles (z=-3.5), 4.4pp at 2-4 and 3.9pp at 4-6, replicated across both arm pairs that flip only this switch. It fails on precisely the archetype that motivated it, for precisely the reason that motivated it: a lead's angular size and a sprite's both carry1/dist, so their ratio isspeedMultiplier / spriteScale= 4.0 for an Edge Case, which means an Edge Case is punished four times as hard when the predicted direction is wrong. It is wrong often, becauseupdateEnemyAireturns early insideATTACK_RADIUS(enemyAi.ts) — an enemy closing at full speed one window stops dead the next to bite. A previous window's velocity is therefore not a predictor of the next one across exactly the state changes that matter. Do not re-attempt this by tuning the lead constant; the constant is right and the predictor is wrong. Two candidate explanations were written down the day this landed and both were withdrawn within hours, so they are recorded here rather than left to be re-derived: the enemy attack-stop cannot be it (ATTACK_RADIUSis 0.5 tiles, and only 7.1% of the affected bucket's shots were fired inside it), and reading the engine's own velocity viagetEnemies()would change nothing, becauserun-balancing-telemetry.mjsruns one 50ms engine frame per pump so the snapshot difference already is that frame's displacement. The remaining hypothesis is that a straight-line prediction of a curving path errs most close in — the enemy re-steers at the player every frame — which matches the observed decay with range. Score the predictor before writing a third fix — but note that doing so needs new instrumentation, not the archived captures: the event log records no enemy-position time series (onlyshotrows withdist/targetArch), so predicted-vs-actual positions have to be logged per decision first. An earlier version of this entry claimed the data was already on disk; it is not. The measured arm is kept switched off rather than deleted so that work has a control. Note also that pooled over all archetypes this reads as a clean null (78.7% vs 79.0%) — the regression is only visible once Edge Cases are separated out, which is the general shape of a real effect hiding in a minority subgroup.CORRECTION, 2026-08-14 — the predictor was never what failed, the third hypothesis above is refuted too, and the gated lead ships ON. The instrumentation the entry above asks for was built (
shotevents now carry the target's two preceding positions plus the shooter's position and facing), and scored against where enemies actually went, leading is more accurate at every range — it beats no-lead on 93.8% of shots and cuts positional error 38-60%. So "a straight-line prediction of a curving path errs most close in" joins the two hypotheses already withdrawn above; all three were plausible and none survived contact with the data. What the ungated lead actually loses to is the re-aim it demands: bearing rate goes asv_perp / dist, so inside 2 tiles the lead asks for 4.6x more turning than the accuracy it buys (against ~1.2x beyond), and the bot's measured facing error there is already 0.0859 rad — above even Casual'sfireAngleEps. It is firing mid-correction and the lead hands it more correction to do; that 4.6 / 1.3 / 1.2 / 1.2 ratio profile tracks the measured -8.9 / -4.4 / -3.9 / -1.5pp closely enough to be the same phenomenon, and the Edge-CasespeedMultiplier / spriteScale= 4.0 factor above still explains why the penalty lands on that archetype.BOT_AIM_LEAD_MIN_DIST = 2.0is read off that cliff, not tuned — setting it to 0 reproduces the rejected arm exactly, which is what keeps the retry single-variable. Result: +2.8pp of single-pellet hit rate at 6-8 tiles (z=3.69) against a +0.3pp measured null control, unchanged at 0-2 by construction, guard clean, and it replicates on a second substrate. Full numbers in history.md. What stays rejected is the ungated lead, and the transferable point is the one this entry got wrong twice: score a predictor before writing a fix for it. -
Making the bot's aim continuous — as a gamepad axis, and then as real mouse look (2026-08-24,
bot/analog-turn-axis, not merged). The bot's view visibly steps rather than sweeps, and the cause is structural: it dispatches only syntheticKeyboardEvents, so a turn is a key held forturnBurstMsand the view is frozen for the rest of the decision. Measured off the shipped board, the camera starts or stops turning 7.2 / 9.4 / 10.4 times a second on Casual / Gamer / Pro with 18% / 26% / 40% of turn bursts lasting a single 60Hz frame — a ladder that followsrotSpeedMultiplier2.0 / 3.5 / 5.0 exactly, becauseturnBurstMsdivides by it. The better the profile, the worse it looks. Three arms were built and measured on Pro/hard, two attempts each:- The gamepad axis alone was inert, and that is the useful half of the finding. The engine has always accepted
gpTurn, and it is reachable by shimmingnavigator.getGamepads(pollGamepadtakes the first non-null pad, wants no connection event, and reads every button through?.pressed). But the turn branches pass the burst as the intent's owndurationMs, sohold === duration, the axis came out at 1.0 on 98.6% of frames, and an axis at full deflection for a shortened decision is bit-for-bit what the key already did. Cadence was identical. The jitter is not the input modality — it is the decision being shortened to the length of the turn, leaving no window to spread rotation across. - Spreading the turn over the full decision helped, but less than it first appeared. Median turn run 1-2 -> 4 frames and one-frame bursts 47-50% -> 37-41%. It cost the thing it was always going to cost: one of two runs reached 6 levels where every other arm reached 15, which is the same shape of failure as the per-key-hold attempt that cost 25pp of Pro/hard qualify rate.
- Mouse look works and is not better.
profiles.mjsclaims mouse-look "isn't available to a Playwright-automated browser"; that is false. A trustedpage.mouse.clickon the canvas takes pointer lock, andonMouseMovegates ondocument.pointerLockElementand not onisTrusted, so the page can synthesisemousemovefrom inside the virtual-clock pump — 8px per frame rotates exactly 0.02 rad per frame, against 0.06-then-nothing for the same movement in one event. It also makesrotSpeedMultipliercancel out of the conversion entirely. Measured, it had the fewest start/stops but the shortest runs, no better per-frame step, and a 1.89 rad artifact from themouseDXthat accumulates while the lock is granted.
Two measurement errors mattered more than the feature, and both were mine.
report-turn-cadence.mjsunderstood turn keys andgpTurnonly, and scored the mouse arm at 0% turn duty across two complete runs — its own doc comment warned about exactly that. And cadence cannot see step size: one pixel of mouse is 0.0025 rad against 0.217 for a full 60Hz frame at Pro's multiplier, ranked identically untilstepP90/stepMaxwere added.The substrate was wrong too, which invalidated the first round outright.
run-balancing-telemetry.mjspasses norecordStepMs, soBotfalls back tostepMsand the harness simulates one 50ms frame per decision — nothing can be spread across frames there. It showed asstepP90 = 0.650(2.6 * 5 * 0.05) for every arm, including ones that were supposed to differ, while the shipped board records at 1/60 and never matched. Re-run withRECORD_STEP_MS=16.6667, the spread arm's headline 73%->43% shrank to 47-50%->37-41%. Any per-frame claim measured through the telemetry harness is measuring the decision rate, not the frame rate.Rejected by the user, 2026-08-24, on the ground that no arm bought enough to be worth a risk to skill:
stepP90never fell in any arm (the peak jump is unchanged), and the only arm that improved smoothness is the only one that dropped a run.report-turn-cadence.mjssurvives as the instrument — the three flags did not. - The gamepad axis alone was inert, and that is the useful half of the finding. The engine has always accepted
-
Biasing key placement toward the gate each key opens (
placeKeys,src/map/generation/doorsKeys.ts). The obvious reading of "level 3's key/door chain is too long" is that the keys are scattered too far from their doors, so this made each key land in the nearest quarter of the region it is hidden in. Measured over the campaign it moved level 3 by only -2.6% and made level 7 +24% worse: distance-to-gate ignores which side the player approaches from, so pulling a key toward its door can push it away from the route. Reverted. The real cost was never placement but the order the bot picked keys up in — see the nearest-key ordering inroutePlanner.mjs. Note this one is not a bot-behaviour change at all; it is recorded here because it was the first attempt at the same symptom, and because a generation change would have invalidated every baked replay while the planner fix invalidates none. -
Moving the door staging point inside the engine's door-open reach.
openDoorAheadprobesradius + 0.15 = 0.35tiles ahead of the player's exact position, and the bot stages at the centre of the tile in front of the door — 0.5 tiles out, so 0.15 beyond the reach for every approach direction. Placing the staging point inside that band instead looks like the obvious correction and is arithmetically sound, but it cost qualifying runs 3/3 -> 0/3 on all three profiles. The band is 0.15 tiles wide while one walk decision covers 0.16, and wall collision rejects a whole step rather than clamping it, so coarse driving cannot land inside the band — the bot ends up stranded approaching the staging point rather than at it. The coarse-approach-then-fine-push split is therefore load-bearing and correct. Note the follow-up this originally recommended — verify the push opened the door and retry — was then implemented and measured to be a no-op: over a full scan the retry fired zero times, because the door does open on the first push every time. The stalls that prompted all of this turned out to be loot detours wedging at doorways, not door opening at all. The staging point being out of reach remains true and unfixed, but it is latent, not the cause of anything observed. -
Pre-emptive engagement (rework plan Stage 6) — letting the bot start fights with enemies that have not aggroed yet. Built, measured at full protocol scale (4 combos x 40 attempts x 2 sides, ~3h), and reverted as an unjustifiable no-op. The gate mirrored the engine's own aggro rule —
d < AGGRO_RADIUS 7.5 && hasLineOfSight(enemyAi.ts) — because that was the justification: inside it the enemy notices the player unprompted, so engaging first only decides who shoots first rather than adding a fight. That justification is also precisely what makes it useless. The engine's aggro condition is the same condition, so any eligible enemy is one the engine will aggro on its very next AI tick: the eligibility window is one tick wide by construction. Instrumented over 24 levels / 35,770 decisions, the gate opened on 0.99% of decisions and the bot acted on 0.66% — about 10 decisions per level, mostly the same enemy across consecutive ticks, so roughly 1-3 enemies per level engaged fractionally early. Against ~14 kills a level that cannot move an aggregate, and the A/B duly measured nothing: every win-metric movement on the three feature-ON combos sat inside the noise floor established by the null control (Casual, gated off on both sides, identical code — which still movedenemyAccuracy+5.2%,minHealthReached-29.5%,trapSpike+57.6%), andkillsObservedmoved -1.0%/+0.3%/-0.6% against the control's +0.2%. Guards passed and all three stage-specific rollbacks were clear; it was neutral, not harmful. The trap for anyone re-attempting it: the only way to make it fire often enough to measure is to engage beyond the engine's aggro radius, i.e. to start fights that would not otherwise have happened — which contradicts both the stated justification and the standing user constraint that the bot must never run a long corridor for an enemy that had no chance of reaching it. Justification and usefulness are mutually exclusive here; pick a different lever. Worth keeping from the exercise: the walking-distance gate was genuinely load-bearing, rejecting 27% of candidates overall and 43 of 45 onstage07_service.rb— enemies plainly visible across a gap the bot would have had to walk around. A straight-line gate would have engaged every one of them, which is exactly the failure the constraint predicted. Two artefacts survive the revert because they are correct independently:pickThreatreceivingthis.tuningat every call site (an omitted argument silently fell back to unmodified defaults, which would have made the baseline side of this very A/B run candidate behaviour), and Pro/normal's permanent place in the A/B combo set. -
Per-profile evasion (profile retune Stage 3) — making dodging and combat-strafing a skill tier. The plan's centrepiece: promote
DODGE_REACTION_FLOOR_SEC(a new deterministic reaction-latency threshold, 0.30/0.18/0.10),DODGE_MIN_LOOKAHEAD_SEC(0.35/0.6/0.9),DODGE_LOOKAHEAD_DECISIONS(2/3/4) andCOMBAT_STRAFE_FLIP_TICKS(4/8/12) from global constants to per-profile tuning, so a higher tier would take less damage rather than merely playing faster. Measured at n=20/profile over 8 levels and reverted. The reaction floor works exactly as designed on the bottom tier — Casual'senemyAccuracyrose 11% (0.499 -> 0.556), i.e. it gets hit more, which is what a worse player should do. But the ladder stayed inverted and Pro got worse on both damage axes despite strictly better dodging (enemyAccuracy0.531 -> 0.542,dmg.enemyRangedPerSec0.439 -> 0.465), every tier took more ranged damage per second of exposure than before, andttkNormal's smallest adjacent step fell 1.161x -> 1.118x, breaching its pre-registered 1.15x floor. The reason, found afterwards by reading rather than measuring — and it is not the one first recorded here. The initial conclusion was "the dodge machinery may not be net-protective". That was wrong, and worth stating plainly so nobody acts on it. The bot does shoot while dodging: the dodge block sits inside theelsearm of the aim gate, i.e. the branch that also setsfire. What it cannot do is dodge while re-aiming. There is exactly one dodge call site and it is in that fire branch, so a bolt is only ever dodged when the bot already has a firing solution —|delta| <= profile.fireAngleEpsand line of sight. The module's own note records the fire branch as "roughly 43% of combat decisions", so for the other ~57% the bot is a stationary target with no evasion available at all. That is why tuning dodge quality per tier moved so little: quality cannot matter in windows the bot never gets, and Pro's longer 0.9s lookahead mostly qualifies bolts that have already landed by the next fire-branch decision. The experiment that follows is a directed dodge in the re-aim branch, and note it is not the one already reverted under "strafe only once aimed": that added blindcombatStrafeKeyoscillation to re-aiming and fought aim convergence on every decision, costing 30pp of qualify rate. A directed dodge fires only when a bolt is genuinely inbound — rare, and exactly when being hit is imminent — so it perturbs aim only in those moments. It must also dropKeyWwhile dodging, or holding forward plus lateral engagesdiagonalScaleand cuts the forward axis 29%, which is the mechanism behind the 0%->72% regression. -
Diagonal strafing applied to every turn-and-move branch. Casual/normal's level-2 death rate went 0% → 72%. Mechanism:
diagonalScale(1/√2) keeps total speed constant but cuts the forward component by 29%, and in the hazard-crossing and critical-health branches the forward axis is the survival axis. Kept in plain navigation only. This is the single largest measured regression in the bot's history and the reason any lateral key added to a survival branch is treated as guilty until proven otherwise. -
Per-key hold durations (letting a movement key outlast the turn beside it). Looks like the obvious fix for a grid-locked gait; measured slower and less accurate on all three combos and cost Pro/hard 25pp of qualify rate. A short decision during heading correction is closed-loop control, not a bug.
-
Extending the combat strafe to the re-aim branch. Cost Gamer 30pp of qualify rate and doubled its stuck count for no accuracy gain — that branch exists to converge on a firing angle, and lateral movement perturbs the very quantity it is nulling out.
-
Removing the stall-strafe turn-burst widening. The widening looks like a bug (it holds a turn key for a full step against a burst computed in microseconds). Removing it raised
enemyAccuracyon all three combos and breached Gamer's qualify-rate guard: it is load-bearing for dodging, not aiming. -
Two attempts at the "oscillation" wobble — first gating the diagonal strafe on a safety check, then pinning the turn direction near
atan2's ±π branch cut. Both made the detector fire more. The diagnosis itself was wrong:detectOscillationwas largely flagging the bot's ordinary stop-turn-sprint gait, and both fixes slightly lengthened convergence, so more ticks accrued per turn. Fixed by correcting the detector instead (a route-progress exemption, findings 38 → 4). -
A threat-independent strafe to unwedge a pinned bot. Both existing anti-stall wiggles (
combatPolicy.mjs's stall-strafe and the cornered-retreat strafe) are gated onthreat != null, so a bot driving withignoreThreatshas no recovery mechanism at all — which reads as an obvious hole, and was the first thing planned whenverify (multiplayer-transition)kept wedging. It was not built, because fixing the budget removed the symptom entirely: withBOT_NAV_STALL_BAIL_TICKSgiving up early enough fordriveTowardWithReplanto still have re-plans left, the fixed configuration produces zero stall/oscillation/heldKeyNoMovement findings over six runs. Re-attempt only with a case the bail cannot clear, and read the 0%→72% diagonal-strafe entry below first — a lateral key is the single most regression-prone thing to add to this bot.
The recurring lesson across all of these: a bot-behaviour change is not validated by the metric it was designed to move, because that metric is itself measured through the bot. Guard metrics (qualifyRate, per-level conditional death rate) are pre-registered in abReport.mjs for exactly this reason — see Balancing Telemetry.
BOT_WALKING_DISTANCE_THREATS and BOT_ANGULAR_FIRE_GATE were both measured against a staged curl campaign and both moved nothing — full numbers in history.md, including the powered three-arm re-run that retracted the one cell which had looked significant. Both stay ON, and that is a decision rather than an oversight, so do not re-open it as "two dead flags nobody removed".
The reasoning, which the measurements do not touch:
- Each replaces a dimensionally wrong quantity with a right one. Straight-line distance is the wrong measure for a decision about travel; a fixed angular tolerance is the wrong measure against a target whose angular width falls off as
1/dist. A null says the wrong quantity happened not to cost anything on this campaign, not that it was the right quantity. - Neither costs anything measurable — no perf signal, and each is a small expression inside a decision that already exists rather than a new mechanism. This is the substantive difference from Stage 6, which was reverted: that added a whole engagement path for a gate that could open on 0.99% of decisions. These two fire constantly (the walking-distance rule changes the answer at 16.1% of candidate-bearing positions, 33.2% on the wall level; the angular gate binds on 45.4% of Casual's shots) and are nearly free.
- The instrument cannot see the reason they exist. The bot is destined to become a deathmatch opponent, where "tracks you through a wall" and "takes a long shot it cannot land" are things a human notices. No bot harness measures how a fight feels, so a null here is silent on that question rather than evidence against it.
What the nulls did retire is the "unproven" label. Neither may be described as awaiting evidence; both are measured neutral on curl/hard, which is weaker than "they work" and stronger than "untested".
A CSS class that sets display silently defeats the hidden attribute. Show/hide in this app is done by toggling the hidden attribute, which works because the UA stylesheet carries [hidden] { display: none }. But that is a user-agent origin rule, and any author-origin declaration beats it — so the moment a class on the same element sets display (flex, grid, block, anything), the element stays visible with hidden set, with no error and nothing in the console to hint at it. The fix is an explicit override per class:
.loading-screen[hidden] { display: none; }
.canvas-area[hidden] { display: none; }Both of those exist in src/style.css because this shipped twice — first .loading-screen (a loading overlay drawn on top of the intro screen while carrying hidden), then .canvas-area later, in exactly the same way. Any new .hidden-toggled element whose class sets display needs the same one-line override, and it will not be caught by the test suite: jsdom applies no UA stylesheet, so a unit test asserting el.hidden === true passes while the real browser still paints the element. .tab-panel[hidden] is the third instance of the same pattern. (multiplayer-game-state-spec.md refers to this as "a documented past bug class in this project" — this is that documentation.)
A flex item wraps instead of growing, because its min-content width is a word, not the string. .tab-btn is flex: 1 (basis 0) with the default min-width: auto, so a tab should never be squeezed below its own content — but "▶ GitHub" has a break opportunity at the space, making its min-content width just GitHub. Flexbox therefore left the tab at its equal share and the label wrapped under the glyph, which is what shipped and had to be fixed. white-space: nowrap on .tab-btn is what makes min-width: auto floor the item at the whole string, so the marked tab takes the width it needs from its neighbours instead. This is invisible to the test suite — jsdom has no layout engine, so wrapping and widths can only be caught in a real browser; the fix was measured across every tab-count × marker-state combination rather than eyeballed on the one that happened to be loaded (checking only the shorter "Demos" label is precisely how it shipped wrapped).
Separately, six tabs (Continue and Multiplayer both showing) never fit the 287px strip even with no marker — the Settings gear sat past the sidebar's right edge, cut off. .launch-tabs is flex-wrap: wrap so that degrades to a second row instead of a tab silently leaving the sidebar; four and five tabs still fit one row.
No canvas overlay hosts an interactive control. Every settings-shaped control in this codebase is DOM in the sidebar (#tab-panel-settings since the Settings tab landed) backed by localStorage; the canvas overlays are strictly transient, non-interactive feedback (toasts, briefings, the pause and end screens). That is an architectural convention rather than a style preference — a control painted into the canvas would need its own hit-testing, focus handling and accessibility story, none of which exist here. Put new settings in the sidebar.
The file tree belongs to the tab the workspace was loaded from — and exactly one workspace is loaded at a time. One #file-tree sitting outside every panel was a shipped bug: the demo campaign's files were listed under the GitHub tab too, and a failed GitHub load relabelled the campaign you were mid-run in. Each source now has its own pane and its own status line (local, github, demo — Local and Continue share one, since both go through pickWorkspace()), and commitWorkspaceSlot is the single writer of the workspace* globals, so every downstream reader (campaign progression, autosave eligibility, highscore source, multiplayer host gating) was untouched by the change.
Keeping several workspaces live at once was built first, and rejected on sight (user call): three tabs each presenting a loaded workspace, when only one of them is what the engine plays, reads as a lie no marker can undo. A load therefore clears every other slot (clearWorkspaceSlot), and a tree on screen always means "this is the loaded workspace", never "one of several you could pick from". That also keeps the model the code already had — one workspace, one campaign — rather than inventing a second kind of state that only the sidebar knew about. The cost is that going back to a workspace means re-loading it (instant for Demos, a re-fetch for GitHub, a re-pick for Local); this was accepted deliberately.
Consequences worth not "simplifying" away:
-
The tab marker exists only while a level does. The rule is marked ⟺ a level is loaded; only then does the glyph distinguish ▶ (advancing) from ❚❚ (not). It used to be marked ⟺ a workspace is loaded, which is a different and much longer-lived condition:
loadedWorkspaceSlotis only ever assigned, never cleared, so once anything had loaded the tab carried a transport glyph for the rest of the page's life — reading ❚❚ after a run ended, offering to resume something that no longer existed.Three things this settled, all of them found by tracing rather than by the original report:
- The flag hangs off
showFileTreePlaceholder, notresetToFileTree. That function is the one place that actually means "no level is on screen", and fourcatchpaths reach it without touchingresetToFileTreeorcommitWorkspaceSlot(which sits inside the sametryblocks): a failed local pick, a repo that 404s, a Continue Run whose workspace is gone, and a replay that cannot be rebuilt. Hooking the other two left the marker stuck over an empty viewport after a failed load — the reported bug, reintroduced by its own fix. - It is an explicit flag, not
currentLevelPath !== null.launchLevelrepaints the marker ~54 lines before it assignscurrentLevelPath, so deriving it there reads the previous level's value, andnullon a session's first launch — the tab would sit unmarked through the whole briefing and pop the glyph in on Start. - The end-of-run overlays used to claim ▶. The engine
stop()s itself before firingonGameOver/onWinandstop()never notifies the freeze handler, so nothing clearedgameIsRunning; the tab said "running" for the entire Kernel Panic or Build Successful screen.onGameOver,onWinandonMultiplayerSessionEndednow clear it. The multiplayer case had been written off in a comment claiming "a session is driven from the Multiplayer tab, which never carries the marker" — that was wrong: all six buttons always render andSLOT_FOR_TABmaps onlymultiplayer/settingstonull, so the source tab is on screen and stale throughout. What a running session's marker should say is a separate question, left open.
- The flag hangs off
-
The intro tour cannot enter the Highscores dialog, and describes it instead. A dialog opened with
showModal()lives in the browser's top layer, which paints above every z-index;.intro-touris a fixedz-index: 50layer. Its popout, highlight and buttons would all render behind the dialog and be unreachable — the same constraint that already made the tour a fixed layer rather than a<dialog>of its own.TourStepalso has noonEnter/onLeavehook,activateTabonly.click()s an id, andstop()restores the tab and focus but could not close a dialog it opened, so an Escape mid-tour would strand it. The highscore step therefore stays anchored on the button and carries the board's contents in prose. Steps that do reach into a tab panel declareactivateTabId; a step anchored on a tab button (Demos, Repo, Multiplayer) needs nothing, which is why the Repo step targetstab-reporather thanrepo-input. -
Switching tabs is display-only. It swaps which pane and name line are visible and never touches the running game — only the tab holding the loaded workspace shows a tree at all, and a marker on that tab says which one it is from anywhere, including the two tabs that never show a tree. That marker doubles as a transport readout (▶ running / ❚❚ not), driven off the engine's existing
onFreezeChangeedge plus the briefing andresetToFileTree. It is also now the only DOM-observable end of that wiring —consoleSidebar.setPausedjust sets an internal flag — so the pause test asserts through it rather than through "did not throw". Multiplayer deliberately does not drive it:sessionEngine.tswiresonFreezeChangeto a no-op on purpose, and a session is driven from the Multiplayer tab, which never carries the marker. -
Every loader claims its own tab before its first
await. In the app that is already true (each button lives inside its own panel), but as an invariant it guarantees the status line and tree a load writes are the ones on screen, and it stopsclearCampaignSave()'s bounce off the Continue tab from yanking the view mid-load. -
A slot's status has a default as well as a current value. Clearing a slot restores the default, which is how the "browser does not support the File System Access API" message survives a GitHub load clearing the local slot — that message outlives any workspace.
-
Panes are toggled, never re-rendered.
renderFileTreewipes its container and rebuilds every row collapsed, so re-rendering on a tab switch would silently discard the tree's expanded folders. And thedisplay/hiddentrap above applies here:.file-tree-panedeliberately declares nodisplay.
Player feedback is presentation, and presentation stays out of the simulation. The locked-door hint (cueLockedDoorHint) is the worked example: a toast counter, a ping window and a target key, all on PlayerState, all ticked in tickEffects(), none of them in SIMULATION_BALANCE, reconciliationSnapshot.ts or rosterSnapshot(). balanceHash.ts explicitly exempts anything that changes only rendering, so no replay is invalidated and defaultHighscore.ts needs no regeneration. Local-player-only, on the same reasoning as cueLocalAcidOverflow: openDoorAhead loops every player, and a teammate shoving a door across a coop level must not toast, log or ping your screen — peers that cue differently stay in lockstep because nothing here reaches simulate().
Three details of that hint are load-bearing and would be easy to "simplify" back into bugs. The ping window doubles as the rate limit — the trigger fires only while keyPingFrames is 0, so a player leaning on W gets one hint per window rather than one per tick; that matters off-screen too, since the playtest bot pushes into doors and run-balancing-telemetry.mjs forwards page console. And the key is chosen by reachability, not Math.hypot — nextKeyStep floods with keyRoute's own passable, which counts a door you already hold the key for as open. PathField cannot be substituted here however convenient it looks: its isWall calls any still-closed DOOR_TILE solid, so it under-reports a key one openable door away. (An earlier version of this entry described nearestReachableKey and PathField.distAt; both were replaced by nextKeyStep on 2026-08-21, which also stopped the hint going silent when the asked-for key was itself locked away — it now deliberately points at the key that unblocks it.) And the two triggers are not peers. Walking past a key arms the same ping (cueNearbyKeyHint), and it yields to a live door hint — but a door bump overrides a live proximity ping rather than deferring to it, tracked by keyPingIsDoorHint. Making that suppression symmetric is the tempting simplification and it is wrong: bumping a door is something the player did, and the early return skips the denial thunk and the toast along with the ping, so a passive hint would swallow the only feedback that the door refused them and the door would read as simply broken.
The proximity hint is ordered cheap-test-first, and that ordering is the design. reachable floods the whole grid, and unlike the door trigger — which runs once per bump — this one lives on a per-frame path. So it filters on geometry first (at most four keys, a rect or radius test each) and only floods once something is genuinely in range and not already latched, with the whole scan throttled to four times a second. That throttle is not about the common case but the worst one: a key that is near but not reachable, sitting in a locked room you are walking past, fails the flood every time it is checked. The trigger shapes are two because the geometry is: measured across the demo campaign, 20 of 25 keys sit inside a map.rooms rect, 4 in corridors and 1 in another rect type, so a room-only rule would never fire for 16% of keys — including a level whose only key lies in a corridor. Hence the room rect where there is one and a radius around the key where there is not. It fires once per key per level rather than on a cooldown: the job is "make sure you noticed", which is done the first time, and a player who walked past and kept going has already decided.
fillText's fourth argument is a clamp, not a layout. It does not wrap and does not overflow — it compresses the glyphs horizontally to fit, so a line that overruns its box still renders, just visibly squashed at a different pitch from every line beside it. That failure mode is invisible to a test asserting what text was drawn, and it shipped: the death screen picked between two fixed widths (420, or 620 once it carried stats rows), and the death screen picked its width before the stats rows existed, so the rollback wording — up to 71 characters, needing 556px — was squeezed into 388. The grouped stat values (Path … · Map … · Lore … · Secrets … · Streaks …, 446px) were squeezed even at 620, because each column only gets boxW / 2 - 24; no choice between two fixed widths could have fixed that one. The worst case is the replay-ended screen, whose reasons interpolate a file path — main.ts's balance-mismatch string reaches 986px, and is the one case long enough to stay squeezed after the fix, since a box that wide runs into the canvas edge first. If that screen ever needs to be genuinely correct rather than merely better, it needs wrapping — which a later pass did add (2026-08-23): body lines wrap instead of being squashed, stat rows split at a content-driven column rather than the box centre, and a value that still does not fit steps 13px to 11px. The balance-mismatch string is the one case long enough to stay imperfect, since a box wide enough for 986px runs into the canvas edge first.
overlayLayout therefore derives a minimum width from the content and keeps 420/620 only as floors, so every screen that already fitted is unmoved. It can do this while staying pure — no measureText, which is what lets show() wire up input on a canvas with no 2D context — because the overlay is monospace end to end: a character count is a width. CHAR_EM is that per-character advance, 0.602em measured in Chromium at every size and weight the overlay uses, rounded up to 0.62 because the ui-monospace fallback differs by platform. The rounding direction is the load-bearing part: over-estimating pads the box by a few pixels, under-estimating puts the squeeze back. maxWidth stays applied underneath as a backstop, so a wrong estimate — or a box already pinned against the w - 48 canvas edge, which is the one place text can still be squeezed — degrades to the old rendering rather than letting text escape its box.
The button label is the piece still drawn with no clamp at all. It fits today (the widest, "Return to file tree", is 149px in a 170px button) and the suite would not report it if that changed. [notes: gameover pun for cheated runs]
Every level used to look the same: one global TextureSet, with a single boolean (GameMap.bonusLevel) flipping walls/floor/ceiling to a cool variant for header-file "restock arena" levels. Stylesets replace that with five named looks (stone, rust, tech, marble, techCool), one per level, fixed for that level's duration. stone and techCool reproduce the two pre-existing palettes byte-for-byte, so the change reads as "three looks were added", not as a re-theme.
Only structural surfaces vary. Wall, floor, door, ceiling tone and minimap wall color are per-styleset; the lore terminal, hazard/acid, teleporter pad and spike-trap flats are built once and shared by reference across all five. Those five carry meaning — green is acid, red is a live spike trap — and a player must never have to re-learn which color hurts them because the level changed palette. The shared-by-reference construction is what enforces it: there is no per-styleset copy to accidentally diverge.
The styleset draws zero values from the map generator's rng. mapGenerator.ts's header records that the order of its generation/* calls is the rng draw sequence, so consuming even one value shifts every downstream placement — which would move every existing map layout, invalidate every recorded replay, and invalidate the shipped defaultHighscore.ts board, all to pick a wall color. generation/styleSet.ts therefore runs its own single-draw mulberry32 seeded off the same seedFrom(parsed) content hash. A dedicated draw rather than seedFrom(parsed) % n: the modulo keys off FNV-1a's least-scrambled low bits, and a real campaign's files (one language, similar entity names) can cluster into one bucket far more often than chance — the demo campaign's spread across ≥3 stylesets is asserted in styleSet.test.ts, and the zero-draw property has its own exactness gate in mapGenerator.test.ts.
Two consequences fall out of deriving it from the parsed source alone, and both are the reason it's done that way rather than by campaign position or per-session roll: multiplayer guests regenerate the identical styleset from the identical file with no new sync, and replay needs no new recorded field — the styleset is a pure function of (seedFrom(parsed), bonusLevel), and playback already re-parses the file under an astHash guard and already records bonusLevel.
Bonus levels keep a dedicated styleset. techCool is reserved for bonusLevel levels and is never in the normal pool, so seeing it still means "restock arena" — the signal survives the added variety instead of being diluted by it.
On the WAD side, a missing slot borrows from a sibling styleset before falling back to programmer art. Resolving each styleset strictly in isolation would let a sparse WAD (the DOOM shareware IWAD, a small texture pack) leave one styleset fully WAD-textured and the next fully procedural — the campaign would visibly flip between "real textures" and "defaults" from level to level, which reads as a bug. Borrowing degrades instead to "two stylesets look alike", which is merely less variety. Measured rather than assumed: npm run report:wad-stylesets shows all five stylesets resolving a distinct wall and floor of their own across all five catalog WADs plus a real Doom IWAD, with marble's GRAY4/GRAYTALL fallbacks added specifically because the shareware IWAD and HACX ship no MARBLE* and marble would otherwise have borrowed stone's.
A 2026-07 perf audit built a repeatable measurement harness (npm run perf:bench/npm run perf:report, real-clock Playwright runs with per-frame phase timing via ?perfDebug=1) rather than tuning off intuition, specifically to answer whether WALL_EDGE_ANTIALIASING_ENABLED and RESPONSIVE_CANVAS_SCALING_ENABLED were worth shipping on by default — both had shipped disabled ("perf impact under measurement") until this audit ran. Interleaved A/B across a benchmark scenario matrix (idle, replay-combat, particle stress, a huge GitHub-repo map, bot play) found antialiasing under the harness's 0.2ms/frame detection floor on demo-scale maps and only ~+0.4ms/frame on a 160×160 monster map, and canvas scaling completely free both headless and in an attended on-screen A/B (the compositor upscale costs nothing measurable) — both flags were flipped to true. See Feature Flags for the current defaults and perf-findings.json (repo root) for the full per-finding writeup.
The same audit investigated a real user-reported bug (drastic framedrops on a large real-world repo, even with no ammo/no shots fired) and found no engine-side hotspot reproduces it — the projection-cost path the report initially seemed to implicate doesn't even run on an out-of-ammo fire() call. Rather than guess further, ?perfDebug=1 shipped as a permanent opt-in diagnostic (per-frame phase breakdown, entity counts, mousemove rate, heap usage) so the next occurrence can be captured directly from the affected machine — still open, tracked in notes.
One finding was closed as a deliberate non-fix rather than left open indefinitely: a WebKit-specific rendering discrepancy found during the audit was scoped out as wontfix — Safari/WebKit is exercised in CI for correctness (verify.yml's browser matrix) but isn't a target for perf-tuning effort, a user decision rather than a technical dead end.
Public highscore boards for public GitHub repos were scoped and scrapped (user, 2026-08-19). highscores.ts writes to localStorage and makes no network call, and that is the intended end state rather than a stage on the way to something.
The consequence is worth stating positively: the app has no outbound user data at all. The only network traffic it originates is the signaling server during multiplayer, which a player opts into by joining a session, and the online WAD catalog's same-origin fetch of files the build already placed in public/wads/. That makes the standing audit item's assertion simple enough to be worth enforcing as an invariant: a fetch to anywhere else is a regression.
Why it was declined rather than deferred, since the pieces to build it genuinely exist and the idea keeps looking cheap:
- Consent could not carry the privacy question. A
HighscoreEntrycarriescampaignName(anowner/repo),codebaseLinesOfCodeandcodebaseComplexity. For a local or private workspace, publishing that discloses the name and the size and complexity of code that is not public — so the feature needs a "is this repo actually public?" gate on top of opt-in, and the failure mode of getting that wrong is disclosing someone's private codebase metrics. - An unverified board is not worth having, and a verified one is a service. Client-submitted scores are trivially forged. This project can do better than most — a run carries a deterministic
replay, anastHashand abalanceHash, andverify-replay.mjsalready re-simulates a run end to end — but recomputing a submitted score means running the engine headless per submission, which is CPU per request and an abuse vector in its own right. - The hosting commitment is the real cost. The signaling server is deliberately ephemeral: in-memory
Maps on a 5-minute TTL, nothing to back up, nothing to moderate. A scores board needs durable storage, backups, and a moderation and takedown path for free-text player names on a public surface. That is a different kind of thing to operate than a stateless relay.
If it is ever revisited, revisit the hosting commitment first — the client-side work is small and the server-side work is permanent.
The multiplayer backend is updated by a systemd timer on its own host running docker/update.sh, installed by update.sh --install. CI is not involved at all (user, 2026-08-19).
Why CI is the wrong side to drive it. The obvious shape — the deploy workflow reaches into the backend host and updates it — means a deploy key held as a GitHub secret that can run arbitrary commands on that machine, plus inbound SSH. That is a permanent remote-code-execution path onto the box, added so a rarely-changing signaling server updates a few hours sooner. Pulling extends trust only to the git remote, which the host already trusts by virtue of running code from it.
Why not the file-drop the backlog originally described (deploy writes .update-backend, backend watches it with inotify): the frontend deploys by FTPS to a webhost, and the backend is a different machine running docker/docker-compose.yml. There is no shared filesystem for inotify to watch. The cross-machine version would be polling a published URL, which needs a version artifact that does not exist — and buys nothing over just pulling, since git pull already answers "is there anything new".
And the container cannot restart itself anyway. docker compose up -d --build runs on the host, from outside the container, so the actor was always going to be a host-level agent rather than the signaling process.
Two constraints any implementation inherits. update.sh runs as the deploy user and escalates only for the docker calls, because a root git pull leaves the repo root-owned — so running it from a timer needs a NOPASSWD sudoers entry scoped to those commands, there being no TTY to prompt on. And restarting signaling drops open lobbies and pending joins (sessions is an in-memory Map on a 5-minute TTL) while leaving established peer connections alone, since peers talk directly after the handshake — which is why the timer belongs at a quiet hour rather than running frequently.
The workspace can load a repository from GitHub, and the backlog long carried "should support any http(s) available git repo". That second half is out of scope, by decision (user, 2026-08-19) — not deferred, not blocked on effort. Adding more hosts (GitLab, Gitea/Forgejo, Bitbucket) stays open and is unaffected.
Why the general form is not a bigger version of the specific one. Loading GitHub is two HTTP calls against a JSON API. "Any git repo over http(s)" is git's smart-HTTP protocol — GET /info/refs?service=git-upload-pack, then POST /git-upload-pack returning packfiles — and it fails in a browser for two independent reasons, either of which is fatal on its own:
- CORS. Git hosts do not send
Access-Control-Allow-Originon those endpoints, sofetchcannot read the response at all. isomorphic-git, the only serious browser git client, ships its own CORS proxy — that is the tell that this is structural rather than incidental. - It needs a real git client: packfile parsing and delta resolution, against a project that hand-rolls a ZIP reader and a WAD decoder rather than take a dependency (see Dependency Minimalism).
The workaround this project already uses does not transfer. The online WAD catalog hit exactly this wall — none of its upstream hosts send usable CORS headers — and solved it by moving acquisition out of the runtime into scripts/fetch-online-wads.mjs at build time. That works because the catalog is fixed and known in advance. A player picks their repo at runtime, so there is nothing to pre-fetch.
Which leaves a proxy the app controls, and that is the part being declined. There is a server already (scripts/multiplayer-server.mjs), but turning it into an open fetch proxy means an SSRF surface, unbounded third-party bandwidth billed to whoever hosts it, and an availability commitment for a feature that is otherwise entirely client-side. The cost is operational and permanent; the benefit over supporting the three or four hosts people actually use is small. Add hosts, do not add a proxy.
GitHub shipped first; GitLab and Codeberg followed on 2026-08-20 behind one auto-detecting "Repo" tab that reads the host off the pasted domain (see history.md for the three traps that pass found). Bitbucket was the last candidate and is declined (user, 2026-08-23), closing the backlog item rather than leaving it open indefinitely.
The reasoning is about this feature rather than about Bitbucket generally.
The adapter reads public repositories over an unauthenticated API, so the
only usage it can serve is someone pasting a public repo to turn it into a
level. Bitbucket's presence in that population is small — it is predominantly
private and team-hosted Atlassian work, and dropping Mercurial hosting in 2020
moved most of the open-source projects that were distinctively there elsewhere.
Against that, the per-host cost is not nominal: a sibling adapter, a URL and
shorthand parser, a tab entry, its own rate-limit inference path, and tests.
The backlog item also recorded that Bitbucket is unproven — its API answered
ACAO: * on the one probe, but the probe itself 404'd, so the tree shape and
pagination were never actually checked and scoping would have started by
redoing that work.
Declined, not walled off: the "Repo" tab dispatches on domain, so if a real request for Bitbucket (or Gitea, or Forgejo) ever arrives, the adapter drops into a slot that already exists.
The rule that outlives the decision — check CORS against the live host before
writing any adapter, not after. It is the only property that decides whether a
host is possible at all from a page, and it cannot be reasoned about from
documentation: GitLab sends ratelimit-remaining but omits it from
Access-Control-Expose-Headers, so a header-gated check is dead code; its
obvious /-/raw/ URL sends no CORS header at all. Self-hosted Gitea and Forgejo
configure CORS per deployment, so "Gitea works" is a claim about one instance,
never about the software.
The project deliberately avoids adding a new npm dependency where a built-in browser API already covers the need: crypto.subtle.digest('SHA-256', ...) is used for the AST/campaign hash shown in highscores instead of a bundled hash library, and highscore-board compression uses CompressionStream/DecompressionStream instead of a bundled compression library. The same principle drove the parser architecture: rather than pull in a second parsing library to cover additional languages, every non-bespoke language (12 of 14) is handled by one data-driven adapter built on the same Tree-sitter grammars already in use — see Architecture. [notes: Task 23, 35, 44]
The DOOM WAD texture importer (src/wad/) followed the same reasoning after an explicit check: every npm package found for WAD parsing was either Node-only (native canvas bindings, unusable in a browser bundle), unmaintained 0.0.x experiments scoped as full map/BSP editors, or pulled in unrelated computational-geometry libraries for sector triangulation — nothing cleanly covered "decode PLAYPAL + PNAMES + TEXTUREx + patches + flats" without dragging in a pile of unrelated transitive risk for a format that's small and fully specified. Hand-rolled with DataView, the same style as storageCompression.ts's byte handling.
scripts/lib/zipReader.mjs — the ZIP reader backing the online WAD/texture-pack catalog's build-time fetch step (scripts/fetch-online-wads.mjs, see Architecture) — got the same treatment: it's a one-shot dev tool that only ever needs to pull one named file out of a handful of known-good, non-encrypted, non-multi-disk ZIPs, so a full-featured unzipper/adm-zip-style dependency would be pulling in far more format surface (zip64, encryption, streaming) than this actually exercises. Hand-rolled the same way: raw buffer reads at explicit offsets, node:zlib.inflateRawSync for the one compression method (DEFLATE) actually used by these sources.
The classic Doom cheat codes (IDDQD/IDKFA/IDCLIP) are implemented but deliberately kept out of the README and out of the game's own UI — no menu entry, nothing in the on-screen controls legend, nothing that surfaces them in play. They are written up in doc/user/controls.md, for the same reason Doom's own were in the manual: an Easter egg you have to go looking for, not one you trip over. Highscore recording is disabled for the rest of a campaign once any cheat has been used, so an Easter egg can't be used to inflate a leaderboard entry. [notes: Task 39]
The rule that keeps cheats out of the UI is about disclosure, not about staying silent once one has fired — the in-play ⚠ CHEATS USED badge already reacts, and the Kernel Panic screen gained a one-line jab on the same grounds. Nothing on either surface is reachable, or even visible, until the player has typed a code they had to go looking for, so neither one is a route to discovering the codes. Two constraints on that line worth keeping: it is read from EngineStats.cheatsUsed rather than main.ts's campaign-scope cheatsUsed, because that is the latch the badge reads and the two surfaces must not contradict each other; and it is deliberately absent from the rollback-offer variant of the screen, which is a decision the player has to read and where every line is required to describe that choice (see the bare-count rule above).
TODO/FIXME-flagged comments deliberately bypass the normal lore-terminal length/multi-line gate (see Game Design), while license/copyright headers are explicitly excluded from becoming lore terminals at all — both are content-filtering decisions about what "authored lore" should mean, made after noticing the literal effect each would otherwise have (most real TODOs are one-liners; nearly every file's own license header would otherwise qualify as a lore comment by length alone). [notes: Task 39, 42]
public/'s favicon set was generated from a single user-supplied 1024x1024 RGBA source image (a "CODE" wordmark logo with a genuinely transparent — not just visually dark — background; confirmed via the PNG's alpha channel before trusting it, since a flattened preview render can look opaque even when the underlying file isn't). Generation cropped tight to the logo's actual content (ignoring a few thousand near-zero-alpha stray pixels scattered outside it — alpha > 40 was the real bounding box) rather than using the full 1024x1024 canvas, since the source had a lot of dead transparent padding that would've made the logo tinier than necessary at small sizes. The transparent PNG variants (favicon-*.png, android-chrome-*.png, favicon.ico) keep the real alpha channel; apple-touch-icon.png and the mstile-*.png Windows tiles are composited onto an opaque #1a1a1f background instead (matching --bg in src/style.css) since iOS/Windows historically render transparent icon regions as solid black rather than see-through. The android-chrome-maskable-*.png pair keeps the logo scaled down further (content within roughly the inner 66% of the canvas) so it survives an OS's own circular/squircle mask without clipping the wordmark. No Safari pinned-tab mask-icon SVG was generated — that needs a proper vector trace of the logo silhouette, and this project has no vectorization tooling; low priority anyway since Safari deprecated pinned tabs. At 16x16 the wordmark is only just legible (an inherent limit of a 4-letter beveled/textured logo at that size, not a generation bug) — 32x32 and up render cleanly.
Static-meta SEO only (meta description, canonical link, Open Graph/Twitter Card tags, a minimal schema.org VideoGame JSON-LD block, robots.txt, sitemap.xml) — deliberately not full SSR/prerendering, which would be a real architectural change (the app is a pure client-side Vite SPA, no vite-plugin-ssr/Astro/similar) rather than an "SEO pass." Modern crawlers (Google/Bing) execute JS and can already index the SPA's content; static meta tags mainly matter for link-preview bots (Slack/Discord/Twitter/etc.) that don't, plus the canonical/robots/sitemap trio every crawler reads regardless. The intro screen's marketing copy (index.html's #intro-screen) is static markup, not JS-injected, so it's already visible to a no-JS crawler without any extra <noscript> duplication needed — only a one-line <noscript> "requires JavaScript" notice was added, for honesty rather than content coverage.
Canonical URL is https://codeenstein3d.mcdope.org/ — the intended hosting domain (no live deploy exists yet at the time this was added; CI at the time only had .github/workflows/verify.yml's test/build gate). Both og:url and sitemap.xml's single entry (this is a one-page SPA — no other indexable routes) use the same URL, one constant to update if the real hosting domain ever changes. .github/workflows/deploy.yml was added later — it builds and FTP-uploads dist/ on a tag push (or a manual workflow_dispatch run, pinned to the most recent tag), using host/user/password from the DEPLOY_HOST/DEPLOY_USER/DEPLOY_PASS GitHub secrets.
The Open Graph/Twitter Card image (public/og-image.png, 1200x630) was generated the same way as the favicon set — same source "CODE" logo, cropped to its real content bbox — but composited onto an opaque #1a1a1f canvas rather than kept transparent, since social platforms render OG image transparency unpredictably (often flattened to black or white depending on the client). The tagline is rendered directly onto the image in LiberationMono-Bold.ttf, matching the game's ui-monospace UI font stack.
main.ts's document.title = "Codeenstein 3D (Build: ...)" (notes Task 58) runs at module load and is intentionally left untouched — it's a live debug aid for catching a stale cached bundle, not page metadata. og:title/twitter:title are separate static meta tags unaffected by that runtime overwrite, so a crawler reading the static <head> (or even one that runs JS, since only document.title changes, not the meta tags) always sees the stable, non-timestamped title.
On 2026-08-23 four user@host SSH lane targets were found in notes and
history.md — copied out of the gitignored ssh-hosts.env, in a public repo,
where they had sat for three weeks. They were then published a second time in
the pull request description of the PR that removed them, minutes after the
rule against doing exactly that had been written down.
The second leak is the one that set the design. Intent, a written rule, and
having just thought about the problem were all present, and none of them worked.
So the control is scripts/check-secrets.mjs on a pre-commit hook, a commit-msg
hook, and a CI job that also reads the PR description.
Three decisions inside it are worth keeping:
- Two layers. Exact matching against the real values has no false positives but needs the values; shape matching needs nothing but has false positives. Neither alone covers both a local commit and a fork PR.
- The scanner never prints the matched text. It reports
path:line [rule]. Reporting the value would put it in a public CI log — the same mistake in a new place. Pinned by a test. - Shape rules can be suppressed per path; exact rules never can. Files whose
purpose is host-shaped examples (
ssh-hosts.env.dist, the coturn ACL config, the multiplayer server's IPv4 fixtures) would otherwise fire forever. No doc or notes file is on that list, deliberately.
VITE_* values are excluded from the denylist: Vite inlines them into the
shipped client bundle, so they are public by construction and matching on them
flagged package.json, index.html and the sitemap.