-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathnotes
More file actions
253 lines (199 loc) · 68.9 KB
/
Copy pathnotes
File metadata and controls
253 lines (199 loc) · 68.9 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
# Playtest notes / backlog
Open work only. When an item is done it **moves** to [`doc/dev/history.md`](doc/dev/history.md) — it does not get a checkmark and stay here, which is how this file grew to 33 completed entries and an item complaining about them.
Durable rationale ("why is the code like this?") belongs in [`doc/dev/decisions.md`](doc/dev/decisions.md); the chronological record, including approaches measured and reverted, is in `history.md`.
**Audited 2026-08-12** (the previous pass was 2026-08-10). Every open item making a checkable claim was re-verified by *running* something — a grep against `master`, a generated map, or the code itself — not by re-reading the item. That distinction is the point: several claims corrected below had been read past repeatedly, and three separate false claims were written into this file and into `decisions.md`/`history.md` in the two days before this audit.
**Corrected in this pass:**
- **Stage 6 was not open work.** It was built, measured at full protocol scale and reverted as a no-op — `decisions.md` had carried that verdict the whole time while this file described it as a stage still to build.
- **The lane hosts are not commented out.** `ssh-hosts.env` lists four, none commented; the MP-campaign item said the opposite. Heat is why that campaign is parked, not lanes. — **STALE WITHIN HOURS (noted 2026-08-15): the laptop lane was commented out at 23:39 the same evening and still is.** Three lanes, not four. This line is left standing because it is what made the MP item's "worth investigating" bullet look credible for two days; probe the file, never a remembered count.
- **Player display names shipped** (2026-08-10, `getPlayerDisplayName` in `engine.ts`); the item still said no per-player identity concept existed. Removed.
- **`pickThreat` is at `combatPolicy.mjs:1282`** as of 2026-08-23 — it has been `~891`, then `~980`, and this warning about stale line numbers had itself gone stale by 300 lines. Grep for the symbol; do not trust a line number here.
- **319 lines of finished multiplayer work (steps 1-7) moved to `history.md`**, where steps 8-11 already were. I moved 8-11 when they landed and left 1-7 behind, so the backlog carried them for three weeks.
**Verified still true**, so they need no re-checking: `pickThreat` filters `e.alive && e.aggroed`, and both gated and ranked on `Math.hypot` — confirming the walking-distance item, which was then *fixed* later the same day (see `history.md`); no threat give-up exists (reverted in `8390c8e`); the bot dispatches zero mouse events, so its aim is still keyboard-quantised; `VIRTUAL_STEP_MS` is still 50; `BOT_AIM_LEAD` is off (**no longer true — it ships ON gated at 2 tiles since 2026-08-14**) and `BOT_ANGULAR_FIRE_GATE` on; `BOT_WALKING_DISTANCE_BLOCKERS` defaults true; `DECORATIONS_ENABLED` is false; `hitFraction` is still `halfWidthPx / scatterPx`; `enemyAi.ts` still has exactly one `spawnProjectile` call site and no weapon concept; `HOST_NAV_DEADLINE_MS` and `perfDebug.ts` exist; the Firefox CI skip is still in place (4 sites).
Items recording a *measurement* rather than a code claim (the demo-campaign difficulty numbers, the damage-source breakdowns, the harness-specific wedge finding) were not re-run — they are history, and their dates say when they were true.
## Open
Grouped by area, and **within each group ordered by dependency** — the first item in a
section is the one whose answer changes the others, not the most urgent. Cross-section
dependencies are named in the items themselves. One-liners that have never been checked
against the code live in [Not yet thought through](#not-yet-thought-through) at the end.
### Bot & instrument
The bot is the measuring instrument for everything else here, so its defects are measurement defects — which is why the balancing, measurement and capture-tooling items sit under this heading rather than in one of their own (merged 2026-08-24). (The old gate here — "nothing bot-shaped is worth scoping until the first item's A/B lands" — was removed on 2026-08-23: that A/B, `turnSplitIntent`, landed 2026-08-09 and the gate had been open ever since.)
**Ordered by gain, then effort — gain meaning "how much closer does this get the bot to being usable as a deathmatch opponent" (user's axis, 2026-08-24).** That replaces the dependency ordering this section carried until now. The first item is the exception: it *defines* the axis rather than scoring on it.
1. **The deathmatch bar** — not work, it is the ranking criterion. Note it holds four promoted defects (no threat give-up, tracking through walls, the wedges, route/replan churn) that are not top-level items and would each rank high if split out.
2. **Make the bot more human-like** — highest gain and the cheapest way in: the analog-axis spike needs no engine change at all, and path smoothing reuses the BFS output and the sight predicate that already exist. Its smoothing bullet is also the *untried* approach to item 3's gait, which is why it goes first.
3. **What the rework did not fix** — the same gain at much higher cost and known risk: the planned gait fix already cost 25pp of qualify rate on Pro/hard, and the first-shot change is rated the riskiest remaining. Attempt after 2, not instead of it.
4. **Three live ghidra defects** — medium gain (an opponent picking the wrong gun at range reads as bad play rather than as inhuman), medium effort, and the sub-item that settles it needs machine time.
5. **`mineRetreat` oscillates** — lowest gain of the bot items, since mines may not feature in a deathmatch level at all, but **the cheapest thing in this section**: one branch, and the item already names the query that measures it.
6. **Demo campaign exhausted as an A/B substrate** — no gain on this axis, but it gates *proving* any of the above: a substrate that can no longer separate two arms cannot validate a bot change either.
7. **Campaign difficulty** — no deathmatch gain, and blocked on manual playtests rather than on effort.
8. **Level-1 drift risk** — no gain, and its own text says to do it only if the staging path misbehaves again.
**Where effort inverts that order**: take 5 ahead of 4 when the appetite is one evening, and 6 ahead of everything when the next change has to be *measured* rather than merely shipped.
Two points the measurement half's own ordering note made, still true after the merge: the ghidra item is about ghidra's damage model *and* its telemetry being wrong, so a capture booked on either would measure the instrument rather than the balance; and the campaign-difficulty item needs no machine time at all — it is answered offline by the solver, and it says what a capture would and would not add.
- [ ] **The bot is destined to become the deathmatch opponent, which changes what "good enough" means for every open bot item (user, 2026-08-09).** Recorded because it reorders the backlog rather than adding to it.
- **The bar changes.** As an *instrument* the bot only has to be unbiased and stable — it may play badly as long as it plays badly *consistently*, which is why several defects were tolerable. As an *opponent* it has to be good and fun, and "consistently mediocre" is the failure mode rather than the escape hatch. `multiplayer-research.md` already files "adapt balancing AI for deathmatch/bots" as its own initiative and warns it is not an adaptation of existing code; this is the reason to believe that.
- **Which open items get promoted.** They stop being measurement noise and become gameplay bugs the player would *see*:
- **Threats have no give-up** — an opponent that locks onto something it cannot reach and stops playing is unmissable in PvP.
- **`pickThreat` uses `visible` mostly as a sort key** — an opponent tracking you through a wall reads as cheating, which is worse than reading as stupid. **Narrowed 2026-08-12, not closed**: an occluded enemy is now filtered out when it is far *by corridor* (the walking-distance fix), so the blatant case is gone, but one occluded a few tiles away through a thin wall is still a live target and still tracked through it.
- **The wedges** (curl 2.0% of level-visits, improved but unfixed) — a bot that stops moving is the single most visible failure a deathmatch opponent can have.
- **Route/replan churn**, still unresolved.
- **Three design choices are instrument-shaped and would be wrong for an opponent.** Worth knowing before anyone treats the current bot as a starting point:
1. **Keyboard-only input.** `dispatchSegment` sends synthetic `KeyboardEvent`s and deliberately never the mouse. Aim is therefore quantised to key-hold durations — today's turn-split got the *granularity* right, but the mechanism is still "hold a turn key for N ms", not "point here". A credible deathmatch opponent wants direct aim.
2. **A 50ms decision window** (`VIRTUAL_STEP_MS`), i.e. 20 decisions/sec, and `MultiplayerBot` raises `minDecisionMs` further for lockstep. Fine for measurement, coarse for duelling.
3. **The profiles are bot-shaped, not player-shaped.** Casual/Gamer/Pro exist to span a *measurement* range (`fireAngleEps`, `ammoThrift`, `selfHarmAversion`); as difficulty tiers a player picks, they would want re-deriving from "how does this feel to fight" rather than "how does this spread telemetry".
- **Measured evidence, 2026-08-21**: the balance pass (enemy HP 25 -> 35, Edge Cases 10-15 -> 25-35, trash healing cut) moved Gamer **+1%** and Pro **+3%** on the regenerated board while **Casual fell 71% and stopped finishing the campaign — 17 levels cleared before, 6 after**. Two profiles absorbed a difficulty increase that the third could not survive at all, which is the ladder failing in exactly the place a *player-shaped* Casual would need to work. Nothing caught it: Casual's board bar is "clear level 3", so a 6-level run qualified and was recorded as a pass, and the pre-merge smoke used Gamer and Pro — the two that turned out fine. **Read as a bot artifact, not a balance problem** (user's call, 2026-08-21): Casual has always sat too close to Gamer to be a real low-skill model, so what this exposed is the profile ladder rather than the game. Worth keeping because it is the first hard number on how compressed the ladder is under load.
- **Today's aim fix is on this path**, not a detour: `turnSplitIntent` is the first change made because the bot should *play* better rather than because a metric was contaminated.
- **The architectural groundwork is done and its throughput half is closed** — see `history.md` (2026-08-19). The bot and engine now share their tile/sight predicates, and the harness's throughput ceiling turned out to be one Node event loop, fixed by sharding lanes rather than by moving decisions into the page. What deathmatch still needs from that thread is the *in-page decision loop*, which was scoped but deliberately **not** built: it buys nothing for throughput now, and a shipped opponent cannot use `__codeensteinTestHooks` at all, since the hooks are gated on `TEST_HOOKS_BUILD_ENABLED && isTestHooksActive()` and a build folds the block away — note the constant is load-bearing: gating on the *call* alone shipped them for a month (see `engine.ts:235-244` and `scripts/check-bundle-hygiene.mjs`). That gap, not threading, is what to scope first.
- [ ] **Make the bot more human-like** (raw list from the user, 2026-08-24, offered as starting points rather than findings). Checked against the code: **two of the five are already shipped, one is right but blames the wrong component, and the genuinely new material is a mechanism none of the existing items name.** Read with the two items above — the goal is already recorded as "the bot is destined to become the deathmatch opponent", and the gait complaint is already recorded as "the grid-locked gait"; this item is the *how*, not a fourth statement of the *what*. **Decide what human-likeness is for before building any of it**: telemetry validity (numbers that transfer to human play) and watchable footage pull in different directions, and only the first justifies re-baselining every stored campaign.
- **"BFS gives implausible movement" — right symptom, wrong culprit.** The search is fine; the *path-to-waypoint* step throws away everything that would make it look human. `pathToWaypoints` (`pathfind.mjs:84-86`) maps every tile of the BFS result to its centre with no collinear collapse and no string-pulling, and `bot.mjs:463` says so in its own words ("planRoute emits one waypoint per tile, so a winding corridor gives a new bearing every tile"). 4-directional BFS plus one waypoint per tile is a Manhattan staircase walked centre-to-centre — a human cuts the corner and never visits the centre of anything. The cheap fix is a funnel/line-of-sight string-pull over the same BFS output (collapse collinear runs, then shortcut any waypoint pair with clear sight), which reuses the bot-side `hasLineOfSight` that already exists; any-angle search (Theta*) is the expensive version of the same idea. Replacing BFS itself buys nothing.
- **"Aiming jitters" — confirmed still present 2026-08-24, and it is the turn *mechanism*, not the aim math and not the decision window.** Measured by decoding the shipped `defaultHighscore.ts` board and counting runs of held `KeyQ`/`KeyE` per recorded frame — repeatable offline, no browser and no campaign. The camera starts or stops turning **7.2 / 9.4 / 10.4 times per second** on Casual / Gamer / Pro, and the share of turn bursts lasting a single 60Hz frame is **18% / 26% / 40%** (median burst 6 / 2 / 2 frames; turn duty 57% / 36% / 30%). The mechanism is confirmed rather than merely consistent: `turnBurstMs` is `|Δ| / (ENGINE_ROT_SPEED × rotSpeedMultiplier)`, so a bigger multiplier shortens every burst, and the one-frame share climbs exactly in step with the recorded multipliers 2.0 / 3.5 / 5.0. **The better the profile, the worse it looks** — which is the wrong way round for a deathmatch opponent, since that would be Pro-shaped. The bot dispatches only synthetic `KeyboardEvent`s and never the mouse (`bot.mjs:2219`), so the view moves only while a key is down and is frozen for the rest of the decision. (An earlier draft of this bullet blamed `WATCH_STEP_MS`'s 130ms headed window; that is the wrong number — these figures come from a 50ms headless recording, and the defect is independent of the window.) **Instrument now exists**: `npm run report:turn-cadence` scores any recorded board — turn duty, start/stops per second, run lengths, one-frame share, and per-frame rotation magnitude — so a future attempt is one run to measure, not a rebuild. Note it must be run against a board recorded at 1/60 (`RECORD_STEP_MS`); the telemetry harness records one 50ms frame per decision and cannot express per-frame smoothness at all.
- **"Steer in velocity space" — tried and rejected 2026-08-24, do not re-attempt without reading why.** The engine does accept a continuous turn axis the bot never used (`gpTurn`, `input.ts:589-591` applied at `engine.ts:4242-4255`), and it is reachable by shimming `navigator.getGamepads`. It buys nothing: the turn branches pass the burst as the intent's own `durationMs`, so the axis sits at **1.0 on 98.6% of frames** and re-encodes the key exactly. **The jitter is the decision being shortened to the length of the turn, not the input modality.** Spreading it over the full window did help — median turn run 1-2 -> 4 frames, one-frame bursts 47-50% -> 37-41% — but one of two runs then reached 6 levels where every other arm reached 15, the same shape as the per-key-hold attempt that cost 25pp. Real mouse look also works (`profiles.mjs`'s "not available to Playwright" claim is false) and measured *worse*. Full write-up, including the two metric errors and the substrate that invalidated the first round, in `decisions.md` under *Playtest-Bot Behaviour: Approaches Measured and Rejected*.
- **"Preferred range, keep LOS, circle-strafe, break LOS to reload" — three of the four already ship.** Standoff: `STANDOFF_MIN_TARGET_HP` 500 / `STANDOFF_DISTANCE` 5 retreat along the escape vector from anything big inside 5 tiles. Circle-strafe: `COMBAT_STRAFE_FLIP_TICKS` 8 reverses the sidestep every 8 engaged decisions, gated by `COMBAT_STRAFE_MIN_DISTANCE` 1.5 and `strafeIsSafe`. LOS: `hasLineOfSight` gates every shot and a broken one drives an approach rather than a fire. **"Seeded reversals" would be a change** — the flip is a fixed 8-tick period today, and any randomisation has to be drawn from a seeded PRNG or replays stop reproducing.
- **Reload is the one genuinely missing piece, and it is a good one.** There is no `KeyR` anywhere in `scripts/lib/` — the bot never reloads deliberately. The engine only starts a reload when a trigger pull finds the magazine short (`engine.ts:5669`), so the bot *always* spends its 1.1-2.0s mid-fight, in the open, at the worst possible moment. `EngineStats` already exposes `reloading`/`reloadRemainingSec` (its doc comment was written for a bot to read), and the bot already holds fire on them — the missing half is *initiating* one during a lull or behind cover. `reload` is already in `InputSnapshot`, already recorded, already networked.
- **"Decouple view from velocity" — already done.** `movementKeysFor` picks the 8-way key combination closest to the desired direction, and its own doc states the bot "can move in any of these eight directions at full speed without turning at all"; `movementVectorFor` exists precisely because, once the bot can strafe and reverse, "ahead" and "where I am going" are different vectors. The residual is the 22.5-degree octant quantisation, which is exactly what the analog axes above would remove.
- **The prior art that constrains all of it.** Widening diagonal movement to every turn-and-move branch was measured and reverted: Casual/normal level-2 death rate went **0% -> 72%**, because `diagonalScale` (1/sqrt2) cuts the *forward* component by 29% and in the hazard-crossing and critical-health branches forward is the survival axis (`combatPolicy.mjs`, `diagonalStrafeKey`'s doc). Velocity-space steering has to re-answer that, not assume it away. Likewise the gait item above: per-key hold durations were the planned fix for exactly this and cost 25pp of qualify rate on Pro/hard, because a short decision during heading correction *is* the control loop.
- **Cost of admission, whatever is built.** Every stored campaign dataset becomes non-comparable, and `generate-default-highscore.mjs` records the shipped board with this bot, so `defaultHighscore.ts` needs regenerating. Grade any candidate with `report-profile-separation.mjs` before believing it — "more human" must not collapse the Casual/Gamer/Pro ladder, which the 2026-08-21 evidence above already shows is compressed under load. And validate the instrument *before* booking a campaign on it, not through one.
- [ ] **Bot: what the rework did *not* fix.** Two structural faults survived, both because the obvious fix was measured and made things worse:
- **The grid-locked gait.** `driveToward` still aims at BFS tile centres a tile apart, and a turn still truncates the step issued with it. Giving each key its own hold duration was the planned fix; it cost 25pp of qualify rate on Pro/hard and raised level times on all three combos, because a short decision during heading correction *is* the control loop. Any retry has to keep re-evaluation frequent while turning. Symptom to watch: headed pure-turn share, measured 30% -> 44% after sprinting landed.
- **It still concedes the first shot** (`pickThreat` filters `.aggroed`). Planned as its own stage and not attempted — the plan rates it the riskiest remaining change, since aggro is sticky and waking enemies that would have stayed asleep can't be undone within a level.
Optional and unattempted: sprint-backpedal retreats instead of turning 180° and running blind (touches `criticalHealth`/`mineRetreat`, the branches the 0%->72% diagonal-strafe regression came from).
- [ ] **Three live ghidra defects, and one capture that would settle the rest (restated 2026-08-23).** The long-form investigation this item used to carry is closed and moved to `history.md` — including the finding that matters most for anyone tuning weapon choice, that the cluster fast-path and not `scoreRangedWeapon` is what selects ghidra and the shotgun. What is left here is only what is still open.
- **The scorer claims 150 per rocket; the engine delivers ~131.** `ROCKET_ENEMY_TRIGGER_RADIUS` is 0.4, so a direct hit detonates up to 0.4 tiles off centre and `rocketDamageAt` applies blast falloff from there. `expectedDamagePerShot` returns a flat `damagePerPellet` inside ~10 tiles, which is exactly the band the bot fires in. **~13% optimistic**, and it changes weapon choice, so it wants its own A/B rather than a quiet edit.
- **`expectedDamagePerShot` also under-values ghidra past ~10 tiles for a reason that does not exist** — it applies `widthOverScatter` to every single-pellet weapon (`combatPolicy.mjs:1786`), discounting a rocket to 0.69/0.49/0.35/0.26 at 11/12/13/14 tiles. A rocket has no cone: `fire()` returns before `resolveShot`, so it flies exactly along `dirX/dirY`. Dormant, not harmless.
- **`report:damage-model`'s ghidra rows measure nothing at all.** It joins `damageDealt` to a `shot` on a shared timestamp, and a rocket's damage lands seconds later at the far end of its flight, so it reports 100% miss and ~0.0 damage/shot in every range bucket. Its header's "a 2026-08-06 capture measured ghidra 93 against 150" does not reproduce from any capture; the retraction sits next to the claim rather than replacing it, because whatever produced 93 is still unaccounted for.
- **The one action: take a schema-3 capture.** `rocketDetonated` exists (schema 3, 2026-08-20) and records `direct`/`dist`/`enemiesHit`/`dmg`/`selfDmg` plus a `hits` array in `damageTaken.by`'s shape. **Every archived capture is `"v":2`**, so the archive cannot answer any rocket question — but nothing else is blocking. **Corrects two things this item used to say**: that the work was blocked "until a detonation event exists" (it exists), and that rockets are too rare to measure — `balancing_capture_dens_null` fired **310 rockets in 105,651 shots, 0.29%**, roughly 3x the 0.105% figure quoted from the older archive, and a usable sample rather than a rounding error.
- [ ] **`mineRetreat` oscillates, and the old "too rare to measure" caveat on it was wrong (re-read 2026-08-23).** The 2026-07-31 backpedal fix and its root cause are in `history.md`; what is open is what the audit found. The branch fires in **39 archived captures across 267 anomaly rows** — not the zero the item used to claim — and **every** occurrence is an oscillation finding, e.g. `balancing_capture_dens_null/logs/Gamer-hard-021.log`: `branch=mineRetreat activity=route travelled=15.4t net=1.95t ratio=7.9x`. So it is measurable now, and what it measures is bad.
- **Do not reuse `KeyS` traffic as the proxy for "did the branch run".** `movementKeysFor` emits `KeyS` for three of eight octants and is called from plain navigation (`combatPolicy.mjs:2942`), so the two are unrelated. Query `branch=mineRetreat` in the anomaly rows instead.
- [ ] **The demo campaign is exhausted as an A/B substrate; use staged curl.** The 2026-08-09 difficulty reading and its 2026-08-12 supersession are closed out in `history.md` — the short version is that a bot fix, not a content change, made the campaign stop killing the bot, so it can no longer show a difference between arms. Staged curl is the replacement substrate.
- **Sequencing trap that survives all of it**: `defaultHighscore.ts` is generated from this campaign, so any content change here means regenerating it — after deciding difficulty, not before.
- [ ] **Campaign difficulty: parked pending manual playtests (2026-08-21).** The measured drift/jaggedness work, the carryover cap that shipped as `CARRYOVER_CAP_MULTIPLE = 3`, and the corpus correction that re-priced every number are all closed out in `history.md`. Two things stay here because they still gate future work.
- **`loot.ts` is not covered by the balance hash.** `simulationBalanceCoverage.test.ts` enumerates `combatConstants.ts` and `traps.ts` only, so a loot constant can move without invalidating a replay — which is either a gap or a deliberate exemption, and nobody has decided which.
- **What to write down while playing**, because it is exactly what the offline work cannot supply: *which level number* a run died on and on which repo (the finding is that difficulty tracks campaign position, not repo); *whether it was ammo or damage* — the model says ammo is comfortable everywhere in the first 20 levels (worst clear ratio 2.07x, median ~20x), so a death for lack of ammo would contradict it and is the single most valuable observation available; *roster size when it felt wrong* (first-20 rosters run median 10, p90 85, max 481); and *where a campaign stopped being interesting*, as a level number.
- [ ] **All three notions of "level 1" now agree — the solver third was ALREADY FIXED (2026-08-07), and this item claimed otherwise for four days. Only a drift risk is left.** Verified 2026-08-11 by three checks: `order()` picks `sorted.find(withMain) ?? sorted[0]` (`stage-campaign.mjs:262`), which is exactly `findEntrypointByScanning`'s `bestWithMain ?? bestOverall`; `complexityByFile` scores with the browser's own `complexityScore` and the zero-complexity filter uses that same number; and the `findEntrypointByName` stage is unreachable for a staged campaign *by construction*, since staging renames every file `${slot}_${path}` and `03_src_main.c` never equals `main.c`. The "models only the `bestOverall` branch" claim described the pre-2026-08-07 world; I repeated it without checking. The rest of the item has moved to `history.md`.
- **The residual is a drift risk, not a bug.** The solver *re-derives* the rule where the planner now *asks* for it (`campaignLevelOrder`). It agrees today; it can silently stop agreeing if `findEntrypointByScanning` or the staging naming scheme changes, and nothing would fail. Making `order()` ask too would close that — worth doing only if the staging path misbehaves again.
### Map generation & gameplay content
One dependency chain runs through this section: the decoration mechanism (a tile kind plus a texture slot) is shared by the three wall items that follow it. There used to be a second — "AST facts unlock the enemy archetypes" — and it was measured and refuted on 2026-08-20; see [decisions.md](doc/dev/decisions.md#enemy-archetypes-are-not-driven-by-richer-ast-facts).
- [ ] **Enemy diversity needs a fourth archetype; density and the weapon table already shipped.** What landed (density `COMPLEXITY_PER_EXTRA_ENEMY` 10 -> 5, per-archetype `ENEMY_WEAPONS` at unchanged mean DPS) is closed out in `history.md`. This is the half that did not.
- **Only `normal` and `edgeCase` are available inside an entity room** — flagging a member `elite` breaks "one Elite per Elite room" and multiplies the room's DPS by its size. The switch-heavy skirmisher swap that did ship reaches only **0.91% of enemies / 3.22% of entity rooms**, so a rule over the existing three archetypes has already been pushed about as far as it goes. The next step is a fourth archetype, not another rule.
- **The facts are language markers more than code markers**, which caps any fact-driven rule: private/protected is meaningful in some languages and absent in others, so a rule keyed on them is uneven across the corpus by construction.
- **The thematic direction is nearly free**: enemies read only `complexityScore` today, and the parser already knows more than that. What is *not* free is recursion, call-graph fan-in, async and deprecation — those need new AST extraction.
- **Pre-register before touching it**: any enemy-scaling change moves `SIMULATION_BALANCE`, so `defaultHighscore.ts` needs regenerating and every shipped replay is invalidated. Batch it with other balance work and pay the hash once.
- [ ] **rooms need deco of some kind, currently an empty wasteland.** Implemented in Task 21, then disabled behind `DECORATIONS_ENABLED = false` in `props.ts:16` (`mapGenerator.ts` only imports it) because billboards do not read as furniture. The WAD/billboard investigation and its conclusion — that the fix is solid, orientation-fixed geometry, not sprite art — are closed out in `history.md`.
- **The solid-tile fix only covers one of the four kinds.** A grid raycaster can express a full-height blocking tile, which suits `pillar`; `desk`/`block` need partial-height geometry, which is the same capability the windows item below prices, and neither item had noticed it depends on the other.
- [ ] **objects on walls (properly sized, not simply "slap a texture on the whole wall")** — scoped 2026-08-19, now **posters and flat fixtures only**; lamps and windows are their own items below. This is the *wall-surface* half of decoration, the floor-standing half being the "rooms need deco" item above.
- **The mechanism, shared with the deco item: a new tile kind plus a texture slot.** `raycaster.ts` already dispatches wall texture by tile kind (`hitTile === DOOR_TILE || BRANCH_DOOR_TILE : LORE_TILE : ...`), so a `POSTER_TILE` drops into that chain exactly as a `RACK_TILE` would for solid props. **And it is what "properly sized" asks for**: the fixture is drawn *within* the wall texture at correct proportions rather than stretched across it — precisely how DOOM does wall fixtures.
- **One caveat to design around:** a tile kind textures *all four faces* of that block, so a poster would appear on every side. The renderer already distinguishes the hit side (it dims y-side walls to fake directional lighting), so per-face selection is reachable — but it has to be built deliberately, not assumed.
- **Posters: showing what?** Source-derived is the thematically right answer and the cheap options are free — the file's own name/path, the entity's signature, the language. **Images from the source tree are cheap for a local workspace and currently impossible for a GitHub one**: `RemoteFileHandle.getFile()` returns `{ text(): Promise<string> }` and nothing else, so a repo-loaded workspace cannot supply bytes without extending that interface. Nothing decodes images today, though `createImageBitmap` needs no dependency. Needs a size cap and a rule for *which* images.
- [ ] **wall lamps / flares.** Split out 2026-08-19 from "objects on walls". **There is no lighting model to hang this on** — the renderer has distance fog plus a fixed dim on y-side walls to fake directionality, and nothing else. No point lights of any kind. So the sub-item is really two, with very different costs:
- **Glow is cheap and worth doing.** An emissive fixture texture plus a local brightness boost, modulating the `fogShade(dist)` scalar the renderer *already computes per column* (`raycaster.ts:104`, applied at `:376`) by proximity to the light. Arithmetic on a value that exists, in a pass that already runs.
- **Shadows are a different renderer.** Real occlusion means testing every light against every column, inside the loop `perf-review-2026-08-02.md` measures at ~93% of frame time. Either decline it outright or scope it alone with a frame budget agreed *before* any code — this is the sub-item most likely to quietly eat the 60fps target.
- Note the two are separable in the fiction as well as the code: a lamp that glows but casts no shadow reads fine, which is what most of this era's games actually did.
- [ ] **windows into the next room, with enemy aggro through them.** Split out 2026-08-19 from "objects on walls". **Constraint from the user: a window must not be a whole wall tile — nothing goes floor-to-ceiling.** That single requirement is what makes this much harder than "a see-through tile", and it is worth understanding before scoping:
- **A grid raycaster draws full-height columns by construction.** A partial-height opening means one screen column has to be *banded* — near wall above, the far surface through the middle, near wall below — which needs **two DDA results per column** (near hit and whatever is behind it), two texture slices, and its own fog distance for the far slice. The DDA today stops at the first solid tile and has no compositing at all.
- **And the depth buffer cannot express it.** `zBuffer` is a `Float64Array` indexed by column — **one depth per column** — and the sprite pass occludes with `proj.depth >= zBuffer[x]`. A banded column has two depths, so an enemy seen through the window is either wrongly hidden (near wall's depth stored) or wrongly drawn over the wall (far depth stored). Fixing it properly means a per-column *span* list rather than a scalar, which changes the sprite pass too. **This is the real cost of the feature, not the transparency.**
- **The simulation half is smaller but genuinely new**: a window is **solid to bodies and bullets, transparent to sight** — a *third* category beyond the two `src/engine/mapPredicates.ts` has (`isWall` = solid to a body or a bullet, `isRouteBlocking` = solid to a route planner). The engine and the bot need it *differently*: an enemy should aggro through a window, while the bot must not try to shoot through one. That module already expresses exactly this kind of divergence as an argument (`sealedCorner`), so it is the right home.
- **Then decide whether the player can shoot through.** If yes it is a firing port and changes tactics; if no, "I can see him but cannot hit him" has to be legible from the art alone.
- **Cheaper things that get most of the feel**, worth pricing against the above before committing: a full-height barred opening (no banding needed, still a window fictionally), or a purely decorative lit alcove with no sight line at all.
- **The span list is not windows-only (2026-08-20).** The same per-column span buffer is what `desk`/`block` in the "rooms need deco" item need in order to stop being floor-to-ceiling, so if this ever gets built it should be priced against both items rather than windows alone. The cost is in the fan-out, not the DDA: `zBuffer` has 94 occurrences across 6 non-test engine modules.
- [ ] **proper collisions for player and enemies (and enemies/enemies)** — today *nothing* collides with anything except the tile grid. `Player#move` (`player.ts:86`) and the enemies' `slideAxes` (`enemyAi.ts:322`) both run the same per-axis AABB slide against walls and nothing else, so players walk through enemies, enemies walk through each other, and a pack spawned at a room centre (`generation/enemies.ts` anchors a pack's first member there) can stand in one spot. The collision boxes exist and differ — player half-width `0.2` (`player.ts` `DEFAULT_CONFIG`), `ENEMY_RADIUS = 0.3` — they are just never tested against each other. **The catch is that this is a level-design change wearing a physics costume**, and the cost lands almost entirely on the map, the bot and the balance baseline rather than on the collision code, which is a dozen lines.
- **MVP: you cannot walk through an enemy that is biting you.** Melee contact blocks the player's step, so an engagement has to be resolved (kill it, or back out) instead of being walked through. Deliberately *only* the player-vs-melee-enemy pair — enemies still pass through each other, and a non-engaged enemy still passes through you. Four properties make this the cheap slice rather than a down-scoped version of the full item:
- **The enemy half already exists.** `updateEnemy` returns before its chase step whenever `dist <= ATTACK_RADIUS` (`enemyAi.ts:183-190`), so an attacking enemy already halts at exactly contact distance and never walks into the player. Only the player side is missing.
- **It needs no new constant.** Use the same Euclidean `dist <= ATTACK_RADIUS` test the melee attack itself uses, so "it can bite you" and "you cannot pass it" are one predicate. `ATTACK_RADIUS` is already inside `COMBAT_BALANCE` and therefore already in `SIMULATION_BALANCE`, so the player-radius hashing gap below does not apply to the MVP at all. (The coincidence is worth noting: player `0.2` + `ENEMY_RADIUS` `0.3` = 0.5 = `ATTACK_RADIUS`, so contact and melee reach are the same threshold either way.)
- **It is dt-invariant for free.** A rejected step is a positional predicate, exactly like `collidesWithWall` — no accumulation, no push, nothing that integrates. None of the dt traps below apply.
- **Never block a retreat.** Reject an axis step only if it would *decrease* separation below contact; always allow one that increases it. This kills the pin-in-a-corner failure the full item has to design around, and keeps the 1-tile-corridor case honest — a melee enemy in a squeeze does block the way forward, which is the point, but you can always back out of it.
- **Where it goes**: `Player#move` takes the wall check as `this.collides` (`player.ts:97`); the blocker predicate wants the same treatment, threaded through `moveForward`/`strafe` from the four call sites in `handleMovement` (`engine.ts:4237-4245`) which already pass `this.map`. Living enemies only — `alive` already excludes corpses, matching multiplayer's own no-corpse-collision rule.
- **What the MVP still costs**: it is a simulation change, so it moves nothing in the hash but does change recorded outcomes — replays and `defaultHighscore.ts` still need the batch treatment below. And it still breaks the bot's documented pass-through assumption (`bot.mjs:667`): a bot pushing into a melee enemy in a corridor can now stall where it previously walked through. Check `ignoreThreats` and the campaign baselines before trusting any number gathered across it.
- **Hard blocking closes most of the map.** Corridor width is rolled per corridor at 60% one tile, 30% two, 10% three (`CORRIDOR_WIDTH_WEIGHTS`, `corridors.ts:118`). In a 1-tile corridor the player's centre is confined to a 0.6-tile band (0.2 clearance from each wall) and an enemy's to 0.4, so the widest separation physically reachable is exactly **0.5** — which is precisely the combined half-width at which they touch. A single enemy standing in a one-tile corridor would be an impassable wall, in 60% of corridors. The MVP above accepts exactly this for a *melee-engaged* enemy and only in the direction of the enemy, which is the mechanic; what is not acceptable is an idle or unaware body doing it, or doing it to a player with nowhere to retreat to. That is the same doorway-blocking failure the multiplayer spec already refused player-player collision over (`multiplayer-game-state-spec.md:848`).
- **So the shape has to be soft separation, not blocking.** Enemies get a push-apart correction applied *after* the wall slide and re-checked against `collidesWithWall` so it can never shove a body into geometry; the player is never blocked by an enemy at all, but displaces one it walks into. Player-always-wins is what keeps "pinned in a corner by two melee enemies" from becoming an unrecoverable death, and it is the only version that leaves every existing level layout passable.
- **Enemy-enemy is the part that changes the fight.** Chasers steer down a shared BFS field (`pathField.ts`, via `nextWaypoint`) and today a whole pack converges on the same waypoint and overlaps into one silhouette. Separate them and a pack conga-lines through a 1-tile corridor instead — the followers jam behind the leader, and whether that reads as tactical (chokepoints matter) or broken (enemies milling in a doorway) is a playtest question, not a code one. Roaming is confined to `withinHome`, so the clumping that is actually visible today is mostly aggroed packs, not idle ones.
- **Two determinism traps.** Separation must be resolved in a fixed order (ascending enemy array index — the `EnemySpatialGrid` contract already exists for exactly this reason) or the same frame resolves differently depending on iteration order. And it must be **dt-invariant**: the balancing harness integrates at 50ms (`MAX_DT`), `generate-default-highscore.mjs` records at 1/60, and a player runs at 8-16ms — a fixed per-frame nudge would separate ~3x faster on the bot than in a real run. `dtInvariance.test.ts` is the existing gate; extend it rather than inventing a check.
- **It moves the balance fingerprint, and exposes a gap in it.** Entity collision is simulation, so every shipped replay is invalidated and `defaultHighscore.ts` needs regenerating — batch it with the enemy-archetype item above and pay the hash once (same pre-registration note). Worth fixing while in there: `SIMULATION_BALANCE` (`engine.ts:487`) covers `COMBAT_BALANCE`, so `ENEMY_RADIUS` is hashed, but the **player's** radius lives in `player.ts` `DEFAULT_CONFIG` and is in nothing — today that is harmless, and the moment it governs entity contact a tweak to it would drift every replay silently. Same class as the `loot.ts` gap noted above.
- **The playtest bot assumes pass-through, in writing.** `bot.mjs:667` documents `ignoreThreats` as safe *because* "the engine's own collision only checks map geometry, never other entities, so a bot that simply never stops to fight can walk straight past/through an enemy" — `verify-multiplayer-transition.mjs`'s god-mode host depends on that. Expect wedges there, and expect every stored campaign dataset to become non-comparable across the change.
- **Multiplayer is already decided and should stay decided.** Player-player collision: none, by construction (`multiplayer-game-state-spec.md:839-849`) — it also makes the spawn-shortfall rule a one-line modulo (`:195`). Dead players leave the world simulation, so no corpse collision either (`:873`). Enemy-vs-player collision in coop would be a new desync surface on top of lag compensation; keep it single-player-shaped until it is proven fun.
- **Perf is the cheap part, with one caveat.** `EnemySpatialGrid` already exists and is O(enemies) to rebuild, but its doc comment's whole premise is that it is built only on the rare frames that need proximity queries (rockets in flight) and "costs nothing at all" otherwise. Entity collision needs it every frame — still cheap, but that comment stops being true and should be rewritten rather than left to mislead.
- [ ] **exit/return often in the middle of the map, sometimes even worse** — the exit tile is labelled `return` in-world, so "exit/return" is one thing, not two. Measured 2026-08-24 over 309 generated levels from 8 corpus repos plus all 17 demo levels, then checked against a real recording (`codeenstein-demo-campaign-replay-1787590063325.webm`, 03:44 = `stage04_widget.js`). **The complaint is correct and the first aggregate answer to it was not — see the retraction below before quoting any centrality number.** The mechanism is in `pickExit` (`spawnExit.ts:66`): it returns the *centre of the room whose centre is farthest from the spawn by straight-line distance*, and straight-line distance is not walking distance.
- **The measurable defect: the exit is in the interior of the footprint, not at a terminus.** Across the corpus the exit is **not** the walking-farthest room on **60% (185/309)** of levels. Walk-to-exit as a share of the deepest point reachable from spawn: p50 **0.74**, p10 **0.46**. Walking given up against the best available room: p50 7 tiles, p90 86, **max 464** — as a share of that room's own walk, p50 10%, p90 50%. **26% of levels leave more than 40% of the map behind you** when you reach the exit. On the 03:44 level specifically: spawn top-centre, exit mid-left, 37 tiles of a 54-tile-deep level, and the entire bottom-right half never on the path.
- **Retracted: "the exit is peripheral, the spawn is the central thing."** That was measured as distance-from-centre over half the bounding-box diagonal, which is a bad proxy for what a player sees — it scores "at the left edge but vertically centred" as 0.48 and "upper-right quadrant" as 0.55. And it was quoted from the corpus (p50 0.19) when the campaign the user actually plays spreads 0.14-0.55 with **6 of 17 levels at 0.42 or above**. The spawn genuinely is central by construction (a corner of room 0, which `mapGenerator.ts:166` jitters ±6 tiles around the map centre), and that is worth keeping — but it does not license the claim that exits are not.
- **Two things make it read as more central than the geometry alone**, both cheap to change and neither touching generation:
- **Both map widgets frame the whole grid, not the built area.** The minimap scales by `Math.max(map.width, map.height)` (`raycaster.ts:1045`) and the automap by `map.width * CELL_PX` (`automap.ts:209`), while the built area is a median **~45%** of grid area across the demo campaign — 33% on the 03:44 level, 24% on `stage16_hardware.h`. The level is drawn as a blob floating in dead rock, so anything inside it reads as mid-widget. Cropping both to the floor bounding box is a display fix with no simulation cost.
- **`pickExit` returns the room *centre***, so the exit always stands mid-room with floor on every side, rather than against a far wall the way a hand-placed exit would.
- **Cheap fix: rank rooms by flood-fill distance instead of straight-line distance.** One BFS from the spawn over the walkable grid, then take the farthest room centre. Still pure geometry drawing nothing from `rng`, exactly like the function it replaces.
- **Deeper fix: move the spawn, not the exit.** Choosing the spawn **room** — rather than only its corner within room 0 — as the one farthest from the exit converts a radial hop out of the middle into a real traversal, without changing map size or room count. `pickMultiplayerSpawns` already does exactly this for multiplayer (maximise minimum distance to the exit and to every other spawn), so single-player is the odd one out rather than this being new machinery.
- **Do not simply maximise, either way.** The walking-farthest room can be a dead-end pocket behind two locked gates, which trades a short level for a key hunt. Prefer the farthest room whose path crosses no more gates than the current one, and keep `pickSafeSpawn`'s enemy-distance safety if the spawn room stops being room 0.
- **Cost of admission**: any generation-side change here changes every map. `pickExit`/`pickSafeSpawn` draw no `rng` themselves, but the exit tile feeds `spawnEnemies`' own exit re-roll, so the draw sequence shifts downstream and the whole level differs — `defaultHighscore.ts` needs regenerating, every stored replay is invalidated (the roster moves, so `balanceHash` catches it), and the balance baselines reset. The two display fixes above cost none of that and can ship alone.
- **Caveat on the numbers**: the flood-fill treats teleporter pads as plain floor, so on levels with many pads (39 on a single vim level against 2 in the whole demo campaign) real travel can be shorter than measured. The "not the farthest room" comparison is unaffected — both sides use the same metric — but "how much of the map lies beyond the exit" is an upper bound.
### Rendering, HUD & assets
- [ ] **HUD art from a loaded WAD — the enhancement layer left over from the DOOM status bar (shipped 2026-08-20).** The bar itself, its vocabulary, `hudLayout.ts` and the face are all done and written up in [history.md](doc/dev/history.md) — read that before re-scoping this, including the three premises of the original item that turned out to be wrong. What is left is only the optional art layer.
- **`STBAR`/`STTNUM*`/`STFACE*` are patch-format lumps and `patch.ts` already decodes that format generically** — no new binary parser, the same finding as the enemy-sprites item below. It needs the player's own WAD, so it can only ever be an enhancement: the bar has to look right with nothing loaded, which it now does.
- **The face needs art only — its behaviour is finished.** The original item said "nothing tracks damage direction today", and that is stale: `hurtDir` is on `PlayerState`, resolved when the damage lands, and `faceKeyFor` already keys on it (`hurt${tier}_${dir}`) alongside health tiers, god mode, death and kill streaks. So this is a sprite swap, not the engine change it was once costed as.
- **`hudLayout.ts` is the hook**, and says so in its own header — a WAD skin replaces what goes in each panel rect, not the geometry.
- **Possible perf upside, to measure not assume.** A pre-composited bar blitted with one `drawImage` may be cheaper than drawing text and shapes every frame — `perf-review-2026-08-02.md` found the expensive class is geometry Skia cannot emit as axis-aligned quads. Directly measurable: the HUD has its own `?ablate=hud` group and `perf:bench` has scenarios. Measure it against the shipped bar, in the same session (see the DVFS note in that write-up).
- **Licensing, same split as the WAD-audio item**: the player's own WAD is settled precedent from textures; catalog WADs in `public/wads/` are served from the deployed site, so shipping their HUD art is redistribution — check `onlineWadCatalog.ts`'s bar for that case.
- [ ] **enemy textures from wad? if yes: what about hitbox?** Investigated 2026-07-16, no code changed; the feasibility work and its conclusion are in `history.md`. Answer: real silhouette-accurate hitboxes are achievable, the opacity data comes free from the patch format, and the blocker is sharing frame/rotation state between the renderer and the hit test. Note it changes melee too, since `findTargetInProjections` is shared.
- **Scope if picked up**: a sprite-lump loading path in `loadWad.ts` (single-patch, no compositing), frame/rotation state on the enemy, and a per-column opacity test in `findTargetInProjections`. Enemies are 100% procedural today — `WadLoadResult` exposes `styles`/`signals` only (`loadWad.ts:64-69`) and no sprite lump is read anywhere in `src/wad/`.
- [ ] **sounds/bgm from wad (local bgm would override it)** — checked 2026-08-19, and the answer to "is that even possible" is **yes for sounds, effectively no for classic music**. Split it.
- **Sound effects are easy, and the plumbing already exists.** DOOM SFX are `DS*` lumps in about the simplest format there is: a short header (format id, sample rate — usually 11025 — and sample count) followed by **unsigned 8-bit PCM**. Decoding is a `DataView` read plus `(b - 128) / 128` into an `AudioBuffer`, the same style `patch.ts` and `flatLump.ts` already use, and `parseLumpDirectory` (`wadFile.ts`) already hands out arbitrary lumps — no new plumbing. `audio.ts` already routes everything through an SFX bus, so sample playback sits beside the procedural path and inherits the volume sliders for free.
- **Music is the hard half, and it is a dependency problem rather than a format one.** DOOM music is `D_*` lumps in **MUS**, a compact MIDI-like event stream. MUS to MIDI is a well-documented, easy conversion — and then you are holding MIDI, which is *notes, not audio*. Browsers have no synthesizer, so playing it needs either a SoundFont player (FluidSynth-in-WASM, megabytes) or an OPL3 emulator to sound like the real thing. That is a very large dependency for a project that hand-rolled a ZIP reader rather than take `unzipper` — see [Dependency Minimalism](doc/dev/decisions.md#dependency-minimalism).
- **But some WADs ship real audio and those are free.** Doom II BFG and many PWADs replace MUS with MP3/OGG lumps, which `bgm.ts`'s existing `<audio>` + `MediaElementAudioSourceNode` path already plays. So "music from WAD" is worth shipping *for WADs that contain actual audio files* and worth declining for classic MUS — detect and skip, do not half-build a synth.
- **The real design work is the mapping, not the decoding.** DOOM's sound names are DOOM's game: `DSPISTOL`, `DSSHOTGN`, `DSSGCOCK`. This one has 20 public `play*` events including `playAcidOverflow`, `playUltraKill`, `playKeyPing`, `playTeleport`, and weapons (gdb, ghidra, Friday Hotfix, Toolchain) with no DOOM equivalent at all. So it needs a curated mapping table with a per-event fallback — the same shape and the same rigour as `textureAllowlist.ts`, which already requires every name to decode against a real WAD.
- **Procedural stays the default and the fallback.** The existing SFX are a deliberate design, not a placeholder; a WAD is an *override*, and any event without a mapping must fall through to the procedural sound rather than going silent.
- **Precedence, per the original note**: local BGM folder > WAD music > silence. `bgm.ts` is already a playlist provider, so a WAD source is a second provider at lower priority rather than a rewrite.
- **Two smaller traps**: `audio.ts` creates its `AudioContext` lazily on the first user gesture (autoplay policy), so lump decoding has to happen after that or be deferred; and catalog WADs from `public/wads/` are *served from the deployed site*, so enabling this for them is redistributing audio rather than reading the user's own file — check `onlineWadCatalog.ts`'s licensing bar for that case specifically, even though the user's-own-WAD case is already settled precedent from textures.
- [ ] **panel seems large enough, change STABIL into STABILITY** — the panel is large enough, by about 2x, and width was never the interesting question. The label sits on its own row (`LABEL_DY` 14 against `VALUE_DY` 44 and `STRIP_DY` 50, `hud.ts:666-671`), so it owns the panel's full width minus padding rather than sharing a line with the percentage. At both shipped presets `layoutHud` is handed exactly 640x400, where the stabil panel works out to **~116px**: `min` 112 plus its 1.2/6.7 share of the bar's 23px surplus (`usable` 625 against a 602 total minimum). Minus `HUD_PAD` on each side that is ~100px of label row. `STABIL` was *measured* at 33px in the 9px `ui-monospace` label font, and a monospace face advances every glyph equally, so `STABILITY` is 33 x 9/6 = **49.5px** — half the room. The precedent already ships: `RELOADING` is nine characters in the same 9px label font, in the AMMO panel, whose minimum (100) is *narrower* than this one's.
- **The real argument is vocabulary, not width.** Every other panel label is a whole word — `AMMO`, `TOOLS`, `SWAP`, `KEYS`, `SCORE` — which makes `STABIL` the only truncation on the label row, and that, not its length, is why it reads wrong. (The ammo table's `BULL`/`SHEL`/`RCKT` are truncations too, but they are a different tier: a 4-char column inside a row, not a panel label.) The inline note at `hudLayout.ts:170` records that `STAB` was rejected for reading as a melee verb; going the other way to the full word settles the same question in the same spirit.
- **No layout constant has to move.** `PANEL_SPECS`'s `stabil: { min: 112 }` is sized for the bar strip under the label, not for the text — its own comment says so. `STABILITY` plus both pads is 65.5px, still comfortably under, so the min-sum-against-640 pin in `hudLayout.test.ts` is untouched and no other panel has to give width back (there are only 23px of surplus on the whole bar, so there would be nothing to give).
- **One caveat worth stating rather than discovering.** `drawLabel` is fixed at 9px with no shrink ladder — unlike `drawNumeral`, which steps down through `NUMERAL_SIZES` via `fitFontSize`. Irrelevant at both shipped presets, but the uniform-squeeze branch narrows every panel, and `STABILITY` would start clipping at roughly 367 design px where `STABIL` survives to ~272. Per `layoutHud`'s own doc that branch is reachable only through a non-8:5 `?renderRes` override, so no player can see it; it is a contract case of a pure function, not a regression.
- **Ships with two doc edits and a changelog line.** `doc/user/hud-and-ui.md:10` spells the label out in its panel list; `LABEL_STABIL` (`hudLayout.ts:178`) is one string by design ("named so a rename moves one string"), and renaming the constant alongside it is optional. Leave `CHANGELOG.md:60` and `history.md` alone — those are records of what shipped, not current claims. This one *is* player-facing, so it earns a real changelog entry.
- **Verification is visual, because nothing pins the string.** No test asserts `"STABIL"` — `hud.test.ts:892` only names it in a comment — so the check is a screenshot of the bar at **both** presets rather than the one that happens to be loaded, which is exactly how a marker shipped wrapping once before.
- [ ] **Spread a kill's drops across the tile — right now they stack and read as one item.** Every drop from a kill is pushed at exactly `enemy.x, enemy.y`: the guaranteed health pack (`engine.ts:6125`), the ammo/swap roll or its Toolchain consolation (`:6152`/`:6159`), the bonus-weapon roll (`:6164`), and both Elite pushes (`lootApply.ts:156`/`:158`/`:167`). So a regular kill can leave **three** drops on identical coordinates and an Elite **two**. The renderer applies no offset — `sprites.ts:778` projects raw `drop.x`/`drop.y` — so they draw as perfectly overlapping billboards and the player sees one pickup.
- **Cosmetic only, which is why it has gone unnoticed.** `collectLoot` walks every drop within `AMMO_PICKUP_RADIUS`, so the whole stack is picked up and nothing is lost; the bug is that you cannot tell there was anything to think about. It also makes the loot-drop screenshots misleading in the same way.
- **The trap for whoever implements it: do not spend an rng draw.** Jittering the position off the shared PRNG stream shifts every subsequent draw and re-rolls the layout of every level ever generated, invalidating stored replays — the same constraint that keeps the styleset on its own salted stream and `planAcidOverflows` at zero draws. Spread deterministically instead: the drops are pushed in a fixed order and already carry a per-enemy sequence number (`pushLootDrop`'s `dropSeq`), so an index-derived offset costs nothing and changes no existing map.
### Multiplayer & backend
The multiplayer *features* are still one-liners in the last section; this is the operational half.
- [ ] **Self-updating multiplayer backend: `docker/update.sh --install` / `--uninstall` sets up a systemd timer.** Decided 2026-08-19 (user) after checking the alternatives — see [decisions.md](doc/dev/decisions.md#the-backend-updates-itself-by-pulling-ci-never-pushes-to-it).
- **Nothing in CI changes, and no marker file or version endpoint is needed.** The git remote *is* the channel: the timer runs `update.sh`, which already `git pull`s and rebuilds. The original sketch's `.update-backend` + inotify cannot work (the frontend deploys by FTPS to a *webhost*; the backend is a different machine, no shared filesystem) and is dropped.
- **The update logic already exists — do not reimplement it.** `update.sh` aborts on a dirty tree, detects the TURN profile from the running containers, pulls `coturn` only when relevant, rebuilds/recreates, prints status, warns on `docker/.env.example` drift. `--install`/`--uninstall` only need to write/remove a `.service` + `.timer` and reload; both must be idempotent.
- **The one real wrinkle is sudo without a TTY.** `update.sh` deliberately runs as the *deploy user* and shells to `sudo` only for the docker calls, because a root `git pull` leaves the repo root-owned. Under systemd there is no TTY to prompt on, so this needs a `NOPASSWD` sudoers entry scoped to those compose commands. `--install` should write it (or verify it and fail loudly), because discovering it via a silently failing 04:00 timer is the worst version of finding out.
- **Cadence: a quiet hour, not hourly.** Restarting signaling drops **open lobbies and pending joins** — `sessions` is an in-memory `Map` with a 5-minute TTL, so it is ephemeral by design, but someone sitting in a lobby gets kicked. Established peer connections survive, since peers talk directly once the handshake is done. So: `OnCalendar` at a quiet local hour, `RandomizedDelaySec`, and `Persistent=true` so a run missed while the host was off fires at next boot.
- **Make failure visible.** A timer that silently stops is the classic form of this bug — the dirty-tree guard aborting loudly only helps if someone sees it. Add an `OnFailure=` unit, or at minimum have `--install` print the `systemctl status` / `journalctl -u` invocations and document them in `multiplayer-deployment.md` §6.6 next to the manual line this replaces.
### Platform, tooling & tests
- [ ] end2end/functional/visual tests (coverage target: all aspects of gameplay), and a dedicated audit confirming **no data leaves the app at all**. That claim used to read "besides highscores", which was forward-looking; public highscore boards were scoped and **scrapped** on 2026-08-19 (see [decisions.md](doc/dev/decisions.md#highscores-stay-on-the-device)), so the assertion is now the simpler and stronger one: `highscores.ts` is `localStorage` only and there is no outbound user data anywhere in the product. Assert it as a hard invariant — an outbound request carrying *user data* is a regression. The old wording — "a `fetch` to anything but the signaling server" — was overtaken on 2026-08-20 by multi-host repo loading: `src/fs/{github,gitlab,codeberg}.ts` and `remoteHost.ts` all fetch, as does WAD-pack loading (`main.ts:534`). Those are user-initiated GETs with no auth token and no payload, so the spirit holds and the testable rule had to change. The unit-test half of this item shipped as Task 92 — see `## Done`. `scripts/verify-demo-campaign.mjs` (structural/map-generation coverage) and `scripts/verify-campaign-playthrough.mjs` (a real headless-Chromium save/highscore/replay mechanism test) are existing starting points to build the broader e2e/visual coverage on top of.
- [ ] **Two test-harness chores left over from "threaded / forked tests to speed them up?"** The speed premise was already satisfied when that item was written, and the measurements now live in [testing.md](doc/dev/testing.md#the-suite-is-already-parallel--there-is-no-win-left-in-pool): vitest 4 defaults to forks at one worker per core, and the wall clock is the browser-driven `verify-*.mjs` scripts, which no `pool` setting can reach. Both leftovers are now *described* in [testing.md](doc/dev/testing.md#running-the-suite); neither is *fixed*.
- **`npm run coverage` is unreadable locally while `balancing_corpus/` exists** (98MB, present after `npm run balancing:corpus`). Vitest's default glob collects its vendored spec files and reports them as failures, so every local run needs `--dir src` — which is not what the gate runs, and which silently skips the `scripts/` tests including doc pins. A `test.exclude` for the corpus closes it; check first that the pattern cannot also swallow a real `src/` file.
- **CI runs `npm run coverage || npm run coverage`** (`verify.yml:79`) — the whole suite twice on any flake. Find out what the retry was papering over before removing it; it may be load-bearing. One data point since: the 2026-08-23 help-ping timeout failed the *job*, and its 5m33s runtime against a ~2m50s clean pass says both attempts ran and both failed — so a hard failure is doubled rather than hidden. A test sitting marginally over the line is the case that would still be masked.
- **If anyone does try `pool: "threads"` anyway**: the suite constructs real `RaycasterEngine`s and jsdom canvases and relies on `restoreMocks` plus per-file module isolation, which `forks` gives for free. Measure it, do not assume it is a drop-in.
- [ ] **Replays are cut short at ~12-15 min, and paused frames are what fills them.** Both halves of the original one-liner confirmed by reading the code. `MAX_REPLAY_FRAMES_PER_LEVEL = 43200` (`replay.ts:153`) counts one frame per *rendered* frame, per level — 12 min at 60fps, 15 at ~48. And `engine.ts:3181` records at the top of `simulate()`, ~60 lines above the pause/lore early return (`3243`/`3275`), so every paused frame is stored; a blur or a pointer-lock release auto-pauses, so clicking another window while the game stays visible burns the cap at full refresh rate. Overflow then drops that level's segment *and* truncates the payload to the prefix before it, so one long level throws away every level after it.
- **Stop recording frozen frames** — skip when the engine was already frozen last tick (`wasFrozen`) *and* the snapshot is fully inert: no keys, no mouseDX/wheel, no one-shot flags, no gamepad axes, `weaponRequest` null. Lossless by construction, since a frozen tick with inert input mutates nothing that survives the frame. The pause press itself still records (`wasFrozen` is false when it is captured), and so does anything eventful during a pause — a cheat code, click-to-resume, `W` held to scroll lore.
- **The cap never delivered the bound its doc comment claims**, so "add compression" is not the missing piece — compression is already aggressive (binary codec, 334 B/frame of JSON down to ~10 B, then gzip+base64; measured 1.48 MB → 0.42 MB). That puts the shipped board's 213,740 frames at ~2 bytes of quota each, so `MAX_ENTRIES`=10 × `MAX_REPLAY_LEVELS`=100 × 43,200 × 2 B ≈ **86 MB** worst case, and ten ordinary 17-level runs ≈ 14 MB. Both are far past a ~5 MB origin budget: the cap was never binding.
- **What actually protects the quota is already there.** `recordHighscore` (`highscores.ts:212-224`) catches `setItem` throwing and retries — drop this run's replay, then every replay, then keep the bare scoreboard. Graceful, at save time, degrading in the right order. The cap fires at *record* time, silently, and amputates unrecoverably.
- **So keep a per-level cap only as a memory backstop, raised.** Frames sit in RAM as JS objects for the length of the run, so a wedged or AFK level does need a ceiling — ~48 min at 60fps, not 12.
- **And change what overflow does**: keep the frames up to the cap, flag the segment truncated, keep every *later* level. Safe because segments are independent — each `ReplayLevelSegment` carries its own `carryover` (`replay.ts:97`), so level N+1 never depends on level N's frames. The viewer ends that level early with an honest "recording ends here" and moves on. The existing comment's worry ("a replay that silently skips a level is worse than a short one") is satisfied: a flagged partial level is not a silent skip. Optional extra rung on the quota ladder while in there — drop the single largest replay before dropping all of them.
- **Verify with `scripts/verify-replay.mjs`**: a recording made across a long pause must replay to the same score and end state as one without.
> **Lane hosts are never written down here.** The SSH lane machines are
> referred to only as *the NAS lane*, *the laptop lane*, *the server lane* and
> *the fourth lane*. Their real `user@host` values live in `ssh-hosts.env`,
> which is gitignored and auto-managed by `scripts/setup-ssh-lane-host.mjs` —
> this repo is public, and four full SSH login targets sat in this file and in
> `history.md` for three weeks before anyone noticed (scrubbed 2026-08-23).
>
> **This is enforced now, not remembered.** `scripts/check-secrets.mjs` runs as a
> pre-commit hook *and* a commit-msg hook (`.githooks/`, wired automatically by
> npm's `prepare`), and as a CI job that also scans the **pull request
> description** — the channel that leaked a second time, hours after the rule was
> written down. `npm run secrets:scan` checks the tree by hand.
## Not yet thought through
One-liners that have never been checked against the code — kept at the end so it
is obvious at a glance which items are ideas rather than plans. Everything above
carries at least a dated finding. Refining one usually means discovering that the
premise moved: of the last six done this way, three turned out to be already
implemented, already refuted, or impossible as written.
- [ ] Add three multiplayer coop demo campaigns
- [ ] 1x finishable in 5-15min
- [ ] 1x finishable in 10-30min
- [ ] 1x finishable in 25-45min
- [ ] should be visually different, each one should have its own set of levels. make sure each one has all features of our "code to level" pipeline.
- [ ] **multiplayer support — coop is shipped; what is left is deathmatch and bot-filling.** Coop for 2-4 players, host-chosen `maxPlayers`, real human guests and session-end handling are all live and CI-gated. Only the child below is open.
- [ ] adapt balancing AI for deathmatch mode / bots? host should define max players. for every missing player a random bot skill profile is launched. when a player wants to join: remove a bot. — still genuinely open/deferred: step 10 added real host-chosen `maxPlayers` (2-4) and real human guests filling those slots, but bot-filling of *empty* slots and any deathmatch (PvP, not coop) mode are both separate, unbuilt features.
- [ ] deathmatch multiplayer needs continous respawns of ammo, weapons, swap (the resource formerly called armor) and health. also health and swap should only be collected till full, else it should remain on the ground for other players. Note the clamp already exists (`engine.ts` caps at `MAX_HEALTH`) — the missing half is *not consuming* the pickup when it would be wasted
- [ ] add dedicated deathmatch demo campaign, levels should have fair balanced layout - similar to UT99s levels
- [ ] coop multiplayer should spawn additional boss enemies, requiring actual coop to kill. having a distinct color, roaming the whole map (not limitted to their room) and having more aggro, movement speed like normal enemies, should be visually larger as well. unique abilities like "casting mines" or weapon usage? possible names: "VPN down", "CI failed", "Useless Meeting", "Git down", "CloudFlare", "AWS", "GCP", "Azure". **Not to be confused with what already exists**: Elites are already a boss tier and already scale with player count (`eliteScalingFor`, `multiplayerScaling.ts`, CI-gated at `playerCount=3`). This item is about *extra, map-roaming* bosses — enemies are hard room-bound today (`enemyAi.ts:266`) — not about coop difficulty scaling
- [ ] create a repo list on the repo tab similiar to the wad selector, instead of the current buttons. with way more repos
- [ ] fields: same as wad, plus approx. size and difficulty