A dev-only tool for automated balance review: three scripted bot profiles (Casual/Gamer/Pro) play the bundled demo-campaign/ across all three difficulties, and their aggregated combat/economy/navigation stats get written to balancing_telemetry.json (gitignored) for a human — or an LLM balance-review pass — to spot HP-curve/drop-rate/pacing problems without replaying the whole campaign nine times by hand. Not CI-wired. Requires a locally running dev server (npm run dev, default http://localhost:5173).
The bot itself is a shared Bot class (scripts/lib/bot.mjs) that both this script and scripts/generate-default-highscore.mjs drive — a virtual clock, window.__codeensteinTestHooks polling, real KeyboardEvents dispatched at the canvas, BFS route planning done entirely in Node before any browser launches. See bot.mjs's own doc comment for the low-level harness rationale, and Shared bot library (scripts/lib/) below for how the pieces fit together.
| Command | What it does |
|---|---|
npm run balancing:telemetry |
Full 9-combo run (Casual/Gamer/Pro × easy/normal/hard), 3 qualifying runs each, writes balancing_telemetry.json. Slow (up to 9 × unbounded attempts × up to 17 levels) — this is the "generate real data" entry point. |
npm run balancing:watch |
Opens one real, visible Chromium window per profile (Casual → Gamer → Pro by default), plays one full campaign attempt at watchable real-time speed, prints a summary, waits for Enter before the next profile. scripts/watch-bot-sessions.mjs; reuses the same profile definitions and per-attempt driving logic (playRun), not a separate bot. npm run balancing:watch -- Gamer Pro to pick a subset/order; CODEENSTEIN_WATCH_DIFFICULTY=hard to change difficulty. |
npm run balancing:scan |
The permanent automated bot-behavior regression check (distinct from balance-data review) — see Anomaly scanning below. Run this before declaring any navigation/combat change to the bot script fixed. |
npm run balancing:campaign |
Large-scale, resumable data-collection orchestrator — see Large-scale campaigns below. Not the same as balancing:telemetry's fixed 3-qualifying-run sweep; this repeatedly spawns it to build up a much bigger sample, keeping every batch as its own file. |
A run only "counts" for balancing:telemetry's aggregate once it clears level 4 (proves it survived the unarmed early game) — a run that dies on level 1-3 is discarded entirely; a qualifying run keeps all its levels' data, 1–3 included. Both the qualifying level (QUALIFY_LEVEL_INDEX) and the qualifying-run target (REQUIRED_QUALIFYING_RUNS, default 3, overridable via CODEENSTEIN_TELEMETRY_QUALIFYING_TARGET) live in run-balancing-telemetry.mjs.
PROFILES (Casual/Gamer/Pro) lives in scripts/lib/profiles.mjs and is re-exported from run-balancing-telemetry.mjs, so every existing importer is unchanged. It was moved there so unit tests can import the real profiles without pulling in Playwright — combatPolicy.test.mjs previously hand-copied Gamer's values into a local fixture, which silently drifted out of sync and no test could detect.
Read the ladder as an ordering, not a list of numbers. profiles.test.mjs pins what several consumers silently depend on: the exact key order (curateMixedProfiles reads tierNames[0] as the weakest tier), strict monotonicity per knob, and a complete ranged fallback chain per tier — the historical Pro-missing-shotgun bug made Pro slower to qualify than Casual, which is what a broken ladder looks like from the outside.
| knob | Casual → Pro | what it controls |
|---|---|---|
fireAngleEps |
0.08 → 0.03 rad | aim tolerance before the bot will pull the trigger |
fireCooldownMs |
220 → 120 ms | minimum gap between semi-auto shots. Dual-role and worth knowing: it also feeds secPerShot in scoreRangedWeapon, so a slow trigger also makes the bot judge semi-autos as worse. Don't draw per-tier weapon-balance conclusions without accounting for that |
rotSpeedMultiplier |
2 → 5 | bot-only turn-speed override (engine.ts), approximating mouse turn speed per tier — real pointer-lock mouse-look rejects outright under Playwright |
healthDetourThreshold |
0.75 → 0.25 | how early the bot breaks off to find health |
ammoThrift |
1.6 → 0.2 | willingness to burn a finite reserve to shorten a fight |
selfHarmAversion |
2.2 → 0.5 | caution with self-splash weapons; only reachable in the ghidra branch |
weaponPriority |
— | membership is a hard filter, order is only a tiebreak. The scoring loop iterates this list, so an absent weapon is never considered at all — which is why Casual omitting ghidra is a real behavioural tier rather than a preference |
engageRadius |
9.5, identical | deliberately not a tier: "low aggression" must never mean skipping a fight |
A profile may also carry a tuning object, deep-merged over DEFAULT_TUNING in the Bot constructor and under any explicit opts.tuning — so CODEENSTEIN_TELEMETRY_TUNING still wins and single-variable A/Bs keep working. That is how a tier can own what would otherwise be a global constant, i.e. express competence rather than only pace.
Difficulty (easy/normal/hard) is wired through localStorage["codeenstein-difficulty"], same as a real player's setting. Keep the two axes distinct: difficulty changes the world (enemy HP, damage, ammo drop rate, enemy aim spread — src/difficulty.ts), a profile changes the player.
node scripts/report-profile-separation.mjs <dir-or-file> [difficulty]Takes one telemetry JSON containing all three profiles, or a directory of per-profile captures, and exits non-zero when the tiers aren't cleanly ordered. abReport.mjs compares two sides of one change; this compares the three profiles within one capture, which is the question that decides whether per-tier balance conclusions mean anything.
It bounds the smallest adjacent step, not the ladder's ends, and that distinction is the point: Casual is cleanly separated from both other tiers, so an ends-based ratio looks healthy while Gamer↔Pro — the step that actually flipped run-to-run across four n=5 scans — is unreadable. Bars are calibrated against a measured n=40 baseline so the axes that already work cannot be quietly traded away.
Measured 2026-08-01, before any retune: ttkNormal and levelTimeSec were correctly ordered, but both damage-avoidance axes were inverted (enemyAccuracy 0.499/0.556/0.531 and ranged damage per second of exposure 0.354/0.430/0.439 — Pro the most hittable). The tiers differed in pace, not competence.
src/engine/defaultHighscore.ts is generated by playing the campaign with these profiles, so any retune leaves it describing a bot that no longer exists — and nothing used to detect that. A segment's astHash covers the parsed source and balanceHash covers the enemy roster; neither knows anything about bot tuning, so every replay stays valid while the shipped board silently misrepresents the bot.
The generator now bakes profilesHash() into the file as PROFILES_HASH, and profiles.test.mjs recomputes and compares — so this is a failing vitest run rather than an invisible drift. Field order inside a profile is ignored (it changes no behaviour); profile order is not (several consumers read it positionally).
Regenerate the board last. A retune is only one of several things that invalidates it — anything in SIMULATION_BALANCE does too, and that record now includes ACID_DECAY_SECONDS. The 2026-08-02 layout rework learned this the expensive way: the board was regenerated mid-branch, a gameplay constant then joined SIMULATION_BALANCE, and the whole two-hour run had to be repeated. Land every simulation change first, then generate once.
The board is a poor instrument for judging a code change, and always has been. Each entry is a maximum over qualifying runs, so its level count swings on luck: Gamer has come out 14/11/14/11/11/17 levels across six regenerations with no cause in the code. Read balancing:scan's anomaly A/B for that question instead — it compares two builds over many attempts, which is what the 3-combo protocol above exists to make honest.
The bot-behavior logic (navigation, combat, hazard/mine handling, loot detours — ~1450 lines) lives in scripts/lib/bot.mjs's Bot class, not duplicated per script. Both run-balancing-telemetry.mjs and generate-default-highscore.mjs construct one Bot per attempt (new Bot(page, profile, opts)), call bot.startLevel(map) per level, then drive it via bot.tick()/bot.driveLegs()/bot.driveToward()/etc.
-
Config is explicit constructor
opts, not ambient module state.opts.realtime/opts.stepMs(headed-vs-headless timing),opts.recordStepMs(only the highscore generator sets this — see its own module doc comment for why replay recording needs a finer step granularity than bot decision-making),opts.logger(a{debugNav, wpDebug, driftDebug, trace}bag of no-op-by-default callbacks, replacing scatteredprocess.env.CODEENSTEIN_WPDEBUG-style checks), andopts.tuning(deep-merged overDEFAULT_TUNING, the ~40 movement/combat constants both scripts used to duplicate).run-balancing-telemetry.mjsstill resolves its own module-levelHEADED/DEBUG_NAV/etc. consts fromprocess.envonce at import (this is what preserveswatch-bot-sessions.mjs's "setprocess.env.CODEENSTEIN_TELEMETRY_HEADEDbefore dynamically importing" trick) and forwards them into theBotconstructor per attempt. -
scripts/lib/qualifyLoop.mjs'srunQualifyLoop()is the generic "retry attempts in concurrent batches until N qualify (or a cap is hit)" loop bothrun-balancing-telemetry.mjs'srunComboandgenerate-default-highscore.mjs's per-profile driver are thin wrappers around — the qualifying predicate, attempt function, and concurrency are all caller-supplied. -
scripts/lib/virtualClock.mjs'sinstallVirtualClock()is the one virtual-clock installer both scripts import, instead of each keeping its own byte-for-byte-identical copy. -
scripts/lib/combatPolicy.mjsholds the decision core.decide(world, memory, config)returns an intent — which keys to hold and for how long, whether to fire, which weapon to switch to — andbot.mjsis the I/O half that dispatches it. Everything it owns (DEFAULT_TUNING,angleDelta,pickThreat,pickRangedWeapon,findDisarmableMine/findDangerousMine, the burst helpers, the weapon indices) is re-exported frombot.mjs, so no consumer had to change an import; new code should import fromcombatPolicy.mjsdirectly.detectAnomalies/detectHeldKeyNoMovementstay inbot.mjs— they analyse recorded traces rather than making decisions.- Two reasons for the split. It makes the decision logic testable (
combatPolicy.test.mjs, 100+ assertions — before this, nothing inscripts/lib/had any). And it is shaped so it can later be lifted intosrc/engine/combatPolicy.tsas the basis of a real in-game deathmatch opponent:src/cannot import fromscripts/, so the move has to be a copy, and a copy is only mechanical if the file never acquires a dependencysrc/engine/couldn't satisfy. Corrected 2026-08-18: that reasoning is sound but was read too widely — the other direction works, and since that datecombatPolicy.mjsimportssrc/engine/mapPredicates.tsandsrc/map/types.tsdirectly under plain Node. Every such import makes the eventual lift smaller, not harder. The rule it must keep is inmapPredicates.ts: no value import without an explicit.tsextension, and no DOM. The module doc comment lists the rules that keep that true (no page, no async, no Node builtins, all tuning injected, noMath.random/Date.now, total sorts). - Injecting tuning also fixed a latent bug:
pickThreat,pickRangedWeapon,rocketAimUnsafe,findDisarmableMineandfindDangerousMineused to readDEFAULT_TUNINGdirectly, so aBot's ownopts.tuningoverride silently never reached them.
- Two reasons for the split. It makes the decision logic testable (
-
The bot has full WASD, and all eight directions are full speed.
moveForwardtranslates along(dirX, dirY)andstrafealong(-dirY, dirX)(player.ts), both scaled by the samestep;diagonalScale(1/√2) is applied per axis when both are held, so two perpendicular components have magnitude exactlystep. Reversing is the same speed as advancing (forwardSignis signed) and sprint applies throughout. So turning is never required in order to move somewhere — worst-case direction error picking the nearest octant is 22.5°, i.e. 92% of the step lands where wanted, against 0% while standing still to turn. This was not written down anywhere and the bot went a long time without using it: it never emittedKeySat all, and stood still at every route corner (23.5% of all decisions were turn-only).movementKeysFor/movementVectorForincombatPolicy.mjsencode it. Note that any safety check must then scan along the movement vector, not the facing one. -
Evasion: the bot does not stand still to shoot. The ranged-fire branch used to hold no keys at all for a whole decision, which made
aiEffectivenessDanger.enemyAccuracya measure of how easy it is to hit a stationary target rather than of how dangerous enemies are. It now sidesteps while firing, and whengetProjectilesreports a bolt actually inbound it steps off that bolt's flight line specifically and sprints while doing so. Measured marginal effect of the directed dodge over blind sidestepping:dmg.enemyRangedPerSec-25.4/-18.5/-7.9%,enemyAccuracy-0.115/-0.095/-0.085, qualifyRate +0/+55/+25pp.- Three constraints, each load-bearing. Lateral only, never alongside
KeyW—engine.ts'sdiagonalScalecuts the forward component 29% when both axes are active, which is the mechanism behind the recorded 0%→72% regression. Sprint only while genuinely dodging, never for the blind sidestep, or the dance becomes drift. Acid, spikes and walls all block, with the other side tried first — a strafe is optional movement, so unlike a committed route leg it is never worth damage. - Only the fire branch strafes, and that is deliberate. Extending it to the re-aim branch was measured and reverted: it cost Gamer 30pp of qualify rate and doubled its stuck count for no accuracy gain, because that branch exists to converge on a firing angle and lateral movement perturbs the very quantity it is nulling out. Roughly 43% of combat decisions reach the fire branch; that ceiling is the price of not fighting the aim loop.
- Three constraints, each load-bearing. Lateral only, never alongside
-
Per-key hold durations and
dispatchSegment. An intent'sholdsis aMapof key → ms, andapplyActionturns it into the sequence of dispatch phases that realises it: a key whose hold ends early simply stops appearing in later phases, and the held-key diff releases it for free. As landed, every branch gave all its keys the same duration, sosegmentsForalways yielded exactly one phase and dispatch was byte-identical to the pre-split single-turnBurstcall — verified by a differential harness over 21,600 dispatches (3 profiles × 3 step sizes × 400 seeds × 6 consecutive decisions), comparing keys, fire/melee/weapon, duration, mutated memory, trace, and the semi-auto fire clock. Letting keys carry different durations is what fixes the bot standing still while it shoots — done 2026-08-09: the combat branch now gives the turn key its ownturnHoldMswhile movement keeps the widened decision (turnSplitIntent), which took nav-warns from 4,327 to 0 and shots/run from 29 to 35 with damage taken unchanged.minPhaseMsis a floor, and it is the reason multiplayer can't be hurt by this mechanism: a phase shorter than the lockstep input delay never lands before the next is issued, soMultiplayerBotsets it toMIN_DECISION_MSand any decision that would split below it collapses back to a single phase — exactly the behaviour it already had.dispatchSegment(notapplyAction) is now the override point for a non-Playwright control surface.MultiplayerBotused to reimplement all ofapplyAction; the timing bugs its doc comment catalogues all trace to those two copies drifting apart.
-
scripts/lib/pathfind.mjsandscripts/lib/routePlanner.mjsare unaffected by any of this — they were already clean, reusable, stateless modulesbot.mjsimports. -
⚠️ The bot keeps three hand-maintained copies ofisWall(), and adding aTilevalue silently breaks all three.player.ts'sisWall()is the engine's single source of truth for solidity, butscripts/lib/is plain.mjsthat cannot import it, so the tile sets are duplicated as literals:Module Constant Contents pathfind.mjsBLOCKED_TILES{1, 3, 6, 7, 8}routePlanner.mjsHARD_BLOCK_TILES{1, 3, 6, 7, BRANCH_DOOR_TILE, TELEPORTER_TILE}combatPolicy.mjsSTRAFE_BLOCKED_TILES{1, 3, 6, 7, 8}Note they are not identical —
routePlanner's also blocks teleporters, deliberately — so this cannot be collapsed to one constant without thought. The failure mode is what makes it dangerous: aSetlookup for an unknown tile value returnsfalse, so a newly added solid tile is treated as ordinary walkable floor by every bot script. No error, no crash — the bot simply plans routes through it and wedges against a wall it believes is open. Nothing catches this:scripts/**is excluded from the coverage denominator, the files are not type-checked, andbalancing:scanonly reports the symptom (a stall) with no hint at the cause. This shipped once already, onBRANCH_DOOR_TILE = 8.When adding a
Tilevalue tosrc/map/types.ts, grepscripts/lib/for the neighbouring tile numbers and decide each set explicitly. The engine-side touchpoints (isWall, the renderer, the automap/minimap colouring) are at least type-adjacent; these three are not. -
scripts/lib/abReport.mjsholds the pure baseline-vs-candidate comparison helpers behindscripts/report-balancing-ab.mjs— see Matched-scale verification. It has a colocatedabReport.test.mjs:scripts/**is excluded from thesrc/coverage denominator but is still executed byvitest run, so a test placed here does run in CI, exactly likescripts/multiplayer-server.test.mjs. That is the only automated coverage anything inscripts/lib/has, and it's worth extending as more of this library becomes pure functions.
All scoping/debug flags are read once at module load, so they must be set in the same process invocation (not exported separately beforehand if using a subshell that re-execs).
| Var | Effect |
|---|---|
CODEENSTEIN_TELEMETRY_PROFILE |
Restrict to one profile (Casual/Gamer/Pro). |
CODEENSTEIN_TELEMETRY_DIFFICULTY |
Restrict to one difficulty. |
CODEENSTEIN_TELEMETRY_LEVEL_LIMIT |
Cap how many campaign levels get planned/played. |
CODEENSTEIN_TELEMETRY_ATTEMPT_CAP |
Cap attempts per combo (default unbounded — retries until 3 qualify). Use for any scoped/smoke run; never rely on the unbounded default finishing quickly. |
CODEENSTEIN_TELEMETRY_CONCURRENCY |
Attempts run concurrently within one combo (separate browser contexts sharing one Chromium process; default 12). Matters for verification, not just speed — see Matched-scale verification. |
CODEENSTEIN_TELEMETRY_VERBOSE |
Per-attempt death detail (fatal=/kills=/dmgBySource=/engaged-enemy TTKs/weapon tallies). |
CODEENSTEIN_TELEMETRY_DEBUG_NAV |
Permanent tick-by-tick nav/combat trace ([nav] pos=... dir=... threat=... -> moveKeys=...). Not a temporary debug flag — kept on purpose for whatever "why is the bot doing that" question comes up next. |
CODEENSTEIN_TELEMETRY_ANOMALY_SCAN |
Enables the stall/health-drain-frozen detector — see below. |
CODEENSTEIN_TELEMETRY_NAV_DIAG |
Extra per-decision trace bookkeeping (superset used alongside anomaly scan trace recording). |
CODEENSTEIN_TELEMETRY_TRACE_DUMP |
Prints the raw per-decision rows of each level's longest oscillation run — position, bearing error and distance to the nav target, keys held, burst, threat/mine distance, branch. Implies the trace. Run with CONCURRENCY=1: concurrent attempts interleave their rows into a misleadingly coherent-looking mess. Added because aggregates over this detector's own findings produced two wrong diagnoses in a row; reading one run end to end settled it in minutes. |
CODEENSTEIN_BOT_TIMING |
Attributes the run's real wall clock across decide(), the engine's own frames, and the Node<->browser round trip — see Where the harness's wall clock goes. Prints one [phase-timing] block and adds meta.phaseTiming. Off by default and reads no clock when off. |
CODEENSTEIN_TELEMETRY_EXTRA_QUERY |
Appended verbatim to the page URL. Exists for &ablate=floor,effects,viewmodel,hud (a real, shipped, deliberately un-DEV-gated engine switch) — see the same section. Never ablate sprites, walls or shade: the sprites branch is what sets this.target from findTargetUnderCrosshair and walls fills the z-buffer it reads, so ablating either stops the bot being able to shoot. |
CODEENSTEIN_TELEMETRY_HEADED |
Real, visible browser + real wall-clock timing instead of the virtual clock. See Headed vs. headless. |
CODEENSTEIN_CONSOLE_FORWARD |
Forwards the browser's own console output to Node ([console] ... lines) — the engine already logs key pickups/door unlocks; often more reliable ground truth than bot-side telemetry when a freeze's cause is ambiguous. |
CODEENSTEIN_WPDEBUG |
Per-waypoint drive-loop trace ([wpdebug] leg-walk wp=... -> result=...). |
CODEENSTEIN_DRIFTDEBUG |
Traces driveTowardWithReplan's off-route drift/re-plan decisions. |
CODEENSTEIN_DEV_URL |
Override the dev server URL. The built-in default is not uniform — http://localhost:5173 here and for the multiplayer verifiers, but http://localhost:5183 for verify:replay and verify:campaign:playthrough, deliberately, so those don't collide with a manual dev session. Testing has the full per-script table and is the authority. |
CODEENSTEIN_TELEMETRY_DEV_PORT |
Port for the dev server this tool starts itself when CODEENSTEIN_DEV_URL is unset (default 5199). Deliberately not 5173, so a run never collides with a dev server you started yourself. run-perf-benchmark.mjs uses CODEENSTEIN_PERF_PORT for the same purpose, with the same default. |
CODEENSTEIN_TELEMETRY_QUALIFYING_TARGET |
Override REQUIRED_QUALIFYING_RUNS (default 3) — how many qualifying runs a combo needs before its retry loop stops. Used by balancing:campaign to set a small per-invocation batch size. Set it to 999 for any A/B, so the loop runs to ATTEMPT_CAP instead of exiting the moment enough runs qualify — that early exit, not concurrency, is what controls the failure-sample denominator. See Matched-scale verification. |
CODEENSTEIN_TELEMETRY_OUTPUT_FILE |
Override the output path (default balancing_telemetry.json at repo root). Used by balancing:campaign so concurrent invocations each write to their own file instead of racing to overwrite the same one. |
Not a telemetry tool, but it drives the same Bot against the same dev server and is the other job that can run unattended for a long time — so its three knobs belong next to the ones above rather than nowhere.
| Var | Effect |
|---|---|
CODEENSTEIN_HIGHSCORE_QUALIFYING_RUNS |
Qualifying runs collected per profile before the highest-scoring one is baked in (default 3). |
CODEENSTEIN_HIGHSCORE_ATTEMPT_CAP |
Cap attempts per profile (default unbounded) — the only way to bound an otherwise open-ended run. |
CODEENSTEIN_HIGHSCORE_CONCURRENCY |
Attempts run concurrently per profile (default 4). |
The lane-parallel capture orchestrator. Every one of these was undocumented until a 2026-08-12 doc audit; the defaults below are read straight from the constant declarations at the top of the script.
| Var | Default | Effect |
|---|---|---|
CODEENSTEIN_CAPTURE_OUT |
balancing_capture |
Output directory, relative to the repo root. All balancing_* paths are gitignored. |
CODEENSTEIN_CAPTURE_PROFILES |
Casual,Gamer,Pro |
Comma-separated profile list. |
CODEENSTEIN_CAPTURE_DIFFICULTIES |
normal,hard |
Comma-separated difficulty list. Note easy is not in the default sweep. |
CODEENSTEIN_CAPTURE_ATTEMPTS |
60 |
Fixed attempts per cell — this is the denominator, not a target. |
CODEENSTEIN_CAPTURE_CHUNK |
20 |
Attempts claimed per lane per scheduling decision. |
CODEENSTEIN_CAPTURE_MIN_CHUNK |
5 |
Floor on that chunk as a cell drains. |
CODEENSTEIN_CAPTURE_TARGET_CHUNK_MIN |
45 |
Target minutes of work per claim, which is what the chunk size is solved for. |
CODEENSTEIN_CAPTURE_CONCURRENCY |
10 |
Attempts in flight per lane. |
CODEENSTEIN_CAPTURE_MAX_INVOCATIONS |
8 |
Hard cap on bot invocations, as a runaway backstop. |
CODEENSTEIN_CAPTURE_WATCHDOG_MS |
7_800_000 (130 min) |
Per-invocation watchdog. |
CODEENSTEIN_CAPTURE_LEVEL_LIMIT |
unset | Cap campaign levels per attempt. |
CODEENSTEIN_CAPTURE_LOCAL_ONLY |
unset ("1" to enable) |
Run without SSH lane hosts. Without it, a run that has hosts configured but reaches none is an error rather than a silent local-only fallback. |
| Var | Default | Effect |
|---|---|---|
CODEENSTEIN_CAMPAIGN_TARGET |
50 |
Qualifying runs wanted per combo. |
CODEENSTEIN_CAMPAIGN_BATCH_SIZE |
5 |
Qualifying runs per bot invocation. |
CODEENSTEIN_CAMPAIGN_ATTEMPT_CAP |
80 |
Attempts per invocation. |
CODEENSTEIN_CAMPAIGN_CONCURRENCY |
8 |
Attempts in flight per lane. |
CODEENSTEIN_CAMPAIGN_LANES |
2 |
Lanes. |
CODEENSTEIN_CAMPAIGN_MAX_INVOCATIONS |
6 |
Runaway backstop. |
CODEENSTEIN_CAMPAIGN_WATCHDOG_MS |
5_400_000 (90 min) |
Per-invocation watchdog. |
CODEENSTEIN_CAMPAIGN_PROFILE |
unset | Restrict the matrix to one profile. |
CODEENSTEIN_CAMPAIGN_DIFFICULTY |
unset | Restrict the matrix to one difficulty. |
run-balancing-campaign-multiplayer.mjs takes the same knobs under a CODEENSTEIN_MP_CAMPAIGN_ prefix, with different defaults. Spelled out in full rather than as a prefix note, so grepping for one of these actually finds it:
| Var | Default | vs. single-player |
|---|---|---|
CODEENSTEIN_MP_CAMPAIGN_TARGET |
10 |
50 |
CODEENSTEIN_MP_CAMPAIGN_BATCH_SIZE |
2 |
5 |
CODEENSTEIN_MP_CAMPAIGN_ATTEMPT_CAP |
30 |
80 |
CODEENSTEIN_MP_CAMPAIGN_CONCURRENCY |
1 |
8 |
CODEENSTEIN_MP_CAMPAIGN_LANES |
1 |
2 |
CODEENSTEIN_MP_CAMPAIGN_MAX_INVOCATIONS |
3 |
6 |
CODEENSTEIN_MP_CAMPAIGN_WATCHDOG_MS |
14_400_000 (4 h) |
5_400_000 (90 min) |
CODEENSTEIN_MP_CAMPAIGN_PROFILE |
unset | same |
CODEENSTEIN_MP_CAMPAIGN_DIFFICULTY |
unset | same |
CODEENSTEIN_MP_CAMPAIGN_PLAYER_COUNTS |
unset | no equivalent — comma-separated list |
CODEENSTEIN_MP_CAMPAIGN_MIN_LEVEL |
4 |
no equivalent — read by verify-multiplayer-campaign.mjs, not the campaign runner |
| Var | Default | Effect |
|---|---|---|
CODEENSTEIN_REPLAY_SPEED |
1 |
Playback speed multiplier for verify:replay. |
CODEENSTEIN_REPLAY_TRACE |
unset ("1") |
Per-frame replay trace. |
CODEENSTEIN_REPLAY_LEVEL_LIMIT / _CONCURRENCY / _ENTRIES |
unset / 2 / 0 |
Scope, parallelism and which board entries to replay. |
CODEENSTEIN_PERF_HEADLESS |
unset (truthy) | Run perf:bench headless. Measurements are not comparable across this flag. |
CODEENSTEIN_PERF_HAR_RECORD |
unset (truthy) | Record a HAR alongside the benchmark. |
CODEENSTEIN_TRANSITION_NAV_DEADLINE_MS |
540_000 (9 min) |
Wall-clock budget for the host-navigation phase of verify:multiplayer-transition. Bounded in wall clock precisely because every other budget on that path is bounded in decisions, which at a real 300-400 ms decision window admits a legal worst case of over two hours. |
CODEENSTEIN_TRANSITION_CLEAR_EXIT_ROOM |
unset ("1") |
Test hook that clears the exit room's enemies. |
CODEENSTEIN_VITE_NO_WATCH |
unset (truthy) | Disables Vite's file watcher (vite.config.ts). For harnesses that hold the tree open for a long time. |
CODEENSTEIN_VERIFY_BROWSER |
chromium |
Browser for the verify:* scripts. |
CODEENSTEIN_WADS_STRICT |
unset | Equivalent to fetch-online-wads.mjs --strict. |
CODEENSTEIN_EXTRA_WADS |
unset | Colon-separated extra WAD paths for report:wad-stylesets. |
CODEENSTEIN_WATCH_DIFFICULTY |
normal |
Difficulty for balancing:watch. |
CODEENSTEIN_MULTIPLAYER_DEBUG_ICE |
unset ("1") |
ICE-candidate trace in the multiplayer verify scripts. |
POC_SEED / POC_ITERATIONS / POC_SAMPLE_EVERY |
0xc0ffee / 500_000 / 500 |
Cross-browser determinism proof. |
The signaling server's own ~25 CODEENSTEIN_MULTIPLAYER_* variables are deliberately not listed here. node scripts/multiplayer-server.mjs --help prints every one of them with its currently-effective value, generated at invocation time — so unlike every table above, it cannot drift. Read that instead, and see Multiplayer Server Deployment for the ones that matter in production.
Two variables validate their input, both in run-balancing-telemetry.mjs: CODEENSTEIN_TELEMETRY_SEED range-checks 0..0xffffffff and exits non-zero naming the bad value, and CODEENSTEIN_TELEMETRY_TUNING rejects both malformed JSON and a non-object. Both guard a value whose corruption would be invisible — a bad seed or a silently-ignored tuning override produces a run that looks fine and measures the wrong thing.
Every other numeric knob above is a bare Number(process.env.X ?? default), so a typo yields NaN and propagates silently — CODEENSTEIN_MULTIPLAYER_PORT=abc does not fail, it just produces a NaN port. Check a value took effect rather than assuming a bad one would have been rejected.
Measured 2026-08-18 with CODEENSTEIN_BOT_TIMING=1, because the notes backlog carried "decouple bot decisions from engine, for more performance" and nothing had ever measured what a bot decision costs. perfDebug.ts has no bot phase, and it cannot acquire one under this harness: installVirtualClock replaces performance.now, so every in-page phase timing reads 0. That is why the installer now stashes window.__realNow before patching — it is the only real clock left inside the page.
Three numbers, all real time, all per decision:
decide() |
engine frame | CDP transport | |
|---|---|---|---|
| what it is | the bot's actual decision logic | simulate(dt) + render(), inside __pumpVirtualTime |
round trip + serialisation across page.evaluate |
| concurrency 1 | 0.024ms (0.3%) | 5.12ms (63%) | 2.88ms (36%) |
| concurrency 12 | 0.028ms (0.2%) | 8.40ms (46%) | 9.91ms (54%) |
The item's premise, read literally, is refuted. The bot's decision costs 0.2-0.3% of the accounted wall clock. There is no version of "decouple bot decisions for more performance" that pays, because the decisions are free. What costs is everything around them.
The engine renders a frame per decision that nobody ever looks at. advance() is unconditionally simulate(dt) then render(), and the harness pumps one rAF per decision — so a full raycast pass (floor cast, walls, sprites, effects, viewmodel, HUD) runs ~30,000 times per attempt into a canvas no one reads. ?ablate=floor,effects,viewmodel,hud (resolveAblations, engine.ts, deliberately not DEV-gated so the switches exist in the exact build being measured) removes about three quarters of it: engine time 5.12ms -> 1.28ms at concurrency 1, 8.40ms -> 1.97ms at 12.
And it is gameplay-neutral — verified, not assumed. At a pinned CODEENSTEIN_TELEMETRY_SEED, the ablated and control arms produced byte-identical telemetry payloads — 18,773 bytes over 4 levels x 3 attempts at concurrency 1, and 81,080 bytes over 8 levels x 12 attempts at concurrency 12 — with identical decision counts (92,724 both arms) on the larger pair. That check is the point: an ablation that changed a single hit would invalidate every number the run produced, and "it only skips drawing" is exactly the kind of claim that is obviously true until it isn't.
Never ablate sprites, walls or shade. The sprites branch is what sets this.target from findTargetUnderCrosshair, and its else sets this.target = null; walls fills the z-buffer that targeting reads. Ablating either stops the bot being able to shoot at all — a silent, total behavioural change wearing the costume of a rendering flag.
End to end it is worth 11.5%, not the 4.2x the engine numbers suggest. Three interleaved reps, seed pinned, identical decision counts (92,724) in every run: control 196.7s (199/192/199), ablated 174.0s (176/178/168). Interleaved on purpose — the governor here is schedutil and this repo has a recorded case of DVFS flipping the sign of a light-duty comparison. Transport grew as the engine got cheaper (11.06 -> 13.17ms/decision): the bottleneck moved onto the shared CDP connection rather than disappearing.
SUPERSEDED 2026-08-19 — the saturated thing was the Node event loop, and it is fixable without touching the bot. Everything above was measured through a single chromium.launch() with N contexts, driven by a single Node process, so it could not distinguish "a CDP round trip is expensive" from "twelve attempts are queueing on one pipe". Three arms of 24 seed-pinned attempts, byte-identically the same work on each (205,128 decisions):
| arm | wall | transport/decision | CPU idle |
|---|---|---|---|
| 1 process x concurrency 12 | 340.5s | 11.90ms | 53.5% |
| 1 process x concurrency 12, 4 browsers | 363.0s | 12.76ms | 43.4% |
| 4 processes x concurrency 3 | 213.0s | 7.14ms | 6.9% |
Four browsers in one process change nothing, so it is neither the CDP socket nor the browser process — it is the one event loop dispatching ~1,200 round trips a second. An A/B/A drift control agreed to within one second. This is why run-balancing-capture.mjs now runs several local lanes (CODEENSTEIN_CAPTURE_LOCAL_LANES, defaulted from core count): each lane is its own Node process, and they share the machine's existing concurrency budget rather than one taking all of it.
Note the scope: only the capture shards. A bare balancing:telemetry/balancing:scan is still one process and still has this ceiling — which is fine for a smoke test and is why the numbers above still describe it accurately.
And concurrency cannot buy past it — the transport is saturated. Swept ablated at 24 attempts per point (n=1 each, so read the spread, not the ranking):
| concurrency | wall | attempts/min | engine ms/dec | transport ms/dec |
|---|---|---|---|---|
| 8 | 319s | 4.51 | 1.96 | 7.94 |
| 12 | 331s | 4.34 | 2.00 | 12.93 |
| 16 | 328s | 4.38 | 2.02 | 14.38 |
| 20 | 341s | 4.22 | 2.02 | 19.47 |
Throughput varies 6.4% across the whole sweep, which is the same size as the run-to-run spread measured on the A/B above (168-178s on identical config, 6%). So this is a null: no readable difference between 8 and 20. Meanwhile per-decision transport rises 2.5x — each round trip gets steadily slower and exactly as many more are in flight, which is what saturation looks like. Do not re-tune CODEENSTEIN_TELEMETRY_CONCURRENCY hoping for throughput; 12 is fine and so is 8.
That is the useful part of the null: the harness is at the ceiling of this architecture, and no setting escapes it. The only remaining lever is issuing fewer round trips, not faster or more concurrent ones.
What this leaves for the in-page-decision-loop idea. Transport is the half an in-page loop would delete, and at production concurrency (12) it is already the larger half — and it grows once the engine gets cheaper, because the bottleneck moves onto the shared CDP connection rather than disappearing. So the decoupling idea survives, but for the transport, not for the decisions, and it should be judged against the ablation as the cheaper alternative that ships today.
report-balancing-ab.mjs's loadSide merges a directory's *.json with { ...merged.profiles[name], ...profile } — a spread at the difficulty level. A capture writes one file per chunk (Casual-hard-001.json, -002, ...), each carrying the same profiles.Casual.hard key, so the last chunk read silently replaces every earlier one rather than pooling them. There is no error and the output looks entirely normal; it is simply computed from a fraction of the data.
It is built for the 4-combo telemetry protocol, where one file per side is the whole side. For a capture, read report-aim-error.mjs (which does pool across capture dirs) and the event log, and run guards from their own small telemetry run.
CODEENSTEIN_TELEMETRY_ANOMALY_SCAN=1 makes tick() record a per-decision trace (position, health, threat/mine distance, branch, waitingOnSpike) and scan it after every level for two patterns:
stall— position anchored within 0.05 tiles for 20+ ticks (excluding legitimate spike-cycle waits).healthDrainFrozen— position anchored for 2+ ticks while health is also dropping.
Findings print as [anomaly] <profile>/<difficulty> level N: <type> (...). The balancing:scan npm script runs all three profiles, normal difficulty, 8 levels, 5 attempts each. Run this (or a scoped subset via the env vars above) before reporting any bot navigation/combat fix as verified — a "few manual traces looked fine" is not sufficient given this bot's history of freezes that only reproduce after hundreds of ticks or under specific map geometry.
This scanner is headless-only. It cannot see bugs that only manifest under real per-frame timing — see the next section.
turnBurstMs (and its movement counterpart moveBurstMs) compute the exact millisecond hold-duration needed to turn/move by a given amount, on the assumption that holding a key for N ms produces exactly N ms worth of rotation/movement. That assumption is only exactly true in headless mode.
- Headless (
CODEENSTEIN_TELEMETRY_HEADEDunset):window.__pumpVirtualTimeadvances the engine's virtual clock by precisely the requested duration in one pumpedrequestAnimationFramecallback. Arbitrarily fine convergence (down to a fraction of a radian) is genuinely achievable. - Headed (
CODEENSTEIN_TELEMETRY_HEADED=1, used bynpm run balancing:watch): the engine only actually rotates/moves once per real rendered frame (~16.7ms at 60fps). Apage.waitForTimeoutwait shorter than roughly one real frame does not reliably produce a proportionally small rotation — real frame/timer granularity dominates. Any convergence epsilon tighter than roughlyENGINE_ROT_SPEED * rotSpeedMultiplier / realFpsis structurally unreachable in headed mode.
Concretely: a fine-alignment epsilon like MINE_REALIGN_EPS (0.01 rad) converges in 1–2 ticks headless, but in headed mode produced dir bouncing between two fixed values forever — position frozen, chasing a target the real frame rate could never resolve. balancing:scan (headless) showed nothing wrong at all; the bug was only visible while actually watching.
If you're chasing a bug reported from watching (balancing:watch) that a headless balancing:scan doesn't flag:
-
Don't assume it's a log artifact or unreproducible — reproduce it directly. Run
scripts/run-balancing-telemetry.mjswithCODEENSTEIN_TELEMETRY_HEADED=1(this bypasseswatch-bot-sessions.mjs's interactive per-profile Enter-press wrapper entirely, so it's scriptable/backgroundable like any other run), plusCODEENSTEIN_TELEMETRY_DEBUG_NAV=1and tight_PROFILE/_LEVEL_LIMIT/_ATTEMPT_CAP=1/_CONCURRENCY=1scoping. Requires a real display (DISPLAYset, e.g. Xvfb). -
To find a genuine freeze (not just ordinary tick-to-tick movement) in the resulting trace, scan for runs of N+ consecutive
[nav]lines with byte-identicalpos=(x,y):awk ' /^\[nav\]/ { match($0, /pos=\(([0-9.-]+),([0-9.-]+)\)/, p); key = p[1] "," p[2]; if (key == prevKey) { run++; } else { if (run >= 15) print "run of " run " ticks frozen at " prevKey " ending line " NR-1; run = 1; prevKey = key; } } ' trace.log
-
If a fix candidate widens a convergence epsilon ("accept close enough" instead of chasing precision), check what the branch actually does once "satisfied" — if it has no fallback action (e.g. mine-targeting deliberately never adds movement, to avoid walking into blast range), the fix can convert "stuck but still trying" into "immediately idle until an unrelated timeout," which is often strictly worse and can have knock-on effects (an abandoned mine stays live and un-avoided by navigation). Prefer a stall-counter/behavioral trigger — matching this codebase's existing
stallStrafeKey/criticalStallTicksidiom — over a static threshold widening.
scripts/run-balancing-campaign.mjs builds up a much bigger sample than balancing:telemetry's fixed 3-qualifying-run sweep — e.g. 50 qualifying full-campaign runs per combo (450 total) for real balance analysis, rather than the small samples used for regression-testing the bot itself. Differences from balancing:telemetry:
- Resumable, not one-shot. Before touching a combo, it sums
qualifyingRunCount(a fieldbuildComboOutputalready returns) across every file already saved for that combo underbalancing_runs/— killing and restarting the campaign picks up exactly where it left off, with no separate progress-tracking state to drift out of sync with what's actually on disk. - Every batch is its own file, kept forever (not overwritten) — each spawned
run-balancing-telemetry.mjsinvocation is scoped to one combo, collects a small batch (CODEENSTEIN_CAMPAIGN_BATCH_SIZE, default 5) viaCODEENSTEIN_TELEMETRY_QUALIFYING_TARGET, and writes directly to its own path viaCODEENSTEIN_TELEMETRY_OUTPUT_FILE(balancing_runs/<profile>-<difficulty>-<NNN>.json) — no shared-file race between concurrently-running combos. - Runs combos as separate OS processes (
child_process.spawn,CODEENSTEIN_CAMPAIGN_LANESat a time, default 2), each wrapped in a wall-clock watchdog (CODEENSTEIN_CAMPAIGN_WATCHDOG_MS, SIGTERM then SIGKILL after a grace period). This is deliberate, not incidental:run-balancing-telemetry.mjshas no internal safety net for a genuinely wedgedpage.evaluate()/virtual-clock pump — every internal "stuck" resolution (tick-count give-up counters,page.waitForFunctiontimeouts) is bounded and resolves into a normal, non-throwing result, but a true hang would leave aPromise.allinsiderunCombowaiting forever with nothing to catch it. Only an external, OS-level kill can actually stop that. - The calibration figure below is not a throughput number, and reading it as one costs hours. Measured 2026-08-04 on the same Ryzen 5800X: a full 17-level Gamer/hard run at
CONCURRENCY=10sustains 0.31 attempts/minute — about 3.2 minutes per attempt, with the machine CPU-saturated (load 14.9 of 16 threads, ~990% CPU). Raising concurrency does not help; it is already the bottleneck. The 5m13s figure below looks four times faster only because those 8 attempts were racing to qualify (clear level 4) and mostly died early rather than playing all 17 levels — and because 8 attempts atCONCURRENCY=8all fit in one wave, where 60 attempts take six. Budget a fixed-denominator sweep at ~2.6 min/attempt. A 9-combo × 60-attempt sweep is ~23 hours, not the ~8 the figure below suggests.- Revised 2026-08-05 against a completed 360-attempt sweep, which corrected this bullet twice. The 0.31 figure was itself measured on the browser that was dying (see the invocation-length note below); fresh 20-attempt invocations sustain 0.37–0.41 attempts/minute, and the full 6-combo × 60 sweep ran 338 attempts in 14.5h at 0.39/min. And the prediction that normal would be slower than hard was wrong in both directions of reasoning: normal cells ran faster (chunks of 20 in 32–51 min against hard's 49–61), because weaker enemies mean less time fighting per level, which outweighs playing more levels before dying.
- Calibrate the watchdog before a real run on new hardware — the default (90 minutes) was derived 2026-07-15 on a Ryzen 5800X from one real production-representative invocation (full 17-level campaign,
CONCURRENCY=8,QUALIFYING_TARGET=5): 5m13s for 8 attempts to reach 5 qualifying (level-4+) runs, extrapolated to a ~50-minute worst case atATTEMPT_CAP=80with headroom on top. Re-run a similar single-combo calibration invocation (noLEVEL_LIMIT) if running on meaningfully different hardware before trustingCODEENSTEIN_CAMPAIGN_WATCHDOG_MS's default. - Cost is bounded per combo, not just per invocation (added 2026-07-30). The retry loop keeps asking for another invocation until the combo reaches its qualifying target, so a combo the bot simply cannot clear would respawn invocations forever — the watchdog bounds one invocation, never the loop around it. This bit the 2026-07-24 multiplayer campaign for real: the Hard cells never converged (Gamer/hard/2p banked 1 qualifying run across 6 invocations) and the run had to be rescued by hand-lowering the target mid-flight.
CODEENSTEIN_CAMPAIGN_MAX_INVOCATIONS(default 6) /CODEENSTEIN_MP_CAMPAIGN_MAX_INVOCATIONS(default 3) cap it: at the cap the combo gives up loudly, the lane moves on, and the partial data stays intact and resumable — the combo just reports short of target. Set either to0for the old unbounded behaviour. Both orchestrators print the active bound in their startup banner, so a running campaign states its own worst case rather than leaving it to be inferred.- Corollary for clear-rate questions (as opposed to per-combo aggregates): a qualifying target is the wrong knob, because it stops early on easy combos and never terminates on impossible ones. For a fixed denominator, invoke the underlying telemetry script directly per combo with a high qualifying target and
ATTEMPT_CAPset to the sample size you want — the same trick single-player uses withCODEENSTEIN_TELEMETRY_QUALIFYING_TARGET=999.
- Corollary for clear-rate questions (as opposed to per-combo aggregates): a qualifying target is the wrong knob, because it stops early on easy combos and never terminates on impossible ones. For a fixed denominator, invoke the underlying telemetry script directly per combo with a high qualifying target and
- Known risk: a SIGKILL (only reached if SIGTERM doesn't land within the grace period) can leave orphaned Chromium subprocesses behind, since it doesn't give Playwright's own shutdown handlers a chance to run. Kills should be rare (the watchdog is a safety net, not the normal exit path) but worth an occasional
ps aux | grep chromiumspot-check on a long unattended run. - One browser does not survive a 60-attempt invocation (measured 2026-08-04).
main()launches a singlechromiumand reuses it for every attempt in the invocation. On a full 17-level Gamer/hard cell it died at attempt ~38, 2h20m in; all 22 remaining attempts failed instantly onbrowser.newContext()— and the run still reportedattempts used: 60and exited "successfully". Cap a long sweep at ~20 attempts per invocation so each gets a fresh browser, and drive the cell as several invocations.- Corollary, and the load-bearing one:
attemptsUsedcounts attempts started, not samples obtained. It read 60 on a cell that produced 38. Any denominator that matters should be counted as distinctrids in the event log instead —ridis${pid}-${random}-${counter}(run-balancing-telemetry.mjs'sEVENT_SESSION_ID), so it is unique per invocation and several invocations can append to one cell's NDJSON without colliding or double-counting.
- Corollary, and the load-bearing one:
- The script finishes without exiting, so a non-zero exit code is not automatically a failure.
main()writes the output JSON, printsTelemetry saved, and resolves — but nothing callsprocess.exit(0)and a dangling Playwright handle can keep the event loop alive indefinitely. A watchdog then kills it and reports failure for a run whose work completed and whose JSON is already on disk. A driver that deletes output on non-zero rc will destroy good results; check the log forTelemetry savedfirst. Observed 2026-08-04: a complete cell was deleted exactly this way. - Tune
CODEENSTEIN_CAMPAIGN_LANES/_CONCURRENCYto the machine — each lane's invocation gets its ownCODEENSTEIN_TELEMETRY_CONCURRENCY-way internal browser-context concurrency (default 8, lower thanbalancing:telemetry's own default of 12, sinceLANESof these run at once), so total concurrent browser contexts is roughlyLANES × CONCURRENCY_PER_LANE.- A capture splits its local share further, across processes rather than contexts (
CODEENSTEIN_CAPTURE_LOCAL_LANES, defaulted from core count, capped at 4). The total local context count is unchanged —CODEENSTEIN_CAPTURE_CONCURRENCYis divided among the local lanes — because the point is to divide the queue, not to ask more of the machine. Worth 1.60x; see the phase-timing section above.
- A capture splits its local share further, across processes rather than contexts (
- The queue/resumability/watchdog engine itself lives in
scripts/lib/laneOrchestrator.mjs, shared withrun-balancing-campaign-multiplayer.mjs(see SSH-host parallelism below) —run-balancing-campaign.mjsitself only supplies the combo list, env vars, and how to read an existing output file's qualifying count; aRunner(localchild_process, or a remote SSH host) is what actually executes an invocation.
balancing:telemetry's and balancing:telemetry-multiplayer's own real-time-costly data collection can be spread across N SSH hosts — not by giving those one-shot scripts an SSH concept of their own (neither has a lane/queue to plug one into), but through the two campaign orchestrators (balancing:campaign, balancing:campaign-multiplayer), which already exist specifically to spawn many instances of the underlying telemetry script in the first place. Both orchestrators can spread their local lanes across N SSH hosts as well, on top of (not instead of) CODEENSTEIN_CAMPAIGN_LANES/_MP_CAMPAIGN_LANES local lanes — useful when one machine's own core count is the bottleneck, or (for the multiplayer campaign specifically) when running more than one local lane isn't possible at all (see below).
- Host list: a gitignored
ssh-hosts.envat repo root, oneuser@hostper line — auto-managed, not normally hand-edited (see the setup step right below, which appends to it automatically). Blank/#-prefixed lines are ignored; a missing or empty file just means "local lanes only," the common case. - Auth is entirely external — whatever a plain
ssh user@hostwould already use (a pre-unlocked key in your localssh-agent, or a~/.ssh/configalias). Neitherssh-hosts.envnorscripts/lib/sshRunner.mjsever touch credentials. - One-time setup per host, then zero sudo forever after. Adding a new host is one command:
node scripts/setup-ssh-lane-host.mjs user@newhost— a real interactive SSH session where you may be prompted for your sudo password, installing git/a modern Node if missing, Playwright's Chromium system dependencies, cloning the repo, and runningnpm ci. Once setup succeeds, the host is appended tossh-hosts.envautomatically (appendHostIfMissing) — no separate manual edit. Run the same script with no arguments to re-run setup on every already-listed host at once (e.g. after a system update wiped one's Node install) — a safe no-op for any host already present. After setup, the automated per-run bootstrap (sshRunner.mjs's ownbootstrapHost(), run by every real campaign invocation) never touches sudo/apt at all — it only checks git/an adequate Node are already there (failing with a pointer back to the setup script if not), then clone-or-fetch, force-checkout the exact localHEADcommit,npm ci, andnpx playwright install chromium(browser binary only). A host that's unreachable or fails any automated step is logged as a warning and simply excluded from that run — one bad host must never wedge the whole orchestrator (the same lesson a real stuck combo already taught: see the multiplayer campaign's ownATTEMPT_CAPdefault below). - The remote clone always uses
https://, never this machine's own configuredoriginURL as-is — confirmed directly as a real failure: this repo's ownoriginis an SSH-stylegit@github.com:...URL (the natural default for an owner who pushes), and shipping that to an arbitrary lane host assumes it has a matching GitHub SSH key too, which a fresh host generally doesn't.toHttpsCloneUrl()(sshRunner.mjs) normalizes both common SSH forms tohttps://before it ever reaches a remote host — a public repo needs no credentials at all overhttps://. - Why this needs a separate one-time script at all, rather than just automating everything: two of its steps genuinely can't be made both unattended and narrowly sudo-scoped. Installing Node needs NodeSource's own
curl | sudo bashsetup script (or the equivalent by hand) — sudoers can only match the executable (bash), never what's piped into its stdin, so a NOPASSWD rule for that is unrestricted passwordless root, not a scoped step. And Playwright's own--with-deps/install-depshas Playwright itself decide andapt-get installan arbitrary, OS/version-state-dependent package list at run time (confirmed directly —npx playwright install-deps --dry-run chromiumreported a different missing-package list on hosts at different patch levels) — there's no fixed command line a sudoers rule could ever pin for that. Splitting setup (real interactive sudo, once) from every automated run (no sudo, ever) sidesteps both instead of trying to scope either. - ARM hosts work fine — nothing in
sshRunner.mjsis architecture-specific, and Playwright's Chromium build (the only engine this whole family ever launches) has genuine Linux ARM64 support. Real per-attempt wall-clock cost can still vary a lot by hardware, same "calibrate before trusting" discipline asCODEENSTEIN_CAMPAIGN_WATCHDOG_MS's own calibration note above. - A remote lane's own result file is pulled back via
scpinto the exact local path the orchestrator's resumability scan expects, so local and remote lanes are indistinguishable from the queue's point of view. - Known gap, not yet solved: a local watchdog timeout kills the local
sshclient, which best-effort propagates (via a forced pseudo-terminal,-tt) to the remote command, but a genuinely dropped connection can still leave an orphaned remote process running — a real fix needs a remote supervisor, out of scope for this first cut. - Multiplayer-specific limitation:
run-balancing-campaign-multiplayer.mjsdefaultsCODEENSTEIN_MP_CAMPAIGN_LANESto 1, not 2 — every local invocation starts its own isolated signaling+dev server pair on the same fixed ports (multiplayerTestServers.mjs, 8788/5174), so two concurrent local lanes would collide today. Real multiplayer parallelism is expected to come from SSH lanes (each its own remote machine, no port conflict) rather than raising the local lane count.
Any change to navigation/combat/movement logic in run-balancing-telemetry.mjs needs more than "the scan came back clean" before it's trustworthy:
- A/B against a baseline worktree (
git worktree add /tmp/bot-ab-base <last-good-sha>— notgit stash, which can't serve both sides against one dev server) at the sameCODEENSTEIN_TELEMETRY_CONCURRENCY/_ATTEMPT_CAP/_LEVEL_LIMITthat will ultimately be trusted. A small or low-concurrency sample has previously masked a real ~4x survival-rate regression (Casual/normal level-2 death rate looked fine atCONCURRENCY=1, but was 72% — vs. the true baseline's 0% — atCONCURRENCY=6/ATTEMPT_CAP=20). diagonalStrafeKey(the bot's diagonal-movement helper, plain-navigation branch only) is the sharpest cautionary example: an earlier change to its usage caused exactly that 72%-vs-0% regression, only caught via the matched-scale A/B above — not bybalancing:scan. It's scoped to plain-nav only for this reason; don't re-add it tohazard/criticalHealth/mineRetreat/ranged-aim branches, and treat even refinements within its current safe usage as needing the same verification bar, not just a scan.
The "low concurrency masked it" story above is real but was misattributed, which matters because it made the fix look like a knob rather than a rule. runQualifyLoop (scripts/lib/qualifyLoop.mjs) is while (qualifyingRuns.length < requiredQualifyingRuns && attempts < attemptCap), running concurrency attempts per batch — and failureReasons only accumulates from attempts actually run. At CONCURRENCY=1 a bot that qualifies 3-of-3 runs exactly three attempts and records zero failures, whatever its true death rate; at CONCURRENCY=6 the first batch always runs six, so failures get recorded. Concurrency was changing the sample size, not the simulation.
So don't rely on concurrency to produce a denominator. Set CODEENSTEIN_TELEMETRY_QUALIFYING_TARGET=999 for any A/B, which makes the loop always run to ATTEMPT_CAP — a guaranteed denominator instead of an incidental one. Keep CONCURRENCY=6/ATTEMPT_CAP=20 as well, to honour the bar the regression above established.
Run once per side, then diff:
for combo in "Casual normal" "Gamer normal" "Pro hard" "Pro normal"; do
set -- $combo
CODEENSTEIN_TELEMETRY_PROFILE=$1 CODEENSTEIN_TELEMETRY_DIFFICULTY=$2 \
CODEENSTEIN_TELEMETRY_LEVEL_LIMIT=8 CODEENSTEIN_TELEMETRY_ATTEMPT_CAP=20 \
CODEENSTEIN_TELEMETRY_CONCURRENCY=6 CODEENSTEIN_TELEMETRY_QUALIFYING_TARGET=999 \
CODEENSTEIN_TELEMETRY_ANOMALY_SCAN=1 \
CODEENSTEIN_TELEMETRY_OUTPUT_FILE=ab/<side>-$1-$2.json \
node scripts/run-balancing-telemetry.mjs
done
node scripts/report-balancing-ab.mjs ab/base ab/cand # dirs or single filesThe four combos are not arbitrary: Casual/normal is the combo the 72% regression showed up on, Gamer/normal is where most quoted telemetry numbers come from, and Pro/hard is where enemyAimSpreadDeg = 0 (perfect enemy aim) makes any dodging/movement change matter most.
Pro/normal is the fourth for a reason worth stating, because it is the kind of gap that stays invisible until it costs a day. A profile and a difficulty are independent axes, and running Pro only on hard leaves every Pro-specific behaviour untested at normal's enemy aim and damage. That is exactly what happened on 2026-07-31: a reproducible ~22s freeze at a fixed tile on level 1, hitting roughly 40% of Pro/normal attempts, survived a full day of A/Bs because no side of any A/B ran that cell — it only surfaced in the wider balancing:scan, which does sweep every profile. Two of the three combos above are normal and the third changes both axes at once, so Pro and hard were confounded: any Pro-only regression was indistinguishable from a hard-only one. Adding this cell makes the profile axis separable at fixed difficulty, and it is the cheapest possible insurance against re-learning that lesson.
The cost is real — a fourth combo is a third more wall clock per side — so if you must drop one for a change that plainly cannot interact with difficulty (a pure navigation or routing change, say), drop Pro/hard and keep this one, not the other way round.
report-balancing-ab.mjs splits the comparison in two on purpose, and the split is the point:
- Guard metrics —
qualifyRate(fromtrueQualifyingCount, never the flooredqualifyingRunCount) and a per-level conditional death rate,died[i] / reached[i]. Raw death counts don't compare across two runs with different reach: a level nobody got to has no deaths. These are attempt-level, so n=20 detects a 0%→72% swing instantly and nothing near 10pp — small guard movements are not readable at this sample size and must not be reported as if they were. Pre-registered rollback thresholds (inabReport.mjs, deliberately fixed in code rather than chosen after seeing the numbers): qualifyRate down >15pp, any level's conditional death rate up >20pp, or any increase in stuck count. A breach means revert, not tune. The CLI exits non-zero on a breach. - Win metrics —
enemyAccuracy,levelTimeSec,distanceTraveled,routeFollowingOverhead, TTK, health, damage-by-source. Per-level-visit aggregates over hundreds of samples, so these do have real resolution at the same n. NoterouteEfficiencyScoreis marked"flat", not"up": see Route efficiency is mostly not a bot metric below.
A guard pass bounds a large regression, not a small one — know the floor before reading one (2026-08-11). The four-arm aim A/B ran the recipe above exactly as written and every arm cleared 20/20 at every level with a 0.000 death rate. That is the demo campaign's own easiness (levels 1-14 clear at ~99% on Hard) rather than anything about the change under test. The guards still fired correctly against that baseline — verified by feeding checkRollback synthetic diffs, not by reading it: 5/20 deaths at a level, 4/20 attempts failing to qualify, 5/20 runs lost between levels, or a single stuck all breach. What a zero base costs is resolution, and at ATTEMPT_CAP=20 the floor is exactly that: 5 deaths, 4 non-qualifying runs, 1 stuck. So "GUARDS: pass" on that run meant "no regression above that floor" — which is what the guards were designed to catch (the 0%→72% diagonal-strafe regression clears it by a mile) and is not the same as "survival is unaffected". Read it as the former. If a change could plausibly cost a handful of runs rather than most of them, the floor is too coarse to see it, and the answer is a bigger sample or a test case with real mortality (deeper than level 8; 15 is the wall) — not a re-reading of the same pass. An earlier version of this paragraph claimed the guards "could not have breached their thresholds whatever the change did" and "carried no information at all"; both are wrong, and the correction is recorded in Development History.
Passing guards means "not obviously worse", never "the change worked" — always read the win metrics against whatever the change was actually supposed to do. If a guard lands ambiguously (5–20pp worse), escalate that one combo to ATTEMPT_CAP=60 before deciding; never ship on the ambiguous reading.
One A/B side is roughly 4x balancing:scan's wall clock, and a full gate is two of them — confirm before launching one.
A JSON object deep-merged over DEFAULT_TUNING for every bot the run builds:
CODEENSTEIN_TELEMETRY_TUNING='{"NEAR_PI_HEADING_EPS":0}' node scripts/run-balancing-telemetry.mjsThis is how a behaviour A/B should be run whenever the change is gated by a constant. The alternative — a worktree at an older commit — makes the whole diff the thing under test rather than the one value, and worse, a worktree predating the metric you want to read cannot emit it at all. That is not hypothetical: the first attempt to grade the atan2 branch-cut fix stalled on exactly that, because the baseline commit had no anomaly tally in its output. With the override, both sides run the same binary and differ only by the value. Invalid JSON is a hard exit rather than a warning, since silently falling back to defaults would produce a baseline-vs-baseline comparison that looks like a real result.
Behaviour changes that are gated on a constant expose one deliberately, so they stay A/B-able against the same binary rather than needing a worktree:
| switch | false restores |
|---|---|
NAV_FULL_WASD |
standing still to turn whenever the heading error exceeds MAX_WALK_WHILE_TURNING_RAD |
NAV_BACKPEDAL_RETREAT |
spinning to face away before fleeing at critical health |
Three more were added 2026-08-03 for the verify (multiplayer-transition) fix. All three are off in DEFAULT_TUNING and on in MultiplayerBot, so single-player telemetry is unchanged by construction rather than by measurement — and so they can be A/B'd into single-player later without a worktree:
| key | default | on |
|---|---|---|
BOT_NAV_STALL_BAIL_TICKS |
0 (off) |
24 — give up on a drive that has not left BOT_NAV_STALL_RADIUS_TILES (0.5) for this many consecutive decisions, so driveTowardWithReplan re-plans while it still has budget. Sits just above the detectors' own STALL_TICKS_THRESHOLD (20), which keeps the invariant "the bail only fires on something the anomaly scan would have reported as a stall anyway". Suppressed while engaged in combat or waiting out a spike trap, mirroring detectAnomalies' mostlyFiring and SPIKE_WAIT_DOMINANCE exemptions. |
BOT_LOOT_ABANDON_ON_STUCK |
false |
true — abandon a loot detour whose waypoint has exhausted its re-plans, instead of driving the rest of a path planned from a tile the bot never reached. |
MAX_TICKS_PER_WAYPOINT |
600 |
40 — 600 was sized against VIRTUAL_STEP_MS (50), i.e. 30 simulated seconds; at MultiplayerBot's 400ms decisions the same number is 240 real seconds for a one-tile waypoint. |
Note the A/B for these has to be run into single-player (turning them on) rather than out of it, since the multiplayer campaign has no baseline corpus.
Multiplayer telemetry semantics changed on 2026-08-03, and stored runs from before it are not comparable. Two things moved: MultiplayerBot now carries the tuning above, and driveOneBot's final approach uses Bot#driveToExit instead of a bare driveToward. The second one moves outcomes directly — an exit held shut by a living exit-room enemy used to record as stuck however perfectly the bot was standing on it, and that was scored against the bot's navigation when the route had worked exactly as planned.
meta.flags did not exist in multiplayer_balancing_telemetry.json until the same date, which meant compareRunFlags returned comparable: false and every cross-run multiplayer comparison silently lost its one guard against this. It now records botTuning (from MULTIPLAYER_TUNING_DEFAULTS, exported so the recorded value cannot drift from the value used) and finalApproach. A comparison spanning the change will now say so instead of quietly reporting a behaviour delta as a balance delta.
What did not change is the outcome vocabulary. driveToExit reports arrived only when this bot stood on the exit tile and saw the exit accepted; a teammate's exit touch still arrives as teleported and is still classified levelAdvanced. That distinction is what trueQualifyingCount rests on, and widening it into reachedExit would have inflated exactly the number the campaign is judged by.
With CODEENSTEIN_TELEMETRY_ANOMALY_SCAN=1, each combo's output carries an anomalySummary — per anomaly type, findings/ticks totals plus findingsPerRun, ticksPerRun and ticksPerKiloDecision. report-balancing-ab.mjs diffs the last of those.
Two normalizations, both learned the hard way:
- Ticks, not findings.
detectOscillationcounts events. A change that makes the bot cover less ground can trip more qualifying windows while behaving better — which is how the first oscillation fix got graded as a +9.6% regression and reverted on a number that didn't mean what it looked like. - Per decision, not per run. Even ticks-per-run is not exposure-independent. Measured on a real comparison: oscillation ticks/run fell 11.7% while
levelTimeSecfell 7.7%, so most of the apparent win was simply less time on the level.ticksPerKiloDecisiondivides that out.
balancing_telemetry.json (repo root, gitignored) holds a meta block (profile definitions), then per-level and campaign-wide aggregates across 7 categories (map density/demographics, combat pacing, AI effectiveness/danger, damage/healing breakdown, weapon efficiency, economy/loot starvation, navigation/map flow), plus deterministic outlier flags and per-profile crossDifficultyFlags. Judgment-call metrics carry a {mean, max, min, samples} spread rather than a bare mean, so a consumer (human or LLM) can see the actual distribution, not just a single number that might hide a bimodal split.
Two raw fields were added 2026-07-27 because the derived ratios they already fed couldn't answer an A/B on their own. combatPacing.levelTimeSec — previously read only as combatVsExplorationRatio's denominator and never emitted, which left the playtest bot's loudest failure mode (being far slower over a route than a human) with no output field at all; a ratio can't substitute, since a bot that is uniformly 2x slow reports an unchanged ratio. And navigationMapFlow.distanceTraveled — the raw counterpart to routeEfficiencyScore, which is 0 whenever shortestPathTiles is null and so can't distinguish "walked a tight route" from "the optimum was unknown". Together they separate "the bot got faster" (time down, distance flat) from "the bot took a shorter route" (both down).
navigationMapFlow.routeEfficiencyScore sat at 0.345 on the demo campaign (Gamer/normal) and was being read as "the bot walks 3x further than it needs to". It does walk 2.9x the theoretical minimum, but only a minority of that is the bot's doing. Decomposed against the planned route and a new per-decision activity attribution (summarizeActivityDistance in bot.mjs, which charges every tile walked to the errand the bot was on at the time), the 2.9x is three near-independent factors multiplying to 2.97x:
| factor | size | whose fault |
|---|---|---|
| planned route vs bare spawn→exit BFS | 1.68x | level design — and entirely the two key/locked-door levels (demo 3 and 7, 3.7x and 3.5x). The other six plan within 5% of optimal. |
| loot detours (24% of all distance walked) | 1.32x | bot policy, and mostly justified — 63% of those tiles are ammo, 19% mandatory keys, and only 3.3% of them (0.8% of total distance) is provably wasted: a health pack grabbed at hp=1.00, where MAX_HEALTH caps the gain at nothing. |
| actually following the plan | 1.30x | the bot outright. 4.2% of distance is covered while engaged with a threat; ~1.2% of simulated time is the known atan2-branch-cut oscillation. Replan retries and the exit-gate fallback contribute ~0% on these eight levels. |
Two consequences, both now encoded in the code and in WIN_METRICS:
routeEfficiencyScoreis a poor A/B win metric for navigation changes — it moves mostly with level layout and loot policy. It is marked"flat"inabReport.mjsrather than"up".navigationMapFlow.routeFollowingOverheadis the metric to judge navigation on. Distance ÷ the route the bot planned for itself (staticAnalysis.plannedRouteTiles), as a multiplier where 1.0 is a perfect walk; measured 1.71 distance-weighted across levels (the campaign rollup's per-run mean reads ~1.64). Level design divides out; loot stays in, deliberately, because detouring is the bot's own decision.nullwhen the route failed to plan and for the campaign-wide rollup.
plannedRouteTiles sums straight-line waypoint-to-waypoint distance, which is exact here rather than an approximation: planRoute emits one waypoint per tile, and a BFS re-measurement of the true walkable path agreed to the tile on all eight demo levels.
Also, at the combo level (alongside weaponFirstOwnedAtLevel): weaponFirstOwnedAtLevel is a min across qualifying runs — it answers "how soon could this profile realistically get it", not "how often did any run get it at all". weaponAcquisitionRate ({ [weaponIndex]: { count, rate } }) answers the second question directly, for every unlockable weapon (gdb/ghidra/Friday Hotfix/Toolchain) uniformly — added 2026-07-15 specifically to verify Toolchain's new miss-chance acquisition path (see below) actually moves the needle, since the min-level metric alone can't distinguish "3% of runs get it, always around level 9" from "60% of runs get it, always around level 9".
lootRolled used to record a flat 1 (an occurrence, not a quantity) for every drop whose LootDrop.amount was unset at roll time — which is most non-Elite drops, since the real amount is only resolved later, at collection (applyLootDrop). This made lootRolled unit-incompatible with consumed (a real-amount total): a report built on comparing the two (an ammo_starvation_* outlier flag) had to be removed rather than fixed as a result. RaycasterEngine.pushLootDrop now records the real, difficulty-scaled amount a drop is worth (defaultLootAmountFor mirrors applyLootDrop's own fallback exactly) for every kind except "weapon", which stays an occurrence count on purpose — a weapon drop's real value (grant vs. an ammo top-up if already owned) depends on ownership state at collection time, which can change between roll and collection, so 1 is the only thing that can honestly be recorded for it regardless of when. lootRolled and consumed.total should now sit within roughly the same order of magnitude per resource, not off by 10-20x.
economyLootStarvation also gained pctRegularKillLootMisses (a {mean, samples} spread, per-level-visit): the fraction of regular (non-Elite) kills whose ammo/swap roll came up empty — see REGULAR_KILL_NO_DROP_CHANCE in src/engine/loot.ts. Not a "desperation" signal on its own (health is a separate, always-on grant now, independent of this roll — see game-design.md's "Weapon and economy intent" for why) — a mechanic-verification stat, letting real telemetry confirm the ~20% design rate empirically instead of trusting the constant alone.
Before 2026-07-15, enemy ranged bolts had zero aim deviation at all — enemyAccuracy (hits/shots fired) was purely a function of the player dodging (movement, walls), never anything the difficulty setting touched. DIFFICULTY_MULTIPLIERS.enemyAimSpreadDeg (10°/4°/0° easy/normal/hard, src/difficulty.ts) now rotates a bolt's aim vector by a random angle up to that cap before firing (spawnProjectile in projectiles.ts) — enemyAccuracy should now show a real, monotonic difficulty curve instead of the flat ~70-77% band across all three tiers that the original balance report flagged as "difficulty makes enemies tougher, not smarter". Verified directly: a Gamer-profile spot-check went 74.9%→45.6% (easy), 73.8%→58.0% (normal), 77.3%→78.0% (hard, unchanged — 0° spread is the same as the old always-perfect aim).
A separate tool, scripts/run-balancing-telemetry-multiplayer.mjs, mirrors this whole toolchain for real multiplayer sessions (2-4 simultaneous players) — step 11 of the multiplayer implementation plan. Full design rationale lives in doc/dev/multiplayer-balancing-telemetry-spec.md; this section is the user-facing "how to run it" reference. Not CI-wired, not fast — run manually, same as balancing:telemetry.
The two are more different than they look at first glance, for one structural reason: multiplayer has no virtual clock. scripts/lib/virtualClock.mjs cannot fast-forward a real Web-Worker-timer-paced multiplayer simulation (scripts/lib/multiplayerBot.mjs's own doc comment states this outright) — every attempt costs genuine wall-clock time. A single combat-heavy level clear for a 2-bot pair has been directly measured at ~4 real minutes. Every default in this tool (sequential attempts, small qualifying targets, one bundled level per run instead of a full campaign) is sized around that cost, not copied from single-player's cheap virtual-time concurrency.
| Command | What it does |
|---|---|
npm run balancing:telemetry-multiplayer |
Full combo sweep across every profile/difficulty/player-count (plus curated mixed-skill combos, see below), 2 qualifying runs per combo by default, writes multiplayer_balancing_telemetry.json. Real-time cost means this can run for a long while — scope it with the env vars below before trusting the unbounded default to finish quickly. One monolithic process: no incremental persistence, so a kill loses everything collected so far — see balancing:campaign-multiplayer below if that matters. |
npm run balancing:scan-multiplayer |
Fast/cheap preset: Casual/normal/2p only, attempt cap 3, qualifying target 1, disconnectIsolation scenario disabled. The multiplayer pre-merge regression gate — mirrors single-player's balancing:scan role. |
npm run balancing:campaign-multiplayer |
Resumable orchestrator (scripts/run-balancing-campaign-multiplayer.mjs) — spawns one OS process per combo (via the new CODEENSTEIN_MP_TELEMETRY_COMBO_PROFILES pin, see below), each writing its own file under balancing_runs_multiplayer/, shares scripts/lib/laneOrchestrator.mjs with balancing:campaign. See Large-scale campaigns/SSH-host parallelism above — same design, CODEENSTEIN_MP_CAMPAIGN_* env vars instead of CODEENSTEIN_CAMPAIGN_*, smaller defaults throughout (_TARGET 10, _BATCH_SIZE 2, _ATTEMPT_CAP 30, _LANES 1, _MAX_INVOCATIONS 3) given the much higher real-time cost per attempt. |
Both always start their own isolated signaling + dev server pair (scripts/lib/multiplayerTestServers.mjs, ports 8788/5174 — deliberately never 8787/5173, a developer's own manual session's default ports) rather than share whatever a developer's own dev session happens to be pointed at. The signaling server's rate limits are per-IP, not per-session (multiplayer-balancing-telemetry-spec.md §7) — running this tool's own traffic against a shared server risks tripping a budget sized for one human's manual testing. There's no CODEENSTEIN_DEV_URL-equivalent override for this reason: the isolated pair is always used.
Reuses run-balancing-telemetry.mjs's own PROFILES (Casual/Gamer/Pro) and DIFFICULTIES (easy/normal/hard) unchanged. Beyond that, the combo matrix is genuinely different from single-player's:
- One bundled demo-campaign level per run, not the full campaign. Multiplayer level transition is already covered on its own by
verify-multiplayer-transition.mjs(one transition, host god-moded) — re-driving that whole sequence for every combo would multiply this tool's already-real-time-only cost for no new signal. A run "qualifies" once every bot reaches the exit tile alive (teamOutcome === "allReachedExit"). Chaining several consecutive real transitions (the gap neither that script nor this one covers) is instead a separate, dedicated functional check — seenpm run verify:multiplayer-campaign(scripts/verify-multiplayer-campaign.mjs), not a balancing-data tool itself. - Player count (2-4) is a real combo dimension, not fixed —
CODEENSTEIN_MP_TELEMETRY_PLAYER_COUNTS(comma-separated, default2,3,4). - Uniform combos (one skill tier for the whole team) run alongside curated mixed-skill combos (
curateMixedProfiles()) when noPROFILEfilter narrows things to one tier: 2p gets only adjacent-tier pairs (Casual+Gamer, Gamer+Pro — not the skip-a-tier Casual+Pro, less representative of a real pairing while costing the same real time as either neighbor); 3p/4p get one weakest+strongest+filler combo each (filler = the middle tier, repeated for 4p). Deliberately not a blind cartesian product across up to 4 slots — that multiplies cost for combos with little new signal over their neighbors. APROFILEfilter disables mixed combos entirely (a filter means "just this one tier").
All CODEENSTEIN_MP_TELEMETRY_* — read once at module load, same "same process invocation" caveat as the single-player table above.
| Var | Effect |
|---|---|
CODEENSTEIN_MP_TELEMETRY_PROFILE |
Restrict to one profile tier — also disables curated mixed-skill combos (see above). |
CODEENSTEIN_MP_TELEMETRY_DIFFICULTY |
Restrict to one difficulty. |
CODEENSTEIN_MP_TELEMETRY_COMBO_PROFILES |
Pins one exact per-slot combo (comma-separated tier names, e.g. Casual,Gamer for a specific 2p mixed pair) — bypasses the uniform+curated-mixed matrix entirely, running just that one combo. Player count is derived from this list's own length (_PLAYER_COUNTS is ignored); requires _DIFFICULTY to also be set. Used by balancing:campaign-multiplayer to scope one spawned invocation to one combo — a bare _PROFILE filter can express a uniform combo but not a specific mixed one. |
CODEENSTEIN_MP_TELEMETRY_PLAYER_COUNTS |
Comma-separated list of player counts to test, each 2-4 (default 2,3,4). |
CODEENSTEIN_MP_TELEMETRY_QUALIFYING_TARGET |
Qualifying runs needed per combo before its retry loop stops (default 2 — deliberately much smaller than single-player's default 3, given the real-time cost per attempt). |
CODEENSTEIN_MP_TELEMETRY_ATTEMPT_CAP |
Cap attempts per combo (default unbounded). Use for any scoped/smoke run. |
CODEENSTEIN_MP_TELEMETRY_CONCURRENCY |
Attempts run concurrently within one combo (default 1, sequential) — several concurrent real multiplayer sessions against one dedicated signaling+dev server pair is a real resource-contention risk this tool hasn't been measured against; raise deliberately. |
CODEENSTEIN_MP_TELEMETRY_VERBOSE |
Per-attempt detail logging. |
CODEENSTEIN_MP_TELEMETRY_ANOMALY_SCAN |
Enables the shared Bot stall/healthDrainFrozen/rotation detectors (see Anomaly scanning above — these work for MultiplayerBot unchanged, they just need the trace collector turned on). |
CODEENSTEIN_MP_TELEMETRY_NAV_DIAG |
Extra per-decision trace bookkeeping (implies ANOMALY_SCAN). |
CODEENSTEIN_MP_TELEMETRY_HEADED |
Real, visible browsers instead of headless. |
CODEENSTEIN_MP_TELEMETRY_DISCONNECT_SCENARIO |
Set to 0 to skip the disconnectIsolation scenario (default on — it's what balancing:scan-multiplayer's own preset disables, since the real detection wait would dominate that fast preset's own runtime budget). |
CODEENSTEIN_MP_TELEMETRY_OUTPUT_FILE |
Override the output path (default multiplayer_balancing_telemetry.json at repo root, gitignored). |
multiplayer_balancing_telemetry.json (repo root, gitignored) is top-level-keyed by combo (meta + combos), each combo holding:
perPlayerTelemetry— the real per-player 7-category breakdown (map density, combat pacing, AI danger, damage/healing, weapon efficiency, economy, navigation), keyed by roster id (host,guest-1, ...). Reusesrun-balancing-telemetry.mjs's ownaggregateLevelRuntime()unchanged:RaycasterEngine.getMultiplayerTelemetrySnapshot(id)'s shape matches single-player's owngetTelemetrySnapshot()field-for-field (both built from the samebuildTelemetrySnapshotFor). One category single-player has that multiplayer doesn't:navigationMapFlow.routeEfficiencyScoreis omitted — each bot spawns at a different tile, so there's no single team-wide shortest-path figure to compare against, and shipping the aggregator's own "not computed" placeholder zeros would read as a real (and misleadingly bad) result.gameplayHealth— coarser team-level signals: outcome tally (allReachedExit/teamWiped/partial/crashed), a team-wide enemies-killed estimate (a before/after alive-count delta — not per-player-attributable the wayperPlayerTelemetry's own kill counts are, since assist vs. finishing-blow can't be told apart from a bare count), and each player's minimum observed health fraction.perf— fps per player, mean tick-skew per peer pair, andtickSkewGrowthByPair: a first-third-vs-last-third mean comparison per qualifying run, flagging a real desync-widening trend (not just a raw mean/max, which can't tell "briefly spiked then settled" apart from "steadily growing" — a "growing" call requires both a ≥5ms absolute delta and a ≥1.5x ratio, so ordinary real-clock sampling noise can't false-positive).netcodeHealth— real RTT per link (RTCPeerConnection.getStats()'s active-candidate-paircurrentRoundTripTime— star topology, host↔each guest, sampled both directions since each side's own view is a genuinely different measurement point), missed-tick fraction per player (TickInputBundle.heldInputFallbacktally), and reconciliation-correction count/magnitude per player (guest-only — the host is authoritative and never applies a snapshot to itself, so its own entries are always{count: 0, avgMagnitudeTiles: 0}, not missing data; a correction only counts once its position magnitude clears a small noise floor, so ordinary cross-peer float drift doesn't register as a false "correction").
Separately, at the report's top level (not part of the combo matrix — the scenario doesn't vary by bot skill or difficulty, so it runs once per invocation, not once per combo): disconnectIsolation — a real, scored version of verify-multiplayer-disconnect.mjs's guest-disconnect scenario. A real RTCPeerConnection teardown (closing the guest's BrowserContext), then measuring how long the host takes to detect it and whether the host keeps ticking/surviving through the disconnect — {guestFinalStatus, detectedWithinMs, hostKeptTicking, hostSurvived, endedDuringDetection}. Read endedDuringDetection before the latency. Unlike the verify script, this scenario god-modes nobody, so the host can die mid-window — and a guest disconnect that then leaves nobody "alive" ends the run in the same tick that marks the guest "disconnected". getPlayerStatus goes straight from "alive" to null across that, so before 2026-08-23 such a run recorded guestFinalStatus: "undetected" with a null latency — indistinguishable from a real detection failure. It now reports the end reason here and recovers the true final status from getLastSessionEnd(). detectedWithinMs stays null in that case on purpose: the run measured no detection latency, and folding it into the series as though it had would be worse than the ambiguity it replaces.
A mine-corridor stall — root-caused and fixed (applies to single-player too). A uniform-Casual 2-player pair reproducibly hit a stall/healthDrainFrozen anomaly sequence around a mined corridor on demo-campaign/main.c (three real mines clustered within a few tiles of each other, ~pos (37.5, 49.3)). Root cause, confirmed via a live position/health/mine-state trace: findDangerousMine's own "retreat now" trigger only fires once a mine is already within its blast radius — but a mine's fuse (MINE_FUSE_SECONDS, 0.9s) ticks in real time regardless of how often the bot re-evaluates, and MultiplayerBot's own real decision window (DEFAULT_STEP_MS, 400ms) is long enough that a mine already armed by the other mine in the cluster (or by this same bot's own earlier approach) could finish its fuse and detonate entirely within one held decision — the bot correctly saw itself as "safe" (beyond blast radius, aiming at a different, farther mine as a disarm target) right up until the explosion it had no chance to react to. Fixed in shared bot.mjs: findDangerousMine now takes a reactionBufferTiles parameter — a real, decision-window-scaled reaction margin (ENGINE_MOVE_SPEED × ENGINE_SPRINT_MULTIPLIER × stepMs) rather than a fixed tile count, mirroring the existing MELEE_CLOSE_MIN_DISTANCE fix's own pattern. At single-player's much shorter decision windows (WATCH_STEP_MS 130ms, VIRTUAL_STEP_MS 50ms) this rounds to well under a tile — a harmless, mostly-no-op widening; multiplayer's much longer window gets a buffer that actually matters (~2.5 tiles). Verified live: the same stall+healthDrainFrozen compound pattern near the mine cluster is gone, and Casual/normal/2p — previously stuck there — qualified 2/2 on the very next full run. Some mine damage in that corridor is still possible (mines are hazards by design) — what's fixed is the bot getting physically stuck there taking damage it had no chance to react to, not mines being risky at all.
A far more severe stall — root-caused and fixed. The same runs also hit one much longer stall — ~596 ticks (vs. ~40-45 for the mine-corridor one above, and suspiciously close to MAX_TICKS_PER_WAYPOINT's own value of 600), at a different position (~pos (15.8, 16.8)), with no mine or threat nearby. Root cause, confirmed via live repro: checkExit() (engine.ts) starts the level-transition countdown the moment any single alive player touches the exit — a real, intended co-op mechanic ("exit touch is a shared simulation event"), not a bug — and once it elapses, the whole roster is carried to the next level's spawn, including a teammate who's still mid-route. That teammate's own Bot instance had no way to notice: it kept walking its pre-planned waypoint list against a live position that had moved to an entirely different level, using its own now-stale map for every navigation decision (which is exactly why (15.8, 16.8) read as solid wall against main.c's own grid — it was never on main.c at all). Fixed in bot.mjs's shared driveLegs/driveTowardWithReplan/maybeDetourForLoot: a mid-route "teleported" result (already detected, previously silently ignored) now stops the walk immediately instead of continuing. driveOneBot maps this into a new "levelAdvanced" outcome — exactly as real a team success as personally reaching the exit tile, and now counted as such in teamOutcome. Verified for real: the exact combo that had never once qualified across every earlier test run in this investigation (Casual/normal/2p) qualified on its very first attempt after the fix.
Status: implemented. This part was written as a design first and then built; the ordering and the reasoning below are kept as-written because they explain why each piece is shaped the way it is, and §6 records which step shipped what. Where a measurement later corrected the design's own assumption, the correction is recorded inline rather than the original quietly edited away — §7.1 is the clearest example.
What runs today:
| Command | What it does |
|---|---|
npm run balancing:budget |
Solve a campaign's budget offline — no browser, no bot. Nine flags: --dir <path> (default demo-campaign), --difficulty <d>, --all-difficulties, --json <path>, --max-levels <n>, --kill-rate <n>, --carryover-cap <n>, --hp-scaled-drops, --hp-scaled-health. It has no --help — passing one is an unknown argument error. Exits non-zero on an enemy that outlasts every round on its level. |
npm run balancing:corpus |
Fetch the pinned corpus of real repositories to solve against. |
npm run balancing:events |
Turn a raw event log into a markdown report. |
CODEENSTEIN_TELEMETRY_EVENT_LOG=<dir> |
Turn on raw event recording for a telemetry run. Off by default. |
CODEENSTEIN_TELEMETRY_SEED=<n> / ?seed= |
Pin the gameplay seed so loot rolls are reproducible. |
The problem this part exists to solve: levels are generated from arbitrary
repositories, so no amount of playtesting the bundled demo-campaign/ produces
confidence about a repo nobody has ever opened. The bot harness above answers "how
does this campaign play"; it cannot answer "is this level clearable at all" for a
level that does not exist yet. That needs a solver which reads a generated level
and the real combat constants and computes the answer without anyone playing.
Two additions, in priority order:
- An offline solver — generate a level, compute its enemy budget, its loot budget, per-weapon TTK and the resulting ratios. No browser, no bot, no play.
- A raw event log — per-occurrence records written alongside (never instead of) the existing aggregates, so a metric can be invented after the data was collected instead of requiring re-instrumentation.
Every field of TelemetryState and TeamTelemetryState (src/engine/telemetry.ts)
is a scalar accumulator: state.damageBySource[source] += amount,
tallyFor(state, weaponIndex).shotsFired += 1. Once recorded, an event's
timestamp, position, actor and identity are gone — damageBySource.enemyMelee = 412
cannot be decomposed into which enemy, when, or how much per bite.
Raw per-occurrence data exists at exactly two points, and neither reaches disk:
ttkRecords(telemetry.ts) is genuinely one record per enemy that ever aggroed, carrying{category, aggroAtLevelTime, deathAtLevelTime}, and it survives into the snapshot verbatim (buildTelemetrySnapshotForinengine.ts). It is then destroyed by the first aggregation step:run-balancing-telemetry.mjs's per-batch reduce collapses each record to a bare duration inttkByCategory[category], discarding enemy identity and absolute timing. What lands inbalancing_telemetry.jsonis an unordered bag of durations.- The bot's per-decision trace (
mineMemory.traceinbot.mjs, appended by#recordTrace) is a real event log, gated off by default, consumed in-process by the anomaly detectors, and thrown away when the nextstartLevel()reallocates it.CODEENSTEIN_TELEMETRY_TRACE_DUMP=1prints a fragment of it as formatted text.
There is no NDJSON, no CSV and no append-only log anywhere in src/ or the
balancing scripts. src/engine/replay.ts is the closest thing that exists — one
ReplayFrame {dt, input} per simulated frame plus gameplaySeed, and by
construction it reproduces a run exactly — but it records inputs, not outcomes,
carries no damage/kill/loot event, lives in localStorage, and no balancing tool
reads it. The proposed event log is its outcome-side counterpart.
Gating. this.telemetryEnabled = (PLAYER_STATS_ENABLED || isTestHooksActive()) && !this.ablated("telemetry")
(engine.ts). Changed 2026-08-23: PLAYER_STATS_ENABLED now ships true,
so shipped play records the same telemetry the bot always did — this section
used to say "in shipped play nothing is recorded at all", which is no longer
true and matters for anyone reasoning about what a player's browser is doing.
?ablate=telemetry is the off switch; under it teamTelemetry and every
PlayerState.telemetry stay undefined and every call site is a guarded no-op,
which is both the escape hatch and what keeps that path testable.
Existing per-frame cost when telemetry is on — worth writing down so a future reader does not attribute it to the event log:
| Cost | Site |
|---|---|
updateMinHealth + updateTelemetryPerFrame, per living player |
engine.ts |
Full this.enemies scan for peakAggroedCount / combatTimeSec |
engine.ts |
Two Object.values(...).reduce(...) over weaponTallies in buildStats(), which runs once per rendered frame |
engine.ts, called from render() |
PLAYER_STATS_ENABLED's own doc comment (playerStats.ts) used to record that
recording on every real playthrough "measurably slowed gameplay down". That claim was
retracted at the source on 2026-08-23: no number was ever taken for it, and two A/Bs
since have failed to reproduce a cost — most recently +0.025ms busy against a 0.16ms
calibrated floor, with the level-end derivation itself at 0.00027ms. The flag ships
true.
The list above is therefore a cost inventory rather than evidence of a measured problem.
It is still the reason §3 adds no per-frame work whatsoever: the path is cheap because
nobody has widened it, and the ~20 record* call sites are what a widening would
multiply.
There is no single constants module, but since 2026-08-19 there is a single
aggregator. SIMULATION_BALANCE (engine.ts) feeds computeBalanceHash and now
folds in whole tables — WEAPONS, COMBAT_BALANCE and TRAP_BALANCE — rather than
a hand-picked list of scalars. The list of five constants this section used to name
as uncovered (ELITE_DAMAGE_MULTIPLIER, EDGE_CASE_SPEED_MULTIPLIER, SPIKE_DPS,
MINE_DAMAGE_FALLOFF_FLOOR, PROJECTILE_SPEED) is covered by those tables today, and
simulationBalanceCoverage.test.ts fails on any value-carrying export that is missing —
so "wholesale" is enforced rather than claimed. See
Design Decisions's "Simulation constants are hashed as tables, not as picked scalars".
| # | name |
Dmg/pellet | Pellets | spreadPx |
Fire interval | Auto | Ammo/shot | Pool | Magazine / reload | Kind | Range limit | Lifesteal |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | echo pistol | 22 | 1 | 0 | 0.15 s | — | 1 | bullets | 9 / 1.1 s | hitscan | — | — |
| 1 | Regex Shotgun | 25 | 7 | 70 | 0.85 s | — | 1 | shells | 2 / 1.2 s | hitscan | — | — |
| 2 | SIGKILL Knife | 40 | 1 | 0 | (default 0.15 s) | — | 0 | none | — | melee | meleeRange 1.5 |
1 |
| 3 | gdb | 16 | 1 | 0 | 0.09 s | yes | 1 | smg | 45 / 2.0 s | hitscan | — | — |
| 4 | ghidra | 150 | 1 | 0 | 1.1 s | — | 1 | rockets | 1 / 1.6 s | projectile | blast 2.6 | — |
| 5 | Friday Hotfix | 8 | 6 | 45 | 0.1 s | yes | 2.5 | gas | none | hitscan | full to 2.5, zero at maxRange 6.5 |
— |
| 6 | Toolchain | 80 | 1 | 0 | 0.35 s | yes | 0 | none | — | melee | meleeRange 1.5 |
3 |
gdb overrides maxConeDeviationPx to 20 (its own WEAPONS entry); everything else
uses the shared MAX_CONE_DEVIATION_PX = 38 (engine.ts).
Magazines and reloading shipped 2026-08-15; the table's last two columns are
them. Every ranged weapon but Friday Hotfix draws from a magazine that empties into a
reload (auto on dry, or R/gamepad X early), and a reload only moves ammo between
the reserve and the magazine — nothing is ever lost, so the pools below still bound
total output exactly as they did. A weapon with no magazineSize never reloads and
never blocks its own trigger (updateFiring treats "no magazine" and "magazine full"
identically). Switching weapons is still instant (consumeWeaponRequest → equip,
no draw timer) and cancels a reload in progress. Nothing caps a pool while a level is being played —
only the IDKFA cheat's CHEAT_MAX_AMMO = 999 bounds anything mid-level, which is why
the status bar's ammo table shows one number per row rather than current/max. There is
one ceiling, and it applies only at the moment of entering a level: see
CARRYOVER_CAP_MULTIPLE below.
Hitscan weapons have no damage falloff. Damage is flat per pellet at any range.
What falls off is accuracy — the Cone of Fire (fire() in engine.ts):
rangeFraction = min(1, zBuffer[col] / FOG_FAR) // FOG_FAR = 14 tiles
deviation = (rng()*2 - 1) * rangeFraction**3 * (weapon.maxConeDeviationPx ?? 38)
Cubic in range, in screen pixels, resolved against the per-column z-buffer. Melee is exempt. This is the single reason hit probability is not cleanly computable offline — see §4.4.
Ghidra is the one projectile: ROCKET_SPEED = 18 t/s, ROCKET_BLAST_RADIUS = 2.6,
damage 150 * max(0.3, 1 - d/2.6), and it damages the firer too (rockets.ts).
Starting ammo (startingAmmo() in ammo.ts), before carryover:
bullets = max(28, round(shotsToClear * 1.7 + enemies.length * 2.5) + 10)
where shotsToClear = Σ ceil(enemy.maxHp / 22) // 22 = pistol damage
shells = 12 rockets = 4 smg = 40 gas = 40
shells is the one flat reserve that is not gated on owning its weapon — the
shotgun is a starting weapon. Its size was derived from what the old shared pool
afforded at 4 bullets a pull, not from a capture; drop amounts are
SHELLS_DROP_AMOUNT = 2 (one full magazine) and ELITE_SHELLS_DROP_AMOUNT = 8.
Note this is computed after difficulty HP scaling (createPlayerState runs after
the roster rescale in the RaycasterEngine constructor), so it scales with
difficulty for free. It also means level 1 already ships with a guaranteed
bullets-only clear margin: the × 1.7 applies to the perfect-accuracy shot count,
the per-enemy ceil rounds each enemy's cost up on top of that, and the miss buffer
and +10 add more — so the real starting ratio lands closer to 2.3–2.4× for a
typical roster, not 1.7×. That is worth knowing before reading any level-1 economy
number as evidence of anything.
From level 2 on, carryover replaces starting ammo (EngineCarryover, applied in the
RaycasterEngine constructor), so the economy is a campaign running balance, not a
per-level independent quantity. The solver must model it that way or every level after
the first will read as far poorer than it plays.
It no longer overrides it entirely, though (d299509, 2026-08-21). Each pool is
clamped on entry: Math.min(carryover[pool], startingAmmo[pool] * CARRYOVER_CAP_MULTIPLE)
with CARRYOVER_CAP_MULTIPLE = 3 (ammo.ts), applied at engine.ts's constructor. Two
consequences for the solver. startingAmmo is now load-bearing on every level rather
than only the first — the ceiling is derived from it, so a level whose fresh-start
formula is small also has a small ceiling. And a pool the level cannot supply fresh
(rockets before ghidra is owned) is capped against its own flat reserve rather than
against zero. --carryover-cap <n> models it offline.
Three archetypes, distinguished by two booleans on Enemy, not by a stat table.
Only HP is computed from the repo; every other combat stat is a shared constant.
| Regular | Elite | Edge Case | |
|---|---|---|---|
| Trigger | function/method entity | same, complexity ≥ 40 |
corridor dressing, and the tail of a switch-heavy entity's pack |
| Pack size | 1 + floor(c/5) |
ceil(min(c*35*2, 8*350) / 350), ≤ 8 |
1 |
| HP | max(35, round(c*35/count)), anchor-weighted in a guarded (private/protected) room |
min(c*35*2, 2800) split across the pack, each member ≤ 350 |
25–35 uniform |
| Melee dmg | 15 | 30 | 6 |
| Bolt dmg | 12 | 48 | 2.4 |
| Bolt speed | 5 | 3.6 | 7.5 |
| Bolt spread | 0° | 0° | 7° |
| Melee interval | 0.8 s | 0.8 s | 0.8 s |
| Ranged interval | 1.2–2.6 s uniform | 2.4–5.2 s | 0.6–1.3 s |
| Aggro radius | 7.5 | 7.5 | 7.5 |
| Melee radius | 0.5 | 0.5 | 0.5 |
| Ranged range | 8 | 8 | 8 |
| Chase speed | 1.7 | 1.7 | 3.74 |
| Roam speed | 0.8 | 0.8 | 1.76 |
Ranged attacks are per-archetype since 2026-08-19 (ENEMY_WEAPONS in
combatConstants.ts): each archetype's bolt carries its own damage, speed, cooldown
window, aim spread and palette (magenta / orange / cyan), sized so mean ranged DPS
per archetype is unchanged — an Elite's shell is 2× damage on a 2× cooldown, an Edge
Case's spray 0.5× on a 0.5× cooldown. That DPS-neutrality is deliberate: it is what let
COMPLEXITY_PER_EXTRA_ENEMY move the next day as the only difficulty lever in its
change.
Sources: combatConstants.ts (AGGRO_RADIUS, MOVEMENT_SPEED, RANGED_RANGE,
FIRE_COOLDOWN_MIN/MAX, ROAM_SPEED, ATTACK_RADIUS, ATTACK_COOLDOWN,
ATTACK_DAMAGE, ELITE_DAMAGE_MULTIPLIER, EDGE_CASE_SPEED_MULTIPLIER,
EDGE_CASE_DAMAGE_MULTIPLIER, PROJECTILE_SPEED/DAMAGE/RADIUS, ENEMY_WEAPONS),
enemies.ts (the generator-side pack maths).
Those constants are no longer module-private. They moved out of enemyAi.ts into
combatConstants.ts and are exported as COMBAT_BALANCE, which is what closed the
solver blocker this section used to record for incoming-DPS and threat-score metrics.
Aggro needs proximity and line of sight (enemyAi.ts) and is sticky;
being shot sets it unconditionally, bypassing LOS (damageEnemy in engine.ts). There is no
telegraph on either attack. Bolts do not home.
MAX_HEALTH 100 (in combatConstants.ts since the 2026-08-19 move, re-exported
through COMBAT_BALANCE), MOVE_SPEED 3.2, SPRINT_MULTIPLIER 2.0, ROT_SPEED
2.6 rad/s, collision radius 0.2 (player.ts).
Strafe is the same speed as forward (player.ts) and reversing is too.
Armour is swap, starts at 0, caps at MAX_SWAP = 100 (loot.ts), and
absorbs 1:1 before health (damage() in engine.ts). There is no percentage
reduction, no resistance and no invulnerability frames — so effective HP is exactly
health + swap.
No in-level respawn exists. Death ends a single-player run; in coop a dead player
revives at the next level transition at REVIVE_HEALTH = 50 (engine.ts).
Environmental damage: HAZARD_DPS 18 (engine.ts), SPIKE_DPS 20
(traps.ts), mine 32 * max(0.35, 1 - d/2.4) (traps.ts), rocket
self-splash at full strength.
hp |
damage |
ammoDropRate |
enemyAimSpreadDeg |
rollbacks | |
|---|---|---|---|---|---|
| easy | 0.7 | 0.85 | 1.3 | 10° | 2 |
| normal | 1 | 1 | 1 | 4° | 1 |
| hard | 1.5 | 1.5 | 0.7 | 0° | 0 |
The rollback budget is a separate table (ROLLBACKS_BY_DIFFICULTY, same file), not a
member of DIFFICULTY_MULTIPLIERS, because it is a retry count consumed by the app shell
rather than a multiplier applied to a simulation quantity — the engine never reads it. It
is listed here because it runs in the same direction as every other axis and therefore
belongs in any difficulty comparison. The bot opts out of it entirely (see
codeenstein-rollbacks-disabled), so a telemetry run's death count is unaffected.
damage covers enemy melee and bolts only — not traps, hazards or rocket
self-damage. ammoDropRate scales both dropped and pre-placed pickup amounts
(scaledLootAmount, engine.ts), which matters for the solver: the
pre-placed budget is difficulty-dependent too.
Multiplayer adds an Elite-only rescale, hp × (1 + 0.5·(players−1)) and
damage × (1 + 0.25·(players−1)) (multiplayerScaling.ts), unbounded in
player count.
A regular kill fires up to four independent rolls (the kill branch of damageEnemy,
engine.ts):
| Roll | Guaranteed? | Rate | Yields |
|---|---|---|---|
| Health top-up | yes, if health < MAX_HEALTH |
100% | max(1, round(20 × maxHp / 100)) — see HEALTH_SCALE_REFERENCE_HP |
| Weighted ammo/swap | no | 80% (REGULAR_KILL_NO_DROP_CHANCE = 0.2) |
one kind from the table below |
| Miss-consolation Toolchain | only on the 20% miss | 5% of misses = 1% of kills | Toolchain, if level ≥ 4 and unowned |
| Bonus weapon | independent, stacks | 1% (NORMAL_KILL_WEAPON_DROP_CHANCE) |
a random still-locked index from [3,4,5] |
An Elite kill takes a completely separate path (lootApply.ts) — the
if (enemy.elite) dropEliteLoot(...) else {...} split is total, so an Elite never
rolls the miss chance:
| Roll | Rate | Yields |
|---|---|---|
| Guaranteed drop | 100% | health 50; or, at full health, a 50/50 between bullets 18 and swap 30 |
| Bonus weapon | 60% (ELITE_BONUS_WEAPON_DROP_CHANCE) |
a still-locked index, plus Toolchain if level ≥ 4 |
Loot-kind weights (loot.ts), re-normalised after filtering:
| Kind | base (easy/hard) | normal | bonus level |
|---|---|---|---|
| bullets | 28 | 32 | 17 |
| shells | 12 | 14 | 7 |
| smg | 18 | 20 | 20 |
| gas | 18 | 20 | 20 |
| rockets | 10 | 12 | 20 |
| health | 16 | 11 | 20 |
| swap | 16 | 11 | 16 |
Filters (loot.ts): rockets/smg/gas are removed entirely unless the
matching weapon is owned, and health is always removed because the engine passes
healthHandledSeparately = true — health is its own unconditional check now.
Base drop amounts (loot.ts): bullets 4, shells 2, rockets 1, smg 21, gas 21,
health 20, swap 11; elite fallbacks bullets 18, shells 8, rockets 6, smg 80, gas 80,
swap 30, health 50.
Those are bases, not what actually drops (f020902/a65948c, 2026-08-21). Every
ammo kind is scaled by the dead enemy's health at the drop site —
max(1, round(base × maxHp / AMMO_SCALE_REFERENCE_HP)) with AMMO_SCALE_REFERENCE_HP = 88,
the corpus mean — and the health top-up by HEALTH_SCALE_REFERENCE_HP = 100. swap is
deliberately the one exception and stays flat, because it was priced that way. The
floor of 1 exists so a rockets drop, whose base is already 1, cannot round away on a
small enemy. --hp-scaled-drops and --hp-scaled-health model both offline; both are on
by default in the solver.
Three things the brief asked about, answered explicitly:
- The kill method does not affect drops. The drop block reads only
enemy.elite,shooter.health,shooter.ownedWeapons,map.bonusLevel,difficultyLevelandthis.rng.weaponIndex,forcedMeleeandlifestealare consumed earlier, for telemetry and the lifesteal heal, and never reach it. Killing with the knife, a rocket or the flamethrower yields identical distributions. - Leftover magazine contents do not carry over. Magazines exist (they shipped 2026-08-15, tabulated above), but they are the player's — enemies carry no inventory at all, so a drop's contents come entirely from the weighted roll.
- Edge Cases take the regular path with no special-casing, and until 2026-08-21
that meant a 10 HP Edge Case had exactly the same expected drop as a 500 HP regular
enemy — the single most suspicious line in this section when it was written, and what
§2.2's self-sustain ratio was built to quantify. Both halves of it have since
moved: Edge Cases are 25-35 HP, and the drop amount is now proportional to
maxHp, so the roll is still uniform across archetypes but the payout is not. The path really is still shared — the scaling happens at the drop site, not in the roll.
| Source | Contents | Cap / rate |
|---|---|---|
Scattered ammo (pickups.ts) |
bullets 11, or rockets 3 at 30% if ghidra owned | Bernoulli(0.22) per non-spawn room; bonus levels 0.65 and ×1.5 |
Secret rooms (secretRooms.ts) |
health 60, rockets 4, swap 40, one weapon | MAX_SECRET_ROOMS = 5 |
Vendor depots (vendorDepots.ts) |
bullets 12, rockets 3, smg 30, gas 30 — each gated on ownership | floor(importCount/4), max 4 |
Exception zones (exceptionZones.ts) |
catch: health 45, swap 35; finally: bullets 14, rockets 3 | MAX_EXCEPTION_ZONES = 3 |
All amounts are scaled by ammoDropRate at collection, not at placement.
Note what pre-placed ammo scales with: room count, i.e. entity count. Not enemy HP, not enemy count, not walkable area. See §7 finding 5.
A file is a level. Folders contribute only ordering (flattenParsableFiles, main.ts flattenParsableFiles). Within a file, one room per CodeEntity, in startLine
order — but only function/method entities produce enemies. Classes, interfaces
and traits get a room and no enemy; globals become acid pools.
Complexity is 1 + decisionPoints + smellBonus (genericParser.ts), where the
smell bonus adds 2 per parameter beyond 5 and 3 per nesting level beyond 3
(astUtils.ts).
The enemy mapping, condensed from spawnEnemies (map/generation/enemies.ts):
const complexity = Math.max(1, room.entity.complexityScore);
const elite = complexity >= ELITE_COMPLEXITY_THRESHOLD; // 40
const eliteTotal = Math.min(
complexity * HP_PER_COMPLEXITY * ELITE_HP_MULTIPLIER, // 35 * 2, raised from 25 on 2026-08-21
ELITE_MAX_MEMBERS * ELITE_MEMBER_HP_CAP, // 8 * 350 = 2800
);
const count = elite
? Math.ceil(eliteTotal / ELITE_MEMBER_HP_CAP) // 350
: 1 + Math.floor(complexity / COMPLEXITY_PER_EXTRA_ENEMY); // 5, was 10 until 2026-08-19
// A guarded (private/protected) room weights the anchor; Elite packs are exempt.
const anchorWeight = !elite && isLockableRoom(room, roomIndex) ? GUARD_ANCHOR_WEIGHT : 1;
const [anchorHp, memberHp] = packHitPoints(elite ? eliteTotal : total, count, anchorWeight);
// every member is additionally clamped to ELITE_MEMBER_HP_CAPThere is no clamp, cap or normalisation on enemy HP. The only clamp() near
complexity is on room geometry (geometry.ts, side capped at 18 tiles),
and the only repo-size normalisation in the whole generator is map dimension
(ROOM_SPREAD × √roomArea + ROCK_RESERVE, clamped 48..160,
mapGenerator.ts) — geometry only. See §7 finding 1 for what that means.
Walkable area is computable but awkward: countWalkableTiles (engine.ts) is
module-private, and staticLevelAnalysis.mjs already keeps a hand-copied
mirror of it and isWalkableTile.
Map generation is fully deterministic and content-addressed. mulberry32(seedFrom(parsed))
(mapGenerator.ts), where seedFrom is FNV-1a over
language:linesOfCode:kind/name/complexityScore,… (seed.ts). Zero
Math.random() exists under src/map/ or src/parser/. One shared stream threaded
through every subsystem in a fixed order, so the order of generation passes is
part of the contract.
Two caveats worth recording: the seed signature omits nestingDepth and comment
text, so it is not a complete fingerprint of what generation actually reads (two
files can share a seed and still generate different maps — determinism holds, the
seed just is not a complete identifier); and GenerateOptions (weapon ownership,
missingWeaponIndices, bonusLevel, maxPlayers) also feed generation, so the
same file at a different campaign position produces different pickups and grid.
Gameplay randomness is seeded but not pinnable. Loot rolls, enemy AI timing,
weapon spread and enemy aim all draw from one createResumablePrng(gameplaySeed)
stream (the engine's own rng field comment, engine.ts). So drops are deterministic given a seed — but the
seed comes from randomSeed() (main.ts), the one sanctioned Math.random()
(prng.ts), and there is no way to supply one: no UI field, no CLI flag,
no env var in any balancing runner. See §7 finding 2.
One direct Math.random() sits on the combat path — baseBloodCount inside damageEnemy
(engine.ts) — sizing a blood-particle burst. It touches no
simulation state, but it is the only one in engine.ts, whose own rng field comment
says "Never Math.random() directly". Recorded here so a future determinism
audit does not have to rediscover that it is benign.
Marker key: A = analytic (computable offline from level data + constants), E = empirical (needs play), A+E = both, and the two should agree — where they disagree, that gap is itself the finding.
Nothing is listed for completeness. Every entry states what you would change if it came out bad; anything without such an answer is in §2.6 instead.
| Metric | Kind | Definition / computation | Needs | Bad when |
|---|---|---|---|---|
ttkAnalytic[w][arch] |
A | (gaps − reloads) × fireIntervalSec + reloads × max(fireIntervalSec, reloadSec). The reload term was added 2026-08-21, six days late; max rather than + because engine.ts ticks the fire cooldown above the reload gate, so the two run concurrently. |
constants + roster | > 8 s vs a normal enemy (matches the existing NORMAL_TTK_HIGH_SEC) |
ttkObserved[arch] |
E | aggro→death window; already collected as avgTtkByCategory |
kill |
≫ analytic (means the weapon is unusable at real range, not merely slow) |
shotsToKill[w][arch] |
A | ceil(hp / (damagePerPellet × pellets)), pellets=1 for ghidra |
constants + roster | any weapon needing > total obtainable ammo for that pool |
overkill[w] |
E | mean −hpAfter on the killing blow ÷ damage dealt |
damageDealt{hpBefore,hpAfter} |
> 30% — the weapon's granularity is wrong for this roster |
dpsSustained[w] |
A | damagePerPellet × pellets / fireIntervalSec |
constants | — (a comparison axis, not a threshold) |
damagePerAmmo[w] |
A | damagePerPellet × pellets / ammoPerShot; ∞ for melee |
constants | a weapon both below the pistol and slower than it has no reason to be fed |
pelletHitRate[w] |
E | pellets connected ÷ pellets fired | shot{pellets}, hit |
— |
triggerHitRate[w] |
E | trigger-pulls landing ≥1 pellet ÷ trigger-pulls | shot, hit |
— |
hitRateByDistance[w][bucket] |
E | hit/shot bucketed by dist (0–2, 2–4, 4–7, 7–10, 10–14 tiles) |
shot{dist}, hit{dist} |
a weapon whose rate collapses inside its own usable range |
switchesPerEngagement |
E | weaponSwitch count between first aggro and last kill of a fight |
weaponSwitch, kill |
— (bot-policy signal, see §2.6 on why the cost half is cut) |
share.shots/damage/kills[w] |
E | per-weapon fraction of each total | shot, damageDealt, kill |
< 2% of kills for an owned weapon = dead content, reported as a first-class line, not a footnote |
damagePerAmmo is the number that decides whether a weapon is worth feeding, and
it is worth precomputing here because the current values are not intuitive:
| Weapon | dmg/trigger | ammo/shot | dmg per ammo unit | dps |
|---|---|---|---|---|
| echo pistol | 22 | 1 | 22.0 | 147 |
| Regex Shotgun | 175 (7×25) | 1 | 175.0 | 206 |
| gdb | 16 | 1 | 16.0 | 178 |
| ghidra | 150 | 1 | 150.0 | 136 |
| Friday Hotfix | 48 (6×8) | 2.5 | 19.2 | 480 |
Read against §1.3's drop weights this already predicts something the solver should confirm per level: every weapon now draws its own pool, so efficiency no longer decides which pool to spend — it decides how far a given pool goes. gdb is the least efficient per round in the game by a wide margin, which is exactly why its own pool is fed by 21-round drops where bullets drop 4. Its 16 damage a round also sits deliberately between the pistol's 22 and the shotgun's 175 a shell: enough rate of fire to out-damage the pistol per second (178 dps against 147), paid for by needing far more rounds to do it. These are the perfect-accuracy figures — §4.4.
Every quantity here is reported three ways: pre-placed only, dropped only, and
combined. Pre-placed is a fixed budget the generator controls; dropped is a
feedback loop that scales with how much you fight. Summing them hides which knob is
wrong, and this game has already been burned by exactly that: a 450-run campaign
found dropped loot was 93–100% of everything actually consumed, every resource
type, every combo (see LOOT_WEIGHTS in loot.ts), which is what motivated the ~30% drop
cut and REGULAR_KILL_NO_DROP_CHANCE.
| Metric | Kind | Definition / computation | Needs | Bad when |
|---|---|---|---|---|
clearRatio[source] |
A | (Σ damage obtainable from ammo of that source + starting loadout) ÷ Σ enemy HP, per weapon and overall. Ammo → damage via damagePerAmmo. |
constants + roster + pickup list + drop model | pre-placed-only < 1.0 means the level is unclearable without farming; combined < 1.2 means no margin for a miss |
selfSustain[arch] |
A | expected damage-worth of one kill's drop ÷ damage needed to kill it. Expected drop = 0.8 × Σ(weightᵢ/Σweight × amountᵢ × damagePerAmmoᵢ), plus the guaranteed health grant valued separately. |
drop tables + weights + constants | > 1.0 = fighting is free ammo and the economy has no floor; ≪ 1.0 = every fight is a net loss |
incomeVsExpenditure |
E | cumulative ammo gained and spent over level time, both curves, split by source | lootCollected, shot |
income curve flat while expenditure climbs — starvation that a total would hide |
starvationEvents |
E | intervals at zero ammo in the preferred pool; count, duration, and the fallback used. Attributed: "no pre-placed nearby" vs "drops came up empty" via the last lootDropped/lootCollected in the window |
shot, lootCollected, lootDropped |
any forced-melee stretch > 10 s |
overflow[source] |
E | amount collected while already at cap. Only health and swap can overflow — ammo has no cap | lootCollected{granted,wasted} |
pre-placed overflow is a placement bug; drop overflow is a rate problem |
uncollectedPrePlaced |
E | pickups still collected === false at level end, valued in damage |
levelStart{prePlaced}, levelEnd |
the generator's effective budget is lower than its nominal one — the gap is the number to tune against |
unrealisedDrops |
A+E | expected drop value of enemies left alive at level end | levelEnd{enemiesAlive} + drop model |
with the line above, gives nominal-vs-actual economy |
scarcity[pool][source] |
A+E | per-pool obtainable damage ÷ damage the pool is asked to deliver | as clearRatio |
tells whether a gun is starved by placement or by drop tables — different fixes |
relianceRatio |
E | share of consumed ammo that came from drops | lootCollected{source} |
> 0.9 means the level is only clearable by engaging optional enemies |
The origin split already half-exists: recordLootCollected(state, origin, kind, amount)
(engine.ts) distinguishes dynamic from static. What it does not carry is
which archetype dropped it, and without that selfSustain cannot be measured
empirically at all — only predicted. That is the one field that must be right from
day one (§3.3).
| Metric | Kind | Definition / computation | Needs | Bad when |
|---|---|---|---|---|
incomingDps[arch] |
A | melee dmg/0.8 at contact + ranged dmg/1.9 (mean of the 1.2–2.6 s window) within 8 tiles, × difficulty damage |
enemyAi.ts constants (currently private) |
— |
survivalWindow(N) |
A | (MAX_HEALTH + swap) ÷ (N × incomingDps) for N = 1…peak observed |
as above | < 3 s at the level's observed peakSimultaneousAggroed |
healthIncomeVsTaken[source] |
A+E | health granted vs damage taken, split pre-placed / dropped | lootCollected, damageTaken |
dropped health ≥ damage taken means fighting is self-sustaining and the level has no attrition |
damageTakenByArchetype |
E | damage attributed to the archetype that dealt it — implemented 2026-08-09 as damageTakenByAttacker (eventMetrics.mjs), which also breaks it down per enemy (lvl:eid), turning "died on level 12" into "one Elite dealt 93% of the damage". Each event's amt is split across its attackers in proportion to their by[].amt shares, since those are pre-multiplier contributions rather than magnitudes |
damageTaken{archetype} |
an Edge Case tier out-damaging regulars means the "harmless nuisance" design intent is not holding |
killRateByHpBand |
E | spawned vs killed, banded by max HP, with median TTK per band — the denominator comes from levelStart.enemies[] so it is exact rather than inferred. Added 2026-08-09; it is what showed 1,332 Elites spawning for 2 kills. Read the contrast between bands, never one rate alone: most of a roster is walked past, so even a healthy band sits well under 100%, and an empty band is itself a finding |
levelStart, kill |
a populated band at ~0% is content that cannot be fought, not content that is hard |
deaths, nearDeath, timeBelowThreshold |
E | deaths per level; dips below 25% that recovered; time below it | damageTaken, playerDeath |
zero near-deaths across a campaign = no stakes (this is exactly how Easy's damage floor got raised) |
damageBySource exists today with six sources (telemetry.ts) but no archetype
attribution — enemyMelee merges an Elite's 30-damage bite with an Edge Case's 6.
The archetype split is new and is what makes threat scoring checkable.
| Metric | Kind | Definition / computation | Needs | Bad when |
|---|---|---|---|---|
enemyCount, enemyHpTotal, enemyDpsTotal |
A | roster sums | roster + constants | — |
| the same three per walkable tile | A | ÷ countWalkableTiles |
roster + grid | > 1.5× the campaign mean (matches the existing DENSITY_OUTLIER_MULTIPLIER) |
threatScore[arch] |
A | incomingDps × √hp × rangeFactor × speedFactor, normalised so a regular enemy = 1.0. rangeFactor = ranged reach ÷ 8; speedFactor = chase speed ÷ 1.7 |
constants | lets archetypes be ranked at all — today Elite-vs-EdgeCase is a judgement call |
combatVsExploration, levelTimeSec, distanceTraveled |
E | already collected | existing | — |
backtracking |
E | tiles walked over an already-visited tile ÷ total | levelEnd + existing distance |
> 0.4 means the layout is asking for re-walks |
difficultyCurve |
A+E | per-level clearRatio and enemyDpsTotal across the campaign, as a series |
per-level solver output | a spike > 2× its neighbours |
This is the group that only matters because levels come from source code, and it is the reason the solver exists.
| Metric | Kind | Definition / computation | Needs | Bad when |
|---|---|---|---|---|
complexityToHpCurve |
A | every entity's (complexityScore, resulting HP, archetype), plotted |
parse + roster | see below |
hpOutliers |
A | entities whose single-enemy HP exceeds the level's total obtainable damage | roster + clearRatio |
any hit is a hard failure — that enemy cannot be killed with everything on the level |
clampEffectiveness |
A | how many entities hit a clamp, and which | parse + generator | a rising share hitting ELITE_MEMBER_HP_CAP means the split is doing the work; zero would mean it is inert |
perLevelBudget |
A | one line per level: enemy HP total, enemy DPS total, ammo damage (pre/drop/combined), health (pre/drop/combined), ratios | all of the above | the tuning table |
corpusDistribution |
A | the same budget report over N repos of varying size and language | corpus | shows the spread rather than one sample |
The curve is the point, and it is now bounded on both branches. Regular packs
self-limit by construction: per-member HP is 35c / (1 + ⌊c/5⌋)
(HP_PER_COMPLEXITY = 35, COMPLEXITY_PER_EXTRA_ENEMY = 5), which asymptotes to 175 as
complexity rises. Elites keep the doubled room budget but no longer put it in one
body — it is split across up to ELITE_MAX_MEMBERS = 8 members each capped at
ELITE_MEMBER_HP_CAP = 350, so a room tops out at 2,800 base and 4,200 on Hard however
complex the function is. Plotting both on one axis is still worth doing, but what it
shows now is a ceiling rather than a ray.
This paragraph described a ray until 2026-08-08, and the history is worth keeping
because it is what the metric was built for. Elite HP was complexity × 25 × 2 in a
single body, linear and unbounded: at c = 39 an entity spawned four enemies of 244 HP
(976 total) and at c = 40 one enemy of 2000 HP — one point of complexity doubling
the entity's total HP and multiplying its single-enemy HP by 8.2×, into a target that
could not be split, kited or partially cleared. vim's complexity-672 function produced
33,600 HP in one body. §7.1 has the measurement that settled it (1,332 Elites spawned,
2 died) and the reasoning behind the two constants.
Standard shooter metrics that do not apply to this game, listed so nobody adds them back for completeness:
| Excluded | Why |
|---|---|
No longer excluded. True until 2026-08-15, when magazines shipped; the solver kept saying it until 2026-08-21, understating a pistol kill on a 504 HP Elite as 3.30s against a real 5.20s. Now charged — see timeToKill. |
|
| Burst DPS vs sustained DPS | Still excluded, but no longer for the stated reason: magazines exist now, so the two genuinely differ. What the solver models is a single sustained kill, which is where its reloads term lands; a burst/sustained split needs an engagement model it does not have. |
| Weapon draw / switch cost | Switching is instant. Switch frequency is kept in §2.1 as a bot-policy signal; the cost half would measure a constant zero. |
| Hitscan damage falloff | Does not exist. Damage is flat with range; accuracy falls off cubically, and §2.1 measures that instead. |
| Armour as a separate pool with its own curve | swap absorbs 1:1 with no reduction, so effective HP is exactly health + swap and a second model adds nothing. |
| Headshot / hit-location stats | There are no hit locations — a pellet either connects with a sprite column or does not. |
| Movement/aim heatmaps | The player here is a bot with a route planner; its position distribution describes routePlanner.mjs, not the level. |
NDJSON — one JSON object per line, appended. Against the alternatives:
- vs. a single JSON document (what
balancing_telemetry.jsonis today, written by oneJSON.stringify(output, null, 2)at process end,run-balancing-telemetry.mjs): a killed campaign loses that file entirely. An event log must survive a SIGKILL — and the campaign orchestrator does SIGKILL invocations, by design (laneOrchestrator.mjs's watchdog). With NDJSON a truncated final line is discarded and everything before it is intact; a truncated JSON array is unparseable. - vs. CSV: events are heterogeneous — a
shotand alevelEndshare almost no fields. A single CSV needs a wide sparse header, and one file per event type loses the interleaved ordering that makes timeline metrics possible. - Dependency cost: zero.
JSON.parseper line,fs.appendFileSyncto write. No parser, no schema library. That matters givendoc/dev/decisions.md's Dependency Minimalism section.
perf_runs/*.json already sets the precedent for keeping raw arrays (rawDeltas,
one float per frame) next to the summaries computed from them.
One file per combo, not per run: <dir>/<profile>-<difficulty>.ndjson, where
<dir> is whatever CODEENSTEIN_TELEMETRY_EVENT_LOG names. writeEventBatches
derives the path from profile and difficulty only, and appendEvents never
truncates, so every attempt of a combo — and every invocation covering that
combo — appends to the same file. That is what makes a chunked or resumed sweep
work; runs stay distinguishable because rid is unique per invocation.
Choose that directory to match .gitignore. An earlier version of this
section said balancing_events/ was "gitignored alongside the existing
balancing_runs*/ entries" — it is not. .gitignore covers
balancing_telemetry.json, balancing_runs/, balancing_runs_*/ and
balancing_corpus/ only, so a capture written to balancing_events/ shows up as
hundreds of megabytes of untracked NDJSON. Use a balancing_runs_*/ prefix, which
also matches the archive convention used for pre-change snapshots.
Every line carries exactly five envelope fields, kept short because they repeat on every record:
| Field | Meaning |
|---|---|
v |
schema version integer — bump on any breaking field change |
e |
event type |
sid |
session id (one process invocation) |
rid |
run id (one campaign attempt) |
lvl |
campaign level index |
t |
levelTime in seconds, monotonic within a level |
The level fingerprint does not repeat per line — it lives once in levelStart,
and it reuses identifiers that already exist rather than inventing a new one:
astHash (parsed source + campaign name), balanceHash
(computeBalanceHash(map, SIMULATION_BALANCE), balanceHash.ts — the enemy
roster plus the simulation constants in force) and seed (the gameplay seed).
Together these answer "same repo, same commit, same constants, same universe?"
without a new mechanism, and balanceHash in particular already exists precisely
to catch the case where a constant moved but no source byte did.
Fields marked † exist only to make a named metric computable; if that metric is dropped, the field goes with it.
levelStart — once per level. Makes the log self-contained, so uncollected-loot
and unrealised-drop analysis needs no re-run of the generator.
file, astHash, balanceHash, seed, difficulty, campaignLevelIndex,
walkableTiles†, // enemy/HP per unit area (§2.4)
ownedWeapons[], startHealth, startSwap, startAmmo{bullets,rockets,smg,gas},
enemies[]: {eid, arch, maxHp, x, y} // roster; unrealised drops (§2.2)
prePlaced[]: {pid, kind, amount, x, y} // pre-placed budget (§2.2)
shot — one per trigger-pull.
w (weapon index), pellets†, ammoAfter, forcedMelee, dist† (range to the
crosshair target, or null).
pellets and dist exist for pelletHitRate and hitRateByDistance (§2.1); note
that separating this from hit is what fixes the >100% accuracy problem in §7.4.
hit — one per pellet that connects.
w, eid (or mine), dist†.
damageDealt — one per HP change on an enemy.
w, eid, arch, amt, hpBefore†, hpAfter†, splash†.
hpBefore/hpAfter exist solely for overkill (§2.1) — on the killing blow,
hpAfter is the pre-clamp negative value, which damageEnemy (engine.ts) currently
computes and immediately discards.
rocketDetonated — one per rocket blast (schema 3+).
w, x, y, direct, dist, enemiesHit, dmg, selfDmg, hits
({eid, arch, amt}[]).
Why it exists. A rocket is the one weapon that emits no shot-time hit
events at all: fire() returns at if (w.isRocket) before resolveShot runs,
so across the entire archive 0 of 3,248,140 hit events are ghidra. Every
rocket question therefore had to be simulated or dropped —
report:damage-model's ghidra rows read 100% miss on every capture on disk, and
the 2026-08-20 finding that the shotgun is the only weapon hitting 2+ enemies
(23.8% of its shots, against 0.0% for pistol and gdb) could not be evaluated for
rockets at all, because splash leaves no per-target trace.
hits is deliberately the same {eid, arch, amt}[] shape damageTaken.by uses,
so "how many enemies did one shot hit, and which archetypes" is the same query
for a rocket as for a shotgun blast.
direct is the field that cannot be reconstructed. It says an enemy stopped
the rocket rather than the wall behind it. Splash lands either way, so damage
dealt cannot distinguish the two — which is exactly the distinction the
2026-08-19 tunnelling analysis had to reproduce in a simulation because the log
could not answer it.
kill — eid, arch, maxHp, w, forcedMelee, aggroAt† (closes the TTK
window without needing the separate ttkRecords array).
damageTaken — src (the six existing DamageSource values), by† (a
{eid, arch, amt}[] per-attacker breakdown, null for traps/hazards/self-splash),
amt, healthAfter, swapAfter. by is what makes damageTakenByArchetype
(§2.3) possible at all.
Read by for attribution and amt for magnitude — they answer different
questions and do not have to agree. One event covers one summed application
of damage, because damage() is deliberately called once per player per frame:
swap absorbs 1:1 before health, so splitting a 30-point call into three 10-point
calls would change the absorption arithmetic. amt is therefore the difficulty-
scaled figure that actually landed, while by[].amt are the pre-multiplier
contributions that produced it.
Each by[].eid indexes the same roster levelStart records, so its arch is
cross-checkable rather than merely trusted — verify-event-log.mjs does exactly
that for damageDealt/kill already. arch survives as an always-null
field so a schema-1 reader finds every key it knew.
lootDropped — at spawn time, so drops that are never collected stay visible.
did, kind, amount, x, y, fromEid†, fromArch†, elite†.
lootCollected — did or pid, kind, source (preplaced | drop),
fromArch† (drops only), amount, granted†, wasted†.
sourceandfromArchare the two fields that must be right from day one. Withoutsourcethe pre-placed/dropped split of §2.2 is not reconstructible after the fact; withoutfromArchthe self-sustain ratio — the most useful number in the whole catalog — can only ever be predicted, never measured. Both are cheap: the origin split already exists in the collect path (recordLootCollected(state, origin, kind, amount),engine.ts), andpushLootDrop(drop, enemy)(engine.ts) already receives the dropping enemy, so the archetype is in scope with no new plumbing.
granted/wastedsplit the amount that actually applied from the amount lost to a cap — needed foroverflow(§2.2), and only ever nonzero for health and swap, since ammo has no cap.
weaponSwitch — from, to, reason (manual | granted | autoEquip).
weaponGranted — w, via (secret | eliteBonus | missChance |
normalBonus | forcedUnlock), duplicate† (whether it fell back to an ammo
top-up).
playerDeath — src, arch†, t.
levelEnd — outcome (cleared | died | abandoned), killCount, score,
healthEnd, swapEnd, ammoEnd{}, distanceTraveled,
enemiesAlive[]: {eid, arch, maxHp}†, prePlacedUncollected[]: {pid, kind, amount}†.
The two † arrays are the "nominal vs actual budget" sweep — together they give the
gap between what the generator placed and what the run actually had access to.
scripts/verify-event-log.mjs runs every consistency check over a log
directory and exits non-zero on failure, so this is repeatable rather than a
spot-check someone remembers to do. It is deliberately a script: "is the
telemetry right?" was asked three times, and two of those turned up a real
defect after an ad-hoc check had already passed. Ad-hoc checks answer the
question you thought to ask.
Three tiers, weakest to strongest, and the script says which is which:
- Structural — parseable lines, nothing dropped by the buffer cap, envelope complete on every record.
- Conservation —
hpBefore - amt == hpAfter; spawned enemies equal kills plus survivors per level-visit; every collected drop references a real spawn. Catches corruption, not a consistently wrong value. - Cross-site agreement — an enemy's archetype is emitted independently at
five call sites, and all five must match the roster
levelStartrecorded. Catches a mislabel at one site, not one shared by all.
Verified against injected faults, not just against good data. Three
mutations — an archetype relabelled on one lootDropped, seven HP added to one
hpAfter, one kill removed — are each caught and named:
lootDropped.fromArch agrees with the roster: 1 failure(s), e.g. {"eid":9, "said":"normal","roster":"edgeCase"}. A checker that has only ever seen
passing input is not known to check anything.
A check that never ran is printed as 0 rather than omitted, since a silently
absent row reads as a pass — targetArch only exists in logs captured after
that field was added.
Internal consistency cannot tell you the roster itself is right; every site agreeing on the same wrong answer looks identical to every site being correct. Two independent sources close that:
- The roster, against a separate runtime.
npm run balancing:budget --jsongenerates the same campaign through Node/esbuild rather than the browser. Measured on the demo campaign: the archetype-and-HP multiset is identical on all 12 captured levels, difficulty scaling included, computed by two different code paths. targetArch, against what was actually hit. For a single-pellet weapon the pellet flies dead-centre, so the crosshair archetype should be the archetype hit: measured 99.9% over 1,135 single-pellet shots and 100% over 749 multi-pellet ones. The single outlier is a gdb shot recorded against an Edge Case that landed on a regular — one frame of staleness, sincethis.targetis set during the previous frame's render and Edge Cases cross the reticle at 3.74 tiles/sec.- The counters, against the events. Shots, hits and kills match exactly for all seven weapons and every damage source matches to the unit — two independent recording paths in the same run.
Cross-checked against the aggregate counters over a 4-attempt Gamer/hard capture,
since the two are recorded by independent code in the same run: shots, hits and
kills match exactly for all seven weapons, and every damage source matches to the
unit. Internal conservation also holds — hpBefore - amt == hpAfter on all 2,649
damage records, spawned enemies equal kills plus survivors on all 64 level-visits,
all 528 kills pair with a lethal damage record, no drop was collected without a
matching spawn, and nothing was lost to the buffer cap.
One real data-loss bug was found this way, after the checks above had already
passed. Three stuck return paths in playRun returned before
pullLevelResult, so that level's buffered events were never drained and were
discarded with the page. The balanced levelStart/levelEnd counts hid it: a
lost level drops both records together, so the totals stay consistent. Proven in
both directions by forcing stuck runs with
CODEENSTEIN_TELEMETRY_TUNING='{"MAX_TICKS_PER_WAYPOINT":1}' — with the fix, two
stuck runs yield 2 levelStart, 11 shots, 2 kills and their loot; without it, no
event file is written at all. Fixed by drainEventsInto, which drains without
pulling a snapshot, since a stuck level has no usable snapshot and must not enter
levelSnapshots.
The bias mattered more than the volume: a stuck run's last level is exactly the level that caused the problem, so the events most worth reading were the ones being dropped. Anything about wedges or hard-level failures collected before this fix is missing its most interesting level.
Three gaps remain, all deliberate and all worth knowing before reading a number:
weaponSwitchandweaponGrantedare never emitted. They appear in the schema above and in §7's blocked-metrics table; nothing currently depends on them.- Multiplayer emits no events at all.
run-balancing-telemetry-multiplayer.mjsnever sets?eventLog=1and never drains, so the whole event stream is single-player only. The engine side would work unchanged — every emission point is shared — but each peer would record its own copy, so a reader would need to de-duplicate by roster id first. - A splash weapon emits no
hit.fire()returns before the pellet loop forisRocket, so a rocket's damage arrives later asdamageDealtfrom the blast and never as a per-pellet hit. Ghidra accordingly showed 17 shots, 15 damage records and 0 hits — a structural zero that an earlier version of the balance review printed as though it measured accuracy.shot.splashnow flags it andweaponUsagereturnsnullrather than a rate. Emitting hits from the blast would not fix it: one round can strike several enemies, which is the same100% class of error that separating
shotfromhitfixed for the shotgun. Judge splash weapons on damage and kill share.
Gate: the existing one, unchanged. Events are pushed behind the already-computed
this.telemetryEnabled boolean (engine.ts). No new flag, no new branch in
shipped play, and no possibility of this reaching a real player's session.
Every emission point is event-rate, not frame-rate. Nothing is added to
simulate(), render(), renderNormalFrame() (all engine.ts), renderScene, handleMovement, updateEnemyAi, collectLoot or any
other per-frame function. The emission sites are exactly:
| Event | Site |
|---|---|
shot |
fire() — engine.ts |
hit, damageDealt, kill |
damageEnemy() — engine.ts |
damageTaken, playerDeath |
damage() and killPlayer() — engine.ts |
lootDropped |
pushLootDrop() — engine.ts |
lootCollected |
both recordLootCollectedEvent call sites in collectLoot — the loot-drop branch (origin: "drop") and the static-pickup branch ("preplaced") |
weaponSwitch, weaponGranted |
the weapon-request handler and grantOrTopUpWeapon |
levelStart, levelEnd |
engine construction and endGame() — engine.ts |
So the render-loop-adjacent diff for this work should be empty. That is a
design property, not an aspiration — if a later step needs a per-frame sample,
that is the moment to stop and re-read playerStats.ts.
Buffer. A plain array on the engine, preallocated at construction only when
telemetry is enabled. Objects are pushed, not serialised — JSON.stringify happens
in Node, off the browser's clock entirely.
Flush. A new __codeensteinTestHooks.drainEvents() returns the buffer and
resets it, mirroring the existing getTelemetrySnapshot() hook (engine.ts).
The bot drains at every level boundary and once more at run end; Node
appendFileSyncs the lines. Draining at a level boundary — not on a timer — keeps
the flush off any hot path and bounds the buffer at one level's events.
On crash: everything up to the last drain is on disk. The current level's undrained tail is lost. That is the deliberate trade — the alternative, flushing per event across the Playwright boundary, would cost a round-trip per shot. A partially-written final line is discarded by the reader, which is the whole reason for NDJSON over a JSON array.
Volume sanity. A campaign level runs roughly 60–200 s with a few hundred kills'
worth of activity; at ~10–40 events/second of combat this is single-digit MB per
run uncompressed. Not a concern, but the reader should stream lines rather than
JSON.parse a whole file.
scripts/lib/levelSolver.mjs (pure, unit-testable, no I/O) plus
scripts/report-level-budget.mjs (CLI, formatting, exit codes) — the same split as
the existing abReport.mjs / report-balancing-ab.mjs and
profileSeparation.mjs / report-profile-separation.mjs pairs.
It does not reimplement anything. scripts/lib/loadEngineModules.mjs already
esbuild-bundles the real src/parser/registry.ts and src/map/mapGenerator.ts for
plain Node — including the ?url grammar-wasm rewrite — and
verify-demo-campaign.mjs and generate-default-highscore.mjs Phase 0 already use
it to produce real GameMaps headlessly. The solver walks a directory, calls
parseFile then MapGenerator.generate() per file, and analyses the result.
It numbers the levels the way the game does, which is not the walk order
(58cad4b, 2026-08-21). entrypointIndex() models findEntrypoint's cascade, and
everything ahead of the pick is dropped from the numbering entirely rather than counted
as levels 1..n. This is not a rounding detail: across the corpus it removed 1,420 of
7,088 solved levels (20%) that a player can never reach, and it moved the headline
answers with them — per-repo drift read 17.9× rather than 22.2×. Any figure in this
document produced before that date is over a corpus that included unreachable levels.
scripts/lib/staticLevelAnalysis.mjs's analyzeStaticLevel(map, route) already
computes enemy counts, per-category tallies, walkable tiles, density and a
pre-placed ammo summary. The solver extends that rather than replacing it — the
existing function keeps its current callers and shape.
This is a hard constraint, and this repo has already been bitten by ignoring it — see §7.3. The rule: the solver imports every number from the module that owns it.
loadEngineModules()'s entry stub gains re-exports from src/engine/loot.ts
(weights, drop amounts, REGULAR_KILL_NO_DROP_CHANCE,
NORMAL_KILL_WEAPON_DROP_CHANCE, ELITE_BONUS_WEAPON_DROP_CHANCE),
src/difficulty.ts (DIFFICULTY_MULTIPLIERS) and src/engine/ammo.ts
(startingAmmo, AMMO_META). It already re-exports the real WEAPONS. All of
these are pure data modules with no DOM dependency, so bundling them costs nothing
— exactly the argument loadEngineModules's own doc comment already makes for
UNLOCKABLE_WEAPONS.
The one gap was src/engine/enemyAi.ts's combat constants — CLOSED 2026-08-19.
ATTACK_DAMAGE, ATTACK_COOLDOWN, FIRE_COOLDOWN_MIN/MAX, AGGRO_RADIUS,
RANGED_RANGE, MOVEMENT_SPEED, ELITE_DAMAGE_MULTIPLIER,
EDGE_CASE_DAMAGE_MULTIPLIER, EDGE_CASE_SPEED_MULTIPLIER were all module-private,
so incomingDps, survivalWindow and threatScore (§2.3, §2.4) could not be
computed without exporting them — and copying them into the solver was never an
option, that being precisely the failure §7.3 documents.
They now live in src/engine/combatConstants.ts and are exported, so the blocker is
gone. Both halves of the sequencing below also happened, in one change rather than
two: the same move folded them into SIMULATION_BALANCE wholesale, as
COMBAT_BALANCE and TRAP_BALANCE tables rather than as named scalars, and paid the
balanceHash move with a defaultHighscore.ts regeneration. What the plan below got
right is that it is a two-part cost; what it got wrong is the framing of the second
part as a list of five constants — naming constants individually is what let this
doc's own list go stale in the first place. See
Design Decisions's "Simulation constants are hashed as tables, not as picked scalars".
The original sequencing plan, kept because the ordering rule still holds for the next
simulation change: exporting is early and cheap, folding into SIMULATION_BALANCE is
last, after every other simulation change has landed — the existing rule from "Land
every simulation change first, then generate once" above, which the 2026-08-02 layout
rework already paid for learning.
Worth reading before trusting any solver output from that window, and before adding a sixth ammo pool.
report-level-budget.mjs and stage-campaign.mjs each built their own literal
dropAmounts map. When the shotgun got its own shells pool, neither gained a
shells key and loadEngineModules.mjs — whose loot.ts re-export is a
hand-maintained string of names — never re-exported SHELLS_DROP_AMOUNT
either. Three copies of the same list, none of them linked to the source.
Nothing failed. dropAmounts.shells was undefined, scaledAmount returned
NaN, the drop budget went NaN, and clearRatio.combined went with it. Two
consequences, the second worse than the first:
- The solver's own guard is
combined < 1.0, andNaN < 1isfalse, so the "NOT clearable" line became unreachable on every repo. Its closing verdict — "OK: every enemy is killable with the damage obtainable on its level" — was vacuous for six days, which is the exact failure mode of a check that cannot fail. - Carryover is computed from the same budget, so
carrystopped accumulating across levels — sinatra read2716, 2677, 2815, …, 0, 0, 600where it now reads2716 … 31436. The campaign running balance, which §1.2 says the solver must model or every level after the first reads as poorer than it plays, was modelling nothing. stage-campaign.mjspicks the level a staged campaign must include withreduce((a, b) => ratio(b) < ratio(a) ? b : a). UnderNaNthat comparison is always false, so it keeps the first element: "include the tightest level" degraded to "include whichever level came first". Any campaign staged in that window was not the level set it reports being.
dropAmountsFrom() (levelSolver.mjs) is now the single builder both entry
points call, and it throws on a loot kind it cannot price rather than
returning a map with a hole in it. levelSolver.test.mjs pins its coverage
against the real weight tables and asserts the re-export list carries each
constant — word-bounded, because the first version of that check used
includes() and ELITE_SHELLS_DROP_AMOUNT contains SHELLS_DROP_AMOUNT, so it
passed while the bug was reintroduced under it. Mutation-tested both ways.
What this invalidates. Nothing from a bot capture: the harness never reads
these constants, and the 2026-08-20 density A/B ran on demo-campaign with no
staging involved. What it invalidates is any solver output between those dates
— every balancing:budget report, and the level selection of any staged
campaign. Re-run rather than re-read.
Per level, from the generated GameMap plus the constants:
- Enemy budget — count and HP total by archetype, DPS total, per walkable tile.
- Loot budget, three ways — pre-placed (walk
map.ammoPickups, convert to damage viadamagePerAmmo), potential drops (expected value over the weight tables, gated by which weapons are owned at that campaign position, timesammoDropRate), and combined; health likewise. - Per-weapon TTK and shots-to-kill against every distinct HP value on the level.
selfSustainper archetype andclearRatioper source.- The complexity→HP curve and its outliers, including the hard-failure check: any single enemy whose HP exceeds the level's total obtainable damage.
Across a campaign it carries ammo forward the way the engine does (startingAmmo
only applies where there is no carryover), so the ratios are a running balance
rather than eight independent levels.
Difficulty is a parameter — the same level solved at easy/normal/hard gives three
budgets, because ammoDropRate and hp both move.
Per-shot hit probability is not cleanly computable offline. The Cone of Fire
deviates by (rng()*2-1) × rangeFraction³ × maxConeDeviationPx in screen pixels
against the per-column z-buffer (fire() in engine.ts), so whether a pellet
connects depends on the raycast geometry of the specific tile the enemy is standing
on and the sprite's projected width at that distance.
So the solver reports perfect-accuracy TTK and perfect-accuracy clearRatio as a
lower bound on cost, labelled as such in the output, and leaves real hit rate to
the empirical half (§2.1). A level whose perfect-accuracy clearRatio is below
1.0 is definitively unclearable; a level above 1.0 is not thereby proven clearable.
That asymmetry is useful and honest, and it is a better contract than a fabricated
accuracy coefficient that would quietly mean nothing.
(There is a partial in-engine model already — the worstCaseDeviation term behind the
bot's effective-range estimate, in engine.ts. It is a bound, not a probability,
and it is not a substitute.)
scripts/fetch-balancing-corpus.mjs, shaped like the existing
fetch-online-wads.mjs: idempotent (skip any destination that already exists),
gitignored destination, no credentials. One deliberate difference — it degrades
on failure instead of process.exit(1). That exact failure mode is already
logged against fetch-online-wads.mjs in notes, where it takes npm run dev and
npm run build down with it; there is no reason to reproduce it.
Selection criteria: pinned commits (so a corpus run is reproducible), permissive licences, and coverage of both axes that matter — size and language, chosen from what the parser already supports (bash, C, C++, C#, Go, Java, JavaScript, Objective-C, PHP, Python, Ruby, Rust, Scala, TypeScript).
| Bucket | Target | Why |
|---|---|---|
| tiny | 1–5 files, a few hundred LOC | the degenerate case — does a 2-room level even have a viable budget |
| small | ~20 files, one language | the common "someone points it at their side project" case |
| medium | ~200 files, mixed | the case the demo campaign approximates |
| large | 1000+ files | where the per-level ratios diverge most, and where an unbounded Elite is most likely |
| pathological | a file with one very high-complexity function | the §7.1 case, deliberately included as a fixture rather than hoped for |
The pathological entry should be a committed fixture under scripts/fixtures/,
not a fetched repo — it is a regression test for the outlier check, and it must not
depend on some upstream project keeping its worst function.
Every number below is real. §5.1 and §5.2 are computed from the constant tables in §1 alone, so they held before anything was built; §5.3 and §5.4 are abridged output from the shipped tools.
Expected ammo value of one kill's drop, as a fraction of the damage that kill cost.
Weights from NORMAL_LOOT_WEIGHTS with health filtered out
(healthHandledSeparately), times the 80% REGULAR_KILL_NO_DROP_CHANCE hit rate,
valued through damagePerAmmo.
Regular enemy, all weapons owned cost 250 dmg → self-sustain 0.56
Regular enemy, pistol + shotgun only cost 250 dmg → self-sustain 0.23
Edge Case, all weapons owned cost 12 dmg → self-sustain 11.6 ⚠
Elite (c=40), player damaged cost 2000 dmg → self-sustain 0.00 ⚠
Elite (c=40), player at full health cost 2000 dmg → self-sustain 0.10 ⚠
Measured against real generated levels once the solver shipped, these predictions
held and sharpened. On the demo campaign (normal): Edge Cases 8.1–13.8×, regular
enemies above 1.0 on 11 of 17 levels peaking at 2.33, Elites 0.00. Across
ripgrep's first 25 levels the regular figure runs 0.67–6.74 — real repositories
are full of trivial functions, which floor at HP_PER_COMPLEXITY (35 HP since 2026-08-21, 25 before) while
still paying a full-sized drop, so the more ordinary code a repo contains the more
free ammo it prints. Carried ammo climbs from 10,456 to 53,293 damage over those 25
levels without ever being spent down.
Three things fall straight out, and all three are actionable:
- Edge Cases are a self-sustain engine at 11.6×. They take the regular drop
path with no special-casing (§1.3), so a 12 HP nuisance yields the same expected
loot as a 250 HP regular enemy. Corridor dressing places 1–3 of them per breakup
room. Whatever ammo scarcity the rest of the design is aiming for, this bypasses
it. The fix is a drop-scaling term tied to
maxHp, or excluding Edge Cases from the ammo roll the wayhealthalready is. SUPERSEDED 2026-08-21 — the first of those two fixes shipped (f020902), along with an Edge Case HP raise to 25-35 that moves the other half of the ratio. The 11.6× figure and the 12 HP premise are both pre-change and must not be quoted as current; re-derive from a fresh report before acting on this bullet. - Elites are a pure ammo sink. The guaranteed drop is health unless you are
already at full, so a 2000 HP fight typically returns no ammo at all. That is
defensible as design — a boss should cost something — but it should be a
deliberate choice, and right now it is a side effect of
dropEliteLoot's health-first branch. - Early game is 2.4× harsher than late game, because
rollLootfilters outrockets/smg/gasuntil the matching weapon is owned and redistributes their share across a much smaller table. That is the intended behaviour (loot.ts), but nothing currently measures its size.
weapon dmg/trigger ammo dmg/ammo dps pool
echo pistol 22 1.0 22.0 147 bullets
Regex Shotgun 175 1.0 175.0 206 shells ← 8.0x the pistol, own pool
gdb 16 1.0 16.0 178 smg ← least efficient per round, by design
ghidra 150 1.0 150.0 136 rockets
Friday Hotfix 48 2.5 19.2 480 gas
SIGKILL Knife 40 -- inf 267 --
Toolchain 80 -- inf 229 --
Perfect accuracy — see §4.4. Melee dps assumes contact is maintained; the knife
defines no fireIntervalSec and is semi-auto, so its figure additionally assumes
mashing at the engine's 0.15 s default floor, while Toolchain genuinely fires
continuously while held.
SUPERSEDED 2026-08-21 — every column below moved and none of it is reproducible.
HP_PER_COMPLEXITY 25 -> 35 changes the HP totals, 507127c charges reload time in the
solver's TTK, f020902/a65948c made drops and heals HP-scaled, d299509 capped
carryover, and 58cad4b renumbered the levels. It is kept as a worked example of the
report's shape — what the columns mean and how to read a ratio — not as data. Re-run
the command for current figures.
npm run balancing:budget, abridged to the budget table:
level file enemies HP tot ammo dmg (carry/pre/drop) ratio (nofarm/comb)
1 main.c 11 550 3894 / 2363 / 1243 11.38 / 13.63
2 stage02_bootstrap.sh 13 640 5706 / 481 / 1469 9.67 / 11.96
7 stage07_service.rb 13 448 9668 / 0 / 1768 21.58 / 25.53
12 stage12_render_engine.cpp 18 2659 11597 / 1966 / 2864 5.10 / 6.18
15 stage15_god_object.java 13 4803 11542 / 1050 / 1853 2.62 / 3.01
16 stage16_hardware.h 11 228 7789 / 1344 / 1787 40.06 / 47.89
17 stage17_the_monolith.php 77 8883 8905 / 3994 / 12467 1.45 / 2.86
X combined clear ratio below 1.0 -- NOT clearable even counting every drop
! combined clear ratio below 1.2 -- no margin for a missed shot
'nofarm' counts only what you carried in plus what is on the floor.
Two things the demo campaign shows immediately. Ratios run 1.45x to 47.9x —
nothing is close to starved, which is the same conclusion the 450-run campaign
reached from the other direction, and it is why loot.ts's drop amounts were cut
~30%. And level 16 (.h, a bonus level) sits at 47.9x while level 17 sits at 2.86x
with 77 enemies, so the curve's shape is set almost entirely by what the source
files happen to contain.
The remaining sections (## Threat and survival, ## Self-sustain by archetype,
## Enemy HP outliers) print alongside it, and --all-difficulties adds a
comparison table.
npm run balancing:events, over a 3-attempt, 6-level capture (2136 events). This
is the half that could not exist before the log did:
| weapon | pulls | pellets fired | pellet hits | pellet hit rate | pulls that hit | kill share |
| echo pistol | 261 | 261 | 156 | 59.8% | 59.8% | 26.4% |
| Regex Shotgun | 39 | 273 | 114 | 41.8% | 87.2% | 14.2% |
| gdb | 193 | 193 | 101 | 52.3% | 52.3% | 23.6% |
| Friday Hotfix | 7 | 42 | 11 | 26.2% | 42.9% | 2.0% |
Hit rate by engagement distance
| Regex Shotgun | 0-2 | 35 | 27 | 77.1% |
| Regex Shotgun | 2-4 | 154 | 59 | 38.3% |
| Friday Hotfix | 0-2 | 18 | 11 | 61.1% |
| Friday Hotfix | 2-4 | 24 | 0 | 0.0% |
Overkill (damage wasted on the killing blow)
| SIGKILL Knife | 45 kills | mean 24.8 | max 38 |
| Toolchain | 5 kills | mean 60.2 | max 68 |
reliance on drops 79.6%
health pickups that granted nothing 19 of 50
drops spawned by archetype edgeCase 93, normal 70
Self-sustain, measured
| normal | 61 kills | mean HP 119 | ratio 1.09 |
| edgeCase | 87 kills | mean HP 13 | ratio 10.57 |
Four readings worth acting on, none of which the aggregate counters can produce:
- The measured self-sustain matches the solver's independent prediction. Observed 1.09 and 10.57; predicted from the drop tables alone, 0.76–2.33 and 8.1–13.8. Two models built on the same constants but from opposite directions — one from the weight tables, one from the rolls that actually happened — agreeing is what makes the Edge Case finding trustworthy rather than an artifact.
- Friday Hotfix is dead content. 2.0% of kills, 0.5% of damage, and a hit rate
that goes 61.1% → 0.0% between the 0–2 and 2–4 tile buckets. Its 3.5-tile
maxRangebites well inside the range the bot actually fires at. (Acted on 2026-08-09: the cutoff is now a decay curve, full damage to 2.5 and zero at 6.5. The acceptance criterion — pellets landing in the 4–6 tile bucket and kill share moving off 5.4%, without total gas damage per level rising — is still unmeasured.) - The shotgun's cone is doing exactly what it was designed to, 77.1% → 38.3% across the same boundary — now measured rather than asserted.
- Toolchain wastes 60.2 of its 80 damage per kill. It is a safety net, so overkill is expected; the size of it is not, and it is the first evidence that the 2x-the-knife damage buys very little against this roster.
Ordered smallest-useful-first, each step its own commit, each leaving the build green, the suite passing and the telemetry path off by default.
| Step | What | State |
|---|---|---|
| 0 | levelSolver.mjs + report-level-budget.mjs + balancing:budget. No engine change at all — reuses loadEngineModules and extends analyzeStaticLevel. |
done |
| 1 | combatConstants.ts: lift the enemy/player/projectile scalars out of enemyAi.ts, projectiles.ts, rockets.ts and engine.ts into a dependency-free module the solver can bundle. Unlocks incoming DPS, survival window, threat score. |
done |
| 2 | ?seed= / CODEENSTEIN_TELEMETRY_SEED — pin the gameplay seed so loot rolls are reproducible. |
done |
| 3 | Fix ROCKET_TRAVEL_SPEED and pin every bot mirror against the engine with constantMirrors.test.mjs. |
done, see §7.3 for the gate |
| 4 | fetch-balancing-corpus.mjs, recursive level collection via the real workspace.ts helpers, and the committed pathological fixture. |
done |
| 5 | events.ts + drainEvents() hook + NDJSON writer; levelStart, damageDealt, kill, levelEnd. |
done |
| 6 | The rest: shot, hit, damageTaken, playerDeath, lootDropped, lootCollected. |
done |
| 7 | eventMetrics.mjs + report-balancing-events.mjs — derive the empirical catalog back out of the log. |
done |
| 8 | Fold the Step-1 constants into SIMULATION_BALANCE, closing the gap its own comment documents. |
not done, deliberately |
Step 8 is left undone on purpose. Folding those constants into the hash moves
balanceHash, which invalidates every shipped replay and demands a multi-hour
defaultHighscore.ts regeneration. That is a release-shaped decision, not a
refactor, and the existing rule above — land every simulation change first, then
generate once — says it goes last. The constants now live somewhere it can be done
cheaply when someone decides to; nothing else depends on it.
The render-loop diff for steps 5 and 6 is empty. Not "small" — empty. No
emission point sits in simulate(), render(), renderNormalFrame(),
handleMovement, updateEnemyAi or collectLoot's per-frame scan. That was the
design property §3.4 committed to, and it is the one worth checking on any future
change to this path.
7.1 — Elite HP has no clamp, and this is the highest-value thing the solver will
find. hp = complexity × 25 × 2 with count = 1 for any complexity ≥ 40
(enemies.ts) is linear and unbounded. A complexity-200 function produces a
10,000 HP enemy dealing double damage: 15,000 on Hard, 37,500 in 4-player coop.
Its room is capped at 18 tiles a side (geometry.ts), so the arena does not
grow with it. Regular packs self-limit — per-member HP asymptotes to 250 — but
Elites do not. This cliff has already caused one live incident: enemies.ts
records a complexity-44 function producing 4400 HP and killing the bot 12/12 runs,
and the fix lowered ELITE_HP_MULTIPLIER 4→2 rather than adding a bound.
Measured, and it corrected this entry's original claim. The committed fixture
scripts/fixtures/pathological-repo/ holds a complexity-805 function, which
becomes a 40,250 HP Elite. As level 1 the solver reports it killable — because
startingAmmo derives the player's bullets from the level's own total enemy HP
(ammo.ts), so an unbounded Elite quietly funds its own counter-play. That
compensation exists only at campaign position 1. From level 2 on, carryover replaces
the starting formula (engine.ts) and nothing scales with what the level
contains: the same function at position 2 is 31.9× all obtainable damage on its
level, a clear ratio of 0.03. The fixture therefore ships a trivial level that
sorts first, so it pins the case that actually bites.
Answered 2026-08-08: an absolute ceiling, set from measured kills, applied
per enemy rather than per room. The Stage C event logs settle it — across seven
repositories 1,332 Elites spawned and 2 died, both ~2,000 HP on normal and
taking 22-24s; none at all on hard, where the largest enemy ever killed is
338 HP in 112,311 kills. Kill rate runs 21-26% up to ~500 runtime HP and
0.15% above 2,000, with the band between them empty by construction: the old
ELITE_HP_MULTIPLIER stacked on un-split HP, so complexity 39 gave 4 × 244 and
complexity 40 gave 1 × 2,000. Every one of the 514 cleared runs on an
Elite-bearing level left the Elite alive. Reproduce any of this with
killRateByHpBand (eventMetrics.mjs), surfaced by balancing:events.
The fix keeps the 2x room budget but caps and re-splits it
(ELITE_MEMBER_HP_CAP = 350, ELITE_MAX_MEMBERS = 8, so a room tops out at
2,800 base / 4,200 on Hard, and vim's complexity-672 function drops from 33,600
in one target to that). Two details worth not re-deriving:
- The ceiling is stated in runtime HP but applied to base HP. 350 base is 525 after Hard's 1.5x, which is the top of the band the data covers. Setting the constant to the measured 500 directly would put Hard at 750 — past every observation, in a range nothing has ever spawned in.
- The cap against obtainable damage was the other candidate and is still not
what shipped, because
spawnEnemiesruns per room and has no view of the level's pickups. It is also no longer the interesting axis:balancing:budgetalready passed its every-enemy-is-killable gate on all 23 repos before this change. The binding constraint was never ammo, it was time-to-kill under fire — which is precisely what splitting one target into a pack addresses and what the solver cannot see (§13: solver output vs observed death rate, ρ = −0.02).
Still unverified, and the narrowed re-run is the test. A 525 HP enemy on Hard sits above the 338 ever killed there — not because 525 was tried and failed, but because the 500-2,000 band never existed to sample. Interpolation says ~5s TTK against 3.4s at 250-499, but that is an extrapolation and should be named as one. Success is not "the fatal levels clear" — a level can clear by walking past its Elite — it is a non-zero kill rate above 500 HP.
7.2 — Drops are seeded but not reproducible. Map generation is fully
deterministic and content-addressed. Gameplay randomness, including every loot roll,
runs through one seeded mulberry32 stream — so drops are reproducible given a
seed. But the seed comes from randomSeed() (main.ts) and nothing can pin
it: no UI field, no CLI flag, no env var in any balancing runner. So today,
"same repo + same seed + same version = same level" holds for the level and not
for the run. Step 2 fixes it.
7.3 — A mirrored constant had already drifted, exactly as predicted — and the
gate for fixing it turned out to be unrunnable.
scripts/lib/combatPolicy.mjs's WEAPON_STATS defined ROCKET_TRAVEL_SPEED: 5 with a comment
saying it mirrors projectiles.ts's PROJECTILE_SPEED — but it is used at
combatPolicy.mjs's rocket-flight estimate to model the player's own ghidra rocket flight time, and
the real player-rocket speed is ROCKET_SPEED = 18 (combatConstants.ts). The constant
is named after the rocket and sourced from the enemy bolt. The bot overestimated
rocket flight time by 3.6×, so rocketDetonationDistanceAfterClosing reported
threats as far closer at detonation than they would be, making the bot more
rocket-shy than the game warrants. Fixed, and every mirrored weapon stat and
engine scalar is now pinned against the real module by
scripts/lib/constantMirrors.test.mjs (verified red against the old value). Adding a Weapon
already warns that "nothing links the two, and nothing fails when they drift" — this is that
failure, realised. It is also the concrete argument for §4.2's rule.
The A/B that should have gated the fix cannot see it, and finding that out is the
more useful result. The documented recipe (LEVEL_LIMIT=8, the constant flipped
via CODEENSTEIN_TELEMETRY_TUNING) was set up and then abandoned once the event
log answered a cheaper question first: does the bot ever fire a rocket at all?
Measured with ?eventLog=1 on a Pro run — the profile whose weaponPriority
leads with ghidra — reaching level 12 across four attempts:
| ghidra owned from | level 8 (forced unlock), 4 rockets in the pool |
| shots fired, levels 1–12 | 1961 |
| ghidra shots | 0 |
| gdb shots over the same span | 1102 |
So rocketDetonationDistanceAfterClosing — the only consumer of
ROCKET_TRAVEL_SPEED — never executes in a single-player bot run. The A/B would
have come back "no significant difference", and that would have meant nothing: the
branch under test never ran. One scoped capture settled it instead of two hours of
A/B wall clock.
Two consequences:
- The fix is correctness-only and cannot be regression-tested by behaviour, so
constantMirrors.test.mjsis the whole guard — which is why it pins every mirrored value rather than only this one. - Ghidra is dead content for the bot, by the same threshold §2.1 defines. Note
the fix should have made it less rocket-shy, and the measurement above was
taken with the fix already in — so something else gates the choice, in
scoreRangedWeapon's ammo-economy or self-harm terms against a 4-rocket reserve. Worth its own investigation; not this one.
7.3a — Why the bot never fires ghidra: root-caused, and one fix attempt failed. Three gates stack, and they are not equally to blame.
| # | Gate | Effect |
|---|---|---|
| 1 | rocketAimUnsafe hard-excludes ghidra below ROCKET_SAFE_DISTANCE (4 tiles), before any score is computed |
The bot fights at a median of 3.64 tiles — measured over 4,825 aimed shots. Only 46.9% of shots are at ≥4 tiles, 36.3% at ≥5 (ROCKET_CLUSTER_MIN_DIST), and 0.7% at ≥8, where self-harm risk reaches zero. |
| 2 | SELF_HARM_PENALTY_SEC = 25, scaled by risk |
scoreRangedWeapon is denominated in seconds-to-kill and a real kill takes 1–3s. A 25s penalty dominates any comparison the moment risk is nonzero: at 5 tiles ghidra scored 21.0 against the pistol's 1.3. |
| 3 | The scorer models ghidra as single-target | Its single-target DPS (136) ties gdb's (133), so the model sees a slower, scarcer, self-damaging weapon with no upside. Splash — its entire reason to exist — has no term. |
The fix attempt, and why it is not in the tree. Gates 2 and 3 were addressed
(SELF_HARM_PENALTY_SEC 25→5; a ROCKET_MAX_EFFECTIVE_TARGETS splash divisor;
the blanket "never rocket an Edge Case" guard made cluster-aware), both tied to one
constant so the A/B was a single-binary flip. Unit-level the change worked exactly
as intended — pickRangedWeapon began returning ghidra for clustered regulars and
for clustered Elites, and left lone weak targets alone.
In play it did nothing. A/B on Gamer/hard, 8 attempts a side, all 17 levels:
| base | cand | |
|---|---|---|
| ghidra shots | 2 | 1 |
| ghidra kills | 0 | 0 |
selfRocket damage |
0 | 0 |
| deaths (all levels) | 7 | 8 |
Reverted. The reason is gate 1, which fires first and which neither change touched: the bot plays at knife range, so ghidra is excluded from over half of all engagements before scoring ever runs. Fixing the scorer while the hard gate does the excluding is fixing the wrong layer — which is only obvious once engagement distance is measured, and that took the event log.
A second attempt also failed, and the pair of failures is the useful record.
The follow-up widened three more constants together — ROCKET_SAFE_DISTANCE 4→3
(the engine's real ROCKET_BLAST_RADIUS is 2.6, past which the firer takes
literally zero damage, so 4 sat 1.4 tiles beyond any danger),
ROCKET_CLUSTER_MIN_DIST 5→4, plus the scoring fix from the first attempt. Before
launching, a probe confirmed the two sides genuinely differed: base and candidate
picked different weapons in the 4–4.5 tile band.
| base | cand | |
|---|---|---|
| ghidra shots | 5 | 11 (of ~4,400 pulls) |
| ghidra kills | 1 | 0 |
| min rocket firing range | 4.95 | 5.13 |
selfRocket damage |
0 | 0 |
| deaths / campaigns completed | 8 / 0 | 8 / 0 |
The min firing range is the whole story: not one rocket was fired in the 4–4.5 band the change existed to open. Reverted.
Instrumented, and it refuted the remaining hypothesis too. shot now carries
targetArch and targetHp alongside dist, and
eventMetrics.mjs's weaponChoiceByTarget turns "what does the bot reach for,
against what, at what range" into a table. One 4-attempt Gamer/hard capture
settled in minutes what two 55-minute A/Bs could not:
| target | range | shots | weapons chosen |
|---|---|---|---|
| normal | 4–7 | 826 | gdb 63%, pistol 34%, shotgun 3%, ghidra 0% |
| normal | 7–10 | 75 | pistol 55%, gdb 44%, ghidra 1% |
| elite | 4–7 | 5 | shotgun 100% |
| elite | 2–4 | 7 | Friday Hotfix 71%, shotgun 29% |
| elite | 0–2 | 87 | knife 86%, Toolchain 9%, Friday Hotfix 5% |
The Edge Case guard was never the blocker. At 4–7 tiles against normal
targets — the largest bucket in the capture — ghidra is chosen 0% of the time, and
that guard does not apply there at all. Both the fast-path's !threat.edgeCase and
the scoring-loop copy of it are irrelevant to why ghidra goes unused.
The answer is in the Elite rows. Of 99 shots taken at an Elite, 88% are
inside 2 tiles, and 86% of those are the SIGKILL Knife. The bot closes to contact
on precisely the targets a rocket is for — so the range at which ghidra is legal is
a range at which it has already stopped shooting Elites. No constant fixes that:
it is what the profile's engageRadius, the route planner and the melee fallback
add up to.
That also re-frames §7.1's level-15 deaths. The bot is not dying because it lacks
firepower; it is dying because it walks up to two 3,500 HP enemies that deal double
damage and stabs them. killsForcedByMelee reads zero throughout — this is chosen
melee, not desperation.
The methodological point, which is the durable part. Two attempts have now been designed
off a synthetic probe of pickRangedWeapon and both were null, because the probe
supplies an idealised threat (a 4-strong cluster of non-Edge-Case enemies) that
real play rarely presents. The most likely remaining blocker is the cluster
fast-path's own !threat.edgeCase test — untouched by either attempt, and Edge
Cases are 62–78% of the roster on exactly these levels while moving 2.2× faster, so
they are usually the selected threat. But that is a hypothesis, and two hypotheses
have already failed.
The cheap way to settle it is to record the threat's archetype and distance on the
shot event (or a dedicated weaponChoice event carrying the rejected
alternatives). Then the question "what actually blocks ghidra, at what range,
against what" is a query rather than a guess — which is what the event log is for,
and what turned the engagement-distance question from an argument into a table.
Also worth stating plainly: it is entirely possible ghidra is simply the wrong weapon for this bot. It fights at a 3.64-tile median; the weapon needs 4–5 tiles to be safe and worth its ammo. Making it useful may be a change to how the bot positions, not to how it scores weapons — and that is a much larger piece of work than any constant.
Both reverted patches are preserved and are correct as far as they go.
7.4 — Cross-weapon hit rate is not comparable today. recordShot counts
trigger-pulls (once per fire()); recordHit counts pellets (once per landed
pellet, in the same fire(), plus one per rocket in advanceRockets). For the
7-pellet shotgun hits/shotsFired can reach 7.0, and accuracyPct
(playerStats.ts) is unclamped, so a shotgun-heavy run reports over 100%
accuracy. No longer harmless in shipped play: PLAYER_STATS_ENABLED became true
on 2026-08-23, so this ratio is now the "Weapon accuracy" row on the Commit Summary and
both run-end screens, and a shotgun-heavy level shows a player a figure over 100%. (The
accuracy score bonus is unaffected — scoring.ts clamps its fraction to 1.) The
player guide states the caveat rather than implying the number is trustworthy;
fixing the counters is still open, and this is now a player-facing defect rather than
a reporting one. weaponEfficiency in the balancing report reads the same ratio, so any
cross-weapon accuracy comparison drawn from it so far is wrong. §3.3's separate
shot/hit events fix it; the existing counters are left alone per the
alongside-not-instead rule, so the old ratio stays wrong and should be read as
"pellets per trigger-pull", not accuracy.
7.5 — Pre-placed ammo does not scale with what it has to kill. Placement is one
Bernoulli(0.22) trial per non-spawn room (placeAmmoPickups, pickups.ts), so it tracks entity
count. Enemies come only from function/method entities, so a file full of
classes and interfaces gets rooms and pickups but no enemies, while a file of dense
functions gets the same pickup rate against far more HP. Meanwhile startingAmmo
does scale with total enemy HP (ammo.ts) — but carryover overrides it from
level 2 on. Open question: should pre-placed ammo scale with the level's enemy HP
the way starting ammo does? The solver's clearRatio (pre-placed) column is
exactly the evidence needed to decide.
7.6 — Terminology, for anyone reading the original brief. "Enemies per
folder/file" — a file is a level; folders contribute only ordering
(main.ts). Classes, interfaces and traits produce a room but no enemy;
globals become acid pools; only functions and methods produce enemies.
7.7 — One Math.random() sits on the combat path. baseBloodCount inside damageEnemy. Cosmetic particle count only, touching no
simulation state — but it is the only direct one in engine.ts, whose own comment
says "Never Math.random() directly". Recorded so a future determinism
audit does not have to rediscover it is benign.
7.8 — countWalkableTiles is module-private (engine.ts), and
staticLevelAnalysis.mjs already hand-mirrors it and isWalkableTile. The
per-unit-area metrics in §2.4 need one of the two. This joins the three
hand-maintained isWall() mirrors documented above — same failure mode, same fix.
| Metric | Blocked by | Status |
|---|---|---|
incomingDps, survivalWindow, threatScore |
enemyAi.ts constants were module-private |
done — moved to combatConstants.ts |
overkill |
damageEnemy discarded the pre-clamp negative HP |
done — damageDealt.hpAfter |
selfSustain measured (rather than predicted) |
no dropping-archetype field | done — lootDropped.fromArch |
hitRateByDistance |
no distance recorded on a shot or hit | done — hit.dist |
| Cross-weapon hit rate | pellets and trigger-pulls were one counter | done — separate shot/hit events |
| Reproducible drop economy | gameplay seed could not be pinned | done — ?seed= |
Anything about repos other than demo-campaign/ |
no corpus | done — balancing:corpus |
damageTakenByArchetype |
melee damage arrives from updateEnemies already summed per player, and a bolt carries no reference to the enemy that fired it |
done (2026-08-05, schema 2) — damageTaken.by, a {eid, arch, amt}[] breakdown. Neither return shape had to change: melee rides the onMeleeAttack hook that already carried the biting enemy, and bolts carry a new Projectile.srcEid set at spawn. damage() is still called once per player per frame with the summed amount — splitting it would change swap-absorption arithmetic and could change who dies on which frame |
weaponSwitch / weaponGranted |
the drop path grants through lootApply.ts, which has no engine reference |
open — low value next to the above; switch frequency is a bot-policy signal, not a balance one |
Should Elite HP be capped, and against what?Answered 2026-08-08 — an absolute per-enemy ceiling from measured kills, with the room's budget split across a pack rather than concentrated. See §7.1.- Should Edge Cases drop like a 250 HP enemy? The 11.6× self-sustain in §5.1 says no; the counter-argument is that they are placed for pacing, not economy, and nerfing their drops makes corridor dressing feel like a chore.
- Should pre-placed ammo scale with enemy HP? (§7.5)
- Should the solver's verdict gate anything? It could exit non-zero on an unclearable level, which would make it usable as a check on a generated campaign — but "unclearable at perfect accuracy" is a hard failure while "clearable at perfect accuracy" proves nothing (§4.4), so only the failing direction is trustworthy as a gate.
- How much does multiplayer change the budget? Elite HP scales with player
count while loot does not obviously scale to match. Out of scope for the first
solver, but the asymmetry is visible in
multiplayerScaling.tsand deserves its own pass.