Skip to content

[Deps] Require laya 0.3.9, lock 0.3.20, and drop the local weight-init skip - #37

Open
cacheline999 wants to merge 3 commits into
ThinkFlowLab:mainfrom
cacheline999:deps/laya-0.3.9
Open

cacheline999 wants to merge 3 commits into
ThinkFlowLab:mainfrom
cacheline999:deps/laya-0.3.9

Conversation

@cacheline999

@cacheline999 cacheline999 commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Why

Closes #28. #22 wrapped laya.load in transformers' no_init_weights to skip the throwaway random weight init. Laya does this itself from 0.3.9, so the wrapper and its test stubs are no longer needed.

What

  • The laya extra is laya>=0.3.9. uv lock --upgrade-package laya moves the lock from 0.3.5 to 0.3.20, the release system1-omni's Laya worker runs. Only the laya entry of the lock changes. 0.3.26 is the latest today; say if you would rather lock that.
  • without_weight_init() and the with block around laya.load are gone from s1a/decision_models/laya.py, and so are the transformers.initialization stubs in the tests.
  • --model laya keeps every request in fp32 on MPS. From 0.3.10 Laya runs a request of five or more questions in fp16 there; LayaModel.from_env raises the agent's mps_amp_min_rows so that does not happen, unless LAYA_MPS_AMP_MIN_ROWS (Laya's own variable) is set. See below for why.
  • The test for an agent without system_one pins the version lookup, so it passes with the laya extra installed as well as on the core install, and a second test covers a laya without package metadata.
  • CHANGELOG, docs/configuration.md and .env.example say the above.

Why fp32 on MPS

Only the browser front asks five or more questions in one request, on pages that offer all four operations. With Laya's default, one episode would mix fp32 and fp16 steps. Measured on an M1 Pro, the same request in fp32 and in fp16 in one process, about 87 pairs per cell:

checkpoint state fp16 against fp32, 5 to 6 questions largest probability change decisions changed
english short 15 to 20 ms slower 0.010 0
english long 6 ms slower to 9 ms faster 0.028 0
multilingual short 14 ms slower 0.051 2 of 180
multilingual long 12 ms slower 0.030 0

The two changed decisions were close to even. On this Mac fp16 pays only for long english states with eight questions (43 ms faster). The one report from an M5 (NandhaKishorM/laya#109, multilingual, short state) has fp16 5 ms faster at five questions and 17 ms at eight; setting LAYA_MPS_AMP_MIN_ROWS=5 brings that back. CUDA and CPU are unaffected.

Verification

Main and this branch were each installed with uv sync --extra dev --extra laya from their own lock, so main ran laya 0.3.5 and this branch 0.3.20, both with torch 2.14.0 and transformers 5.17.0, on an M1 Pro with macOS 26.1. The checkpoint is convaiinnovations/laya at 7b928d8. The machine had a load average of 5 to 10 from other work; main and the branch were run alternately under it.

  • ruff format --check, ruff check, ty check: all pass.

  • pytest -q with the laya extra: 490 passed, 41 skipped, 1 failed. The failure is test_s1a_home_from_the_dotenv_file_places_the_logs_and_the_runtime_root, which fails the same way on main in this setup (a checkout under a temporary directory) and passes in CI. On main with the extra, the system_one test fails as well; here it passes.

  • pytest -q on the core install: CI. CI installs no laya extra, so the runs with the extra are local only.

  • Main against this branch through LayaModel.from_env(). Each cell is three fresh processes per side. Each process loads the model, then for each of the 30 ticket-router tickets asks one choice, one noul, and one request of six questions.

    checkpoint device load, main (s) load, branch (s) one question p50, main / branch (ms) six questions p50, main / branch (ms) answers
    english CPU 49.5 / 20.0 / 16.5 25.4 / 17.0 / 16.2 243 / 255 925 / 957 identical
    english MPS 18.5 / 20.1 / 18.1 18.0 / 18.6 / 17.8 100 / 96 464 / 436 identical
    multilingual (LAYA_SUBFOLDER) CPU 18.4 / 18.2 / 18.7 20.9 / 17.4 / 24.5 114 / 107 395 / 380 identical
    multilingual (LAYA_SUBFOLDER) MPS 24.9 / 18.3 / 19.2 20.2 / 17.9 / 19.4 41 / 45 163 / 180 identical

    Load is from process start, imports included; the first english CPU run of each side read the files from disk. The p50 columns give the median of the three runs. Identical means every choice, probability, confidence and noul value is equal across the six runs of a cell.

  • Control, this branch with LAYA_MPS_AMP_MIN_ROWS=5 on MPS: one-question answers still equal main's; six-question answers differ by up to 0.010 (english) and 0.051 with two changed decisions (multilingual). So the default above is what keeps the answers.

  • Control, main with its weight-init skip turned off: 52.6 s (english) and 46.0 s (multilingual) to load on CPU, one run each. So the skip works for both checkpoints, and laya's own keeps it.

Not covered: the typed-decisions checkpoint and a local checkpoint path (neither is on this machine), transformers 4.x (laya has the fallback import; only 5.17.0 was run), CUDA, and Apple chips other than the M1 Pro.

laya 0.3.9 builds the encoder under transformers' no_init_weights inside
laya.load and loads the checkpoint with strict=True, so the
without_weight_init() wrapper from ThinkFlowLab#22 and its test stubs are no longer
needed. The lock moves laya from 0.3.5 to 0.3.9, the first release with
the skip; later releases change MPS precision and are left for their
own bump.

The system_one config-error test pins the version lookup, so it passes
with the laya extra installed as well as on the core install.

Closes ThinkFlowLab#28
uv lock --upgrade-package laya, as ThinkFlowLab#28 asks, pinned to 0.3.20, the release
system1-omni's worker runs. Only the laya entry of the lock changes.

From 0.3.10 laya runs a request of five or more questions in fp16 on MPS.
Against fp32 that moved probabilities by up to 0.05 on the multilingual
checkpoint and flipped 2 of 180 decisions, and on an M1 Pro it was slower for
short states. LayaModel.from_env now raises the agent's mps_amp_min_rows so
such requests stay in fp32, unless LAYA_MPS_AMP_MIN_ROWS is set. With that,
answers match main on CPU and MPS for both cached checkpoints.

Also tests the 'laya unknown' branch, which the pinned version lookup had
stopped covering.
@cacheline999 cacheline999 changed the title [Deps] Require laya 0.3.9 and drop the local weight-init skip [Deps] Require laya 0.3.9, lock 0.3.20, and drop the local weight-init skip Oct 4, 2026

Copy link
Copy Markdown
Contributor

@cacheline999 could you add a Demo / evidence section following the self-review guidance? The CPU/MPS measurements already answer the right questions; the missing piece is a reproducible artifact behind the tables.

Please link the comparison script and sanitized per-run outputs for the same ticket inputs on baseline/head: model loading, one- and six-question answers, and the default MPS setting versus LAYA_MPS_AMP_MIN_ROWS=5. Include the exact source/checkpoint revisions and invocation, and retain the two changed multilingual decisions in the override results. Keep the existing cold-load/background-load notes and untested configurations explicit.

Existing logs or a terminal trace are enough; no video or broader hardware sweep is needed. Please use safe sample data and redact credentials/private paths from anything shared.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Deps] Require laya 0.3.9 and drop the local weight-init skip

2 participants