[Feat] Served Laya as a decision model (--model laya-served) - #35
cacheline999 wants to merge 20 commits into
Conversation
laya 0.3.9 builds the encoder under transformers' no_init_weights inside laya.load and loads the checkpoint with strict=True, so the without_weight_init() wrapper from ThinkFlowLab#22 and its test stubs are no longer needed. The lock moves laya from 0.3.5 to 0.3.9, the first release with the skip; later releases change MPS precision and are left for their own bump. The system_one config-error test pins the version lookup, so it passes with the laya extra installed as well as on the core install. Closes ThinkFlowLab#28
system1-omni's Laya worker pins laya[serve]==0.3.20. Comparing in-process Laya with the served one (ThinkFlowLab#20) needs the same library on both sides, so the in-process extra moves from 0.3.9 to 0.3.20. 0.3.20 autocasts to fp16 on MPS for requests with five or more questions; CPU runs and smaller requests are unchanged.
docs/served-laya.md designs the served Laya decision model: requirements, the target interface, what today's server lacks and where each part gets built, the client (--model laya-served), error handling and trade-offs. docs/api/laya-systemone.openapi.yaml is the target interface as OpenAPI 3.1 (RFC 9457 errors with stable codes, request ids, Server-Timing, /livez and /readyz, 503 with Retry-After, served_by per response); each item is marked implemented or planned. laya-systemone.current.openapi.yaml specifies today's server (laya-serve 0.3.20 behind the system1-omni worker and frontend), checked against traffic from a running worker.
ServedLayaModel asks a Laya served over HTTP (system1-omni's worker, its omni-jev frontend, or plain laya-serve) through POST /v1/systemone, with the body the in-process model builds (laya_question). The client holds one deadline per decision, retries once on a dropped connection, 502 or 504, and after Retry-After on 503, sends one X-Request-Id per decision, and maps every other status to MODEL_CALL_FAILED or MODEL_SERVICE_CONFIG_ERROR. Decision.model is the served checkpoint and revision, read from the response's served_by when a server sends it, else from /health (refreshed after 30 s, since the worker reports a CPU fallback there), else from routing.repo against plain laya-serve. The window check is shared with LayaModel (check_window). Tests run over httpx.MockTransport: the shared contract, the error and retry mapping, identity, and responses recorded from a real worker, the frontend and laya-serve, which are checked against the as-implemented OpenAPI spec (jsonschema, pyyaml and referencing join the dev extra).
--model laya-served on every agent, on decide, on the rails and in MCP decide. A test finds any name list, match or Literal that has laya without laya-served, so a new front cannot miss it.
…nkFlowLab#20) A tool-front tick keeps Decision.model as model and, from a served model, served_by, url, request_id and server_timing (Decision.provenance), so a run's artifacts show the checkpoint, revision and device of each step.
decision-models.md and configuration.md list laya-served and its variables; served-laya.md gains a run section (start the worker once, use it from the CLI and MCP) and matches system1-omni#30 on /health freshness and GPU-only worker options.
A refused connection after the retry now says the worker listens only once warm and points to system1-omni's recipe; a first /health read that fails no longer logs about a previous reading that does not exist. The --model help of every front names laya-served.
90 decisions per configuration on an M1 Pro: in-process Laya, the system1-omni worker (compile + fp16) and the same worker behind omni-jev routed every (seed, ticket) pair the same and scored 63/90 each; p50 101, 77 and 82 ms. compare_served.py builds the table from the job dirs and keeps every decision in served_laya_records.json.
The worker now starts as `python -m frontend.laya_mps --compile --weights fp16` and its /health reports `compile.enabled` instead of `compile.mode`. served_by records `compiled` (true/false); both specs, the run section and the not-up message follow. Worker fixtures were recorded again at 3d6cb57; the other responses came out byte for byte the same. The ticket-router comparison was rerun at 3d6cb57: 63/90 in every configuration, all 90 decisions routed the same, p50 88 ms in process and 72 ms served, direct and through omni-jev.
…FlowLab#20) The registration check read s1a's sources with the platform encoding, cp1252 on Windows, and reported paths with backslashes there. Every new text read and write now names utf-8, and paths are reported in POSIX form.
…stopped episodes (ThinkFlowLab#20) The identity refresh every 30 s ran before the decision's deadline began, so a stalled /health could stretch one decision to LAYA_SERVED_TIMEOUT_S plus 2 s. The refresh now shares the decision's deadline and takes at most half of what is left. compare_served.py paired the batch's ticket ids with the ticks, so a job with an episode that stopped early raised in zip(). It now pairs the processed routes with the ticks and checks that each pair agrees.
…rors (ThinkFlowLab#20) The not-up error quoted the M1 Pro ready time with --compile and fp16, which misleads on a CPU worker or another Mac; it now says the worker may still be starting and points to the recipe. The timeout comment gives its reason instead of a measured latency range, and the browser comment names served Laya as the HTTP one. The run section quotes the real message.
hsliuustc0106
left a comment
There was a problem hiding this comment.
Found three P2 issues in the served-Laya client: the total decision deadline is not enforced, a missing routed-model health entry can attribute an answer to a different checkpoint, and CPU execution can be recorded as compiled.
Reviewed commit: 3a1712b (head rechecked before posting).
Validation: 335 affected tests passed and 6 skipped across two targeted runs; ruff format/check, ty check, uv lock --check --offline, and scripts/smoke.sh passed using the existing Python 3.12 environment. The deadline issue was reproduced against a local HTTP server; the identity issues were reproduced using the recorded health fixture and checked against the referenced worker implementation at system1-omni 3d6cb57. No real-model/GPU or browser integration tests were run.
…escribes (ThinkFlowLab#20) httpx's timeout bounds each read, so a server trickling its body could outlast the decision deadline and still be accepted, and a /health read could outlast its share. Both requests now run under asyncio.wait_for with their budget; the decision's expiry is the usual timeout error. When routing.model is not in /health's models (the worker lists only the models it loaded at startup; laya loads others on first use), the client used the worker's top-level identity, which belongs to its primary model. The top level is now used only when its checkpoint is routing.repo; otherwise the record keeps routing.repo with revision and device unknown. compiled was the --compile flag; the worker runs models on the CPU uncompiled. It is now true only for a model on the GPU, and unknown when /health does not say.
…FlowLab#20) served_by_from_health's docstring now states the identity rule as the code applies it, in place of two inline comments that repeated it. The wait_for rationale sits once in the client's docstring; docstrings that restated identity() and deadline() and the PROVENANCE_KEYS comment are gone, and the timeout comment fits on its constant's line.
Thanks, all three fixed in 9c90f13, each with a test that fails on 3a1712b. b4085a2 only trims comments. |
…ields, results at 58b8cbe (ThinkFlowLab#20) system1-omni#30 merged as 58b8cbe. Its worker prepares a checkpoint loaded while serving inside the request that asked for it and holds other requests meanwhile, which outlasts the client's deadline: the timeout error and the design doc now say so and how to load such a checkpoint ahead. /health lists late-loaded checkpoints and adds preparing and compile.active; the spec, the two /health fixtures and the identity docs follow. The run section links the recipe on main. compare_served.py prints a row for a configuration that recorded no decision. Fixtures and the ticket-router comparison were run again on 58b8cbe: the other 25 fixtures are byte-identical, 63/90 in all three configurations, 90/90 routes the same.
…t; name the run's commit (ThinkFlowLab#20)
…auses a late load (ThinkFlowLab#20) No server reads LAYA_MAX_LEN: the window is the checkpoint's (english 512, multilingual 1024). The filled-window hint and the docs told users to raise it on the worker together with LAYA_SERVED_MAX_LEN, which would let a cut state through. They now say to match the checkpoint, and the missing window in /health is listed as a server gap. The client names its model in every request, so a state never changes the checkpoint; the late-load notes said it could.
Why
Closes #20.
--model layaloads Laya in every process, so each CLI run and each MCP call pays the load and nothing shares a warm model. system1-omni can now keep Laya warm in a worker (ThinkFlowLab/system1-omni#30), and this PR lets agents use it.How
s1a/decision_models/served.pyis the new backend. Open it first.ServedLayaClientposts to/v1/systemoneand owns the connection handling:LAYA_SERVED_TIMEOUT_S, default 5 s);Retry-Afteron a 503;X-Request-Idper decision, reused on the retry;MODEL_CALL_FAILEDorMODEL_SERVICE_CONFIG_ERROR.ServedLayaModelbuilds the same body as in-process Laya withlaya_question(). It shares the full-window check withLayaModel, nowcheck_window()inlaya.py.It also records who answered:
Decision.modelis<checkpoint>@<revision>. The details come from the response'sserved_bywhen a server sends one. Otherwise they come from the worker's/health, re-read after 30 s, since the worker reports there when Laya falls back to the CPU. Plain laya-serve reports less, so the record keepsrouting.repofor it.The contract is written down in
docs/served-laya.md, as design and trade-offs. There are two OpenAPI 3.1 specs indocs/api/:laya-systemone.current.openapi.yamlis today's servers, checked against recorded traffic;laya-systemone.openapi.yamlis the target, with problem+json errors,served_by,Server-Timingand/readyz.The gap between the two belongs to system1-omni; section 4 of the doc lists where each part would go.
What
--model laya-servedworks on every agent, ondecide, on the rails and in MCPdecide.tests/test_served_laya_registration.pyfails if a new list offerslayawithout it.LAYA_SERVED_URLis required: the worker,omni-jevor laya-serve.LAYA_SERVED_MODEL,LAYA_SERVED_API_KEY,LAYA_SERVED_TIMEOUT_S,LAYA_SERVED_MAX_LEN.modelfor every decision model. A served model's ticks also keepserved_by,url,request_idandserver_timing.decision-models.md,configuration.md,.env.example, the CHANGELOG, and a run section inserved-laya.md.jsonschema,pyyamlandreferencingfor the spec check (pyproject.toml,uv.lock, CONTRIBUTING). They were already in the environment throughmcpandopenjiuwen.layaextra to 0.3.20, the version the worker runs.The worker features used here (warm before listening, a live
/healthwith checkpoint, revision and device) come from ThinkFlowLab/system1-omni#30, now merged. Against plain laya-serve everything works, with the checkpoint as the only identity.The merged worker prepares a checkpoint it loads while serving inside the request that asked for it, and holds other requests meanwhile. That takes longer than the 5 s deadline, so decisions time out until it is ready; the error says so, and the run section shows how to load such a checkpoint ahead.
Results
ticket_router on an M1 Pro, 3 seeds × 30 tickets per configuration. Laya is checkpoint
55cf4c4with laya 0.3.20. The worker is system1-omni58b8cbe(the merge of ThinkFlowLab/system1-omni#30), started with--compile --weights fp16. Details, commands and every decision are inevals/ticket_router/SERVED_LAYA.md.--model laya, MPS, fp32)omni-jevAll three routed every one of the 90 decisions the same. In-process and worker probabilities differ by at most 0.001. In-process Laya is slower here because it runs fp32 and uncompiled, not because of HTTP.
Open questions
laya-served, chosen so records tell served runs from in-process ones.docs/api/laya-systemone.openapi.yaml. If it looks right, the worker-side parts belong in system1-omni.Verification
uv run ruff format --check . && uv run ruff check . && uv run ty check: all pass.uv run pytest -qandscripts/smoke.sh: 595 passed and 43 skipped with the laya extra installed;smoke: ok. CI's core install runs the new tests without torch or laya.CHANGELOG.mdand the docs say what the code does now.New tests run over
httpx.MockTransport:DecisionModelContract;tests/data/served_laya/holds 27 responses recorded from the worker, fromomni-jev(with the worker up and down) and from laya-serve. They validate against the as-is spec. On the merge only the two worker/healthresponses changed (preparing,compile.active); the other 25 are byte-identical.Live on the same M1 Pro, at the [RFC] Local Cua-S1 4B and screenshot targets for #14 and #15 #30 head
3d6cb57. The ticket_router comparison above and the last row were run on the merge:decidevia the CLI and over MCP (stdio)injection_guardrailomni-jevomni-jev--device cpu) withLAYA_API_KEYsetLAYA_SERVED_API_KEY; with it, 21/30 and ticks showcpu,float32, not compiledmultilingualrequest while onlyenglishis loaded (--compile)