Community help wanted
We are looking for community help connecting system1-agents to persistent Laya serving in system1-omni and evaluating it on Apple Silicon. Python integration, API review, tests, documentation, and benchmark contributions are all welcome.
Useful ways to help:
- Implement explicit local-endpoint configuration and the served-model adapter using existing components where appropriate.
- Review the shared API contract with system1-omni and add compatibility and failure-handling tests.
- Try CLI/MCP workflows on a Mac and contribute reproducible task results, including correctness and latency.
- Write or test setup instructions and minimal examples.
Please comment with a small piece you would like to own, questions about the contract, or the Mac hardware you can test on. You do not need to take on the full integration. Focused PRs and reproducible bug reports are welcome; contributors working on backend acceleration should coordinate in ThinkFlowLab/system1-omni#3.
Goal
Coordinate with system1-omni so system1-agents can use a persistent, locally served Laya model, including an accelerated Apple Silicon backend when validated. Keep model execution optimizations in system1-omni/Laya and agent integration and task evaluation here.
Related work:
Current integration gaps
LayaModel.from_env() currently loads Laya in process and supports LAYA_DEVICE, including MPS. Separate CLI invocations load separate model instances.
The existing HTTP decision transport is configured around TypeSafe/OpenRouter, requires an API key, and identifies its adapter as Jev. A URL override alone does not provide a clear, validated local-Laya integration. The current question interface exposes choice and noul; the serving proposal also includes score, which should not be assumed to be supported by this client.
Collaboration scope
system1-omni / Laya side
- Provide the persistent worker, frontend, readiness behavior, and documented endpoint contract from the linked issues.
- Publish supported checkpoints, devices, decision types, and backend benchmark artifacts.
- Own model loading, warmup, inference optimization, and accelerator resource lifecycle.
system1-agents side
- Add explicit configuration for a served decision model: endpoint, model identity, optional authentication, and bounded request timeouts. Preserve the current in-process Laya path.
- Reuse existing decision types, answer validation, and transport components where suitable; avoid requiring a dummy cloud API key or labelling local Laya runs as Jev.
- Verify
choice, noul, and multi-question requests against the server. Explicitly document score as outside the initial client scope unless a concrete agent use case requires it.
- Agree on question serialization, response/probability fields, health/readiness, errors, and retry ownership across the two repositories. In particular, preserve the Laya-specific
noul instruction mapping rather than assuming every Jev wire representation is interchangeable.
- Record model/checkpoint and serving backend identity in run artifacts, and distinguish client round-trip time from backend-reported inference time when available.
- Document starting the server once and using it from CLI/MCP workflows, with a minimal local setup example.
Joint validation
Start with CPU contract/smoke tests, then validate on an Apple Silicon Mac. Compare in-process Laya, directly served Laya, and Laya through the system1-omni frontend using a fixed checkpoint and identical observations/questions. Use the existing ticket-router task with a fixed seed/dataset and --rethink off for an initial agent-level evaluation, plus focused noul contract cases.
Before performance runs, declare the hypothesis, isolated variable, controls, success criterion, and stop condition. Separate downloads, loading, warmup, and readiness from warm request latency. Use one feasibility run and two measured runs per configuration by default; record sufficient requests for p50/p95, variability, commands, raw results, hardware/runtime versions, and both repository SHAs. Keep the same server alive per configuration. Follow reservation requirements on scheduler-managed GPU hosts.
Check decision agreement and probability differences within predeclared tolerances, task score, per-request latency, and total task time. A persistent server may amortize startup without improving warm inference; report those outcomes separately and do not attribute backend acceleration to the Rust frontend.
Acceptance criteria
Community help wanted
We are looking for community help connecting system1-agents to persistent Laya serving in system1-omni and evaluating it on Apple Silicon. Python integration, API review, tests, documentation, and benchmark contributions are all welcome.
Useful ways to help:
Please comment with a small piece you would like to own, questions about the contract, or the Mac hardware you can test on. You do not need to take on the full integration. Focused PRs and reproducible bug reports are welcome; contributors working on backend acceleration should coordinate in ThinkFlowLab/system1-omni#3.
Goal
Coordinate with system1-omni so system1-agents can use a persistent, locally served Laya model, including an accelerated Apple Silicon backend when validated. Keep model execution optimizations in system1-omni/Laya and agent integration and task evaluation here.
Related work:
/v1/systemonetransport.Current integration gaps
LayaModel.from_env()currently loads Laya in process and supportsLAYA_DEVICE, including MPS. Separate CLI invocations load separate model instances.The existing HTTP decision transport is configured around TypeSafe/OpenRouter, requires an API key, and identifies its adapter as Jev. A URL override alone does not provide a clear, validated local-Laya integration. The current question interface exposes
choiceandnoul; the serving proposal also includesscore, which should not be assumed to be supported by this client.Collaboration scope
system1-omni / Laya side
system1-agents side
choice,noul, and multi-question requests against the server. Explicitly documentscoreas outside the initial client scope unless a concrete agent use case requires it.noulinstruction mapping rather than assuming every Jev wire representation is interchangeable.Joint validation
Start with CPU contract/smoke tests, then validate on an Apple Silicon Mac. Compare in-process Laya, directly served Laya, and Laya through the system1-omni frontend using a fixed checkpoint and identical observations/questions. Use the existing ticket-router task with a fixed seed/dataset and
--rethink offfor an initial agent-level evaluation, plus focusednoulcontract cases.Before performance runs, declare the hypothesis, isolated variable, controls, success criterion, and stop condition. Separate downloads, loading, warmup, and readiness from warm request latency. Use one feasibility run and two measured runs per configuration by default; record sufficient requests for p50/p95, variability, commands, raw results, hardware/runtime versions, and both repository SHAs. Keep the same server alive per configuration. Follow reservation requirements on scheduler-managed GPU hosts.
Check decision agreement and probability differences within predeclared tolerances, task score, per-request latency, and total task time. A persistent server may amortize startup without improving warm inference; report those outcomes separately and do not attribute backend acceleration to the Rust frontend.
Acceptance criteria