Skip to content

[Help wanted] Collaborate on system1-omni Laya integration and Mac evaluation #20

Description

@hsliuustc0106

Community help wanted

We are looking for community help connecting system1-agents to persistent Laya serving in system1-omni and evaluating it on Apple Silicon. Python integration, API review, tests, documentation, and benchmark contributions are all welcome.

Useful ways to help:

  • Implement explicit local-endpoint configuration and the served-model adapter using existing components where appropriate.
  • Review the shared API contract with system1-omni and add compatibility and failure-handling tests.
  • Try CLI/MCP workflows on a Mac and contribute reproducible task results, including correctness and latency.
  • Write or test setup instructions and minimal examples.

Please comment with a small piece you would like to own, questions about the contract, or the Mac hardware you can test on. You do not need to take on the full integration. Focused PRs and reproducible bug reports are welcome; contributors working on backend acceleration should coordinate in ThinkFlowLab/system1-omni#3.

Goal

Coordinate with system1-omni so system1-agents can use a persistent, locally served Laya model, including an accelerated Apple Silicon backend when validated. Keep model execution optimizations in system1-omni/Laya and agent integration and task evaluation here.

Related work:

Current integration gaps

LayaModel.from_env() currently loads Laya in process and supports LAYA_DEVICE, including MPS. Separate CLI invocations load separate model instances.

The existing HTTP decision transport is configured around TypeSafe/OpenRouter, requires an API key, and identifies its adapter as Jev. A URL override alone does not provide a clear, validated local-Laya integration. The current question interface exposes choice and noul; the serving proposal also includes score, which should not be assumed to be supported by this client.

Collaboration scope

system1-omni / Laya side

  • Provide the persistent worker, frontend, readiness behavior, and documented endpoint contract from the linked issues.
  • Publish supported checkpoints, devices, decision types, and backend benchmark artifacts.
  • Own model loading, warmup, inference optimization, and accelerator resource lifecycle.

system1-agents side

  • Add explicit configuration for a served decision model: endpoint, model identity, optional authentication, and bounded request timeouts. Preserve the current in-process Laya path.
  • Reuse existing decision types, answer validation, and transport components where suitable; avoid requiring a dummy cloud API key or labelling local Laya runs as Jev.
  • Verify choice, noul, and multi-question requests against the server. Explicitly document score as outside the initial client scope unless a concrete agent use case requires it.
  • Agree on question serialization, response/probability fields, health/readiness, errors, and retry ownership across the two repositories. In particular, preserve the Laya-specific noul instruction mapping rather than assuming every Jev wire representation is interchangeable.
  • Record model/checkpoint and serving backend identity in run artifacts, and distinguish client round-trip time from backend-reported inference time when available.
  • Document starting the server once and using it from CLI/MCP workflows, with a minimal local setup example.

Joint validation

Start with CPU contract/smoke tests, then validate on an Apple Silicon Mac. Compare in-process Laya, directly served Laya, and Laya through the system1-omni frontend using a fixed checkpoint and identical observations/questions. Use the existing ticket-router task with a fixed seed/dataset and --rethink off for an initial agent-level evaluation, plus focused noul contract cases.

Before performance runs, declare the hypothesis, isolated variable, controls, success criterion, and stop condition. Separate downloads, loading, warmup, and readiness from warm request latency. Use one feasibility run and two measured runs per configuration by default; record sufficient requests for p50/p95, variability, commands, raw results, hardware/runtime versions, and both repository SHAs. Keep the same server alive per configuration. Follow reservation requirements on scheduler-managed GPU hosts.

Check decision agreement and probability differences within predeclared tolerances, task score, per-request latency, and total task time. A persistent server may amortize startup without improving warm inference; report those outcomes separately and do not attribute backend acceleration to the Rust frontend.

Acceptance criteria

  • Both repositories agree on and document the integration contract and initial supported question types.
  • A system1-agents CLI/MCP workflow can use a persistent local Laya endpoint without a cloud key.
  • Existing in-process Laya behavior remains available.
  • Contract tests cover answer validation, model identity, optional authentication, timeouts, and unavailable-backend errors.
  • First inference after server readiness succeeds and repeated requests reuse the loaded worker.
  • A reproducible task evaluation records correctness and timing with links to backend benchmark artifacts.
  • Setup instructions and cross-repository links are published; measured limitations are stated.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions