Skip to content

[RFC]: Write down which parts of the /v1/systemone contract all workers share #61

Description

@twu3202

Motivation

The frontend forwards /v1/systemone and /health and leaves the request and response to each worker. Its README links to the Jev API reference, and each model writes down its own contract, as src/models/cua_s1/README.md does. Nothing says what every worker behind the frontend should do.

TypeSafe also publishes its System One API as an OpenAPI document, version 0.2.0, called System One 0.2.0 below. It adds GET /v1/models and the shape of validation errors. The example on its Confidence page computes Choice confidence as (n * p_max - 1) / (n - 1), and its system-one-adapter-python has the same Choice formula and one for Score. None of these says much about errors beyond validation, a busy worker, health, or how usage is counted.

Workers here differ on those points. Some of that comes from the reference servers being ported, and some from choices made in this repo. The Cua-S1 contract is mine: in #11 I took LAYA's normalized entropy for confidence, while Cua's own chooser reports p_max. So the Cua-S1 workers match neither their reference nor TypeSafe.

Checked at main d2665e1 and the open PR heads on 2026-10-02:

Workers on main and in open PRs TypeSafe
Choice confidence Normalized entropy in LAYA's runtime, the Cua-S1 workers (#11) and #32. TypeSafe's formula in #55, as the Open-Jev server does, and in Decider's runtime (#57). (n * p_max - 1) / (n - 1)
Error body {"detail": "..."} in LAYA, the Cua-S1 workers and #32. {"error": "..."} in #55, as the Open-Jev server returns, and in #34's in-process mode. Plain text for the frontend's own 502 and 504. {"detail": [...]} for 422
Malformed JSON 400 in LAYA, the Cua-S1 workers and #32. 422 in #55. #34 maps every invalid request its engine reports to 400. Not named; 422 is the only status listed for a bad request
Worker busy The Cua-S1 text and native workers wait with no limit. The Cua-S1 multimodal worker returns 503 at once, without Retry-After. #34 maps an engine's Busy to 503. 529 Overloaded on the API reference page (429 is for rate limits)
usage.input_tokens Summed over every question's prompt in LAYA, Cua-S1 and #32. Summed over every candidate's prompt in #55, as the Open-Jev server does. Shared state counted once in Decider's runtime (#57). CLM's own server, run behind the frontend in #23, counts encoder cache misses, so identical requests get different bodies; #23 already asks which fields may differ. "billable input tokens"
Response model cua-ai/cua-s1-4b-0.2@<revision>:<modality> in Cua-S1, decider-2b-v11 in Decider's runtime (#57), Qwen/Qwen3.8-27B in #55 as the Open-Jev server returns. The model that answered
GET /v1/models Not routed by the frontend, so it returns 404 even when the worker serves it, as the Open-Jev and Decider servers do. Defined
/health {"status": "ready", ...} in the Cua-S1 workers and #55, {"status": "ok"} in #32. LAYA's server answers {"status": "ok", ...} once its checkpoints are loaded, without a warmup request. Not defined

Two of these affect clients. For a two-option answer at 0.8 and 0.2, TypeSafe's formula gives 0.60 and normalized entropy gives 0.28, so a client that sends anything below 0.5 to a human, as in the example on TypeSafe's Confidence page, acts on one worker's answer and not on the other's. The system1-agents client gives up after 5 seconds (wire.py), so a worker that queues without a limit can keep computing answers nobody reads. And once #34 lands, the frontend itself turns an in-process engine's errors into status codes, so it needs a rule for them anyway.

More workers are on the way (#32, #55 and the one proposed in #57), so this is cheaper to settle now than after they merge.

Proposed Change

A page on the docs site, linked from the HTTP interface section of src/frontend/README.md, in two parts.

What every worker in this repo does, with System One 0.2.0 as the base. These are suggested defaults, and the questions below are where I'm least sure:

What each model states in its README: the confidence definition, what usage.input_tokens counts, the response model, the request model names it accepts, its question types and limits, and extra answer fields such as LAYA's answer_confidence or Decider's certainty. Whether the first three follow the model's reference or System One 0.2.0 is question 1. The page collects these into one table.

External servers used as workers, such as laya-serve in the LAYA recipe, are listed as they are.

Questions:

  1. Where a model's reference implementation defines confidence, usage or the response model, should the worker follow it, as [New Model]: Native Rust/CUDA support for Open-Jev-27B-v1.1 #54 and [RFC]: Native Rust/CUDA inference for Mapika/decider-2b #57 do, with TypeSafe's definitions only where the reference has none? Or should every worker match System One 0.2.0? I lean towards the first. LAYA's runtime reports normalized entropy as confidence and p_max as answer_confidence. Cua-S1 changes either way: to p_max like Cua's chooser, or to TypeSafe's formula.
  2. Malformed JSON: 400, as LAYA, the Cua-S1 workers, Add LFM2.5-350M choice worker with candidate batching #32 and feat(frontend): add the in-process engine contract and HTTP service #34 do, or 422, as open_jev: add native Rust/CUDA text worker #55, the Open-Jev server and Decider's server do, and System One lists for validation errors? And should detail be a string, as LAYA and the Cua-S1 workers send, or a list, as System One and Decider's server send for 422?
  3. Should a busy worker answer at once, or is a limit at the ingress enough, as the frontend README allows? If it answers, 503 as in feat(frontend): add the in-process engine contract and HTTP service #34 and Decider's server, or 529 as on the API reference page? TypeSafe's SDK and system1-agents retry both, and both honour Retry-After.

Alternatives considered:

Impact: no performance or memory change. On the API side, some workers change status codes and error bodies; #55 would then differ from the Open-Jev server on errors. Cua-S1's confidence values change whatever question 1 decides.

Feedback Period

One week, until 2026-10-09.

Anything else

If this is agreed, I'll write the page and send two small PRs. In the first, the frontend forwards GET /v1/models and returns JSON for its own 502 and 504, including the one #42 adds. In the second, the Cua-S1 text and native workers follow the page. The multimodal worker follows the same contract, so I can include it too, or leave it to its own open PRs if its author prefers. Other workers can change when their owners next touch them, and no open PR needs to wait for this. A script that sends valid and invalid requests to a running worker could come later, as its own issue.

Before submitting a new issue

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions