You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The frontend forwards /v1/systemone and /health and leaves the request and response to each worker. Its README links to the Jev API reference, and each model writes down its own contract, as src/models/cua_s1/README.md does. Nothing says what every worker behind the frontend should do.
TypeSafe also publishes its System One API as an OpenAPI document, version 0.2.0, called System One 0.2.0 below. It adds GET /v1/models and the shape of validation errors. The example on its Confidence page computes Choice confidence as (n * p_max - 1) / (n - 1), and its system-one-adapter-python has the same Choice formula and one for Score. None of these says much about errors beyond validation, a busy worker, health, or how usage is counted.
Workers here differ on those points. Some of that comes from the reference servers being ported, and some from choices made in this repo. The Cua-S1 contract is mine: in #11 I took LAYA's normalized entropy for confidence, while Cua's own chooser reports p_max. So the Cua-S1 workers match neither their reference nor TypeSafe.
Checked at maind2665e1 and the open PR heads on 2026-10-02:
Workers on main and in open PRs
TypeSafe
Choice confidence
Normalized entropy in LAYA's runtime, the Cua-S1 workers (#11) and #32. TypeSafe's formula in #55, as the Open-Jev server does, and in Decider's runtime (#57).
(n * p_max - 1) / (n - 1)
Error body
{"detail": "..."} in LAYA, the Cua-S1 workers and #32. {"error": "..."} in #55, as the Open-Jev server returns, and in #34's in-process mode. Plain text for the frontend's own 502 and 504.
{"detail": [...]} for 422
Malformed JSON
400 in LAYA, the Cua-S1 workers and #32. 422 in #55. #34 maps every invalid request its engine reports to 400.
Not named; 422 is the only status listed for a bad request
Worker busy
The Cua-S1 text and native workers wait with no limit. The Cua-S1 multimodal worker returns 503 at once, without Retry-After. #34 maps an engine's Busy to 503.
529 Overloaded on the API reference page (429 is for rate limits)
usage.input_tokens
Summed over every question's prompt in LAYA, Cua-S1 and #32. Summed over every candidate's prompt in #55, as the Open-Jev server does. Shared state counted once in Decider's runtime (#57). CLM's own server, run behind the frontend in #23, counts encoder cache misses, so identical requests get different bodies; #23 already asks which fields may differ.
"billable input tokens"
Response model
cua-ai/cua-s1-4b-0.2@<revision>:<modality> in Cua-S1, decider-2b-v11 in Decider's runtime (#57), Qwen/Qwen3.8-27B in #55 as the Open-Jev server returns.
The model that answered
GET /v1/models
Not routed by the frontend, so it returns 404 even when the worker serves it, as the Open-Jev and Decider servers do.
Defined
/health
{"status": "ready", ...} in the Cua-S1 workers and #55, {"status": "ok"} in #32. LAYA's server answers {"status": "ok", ...} once its checkpoints are loaded, without a warmup request.
Not defined
Two of these affect clients. For a two-option answer at 0.8 and 0.2, TypeSafe's formula gives 0.60 and normalized entropy gives 0.28, so a client that sends anything below 0.5 to a human, as in the example on TypeSafe's Confidence page, acts on one worker's answer and not on the other's. The system1-agents client gives up after 5 seconds (wire.py), so a worker that queues without a limit can keep computing answers nobody reads. And once #34 lands, the frontend itself turns an in-process engine's errors into status codes, so it needs a rule for them anyway.
More workers are on the way (#32, #55 and the one proposed in #57), so this is cheaper to settle now than after they merge.
Proposed Change
A page on the docs site, linked from the HTTP interface section of src/frontend/README.md, in two parts.
What every worker in this repo does, with System One 0.2.0 as the base. These are suggested defaults, and the questions below are where I'm least sure:
Errors are JSON with one key, detail, as LAYA, Decider's server and the Cua-S1 workers already return. The frontend's own 502 and 504 use the same body.
422 for a well-formed request the worker can't answer. 413 for a request over one of the worker's size limits.
The frontend forwards GET /v1/models like /health, for workers that serve it.
What each model states in its README: the confidence definition, what usage.input_tokens counts, the response model, the request model names it accepts, its question types and limits, and extra answer fields such as LAYA's answer_confidence or Decider's certainty. Whether the first three follow the model's reference or System One 0.2.0 is question 1. The page collects these into one table.
External servers used as workers, such as laya-serve in the LAYA recipe, are listed as they are.
Questions:
Where a model's reference implementation defines confidence, usage or the response model, should the worker follow it, as [New Model]: Native Rust/CUDA support for Open-Jev-27B-v1.1 #54 and [RFC]: Native Rust/CUDA inference for Mapika/decider-2b #57 do, with TypeSafe's definitions only where the reference has none? Or should every worker match System One 0.2.0? I lean towards the first. LAYA's runtime reports normalized entropy as confidence and p_max as answer_confidence. Cua-S1 changes either way: to p_max like Cua's chooser, or to TypeSafe's formula.
Should a busy worker answer at once, or is a limit at the ingress enough, as the frontend README allows? If it answers, 503 as in feat(frontend): add the in-process engine contract and HTTP service #34 and Decider's server, or 529 as on the API reference page? TypeSafe's SDK and system1-agents retry both, and both honour Retry-After.
Impact: no performance or memory change. On the API side, some workers change status codes and error bodies; #55 would then differ from the Open-Jev server on errors. Cua-S1's confidence values change whatever question 1 decides.
Feedback Period
One week, until 2026-10-09.
Anything else
If this is agreed, I'll write the page and send two small PRs. In the first, the frontend forwards GET /v1/models and returns JSON for its own 502 and 504, including the one #42 adds. In the second, the Cua-S1 text and native workers follow the page. The multimodal worker follows the same contract, so I can include it too, or leave it to its own open PRs if its author prefers. Other workers can change when their owners next touch them, and no open PR needs to wait for this. A script that sends valid and invalid requests to a running worker could come later, as its own issue.
Before submitting a new issue
Make sure you already searched for relevant issues in the issue tracker, and looked for the answer in the documentation.
Motivation
The frontend forwards
/v1/systemoneand/healthand leaves the request and response to each worker. Its README links to the Jev API reference, and each model writes down its own contract, assrc/models/cua_s1/README.mddoes. Nothing says what every worker behind the frontend should do.TypeSafe also publishes its System One API as an OpenAPI document, version 0.2.0, called System One 0.2.0 below. It adds
GET /v1/modelsand the shape of validation errors. The example on its Confidence page computes Choice confidence as(n * p_max - 1) / (n - 1), and its system-one-adapter-python has the same Choice formula and one for Score. None of these says much about errors beyond validation, a busy worker, health, or howusageis counted.Workers here differ on those points. Some of that comes from the reference servers being ported, and some from choices made in this repo. The Cua-S1 contract is mine: in #11 I took LAYA's normalized entropy for
confidence, while Cua's own chooser reportsp_max. So the Cua-S1 workers match neither their reference nor TypeSafe.Checked at
maind2665e1 and the open PR heads on 2026-10-02:mainand in open PRsconfidence(n * p_max - 1) / (n - 1){"detail": "..."}in LAYA, the Cua-S1 workers and #32.{"error": "..."}in #55, as the Open-Jev server returns, and in #34's in-process mode. Plain text for the frontend's own 502 and 504.{"detail": [...]}for 422Retry-After. #34 maps an engine'sBusyto 503.529 Overloadedon the API reference page (429 is for rate limits)usage.input_tokensmodelcua-ai/cua-s1-4b-0.2@<revision>:<modality>in Cua-S1,decider-2b-v11in Decider's runtime (#57),Qwen/Qwen3.8-27Bin #55 as the Open-Jev server returns.GET /v1/models/health{"status": "ready", ...}in the Cua-S1 workers and #55,{"status": "ok"}in #32. LAYA's server answers{"status": "ok", ...}once its checkpoints are loaded, without a warmup request.Two of these affect clients. For a two-option answer at 0.8 and 0.2, TypeSafe's formula gives 0.60 and normalized entropy gives 0.28, so a client that sends anything below 0.5 to a human, as in the example on TypeSafe's Confidence page, acts on one worker's answer and not on the other's. The system1-agents client gives up after 5 seconds (
wire.py), so a worker that queues without a limit can keep computing answers nobody reads. And once #34 lands, the frontend itself turns an in-process engine's errors into status codes, so it needs a rule for them anyway.More workers are on the way (#32, #55 and the one proposed in #57), so this is cheaper to settle now than after they merge.
Proposed Change
A page on the docs site, linked from the HTTP interface section of
src/frontend/README.md, in two parts.What every worker in this repo does, with System One 0.2.0 as the base. These are suggested defaults, and the questions below are where I'm least sure:
detail, as LAYA, Decider's server and the Cua-S1 workers already return. The frontend's own 502 and 504 use the same body./healthreturns 200 only once the worker has loaded and answered a warmup request. A worker that listens earlier returns 503 until then. Both carry astatusfield, as in feat(frontend): add the in-process engine contract and HTTP service #34 and feat(frontend): own the engine in a worker thread with a bounded queue #35.GET /v1/modelslike/health, for workers that serve it.What each model states in its README: the
confidencedefinition, whatusage.input_tokenscounts, the responsemodel, the requestmodelnames it accepts, its question types and limits, and extra answer fields such as LAYA'sanswer_confidenceor Decider'scertainty. Whether the first three follow the model's reference or System One 0.2.0 is question 1. The page collects these into one table.External servers used as workers, such as
laya-servein the LAYA recipe, are listed as they are.Questions:
confidence,usageor the responsemodel, should the worker follow it, as [New Model]: Native Rust/CUDA support for Open-Jev-27B-v1.1 #54 and [RFC]: Native Rust/CUDA inference for Mapika/decider-2b #57 do, with TypeSafe's definitions only where the reference has none? Or should every worker match System One 0.2.0? I lean towards the first. LAYA's runtime reports normalized entropy asconfidenceandp_maxasanswer_confidence. Cua-S1 changes either way: top_maxlike Cua's chooser, or to TypeSafe's formula.detailbe a string, as LAYA and the Cua-S1 workers send, or a list, as System One and Decider's server send for 422?Retry-After.Alternatives considered:
confidence,usageandmodel. Errors, busy status and health aren't model behavior, though, and a client shouldn't need to read every README to handle them.Impact: no performance or memory change. On the API side, some workers change status codes and error bodies; #55 would then differ from the Open-Jev server on errors. Cua-S1's
confidencevalues change whatever question 1 decides.Feedback Period
One week, until 2026-10-09.
Anything else
If this is agreed, I'll write the page and send two small PRs. In the first, the frontend forwards
GET /v1/modelsand returns JSON for its own 502 and 504, including the one #42 adds. In the second, the Cua-S1 text and native workers follow the page. The multimodal worker follows the same contract, so I can include it too, or leave it to its own open PRs if its author prefers. Other workers can change when their owners next touch them, and no open PR needs to wait for this. A script that sends valid and invalid requests to a running worker could come later, as its own issue.Before submitting a new issue