docs: add Cua-S1 4B 0.2 inference contract - #11
Merged
Merged
Conversation
Record the pinned upstream revisions, the inference contract, the /v1/systemone request mapping and the comparison tolerances for Cua-S1 4B 0.2 in src/models/cua_s1/README.md, and list the directory in the README layout table. Part of ThinkFlowLab#10. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
State the error body and split the rejection rules into 400 for bodies that are not usable JSON (including repeated keys) and 422 for requests the model cannot answer, so the text and multimodal workers behave the same way. Part of ThinkFlowLab#10. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
twu3202
marked this pull request as ready for review
September 27, 2026 12:44
This was referenced Sep 27, 2026
| | --- | --- | | ||
| | [`src/frontend/`](src/frontend/) | Rust serving code and the small engine interface. | | ||
| | [`src/models/laya/`](src/models/laya/) | LAYA preprocessing, batching, state, execution, and output processing. | | ||
| | [`src/models/cua_s1/`](src/models/cua_s1/) | Cua-S1 4B 0.2 inference contract, request mapping, and execution. | |
Contributor
There was a problem hiding this comment.
I suggest not to add third level folder name here and revise it as src/models/ for model implementation
Contributor
Author
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
twu3202
added a commit
to twu3202/system1-omni
that referenced
this pull request
Sep 29, 2026
Bring in the README change from the ThinkFlowLab#11 review: one src/models/ row in the layout table instead of one row per model. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
twu3202
added a commit
to twu3202/system1-omni
that referenced
this pull request
Sep 29, 2026
Bring in the README change from the ThinkFlowLab#11 review: one src/models/ row in the layout table instead of one row per model. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
4 tasks
hsliuustc0106
approved these changes
Sep 29, 2026
hsliuustc0106
left a comment
Contributor
There was a problem hiding this comment.
Reviewed the inference contract against the pinned upstream source and model metadata. No actionable findings. Also checked the latest README layout-table update at a59d5ea. Validation was source-based; model inference and GPU comparisons were not rerun.
twu3202
added a commit
to twu3202/system1-omni
that referenced
this pull request
Sep 30, 2026
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
1 task done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Part of #10. This records the pinned upstream revisions, the inference contract, the
/v1/systemonerequest mapping and the comparison tolerances for Cua-S1 4B 0.2, before any worker or engine code lands. It addssrc/models/cua_s1/README.mdand a row for that directory in the README layout table.The first target is the
textadapter on CUDA, loaded directly through Hugging Face Transformers and PEFT. Themultimodaladapter,scoreandnoulquestions, and Metal are listed as not covered yet.Three choices in the mapping are worth a look:
confidenceis the normalized entropy1 - H(p) / ln(n), the same as the LAYA worker reports, so both models behind the frontend use one definition. The Jev API does not publish its formula, and upstream's chooser reportsp_max.score,noul, more than 26 options) return422instead of being mapped onto letters.The second commit defines the error body (
{"detail": ...}, as the LAYA worker returns) and splits rejections into400for bodies that are not usable JSON, including repeated keys, and422for requests the model cannot answer. This keeps the text worker and the multimodal worker in #12 consistent.Test Plan
I checked the factual statements against the pinned revisions in upstream's
four-benvironment (Python 3.12, torch 2.14.0, Transformers 5.17.0, PEFT 0.21.0):ci/fetch_pinned_weights.py --verify-only.FourBModelwith thetextadapter on CPU and scored upstream's two fixtures through upstream's chooser, once with the default template and once withenable_thinking=False.FourBModelon CUDA (one RTX 6000 Ada, torch 2.14.0+cu130) in bfloat16.System1-Omni Version / Commit:
7aa2c30, based on66f689f.Test Result
textadapter attaches 128 LoRA modules (21.2M parameters) toQwen3_5ForCausalLM, which matches the layers listed in the README.AtoZare single tokens with ids 32 to 57. The positive fixture prompt is 218 tokens and ends with<|im_start|>assistant\n<think>\n.submit-form)abstain)enable_thinking=False, the expected option's probability shifts by 0.007 to 0.031 (float32: positive 0.985, negative 0.455), so the suffix belongs in the contract.