Skip to content

docs: add Cua-S1 4B 0.2 inference contract - #11

Merged
hsliuustc0106 merged 3 commits into
ThinkFlowLab:mainfrom
twu3202:cua-s1-contract-docs
Sep 29, 2026
Merged

hsliuustc0106 merged 3 commits into
ThinkFlowLab:mainfrom
twu3202:cua-s1-contract-docs

Conversation

@twu3202

@twu3202 twu3202 commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Part of #10. This records the pinned upstream revisions, the inference contract, the /v1/systemone request mapping and the comparison tolerances for Cua-S1 4B 0.2, before any worker or engine code lands. It adds src/models/cua_s1/README.md and a row for that directory in the README layout table.

The first target is the text adapter on CUDA, loaded directly through Hugging Face Transformers and PEFT. The multimodal adapter, score and noul questions, and Metal are listed as not covered yet.

Three choices in the mapping are worth a look:

  • confidence is the normalized entropy 1 - H(p) / ln(n), the same as the LAYA worker reports, so both models behind the frontend use one definition. The Jev API does not publish its formula, and upstream's chooser reports p_max.
  • Questions the adapters cannot answer (score, noul, more than 26 options) return 422 instead of being mapped onto letters.
  • Ties go to the earliest option. Upstream's chooser returns an error on a tie instead.

The second commit defines the error body ({"detail": ...}, as the LAYA worker returns) and splits rejections into 400 for bodies that are not usable JSON, including repeated keys, and 422 for requests the model cannot answer. This keeps the text worker and the multimodal worker in #12 consistent.

Test Plan

I checked the factual statements against the pinned revisions in upstream's four-b environment (Python 3.12, torch 2.14.0, Transformers 5.17.0, PEFT 0.21.0):

  • Downloaded both model revisions and verified every file with upstream's ci/fetch_pinned_weights.py --verify-only.
  • Read the adapter safetensors headers for target modules, layers, rank, alpha and dtype.
  • Loaded the tokenizer and chat template to check the letter token ids, special tokens, the prompt suffix, and the prompt length for upstream's positive fixture.
  • Loaded FourBModel with the text adapter on CPU and scored upstream's two fixtures through upstream's chooser, once with the default template and once with enable_thinking=False.
  • Scored the same two fixtures with FourBModel on CUDA (one RTX 6000 Ada, torch 2.14.0+cu130) in bfloat16.

System1-Omni Version / Commit: 7aa2c30, based on 66f689f.

Test Result

  • All 20 files match the lock: 9.34 GB for the base model and 187 MB for the adapters.
  • The text adapter attaches 128 LoRA modules (21.2M parameters) to Qwen3_5ForCausalLM, which matches the layers listed in the README.
  • Letters A to Z are single tokens with ids 32 to 57. The positive fixture prompt is 218 tokens and ends with <|im_start|>assistant\n<think>\n.
  • The fixture probabilities are close to upstream's published results. Upstream measured on MPS, so these numbers are context rather than a pass target:
Fixture (expected option) CPU bfloat16 CPU float32 CUDA bfloat16 Upstream, MPS bfloat16 Upstream, MPS float16
Positive (submit-form) 0.976 0.975 0.973 0.973 0.975
Negative (abstain) 0.425 0.431 0.435 0.455 0.430
  • With enable_thinking=False, the expected option's probability shifts by 0.007 to 0.031 (float32: positive 0.985, negative 0.455), so the suffix belongs in the contract.
  • The full CUDA comparison over the fixed input set comes with the text worker PR.

Record the pinned upstream revisions, the inference contract, the
/v1/systemone request mapping and the comparison tolerances for
Cua-S1 4B 0.2 in src/models/cua_s1/README.md, and list the directory
in the README layout table.

Part of ThinkFlowLab#10.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
State the error body and split the rejection rules into 400 for bodies
that are not usable JSON (including repeated keys) and 422 for requests
the model cannot answer, so the text and multimodal workers behave the
same way.

Part of ThinkFlowLab#10.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Comment thread README.md Outdated
| --- | --- |
| [`src/frontend/`](src/frontend/) | Rust serving code and the small engine interface. |
| [`src/models/laya/`](src/models/laya/) | LAYA preprocessing, batching, state, execution, and output processing. |
| [`src/models/cua_s1/`](src/models/cua_s1/) | Cua-S1 4B 0.2 inference contract, request mapping, and execution. |

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suggest not to add third level folder name here and revise it as src/models/ for model implementation

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in a59d5ea: the layout table now has one src/models/ row for model implementations, in place of the Laya and Cua-S1 rows. #13 and #19 include the change.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
twu3202 added a commit to twu3202/system1-omni that referenced this pull request Sep 29, 2026
Bring in the README change from the ThinkFlowLab#11 review: one src/models/ row in
the layout table instead of one row per model.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
twu3202 added a commit to twu3202/system1-omni that referenced this pull request Sep 29, 2026
Bring in the README change from the ThinkFlowLab#11 review: one src/models/ row in
the layout table instead of one row per model.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the inference contract against the pinned upstream source and model metadata. No actionable findings. Also checked the latest README layout-table update at a59d5ea. Validation was source-based; model inference and GPU comparisons were not rerun.

@hsliuustc0106
hsliuustc0106 merged commit 7c808f6 into ThinkFlowLab:main Sep 29, 2026
1 check passed
twu3202 added a commit to twu3202/system1-omni that referenced this pull request Sep 30, 2026
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants