Skip to content

cua_s1: add a text worker that loads Qwen3.5-4B through Transformers - #13

Merged
hsliuustc0106 merged 8 commits into
ThinkFlowLab:mainfrom
twu3202:cua-s1-text-worker
Sep 30, 2026
Merged

hsliuustc0106 merged 8 commits into
ThinkFlowLab:mainfrom
twu3202:cua-s1-text-worker

Conversation

@twu3202

@twu3202 twu3202 commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Part of #10, on the contract from #11. It adds a /v1/systemone worker for the Cua-S1 4B 0.2 text adapter that loads Qwen/Qwen3.5-4B and the adapter through Transformers and PEFT and answers choice questions. Following the direction on #12, it is the correctness reference that the native worker (#19) is checked against, rather than the production path.

  • src/models/cua_s1/text/contract.py: request mapping, prompts and answers, following src/models/cua_s1/README.md. Its header carries the upstream MIT notice for the copied prompt text.
  • src/models/cua_s1/text/model.py: loading and the answer-letter readout, as upstream FourBModel does them.
  • src/frontend/cua_s1_text.py: the HTTP worker, placed as Add Cua-S1 0.2 multimodal CUDA worker with upstream parity #12 places its own.
  • tests/cua_s1/: tests that need no weights or GPU, and a tokenizer check.
  • recipe/cua_s1/text.md: install, download and launch.
  • src/models/cua_s1/README.md: the status now describes this worker as the reference for native execution.

The parity and latency scripts and the fixed input set behind the results below are not part of the change. They are kept at 951d707.

Test Plan

  • PYTHONPATH=src pytest tests/cua_s1 with CUA_S1_BASE set so the tokenizer check runs, and ruff 0.16.8 (check --select E4,E7,E9,F,I and format --check).
  • Parity: compare_text_with_upstream.py from the scripts linked above scores every question in the fixed input set (14 requests, 16 questions) with the worker's model and with FourBModel, on CUDA in bfloat16 and in float32 with TF32 off.
  • The worker on CUDA behind the frontend, with bench_text.py from the same place: 3 warmup and 20 measured requests per case and path, one at a time.
  • After trimming the worker: the previous (baa80d6) and the current worker on the same 46 requests (the input set, a case for every error), in bfloat16 and in float32.

Environment: one RTX 6000 Ada (48 GB, sm_89), driver 595.91.07, Python 3.12, torch 2.14.0+cu130, Transformers 5.17.0 and PEFT 0.21.0, with the pinned revisions.

System1-Omni Version / Commit: cfbccc4 on main 7c808f6. Parity, frontend and latency were measured at 20918aa; the model code has only been moved and trimmed since then, which the last check covers.

Test Result

  • 21 tests pass and ruff is clean.
  • Parity: all 16 questions have identical prompt token ids and bitwise-identical fp32 probabilities in bfloat16 and in float32.
  • bfloat16 vs float32: the largest per-option difference is 0.0145 and no top option changes, which sets the native engine's allowance at 2 × 0.0145 + 0.01 = 0.039.
  • Frontend: all 14 requests return identical status, content type and body bytes directly and through the frontend.
  • After trimming: the previous and the current worker return byte-identical responses to all 18 answered requests and the same status for every error, in both dtypes. A body over 4 MiB now gets its 413 after the worker reads past the limit, where the previous worker reset the connection.
  • Startup: 5.6 s to a ready /health; the first request after that took 53 ms. The card peaked at 21.3 GiB in use while serving, driven by the 15,446-token input, because the reference path computes logits for every position.
  • Warm latency in bfloat16, direct p50: 46.2 ms at 139 tokens, 49.2 ms at 218, 52.7 ms at 292, 109.3 ms at 712 and 3.08 s at 15,446. Through the frontend p50 stays within 0.7 ms except at 15,446 tokens.

Not covered here: the multimodal adapter (#12), score and noul questions, and more than one request at a time.

Self-review

Before marking this PR ready for review or requesting maintainer review, complete
the self-review checklist.
Keep the PR in draft while this work is incomplete.
For agent assistance, use the optional precheck-pr skill.

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

Add a /v1/systemone worker for the Cua-S1 4B 0.2 text adapter in
src/models/cua_s1/text/, next to the multimodal worker proposed in ThinkFlowLab#12.
It loads the base model and the PEFT adapter directly through
Transformers and PEFT, follows the contract in
src/models/cua_s1/README.md, and answers choice questions only.

Add the fixed input set and tests in tests/cua_s1/ (contract and HTTP
tests that need no weights, and tokenizer checks), and
recipe/cua_s1/text.md with setup, launch, a parity check against
upstream FourBModel and a latency script. Ignore the recipe's weights/
and .venv/ with the same .gitignore lines as ThinkFlowLab#12.

Part of ThinkFlowLab#10.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
FastAPI runs a returned dict through jsonable_encoder, which drops every
key that starts with "_sa". Question names and option keys come from the
request, so a question named "_sample" or an option named "_save" was
missing from the answer while its tokens still counted in usage, and the
choice could name an option that was not in the probabilities.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Resolve the .gitignore and recipe/README.md conflicts with the frontend
merge (ThinkFlowLab#2): use the same Python ignores as ThinkFlowLab#12, and list the Cua-S1 text
recipe next to the Laya one.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
The frontend is merged (ThinkFlowLab#2), so the recipe builds it from the repository
root instead of the pull request branch.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
@twu3202

twu3202 commented Sep 28, 2026

Copy link
Copy Markdown
Contributor Author

Pushed a fix and merged main. The worker now returns answers as a JSONResponse: returning the dict let FastAPI run it through jsonable_encoder, which drops every key that starts with _sa, so a question named _sample or an option named _save went missing from the answer. There is a test for it now. With #2 merged, the recipe builds the frontend from the repository root.

@twu3202 twu3202 mentioned this pull request Sep 28, 2026
4 tasks done
Bring in the README change from the ThinkFlowLab#11 review: one src/models/ row in
the layout table instead of one row per model.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
twu3202 added a commit to twu3202/system1-omni that referenced this pull request Sep 29, 2026
Bring in the README change from the ThinkFlowLab#11 review: one src/models/ row in
the layout table instead of one row per model.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed commit 7119b8ce14edecadc696dfe04b345e2e21274f0f.

No actionable findings. Reviewed contract, adapter selection, model scoring, HTTP lifecycle, and parity harness. CPU: 32 passed, 3 skipped (HTTP dependencies/tokenizer assets unavailable).

@hsliuustc0106

Copy link
Copy Markdown
Contributor

fix conflicts

Comment thread recipe/cua_s1/bench_text.py Outdated
@@ -0,0 +1,122 @@
"""Send the fixed input set to a running worker, directly and through the frontend.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why we need this?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're right, it isn't needed here. It only produced the latency numbers in the description, so I removed it in baa80d6. The script is still at 951d707 if the numbers ever need rerunning.

@@ -0,0 +1,204 @@
"""Compare the Cua-S1 worker with upstream `FourBModel` on the fixed input set.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

same here

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same here: it only produced the parity result in the description (identical token ids and fp32 probabilities against FourBModel). Removed in baa80d6, and kept at 951d707.

@@ -0,0 +1,23 @@
The system message, prompt layout and fixed values in `contract.py`, and the two upstream fixtures converted in `tests/cua_s1/data/text_inputs.json`, come from [trycua/cua](https://github.com/trycua/cua) at `0e75660ce4c2edda519e0c795fa3ad98abf4e76f` under the following license.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

remove this

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done, removed in baa80d6. The MIT notice is now in the header of contract.py, the way #12 does it.

Comment thread src/models/cua_s1/text/engine.py Outdated
@@ -0,0 +1,84 @@
"""Load Qwen3.5-4B with the Cua-S1 `text` adapter and score one prompt.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why we need the enegine under mdoels/

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fair question. It was really the model code, and the name made it look like more than that. src/models/cua_s1/text/ now holds only contract.py and model.py (loading and the letter readout, from engine.py and adapter.py), and the HTTP worker moved to src/frontend/cua_s1_text.py, following #12.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
twu3202 added a commit to twu3202/system1-omni that referenced this pull request Sep 30, 2026
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Move the HTTP worker to src/frontend/cua_s1_text.py, so src/models/cua_s1/text/
holds only the model: contract.py (request mapping, prompts, answers) and
model.py (adapter checks, loading, readout; formerly engine.py and adapter.py).
The upstream MIT notice moves into contract.py. The upstream comparison and
latency scripts are no longer part of the change; the recipe keeps setup and
launch only.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
twu3202 added a commit to twu3202/system1-omni that referenced this pull request Sep 30, 2026
Keep the model, its HTTP worker and the tests that exercise them. The
checked-in input set, revision detection, extra flags, bearer auth and
duplicated validation go; the limits become constants. The README now
describes the worker as the correctness reference for native execution.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
@twu3202

twu3202 commented Sep 30, 2026

Copy link
Copy Markdown
Contributor Author

Now that #12 is in, src/models/cua_s1/multimodal/protocol.py and this PR's text/contract.py repeat the same pieces: the system prompt with its MIT notice, the JSON checks (duplicate keys, non-finite numbers, lone surrogates) and the choice answer with its confidence. I could move them into one small module under src/models/cua_s1/ that both import, here or in a separate PR once the follow-up fixes to #12 are in, so it doesn't get in their way.

@hsliuustc0106
hsliuustc0106 merged commit fc879d8 into ThinkFlowLab:main Sep 30, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants