cua_s1: add a text worker that loads Qwen3.5-4B through Transformers - #13
Conversation
Add a /v1/systemone worker for the Cua-S1 4B 0.2 text adapter in src/models/cua_s1/text/, next to the multimodal worker proposed in ThinkFlowLab#12. It loads the base model and the PEFT adapter directly through Transformers and PEFT, follows the contract in src/models/cua_s1/README.md, and answers choice questions only. Add the fixed input set and tests in tests/cua_s1/ (contract and HTTP tests that need no weights, and tokenizer checks), and recipe/cua_s1/text.md with setup, launch, a parity check against upstream FourBModel and a latency script. Ignore the recipe's weights/ and .venv/ with the same .gitignore lines as ThinkFlowLab#12. Part of ThinkFlowLab#10. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
FastAPI runs a returned dict through jsonable_encoder, which drops every key that starts with "_sa". Question names and option keys come from the request, so a question named "_sample" or an option named "_save" was missing from the answer while its tokens still counted in usage, and the choice could name an option that was not in the probabilities. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Resolve the .gitignore and recipe/README.md conflicts with the frontend merge (ThinkFlowLab#2): use the same Python ignores as ThinkFlowLab#12, and list the Cua-S1 text recipe next to the Laya one. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
The frontend is merged (ThinkFlowLab#2), so the recipe builds it from the repository root instead of the pull request branch. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
|
Pushed a fix and merged main. The worker now returns answers as a |
Bring in the README change from the ThinkFlowLab#11 review: one src/models/ row in the layout table instead of one row per model. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Bring in the README change from the ThinkFlowLab#11 review: one src/models/ row in the layout table instead of one row per model. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
hsliuustc0106
left a comment
There was a problem hiding this comment.
Reviewed commit 7119b8ce14edecadc696dfe04b345e2e21274f0f.
No actionable findings. Reviewed contract, adapter selection, model scoring, HTTP lifecycle, and parity harness. CPU: 32 passed, 3 skipped (HTTP dependencies/tokenizer assets unavailable).
|
fix conflicts |
| @@ -0,0 +1,122 @@ | |||
| """Send the fixed input set to a running worker, directly and through the frontend. | |||
| @@ -0,0 +1,204 @@ | |||
| """Compare the Cua-S1 worker with upstream `FourBModel` on the fixed input set. | |||
| @@ -0,0 +1,23 @@ | |||
| The system message, prompt layout and fixed values in `contract.py`, and the two upstream fixtures converted in `tests/cua_s1/data/text_inputs.json`, come from [trycua/cua](https://github.com/trycua/cua) at `0e75660ce4c2edda519e0c795fa3ad98abf4e76f` under the following license. | |||
| @@ -0,0 +1,84 @@ | |||
| """Load Qwen3.5-4B with the Cua-S1 `text` adapter and score one prompt. | |||
There was a problem hiding this comment.
why we need the enegine under mdoels/
There was a problem hiding this comment.
Fair question. It was really the model code, and the name made it look like more than that. src/models/cua_s1/text/ now holds only contract.py and model.py (loading and the letter readout, from engine.py and adapter.py), and the HTTP worker moved to src/frontend/cua_s1_text.py, following #12.
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Move the HTTP worker to src/frontend/cua_s1_text.py, so src/models/cua_s1/text/ holds only the model: contract.py (request mapping, prompts, answers) and model.py (adapter checks, loading, readout; formerly engine.py and adapter.py). The upstream MIT notice moves into contract.py. The upstream comparison and latency scripts are no longer part of the change; the recipe keeps setup and launch only. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Keep the model, its HTTP worker and the tests that exercise them. The checked-in input set, revision detection, extra flags, bearer auth and duplicated validation go; the limits become constants. The README now describes the worker as the correctness reference for native execution. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
|
Now that #12 is in, |
Purpose
Part of #10, on the contract from #11. It adds a
/v1/systemoneworker for the Cua-S1 4B 0.2textadapter that loadsQwen/Qwen3.5-4Band the adapter through Transformers and PEFT and answerschoicequestions. Following the direction on #12, it is the correctness reference that the native worker (#19) is checked against, rather than the production path.src/models/cua_s1/text/contract.py: request mapping, prompts and answers, followingsrc/models/cua_s1/README.md. Its header carries the upstream MIT notice for the copied prompt text.src/models/cua_s1/text/model.py: loading and the answer-letter readout, as upstreamFourBModeldoes them.src/frontend/cua_s1_text.py: the HTTP worker, placed as Add Cua-S1 0.2 multimodal CUDA worker with upstream parity #12 places its own.tests/cua_s1/: tests that need no weights or GPU, and a tokenizer check.recipe/cua_s1/text.md: install, download and launch.src/models/cua_s1/README.md: the status now describes this worker as the reference for native execution.The parity and latency scripts and the fixed input set behind the results below are not part of the change. They are kept at 951d707.
Test Plan
PYTHONPATH=src pytest tests/cua_s1withCUA_S1_BASEset so the tokenizer check runs, and ruff 0.16.8 (check --select E4,E7,E9,F,Iandformat --check).compare_text_with_upstream.pyfrom the scripts linked above scores every question in the fixed input set (14 requests, 16 questions) with the worker's model and withFourBModel, on CUDA in bfloat16 and in float32 with TF32 off.bench_text.pyfrom the same place: 3 warmup and 20 measured requests per case and path, one at a time.baa80d6) and the current worker on the same 46 requests (the input set, a case for every error), in bfloat16 and in float32.Environment: one RTX 6000 Ada (48 GB, sm_89), driver 595.91.07, Python 3.12, torch 2.14.0+cu130, Transformers 5.17.0 and PEFT 0.21.0, with the pinned revisions.
System1-Omni Version / Commit:
cfbccc4on main7c808f6. Parity, frontend and latency were measured at20918aa; the model code has only been moved and trimmed since then, which the last check covers.Test Result
413after the worker reads past the limit, where the previous worker reset the connection./health; the first request after that took 53 ms. The card peaked at 21.3 GiB in use while serving, driven by the 15,446-token input, because the reference path computes logits for every position.Not covered here: the
multimodaladapter (#12),scoreandnoulquestions, and more than one request at a time.Self-review
Before marking this PR ready for review or requesting maintainer review, complete
the self-review checklist.
Keep the PR in draft while this work is incomplete.
For agent assistance, use the optional precheck-pr skill.