Conversation
Add a /v1/systemone worker for the Cua-S1 4B 0.2 text adapter in src/models/cua_s1/text/, next to the multimodal worker proposed in ThinkFlowLab#12. It loads the base model and the PEFT adapter directly through Transformers and PEFT, follows the contract in src/models/cua_s1/README.md, and answers choice questions only. Add the fixed input set and tests in tests/cua_s1/ (contract and HTTP tests that need no weights, and tokenizer checks), and recipe/cua_s1/text.md with setup, launch, a parity check against upstream FourBModel and a latency script. Ignore the recipe's weights/ and .venv/ with the same .gitignore lines as ThinkFlowLab#12. Part of ThinkFlowLab#10. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
FastAPI runs a returned dict through jsonable_encoder, which drops every key that starts with "_sa". Question names and option keys come from the request, so a question named "_sample" or an option named "_save" was missing from the answer while its tokens still counted in usage, and the choice could name an option that was not in the probabilities. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Resolve the .gitignore and recipe/README.md conflicts with the frontend merge (ThinkFlowLab#2): use the same Python ignores as ThinkFlowLab#12, and list the Cua-S1 text recipe next to the Laya one. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
The frontend is merged (ThinkFlowLab#2), so the recipe builds it from the repository root instead of the pull request branch. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
A /v1/systemone worker for the text adapter in Rust (src/models/cua_s1/native, crate omni-cua-s1-native). It answers every request as the reference worker in src/models/cua_s1/text does, with the same validation, error bodies, prompt token ids and answer format, and runs the Qwen3.5-4B forward pass on its own CUDA kernels in src/backends/cuda/qwen3_5. The kernels are built into libqwen3_5_cuda.so by build.sh in that directory and loaded at run time, so building the workspace needs no CUDA toolkit. The adapter is merged into the bfloat16 weights beforehand, by recipe/cua_s1/export_text_merged.py. Prompts up to 2048 tokens run as CUDA graphs captured per exact length, bitwise identical to the eager pass. GEMMs go through cuBLASLt with algorithms tuned per GPU and kept in a file that records the GPU and cuBLASLt version they were tuned for. Attention and the chunked Gated DeltaNet prefill run on tensor cores. recipe/cua_s1/native.md covers the build, the export, launching, and the checks against the float32 reference worker and against the reference worker over HTTP. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Cs1GemmPlan now records the cuBLASLt version a plan was tuned with, and import refuses a plan from another version, so the check no longer depends on the caller (ABI version 2; the plans file format is unchanged). A shape that was not tuned borrows a tuned algorithm only from an M at most twice its own, or from the largest tuned M when nothing larger was tuned, and otherwise takes cuBLASLt's first choice. An exhaustive tune fails when cuBLASLt cannot list its algorithms instead of timing only the shortlist. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Bring in the README change from the ThinkFlowLab#11 review: one src/models/ row in the layout table instead of one row per model. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Bring in the README change from the ThinkFlowLab#11 review: one src/models/ row in the layout table instead of one row per model. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Rust 1.98 adds clippy's chunks_exact_to_as_chunks, which fails the three bfloat16 decodes under -D warnings; they now use as_chunks. The recipe and test scripts pass ruff check (E4, E7, E9, F, I) and ruff format, as the multimodal workflow runs them over recipe/cua_s1. The reformatted scripts produce the same corpus, float vectors and printable table as before. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Move the HTTP worker to src/frontend/cua_s1_text.py, so src/models/cua_s1/text/ holds only the model: contract.py (request mapping, prompts, answers) and model.py (adapter checks, loading, readout; formerly engine.py and adapter.py). The upstream MIT notice moves into contract.py. The upstream comparison and latency scripts are no longer part of the change; the recipe keeps setup and launch only. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Drop the comparison and benchmark tooling (diff_corpus.py, diff_workers.py, check_native.py, and the worker's --score-all, --bench and --encode-only modes with the eager/served switch they used), the float-vector generator, and the Unicode table that made error messages escape non-ASCII names exactly as Python's repr does, along with code only they used. The export record and tokenizer are now checked once at startup, and the export script refuses weights without download metadata, which the worker would refuse later. The upstream MIT notice moves into contract.rs, and the README and recipe keep setup, launch and tests. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Keep the model, its HTTP worker and the tests that exercise them. The checked-in input set, revision detection, extra flags, bearer auth and duplicated validation go; the limits become constants. The README now describes the worker as the correctness reference for native execution. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Signed-off-by: Tianyao Wu <rayroy31@gmail.com> # Conflicts: # src/models/cua_s1/README.md
Run each prompt eagerly with cuBLASLt's first-choice GEMM algorithms; CUDA Graphs and GEMM tuning move to a follow-up. Parse requests with serde_json instead of emulating CPython's json module, serve from main.rs without the extra configuration, and check the attention kernel against a float64 reference instead of a second kernel. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Prompts up to 2,048 tokens run as a CUDA graph captured for their length, and the GEMM algorithms are timed at startup among cuBLASLt's candidates for a set of lengths that others borrow from. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
This was referenced Sep 30, 2026
Contributor
Author
|
Closing this one: the graph part landed through #52, and I'd rather not add the startup GEMM timing on its own. It adds about 4 s to startup and makes answers depend on which kernels a given start picked (148 of 276 answers moved by up to 0.021 between restarts in my runs), which doesn't fit the reproducible-by-default path the native worker has now. The branch stays on my fork if it's wanted later. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
The follow-up noted on #19: CUDA Graphs and GEMM algorithm timing for the native Cua-S1 worker. It builds on #19 (and #13), whose commits are included here until they merge; this PR's own change is 7 files, +363/-32. Kept in draft until #19 merges.
Choices worth a look:
The comparison scripts behind the first results are the ones from #19, kept at f9ab3fc.
Test Plan
bench_text.pyover HTTP, one request at a time, 3 warmup and 20 measured requests per case, alternating with cua_s1: add a native CUDA text worker #19, two runs each.Environment: one RTX 6000 Ada (48 GB, sm_89), CUDA 13.2, driver 595.91.07, Rust 1.98.1, with the pinned revisions.
System1-Omni Version / Commit:
fc5a59don #19 (8367333).Test Result
/health(1.9 s for cua_s1: add a native CUDA text worker #19); the first request after that took 16 ms. The card peaked at 12.8 GiB during the timing at startup, and at 13.1 GiB while serving with 128 graphs kept and after a 16,384-token prompt (12.1 GiB for cua_s1: add a native CUDA text worker #19).On the second set:
Not covered here: GPUs other than sm_89 and CUDA versions other than 13.2.
Self-review
Before marking this PR ready for review or requesting maintainer review, complete
the self-review checklist.
Keep the PR in draft while this work is incomplete.
For agent assistance, use the optional precheck-pr skill.