cua_s1: add a native CUDA text worker - #19
Conversation
Add a /v1/systemone worker for the Cua-S1 4B 0.2 text adapter in src/models/cua_s1/text/, next to the multimodal worker proposed in ThinkFlowLab#12. It loads the base model and the PEFT adapter directly through Transformers and PEFT, follows the contract in src/models/cua_s1/README.md, and answers choice questions only. Add the fixed input set and tests in tests/cua_s1/ (contract and HTTP tests that need no weights, and tokenizer checks), and recipe/cua_s1/text.md with setup, launch, a parity check against upstream FourBModel and a latency script. Ignore the recipe's weights/ and .venv/ with the same .gitignore lines as ThinkFlowLab#12. Part of ThinkFlowLab#10. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
FastAPI runs a returned dict through jsonable_encoder, which drops every key that starts with "_sa". Question names and option keys come from the request, so a question named "_sample" or an option named "_save" was missing from the answer while its tokens still counted in usage, and the choice could name an option that was not in the probabilities. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Resolve the .gitignore and recipe/README.md conflicts with the frontend merge (ThinkFlowLab#2): use the same Python ignores as ThinkFlowLab#12, and list the Cua-S1 text recipe next to the Laya one. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
The frontend is merged (ThinkFlowLab#2), so the recipe builds it from the repository root instead of the pull request branch. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
A /v1/systemone worker for the text adapter in Rust (src/models/cua_s1/native, crate omni-cua-s1-native). It answers every request as the reference worker in src/models/cua_s1/text does, with the same validation, error bodies, prompt token ids and answer format, and runs the Qwen3.5-4B forward pass on its own CUDA kernels in src/backends/cuda/qwen3_5. The kernels are built into libqwen3_5_cuda.so by build.sh in that directory and loaded at run time, so building the workspace needs no CUDA toolkit. The adapter is merged into the bfloat16 weights beforehand, by recipe/cua_s1/export_text_merged.py. Prompts up to 2048 tokens run as CUDA graphs captured per exact length, bitwise identical to the eager pass. GEMMs go through cuBLASLt with algorithms tuned per GPU and kept in a file that records the GPU and cuBLASLt version they were tuned for. Attention and the chunked Gated DeltaNet prefill run on tensor cores. recipe/cua_s1/native.md covers the build, the export, launching, and the checks against the float32 reference worker and against the reference worker over HTTP. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
CUDA build paths are now in flight and diverging: Laya generates CUDA from TileLang, Cua-S1 hand-writes CUDA C++ with cuBLASLt, and the multimodal worker plans Triton. Nothing links against anything else yet, so the divergence is invisible today and blocking as soon as one model reuses another's kernels — which ThinkFlowLab#9 already plans, since a Kev engine and Cua-S1 share a Qwen3.5 backbone. This adds the contract and the checker, and nothing else: - contract.md: what backends must agree on — a JSON manifest per backend, four C ABI rules, the numerics that must be declared rather than discovered in a parity failure, and the build-script interface. - check_contract.py: reads the manifests and checks them against that contract. - build_script.py: reads a build script as text and reports what it declares. Never executed, because CI must not run repository code to decide whether a manifest is honest. A backend may be a subdirectory named after itself or sit directly under src/backends/cuda/, which is the layout Laya's kernels/ and tools/ use. The repository root is passed down explicitly rather than inferred by walking up from the backend, because the flat layout makes that guess one level wrong. The compile and parity tiers need hardware and are described in contract.md but deliberately not wired up here. This change needs no CUDA toolkit and no GPU, which is why it can land ahead of them. 53 tests pass. The checker was run against the real qwen3_5/ files from ThinkFlowLab#19, which it accepts, and against the real tools/ from ThinkFlowLab#16, which it rejects until that backend exposes an executable build script — the divergence it exists to surface.
|
Read What is right, and it is not obvious
Three things I would change1. if (above && usable(g, p, above->algo)) { p.algo = above->algo; }
else if (below && usable(g, p, below->algo)) { p.algo = below->algo; }
2. int acc = 0;
for (...) acc ^= p[i].x ^ p[i].w;
if (acc == 0x7fffffff) *sink = acc;The store is conditional on a value that essentially never occurs, so the compiler can prove the loop has no observable effect and delete it — at which point 3.
Smaller notes:
I have not run any of this — these are from reading, and the two that I would act on first are 1 and 2. |
Both models that need this readout share the shape of it: Kev (ThinkFlowLab#27) scores scale * dot(k_proj(c_i), q_proj(s)) and softmaxes over a question's candidates; CLM (ThinkFlowLab#9) scores exp(logit_scale) * cos(state_head(s), action_head(c)) and does the same. ThinkFlowLab#9 asks for exactly this fusion, and ThinkFlowLab#19's gemm.cu does not have it. The fusion is more than fewer launches. For a cosine the candidate's norm and its dot with the query need the same elements, so both accumulate in one read of C -- for Kev, 255 rows of 2560 floats read once instead of twice. The softmax then runs in-block over shared memory, so there is no second kernel, no atomics and no global round trip for the similarities. A batch is a grid over questions rather than a loop of launches. candidate_scoring.cu has NEVER BEEN COMPILED. The authoring machine has no CUDA toolkit and no NVIDIA GPU, so nothing here claims it compiles, runs or is fast. What is established is the arithmetic it performs: - reference.py is the float64 oracle, and its invariants are tested: a distribution, 1/K for indistinguishable candidates, no NaN for a zero candidate, no overflow where a naive softmax would produce inf. - kernel_simulation.py reproduces the kernel's algorithm in numpy -- float32 accumulation, the max-subtracted softmax, the norm floored as max(sqrt(sum), eps) and not sqrt(max(sum, eps)) -- and agrees with the oracle to a worst absolute difference of 1.8e-07 over the fixed cases. The simulation rounds twice per multiply-add where the kernel uses fmaf, so that is an upper bound on the kernel's error. - The reduction is a tree, matching __shfl_down_sync, tested structurally and shown to differ from a sequential sum on a concrete input. Tolerance is declared at 1e-4, three orders of magnitude above the simulated worst case, so a hardware failure is a failure rather than tolerance noise. gpu_parity.py checks a compiled library against the reference and skips with a reason when there is no GPU or no library. The inputs are not committed: they are deterministic from a fixed seed and would be about a megabyte of generated floats. 17 tests pass. The backend satisfies the contract from ThinkFlowLab#25, which caught the manifest over-claiming architectures the build script does not reach by default.
Both models that need this readout share the shape of it: Kev (ThinkFlowLab#27) scores scale * dot(k_proj(c_i), q_proj(s)) and softmaxes over a question's candidates; CLM (ThinkFlowLab#9) scores exp(logit_scale) * cos(state_head(s), action_head(c)) and does the same. ThinkFlowLab#9 asks for exactly this fusion and ThinkFlowLab#19's gemm.cu does not have it. The fusion is more than fewer launches. For a cosine the candidate's norm and its dot with the query need the same elements, so both accumulate in one read of C -- for Kev, 255 rows of 2560 floats read once instead of twice. The softmax runs in-block over shared memory, so there is no second kernel, no atomics and no global round trip for the similarities. A batch is a grid over questions. MEASURED, not just intended. RTX 4090 (sm_89), driver 595.58.03, CUDA 13.0 (V13.0.88): worst absolute difference: 1.835e-07 (tolerance 0.0001) all cases within tolerance All 13 fixed cases pass, including a single candidate, identical candidates, a zero candidate, K=255, and logits whose raw exponential overflows. Three defects were found and fixed; the first two only a GPU could show. 1. The parity harness passed host pointers where device pointers were expected, so the kernel dereferenced host memory as device memory. compute-sanitizer located it as an invalid global read with consecutive lanes on consecutive addresses. Fixed by allocating, copying and synchronising through cudart. 2. The query norm was summed across warps that had each computed the same partial: the inner loop already covers all of D within one warp, so the norm came out sqrt(8) too large on a 256-thread block. Every unnormalized case sat at float32 rounding while the normalize cases were off by ~1e-2, which is what pointed at it. Warp 0 now computes it alone. kernel_simulation.py did NOT catch this, because it was written from the intent rather than transcribed from the .cu. A simulation is only as good as its fidelity; hardware is what settles it. 3. The manifest claimed sm_89 while build.sh defaults to sm_80. The contract checker from ThinkFlowLab#25 caught that, as it caught the same class of over-claim in the Laya backend. The manifest is now status=validated with a declared tolerance and a reference entrypoint, which is what the contract requires of a backend making a parity claim. The reference is float64 in reference.py, and gpu_parity.py skips with a reason rather than passing when there is no GPU or no library. Not yet exercised: the batch path beyond questions=1, the K>255 rejection, and any architecture or CUDA version other than sm_89 on 13.0.
Both models that need this readout share the shape of it: Kev (ThinkFlowLab#27) scores scale * dot(k_proj(c_i), q_proj(s)) and softmaxes over a question's candidates; CLM (ThinkFlowLab#9) scores exp(logit_scale) * cos(state_head(s), action_head(c)) and does the same. ThinkFlowLab#9 asks for exactly this fusion and ThinkFlowLab#19's gemm.cu does not have it. The fusion is more than fewer launches. For a cosine the candidate's norm and its dot with the query need the same elements, so both accumulate in one read of C -- for Kev, 255 rows of 2560 floats read once instead of twice. The softmax runs in-block over shared memory, so there is no second kernel, no atomics and no global round trip for the similarities. A batch is a grid over questions. MEASURED, not just intended. RTX 4090 (sm_89), driver 595.58.03, CUDA 13.0 (V13.0.88): worst absolute difference: 1.835e-07 (tolerance 0.0001) all cases within tolerance All 13 fixed cases pass, including a single candidate, identical candidates, a zero candidate, K=255, and logits whose raw exponential overflows. Three defects were found and fixed; the first two only a GPU could show. 1. The parity harness passed host pointers where device pointers were expected, so the kernel dereferenced host memory as device memory. compute-sanitizer located it as an invalid global read with consecutive lanes on consecutive addresses. Fixed by allocating, copying and synchronising through cudart. 2. The query norm was summed across warps that had each computed the same partial: the inner loop already covers all of D within one warp, so the norm came out sqrt(8) too large on a 256-thread block. Every unnormalized case sat at float32 rounding while the normalize cases were off by ~1e-2, which is what pointed at it. Warp 0 now computes it alone. kernel_simulation.py did NOT catch this, because it was written from the intent rather than transcribed from the .cu. A simulation is only as good as its fidelity; hardware is what settles it. 3. The manifest claimed sm_89 while build.sh defaults to sm_80. The contract checker from ThinkFlowLab#25 caught that, as it caught the same class of over-claim in the Laya backend. The manifest is now status=validated with a declared tolerance and a reference entrypoint, which is what the contract requires of a backend making a parity claim. The reference is float64 in reference.py, and gpu_parity.py skips with a reason rather than passing when there is no GPU or no library. Not yet exercised: the batch path beyond questions=1, the K>255 rejection, and any architecture or CUDA version other than sm_89 on 13.0.
Both models that need this readout share the shape of it: Kev (ThinkFlowLab#27) scores scale * dot(k_proj(c_i), q_proj(s)) and softmaxes over a question's candidates; CLM (ThinkFlowLab#9) scores exp(logit_scale) * cos(state_head(s), action_head(c)) and does the same. ThinkFlowLab#9 asks for exactly this fusion and ThinkFlowLab#19's gemm.cu does not have it. The fusion is more than fewer launches. For a cosine the candidate's norm and its dot with the query need the same elements, so both accumulate in one read of C -- for Kev, 255 rows of 2560 floats read once instead of twice. The softmax runs in-block over shared memory, so there is no second kernel, no atomics and no global round trip for the similarities. A batch is a grid over questions. MEASURED, not just intended. RTX 4090 (sm_89), driver 595.58.03, CUDA 13.0 (V13.0.88): worst absolute difference: 1.835e-07 (tolerance 0.0001) all cases within tolerance All 13 fixed cases pass, including a single candidate, identical candidates, a zero candidate, K=255, and logits whose raw exponential overflows. Three defects were found and fixed; the first two only a GPU could show. 1. The parity harness passed host pointers where device pointers were expected, so the kernel dereferenced host memory as device memory. compute-sanitizer located it as an invalid global read with consecutive lanes on consecutive addresses. Fixed by allocating, copying and synchronising through cudart. 2. The query norm was summed across warps that had each computed the same partial: the inner loop already covers all of D within one warp, so the norm came out sqrt(8) too large on a 256-thread block. Every unnormalized case sat at float32 rounding while the normalize cases were off by ~1e-2, which is what pointed at it. Warp 0 now computes it alone. kernel_simulation.py did NOT catch this, because it was written from the intent rather than transcribed from the .cu. A simulation is only as good as its fidelity; hardware is what settles it. 3. The manifest claimed sm_89 while build.sh defaults to sm_80. The contract checker from ThinkFlowLab#25 caught that, as it caught the same class of over-claim in the Laya backend. The manifest is now status=validated with a declared tolerance and a reference entrypoint, which is what the contract requires of a backend making a parity claim. The reference is float64 in reference.py, and gpu_parity.py skips with a reason rather than passing when there is no GPU or no library. Not yet exercised: the batch path beyond questions=1, the K>255 rejection, and any architecture or CUDA version other than sm_89 on 13.0.
Cs1GemmPlan now records the cuBLASLt version a plan was tuned with, and import refuses a plan from another version, so the check no longer depends on the caller (ABI version 2; the plans file format is unchanged). A shape that was not tuned borrows a tuned algorithm only from an M at most twice its own, or from the largest tuned M when nothing larger was tuned, and otherwise takes cuBLASLt's first choice. An exhaustive tune fails when cuBLASLt cannot list its algorithms instead of timing only the shortlist. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
|
Thanks, this is useful, especially with Kev planning to load the same library.
Also taken: an exhaustive tune now returns an error when cuBLASLt can't list its algorithms. Pushed as 637ced1. With the existing plans file the scores are bitwise unchanged, and a fresh |
Bring in the README change from the ThinkFlowLab#11 review: one src/models/ row in the layout table instead of one row per model. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Bring in the README change from the ThinkFlowLab#11 review: one src/models/ row in the layout table instead of one row per model. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Rust 1.98 adds clippy's chunks_exact_to_as_chunks, which fails the three bfloat16 decodes under -D warnings; they now use as_chunks. The recipe and test scripts pass ruff check (E4, E7, E9, F, I) and ruff format, as the multimodal workflow runs them over recipe/cua_s1. The reformatted scripts produce the same corpus, float vectors and printable table as before. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
hsliuustc0106
left a comment
There was a problem hiding this comment.
Reviewed commit 6104d17cc0d879c726002be64f6a8ac3e91e00e6.
No actionable findings in native-worker source review (request mapping, tokenizer export, loading, buffer layout, graph ownership, launch boundaries). Rust library tests: 9 passed, 1 ignored (external float vectors). CUDA kernels and model parity not executed; source review is not numerical certification.
|
cleanup it |
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Move the HTTP worker to src/frontend/cua_s1_text.py, so src/models/cua_s1/text/ holds only the model: contract.py (request mapping, prompts, answers) and model.py (adapter checks, loading, readout; formerly engine.py and adapter.py). The upstream MIT notice moves into contract.py. The upstream comparison and latency scripts are no longer part of the change; the recipe keeps setup and launch only. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Drop the comparison and benchmark tooling (diff_corpus.py, diff_workers.py, check_native.py, and the worker's --score-all, --bench and --encode-only modes with the eager/served switch they used), the float-vector generator, and the Unicode table that made error messages escape non-ASCII names exactly as Python's repr does, along with code only they used. The export record and tokenizer are now checked once at startup, and the export script refuses weights without download metadata, which the worker would refuse later. The upstream MIT notice moves into contract.rs, and the README and recipe keep setup, launch and tests. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
|
Thanks, cleaned up in a086a31. It drops the comparison and benchmark tooling ( If it helps, once #34 and #35 land I can move serving into the frontend's in-process engine, which would take the HTTP code out of this crate. |
Keep the model, its HTTP worker and the tests that exercise them. The checked-in input set, revision detection, extra flags, bearer auth and duplicated validation go; the limits become constants. The README now describes the worker as the correctness reference for native execution. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
|
is this ready for review? |
Signed-off-by: Tianyao Wu <rayroy31@gmail.com> # Conflicts: # src/models/cua_s1/README.md
Run each prompt eagerly with cuBLASLt's first-choice GEMM algorithms; CUDA Graphs and GEMM tuning move to a follow-up. Parse requests with serde_json instead of emulating CPython's json module, serve from main.rs without the extra configuration, and check the attention kernel against a float64 reference instead of a second kernel. Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
|
Yes. I've just pushed the trimmed version: the minimal eager path, with CUDA Graphs and GEMM timing left for a follow-up PR, and the description now has its results. |
|
For the follow-up: the previous version of this PR captured CUDA Graphs and timed cuBLASLt's candidate GEMM algorithms at startup, without the exhaustive search and plan files. Measured with the same bench on the same GPU, in a separate run, p50 in milliseconds:
Startup goes from 1.9 s to about 6 s with the timing. I'll rewrite that part on top of this PR, test it again, and open it as a separate PR. |
|
generally lgtm, please fix the minors |
|
please help review PR #52 as well |
- src/frontend/laya_mps.py: the HTTP worker, started with flags (PYTHONPATH=src python -m frontend.laya_mps --device mps --model english [--compile] [--weights fp16] [--require-device]) instead of LAYA_WORKER_* environment variables. laya-serve's own variables (LAYA_API_KEY, ...) still apply. - src/models/laya/: model-side code only. engine.py holds the warmup, the per-model description /health returns and the revision record; optimize.py is unchanged. - tests/laya/: the unit and contract tests, run with PYTHONPATH=src. - recipe/laya/requirements-mps.txt pins the validated environment. - recipe/laya/bench/: 10 scripts become 6. The answer comparison is a section of report.py, the frontend-overhead probe is paired.py's --a-url/--b-url mode, and the workload generator and token check are replaced by the committed workloads.jsonl plus its token counts in the README. The five result tables leave the repository and join the raw data as a release asset. - README: LAYA's row in the models table points at the Apple Silicon worker. No behaviour change: 26 unit and 13 contract tests pass, and a feasibility pass of every script gives the same numbers as before (68-token request 28 ms with both options, paired ratio 0.64, answers within 0.0031, frontend overhead ratio 1.00-1.01).
- src/frontend/laya_mps.py: the HTTP worker, started with flags (PYTHONPATH=src python -m frontend.laya_mps --device mps --model english [--compile] [--weights fp16] [--require-device]) instead of LAYA_WORKER_* environment variables. laya-serve's own variables (LAYA_API_KEY, ...) still apply. - src/models/laya/: model-side code only. engine.py holds the warmup, the per-model description /health returns and the revision record; optimize.py is unchanged. - tests/laya/: the unit and contract tests, run with PYTHONPATH=src. - recipe/laya/requirements-mps.txt pins the validated environment. - recipe/laya/bench/: 10 scripts become 6. The answer comparison is a section of report.py, the frontend-overhead probe is paired.py's --a-url/--b-url mode, and the workload generator and token check are replaced by the committed workloads.jsonl plus its token counts in the README. The five result tables leave the repository and join the raw data as a release asset. - README: LAYA's row in the models table points at the Apple Silicon worker. No behaviour change: 26 unit and 13 contract tests pass, and a feasibility pass of every script gives the same numbers as before (68-token request 28 ms with both options, paired ratio 0.64, answers within 0.0031, frontend overhead ratio 1.00-1.01).
All three come from review on ThinkFlowLab#25. 1. reference.entrypoint was used as a path without checking its type. A list reached os.path.isabs, whose TypeError escaped check_manifest: the CLI printed no structured report at all, not even under --json, and every later backend went unchecked. Non-strings are now an Issue and checking continues. 2. A directory holding any *.backend.json was exempt from the undeclared-backend check, but discovery only loads <dir>/<dir>.backend.json. A manifest named typo.backend.json therefore bypassed every check for that backend: exit 0, checked: [], no errors and no warnings. A manifest that is present but not under the expected name is now reported. 3. Architectures read from `arch=` variables and targets written literally in a -gencode flag were combined with `or`, so the literals were discarded whenever a variable existed. A script assigning `arch=${2:-89}` and then compiling `-gencode arch=compute_90,code=sm_90` builds sm_90 only, yet a manifest declaring [89] passed with no findings. build_script.py now reports gencode_architectures: the targets that actually reach nvcc, resolved against the variable environment in force at each line, so a loop that reassigns `arch` still yields both targets. The checker treats those as the authority and reports an assigned value that never reaches a -gencode flag. 62 tests, seven of them new and one per defect. Verified against ThinkFlowLab#19's real build.sh, which still parses to variables [89] / gencode [89] and passes with no findings.
Purpose
Part of #10: the minimal native text/prefill path for Cua-S1 4B 0.2, as discussed on #12. It builds on #13, whose commits are included here until it merges; the #13 worker is the correctness reference it is checked against.
It serves the
textadapter on/v1/systemonelike the #13 worker (validation, status codes, prompt token ids, answer format) and runs the forward pass on its own CUDA kernels, without Python or PyTorch:src/models/cua_s1/native/: the Rust crateomni-cua-s1-native: request handling, tokenization, the Qwen3.5-4B forward pass (weights, buffers, the layer loop) and the 26-row head.src/backends/cuda/qwen3_5/: the operations in CUDA C++: RMSNorm variants, the Gated DeltaNet convolution, gates and chunked prefill, rotary embedding, a FlashAttention-2 style attention kernel on tensor cores, and bfloat16 GEMMs through cuBLASLt.build.shbuildslibqwen3_5_cuda.so, which the worker loads at run time, so the workspace builds and tests without a CUDA toolkit.recipe/cua_s1/native.mdandexport_text_merged.py: build, export the merged weights, launch.Choices worth a look:
textadapter is merged into the bfloat16 weights beforehand. Norms and elementwise operations round to bfloat16 where Transformers does; attention and the Gated DeltaNet prefill keep some intermediate results in bfloat16, as FlashAttention and flash-linear-attention do.The comparison scripts behind the results below are not part of the change. They are kept at f9ab3fc.
Test Plan
cargo fmt --all --check,cargo clippy --workspace --locked --all-targets -- -D warnings,cargo test --workspace --lockedandcargo build --workspace --release --locked.CUA_S1_CUDA_LIB=$PWD/target/release/libqwen3_5_cuda.so cargo test --release -p omni-cua-s1-native --test kernels -- --ignored): attention at 1 to 2,048 tokens and the chunked Gated DeltaNet prefill, each against a float64 reference.bench_text.pyover HTTP, one request at a time, 3 warmup and 20 measured requests per case, two runs each.Environment: one RTX 6000 Ada (48 GB, sm_89), CUDA 13.2, driver 595.91.07, Rust 1.98.1, with the pinned revisions.
System1-Omni Version / Commit:
8367333on #13 (cfbccc4).Test Result
413where the previous version reset the connection./health; the first request after that took 14 ms. The card peaked at 12.1 GiB in use while serving.Not covered here:
scoreandnoulquestions and themultimodaladapter; more than one request at a time; GPUs other than sm_89 and CUDA versions other than 13.2.Self-review
Before marking this PR ready for review or requesting maintainer review, complete
the self-review checklist.
Keep the PR in draft while this work is incomplete.
For agent assistance, use the optional precheck-pr skill.