Skip to content

cua_s1: CUDA graphs and GEMM timing for the native worker - #50

Closed
twu3202 wants to merge 18 commits into
ThinkFlowLab:mainfrom
twu3202:cua-s1-native-graphs
Closed

twu3202 wants to merge 18 commits into
ThinkFlowLab:mainfrom
twu3202:cua-s1-native-graphs

Conversation

@twu3202

@twu3202 twu3202 commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

The follow-up noted on #19: CUDA Graphs and GEMM algorithm timing for the native Cua-S1 worker. It builds on #19 (and #13), whose commits are included here until they merge; this PR's own change is 7 files, +363/-32. Kept in draft until #19 merges.

  • Prompts up to 2,048 tokens run as a CUDA graph captured for their exact length on first use, keeping the 128 most recently used lengths. Longer prompts run eagerly.
  • At startup, the worker times cuBLASLt's candidate algorithms (its heuristic shortlist, up to 16) for each GEMM shape at 19 prompt lengths from 64 to 16,384 tokens, and keeps the fastest when it beats the first choice by more than 3%. Other lengths borrow the choice made for a nearby length.

Choices worth a look:

  • The timing runs at every start. When candidates are within a few percent of each other, a restart can pick a different one, which moves the probabilities slightly (by up to 0.021 on the second prompt set below, with no top option changing). Saving the choices to a file would make restarts reproducible, for about 150 more lines; I left that out to keep this small.
  • There are no new flags: the graph length limit, the number of kept graphs and the timed lengths are constants.

The comparison scripts behind the first results are the ones from #19, kept at f9ab3fc.

Test Plan

Environment: one RTX 6000 Ada (48 GB, sm_89), CUDA 13.2, driver 595.91.07, Rust 1.98.1, with the pinned revisions.

System1-Omni Version / Commit: fc5a59d on #19 (8367333).

Test Result

  • CI's checks and the kernel tests pass.
  • Accuracy: in each of the three starts, the largest difference from the float32 worker is 0.0126 over the 16 questions, against the allowance of 0.039, and no top option changes.
  • Between starts, one or two of the 18 answered requests changed, by at most 0.0083 in a probability, and every status is the same. Against cua_s1: add a native CUDA text worker #19, every status is the same and probabilities differ by at most 0.0103.
  • On the second set, both starts pass the same rule: the largest difference from float32 is 0.0187 (0.0172 for cua_s1: add a native CUDA text worker #19) against an allowance of 0.073, and no top option changes. Between the two starts, 148 of the 276 answers changed, by at most 0.021, and no top option changed. cua_s1: add a native CUDA text worker #19 gave byte-identical answers in both starts.
  • Startup: 6.2 s to a ready /health (1.9 s for cua_s1: add a native CUDA text worker #19); the first request after that took 16 ms. The card peaked at 12.8 GiB during the timing at startup, and at 13.1 GiB while serving with 128 graphs kept and after a 16,384-token prompt (12.1 GiB for cua_s1: add a native CUDA text worker #19).
  • Latency, p50 in milliseconds:
Prompt tokens #19 This PR
139 18.2 16.1
218 24.6 21.1
292 32.3 28.8
712 68.8 60.1
15,446 1586.0 1477.6

On the second set:

Not covered here: GPUs other than sm_89 and CUDA versions other than 13.2.

Self-review

Before marking this PR ready for review or requesting maintainer review, complete
the self-review checklist.
Keep the PR in draft while this work is incomplete.
For agent assistance, use the optional precheck-pr skill.

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

Add a /v1/systemone worker for the Cua-S1 4B 0.2 text adapter in
src/models/cua_s1/text/, next to the multimodal worker proposed in ThinkFlowLab#12.
It loads the base model and the PEFT adapter directly through
Transformers and PEFT, follows the contract in
src/models/cua_s1/README.md, and answers choice questions only.

Add the fixed input set and tests in tests/cua_s1/ (contract and HTTP
tests that need no weights, and tokenizer checks), and
recipe/cua_s1/text.md with setup, launch, a parity check against
upstream FourBModel and a latency script. Ignore the recipe's weights/
and .venv/ with the same .gitignore lines as ThinkFlowLab#12.

Part of ThinkFlowLab#10.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
FastAPI runs a returned dict through jsonable_encoder, which drops every
key that starts with "_sa". Question names and option keys come from the
request, so a question named "_sample" or an option named "_save" was
missing from the answer while its tokens still counted in usage, and the
choice could name an option that was not in the probabilities.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Resolve the .gitignore and recipe/README.md conflicts with the frontend
merge (ThinkFlowLab#2): use the same Python ignores as ThinkFlowLab#12, and list the Cua-S1 text
recipe next to the Laya one.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
The frontend is merged (ThinkFlowLab#2), so the recipe builds it from the repository
root instead of the pull request branch.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
A /v1/systemone worker for the text adapter in Rust
(src/models/cua_s1/native, crate omni-cua-s1-native). It answers every
request as the reference worker in src/models/cua_s1/text does, with the
same validation, error bodies, prompt token ids and answer format, and
runs the Qwen3.5-4B forward pass on its own CUDA kernels in
src/backends/cuda/qwen3_5.

The kernels are built into libqwen3_5_cuda.so by build.sh in that
directory and loaded at run time, so building the workspace needs no
CUDA toolkit. The adapter is merged into the bfloat16 weights
beforehand, by recipe/cua_s1/export_text_merged.py. Prompts up to 2048
tokens run as CUDA graphs captured per exact length, bitwise identical
to the eager pass. GEMMs go through cuBLASLt with algorithms tuned per
GPU and kept in a file that records the GPU and cuBLASLt version they
were tuned for. Attention and the chunked Gated DeltaNet prefill run on
tensor cores.

recipe/cua_s1/native.md covers the build, the export, launching, and the
checks against the float32 reference worker and against the reference
worker over HTTP.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Cs1GemmPlan now records the cuBLASLt version a plan was tuned with, and
import refuses a plan from another version, so the check no longer
depends on the caller (ABI version 2; the plans file format is
unchanged). A shape that was not tuned borrows a tuned algorithm only
from an M at most twice its own, or from the largest tuned M when
nothing larger was tuned, and otherwise takes cuBLASLt's first choice.
An exhaustive tune fails when cuBLASLt cannot list its algorithms
instead of timing only the shortlist.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Bring in the README change from the ThinkFlowLab#11 review: one src/models/ row in
the layout table instead of one row per model.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Bring in the README change from the ThinkFlowLab#11 review: one src/models/ row in
the layout table instead of one row per model.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Rust 1.98 adds clippy's chunks_exact_to_as_chunks, which fails the three
bfloat16 decodes under -D warnings; they now use as_chunks. The recipe and
test scripts pass ruff check (E4, E7, E9, F, I) and ruff format, as the
multimodal workflow runs them over recipe/cua_s1. The reformatted scripts
produce the same corpus, float vectors and printable table as before.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Move the HTTP worker to src/frontend/cua_s1_text.py, so src/models/cua_s1/text/
holds only the model: contract.py (request mapping, prompts, answers) and
model.py (adapter checks, loading, readout; formerly engine.py and adapter.py).
The upstream MIT notice moves into contract.py. The upstream comparison and
latency scripts are no longer part of the change; the recipe keeps setup and
launch only.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Drop the comparison and benchmark tooling (diff_corpus.py, diff_workers.py,
check_native.py, and the worker's --score-all, --bench and --encode-only
modes with the eager/served switch they used), the float-vector generator,
and the Unicode table that made error messages escape non-ASCII names
exactly as Python's repr does, along with code only they used. The export
record and tokenizer are now checked once at startup, and the export script
refuses weights without download metadata, which the worker would refuse
later. The upstream MIT notice moves into contract.rs, and the README and
recipe keep setup, launch and tests.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Keep the model, its HTTP worker and the tests that exercise them. The
checked-in input set, revision detection, extra flags, bearer auth and
duplicated validation go; the limits become constants. The README now
describes the worker as the correctness reference for native execution.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>

# Conflicts:
#	src/models/cua_s1/README.md
Run each prompt eagerly with cuBLASLt's first-choice GEMM algorithms; CUDA
Graphs and GEMM tuning move to a follow-up. Parse requests with serde_json
instead of emulating CPython's json module, serve from main.rs without the
extra configuration, and check the attention kernel against a float64
reference instead of a second kernel.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Prompts up to 2,048 tokens run as a CUDA graph captured for their length,
and the GEMM algorithms are timed at startup among cuBLASLt's candidates
for a set of lengths that others borrow from.

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
@twu3202

twu3202 commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

Closing this one: the graph part landed through #52, and I'd rather not add the startup GEMM timing on its own. It adds about 4 s to startup and makes answers depend on which kernels a given start picked (148 of 276 answers moved by up to 0.021 between restarts in my runs), which doesn't fit the reproducible-by-default path the native worker has now. The branch stays on my fork if it's wanted later.

@twu3202 twu3202 closed this Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant