Skip to content

cuda: add a fused candidate-scoring kernel for CLM and Kev - #69

Draft
xiaoyu-xyz wants to merge 7 commits into
ThinkFlowLab:mainfrom
xiaoyu-xyz:cuda-scoring-kernel
Draft

xiaoyu-xyz wants to merge 7 commits into
ThinkFlowLab:mainfrom
xiaoyu-xyz:cuda-scoring-kernel

Conversation

@xiaoyu-xyz

@xiaoyu-xyz xiaoyu-xyz commented Oct 3, 2026 •

Copy link
Copy Markdown

One CUDA kernel that turns a query vector and a matrix of candidate vectors into a probability distribution, without materialising the similarities in between.

Two models need this readout and neither has it:

Model Similarity
CLM (#9) "exp(logit_scale) * cos(state_head(s), action_head(c))", answered by "the softmax over the question's candidates"
Kev (#27) "scale * (k(h_opts) @ q(h_decide)) with scale = 1/sqrt(256), softmaxed over that question's candidates"

Neither is in #19's gemm.cu, which is a bf16 cuBLASLt GEMM. The two differ in the similarity and in whether a projection happens first; the kernel takes the shape they share and leaves the projections to the caller, which is where #6 puts kernel selection.

Core code: 123 lines (candidate_scoring.cu). The oracle, the simulation, the parity harness, the benchmark and the tests are separate.

What is fused, precisely

Not merely "fewer launches". When the similarity is a cosine — CLM's case, not Kev's, which is a plain dot — each candidate's norm and its dot with the query need the same elements, so both accumulate in a single read of C. The unfused version composes a normalisation pass with a scoring pass and reads C twice. Beyond that, one thread block per question with the softmax over shared memory means no second kernel, no atomics and no global round trip for the similarities; K ≤ 255 is what makes the in-block softmax possible at all, and a batch is a grid over questions rather than a loop of launches.

What it is worth

bench.cu is the unfused baseline: the same block with the row norm loaded from a staged array instead of accumulated beside the dot, plus a separate norm pass. The protocol and both result sets are in docs/benchmarks/scoring-fusion/; the short version is that the benchmark found a defect the correctness checks could not.

At the 256 threads this kernel was first written with, the fusion lost: 0.82x at 255 candidates of 512 and 0.57x on Kev's shape. Achieved bandwidth was 1 to 190 GB/s against a device peak near 1000, so nothing was bandwidth-bound and there was no read to save — and a single question and a batch of eight took the same 53 µs, which is only possible if each question sits on its own SM. One thread block per question means a one-question request, the common serving case, occupies one of 128 SMs, with the norm accumulation on that block's critical path.

Raising the block to 1024 threads, the maximum on sm_89, reverses every result:

shape similarity K D questions fused unfused ratio
clm-cosine-5x512 cosine 5 512 1 9.2 µs 11.3 µs 1.22x
clm-cosine-64x512 cosine 64 512 1 12.1 µs 18.3 µs 1.51x
clm-cosine-255x512 cosine 255 512 1 23.6 µs 42.2 µs 1.79x
clm-cosine-batch8-255x512 cosine 255 512 8 23.8 µs 43.8 µs 1.84x
kev-dot-255x2560 dot 255 2560 1 55.3 µs 100.3 µs 1.81x

Kev is the control: it has no norm to fuse, so its gain is entirely the second kernel and the global round trip it does not need. That it lands at the same ~1.8x is the useful part.

Through a CUDA Graph. The ABI is written for capture, so the same launch was also replayed from a graph. It saves a flat ~3 µs — 1.49x of the 9 µs smallest shape and 1.06x of the 52 µs largest — which explains the launch-bound floor at K = 5 but does not change which configuration wins at any shape measured. A graph and the fused kernel are complementary; the two effects multiply.

Verified on hardware

RTX 4090 (sm_89), CUDA 13.0 (V13.0.88), on two driver stacks, build.sh ./out 89:

driver worst absolute difference (tolerance 1e-4)
595.58.03 1.835e-07
580.76.05 1.835e-07

gpu_parity.py runs four groups and all pass:

  • 13 fixed cases, including the degenerate ones — a single candidate, identical candidates, a zero candidate — and the one whose raw logits overflow a naive softmax.
  • The batch path: 1, 8, 4 and 3 questions with K up to 255 and D up to 256, each row compared with the reference and checked to sum to 1. The single-question call goes through cs_score_candidates and the rest through cs_score_candidates_batch, so the two entry points have to agree.
  • Rejected arguments: K above 255, K of zero, D of zero, a zero scale, a negative temperature and a null query each return an error rather than a scored answer. K above 255 is the one that matters — the shared array is sized MAX_K, so scoring the first 255 of a larger set would be a confident wrong answer instead of a failure.
  • CUDA Graph capture and replay, which the ABI was written for (it queues on the caller's stream and never synchronizes) but had never been checked: exactly one node is captured, and the replay matches a direct call exactly, not merely within tolerance.

compute-sanitizer --tool memcheck reports ERROR SUMMARY: 0 errors. test_scoring.py is 17/17 on CPU. gpu_parity.py and bench.py skip rather than fail when there is no library or no device, so a machine without a GPU reports that it could not check rather than that it passed.

Defects the GPU found, in order

  1. The parity harness passed host pointers. A numpy buffer is host memory, so the kernel dereferenced it as device memory — cudaErrorIllegalAddress, not a wrong answer. compute-sanitizer located it as an invalid global read with consecutive lanes on consecutive addresses, which is the coalesced load pattern.
  2. The query norm was counted warps times. The inner loop d = lane; d < D; d += WARP already covers every element of D across one warp's 32 lanes, but every warp computed the same partial and the partials were then summed across warps — so the norm came out sqrt(8) too large on the 8-warp block of the time. Every unnormalized case sat at float32 rounding while the normalize cases were off by ~1e-2, which is what pointed at it.
  3. The occupancy defect above. Found by the benchmark, invisible to every correctness check, and fixed by a one-line change.

The second is worth recording: the simulation did not catch it. kernel_simulation.py computed the query norm correctly because it was written from the intent rather than transcribed from the .cu. A simulation is only as good as its fidelity, and the hardware is what settles it. The third is the same lesson one level up: being correct says nothing about being fast, and only a measurement says which.

The manifest, and CI

scoring.backend.json declares status: validated, abi_version: 1 and max_abs: 1e-4 — 545x of headroom over the measured worst case, so a failure is a real failure rather than tolerance noise.

It is written to the contract #25 defines, which is not on main yet. Checked against that branch's checker and drivers: check_contract.py reports 0 errors and 0 warnings, discover_buildable.py picks the backend up for the compile matrix, and run_gpu_parity.py selects it by status: validated and runs the reference.entrypoint the manifest declares. So the CUDA workflow in that stack covers this backend with no changes here; until it merges, these checks are run by hand and the repository's ci.yml, which is Rust-only, does not run them.

Both models that need this readout share the shape of it: Kev (ThinkFlowLab#27) scores
scale * dot(k_proj(c_i), q_proj(s)) and softmaxes over a question's candidates;
CLM (ThinkFlowLab#9) scores exp(logit_scale) * cos(state_head(s), action_head(c)) and does
the same. ThinkFlowLab#9 asks for exactly this fusion and ThinkFlowLab#19's gemm.cu does not have it.

The fusion is more than fewer launches. For a cosine the candidate's norm and its
dot with the query need the same elements, so both accumulate in one read of C --
for Kev, 255 rows of 2560 floats read once instead of twice. The softmax runs
in-block over shared memory, so there is no second kernel, no atomics and no
global round trip for the similarities. A batch is a grid over questions.

MEASURED, not just intended. RTX 4090 (sm_89), driver 595.58.03, CUDA 13.0
(V13.0.88):

    worst absolute difference: 1.835e-07 (tolerance 0.0001)
    all cases within tolerance

All 13 fixed cases pass, including a single candidate, identical candidates, a
zero candidate, K=255, and logits whose raw exponential overflows.

Three defects were found and fixed; the first two only a GPU could show.

1. The parity harness passed host pointers where device pointers were expected,
   so the kernel dereferenced host memory as device memory. compute-sanitizer
   located it as an invalid global read with consecutive lanes on consecutive
   addresses. Fixed by allocating, copying and synchronising through cudart.

2. The query norm was summed across warps that had each computed the same
   partial: the inner loop already covers all of D within one warp, so the norm
   came out sqrt(8) too large on a 256-thread block. Every unnormalized case sat
   at float32 rounding while the normalize cases were off by ~1e-2, which is what
   pointed at it. Warp 0 now computes it alone.

   kernel_simulation.py did NOT catch this, because it was written from the
   intent rather than transcribed from the .cu. A simulation is only as good as
   its fidelity; hardware is what settles it.

3. The manifest claimed sm_89 while build.sh defaults to sm_80. The contract
   checker from ThinkFlowLab#25 caught that, as it caught the same class of over-claim in the
   Laya backend.

The manifest is now status=validated with a declared tolerance and a reference
entrypoint, which is what the contract requires of a backend making a parity
claim. The reference is float64 in reference.py, and gpu_parity.py skips with a
reason rather than passing when there is no GPU or no library.

Not yet exercised: the batch path beyond questions=1, the K>255 rejection, and
any architecture or CUDA version other than sm_89 on 13.0.
…ns run

The README said the `.cu` had never been compiled and singled out the
`K > 255` rejection and the batch path as unproven. All three have now run:
the kernel builds on CUDA 13.0 and matches the reference on a second driver
stack, the two entry points agree, and every rejected argument returns an
error instead of a scored answer.

Verified on an RTX 4090 (sm_89, CUDA 13.0 V13.0.88):

  driver 595.58.03   worst absolute difference 1.835e-07
  driver 580.76.05   worst absolute difference 1.835e-07   (tolerance 1e-4)

13 fixed cases, 4 batch configurations (1, 8, 4 and 3 questions, K up to 255,
D up to 256) and 6 rejected arguments all pass; `test_scoring.py` is 17/17.
…d of a phrase neither contains

The README put a sentence in ThinkFlowLab#9's mouth that it does not contain. Both issues
describe the readout precisely enough to quote directly, and the same sweep
confirms the two numbers this kernel is shaped around: ThinkFlowLab#27 gives the option
count as 1-255 and Kev's hidden size as 2560.
The tolerance note named one driver and only the fixed cases, and said nothing
about sm_89 being the only capability exercised while the manifest's own
default_arch is 80.
The kernel argued that fusing the norm into the scoring loop saves a read of
the candidate matrix. Measured against an unfused baseline that reads it twice,
it did not: 0.82x at CLM's widest shape and 0.57x on Kev's, because a
one-question request occupies one of the 128 SMs and the norm accumulation sat
on that block's critical path. Achieved bandwidth was 1 to 190 GB/s against a
device peak near 1000, so there was no read to save.

Raising the block from 256 to 1024 threads moves the work onto 32 warps instead
of 8 and reverses every result: 1.79x at 255 candidates of 512, 1.84x batched,
1.81x on Kev's plain dot. Correctness is unchanged -- the same 1.835e-07 worst
case, and compute-sanitizer reports 0 errors.

Also adds the CUDA Graph check the ABI was written for but never had: one node
captured, and the replay matches a direct call exactly.
A graph replay saves a flat ~3us of launch, which is 1.49x of the smallest
shape and 1.06x of the largest, so it does not change which configuration wins
anywhere measured -- but it explains the floor at K = 5.

Every timing in a row is now taken on one created stream. Measuring the direct
launch on the legacy default stream added about 5us at K = 255, which is larger
than the effect being measured.
…ed against

`cs_score_abi_version()` returned a literal and nothing else in the source stated
the number, so `scoring.backend.json`'s `abi_version` had no counterpart to
disagree with. ThinkFlowLab#25 now reads `#define <PREFIX>_ABI_VERSION N` out of a backend's
declared sources and requires the manifest to match; this backend was the one it
reported a warning for.

The macro is defined once and the function returns it, so the two cannot drift:
the manifest says 1, the macro says 1, and a manifest saying anything else is now
reported as `abi_version is 2 but candidate_scoring.cu defines
CS_SCORE_ABI_VERSION 1`.

No behaviour change: the function body expands to `return 1;` exactly as before,
so the compiled library is unchanged apart from line numbers. The file has not
been recompiled here -- there is no CUDA toolkit on this machine -- but the
preprocessed result is what it was.

17 tests still pass.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant