cuda: add a fused candidate-scoring kernel for CLM and Kev - #69
Draft
xiaoyu-xyz wants to merge 7 commits into
Draft
xiaoyu-xyz wants to merge 7 commits into
xiaoyu-xyz wants to merge 7 commits into
Conversation
Both models that need this readout share the shape of it: Kev (ThinkFlowLab#27) scores scale * dot(k_proj(c_i), q_proj(s)) and softmaxes over a question's candidates; CLM (ThinkFlowLab#9) scores exp(logit_scale) * cos(state_head(s), action_head(c)) and does the same. ThinkFlowLab#9 asks for exactly this fusion and ThinkFlowLab#19's gemm.cu does not have it. The fusion is more than fewer launches. For a cosine the candidate's norm and its dot with the query need the same elements, so both accumulate in one read of C -- for Kev, 255 rows of 2560 floats read once instead of twice. The softmax runs in-block over shared memory, so there is no second kernel, no atomics and no global round trip for the similarities. A batch is a grid over questions. MEASURED, not just intended. RTX 4090 (sm_89), driver 595.58.03, CUDA 13.0 (V13.0.88): worst absolute difference: 1.835e-07 (tolerance 0.0001) all cases within tolerance All 13 fixed cases pass, including a single candidate, identical candidates, a zero candidate, K=255, and logits whose raw exponential overflows. Three defects were found and fixed; the first two only a GPU could show. 1. The parity harness passed host pointers where device pointers were expected, so the kernel dereferenced host memory as device memory. compute-sanitizer located it as an invalid global read with consecutive lanes on consecutive addresses. Fixed by allocating, copying and synchronising through cudart. 2. The query norm was summed across warps that had each computed the same partial: the inner loop already covers all of D within one warp, so the norm came out sqrt(8) too large on a 256-thread block. Every unnormalized case sat at float32 rounding while the normalize cases were off by ~1e-2, which is what pointed at it. Warp 0 now computes it alone. kernel_simulation.py did NOT catch this, because it was written from the intent rather than transcribed from the .cu. A simulation is only as good as its fidelity; hardware is what settles it. 3. The manifest claimed sm_89 while build.sh defaults to sm_80. The contract checker from ThinkFlowLab#25 caught that, as it caught the same class of over-claim in the Laya backend. The manifest is now status=validated with a declared tolerance and a reference entrypoint, which is what the contract requires of a backend making a parity claim. The reference is float64 in reference.py, and gpu_parity.py skips with a reason rather than passing when there is no GPU or no library. Not yet exercised: the batch path beyond questions=1, the K>255 rejection, and any architecture or CUDA version other than sm_89 on 13.0.
…ns run The README said the `.cu` had never been compiled and singled out the `K > 255` rejection and the batch path as unproven. All three have now run: the kernel builds on CUDA 13.0 and matches the reference on a second driver stack, the two entry points agree, and every rejected argument returns an error instead of a scored answer. Verified on an RTX 4090 (sm_89, CUDA 13.0 V13.0.88): driver 595.58.03 worst absolute difference 1.835e-07 driver 580.76.05 worst absolute difference 1.835e-07 (tolerance 1e-4) 13 fixed cases, 4 batch configurations (1, 8, 4 and 3 questions, K up to 255, D up to 256) and 6 rejected arguments all pass; `test_scoring.py` is 17/17.
…d of a phrase neither contains The README put a sentence in ThinkFlowLab#9's mouth that it does not contain. Both issues describe the readout precisely enough to quote directly, and the same sweep confirms the two numbers this kernel is shaped around: ThinkFlowLab#27 gives the option count as 1-255 and Kev's hidden size as 2560.
The tolerance note named one driver and only the fixed cases, and said nothing about sm_89 being the only capability exercised while the manifest's own default_arch is 80.
The kernel argued that fusing the norm into the scoring loop saves a read of the candidate matrix. Measured against an unfused baseline that reads it twice, it did not: 0.82x at CLM's widest shape and 0.57x on Kev's, because a one-question request occupies one of the 128 SMs and the norm accumulation sat on that block's critical path. Achieved bandwidth was 1 to 190 GB/s against a device peak near 1000, so there was no read to save. Raising the block from 256 to 1024 threads moves the work onto 32 warps instead of 8 and reverses every result: 1.79x at 255 candidates of 512, 1.84x batched, 1.81x on Kev's plain dot. Correctness is unchanged -- the same 1.835e-07 worst case, and compute-sanitizer reports 0 errors. Also adds the CUDA Graph check the ABI was written for but never had: one node captured, and the replay matches a direct call exactly.
A graph replay saves a flat ~3us of launch, which is 1.49x of the smallest shape and 1.06x of the largest, so it does not change which configuration wins anywhere measured -- but it explains the floor at K = 5. Every timing in a row is now taken on one created stream. Measuring the direct launch on the legacy default stream added about 5us at K = 255, which is larger than the effect being measured.
…ed against `cs_score_abi_version()` returned a literal and nothing else in the source stated the number, so `scoring.backend.json`'s `abi_version` had no counterpart to disagree with. ThinkFlowLab#25 now reads `#define <PREFIX>_ABI_VERSION N` out of a backend's declared sources and requires the manifest to match; this backend was the one it reported a warning for. The macro is defined once and the function returns it, so the two cannot drift: the manifest says 1, the macro says 1, and a manifest saying anything else is now reported as `abi_version is 2 but candidate_scoring.cu defines CS_SCORE_ABI_VERSION 1`. No behaviour change: the function body expands to `return 1;` exactly as before, so the compiled library is unchanged apart from line numbers. The file has not been recompiled here -- there is no CUDA toolkit on this machine -- but the preprocessed result is what it was. 17 tests still pass.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
One CUDA kernel that turns a query vector and a matrix of candidate vectors into a probability distribution, without materialising the similarities in between.
Two models need this readout and neither has it:
exp(logit_scale) * cos(state_head(s), action_head(c))", answered by "the softmax over the question's candidates"scale * (k(h_opts) @ q(h_decide))withscale = 1/sqrt(256), softmaxed over that question's candidates"Neither is in #19's
gemm.cu, which is a bf16 cuBLASLt GEMM. The two differ in the similarity and in whether a projection happens first; the kernel takes the shape they share and leaves the projections to the caller, which is where #6 puts kernel selection.Core code: 123 lines (
candidate_scoring.cu). The oracle, the simulation, the parity harness, the benchmark and the tests are separate.What is fused, precisely
Not merely "fewer launches". When the similarity is a cosine — CLM's case, not Kev's, which is a plain dot — each candidate's norm and its dot with the query need the same elements, so both accumulate in a single read of
C. The unfused version composes a normalisation pass with a scoring pass and readsCtwice. Beyond that, one thread block per question with the softmax over shared memory means no second kernel, no atomics and no global round trip for the similarities;K ≤ 255is what makes the in-block softmax possible at all, and a batch is a grid over questions rather than a loop of launches.What it is worth
bench.cuis the unfused baseline: the same block with the row norm loaded from a staged array instead of accumulated beside the dot, plus a separate norm pass. The protocol and both result sets are indocs/benchmarks/scoring-fusion/; the short version is that the benchmark found a defect the correctness checks could not.At the 256 threads this kernel was first written with, the fusion lost: 0.82x at 255 candidates of 512 and 0.57x on Kev's shape. Achieved bandwidth was 1 to 190 GB/s against a device peak near 1000, so nothing was bandwidth-bound and there was no read to save — and a single question and a batch of eight took the same 53 µs, which is only possible if each question sits on its own SM. One thread block per question means a one-question request, the common serving case, occupies one of 128 SMs, with the norm accumulation on that block's critical path.
Raising the block to 1024 threads, the maximum on sm_89, reverses every result:
clm-cosine-5x512clm-cosine-64x512clm-cosine-255x512clm-cosine-batch8-255x512kev-dot-255x2560Kev is the control: it has no norm to fuse, so its gain is entirely the second kernel and the global round trip it does not need. That it lands at the same ~1.8x is the useful part.
Through a CUDA Graph. The ABI is written for capture, so the same launch was also replayed from a graph. It saves a flat ~3 µs — 1.49x of the 9 µs smallest shape and 1.06x of the 52 µs largest — which explains the launch-bound floor at K = 5 but does not change which configuration wins at any shape measured. A graph and the fused kernel are complementary; the two effects multiply.
Verified on hardware
RTX 4090 (sm_89), CUDA 13.0 (V13.0.88), on two driver stacks,
build.sh ./out 89:gpu_parity.pyruns four groups and all pass:Kup to 255 andDup to 256, each row compared with the reference and checked to sum to 1. The single-question call goes throughcs_score_candidatesand the rest throughcs_score_candidates_batch, so the two entry points have to agree.Kabove 255,Kof zero,Dof zero, a zero scale, a negative temperature and a null query each return an error rather than a scored answer.Kabove 255 is the one that matters — the shared array is sizedMAX_K, so scoring the first 255 of a larger set would be a confident wrong answer instead of a failure.compute-sanitizer --tool memcheckreportsERROR SUMMARY: 0 errors.test_scoring.pyis 17/17 on CPU.gpu_parity.pyandbench.pyskip rather than fail when there is no library or no device, so a machine without a GPU reports that it could not check rather than that it passed.Defects the GPU found, in order
cudaErrorIllegalAddress, not a wrong answer.compute-sanitizerlocated it as an invalid global read with consecutive lanes on consecutive addresses, which is the coalesced load pattern.warpstimes. The inner loopd = lane; d < D; d += WARPalready covers every element ofDacross one warp's 32 lanes, but every warp computed the same partial and the partials were then summed across warps — so the norm came outsqrt(8)too large on the 8-warp block of the time. Every unnormalized case sat at float32 rounding while the normalize cases were off by ~1e-2, which is what pointed at it.The second is worth recording: the simulation did not catch it.
kernel_simulation.pycomputed the query norm correctly because it was written from the intent rather than transcribed from the.cu. A simulation is only as good as its fidelity, and the hardware is what settles it. The third is the same lesson one level up: being correct says nothing about being fast, and only a measurement says which.The manifest, and CI
scoring.backend.jsondeclaresstatus: validated,abi_version: 1andmax_abs: 1e-4— 545x of headroom over the measured worst case, so a failure is a real failure rather than tolerance noise.It is written to the contract #25 defines, which is not on
mainyet. Checked against that branch's checker and drivers:check_contract.pyreports 0 errors and 0 warnings,discover_buildable.pypicks the backend up for the compile matrix, andrun_gpu_parity.pyselects it bystatus: validatedand runs thereference.entrypointthe manifest declares. So the CUDA workflow in that stack covers this backend with no changes here; until it merges, these checks are run by hand and the repository'sci.yml, which is Rust-only, does not run them.