You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
jaredpalmer/kev-4b
(sha 139fdd94f1b6a6ad80cc15e08fcb99cac885a101, Apache-2.0, 9,662 downloads as
of 2026-09-28), code at jaredpalmer/kev.
The family spans 0.5B to 9B; this request covers 4B first.
Unlike LAYA and Cua-S1, the published repository is a PEFT adapter, not a full
checkpoint: it has no config.json and no model.safetensors, only adapter_model.safetensors (fp32, 129.9 MB, r=16, α=32, 496 tensors), head.pt (fp32, 5.25 MB) and the tokenizer. The base model is Qwen/Qwen3.5-4B-Base at 1001bb4d826a52d1f399e183466143f4da7b741b. Weights are bf16, 9.32 GB total,
8.41 GB for the text backbone.
The closest planned or implemented model
Cua-S1 4B 0.2 (#10) is
the closest, and deliberately so. Both use a Qwen3.5 backbone whose text configs
are identical value-for-value: 32 layers, full_attention at indices [3,7,11,15,19,23,27,31] (8 layers) and 24 Gated DeltaNet layers, head_dim 256,
16 query / 4 KV heads, linear_conv_kernel_dim 4, partial rotary 0.25, hidden
2560, intermediate 9216.
So this is the case #9 asks about: a second model shaped differently from the
first two, sharing kernels rather than duplicating them. Reading PR #19, the qwen3_5 kernels are consumed through libloading and a C ABI that knows nothing
about which model it serves, so I expect all nine kernel files to transfer
unchanged. What does not transfer is the engine and the readout.
Three differences from Cua-S1 that matter:
The adapter covers the Gated DeltaNet layers. Cua-S1's LoRA targets 7
modules (q/k/v/o_proj + 3 MLPs), which is why [New Model]: Support Cua-S1 4B 0.2 — community help wanted #10 and cua_s1: add a native CUDA text worker #19 can say the GDN
layers run on base weights. Kev's targets 12, including all five GDN
projections (in_proj_qkv, in_proj_a, in_proj_z, in_proj_b, out_proj), verified across all 24 linear layers in the adapter header. The
merge is mandatory there; copying the "GDN runs on base weights" assumption
would silently produce wrong decisions.
The readout is a trained bilinear pointer head, not a letter-row argmax. head.pt holds q and k as [256, 2560] fp32 plus biases (1,311,232
parameters, temperature 2.406050072164233 embedded): logits are scale * (k(h_opts) @ q(h_decide)) with scale = 1/sqrt(256), softmaxed over
that question's candidates. With tie_word_embeddings: true, Cua-S1's 26
"letter rows" are sliced from the base model's own embedding rather than
trained, so the two readouts are not variations of one mechanism.
Several questions share one prefill.is_hybrid() returns true for
Qwen3.5, which makes the packed block-causal layout unavailable and the row
form mandatory: state once, then one branch per question with positions
restarting at the state length. cs1_attention has no mask parameter, so this
is a constraint, not a preference. It also means the state is re-prefilled per
question, which is the cost Gated DeltaNet prefix-state reuse would remove.
Also: the option count is per-request (1–255, score is L levels, noul is ["false","true"]), against Cua-S1's fixed 26, and Kev has no chat template — the
prompt is assembled in kev/model.py:encode() from five reused Qwen special
tokens, with <|name|> rewritten to <¦name¦> so a caller cannot forge option
boundaries.
The delimiter ids must come from the base tokenizer, not the adapter repo
jaredpalmer/kev-4b/added_tokens.json maps the five delimiters to 151648–151661, and that range is wrong: it is the Qwen2.5 special-token
range, and Qwen/Qwen3.5-4B-Base has no token at all between 151600 and
151700. Anyone implementing this from the adapter repository alone would index
unrelated embedding rows and get silently wrong decisions, with no error.
The ids the code actually resolves, via AutoTokenizer.from_pretrained:
Delimiter
id
Role
`<
box_start
>`
`<
box_end
>`
`<
fim_prefix
>`
`<
fim_middle
>`
`<
fim_suffix
>`
Confirmed against two independent files that agree — Qwen/Qwen3.5-4B-Base's tokenizer.jsonadded_tokens and its tokenizer_config.json added_tokens_decoder. kev-4b's own tokenizer.json agrees too, but its tokenizer_config.json declares no added tokens at all, so the adapter
repository cannot be treated as authoritative. I would pin these five ids in the
implementation and assert them at load time, because the failure mode is silent.
What's your difficulty of supporting the model you want?
Lower than #10's, because the kernels already exist. No new operators and no new
CUDA backend are needed. The work is a model engine under src/models/kev/ plus
a tokenizer/prompt layer: apply the merge, build the row form, read hidden states
at the <|box_end|> / <|fim_suffix|> delimiter rows, and implement the
1.31M-parameter pointer head. The head is three orders of magnitude smaller than
the backbone and is expressible with the existing cs1_gemm, so it does not
justify a kernel; it should stay fp32, following #19's precedent for the readout.
Two things I would want agreed before code: whether qwen3_5 lists both
consumers in a backend manifest (models: ["src/models/cua_s1/", "src/models/kev/"]), and which revisions are pinned as the parity reference,
since Kev's own kev/model.py has to serve as it, the way #19 used FourBModel.
Use case and motivation
Kev is a Jev-like decision model family that runs on your own hardware, answering
the same choice / score / noul questions /v1/systemone already serves, and
it is the most-downloaded Qwen3.5-based decision model in #9's table. It is the
right second test of the model/backend split because it stresses exactly the parts
LAYA and Cua-S1 do not: cross-question state, a dynamic candidate set, and a
readout that is neither a classifier head nor a tied-embedding slice. If one
backend directory can serve both Cua-S1 and Kev, the reuse story is demonstrated
rather than asserted.
Related: #9 (umbrella), #10 (Cua-S1 4B 0.2), #19 (qwen3_5 kernels), #6 (layout).
Upstream context: EricLBuehler/mistral.rs issue #2443 proposes a hidden-states
prefill path and Gated DeltaNet prefix-state caching for Kev — an issue, not
merged code, and its accuracy claims are unverified by me.
The model to consider
jaredpalmer/kev-4b(sha
139fdd94f1b6a6ad80cc15e08fcb99cac885a101, Apache-2.0, 9,662 downloads asof 2026-09-28), code at
jaredpalmer/kev.The family spans 0.5B to 9B; this request covers 4B first.
Unlike LAYA and Cua-S1, the published repository is a PEFT adapter, not a full
checkpoint: it has no
config.jsonand nomodel.safetensors, onlyadapter_model.safetensors(fp32, 129.9 MB, r=16, α=32, 496 tensors),head.pt(fp32, 5.25 MB) and the tokenizer. The base model isQwen/Qwen3.5-4B-Baseat1001bb4d826a52d1f399e183466143f4da7b741b. Weights are bf16, 9.32 GB total,8.41 GB for the text backbone.
The closest planned or implemented model
Cua-S1 4B 0.2 (#10) is
the closest, and deliberately so. Both use a Qwen3.5 backbone whose text configs
are identical value-for-value: 32 layers,
full_attentionat indices[3,7,11,15,19,23,27,31](8 layers) and 24 Gated DeltaNet layers,head_dim256,16 query / 4 KV heads,
linear_conv_kernel_dim4, partial rotary0.25, hidden2560, intermediate 9216.
So this is the case
#9asks about: a second model shaped differently from thefirst two, sharing kernels rather than duplicating them. Reading PR #19, the
qwen3_5kernels are consumed throughlibloadingand a C ABI that knows nothingabout which model it serves, so I expect all nine kernel files to transfer
unchanged. What does not transfer is the engine and the readout.
Three differences from Cua-S1 that matter:
modules (
q/k/v/o_proj+ 3 MLPs), which is why [New Model]: Support Cua-S1 4B 0.2 — community help wanted #10 and cua_s1: add a native CUDA text worker #19 can say the GDNlayers run on base weights. Kev's targets 12, including all five GDN
projections (
in_proj_qkv,in_proj_a,in_proj_z,in_proj_b,out_proj), verified across all 24 linear layers in the adapter header. Themerge is mandatory there; copying the "GDN runs on base weights" assumption
would silently produce wrong decisions.
head.ptholdsqandkas[256, 2560]fp32 plus biases (1,311,232parameters,
temperature2.406050072164233 embedded): logits arescale * (k(h_opts) @ q(h_decide))withscale = 1/sqrt(256), softmaxed overthat question's candidates. With
tie_word_embeddings: true, Cua-S1's 26"letter rows" are sliced from the base model's own embedding rather than
trained, so the two readouts are not variations of one mechanism.
is_hybrid()returns true forQwen3.5, which makes the packed block-causal layout unavailable and the row
form mandatory: state once, then one branch per question with positions
restarting at the state length.
cs1_attentionhas no mask parameter, so thisis a constraint, not a preference. It also means the state is re-prefilled per
question, which is the cost Gated DeltaNet prefix-state reuse would remove.
Also: the option count is per-request (1–255,
scoreis L levels,noulis["false","true"]), against Cua-S1's fixed 26, and Kev has no chat template — theprompt is assembled in
kev/model.py:encode()from five reused Qwen specialtokens, with
<|name|>rewritten to<¦name¦>so a caller cannot forge optionboundaries.
The delimiter ids must come from the base tokenizer, not the adapter repo
jaredpalmer/kev-4b/added_tokens.jsonmaps the five delimiters to151648–151661, and that range is wrong: it is the Qwen2.5 special-token
range, and
Qwen/Qwen3.5-4B-Basehas no token at all between 151600 and151700. Anyone implementing this from the adapter repository alone would index
unrelated embedding rows and get silently wrong decisions, with no error.
The ids the code actually resolves, via
AutoTokenizer.from_pretrained:Confirmed against two independent files that agree —
Qwen/Qwen3.5-4B-Base'stokenizer.jsonadded_tokensand itstokenizer_config.jsonadded_tokens_decoder. kev-4b's owntokenizer.jsonagrees too, but itstokenizer_config.jsondeclares no added tokens at all, so the adapterrepository cannot be treated as authoritative. I would pin these five ids in the
implementation and assert them at load time, because the failure mode is silent.
What's your difficulty of supporting the model you want?
Lower than #10's, because the kernels already exist. No new operators and no new
CUDA backend are needed. The work is a model engine under
src/models/kev/plusa tokenizer/prompt layer: apply the merge, build the row form, read hidden states
at the
<|box_end|>/<|fim_suffix|>delimiter rows, and implement the1.31M-parameter pointer head. The head is three orders of magnitude smaller than
the backbone and is expressible with the existing
cs1_gemm, so it does notjustify a kernel; it should stay fp32, following #19's precedent for the readout.
Two things I would want agreed before code: whether
qwen3_5lists bothconsumers in a backend manifest (
models: ["src/models/cua_s1/", "src/models/kev/"]), and which revisions are pinned as the parity reference,since Kev's own
kev/model.pyhas to serve as it, the way #19 usedFourBModel.Use case and motivation
Kev is a Jev-like decision model family that runs on your own hardware, answering
the same
choice/score/noulquestions/v1/systemonealready serves, andit is the most-downloaded Qwen3.5-based decision model in #9's table. It is the
right second test of the model/backend split because it stresses exactly the parts
LAYA and Cua-S1 do not: cross-question state, a dynamic candidate set, and a
readout that is neither a classifier head nor a tied-embedding slice. If one
backend directory can serve both Cua-S1 and Kev, the reuse story is demonstrated
rather than asserted.
Related: #9 (umbrella), #10 (Cua-S1 4B 0.2), #19 (
qwen3_5kernels), #6 (layout).Upstream context:
EricLBuehler/mistral.rsissue #2443 proposes a hidden-statesprefill path and Gated DeltaNet prefix-state caching for Kev — an issue, not
merged code, and its accuracy claims are unverified by me.