Skip to content

[New Model]: Support Kev (0.5B–9B) — second Qwen3.5 decision model, reusing the qwen3_5 CUDA backend #27

Description

@xiaoyu-xyz

The model to consider

jaredpalmer/kev-4b
(sha 139fdd94f1b6a6ad80cc15e08fcb99cac885a101, Apache-2.0, 9,662 downloads as
of 2026-09-28), code at jaredpalmer/kev.
The family spans 0.5B to 9B; this request covers 4B first.

Unlike LAYA and Cua-S1, the published repository is a PEFT adapter, not a full
checkpoint
: it has no config.json and no model.safetensors, only
adapter_model.safetensors (fp32, 129.9 MB, r=16, α=32, 496 tensors),
head.pt (fp32, 5.25 MB) and the tokenizer. The base model is
Qwen/Qwen3.5-4B-Base at
1001bb4d826a52d1f399e183466143f4da7b741b. Weights are bf16, 9.32 GB total,
8.41 GB for the text backbone.

The closest planned or implemented model

Cua-S1 4B 0.2 (#10) is
the closest, and deliberately so.
Both use a Qwen3.5 backbone whose text configs
are identical value-for-value: 32 layers, full_attention at indices
[3,7,11,15,19,23,27,31] (8 layers) and 24 Gated DeltaNet layers, head_dim 256,
16 query / 4 KV heads, linear_conv_kernel_dim 4, partial rotary 0.25, hidden
2560, intermediate 9216.

So this is the case #9 asks about: a second model shaped differently from the
first two, sharing kernels rather than duplicating them.
Reading PR #19, the
qwen3_5 kernels are consumed through libloading and a C ABI that knows nothing
about which model it serves, so I expect all nine kernel files to transfer
unchanged. What does not transfer is the engine and the readout.

Three differences from Cua-S1 that matter:

  1. The adapter covers the Gated DeltaNet layers. Cua-S1's LoRA targets 7
    modules (q/k/v/o_proj + 3 MLPs), which is why [New Model]: Support Cua-S1 4B 0.2 — community help wanted #10 and cua_s1: add a native CUDA text worker #19 can say the GDN
    layers run on base weights. Kev's targets 12, including all five GDN
    projections
    (in_proj_qkv, in_proj_a, in_proj_z, in_proj_b,
    out_proj), verified across all 24 linear layers in the adapter header. The
    merge is mandatory there; copying the "GDN runs on base weights" assumption
    would silently produce wrong decisions.
  2. The readout is a trained bilinear pointer head, not a letter-row argmax.
    head.pt holds q and k as [256, 2560] fp32 plus biases (1,311,232
    parameters, temperature 2.406050072164233 embedded): logits are
    scale * (k(h_opts) @ q(h_decide)) with scale = 1/sqrt(256), softmaxed over
    that question's candidates. With tie_word_embeddings: true, Cua-S1's 26
    "letter rows" are sliced from the base model's own embedding rather than
    trained, so the two readouts are not variations of one mechanism.
  3. Several questions share one prefill. is_hybrid() returns true for
    Qwen3.5, which makes the packed block-causal layout unavailable and the row
    form mandatory: state once, then one branch per question with positions
    restarting at the state length. cs1_attention has no mask parameter, so this
    is a constraint, not a preference. It also means the state is re-prefilled per
    question, which is the cost Gated DeltaNet prefix-state reuse would remove.

Also: the option count is per-request (1–255, score is L levels, noul is
["false","true"]), against Cua-S1's fixed 26, and Kev has no chat template — the
prompt is assembled in kev/model.py:encode() from five reused Qwen special
tokens, with <|name|> rewritten to <¦name¦> so a caller cannot forge option
boundaries.

The delimiter ids must come from the base tokenizer, not the adapter repo

jaredpalmer/kev-4b/added_tokens.json maps the five delimiters to
151648–151661, and that range is wrong: it is the Qwen2.5 special-token
range, and Qwen/Qwen3.5-4B-Base has no token at all between 151600 and
151700
. Anyone implementing this from the adapter repository alone would index
unrelated embedding rows and get silently wrong decisions, with no error.

The ids the code actually resolves, via AutoTokenizer.from_pretrained:

Delimiter id Role
`< box_start >`
`< box_end >`
`< fim_prefix >`
`< fim_middle >`
`< fim_suffix >`

Confirmed against two independent files that agree — Qwen/Qwen3.5-4B-Base's
tokenizer.json added_tokens and its tokenizer_config.json
added_tokens_decoder. kev-4b's own tokenizer.json agrees too, but its
tokenizer_config.json declares no added tokens at all, so the adapter
repository cannot be treated as authoritative. I would pin these five ids in the
implementation and assert them at load time, because the failure mode is silent.

What's your difficulty of supporting the model you want?

Lower than #10's, because the kernels already exist. No new operators and no new
CUDA backend are needed.
The work is a model engine under src/models/kev/ plus
a tokenizer/prompt layer: apply the merge, build the row form, read hidden states
at the <|box_end|> / <|fim_suffix|> delimiter rows, and implement the
1.31M-parameter pointer head. The head is three orders of magnitude smaller than
the backbone and is expressible with the existing cs1_gemm, so it does not
justify a kernel; it should stay fp32, following #19's precedent for the readout.

Two things I would want agreed before code: whether qwen3_5 lists both
consumers in a backend manifest (models: ["src/models/cua_s1/", "src/models/kev/"]), and which revisions are pinned as the parity reference,
since Kev's own kev/model.py has to serve as it, the way #19 used FourBModel.

Use case and motivation

Kev is a Jev-like decision model family that runs on your own hardware, answering
the same choice / score / noul questions /v1/systemone already serves, and
it is the most-downloaded Qwen3.5-based decision model in #9's table. It is the
right second test of the model/backend split because it stresses exactly the parts
LAYA and Cua-S1 do not: cross-question state, a dynamic candidate set, and a
readout that is neither a classifier head nor a tied-embedding slice. If one
backend directory can serve both Cua-S1 and Kev, the reuse story is demonstrated
rather than asserted.

Related: #9 (umbrella), #10 (Cua-S1 4B 0.2), #19 (qwen3_5 kernels), #6 (layout).
Upstream context: EricLBuehler/mistral.rs issue #2443 proposes a hidden-states
prefill path and Gated DeltaNet prefix-state caching for Kev — an issue, not
merged code, and its accuracy claims are unverified by me.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions