Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
61 commits
Select commit Hold shift + click to select a range
62e388f
dasLLAMA: IQ4_XS native tier, CPU slice - KqFmt.iq4xs (schema id 44),…
borisbat Aug 30, 2026
5a48b4a
dasLLAMA: IQ4_XS end to end on CPU - the repack-mr freeze and worker-…
borisbat Aug 30, 2026
2de34d9
dasLLAMA: IQ4_XS JIT emitter - emit_block_iq4xs (mx4 LUT decode + sig…
borisbat Aug 30, 2026
7d27711
dasLLAMA: IQ4_XS on the Vulkan tier - KqGemvIq4xs + KqBatchIq4xs (cod…
borisbat Aug 30, 2026
f80fe16
dasLLAMA HOW_TO: section 7 (Metal) rewritten from the census - the ti…
borisbat Aug 30, 2026
3f771db
dasLLAMA: IQ4_XS on the Metal tier - the k6 split scale form, iq4_lut…
borisbat Aug 30, 2026
0373640
dasLLAMA: Q3_K native tier, CPU slice - KqFmt.k3 (id 3), the k6-shape…
borisbat Aug 30, 2026
61d7731
dasLLAMA: Q3_K JIT emitter - a k3 flag through emit_block_kqv2's k6 a…
borisbat Aug 30, 2026
1cfcda5
dasLLAMA: Q3_K on the Vulkan tier - KqGemvK3 + KqBatchK3 (k6's classe…
borisbat Aug 30, 2026
0f60e3d
dasLLAMA: Q3_K on the Metal tier - the k6 split scale form verbatim, …
borisbat Aug 30, 2026
a72e4c5
dasLLAMA: IQ4_XS Vulkan GEMV at parity - the codebook as four packed …
borisbat Aug 30, 2026
d08d415
dasLLAMA: IQ4_XS Metal GEMV at parity - the codebook in threadgroup m…
borisbat Aug 30, 2026
749525d
dasLLAMA: IQ4_XS Metal prefill above parity - the mul_mm arm reads th…
borisbat Aug 30, 2026
05d83f3
dasLLAMA: the nine cm2 tiles stamp from one class template (followup_…
borisbat Aug 30, 2026
bad0b77
dasLLAMA: lint - the iq4xs GEMV header trimmed to the comment cap; ST…
borisbat Aug 30, 2026
d19602e
dasLLAMA: iq4xs/k3/k5/q40 join the cm2 f16 feed - four decode methods…
borisbat Aug 30, 2026
6c7ff33
dasLLAMA: IQ3_S joins the kq lattice (CPU slice) - 64/64 greedy ids v…
borisbat Aug 30, 2026
f063e55
dasLLAMA: the IQ3_S JIT emitter - panel-route tile + gather gemv, pp5…
borisbat Aug 30, 2026
d90e3a8
dasLLAMA: IQ3_S on the Vulkan tier - workgroup-staged grid, 64/64 ids…
borisbat Aug 30, 2026
52d88d2
dasLLAMA: the IQ3_S cm2 tile - a gated grid axis on the template, pp5…
borisbat Aug 30, 2026
cf8c3b6
dasMetal: literal fixed-array locals hoist to program-scope constant …
borisbat Aug 31, 2026
cda7a2c
dasLLAMA: IQ3_S on the Metal tier - constant-table grid, 4-row f4 GEM…
borisbat Aug 31, 2026
c208d0f
dasLLAMA: IQ3_XXS on the CPU tier - the exact halving fold, the share…
borisbat Aug 31, 2026
0e3a8de
dasLLAMA: the IQ3_XXS JIT emitter - a gather arm on the shared panel …
borisbat Aug 31, 2026
0a5add6
dasLLAMA: IQ3_XXS on the Vulkan tier - halved grid, parity signs, no …
borisbat Aug 31, 2026
fd1bf95
dasLLAMA: IQ3_XXS on the Metal tier - the iq3s shapes over the halved…
borisbat Aug 31, 2026
757f3ec
dasLLAMA: IQ4_NL on the CPU tier + JIT - q40's planes, the codebook, …
borisbat Aug 31, 2026
6db1e2c
dasLLAMA: IQ4_NL on the Vulkan tier - pure composition, zero new decodes
borisbat Aug 31, 2026
a6945e2
dasLLAMA: IQ4_NL on the Metal tier - the iq4xs kernels with the scale…
borisbat Aug 31, 2026
95ec35f
dasLLAMA: Q2_K native tier, CPU slice - KqFmt.k2 (id 2, stream code 2…
borisbat Aug 31, 2026
89059a5
dasLLAMA: Q2_K JIT emitter - a fourth kqv2 arm (per-16 nibble scales,…
borisbat Aug 31, 2026
a82599e
dasLLAMA: Q2_K on the Vulkan tier - pair-byte nibble scales on the k3…
borisbat Aug 31, 2026
6f8e45d
dasLLAMA: Q2_K on the Metal tier - the k3 shells minus the hmask, for…
borisbat Aug 31, 2026
d70a4af
dasLLAMA: drop the stale STYLE038 tag on metal_kq_mv_k2 (the hmask-fr…
borisbat Aug 31, 2026
0c65165
dasLLAMA: IQ2_S joins the kq lattice (CPU slice) - the u64-grid tier …
borisbat Aug 31, 2026
9aa0bec
dasLLAMA: IQ2_S Phase B - the JIT emitter rides the panel route, the …
borisbat Aug 31, 2026
48f8516
dasLLAMA: IQ2_S Phase C - the u64 grid crosses to Vulkan, the first 8…
borisbat Aug 31, 2026
d99ceda
dasLLAMA: IQ2_S Phase D - Metal closes the format, the 8 KB grid ride…
borisbat Aug 31, 2026
068b978
dasLLAMA: IQ2_XS joins the kq lattice (CPU slice) - and the dead k2 r…
borisbat Aug 31, 2026
141944c
dasLLAMA: IQ2_XS Phase B - the JIT emitter rides the fmt-23 panel route
borisbat Aug 31, 2026
9356439
llvm_tune: --tune-only <tokens> - the one-family re-mint
borisbat Aug 31, 2026
f2cf3a2
dasLLAMA: adopt the fast JIT dev loop; ledger the cache-invalidation …
borisbat Aug 31, 2026
c8d912b
dasLLAMA: IQ2_XS Phase C - the u64 grid crosses to Vulkan on ksigns-b…
borisbat Aug 31, 2026
9b49748
dasLLAMA: IQ2_XS Phase D - Metal closes the format on all four tiers
borisbat Aug 31, 2026
810c539
dasLLAMA: IQ2_XXS joins the kq lattice (CPU slice) - the LAST format
borisbat Aug 31, 2026
e4190da
dasLLAMA: IQ2_XXS Phase B - the panel route absorbs the aux32 form wi…
borisbat Aug 31, 2026
5615454
dasLLAMA: trim the iq2xxs gather comment to the 3-line cap (STYLE014)
borisbat Aug 31, 2026
80070e2
dasLLAMA: IQ2_XXS Phase C - the iq3xxs Vulkan shell over the two-word…
borisbat Aug 31, 2026
66b088d
dasLLAMA: IQ2_XXS Phase D - Metal closes the format, and the format l…
borisbat Aug 31, 2026
ce7d8d8
dasLLAMA: the tune race keeps one seat per ISA tier (unquirk pass, B1)
borisbat Sep 1, 2026
1e40679
dasLLAMA: followup_metal.md opens with the dark smmla leg
borisbat Sep 1, 2026
d736dfc
llvm_tune: shipped defaults profiles + DAS_TUNE_POLICY=reference (unq…
borisbat Sep 1, 2026
4d1c85b
dasLLAMA: the SDK bundle carries the shipped defaults profiles
borisbat Sep 1, 2026
28b883f
dasLLAMA: arm-neon + x86-vnni512 profiles, the tune_kernels Metal-arm…
borisbat Sep 1, 2026
8b241be
dasLLAMA: the .dlim identity carries the pack-code version; JIT stagi…
borisbat Sep 1, 2026
b1fb8c1
PR-1 comment harvest: rules, facts and citations land; the tri-state …
borisbat Sep 1, 2026
c0f8f3f
PR-1 audit round: the Metal-blob split-plane reads, the f32 fallback,…
borisbat Sep 1, 2026
e86c139
the fresh dragon's wording repairs across the six checklists
borisbat Sep 1, 2026
d9a3b7e
PR-1 gate round: the emitter's renamed-table probe, the IQ3_XXS trim …
borisbat Sep 1, 2026
7dbbf2e
the benchmarks checklist's last dragon pass: the zero-rows rule keyed…
borisbat Sep 1, 2026
0c29ca0
the constant-table hoist names what it accepts: a fixed_array of exac…
borisbat Sep 1, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -162,6 +162,7 @@ Task-specific instructions are split into skill files under `skills/`. You MUST
| `skills/daslang/references/queries.md` | Filter/map/sort/group/aggregate transforms - comprehension -> linq_boost -> plain `for`; avoid `daslib/functional` for new code |
| `skills/decs.md` | Programming with `daslib/decs` / `decs_boost` - entities, components, queries, `[decs_template]`, stages |
| `skills/internal/aot_hash_desync_debugging.md` | `error[50101]: AOT link failed` - semantic-hash desync diagnostics |
| `modules/dasLLAMA/CLAUDE.md` | Any work under `modules/dasLLAMA/` - the module's HOW_TO series (`HOW_TO_ADD_A_FORMAT.md` for a new weight format) and its architecture/review set |

Multiple skill files may apply to one task: creating a new daslib module needs `skills/das_formatting.md`, `skills/daslib_modules.md`, and possibly `skills/internal/documentation_rst.md`.

Expand Down
2 changes: 1 addition & 1 deletion modules/dasLLAMA/ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,7 +38,7 @@ re-transcoding `$LCPP/src/unicode-data.cpp`).

- `ARCHITECTURE_IMAGE.md` - sec.2.1-2.1i: the prepared-image rail, the baked dev-W f16 plane,
and the baked tower twin-W plane.
- `ARCHITECTURE_GPU.md` - sec.2.2b, 2.2w-2.2x: the tensor-GEMM and fused-attention shapes that
- `ARCHITECTURE_GPU.md` - sec.2.2b, 2.2w-2.2y: the tensor-GEMM and fused-attention shapes that
measured out, the tower attention routes, and the tower driver's encode chains.
- `ARCHITECTURE_GPU_PREFILL.md` - sec.2.2c-2.2i, 2.2u-2.2v: the Metal prefill driver's GEMM form
ladder, dev-W knee map, attention slab, MoE bucket rail, chunked submission, the f16 twin
Expand Down
5 changes: 4 additions & 1 deletion modules/dasLLAMA/ARCHITECTURE_ENGINE.md
Original file line number Diff line number Diff line change
Expand Up @@ -78,7 +78,10 @@ stay the reviewer's. A mis-numbered arm dispatches, reads the wrong buffer, and
image plus its element count; the image owns the bytes, a carrier owns nothing but its backing.
Requires nothing in dasllama - the image rail binds planes, every carrier holds them.
- **`dasllama_kqformat.das`** - format IDENTITY: the `KqFmt` enum, the per-format descriptor table
(plane strides, block geometry, stream codes), format predicates. It requires nothing else in
(plane strides, block geometry, stream codes), format predicates, and the shared decode
tables the grid and codebook formats key off - each as a builder function (`iq3s_grid()`,
`iq4nl_lut()`) for kernels that may run on a team lane, plus a global twin for tests,
oracles and the emitter's constant bake. It requires nothing else in
dasllama, because it is the taxonomy everything keys off. ONE id space - the enum; integer ids
exist only at the IR/kernel-param boundary. `kq_sb` is the superblock-lattice predicate: a
`fmt != q8` test does not imply the lattice, so branch on the predicate.
Expand Down
16 changes: 16 additions & 0 deletions modules/dasLLAMA/ARCHITECTURE_GPU.md
Original file line number Diff line number Diff line change
Expand Up @@ -256,3 +256,19 @@ specification and the CPU-vs-GPU transcript cells are its parity instrument. Eve
best-effort: it answers false (or -1) on any shape, knob, quant-mode or device decline, and the
CPU chain serves that encode. Engage is read from counter deltas (`metal_tower_stats`,
`metal_tower_f16_encodes`), never from "the model ran".

### 2.2y The Metal kq split scale plane {#metal-kq-split-scale-plane}

Every superblock format but k4, k5, q40 and iq4nl stores its Metal-blob scale row SPLIT into two
regions of one buffer: the 16-byte sub-scale strips of every superblock first, then the packed
per-superblock d tail. A kernel binds that one buffer twice - the strips at `soff = sb0 * 16` and
the tail at `doff = nsb * 16 + sb0 * 2` - so the two reads stride independently and the strip
read stays 16-byte aligned. k2 is the one shape variation: its tail is 4 bytes per superblock
(`nsb * 16 + sb0 * 4`), because it carries d and dmin. `kq_scales_of` builds the pair;
`metal_blob_scale_plane` mints it at bake time, folding each format's 20-byte decoded row into
`[16B strips][2B d]` (k3's row is 18 bytes and is already in that shape). The 2-byte tail is why
a region's bind offset must be a multiple of 512 elements - the `(off/256)*2` d-plane bind is
4-byte aligned only then - which is what `metal_blob_off_ok` and `moe_site_ok` check. iq4nl is
the exception: it reuses q40's 16-byte plane of eight f16 d per superblock, binds once, and
ignores `doff`. The Vulkan tier does not use this form - it binds the decoded 20-byte row as five
uints per superblock.
26 changes: 22 additions & 4 deletions modules/dasLLAMA/ARCHITECTURE_GPU_VULKAN.md
Original file line number Diff line number Diff line change
Expand Up @@ -76,8 +76,11 @@ compiler pattern-matches only one spelling into that path: a 16-bit load (`int16
members) followed by `unpack8(w)[i & 1u]` - a byte2 lane select - with sub-fields pulled out by
shift and mask. A 32-bit word with a variable shift runs slower; an `unpack8` of a 32-bit word
indexed by a runtime value (a byte4 dynamic select) drops the whole kernel off the block-load
path, to about a third of the rate. Every cm2 decode - q8, Q4_K, Q6_K - is spelled the 16-bit
way, which is why the block structs are `int16` arrays over the same bytes.
path, to about a third of the rate. Every cm2 decode - q8 and the six kq superblock formats -
is spelled the 16-bit way, which is why the block structs are `int16` arrays over the same
bytes. The IQ4_XS codebook is the one runtime-indexed read a decode makes: it is staged into a
16-entry `@workgroup` f16 table ahead of the tile loop (the reference exe's shared-memory table-staging form),
never selected out of a register vector per element.

### 2.2l The cm2 tile pick and the coopmat default ladder {#cm2-tile-pick-and-default}

Expand All @@ -98,8 +101,8 @@ and clamps only the store, so every f16 plane the chain feeds it - the gathered
image and the hidden plane - is sized with 32 rows of slack past its last region
(`ffn_cm2_chunk_rows`).

**The f16 feed admits exactly three weight formats - q8, Q4_K and Q6_K** - the same set the cm2
decode callbacks cover (sec.2.2k) - and each (format, tile) pair has ONE generated class. The
**The f16 feed admits q8 and every kq superblock format** (`kq_sb`) - the set the cm2 decode
callbacks cover (sec.2.2k) - and each (format, tile) pair has ONE stamped class. The
prefill driver reaches them through one dispatcher per stage (`cm2_cls_ensure`, `cm2_cls_set`,
`cm2_cls_enc`), all three keyed on the same `(fmt, ml)` pair, so the pipeline a role ensures,
the set it binds and the kernel it encodes can never be three different classes. The decode
Expand All @@ -111,6 +114,21 @@ NV_cooperative_matrix2, else mm where it has KHR_cooperative_matrix, else sdot4;
the extension lands on mm. The same resolver stamps the mode into the `.dlim` flavor
configuration, so the recorded mode and the running mode cannot drift.

**The tile's fast path is what makes the loads unclamped.** It runs when the weight tile is
whole (`m0 + 128 <= d`), the token column is whole or stamped s, and K is a whole number of BK
steps; the layouts are then created clamp-Undefined and the B and output strides are masked to
a multiple of 8 f16 (`stride &= ~7`). The mask is an identity on today's shapes - `n` and `d`
are 32-multiples - and it exists to make the alignment PROVABLE to the driver's address
analysis, which is what keeps the loads on the wide path. The s column gates only the weight
tile: its partial token column loads unclamped and its store clamps. Everything else takes the
edge path with clamped layouts.

**The no-split arm keeps literal loop bounds and a literal store base.** Where `ksplit` is zero
the k loop runs the literal `0 .. n` with the store at the row base rather than the general
`k0`/`k1`/`ybase` form, although those values are exactly `0`, `n` and `0` on that path: the
general spelling cost 27% of prefill throughput (`benchmarks/lcpp_bench.das` pp512, 5060 Ti).
The split arm keeps the general form.

### 2.2m Class-pipeline creation is the Vulkan tier's one shader A/B seat {#vk-class-pipeline-build}

`vkd_class_pipe` is the single place a class kernel's SPIR-V becomes a pipeline, so both shader
Expand Down
9 changes: 9 additions & 0 deletions modules/dasLLAMA/ARCHITECTURE_MEASUREMENT.md
Original file line number Diff line number Diff line change
Expand Up @@ -108,3 +108,12 @@ same math differ in float terms - while one that changes only WHEN work happens
a CLI flag is never an override (it is the run's own command line, visible where the run is
launched).


### Re-stamping inside the content-addressed archive

A sidecar archived as `records/<box>.tune.<sha12>.json` is content-addressed: its filename
carries the hash of its bytes. Re-stamping such a file's `provenance.engine_sha` to a reachable
commit (the remedy `performance/REVIEW.md` allows when the measured `modules/dasLLAMA/` tree is
byte-identical) therefore re-hashes and renames the file, and every `records/<box>.json` row
whose `tune_sha` named the old file is repointed in the same change - a row left on the old
name points at a file that no longer exists.
36 changes: 36 additions & 0 deletions modules/dasLLAMA/CLAUDE.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,36 @@
# dasLLAMA module instructions

dasLLAMA is the daslang LLM / ASR / vision engine, in-tree at `modules/dasLLAMA/`. **How it is
built and why is the `ARCHITECTURE*.md` set beside this file** (`ARCHITECTURE.md` routes to the
engine, image, GPU, Vulkan, Metal, measurement and media companions) - read the section you
are about to work in before writing code here. The rules binding a diff are `REVIEW*.md`;
`ENVIRONMENT.md` lists every knob; `followup_general.md` / `followup_vulkan.md` are the ledgers;
`PERF_LEDGER.md` is the measured record; `tests/CLAUDE.md` is the test discipline (run suites
ONLY through `tests/run.das`).

Follow the daslang **gen2** conventions - the root `CLAUDE.md` rules apply to every `.das` file
here.

## HOW_TO documents (REQUIRED for the task they name)

A HOW_TO is a procedure: imperative, ordered, each step citing the architecture section that
owns it, validated by execution, with a QUIRKS ledger of every place the pattern broke so a
follow-up arc can unquirk it. Read the one that matches your task before the first edit, and
fix it in the same session when a step turns out wrong.

| Document | Read BEFORE... |
|---|---|
| `HOW_TO_ADD_A_FORMAT.md` | Adding a weight format (a new `KqFmt`): GGUF type -> planes -> CPU kernels -> tune family -> Vulkan -> Metal -> tests |
| `BRINGUP.md` | Bringing a profiling box up from zero (the records rig; `METHODOLOGY.md` is the published method) |

Planned entries in the series: adding a model family, a vision tower, an audio tower, TTS.

## Skill files (REQUIRED)

| Skill file | Read BEFORE... |
|---|---|
| `skills/tune.md` | Touching any `[tune]` / `[tune_perm]` kernel family or the sidecar |
| `skills/internal/llvm_tune_internals.md` | Editing the tune framework itself |
| `skills/perf_lint.md` / `skills/style_lint.md` | Suppressing any lint finding here |
| `skills/internal/tests_in_repo.md` | Adding a test (the deep-engine rules: `options stack`, `T?`-free helpers) |
| `skills/writing_benchmarks.md` / `skills/internal/benchmarks_in_repo.md` | Anything under `benchmarks/` or `performance/` |
8 changes: 8 additions & 0 deletions modules/dasLLAMA/CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,14 @@ IF(NOT DAS_LLAMA_INCLUDED)
install(FILES ${PROJECT_SOURCE_DIR}/modules/dasLLAMA/performance/model_specs.das
DESTINATION ${DAS_INSTALL_MODULESDIR}/dasLLAMA/performance
)
# shipped defaults profiles ([tune_scope(defaults = ...)]): an untuned SDK box adopts its
# CPU class's winners instead of racing — the scope resolves this dir against the installed
# dasllama_math_gen.das, so the bundle must carry it
install(DIRECTORY ${PROJECT_SOURCE_DIR}/modules/dasLLAMA/performance/defaults
DESTINATION ${DAS_INSTALL_MODULESDIR}/dasLLAMA/performance
FILES_MATCHING
PATTERN "*.tune-defaults.json"
)
# third-party notices: the ported reference implementations (MIT) + the weights terms
install(FILES ${PROJECT_SOURCE_DIR}/modules/dasLLAMA/LICENSE.PARAKEET DESTINATION ${DAS_INSTALL_DOCDIR} RENAME PARAKEET.LICENSE)
install(FILES ${PROJECT_SOURCE_DIR}/modules/dasLLAMA/LICENSE.SILERO DESTINATION ${DAS_INSTALL_DOCDIR} RENAME SILERO.LICENSE)
Expand Down
Loading
Loading