From 8bc2700d152391431c3b707b7dcd268ce1efcdc9 Mon Sep 17 00:00:00 2001 From: Tianyao Wu Date: Sun, 27 Sep 2026 19:09:00 +0800 Subject: [PATCH 1/3] docs: add Cua-S1 4B 0.2 inference contract Record the pinned upstream revisions, the inference contract, the /v1/systemone request mapping and the comparison tolerances for Cua-S1 4B 0.2 in src/models/cua_s1/README.md, and list the directory in the README layout table. Part of #10. Signed-off-by: Tianyao Wu --- README.md | 1 + src/models/cua_s1/README.md | 108 ++++++++++++++++++++++++++++++++++++ 2 files changed, 109 insertions(+) create mode 100644 src/models/cua_s1/README.md diff --git a/README.md b/README.md index 6cc42ab..f040cfa 100644 --- a/README.md +++ b/README.md @@ -27,6 +27,7 @@ Implementation code lives under `src/`; recipes and documentation stay at the re | --- | --- | | [`src/frontend/`](src/frontend/) | Rust serving code and the small engine interface. | | [`src/models/laya/`](src/models/laya/) | LAYA preprocessing, batching, state, execution, and output processing. | +| [`src/models/cua_s1/`](src/models/cua_s1/) | Cua-S1 4B 0.2 inference contract, request mapping, and execution. | | [`src/backends/cuda/`](src/backends/cuda/) | NVIDIA GPU operations and kernel integration. | | [`src/backends/metal/`](src/backends/metal/) | Apple GPU operations and kernel integration. | | [`recipe/`](recipe/) | Model setup instructions, launch commands, configuration examples, and example requests. | diff --git a/src/models/cua_s1/README.md b/src/models/cua_s1/README.md new file mode 100644 index 0000000..6bb1ecf --- /dev/null +++ b/src/models/cua_s1/README.md @@ -0,0 +1,108 @@ +# Cua-S1 4B 0.2 model engine + +This directory owns Cua-S1 4B 0.2 ([#10](https://github.com/ThinkFlowLab/system1-omni/issues/10)): request mapping, prompt construction, adapter selection, execution, and the answer-letter readout. This page records the pinned upstream revisions, the inference contract an implementation must match, and how its outputs will be compared with the upstream reference. + +Status: planned; nothing is implemented or validated yet. The first target is the `text` adapter on CUDA, starting with a worker that loads the model directly through Hugging Face Transformers and PEFT. The `multimodal` adapter is deferred; see [Not covered yet](#not-covered-yet). + +## Pinned revisions + +| Artifact | Revision | +| --- | --- | +| Reference code: [`trycua/cua`](https://github.com/trycua/cua/tree/0e75660ce4c2edda519e0c795fa3ad98abf4e76f/libs/cua-s1) `libs/cua-s1` | `0e75660ce4c2edda519e0c795fa3ad98abf4e76f` | +| Base model: [`Qwen/Qwen3.5-4B`](https://huggingface.co/Qwen/Qwen3.5-4B) | `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a` | +| Adapters: [`cua-ai/cua-s1-4b-0.2`](https://huggingface.co/cua-ai/cua-s1-4b-0.2) | `16818868b0cc7813808aae4e87b417657046ab79` | + +The model revisions are the ones upstream pins in `libs/cua-s1/ci/weights.lock.json`, which lists the size and SHA-256 of every file. Downloads are checked against that file with upstream's `ci/fetch_pinned_weights.py --verify-only`. + +The reference implementation is `cua_s1.four_b.FourBModel`. The reference environment is upstream's `four-b` lock (`libs/cua-s1/python/uv.lock`) on Python 3.12, as upstream measured: torch 2.14.0, Transformers 5.17.0, PEFT 0.21.0 and torchvision 0.29.0. It has neither `flash-linear-attention` nor `causal-conv1d`, so Transformers runs the reference PyTorch implementations of those operations. + +## Model + +- **Base.** Qwen3.5-4B has 32 decoder layers: 24 Gated DeltaNet (linear attention) layers and 8 full-attention layers (layers 3, 7, ..., 31, counting from 0). Hidden size is 2560, the MLP intermediate size is 9216, the embedding has 248,320 rows, and the input and output embeddings are tied. The checkpoint also contains one MTP layer, which is not used. +- **Adapters.** `text/` and `multimodal/` are two independently trained PEFT LoRA adapters, both rank 16 and alpha 32 (scale 2), stored in fp32. + - `text/` targets `q_proj`, `k_proj`, `v_proj` and `o_proj` in the 8 full-attention layers, and `gate_proj`, `up_proj` and `down_proj` in all 32 MLPs. The Gated DeltaNet attention blocks keep their base weights; only the MLPs in those layers are adapted. + - `multimodal/` targets the same language modules, plus `linear_fc1` and `linear_fc2` in the 24 vision blocks and in the vision merger. Its keys follow the image-text model layout (`model.language_model.*`, `model.visual.*`), so the two adapters are not interchangeable. +- **Model classes.** Upstream loads `text` with `AutoModelForCausalLM` (`Qwen3_5ForCausalLM`) and `multimodal` with `AutoModelForImageTextToText` (`Qwen3_5ForConditionalGeneration`). `FourBModel` defaults to bfloat16. PEFT keeps the adapter weights in fp32 and, while the adapter is not merged, computes each LoRA branch in fp32 and casts the sum back to bfloat16. + +## Inference contract + +A decision is one forward pass over one prompt, with no decoding. + +1. Options get the letters `A` to `Z` in request order, so a question has at most 26 options. +2. The prompt has a fixed system message and a user message in this layout (upstream `build_prompt`, text modality): + + ```text + Goal: + + App: + Task family: + + Accessibility tree: + + + Options: + A. "