From 2e33ef7dfdb92a830bbe4a9e3e60820d86486bcf Mon Sep 17 00:00:00 2001 From: Tianyao Wu Date: Thu, 1 Oct 2026 19:37:47 +0800 Subject: [PATCH] docs: add a supported models and hardware page A compatibility table for each model and worker on main: CPU, NVIDIA CUDA and Apple Metal, with requirements and validation status. Add the page and the two Cua-S1 recipes to the site nav, link the page from the README, and update the Cua-S1 status that still called the multimodal adapter deferred. Signed-off-by: Tianyao Wu --- README.md | 2 +- docs/supported-models.md | 17 +++++++++++++++++ mkdocs.yml | 3 +++ src/models/cua_s1/README.md | 4 ++-- 4 files changed, 23 insertions(+), 3 deletions(-) create mode 100644 docs/supported-models.md diff --git a/README.md b/README.md index c1e1c8b..1dca9f7 100644 --- a/README.md +++ b/README.md @@ -60,7 +60,7 @@ LAYA can run as an external Python worker for text requests; its in-repository m | LAYA | [External worker](recipe/laya/README.md); model engine planned | | Cua-S1 4B 0.2 (`text` adapter) | [Python worker](recipe/cua_s1/text.md); [native worker](recipe/cua_s1/native.md), CUDA, run on sm_89 | -CUDA and Metal coverage will be documented per model as implementations are added and validated. +[Supported models and hardware](docs/supported-models.md) lists the devices and where each worker has been run. ## Benchmarks diff --git a/docs/supported-models.md b/docs/supported-models.md new file mode 100644 index 0000000..26e4ca6 --- /dev/null +++ b/docs/supported-models.md @@ -0,0 +1,17 @@ +# Supported models and hardware + +This page covers what runs from `main`. Models that are being added are tracked in issues labeled [new model](https://github.com/ThinkFlowLab/system1-omni/issues?q=is%3Aissue%20state%3Aopen%20label%3A%22new%20model%22). + +| Model | Worker | CPU | NVIDIA CUDA | Apple Metal | Requirements | +| --- | --- | --- | --- | --- | --- | +| LAYA, English checkpoint | [External worker](../recipe/laya/README.md) running the upstream Laya runtime, for text requests | Validated ([#2](https://github.com/ThinkFlowLab/system1-omni/pull/2)) | Unverified ([#39](https://github.com/ThinkFlowLab/system1-omni/issues/39)) | Unverified ([#3](https://github.com/ThinkFlowLab/system1-omni/issues/3)) | Python 3.12, `laya[serve]==0.3.20` | +| LAYA | In-repository model engine | Planned ([#14](https://github.com/ThinkFlowLab/system1-omni/issues/14)) | Planned ([#14](https://github.com/ThinkFlowLab/system1-omni/issues/14)) | Planned ([#3](https://github.com/ThinkFlowLab/system1-omni/issues/3)) | | +| Cua-S1 4B 0.2, `text` adapter | [Reference worker](../recipe/cua_s1/text.md) on Transformers and PEFT | Unverified | Validated ([#13](https://github.com/ThinkFlowLab/system1-omni/pull/13)) | Unverified | Python 3.12, the versions in `requirements-text.txt` | +| Cua-S1 4B 0.2, `text` adapter | [Native Rust worker](../recipe/cua_s1/native.md) on the [Qwen3.5 CUDA kernels](../src/backends/cuda/qwen3_5/README.md) | Not supported | Validated on compute capability 8.9 ([#19](https://github.com/ThinkFlowLab/system1-omni/pull/19), [#52](https://github.com/ThinkFlowLab/system1-omni/pull/52)) | Not supported | Compute capability 8.0 or newer, the CUDA toolkit to build, weights merged with `export_text_merged.py` | +| Cua-S1 4B 0.2, `multimodal` adapter | Reference worker on Transformers and PEFT, [`src/frontend/cua_s1.py`](../src/frontend/cua_s1.py); no recipe yet | Not supported | Validated ([#17](https://github.com/ThinkFlowLab/system1-omni/pull/17), [#18](https://github.com/ThinkFlowLab/system1-omni/pull/18)) | Not supported | The state is one PNG or JPEG image; upstream's `weights.lock.json` next to the base weights | + +- **Validated:** covered by the recipe on `main` or by the checks in the linked merged pull request. +- **Unverified:** the worker accepts this device, but no recipe or merged pull request covers it. +- **Planned:** not implemented yet; the linked issue tracks it. + +The Cua-S1 workers answer `choice` questions only. diff --git a/mkdocs.yml b/mkdocs.yml index 2d63d2e..3927a55 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -46,8 +46,11 @@ markdown_extensions: nav: - Overview: README.md + - Supported models and hardware: docs/supported-models.md - Frontend: src/frontend/README.md - Recipes: - recipe/README.md - Laya text worker: recipe/laya/README.md + - Cua-S1 text worker: recipe/cua_s1/text.md + - Cua-S1 native text worker: recipe/cua_s1/native.md - Contributing: CONTRIBUTING.md diff --git a/src/models/cua_s1/README.md b/src/models/cua_s1/README.md index 999ae0a..40142c8 100644 --- a/src/models/cua_s1/README.md +++ b/src/models/cua_s1/README.md @@ -2,7 +2,7 @@ This directory owns Cua-S1 4B 0.2 ([#10](https://github.com/ThinkFlowLab/system1-omni/issues/10)): request mapping, prompt construction, adapter selection, execution, and the answer-letter readout. This page records the pinned upstream revisions, the inference contract an implementation must match, and how its outputs will be compared with the upstream reference. -Status: a reference worker for the `text` adapter loads the model through Hugging Face Transformers and PEFT: [`text/`](text/), served by [`src/frontend/cua_s1_text.py`](../../frontend/cua_s1_text.py), with setup in [`recipe/cua_s1/text.md`](../../../recipe/cua_s1/text.md). It is the correctness reference for the native worker in [`native/`](native/): Rust, with the Qwen3.5 forward pass on the CUDA kernels in [`src/backends/cuda/qwen3_5/`](../../backends/cuda/qwen3_5/), set up as in [`recipe/cua_s1/native.md`](../../../recipe/cua_s1/native.md). The `multimodal` adapter is deferred; see [Not covered yet](#not-covered-yet). +Status: a reference worker for the `text` adapter loads the model through Hugging Face Transformers and PEFT: [`text/`](text/), served by [`src/frontend/cua_s1_text.py`](../../frontend/cua_s1_text.py), with setup in [`recipe/cua_s1/text.md`](../../../recipe/cua_s1/text.md). It is the correctness reference for the native worker in [`native/`](native/): Rust, with the Qwen3.5 forward pass on the CUDA kernels in [`src/backends/cuda/qwen3_5/`](../../backends/cuda/qwen3_5/), set up as in [`recipe/cua_s1/native.md`](../../../recipe/cua_s1/native.md). A reference worker for the `multimodal` adapter is in [`multimodal/`](multimodal/), served by [`src/frontend/cua_s1.py`](../../frontend/cua_s1.py); native execution of that adapter is not covered yet. ## Pinned revisions @@ -105,7 +105,7 @@ The bfloat16 worker's own difference from the fp32 worker is reported next to ea ## Not covered yet -- The `multimodal` adapter: image preprocessing, the vision tower and the vision LoRA. This is tracked in [#10](https://github.com/ThinkFlowLab/system1-omni/issues/10). +- Native execution of the `multimodal` adapter: image preprocessing, the vision tower and the vision LoRA. This is tracked in [#10](https://github.com/ThinkFlowLab/system1-omni/issues/10). - `score` and `noul` questions. - More than 26 options per question. - The Metal backend.