Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -60,7 +60,7 @@ LAYA can run as an external Python worker for text requests; its in-repository m
| LAYA | [External worker](recipe/laya/README.md); model engine planned |
| Cua-S1 4B 0.2 (`text` adapter) | [Python worker](recipe/cua_s1/text.md); [native worker](recipe/cua_s1/native.md), CUDA, run on sm_89 |

CUDA and Metal coverage will be documented per model as implementations are added and validated.
[Supported models and hardware](docs/supported-models.md) lists the devices and where each worker has been run.

## Benchmarks

Expand Down
17 changes: 17 additions & 0 deletions docs/supported-models.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
# Supported models and hardware

This page covers what runs from `main`. Models that are being added are tracked in issues labeled [new model](https://github.com/ThinkFlowLab/system1-omni/issues?q=is%3Aissue%20state%3Aopen%20label%3A%22new%20model%22).

| Model | Worker | CPU | NVIDIA CUDA | Apple Metal | Requirements |
| --- | --- | --- | --- | --- | --- |
| LAYA, English checkpoint | [External worker](../recipe/laya/README.md) running the upstream Laya runtime, for text requests | Validated ([#2](https://github.com/ThinkFlowLab/system1-omni/pull/2)) | Unverified ([#39](https://github.com/ThinkFlowLab/system1-omni/issues/39)) | Unverified ([#3](https://github.com/ThinkFlowLab/system1-omni/issues/3)) | Python 3.12, `laya[serve]==0.3.20` |
| LAYA | In-repository model engine | Planned ([#14](https://github.com/ThinkFlowLab/system1-omni/issues/14)) | Planned ([#14](https://github.com/ThinkFlowLab/system1-omni/issues/14)) | Planned ([#3](https://github.com/ThinkFlowLab/system1-omni/issues/3)) | |
| Cua-S1 4B 0.2, `text` adapter | [Reference worker](../recipe/cua_s1/text.md) on Transformers and PEFT | Unverified | Validated ([#13](https://github.com/ThinkFlowLab/system1-omni/pull/13)) | Unverified | Python 3.12, the versions in `requirements-text.txt` |
| Cua-S1 4B 0.2, `text` adapter | [Native Rust worker](../recipe/cua_s1/native.md) on the [Qwen3.5 CUDA kernels](../src/backends/cuda/qwen3_5/README.md) | Not supported | Validated on compute capability 8.9 ([#19](https://github.com/ThinkFlowLab/system1-omni/pull/19), [#52](https://github.com/ThinkFlowLab/system1-omni/pull/52)) | Not supported | Compute capability 8.0 or newer, the CUDA toolkit to build, weights merged with `export_text_merged.py` |
| Cua-S1 4B 0.2, `multimodal` adapter | Reference worker on Transformers and PEFT, [`src/frontend/cua_s1.py`](../src/frontend/cua_s1.py); no recipe yet | Not supported | Validated ([#17](https://github.com/ThinkFlowLab/system1-omni/pull/17), [#18](https://github.com/ThinkFlowLab/system1-omni/pull/18)) | Not supported | The state is one PNG or JPEG image; upstream's `weights.lock.json` next to the base weights |

- **Validated:** covered by the recipe on `main` or by the checks in the linked merged pull request.
- **Unverified:** the worker accepts this device, but no recipe or merged pull request covers it.
- **Planned:** not implemented yet; the linked issue tracks it.

The Cua-S1 workers answer `choice` questions only.
3 changes: 3 additions & 0 deletions mkdocs.yml
Original file line number Diff line number Diff line change
Expand Up @@ -46,8 +46,11 @@ markdown_extensions:

nav:
- Overview: README.md
- Supported models and hardware: docs/supported-models.md
- Frontend: src/frontend/README.md
- Recipes:
- recipe/README.md
- Laya text worker: recipe/laya/README.md
- Cua-S1 text worker: recipe/cua_s1/text.md
- Cua-S1 native text worker: recipe/cua_s1/native.md
- Contributing: CONTRIBUTING.md
4 changes: 2 additions & 2 deletions src/models/cua_s1/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

This directory owns Cua-S1 4B 0.2 ([#10](https://github.com/ThinkFlowLab/system1-omni/issues/10)): request mapping, prompt construction, adapter selection, execution, and the answer-letter readout. This page records the pinned upstream revisions, the inference contract an implementation must match, and how its outputs will be compared with the upstream reference.

Status: a reference worker for the `text` adapter loads the model through Hugging Face Transformers and PEFT: [`text/`](text/), served by [`src/frontend/cua_s1_text.py`](../../frontend/cua_s1_text.py), with setup in [`recipe/cua_s1/text.md`](../../../recipe/cua_s1/text.md). It is the correctness reference for the native worker in [`native/`](native/): Rust, with the Qwen3.5 forward pass on the CUDA kernels in [`src/backends/cuda/qwen3_5/`](../../backends/cuda/qwen3_5/), set up as in [`recipe/cua_s1/native.md`](../../../recipe/cua_s1/native.md). The `multimodal` adapter is deferred; see [Not covered yet](#not-covered-yet).
Status: a reference worker for the `text` adapter loads the model through Hugging Face Transformers and PEFT: [`text/`](text/), served by [`src/frontend/cua_s1_text.py`](../../frontend/cua_s1_text.py), with setup in [`recipe/cua_s1/text.md`](../../../recipe/cua_s1/text.md). It is the correctness reference for the native worker in [`native/`](native/): Rust, with the Qwen3.5 forward pass on the CUDA kernels in [`src/backends/cuda/qwen3_5/`](../../backends/cuda/qwen3_5/), set up as in [`recipe/cua_s1/native.md`](../../../recipe/cua_s1/native.md). A reference worker for the `multimodal` adapter is in [`multimodal/`](multimodal/), served by [`src/frontend/cua_s1.py`](../../frontend/cua_s1.py); native execution of that adapter is not covered yet.

## Pinned revisions

Expand Down Expand Up @@ -105,7 +105,7 @@ The bfloat16 worker's own difference from the fp32 worker is reported next to ea

## Not covered yet

- The `multimodal` adapter: image preprocessing, the vision tower and the vision LoRA. This is tracked in [#10](https://github.com/ThinkFlowLab/system1-omni/issues/10).
- Native execution of the `multimodal` adapter: image preprocessing, the vision tower and the vision LoRA. This is tracked in [#10](https://github.com/ThinkFlowLab/system1-omni/issues/10).
- `score` and `noul` questions.
- More than 26 options per question.
- The Metal backend.
Expand Down
Loading