The model to consider
Add support for Cua-S1 4B 0.2 in System1-Omni.
The release provides independently trained text/ and multimodal/ LoRA adapters for closed-option computer-use element/action decisions. The upstream model card loads them through cua_s1.four_b.FourBModel, selecting the adapter by modality. Version 0.2 is published separately from 0.1; this request specifically targets 0.2.
The closest planned or implemented model
System1-Omni is still in its design stage: no model engines or GPU backends are implemented, and LAYA is the first planned model. This is a request for a new integration, not a claim of existing support or a change to the first-model priority.
The broader model roadmap in #9 already lists Cua-S1 4B as a candidate. This issue tracks its implementation separately. Agent-side integration is tracked in ThinkFlowLab/system1-agents#14; that is distinct from engine support here.
What's needed to support it
Start by auditing and pinning the upstream reference implementation, then document the exact preprocessing, adapter loading, decision/scoring path, and output semantics for both modalities. Confirm the prefill-only execution requirements and compatibility with the frontend proposal in #1 before choosing the implementation approach.
Following the repository architecture, the model engine should own preprocessing, batching, execution, state, and postprocessing. Hardware operations belong in the CUDA or Metal backend. Potential work includes Qwen3.5 execution, LoRA loading or merging, image processing for the multimodal adapter, and matching upstream candidate scoring. The initial implementation may target one backend, with coverage and limitations stated explicitly.
Use case and motivation
Serve closed-option GUI element/action decisions through System1-Omni, including screenshot-conditioned decisions with the multimodal adapter. This would extend the planned model coverage to a Qwen3.5-based computer-use decision model.
Community help wanted
We welcome contributors who can help with any of the following:
- Audit the upstream inference path and propose a small, concrete integration plan.
- Implement model loading and text decision inference.
- Add multimodal preprocessing and screenshot-conditioned inference.
- Implement or integrate the required CUDA or Metal operations.
- Validate outputs against the upstream reference and contribute reproducible setup and usage instructions.
If you would like to help, please comment with the area you can take on and your available hardware. Upstream maintainers and users familiar with FourBModel are especially welcome to clarify inference requirements and useful validation cases.
Acceptance criteria
Performance claims, if contributed, should include the hardware, revisions, commands, raw results, and comparison methodology, with loading and warmup reported separately.
Before submitting
The model to consider
Add support for Cua-S1 4B 0.2 in System1-Omni.
Qwen/Qwen3.5-4BThe release provides independently trained
text/andmultimodal/LoRA adapters for closed-option computer-use element/action decisions. The upstream model card loads them throughcua_s1.four_b.FourBModel, selecting the adapter by modality. Version 0.2 is published separately from 0.1; this request specifically targets 0.2.The closest planned or implemented model
System1-Omni is still in its design stage: no model engines or GPU backends are implemented, and LAYA is the first planned model. This is a request for a new integration, not a claim of existing support or a change to the first-model priority.
The broader model roadmap in #9 already lists Cua-S1 4B as a candidate. This issue tracks its implementation separately. Agent-side integration is tracked in ThinkFlowLab/system1-agents#14; that is distinct from engine support here.
What's needed to support it
Start by auditing and pinning the upstream reference implementation, then document the exact preprocessing, adapter loading, decision/scoring path, and output semantics for both modalities. Confirm the prefill-only execution requirements and compatibility with the frontend proposal in #1 before choosing the implementation approach.
Following the repository architecture, the model engine should own preprocessing, batching, execution, state, and postprocessing. Hardware operations belong in the CUDA or Metal backend. Potential work includes Qwen3.5 execution, LoRA loading or merging, image processing for the multimodal adapter, and matching upstream candidate scoring. The initial implementation may target one backend, with coverage and limitations stated explicitly.
Use case and motivation
Serve closed-option GUI element/action decisions through System1-Omni, including screenshot-conditioned decisions with the multimodal adapter. This would extend the planned model coverage to a Qwen3.5-based computer-use decision model.
Community help wanted
We welcome contributors who can help with any of the following:
If you would like to help, please comment with the area you can take on and your available hardware. Upstream maintainers and users familiar with
FourBModelare especially welcome to clarify inference requirements and useful validation cases.Acceptance criteria
Performance claims, if contributed, should include the hardware, revisions, commands, raw results, and comparison methodology, with loading and warmup reported separately.
Before submitting