Skip to content

[New Model]: Support Cua-S1 4B 0.2 — community help wanted #10

Description

@hsliuustc0106

The model to consider

Add support for Cua-S1 4B 0.2 in System1-Omni.

The release provides independently trained text/ and multimodal/ LoRA adapters for closed-option computer-use element/action decisions. The upstream model card loads them through cua_s1.four_b.FourBModel, selecting the adapter by modality. Version 0.2 is published separately from 0.1; this request specifically targets 0.2.

The closest planned or implemented model

System1-Omni is still in its design stage: no model engines or GPU backends are implemented, and LAYA is the first planned model. This is a request for a new integration, not a claim of existing support or a change to the first-model priority.

The broader model roadmap in #9 already lists Cua-S1 4B as a candidate. This issue tracks its implementation separately. Agent-side integration is tracked in ThinkFlowLab/system1-agents#14; that is distinct from engine support here.

What's needed to support it

Start by auditing and pinning the upstream reference implementation, then document the exact preprocessing, adapter loading, decision/scoring path, and output semantics for both modalities. Confirm the prefill-only execution requirements and compatibility with the frontend proposal in #1 before choosing the implementation approach.

Following the repository architecture, the model engine should own preprocessing, batching, execution, state, and postprocessing. Hardware operations belong in the CUDA or Metal backend. Potential work includes Qwen3.5 execution, LoRA loading or merging, image processing for the multimodal adapter, and matching upstream candidate scoring. The initial implementation may target one backend, with coverage and limitations stated explicitly.

Use case and motivation

Serve closed-option GUI element/action decisions through System1-Omni, including screenshot-conditioned decisions with the multimodal adapter. This would extend the planned model coverage to a Qwen3.5-based computer-use decision model.

Community help wanted

We welcome contributors who can help with any of the following:

  • Audit the upstream inference path and propose a small, concrete integration plan.
  • Implement model loading and text decision inference.
  • Add multimodal preprocessing and screenshot-conditioned inference.
  • Implement or integrate the required CUDA or Metal operations.
  • Validate outputs against the upstream reference and contribute reproducible setup and usage instructions.

If you would like to help, please comment with the area you can take on and your available hardware. Upstream maintainers and users familiar with FourBModel are especially welcome to clarify inference requirements and useful validation cases.

Acceptance criteria

  • Record the pinned upstream code and model revisions, inference contract, and initial backend scope.
  • Load the 0.2 adapters and validate text and multimodal decisions against the upstream reference on fixed inputs, with comparison tolerances declared in advance. Track any deferred modality explicitly.
  • Exercise the supported request path end to end and document unsupported input or output features.
  • Provide reproducible setup, launch, and example-request instructions.
  • Update the supported-model table with only the modalities and backends actually validated.

Performance claims, if contributed, should include the hardware, revisions, commands, raw results, and comparison methodology, with loading and warmup reported separately.

Before submitting

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

help wantedExtra attention is needednew modelRequests to support a new model

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions