cua_s1: complete native vision and screenshot inference with GPU parity - #64
Draft
Levius-Fubuki wants to merge 8 commits into
Draft
Levius-Fubuki wants to merge 8 commits into
Levius-Fubuki wants to merge 8 commits into
Conversation
# Conflicts: # recipe/cua_s1/native.md # src/models/cua_s1/native/Cargo.toml
…e-vision # Conflicts: # recipe/cua_s1/native.md # src/models/cua_s1/native/src/lib.rs
This was referenced Oct 2, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Complete native screenshot inference: PNG/JPEG → RGB preprocessing → 24-block CUDA vision encoder and merger → image feature insertion and 3D positions → language execution → candidate decisions.
omni-cua-s1-visionserves the screenshot HTTP contract and reuses vision features across questions in one request.Includes the implementations from #59, #63 and #56. Vision keeps BF16 base weights and all 50 FP32 LoRA pairs separate; language uses a merged BF16 export. Startup verifies source identity and export hashes. PNG decoding is bounded, 16-bit PNG conversion follows Pillow, and complete request validation precedes inference. CUDA ABI is 4; rebuild both library and worker. Multimodal execution is eager.
The diff retains runtime sources, build dependencies, required upstream notices/licenses and the one-time language exporter.
Build and launch
Use the pinned base/adapter downloads from
recipe/cua_s1/text.md. In its Python environment, install the additional image dependencies and obtain the trusted upstream lock beside the base directory:.venv/bin/python -m pip install torchvision==0.29.0 numpy==2.5.3 Pillow==11.3.0 curl --fail -L https://raw.githubusercontent.com/trycua/cua/0e75660ce4c2edda519e0c795fa3ad98abf4e76f/libs/cua-s1/ci/weights.lock.json -o weights/weights.lock.json PYTHONPATH=src .venv/bin/python recipe/cua_s1/export_multimodal_language.py \ --base weights/Qwen3.5-4B --adapter weights/cua-s1-4b-0.2/multimodal \ --out weights/cua-s1-multimodal-language src/backends/cuda/qwen3_5/build.sh target/release 89 cargo build --release --locked -p omni-cua-s1-native --bins CUA_S1_BASE=weights/Qwen3.5-4B \ CUA_S1_VISION_ADAPTER=weights/cua-s1-4b-0.2/multimodal \ CUA_S1_MODEL=weights/cua-s1-multimodal-language \ CUA_S1_CUDA_LIB=$PWD/target/release/libqwen3_5_cuda.so \ target/release/omni-cua-s1-visionThe worker defaults to
127.0.0.1:8000;/healthreports readiness. The local export manifest is trusted provenance, not an external signature. Keep checkpoint files immutable during inference.Validation
After cleanup: workspace formatting, strict all-target Clippy, workspace tests and the locked release build passed. 25 tests passed; 3 existing GPU tests and 2 external-checkpoint tests skipped. Runtime source was compared against the previous head: only new trailing test modules were removed; production code is unchanged. Existing baseline tests remain.
Feature-specific tests, examples, documentation and validation assets were removed from this PR diff to keep it focused on core implementation. They remain in the pre-cleanup commit and a local archive.
Before cleanup, RTX 4090 validation passed all five native GPU regressions. Standard cases matched 8/8 choices (maximum probability error 0.00331324, allowance 0.02052773); boundary cases matched 11/11 (maximum error 0.07950398, allowance 0.20292068). Live HTTP matched direct native inference for 11 requests / 19 questions, and nine invalid-input cases passed per set. These are historical results from the linked pre-cleanup commit, not a new GPU run. PNG pixels matched Pillow; JPEG decoding differed by up to three intensity levels in the validation image. The finite corpus establishes neither bitwise vision equivalence nor general accuracy or performance guarantees.
Self-review
Cleanup reviewed for unchanged runtime code, retained licensing and valid build targets. Draft status is retained for contributor/maintainer review.