Skip to content

cua_s1: complete native vision and screenshot inference with GPU parity - #64

Draft
Levius-Fubuki wants to merge 8 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-native-vision
Draft

Levius-Fubuki wants to merge 8 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-native-vision

Conversation

@Levius-Fubuki

@Levius-Fubuki Levius-Fubuki commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

Complete native screenshot inference: PNG/JPEG → RGB preprocessing → 24-block CUDA vision encoder and merger → image feature insertion and 3D positions → language execution → candidate decisions. omni-cua-s1-vision serves the screenshot HTTP contract and reuses vision features across questions in one request.

Includes the implementations from #59, #63 and #56. Vision keeps BF16 base weights and all 50 FP32 LoRA pairs separate; language uses a merged BF16 export. Startup verifies source identity and export hashes. PNG decoding is bounded, 16-bit PNG conversion follows Pillow, and complete request validation precedes inference. CUDA ABI is 4; rebuild both library and worker. Multimodal execution is eager.

The diff retains runtime sources, build dependencies, required upstream notices/licenses and the one-time language exporter.

Build and launch

Use the pinned base/adapter downloads from recipe/cua_s1/text.md. In its Python environment, install the additional image dependencies and obtain the trusted upstream lock beside the base directory:

.venv/bin/python -m pip install torchvision==0.29.0 numpy==2.5.3 Pillow==11.3.0
curl --fail -L https://raw.githubusercontent.com/trycua/cua/0e75660ce4c2edda519e0c795fa3ad98abf4e76f/libs/cua-s1/ci/weights.lock.json -o weights/weights.lock.json
PYTHONPATH=src .venv/bin/python recipe/cua_s1/export_multimodal_language.py \
  --base weights/Qwen3.5-4B --adapter weights/cua-s1-4b-0.2/multimodal \
  --out weights/cua-s1-multimodal-language
src/backends/cuda/qwen3_5/build.sh target/release 89
cargo build --release --locked -p omni-cua-s1-native --bins
CUA_S1_BASE=weights/Qwen3.5-4B \
CUA_S1_VISION_ADAPTER=weights/cua-s1-4b-0.2/multimodal \
CUA_S1_MODEL=weights/cua-s1-multimodal-language \
CUA_S1_CUDA_LIB=$PWD/target/release/libqwen3_5_cuda.so \
  target/release/omni-cua-s1-vision

The worker defaults to 127.0.0.1:8000; /health reports readiness. The local export manifest is trusted provenance, not an external signature. Keep checkpoint files immutable during inference.

Validation

After cleanup: workspace formatting, strict all-target Clippy, workspace tests and the locked release build passed. 25 tests passed; 3 existing GPU tests and 2 external-checkpoint tests skipped. Runtime source was compared against the previous head: only new trailing test modules were removed; production code is unchanged. Existing baseline tests remain.

Feature-specific tests, examples, documentation and validation assets were removed from this PR diff to keep it focused on core implementation. They remain in the pre-cleanup commit and a local archive.

Before cleanup, RTX 4090 validation passed all five native GPU regressions. Standard cases matched 8/8 choices (maximum probability error 0.00331324, allowance 0.02052773); boundary cases matched 11/11 (maximum error 0.07950398, allowance 0.20292068). Live HTTP matched direct native inference for 11 requests / 19 questions, and nine invalid-input cases passed per set. These are historical results from the linked pre-cleanup commit, not a new GPU run. PNG pixels matched Pillow; JPEG decoding differed by up to three intensity levels in the validation image. The finite corpus establishes neither bitwise vision equivalence nor general accuracy or performance guarantees.

Self-review

Cleanup reviewed for unchanged runtime code, retained licensing and valid build targets. Draft status is retained for contributor/maintainer review.

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant