Skip to content

Add instance-local rule buckets to the Cua-S1 multimodal worker - #37

Open
Levius-Fubuki wants to merge 100 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-bucket-worker
Open

Levius-Fubuki wants to merge 100 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-bucket-worker

Conversation

@Levius-Fubuki

@Levius-Fubuki Levius-Fubuki commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

Add opt-in rule-bucket execution inside the model-owned Python worker. Nearby lengths can reuse a bucket after strict full-vocabulary-logit checks; per-length gates, resource ownership and shutdown remain instance-local. The latest follow-up integrates the graph lifetime fixes from #22/#33.

Stacked on #33. Incremental runtime comparison. The PR diff remains limited to runtime Python source; validation helpers, tests and generated evidence are preserved separately. This optimizes the Python multimodal worker, not the native Rust/CUDA worker.

Test Plan

System1-Omni Version / Commit: d598de3ce03f57dd20ff9b730279a4b4979623d4.

Fresh validation started 2026-09-30 on RTX 4090 24 GiB, driver 595.71.05, Torch 2.14.0+cu130, Transformers 5.17.0 and PEFT 0.21.0; BF16 base with unmerged adapter, without FLA/causal-conv1d.

Test Result

  • CPU suite: 227 passed in 21.31s.
  • Fresh full-checkpoint validation passes 136 full-vocabulary-logit comparisons and 20 complete responses, plus 255/256/257/319/320/321-token boundary, eviction and budget checks. Real HTTP changed-input parity, independent-model isolation and shutdown pass.
  • The additional full-model eviction checker confirms three surviving-graph replays after actual eviction with changed contents and retained references.
  • Two cold-cache runs across four hot lengths and twelve rotating lengths give 720 timed predictions, all with complete response equality. Bucket total time is 27.3–27.9% below eager on these samples. It is 1.5–1.9% below exact for hot lengths and about 25.0–25.1% below default exact for the rotation; rotation p95 is slightly above eager.
  • Default exact makes no captures in the twelve-length rotation because the revisit interval exceeds its eight-request admission window. Add automatic Graph selection with shared Cua-S1 worker budgets #38 includes a controlled exact-window-32 comparison. These synthetic results support explicit opt-in, not general adoption or concurrency claims.
  • GitHub rust CI passed on the above runtime head.

Full evidence and tradeoffs · Raw primary report · Independent verification summary.

Timing is serial and synthetic on one GPU. Cold-cache totals include capture and parity-check costs; model loading, HTTP parsing and final cleanup are excluded. Two runs are not confidence intervals or proof of production workload behavior.

Self-review

Agent-assisted full-diff review and validation limits found no actionable correctness or architecture defect. Runtime scope, commands and claims were checked against the recorded sources/results. The contributor remains responsible for understanding the changes; this does not replace maintainer review.

Remove archived experiment outputs and assistant planning notes. Keep runtime code, reproducible tools, and model verification metadata unchanged. Link historical reports to an immutable commit and ignore future local artifacts. Remove only tests tied to the deleted historical evidence bundles.
Remove archived experiment outputs and assistant planning notes. Keep runtime code, reproducible tools, and model verification metadata unchanged. Link historical reports to an immutable commit and ignore future local artifacts. Remove only tests tied to the deleted historical evidence bundles.
Remove archived experiment outputs and assistant planning notes. Keep runtime code, reproducible tools, and model verification metadata unchanged. Link historical reports to an immutable commit and ignore future local artifacts. Remove only tests tied to the deleted historical evidence bundles.
Remove archived experiment outputs and assistant planning notes. Keep runtime code, reproducible tools, and model verification metadata unchanged. Link historical reports to an immutable commit and ignore future local artifacts. Remove only tests tied to the deleted historical evidence bundles.
Remove archived experiment outputs and assistant planning notes. Keep runtime code, reproducible tools, and model verification metadata unchanged. Link historical reports to an immutable commit and ignore future local artifacts. Remove only tests tied to the deleted historical evidence bundles.
Remove archived experiment outputs and assistant planning notes. Keep runtime code, reproducible tools, and model verification metadata unchanged. Link historical reports to an immutable commit and ignore future local artifacts. Remove only tests tied to the deleted historical evidence bundles.
Remove archived experiment outputs and assistant planning notes. Keep runtime code, reproducible tools, and model verification metadata unchanged. Link historical reports to an immutable commit and ignore future local artifacts. Remove only tests tied to the deleted historical evidence bundles.
Remove archived experiment outputs and assistant planning notes. Keep runtime code, reproducible tools, and model verification metadata unchanged. Link historical reports to an immutable commit and ignore future local artifacts. Remove only tests tied to the deleted historical evidence bundles.
Remove supplemental tests, recipe tools, documentation, CI, and configuration changes from the PR diff. Runtime Python source and inline third-party notice are unchanged. Validation uses the pre-cleanup test/tool snapshot outside the checkout.
Remove supplemental tests, recipe tools, documentation, CI, and configuration changes from the PR diff. Runtime Python source and inline third-party notice are unchanged. Validation uses the pre-cleanup test/tool snapshot outside the checkout.
Remove supplemental tests, recipe tools, documentation, CI, and configuration changes from the PR diff. Runtime Python source and inline third-party notice are unchanged. Validation uses the pre-cleanup test/tool snapshot outside the checkout.
Remove supplemental tests, recipe tools, documentation, CI, and configuration changes from the PR diff. Runtime Python source and inline third-party notice are unchanged. Validation uses the pre-cleanup test/tool snapshot outside the checkout.
Remove supplemental tests, recipe tools, documentation, CI, and configuration changes from the PR diff. Runtime Python source and inline third-party notice are unchanged. Validation uses the pre-cleanup test/tool snapshot outside the checkout.
Remove supplemental tests, recipe tools, documentation, CI, and configuration changes from the PR diff. Runtime Python source and inline third-party notice are unchanged. Validation uses the pre-cleanup test/tool snapshot outside the checkout.
Remove supplemental tests, recipe tools, documentation, CI, and configuration changes from the PR diff. Runtime Python source and inline third-party notice are unchanged. Validation uses the pre-cleanup test/tool snapshot outside the checkout.
Remove supplemental tests, recipe tools, documentation, CI, and configuration changes from the PR diff. Runtime Python source and inline third-party notice are unchanged. Validation uses the pre-cleanup test/tool snapshot outside the checkout.
Remove supplemental tests, recipe tools, documentation, CI, and configuration changes from the PR diff. Runtime Python source and inline third-party notice are unchanged. Validation uses the pre-cleanup test/tool snapshot outside the checkout.

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review of 3254e902eb86 (incremental stack changes).

Reviewed bucket padding/trimming, the pinned Transformers forward integration, per-real-length parity checks, cache growth and worker shutdown. No independent actionable correctness defect found in the inspected paths.

I would hold adoption until real GPU parity and controlled performance measurements establish the benefit of rule bucketing over the exact-length/eager paths, including varying lengths within a bucket and eviction/replay. This adds execution logic to the Python multimodal worker and does not accelerate #19's native Rust/CUDA worker.

Validation: 227 archived CPU tests passed against this head's runtime source using Transformers 5.17.0 and tokenizers 0.23.2. Six initial failures with Transformers 5.14.1 were version-guard failures and disappeared with the pinned version. Local Torch remains 2.13 rather than pinned 2.14. Tiny real model tensors were exercised on CPU; CUDA capture/stream behavior was mocked. No full-checkpoint GPU parity or new latency measurement was performed.

The validation suites were retrieved from the immutable archived revisions linked in the PR descriptions; they are not retained in the current PR diffs.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants