Add automatic Graph selection with shared Cua-S1 worker budgets - #38
Levius-Fubuki wants to merge 116 commits into
Conversation
Remove archived experiment outputs and assistant planning notes. Keep runtime code, reproducible tools, and model verification metadata unchanged. Link historical reports to an immutable commit and ignore future local artifacts. Remove only tests tied to the deleted historical evidence bundles.
Remove archived experiment outputs and assistant planning notes. Keep runtime code, reproducible tools, and model verification metadata unchanged. Link historical reports to an immutable commit and ignore future local artifacts. Remove only tests tied to the deleted historical evidence bundles.
Remove archived experiment outputs and assistant planning notes. Keep runtime code, reproducible tools, and model verification metadata unchanged. Link historical reports to an immutable commit and ignore future local artifacts. Remove only tests tied to the deleted historical evidence bundles.
Remove archived experiment outputs and assistant planning notes. Keep runtime code, reproducible tools, and model verification metadata unchanged. Link historical reports to an immutable commit and ignore future local artifacts. Remove only tests tied to the deleted historical evidence bundles.
Remove supplemental tests, recipe tools, documentation, CI, and configuration changes from the PR diff. Runtime Python source and inline third-party notice are unchanged. Validation uses the pre-cleanup test/tool snapshot outside the checkout.
Remove supplemental tests, recipe tools, documentation, CI, and configuration changes from the PR diff. Runtime Python source and inline third-party notice are unchanged. Validation uses the pre-cleanup test/tool snapshot outside the checkout.
Remove supplemental tests, recipe tools, documentation, CI, and configuration changes from the PR diff. Runtime Python source and inline third-party notice are unchanged. Validation uses the pre-cleanup test/tool snapshot outside the checkout.
Remove supplemental tests, recipe tools, documentation, CI, and configuration changes from the PR diff. Runtime Python source and inline third-party notice are unchanged. Validation uses the pre-cleanup test/tool snapshot outside the checkout.
Remove supplemental tests, recipe tools, documentation, CI, and configuration changes from the PR diff. Runtime Python source and inline third-party notice are unchanged. Validation uses the pre-cleanup test/tool snapshot outside the checkout.
Remove supplemental tests, recipe tools, documentation, CI, and configuration changes from the PR diff. Runtime Python source and inline third-party notice are unchanged. Validation uses the pre-cleanup test/tool snapshot outside the checkout.
Remove supplemental tests, recipe tools, documentation, CI, and configuration changes from the PR diff. Runtime Python source and inline third-party notice are unchanged. Validation uses the pre-cleanup test/tool snapshot outside the checkout.
Remove supplemental tests, recipe tools, documentation, CI, and configuration changes from the PR diff. Runtime Python source and inline third-party notice are unchanged. Validation uses the pre-cleanup test/tool snapshot outside the checkout.
Remove supplemental tests, recipe tools, documentation, CI, and configuration changes from the PR diff. Runtime Python source and inline third-party notice are unchanged. Validation uses the pre-cleanup test/tool snapshot outside the checkout.
Remove supplemental tests, recipe tools, documentation, CI, and configuration changes from the PR diff. Runtime Python source and inline third-party notice are unchanged. Validation uses the pre-cleanup test/tool snapshot outside the checkout.
hsliuustc0106
left a comment
There was a problem hiding this comment.
Review of 164a3457f94b (incremental stack changes).
Reviewed exact/bucket selection, shared cache/admission/request accounting, bounded timing/history state and cross-mode eviction. No independent actionable correctness defect found in the inspected paths.
I would hold adoption until representative workload measurements demonstrate that automatic selection improves on explicit exact/bucket modes after including capture and parity-check costs. Also validate cross-mode eviction and surviving-graph replay on real GPUs. The heuristic is not itself evidence that auto chooses the fastest mode. This optimizes the Python multimodal worker, not #19's native Rust/CUDA path.
Validation: 264 archived CPU tests passed against this head's runtime source with Transformers 5.17.0 and tokenizers 0.23.2. Local Torch remains 2.13 rather than pinned 2.14. CUDA capture/stream/event behavior was mocked; no full-checkpoint GPU parity or new latency measurement was performed.
The validation suites were retrieved from the immutable archived revisions linked in the PR descriptions; they are not retained in the current PR diffs.
Purpose
Add opt-in automatic selection among eager, exact Graph and rule-bucket execution, sharing one request clock, capture ledger, cache and memory budget. The latest follow-up integrates the graph lifetime fixes from the preceding PRs. Eager remains the default; automatic selection does not guarantee the fastest mode.
Stacked on #37. Incremental runtime comparison. The PR diff remains limited to runtime Python source; validation helpers, tests and generated evidence are preserved separately. This optimizes the Python multimodal worker, not the native Rust/CUDA worker.
Test Plan
System1-Omni Version / Commit:
10b1522e483999e1830f715bb9970128cfaf323f.Fresh validation started 2026-09-30 on RTX 4090 24 GiB, driver 595.71.05, Torch 2.14.0+cu130, Transformers 5.17.0 and PEFT 0.21.0; BF16 base with unmerged adapter, without FLA/causal-conv1d.
PYTHONPATH=src:recipe/cua_s1 python -m pytest tests/cua_s1 -q.git diff <runtime> <validation> -- srcis empty. Source hashes and numerical/performance reports are independently verified.Test Result
rustCI passed on the above runtime head.Full evidence and tradeoffs · Raw primary report · Independent verification summary.
Timing is serial and synthetic on one GPU. Cold-cache totals include capture and parity-check costs; model loading, HTTP parsing and final cleanup are excluded. Two runs are not confidence intervals or proof of production workload behavior.
Self-review
Agent-assisted full-diff review and validation limits found no actionable correctness or architecture defect. Runtime scope, commands and claims were checked against the recorded sources/results. The contributor remains responsible for understanding the changes; this does not replace maintainer review.