Reuse Cua-S1 image preprocessing and vision features per request - #17
Conversation
hsliuustc0106
left a comment
There was a problem hiding this comment.
Reviewed original snapshot 95713a5 and the follow-up diff through 642d14c754785de5f2c2768fe537cef1d4d690ab. No new actionable findings.
The follow-up fixes the inherited extreme-aspect-ratio validation and lock-cleanup test race from #12. Reviewed the protocol/test changes and documentation updates. Protocol and HTTP tests on this head: 61 passed in 14.73s. No accelerator execution was performed.
Validation and scope of the original snapshot review:
No new findings in image-reuse delta. 15 passed, 1 skipped; tensor reuse test later passes on #18 in torch environment.
|
are you serious? 70K+ LoC |
Remove archived experiment outputs and assistant planning notes. Keep runtime code, reproducible tools, and model verification metadata unchanged. Link historical reports to an immutable commit and ignore future local artifacts. Remove only tests tied to the deleted historical evidence bundles.
Remove archived experiment outputs and assistant planning notes. Keep runtime code, reproducible tools, and model verification metadata unchanged. Link historical reports to an immutable commit and ignore future local artifacts. Remove only tests tied to the deleted historical evidence bundles.
Remove archived experiment outputs and assistant planning notes. Keep runtime code, reproducible tools, and model verification metadata unchanged. Link historical reports to an immutable commit and ignore future local artifacts. Remove only tests tied to the deleted historical evidence bundles.
Remove supplemental tests, recipe tools, documentation, CI, and configuration changes from the PR diff. Runtime Python source and inline third-party notice are unchanged. Validation uses the pre-cleanup test/tool snapshot outside the checkout.
Remove supplemental tests, recipe tools, documentation, CI, and configuration changes from the PR diff. Runtime Python source and inline third-party notice are unchanged. Validation uses the pre-cleanup test/tool snapshot outside the checkout.
Remove supplemental tests, recipe tools, documentation, CI, and configuration changes from the PR diff. Runtime Python source and inline third-party notice are unchanged. Validation uses the pre-cleanup test/tool snapshot outside the checkout.
hsliuustc0106
left a comment
There was a problem hiding this comment.
Review of 1a339e1b6363 (incremental stack changes).
Reviewed the request-local preprocessing/image-feature reuse, question-specific embeddings and positions, and token-limit handling. No actionable correctness defect found in the inspected paths. This is a reasonable optimization for the existing Python multimodal reference worker; I recommend accepting it.
Validation: 115 archived CPU tests passed against this head's runtime source. No full-checkpoint GPU parity or new latency measurement was performed. Local environment used Torch 2.13 and Transformers 5.14.1, rather than the recipe's Torch 2.14 / Transformers 5.17 pins.
The validation suites were retrieved from the immutable archived revisions linked in the PR descriptions; they are not retained in the current PR diffs.
|
@hsliuustc0106 Following up on your recommendation to accept Could you approve/merge #17 first, followed by #18? Both remain conflict-free against current Immutable GPU evidence and limitations · raw #17 report. The evidence/source verifier was rerun successfully on October 1; this follow-up adds no new performance claim. |
Purpose
Reuse preprocessing and adapted vision features across questions within one request while preserving question-specific language execution and the single-question reference path.
Stacked on #12. Incremental runtime comparison. The PR diff remains limited to runtime Python source; validation helpers, tests and generated evidence are preserved separately. This optimizes the Python multimodal worker, not the native Rust/CUDA worker.
Test Plan
System1-Omni Version / Commit:
1a339e1b63631271c0a180cb3c848d8698d61bfa.Fresh validation started 2026-09-30 on RTX 4090 24 GiB, driver 595.71.05, Torch 2.14.0+cu130, Transformers 5.17.0 and PEFT 0.21.0; BF16 base with unmerged adapter, without FLA/causal-conv1d.
PYTHONPATH=src:recipe/cua_s1 python -m pytest tests/cua_s1 -q.git diff <runtime> <validation> -- srcis empty. Source hashes and numerical/performance reports are independently verified.Test Result
Full evidence and tradeoffs · Raw primary report · Independent verification summary.
This run validates correctness on synthetic inputs; it makes no new latency or throughput claim.
Self-review
Agent-assisted full-diff review and validation limits found no actionable correctness or architecture defect. Runtime scope, commands and claims were checked against the recorded sources/results. The contributor remains responsible for understanding the changes; this does not replace maintainer review.