Skip to content

Keep Laya weights and workspace on CUDA - #48

Draft
linear3735 wants to merge 17 commits into
ThinkFlowLab:mainfrom
linear3735:codex/laya-workspace
Draft

linear3735 wants to merge 17 commits into
ThinkFlowLab:mainfrom
linear3735:codex/laya-workspace

Conversation

@linear3735

@linear3735 linear3735 commented Sep 30, 2026 •

Copy link
Copy Markdown

Purpose

Upload Laya's 205 used tensors once and allocate 17 reusable workspace buffers. Validate shapes before allocation and release partial allocations on failure.

Part of #14. Depends on #21 and #47. Kept in draft until those merge.

The diff against main currently includes those dependencies. Review this increment: 190 core/configuration lines for residency and workspace.

Tests now live under the repository-root tests/ directory. Cargo target names and test coverage are unchanged.

Test Plan

cargo fmt --all --check
cargo clippy --workspace --locked --all-targets -- -D warnings
cargo test --workspace --locked
cargo build --workspace --release --locked

On H800, compare uploaded tensors with the Torch conversion oracle and check workspace capacity and reuse. Reproduction commands are in src/models/laya/README.md.

System1-Omni Version / Commit: 2f5ff4b; incremental base 9ef493e.

Test Result

fmt, strict Clippy, release build passed locally; 37 CPU tests passed and 7 checkpoint/GPU tests were skipped. The earlier H800 validation below uses the same Laya and CUDA source; GPU tests were not rerun for this test-directory update.

  • All 205 uploaded tensor hashes matched the oracle.
  • All 17 buffers passed two write/read rounds at three shapes, including 16×512.
  • Weight allocations: 803.72 MiB. Largest workspace: 225.11 MiB. Both exclude CUDA and library overhead.

Rust CI, Docs build and benchmark harness tests passed for this update.

Self-review

Before marking this PR ready for review or requesting maintainer review, complete
the self-review checklist.
Keep the PR in draft while this work is incomplete.
For agent assistance, use the optional precheck-pr skill.

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

linear3735 and others added 6 commits September 28, 2026 15:45
The inventory mismatch reported only that the sets were unequal, so the
most likely failure -- pointing the engine at a checkpoint that is not the
frozen one -- said nothing about which tensor was wrong or in which
direction. Both sets were already in scope.

Report the expected-only names as "missing" and the checkpoint-only names
as "unexpected", sorted, so the message is deterministic. The neighbouring
errors already name their tensor (duplicate expected tensor, shape mismatch,
unsupported dtype); this was the one that did not.

Adds a CPU test over the synthetic safetensors fixture that asserts both
directions and the ordering.

fmt, clippy -D warnings, and the workspace tests pass; the checkpoint test
that exercises this path still passes against the frozen checkpoint.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant