cua_s1: accept image embeddings and 3D positions in native language model - #56
Levius-Fubuki wants to merge 2 commits into
Conversation
|
Reviewed The HTTP text worker gives the same answers as before on this card:
Today's worker never calls
The 139-token request itself takes about 12.8 ms. On the 712-token request the extra time was smaller and noisier, and I didn't look into why. Text outputs after a multimodal call matched the text-only ones every time. Keeping a device copy of the text tables and copying it back would remove nearly all of this. #52 is on main now, so this needs a rebase, and the Minor:
|
93a6cb0 to
dea93f2
Compare
dea93f2 to
75a80cb
Compare
|
@twu3202 Rebased onto main
Added a full-model GPU regression for first-use misses, hits with changed IDs, eviction after eight cached lengths, mixed multimodal/text calls and scratch growth, comparing hidden states against eager execution and asserting capture succeeds. The existing feature-insertion regression remains. Local format, Clippy, release build, all 25 Rust CPU tests, all 7 benchmark tests, smoke-manifest validation and strict Docs build pass. Five GPU tests are ignored by default. I could not execute the new/rebased GPU checks: the server used for the original evidence is shut down and SSH refuses connections. I have therefore changed the PR to draft pending fresh ABI-3 GPU parity/replay checks. The PR description distinguishes current CPU results from the historical Update: GitHub Rust, benchmark and Docs CI all pass on |
Purpose
Add
Model::forward_multimodalto accept adapted BF16 image features and explicit int64 T/H/W positions ([3,1,S]). Validate the unpadded prompt, insert image features into ordered placeholders, and execute the language layers with interleaved MRoPE. Accept standalone merged-language weights and preserve the text CUDA Graph cache with separate immutable text position tables.The diff contains only
inputs.rs,lib.rs, andmodel.rs. CUDA ABI remains 3; the HTTP worker remains text-only.Validation
After cleanup: workspace formatting, strict all-target Clippy, workspace tests and the locked release build passed. 19 tests passed; 3 existing GPU tests skipped. Runtime source was compared against the previous head: only new trailing test modules were removed; production code is unchanged. Existing baseline tests remain.
Feature-specific tests, examples, documentation and validation assets were removed from this PR diff to keep it focused on core implementation. They remain in the pre-cleanup commit and a local archive.
No new standalone GPU run was performed for this cleanup. The integrated language/Graph path was exercised as part of #64; that evidence applies to its combined runtime.
Self-review
Cleanup reviewed for unchanged runtime code, retained licensing and valid build targets. Draft status is retained for contributor/maintainer review.