Add native Laya inference with Rust and CUDA - #16
Closed
linear3735 wants to merge 13 commits into
Closed
linear3735 wants to merge 13 commits into
linear3735 wants to merge 13 commits into
Conversation
added 13 commits
September 27, 2026 19:47
4 tasks done
4 tasks
xiaoyu-xyz
added a commit
to xiaoyu-xyz/system1-omni
that referenced
this pull request
Sep 28, 2026
CUDA build paths are now in flight and diverging: Laya generates CUDA from TileLang, Cua-S1 hand-writes CUDA C++ with cuBLASLt, and the multimodal worker plans Triton. Nothing links against anything else yet, so the divergence is invisible today and blocking as soon as one model reuses another's kernels — which ThinkFlowLab#9 already plans, since a Kev engine and Cua-S1 share a Qwen3.5 backbone. This adds the contract and the checker, and nothing else: - contract.md: what backends must agree on — a JSON manifest per backend, four C ABI rules, the numerics that must be declared rather than discovered in a parity failure, and the build-script interface. - check_contract.py: reads the manifests and checks them against that contract. - build_script.py: reads a build script as text and reports what it declares. Never executed, because CI must not run repository code to decide whether a manifest is honest. A backend may be a subdirectory named after itself or sit directly under src/backends/cuda/, which is the layout Laya's kernels/ and tools/ use. The repository root is passed down explicitly rather than inferred by walking up from the backend, because the flat layout makes that guess one level wrong. The compile and parity tiers need hardware and are described in contract.md but deliberately not wired up here. This change needs no CUDA toolkit and no GPU, which is why it can land ahead of them. 53 tests pass. The checker was run against the real qwen3_5/ files from ThinkFlowLab#19, which it accepts, and against the real tools/ from ThinkFlowLab#16, which it rejects until that backend exposes an executable build script — the divergence it exists to surface.
This was referenced Sep 29, 2026
xiaoyu-xyz
added a commit
to xiaoyu-xyz/system1-omni
that referenced
this pull request
Sep 30, 2026
The previous change defined the Engine contract and the HTTP service behind it. This makes them usable, and is where #3's acceptance criteria for persistent serving live: readiness that follows loading and warmup, a queue that refuses rather than grows, a request budget that can expire before work starts, and a shutdown that drains accepted work. - worker.rs: the thread that owns the engine, admission, readiness, the budget and drain. spawn returns only after load and warmup, so a handle cannot observe a half-loaded engine. A panic in model code answers its caller and retires the engine instead of unwinding the worker; bad input retires nothing. - main.rs: mode selection and shutdown wiring. The forwarding path is unchanged. - passthrough.rs: a stand-in engine that echoes requests, so the native path can be run and tested with no model and no GPU. Review follow-up, both with a test that fails against the reviewed commit: - Queue depth is reserved before the job is published, and rolled back when the send fails. The worker does not take the admission lock, so a job in the queue can be dequeued, answered and released before the producer counted it, which released a slot that had not been taken and wrapped the counter. - The drain budget starts when shutdown is signalled, not when the listener finally stops. Admission closes inside the shutdown future and one absolute deadline bounds every step. Before this, a request whose body never arrived held the process open for its whole request budget: 3.009 s to exit with a 200 ms drain budget. Five behavioural differences from the closed ThinkFlowLab#16, each a case it does not cover: /health publishes queue depth and rejections rather than a bool; SIGTERM drains and joins the engine thread instead of returning while it runs; admission and shutdown share a lock, so a request cannot be accepted with no worker left to answer it; the deadline reaches the engine, so it can drop work whose client is gone; model code is one function rather than its own worker loop. 647 lines of core Rust across three files, no test helpers. Cargo.lock unchanged. Verified without hardware: fmt and clippy -D warnings clean, 34 tests pass six times, and the built binary in native mode serves, refuses over-size bodies, reports readiness and exits 0 on SIGTERM within its drain budget.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Superseded by #21. The first small PR covers CPU checkpoint loading; model execution and CUDA changes will follow separately. The full prototype remains on
codex/laya-native.Purpose
Refs #14.
Run the English Laya model in Rust and CUDA using the original weights. Support
choice,score, andnoulthrough the HTTP interface, with optimized RoPE, CUDA Graphs, and a bounded single-worker queue. Serving needs no Python worker.This PR targets Hopper BF16. Model execution lives in
src/models/laya/, CUDA code insrc/backends/cuda/, and build and validation scripts inrecipe/laya/native/.Test Plan
Run formatting, strict Clippy, all-feature workspace tests, and a release build. Compare eager/Graph outputs and HTTP behavior against official Laya; measure warmed engine latency on H800 using the same RoPE in both paths.
System1-Omni Version / Commit:
5e4dd42, based on upstream30622438.Test Result
Validation details, measured source hashes, and samples.