Skip to content

Add native Laya inference with Rust and CUDA - #16

Closed
linear3735 wants to merge 13 commits into
ThinkFlowLab:mainfrom
linear3735:codex/laya-native
Closed

linear3735 wants to merge 13 commits into
ThinkFlowLab:mainfrom
linear3735:codex/laya-native

Conversation

@linear3735

@linear3735 linear3735 commented Sep 28, 2026 •

Copy link
Copy Markdown

Superseded by #21. The first small PR covers CPU checkpoint loading; model execution and CUDA changes will follow separately. The full prototype remains on codex/laya-native.

Purpose

Refs #14.

Run the English Laya model in Rust and CUDA using the original weights. Support choice, score, and noul through the HTTP interface, with optimized RoPE, CUDA Graphs, and a bounded single-worker queue. Serving needs no Python worker.

This PR targets Hopper BF16. Model execution lives in src/models/laya/, CUDA code in src/backends/cuda/, and build and validation scripts in recipe/laya/native/.

Test Plan

Run formatting, strict Clippy, all-feature workspace tests, and a release build. Compare eager/Graph outputs and HTTP behavior against official Laya; measure warmed engine latency on H800 using the same RoPE in both paths.

System1-Omni Version / Commit: 5e4dd42, based on upstream 30622438.

Test Result

  • CPU CI passed: formatting, strict Clippy, 22 tests, and the release build. Four checkpoint/oracle tests were ignored.
  • GPU validation on 2026-09-27: eager/Graph responses matched on 12 requests each. Boundary error was at most 0.0014 within the 0.002 limit, with unchanged decisions. HTTP concurrency and lifecycle checks passed.
  • H800 BF16 short-request engine wall time: official fast 2.84 ms, native 1.81 ms. Both used the same RoPE. Native HTTP C1 median: 2.00 ms.

Validation details, measured source hashes, and samples.

@twu3202 twu3202 mentioned this pull request Sep 28, 2026
4 tasks done
@linear3735 linear3735 closed this Sep 28, 2026
xiaoyu-xyz added a commit to xiaoyu-xyz/system1-omni that referenced this pull request Sep 28, 2026
CUDA build paths are now in flight and diverging: Laya generates CUDA from
TileLang, Cua-S1 hand-writes CUDA C++ with cuBLASLt, and the multimodal worker
plans Triton. Nothing links against anything else yet, so the divergence is
invisible today and blocking as soon as one model reuses another's kernels —
which ThinkFlowLab#9 already plans, since a Kev engine and Cua-S1 share a Qwen3.5 backbone.

This adds the contract and the checker, and nothing else:

- contract.md: what backends must agree on — a JSON manifest per backend, four
  C ABI rules, the numerics that must be declared rather than discovered in a
  parity failure, and the build-script interface.
- check_contract.py: reads the manifests and checks them against that contract.
- build_script.py: reads a build script as text and reports what it declares.
  Never executed, because CI must not run repository code to decide whether a
  manifest is honest.

A backend may be a subdirectory named after itself or sit directly under
src/backends/cuda/, which is the layout Laya's kernels/ and tools/ use. The
repository root is passed down explicitly rather than inferred by walking up from
the backend, because the flat layout makes that guess one level wrong.

The compile and parity tiers need hardware and are described in contract.md but
deliberately not wired up here. This change needs no CUDA toolkit and no GPU,
which is why it can land ahead of them.

53 tests pass. The checker was run against the real qwen3_5/ files from ThinkFlowLab#19,
which it accepts, and against the real tools/ from ThinkFlowLab#16, which it rejects until
that backend exposes an executable build script — the divergence it exists to
surface.
xiaoyu-xyz added a commit to xiaoyu-xyz/system1-omni that referenced this pull request Sep 30, 2026
The previous change defined the Engine contract and the HTTP service behind it.
This makes them usable, and is where #3's acceptance criteria for persistent
serving live: readiness that follows loading and warmup, a queue that refuses
rather than grows, a request budget that can expire before work starts, and a
shutdown that drains accepted work.

- worker.rs: the thread that owns the engine, admission, readiness, the budget
  and drain. spawn returns only after load and warmup, so a handle cannot observe
  a half-loaded engine. A panic in model code answers its caller and retires the
  engine instead of unwinding the worker; bad input retires nothing.
- main.rs: mode selection and shutdown wiring. The forwarding path is unchanged.
- passthrough.rs: a stand-in engine that echoes requests, so the native path can
  be run and tested with no model and no GPU.

Review follow-up, both with a test that fails against the reviewed commit:

- Queue depth is reserved before the job is published, and rolled back when the
  send fails. The worker does not take the admission lock, so a job in the queue
  can be dequeued, answered and released before the producer counted it, which
  released a slot that had not been taken and wrapped the counter.
- The drain budget starts when shutdown is signalled, not when the listener
  finally stops. Admission closes inside the shutdown future and one absolute
  deadline bounds every step. Before this, a request whose body never arrived
  held the process open for its whole request budget: 3.009 s to exit with a
  200 ms drain budget.

Five behavioural differences from the closed ThinkFlowLab#16, each a case it does not cover:
/health publishes queue depth and rejections rather than a bool; SIGTERM drains
and joins the engine thread instead of returning while it runs; admission and
shutdown share a lock, so a request cannot be accepted with no worker left to
answer it; the deadline reaches the engine, so it can drop work whose client is
gone; model code is one function rather than its own worker loop.

647 lines of core Rust across three files, no test helpers. Cargo.lock unchanged.
Verified without hardware: fmt and clippy -D warnings clean, 34 tests pass six
times, and the built binary in native mode serves, refuses over-size bodies,
reports readiness and exits 0 on SIGTERM within its drain budget.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant