Motivation
#2 added the Rust frontend with an external model worker. This proposal adds an in-repository Laya engine that loads the original weights and runs inference directly through CUDA.
Delivery
Implementation will land in small PRs, with at most 500 core changed lines per PR. #21 starts with CPU checkpoint loading and validation. Input processing, model execution, CUDA optimizations and HTTP integration follow as separate changes. The full prototype from #16 remains available for reference.
Design
Start with Laya 0.3.20's English model and the BF16 fast path on Hopper. Support choice, score, and noul through the existing HTTP API.
HTTP → Rust preprocessing → GPU worker → CUDA inference → Rust response
Follow the layout in #6:
| Directory |
Responsibility |
src/frontend/ |
HTTP, request lifecycle, and the small engine interface |
src/models/laya/ |
Tokenizer, weights, model execution, and output decoding |
src/backends/cuda/ |
CUDA resources, Graphs, and kernels |
recipe/laya/native/ |
Build, run, and validation scripts |
- One GPU worker processes requests sequentially. The queue holds 32 requests and returns HTTP 503 when full. A 30-second deadline includes upload and queueing. Readiness follows loading and warmup. There is no cross-request batching.
- CUDA Graphs cover the encoder and decision transformer. A new shape allocates buffers, runs two warmups, and captures synchronously. An LRU cache retains up to four shapes and 512 MiB of workspaces. The variable-size scorer stays outside Graph.
- Use the optimized RoPE kernel, with separate QKV, RoPE, and attention launches.
- Python and TileLang generate CUDA sources; nvcc builds a reusable
liblaya_cuda.so bundle. Compiled binaries are not checked in. Deployment uses the Rust binary, checkpoint, bundle, and CUDA/cuBLAS libraries.
Implementation and checks
Implementation · Build and validation · Results and samples
Checks cover tokenizer and weight parity, eager/Graph outputs, and HTTP concurrency and lifecycle. Boundary cases produced the same decisions with a maximum numeric difference of 0.0014 against a 0.002 limit.
H800 BF16 measurements on 2026-09-27: warmed short-request engine wall time was 2.84 ms for official fast and 1.81 ms for native, both using the same RoPE. Native HTTP C1 median was 2.00 ms. The validation report records the measured source hashes and timing method.
Motivation
#2 added the Rust frontend with an external model worker. This proposal adds an in-repository Laya engine that loads the original weights and runs inference directly through CUDA.
Delivery
Implementation will land in small PRs, with at most 500 core changed lines per PR. #21 starts with CPU checkpoint loading and validation. Input processing, model execution, CUDA optimizations and HTTP integration follow as separate changes. The full prototype from #16 remains available for reference.
Design
Start with Laya 0.3.20's English model and the BF16 fast path on Hopper. Support
choice,score, andnoulthrough the existing HTTP API.Follow the layout in #6:
src/frontend/src/models/laya/src/backends/cuda/recipe/laya/native/liblaya_cuda.sobundle. Compiled binaries are not checked in. Deployment uses the Rust binary, checkpoint, bundle, and CUDA/cuBLAS libraries.Implementation and checks
Implementation · Build and validation · Results and samples
Checks cover tokenizer and weight parity, eager/Graph outputs, and HTTP concurrency and lifecycle. Boundary cases produced the same decisions with a maximum numeric difference of 0.0014 against a 0.002 limit.
H800 BF16 measurements on 2026-09-27: warmed short-request engine wall time was 2.84 ms for official fast and 1.81 ms for native, both using the same RoPE. Native HTTP C1 median was 2.00 ms. The validation report records the measured source hashes and timing method.