Skip to content

[RFC]: Native Laya inference with Rust and CUDA #14

Description

@linear3735

Motivation

#2 added the Rust frontend with an external model worker. This proposal adds an in-repository Laya engine that loads the original weights and runs inference directly through CUDA.

Delivery

Implementation will land in small PRs, with at most 500 core changed lines per PR. #21 starts with CPU checkpoint loading and validation. Input processing, model execution, CUDA optimizations and HTTP integration follow as separate changes. The full prototype from #16 remains available for reference.

Design

Start with Laya 0.3.20's English model and the BF16 fast path on Hopper. Support choice, score, and noul through the existing HTTP API.

HTTP → Rust preprocessing → GPU worker → CUDA inference → Rust response

Follow the layout in #6:

Directory Responsibility
src/frontend/ HTTP, request lifecycle, and the small engine interface
src/models/laya/ Tokenizer, weights, model execution, and output decoding
src/backends/cuda/ CUDA resources, Graphs, and kernels
recipe/laya/native/ Build, run, and validation scripts
  • One GPU worker processes requests sequentially. The queue holds 32 requests and returns HTTP 503 when full. A 30-second deadline includes upload and queueing. Readiness follows loading and warmup. There is no cross-request batching.
  • CUDA Graphs cover the encoder and decision transformer. A new shape allocates buffers, runs two warmups, and captures synchronously. An LRU cache retains up to four shapes and 512 MiB of workspaces. The variable-size scorer stays outside Graph.
  • Use the optimized RoPE kernel, with separate QKV, RoPE, and attention launches.
  • Python and TileLang generate CUDA sources; nvcc builds a reusable liblaya_cuda.so bundle. Compiled binaries are not checked in. Deployment uses the Rust binary, checkpoint, bundle, and CUDA/cuBLAS libraries.

Implementation and checks

Implementation · Build and validation · Results and samples

Checks cover tokenizer and weight parity, eager/Graph outputs, and HTTP concurrency and lifecycle. Boundary cases produced the same decisions with a maximum numeric difference of 0.0014 against a 0.002 limit.

H800 BF16 measurements on 2026-09-27: warmed short-request engine wall time was 2.84 ms for official fast and 1.81 ms for native, both using the same RoPE. Native HTTP C1 median was 2.00 ms. The validation report records the measured source hashes and timing method.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions