Skip to content

[Help wanted] Accelerated Laya serving on Apple Silicon #3

Description

@hsliuustc0106

Community help wanted

We are looking for community contributors to help build and evaluate faster local Laya serving on Apple Silicon. Contributions to one part of the work are welcome; you do not need to implement the whole backend.

Useful ways to help:

  • Profile Laya on PyTorch MPS and identify where inference time is spent.
  • Implement and evaluate targeted optimizations; bring MLX, Core ML, or Metal expertise where measurements justify an alternative backend.
  • Run reproducible benchmarks on available M-series Macs and share hardware details, commands, raw results, and output-parity checks.
  • Help with persistent serving, readiness checks, tests, and setup documentation.

Please comment with the part you would like to tackle, your available hardware if relevant, and any proposed approach so contributors can coordinate before overlapping work. Baseline measurements, unsuccessful optimization results, and small focused PRs are all useful.

Agent-side integration is coordinated in ThinkFlowLab/system1-agents#20.

Problem

We want to serve Laya locally on Apple Silicon Macs with lower decision latency and a model that stays loaded across requests. Issue #1 introduces a Rust HTTP frontend, but leaves tokenization and inference in the backend. That frontend alone does not establish a model inference speedup.

Proposed scope

Add and validate an Apple Silicon inference backend for Laya behind the serving interface from #1. Start with the existing Laya PyTorch MPS implementation as the reference, profile its bottlenecks, and choose the smallest optimization that produces a measured improvement. Evaluate alternatives such as MLX or Core ML only if profiling justifies the additional implementation and maintenance.

  • Keep the selected checkpoint resident across requests and perform a representative warmup before reporting readiness.
  • Support Laya text decisions (choice, score, and noul) through the common /v1/systemone interface, preserving response fields and probability semantics.
  • Make the selected device/backend explicit and observable; do not silently report CPU execution as GPU acceleration.
  • Document supported Mac hardware, macOS/runtime versions, checkpoint, installation, startup, and a reproducible request example.
  • Keep custom batching and scheduling out of the initial scope unless measurements identify them as necessary for the stated workload.

Validation

Before implementation benchmarking, define the hypothesis, independent variable, fixed controls, success criterion, and stop condition. Compare the reference PyTorch MPS backend with the proposed backend using the same Mac, checkpoint, inputs, question types, concurrency, and warmup/cache policy.

Report separately:

  • Checkpoint download, model loading, warmup, and process-to-readiness time.
  • Warm inference latency and end-to-end HTTP latency (p50/p95), including the frontend overhead.
  • Throughput and memory use at a declared concurrency.
  • Decision agreement and probability differences on a fixed set spanning all three decision types, input lengths, and option counts. Define numerical tolerances before comparing results.

Use one feasibility run, then two measured runs per configuration by default, with enough requests per run to support percentile estimates. Record variability, repository SHAs, runtime versions, commands, and raw results. If the run budget does not establish an improvement, report that result rather than claiming acceleration. GPU tests on scheduler-managed hosts must follow their reservation policy.

Acceptance criteria

  • A persistent Laya backend runs on an Apple Silicon Mac through the serving API.
  • Readiness follows model loading and warmup; the first inference request after readiness succeeds.
  • Contract and output-parity checks cover choice, score, and noul.
  • Reproducible benchmarks separate backend inference improvements from HTTP frontend overhead.
  • Results show whether the predeclared latency target was met without unacceptable output differences; limitations are documented.
  • Setup and benchmark instructions are published.

Related: #1. This issue covers local model execution and measured Apple Silicon acceleration; #1 covers the common frontend and transport.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions