Community help wanted
We are looking for community contributors to help build and evaluate faster local Laya serving on Apple Silicon. Contributions to one part of the work are welcome; you do not need to implement the whole backend.
Useful ways to help:
- Profile Laya on PyTorch MPS and identify where inference time is spent.
- Implement and evaluate targeted optimizations; bring MLX, Core ML, or Metal expertise where measurements justify an alternative backend.
- Run reproducible benchmarks on available M-series Macs and share hardware details, commands, raw results, and output-parity checks.
- Help with persistent serving, readiness checks, tests, and setup documentation.
Please comment with the part you would like to tackle, your available hardware if relevant, and any proposed approach so contributors can coordinate before overlapping work. Baseline measurements, unsuccessful optimization results, and small focused PRs are all useful.
Agent-side integration is coordinated in ThinkFlowLab/system1-agents#20.
Problem
We want to serve Laya locally on Apple Silicon Macs with lower decision latency and a model that stays loaded across requests. Issue #1 introduces a Rust HTTP frontend, but leaves tokenization and inference in the backend. That frontend alone does not establish a model inference speedup.
Proposed scope
Add and validate an Apple Silicon inference backend for Laya behind the serving interface from #1. Start with the existing Laya PyTorch MPS implementation as the reference, profile its bottlenecks, and choose the smallest optimization that produces a measured improvement. Evaluate alternatives such as MLX or Core ML only if profiling justifies the additional implementation and maintenance.
- Keep the selected checkpoint resident across requests and perform a representative warmup before reporting readiness.
- Support Laya text decisions (
choice, score, and noul) through the common /v1/systemone interface, preserving response fields and probability semantics.
- Make the selected device/backend explicit and observable; do not silently report CPU execution as GPU acceleration.
- Document supported Mac hardware, macOS/runtime versions, checkpoint, installation, startup, and a reproducible request example.
- Keep custom batching and scheduling out of the initial scope unless measurements identify them as necessary for the stated workload.
Validation
Before implementation benchmarking, define the hypothesis, independent variable, fixed controls, success criterion, and stop condition. Compare the reference PyTorch MPS backend with the proposed backend using the same Mac, checkpoint, inputs, question types, concurrency, and warmup/cache policy.
Report separately:
- Checkpoint download, model loading, warmup, and process-to-readiness time.
- Warm inference latency and end-to-end HTTP latency (p50/p95), including the frontend overhead.
- Throughput and memory use at a declared concurrency.
- Decision agreement and probability differences on a fixed set spanning all three decision types, input lengths, and option counts. Define numerical tolerances before comparing results.
Use one feasibility run, then two measured runs per configuration by default, with enough requests per run to support percentile estimates. Record variability, repository SHAs, runtime versions, commands, and raw results. If the run budget does not establish an improvement, report that result rather than claiming acceleration. GPU tests on scheduler-managed hosts must follow their reservation policy.
Acceptance criteria
Related: #1. This issue covers local model execution and measured Apple Silicon acceleration; #1 covers the common frontend and transport.
Community help wanted
We are looking for community contributors to help build and evaluate faster local Laya serving on Apple Silicon. Contributions to one part of the work are welcome; you do not need to implement the whole backend.
Useful ways to help:
Please comment with the part you would like to tackle, your available hardware if relevant, and any proposed approach so contributors can coordinate before overlapping work. Baseline measurements, unsuccessful optimization results, and small focused PRs are all useful.
Agent-side integration is coordinated in ThinkFlowLab/system1-agents#20.
Problem
We want to serve Laya locally on Apple Silicon Macs with lower decision latency and a model that stays loaded across requests. Issue #1 introduces a Rust HTTP frontend, but leaves tokenization and inference in the backend. That frontend alone does not establish a model inference speedup.
Proposed scope
Add and validate an Apple Silicon inference backend for Laya behind the serving interface from #1. Start with the existing Laya PyTorch MPS implementation as the reference, profile its bottlenecks, and choose the smallest optimization that produces a measured improvement. Evaluate alternatives such as MLX or Core ML only if profiling justifies the additional implementation and maintenance.
choice,score, andnoul) through the common/v1/systemoneinterface, preserving response fields and probability semantics.Validation
Before implementation benchmarking, define the hypothesis, independent variable, fixed controls, success criterion, and stop condition. Compare the reference PyTorch MPS backend with the proposed backend using the same Mac, checkpoint, inputs, question types, concurrency, and warmup/cache policy.
Report separately:
Use one feasibility run, then two measured runs per configuration by default, with enough requests per run to support percentile estimates. Record variability, repository SHAs, runtime versions, commands, and raw results. If the run budget does not establish an improvement, report that result rather than claiming acceleration. GPU tests on scheduler-managed hosts must follow their reservation policy.
Acceptance criteria
choice,score, andnoul.Related: #1. This issue covers local model execution and measured Apple Silicon acceleration; #1 covers the common frontend and transport.