Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
22 commits
Select commit Hold shift + click to select a range
c92919b
Serve Laya on Apple Silicon: worker, benchmarks and recipe
cacheline999 Sep 28, 2026
a2e3609
Add measured Laya reports and a paired frontend-overhead probe
cacheline999 Sep 28, 2026
630d13b
Warm every loaded Laya model and harden the HTTP benchmark
cacheline999 Sep 28, 2026
3bf7c3b
Move the Laya benchmark scripts under recipe/laya/bench
cacheline999 Sep 28, 2026
bc42882
Add LAYA_WORKER_WEIGHTS=fp16 for the compiled worker
cacheline999 Sep 30, 2026
836f5ac
Compile the encoder for multi-question batches; make compile on/off
cacheline999 Sep 30, 2026
2af396f
Recipe: M5 results, test dependencies, first-request note
cacheline999 Sep 30, 2026
37f22e0
Say that fp16 weights are for MPS and warn on CPU
cacheline999 Sep 30, 2026
a07dcfd
Read /health from the live model; correct the Apple Silicon recipe
cacheline999 Sep 30, 2026
4785a39
Run plain fp32 laya after a fallback to the CPU
cacheline999 Sep 30, 2026
a5be4f8
Split the GPU optimizations out of worker.py; correct the Laya README
cacheline999 Sep 30, 2026
5f43fc4
Lay the Laya worker out like the Cua-S1 worker (#19)
cacheline999 Oct 1, 2026
f883827
Laya worker: keep measurements out of comments, report compile.active…
cacheline999 Oct 1, 2026
351b8ec
Bench scripts: mark the imports that follow the sys.path setup (E402)
cacheline999 Oct 1, 2026
235774a
Laya worker: prepare checkpoints loaded while serving; fix report par…
cacheline999 Oct 2, 2026
02c68c5
Laya worker: restore checkpoints evicted for a failed late load; say …
cacheline999 Oct 2, 2026
9a8ee98
Laya recipe: say what a late checkpoint load costs through the fronte…
cacheline999 Oct 2, 2026
64c965f
Laya worker: drop comments that repeat the code or the test name
cacheline999 Oct 2, 2026
8d15710
Laya: third review fixes, requirement-derived tests on CPU and GPU, a…
cacheline999 Oct 2, 2026
fc0a526
Laya: keep profile runs out of the parity section; header test off ma…
cacheline999 Oct 2, 2026
ec48a04
Laya bench and tests: two-sided answer check in paired summaries; rea…
cacheline999 Oct 2, 2026
539f06d
Laya: move the benchmark scripts from recipe/laya/bench to benchmarks…
cacheline999 Oct 3, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -57,7 +57,7 @@ LAYA can run as an external Python worker for text requests; its in-repository m

| Model | Status |
| --- | --- |
| LAYA | [External worker](recipe/laya/README.md); [CPU checkpoint reader](src/models/laya/README.md); model execution planned |
| LAYA | [External worker](recipe/laya/README.md); [Python worker on Apple Silicon (MPS) and CPU](recipe/laya/apple-silicon.md); [CPU checkpoint reader](src/models/laya/README.md); model execution planned |
| Cua-S1 4B 0.2 (`text` adapter) | [Python worker](recipe/cua_s1/text.md); [native worker](recipe/cua_s1/native.md), CUDA, run on sm_89 |

[Supported models and hardware](docs/supported-models.md) lists the devices and where each worker has been run.
Expand Down
6 changes: 6 additions & 0 deletions benchmarks/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -169,3 +169,9 @@ result is not a native-engine speedup. Metal and vision benchmarks follow CUDA.
```sh
python -m unittest discover -s tests/benchmarks -p 'test_*.py' -v
```

## Laya on Apple Silicon

[`laya_mps/`](laya_mps/README.md) holds the scripts behind the numbers of the
[Apple Silicon recipe](../recipe/laya/apple-silicon.md): in-process and HTTP latency, paired comparisons,
profiling and the report with its output-parity section.
60 changes: 60 additions & 0 deletions benchmarks/laya_mps/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,60 @@
# Laya on Apple Silicon: benchmark scripts

Scripts behind the numbers in the [Apple Silicon recipe](../../recipe/laya/apple-silicon.md). Each run writes raw
JSONL to `results/` (kept out of the repository); `report.py` builds the tables from it.

| file | purpose |
| --- | --- |
| `workloads.jsonl` | the fixed inputs: W1–W6 are timed, P* are for answer comparison only. Tokens per row with Laya's tokenizer: W1 68, W2 198, W3 484, W4 68/48/47, W5 40–68, W6 47 |
| `bench_inproc.py` | Laya in-process (no HTTP): load, warmup, first request, warm latency, memory |
| `bench_http.py` | a `/v1/systemone` server, optionally started by the script and optionally behind the frontend: time to ready, first request, warm latency, throughput |
| `paired.py` | two configurations compared request by request, both alive at once: two worker flag sets, or two running servers (e.g. a worker directly and through the frontend) |
| `profile_mps.py` | where a request's time goes on MPS |
| `report.py` | tables from the JSONL, including the run-to-run gate and the answer comparison against a reference config |
| `env.py` | shared: versions, checkpoint, hardware, power and load recorded with each run; memory footprint |

## Run

From the repository root, in the recipe's environment (`.venv`), with the frontend built:

```sh
python benchmarks/laya_mps/bench_inproc.py --device cpu --config C1 --run m1
python benchmarks/laya_mps/bench_inproc.py --device mps --config C2 --run m1
python benchmarks/laya_mps/bench_http.py --config C3 --run m1 --spawn .venv/bin/laya-serve
python benchmarks/laya_mps/bench_http.py --config C4 --run m1 --url http://127.0.0.1:8080 \
--frontend target/release/omni-jev --spawn .venv/bin/laya-serve
python benchmarks/laya_mps/bench_http.py --config C3w --run m1 \
--spawn .venv/bin/python -m frontend.laya_mps --device {device} --model {model} --port {port}
python benchmarks/laya_mps/bench_http.py --config C3o --run m1 \
--spawn .venv/bin/python -m frontend.laya_mps --device {device} --model {model} --compile --weights fp16 --port {port}
python benchmarks/laya_mps/report.py benchmarks/laya_mps/results/*_m[0-9].jsonl --ref C1
```

Repeat with `--run m2` for a second measured run. Runs refuse to start on battery power or above a
1-minute load average of `--max-load` (default 2) unless labelled `--run feasibility`. Memory is the
process's physical footprint, which on Apple Silicon includes MPS allocations.

Two configurations compared request by request, which holds up under background load better than
separate runs:

```sh
python benchmarks/laya_mps/paired.py --run p1 --a "" --b "--compile --weights fp16"
python benchmarks/laya_mps/paired.py --run f1 --a-url http://127.0.0.1:8000 --b-url http://127.0.0.1:8080
python benchmarks/laya_mps/paired.py --summarize benchmarks/laya_mps/results/paired_p1.jsonl
```

## Results

The measured runs on an M1 Pro are published as assets of one release on the fork,
<https://github.com/cacheline999/system1-omni/releases/tag/laya-mps-results-2026-09-28>:

| asset | contents | sha256 |
| --- | --- | --- |
| `laya-mps-reports-2026-10-01.tar.gz` | the tables: baseline report and parity, frontend overhead, paired fp16, paired all optimizations | `e857da5082da4983e07a91104b20f84fb1ffd7994c56123e0a832e0d7a870cea` |
| `laya-mps-results-2026-09-28.tar.gz` | raw JSONL of the baseline runs (C1–C4, C3w, C3s) | `611ed30707ac8c98875b5aa5382360b5a7d760da166d61c626eb07ebe1ee6404` |
| `laya-mps-paired-fp16-2026-09-30.tar.gz` | raw JSONL of the paired fp16 runs | `cc0d6f5bda6e3e0ee1f40c6966f84429e902a2f585b8e1ee33658a9be139326e` |
| `laya-mps-paired-all-2026-09-30.tar.gz` | raw JSONL of the paired all-optimizations runs | `25d1b7bc9d6dff7173f9b972ebd0204f8fb4e2089fe2d27924d48ba6e589780c` |

Extract the raw JSONL into `results/` and run `report.py` or `paired.py --summarize` on it to rebuild
the tables. `C3s` in the baseline runs is an earlier compile mode that compiled one-question requests
only; `--compile` does the same for them and adds the encoder for several questions.
Loading
Loading