Why
#22 found that --model laya spent about 33 s of a 35 s load drawing random weights that the checkpoint then replaced. The "Adding a backend" steps in docs/decision-models.md do not include timing the load.
Julia-1 (#25) is the next in-process backend: a 144M ModernBERT encoder (mmBERT-small) that runs on CPU. Whether it pays the same cost depends on how its package builds the encoder; nobody has checked yet. #25 already asks for load time as a separate measurement. CLM (#24) runs behind its own server and is out of scope.
What
Add a step 5 to "Adding a backend" in docs/decision-models.md:
- For an in-process backend, time a cold
from_env() in a fresh process and put the number in the PR. Above 5 s, profile it. A loader that builds the model from its config (AutoModel.from_config or a bare nn.Module) and then loads a checkpoint over every weight can skip the random init with transformers' no_init_weights, provided the load fails on missing keys (strict=True or an equivalent check). Fix it in the model's package where possible; laya did so in 0.3.9.
Verification
The rendered docs/decision-models.md shows step 5, and #25's PR reports a cold load time.
Why
#22 found that
--model layaspent about 33 s of a 35 s load drawing random weights that the checkpoint then replaced. The "Adding a backend" steps indocs/decision-models.mddo not include timing the load.Julia-1 (#25) is the next in-process backend: a 144M ModernBERT encoder (mmBERT-small) that runs on CPU. Whether it pays the same cost depends on how its package builds the encoder; nobody has checked yet. #25 already asks for load time as a separate measurement. CLM (#24) runs behind its own server and is out of scope.
What
Add a step 5 to "Adding a backend" in
docs/decision-models.md:Verification
The rendered
docs/decision-models.mdshows step 5, and #25's PR reports a cold load time.