Skip to content

[Docs] Time a cold model load when adding an in-process backend #29

Description

@Yunaik

Why

#22 found that --model laya spent about 33 s of a 35 s load drawing random weights that the checkpoint then replaced. The "Adding a backend" steps in docs/decision-models.md do not include timing the load.

Julia-1 (#25) is the next in-process backend: a 144M ModernBERT encoder (mmBERT-small) that runs on CPU. Whether it pays the same cost depends on how its package builds the encoder; nobody has checked yet. #25 already asks for load time as a separate measurement. CLM (#24) runs behind its own server and is out of scope.

What

Add a step 5 to "Adding a backend" in docs/decision-models.md:

  1. For an in-process backend, time a cold from_env() in a fresh process and put the number in the PR. Above 5 s, profile it. A loader that builds the model from its config (AutoModel.from_config or a bare nn.Module) and then loads a checkpoint over every weight can skip the random init with transformers' no_init_weights, provided the load fails on missing keys (strict=True or an equivalent check). Fix it in the model's package where possible; laya did so in 0.3.9.

Verification

The rendered docs/decision-models.md shows step 5, and #25's PR reports a cold load time.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    documentationImprovements or additions to documentation

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions