Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -26,5 +26,17 @@ jobs:
run: cargo clippy --workspace --locked --all-targets -- -D warnings
- name: Test
run: cargo test --workspace --locked
- name: Official Laya packing parity
env:
LAYA_TOKENIZER: ${{ runner.temp }}/laya-tokenizer.json
LAYA_PACKING_ORACLE: ${{ runner.temp }}/laya-packing.json
run: |
curl --fail --location --retry 3 \
https://huggingface.co/convaiinnovations/laya/resolve/55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851/tokenizer/tokenizer.json \
--output "$LAYA_TOKENIZER"
curl --fail --location --retry 3 \
https://raw.githubusercontent.com/linear3735/system1-omni/5e4dd4215c925ebd93bb9ce4097b27bd6375f7c0/recipe/laya/native/packing-golden.json \
--output "$LAYA_PACKING_ORACLE"
cargo test --locked -p omni-laya --test packing -- --ignored
- name: Build
run: cargo build --workspace --release --locked
161 changes: 158 additions & 3 deletions Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

2 changes: 1 addition & 1 deletion Cargo.toml
Original file line number Diff line number Diff line change
@@ -1,3 +1,3 @@
[workspace]
members = ["src/frontend", "src/models/cua_s1/native"]
members = ["src/frontend", "src/models/cua_s1/native", "src/models/laya"]
resolver = "3"
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -49,15 +49,15 @@ Implementation code lives under `src/`; recipes and documentation stay at the re
| [`recipe/`](recipe/) | Model setup instructions, launch commands, configuration examples, and example requests. |
| [`docs/`](docs/) | Project documentation and architecture assets. |

The frontend and the Cua-S1 native worker are Cargo workspace members. The other model and backend directories currently document planned work; they do not prescribe process boundaries.
The frontend, Cua-S1 native worker and Laya checkpoint reader are Cargo workspace members. The other model and backend directories currently document planned work; they do not prescribe process boundaries.

## Supported models

LAYA can run as an external Python worker for text requests; its in-repository model engine is still planned. The Cua-S1 4B 0.2 `text` adapter runs as a Python worker or as a native worker on CUDA:

| Model | Status |
| --- | --- |
| LAYA | [External worker](recipe/laya/README.md); model engine planned |
| LAYA | [External worker](recipe/laya/README.md); [CPU checkpoint reader](src/models/laya/README.md); model execution planned |
| Cua-S1 4B 0.2 (`text` adapter) | [Python worker](recipe/cua_s1/text.md); [native worker](recipe/cua_s1/native.md), CUDA, run on sm_89 |

CUDA and Metal coverage will be documented per model as implementations are added and validated.
Expand Down
27 changes: 27 additions & 0 deletions recipe/laya/native/export_weights.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
"""CPU reference hashes for each tensor's FP32, FP16 and BF16 conversions."""

import argparse, hashlib, json
from pathlib import Path
import torch
from safetensors import safe_open

p = argparse.ArgumentParser()
p.add_argument("checkpoint", type=Path)
p.add_argument("output", type=Path)
a = p.parse_args()
torch.set_num_threads(4)
rows = []
with safe_open(a.checkpoint / "model.safetensors", framework="pt", device="cpu") as f:
for name in f.keys():
x = f.get_tensor(name)
row = {"name": name, "shape": list(x.shape), "source_dtype": str(x.dtype)}
for key, dtype in [
("f32", torch.float32),
("f16", torch.float16),
("bf16", torch.bfloat16),
]:
y = x.to(torch.float32).to(dtype).contiguous()
row[key] = hashlib.sha256(y.view(torch.uint8).numpy().tobytes()).hexdigest()
rows.append(row)
a.output.write_text(json.dumps(rows, indent=2))
print("WEIGHT_ORACLE", len(rows), flush=True)
2 changes: 1 addition & 1 deletion src/models/cua_s1/native/Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@ libloading = "0.8"
memmap2 = "0.9.9"
safetensors = "0.8.0"
serde = "1"
serde_json = { version = "1.0.149", features = ["float_roundtrip", "preserve_order"] }
serde_json = { version = "1.0.149", features = ["float_roundtrip", "preserve_order", "raw_value"] }
# the onig regex backend, as in the Python tokenizers wheel
tokenizers = { version = "=0.22.2", default-features = false, features = ["onig"] }
tokio = { version = "1.49.0", features = ["macros", "net", "rt-multi-thread", "sync"] }
Loading
Loading