Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
93 changes: 91 additions & 2 deletions Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

2 changes: 1 addition & 1 deletion Cargo.toml
Original file line number Diff line number Diff line change
@@ -1,3 +1,3 @@
[workspace]
members = ["src/frontend", "src/models/cua_s1/native"]
members = ["src/frontend", "src/models/cua_s1/native", "src/models/laya"]
resolver = "3"
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -49,15 +49,15 @@ Implementation code lives under `src/`; recipes and documentation stay at the re
| [`recipe/`](recipe/) | Model setup instructions, launch commands, configuration examples, and example requests. |
| [`docs/`](docs/) | Project documentation and architecture assets. |

The frontend and the Cua-S1 native worker are Cargo workspace members. The other model and backend directories currently document planned work; they do not prescribe process boundaries.
The frontend, Cua-S1 native worker and Laya checkpoint reader are Cargo workspace members. The other model and backend directories currently document planned work; they do not prescribe process boundaries.

## Supported models

LAYA can run as an external Python worker for text requests; its in-repository model engine is still planned. The Cua-S1 4B 0.2 `text` adapter runs as a Python worker or as a native worker on CUDA:

| Model | Status |
| --- | --- |
| LAYA | [External worker](recipe/laya/README.md); model engine planned |
| LAYA | [External worker](recipe/laya/README.md); [CPU checkpoint reader](src/models/laya/README.md); model execution planned |
| Cua-S1 4B 0.2 (`text` adapter) | [Python worker](recipe/cua_s1/text.md); [native worker](recipe/cua_s1/native.md), CUDA, run on sm_89 |

CUDA and Metal coverage will be documented per model as implementations are added and validated.
Expand Down
27 changes: 27 additions & 0 deletions recipe/laya/native/export_weights.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
"""CPU reference hashes for each tensor's FP32, FP16 and BF16 conversions."""

import argparse, hashlib, json
from pathlib import Path
import torch
from safetensors import safe_open

p = argparse.ArgumentParser()
p.add_argument("checkpoint", type=Path)
p.add_argument("output", type=Path)
a = p.parse_args()
torch.set_num_threads(4)
rows = []
with safe_open(a.checkpoint / "model.safetensors", framework="pt", device="cpu") as f:
for name in f.keys():
x = f.get_tensor(name)
row = {"name": name, "shape": list(x.shape), "source_dtype": str(x.dtype)}
for key, dtype in [
("f32", torch.float32),
("f16", torch.float16),
("bf16", torch.bfloat16),
]:
y = x.to(torch.float32).to(dtype).contiguous()
row[key] = hashlib.sha256(y.view(torch.uint8).numpy().tobytes()).hexdigest()
rows.append(row)
a.output.write_text(json.dumps(rows, indent=2))
print("WEIGHT_ORACLE", len(rows), flush=True)
25 changes: 25 additions & 0 deletions src/models/laya/Cargo.toml
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
[package]
name = "omni-laya"
version = "0.1.0"
edition = "2024"
publish = false

[dependencies]
anyhow = "1"
half = "2"
memmap2 = "0.9"
safetensors = "0.6"
serde = { version = "1", features = ["derive"] }
serde_json = "1"

[dev-dependencies]
sha2 = "0.10"
tempfile = "3"

[[test]]
name = "checkpoint"
path = "../../../tests/laya/checkpoint.rs"

[[test]]
name = "weights"
path = "../../../tests/laya/weights.rs"
19 changes: 18 additions & 1 deletion src/models/laya/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,4 +4,21 @@ LAYA is the first planned System1-Omni model. This directory owns its complete r

GPU operations and kernel implementations belong in [`backends/cuda/`](../../backends/cuda/) and [`backends/metal/`](../../backends/metal/). Setup and usage examples belong in the top-level [`recipe/`](../../../recipe/) directory.

Status: planned; no model implementation or validated GPU backend support yet.
The `omni-laya` crate currently reads and checks the English Laya 0.3.20 checkpoint. `Config::load` validates the architecture and temperatures; `Weights` checks tensor names and shapes and converts FP32, FP16 and BF16 values. `checkpoint_tensors()` lists the 206 expected tensors. Each backend chooses its own storage precision.

Keep checkpoint files unchanged while `Weights` holds a read-only memory mapping. This crate does not yet execute inference.

## CPU checks

The normal workspace tests cover configuration errors, malformed tensors, inventory mismatches and conversion boundaries without downloading weights.

To check the complete checkpoint, use `convaiinnovations/laya` revision `55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851` and a Python environment with PyTorch, safetensors and NumPy:

```sh
export LAYA_CHECKPOINT=/path/to/laya/snapshot
export LAYA_WEIGHT_ORACLE=/tmp/laya-weight-oracle.json
python recipe/laya/native/export_weights.py "$LAYA_CHECKPOINT" "$LAYA_WEIGHT_ORACLE"
cargo test --release --locked -p omni-laya --test weights -- --ignored
```

These two CPU tests check all 206 tensor names and shapes, 618 conversion hashes, and the legacy temperature buffer. The normal CI job skips them because it does not download the full checkpoint.
91 changes: 91 additions & 0 deletions src/models/laya/src/config.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,91 @@
use anyhow::{Context, Result, ensure};
use serde::Deserialize;
use std::{collections::HashMap, fs, path::Path};

#[derive(Debug, Deserialize)]
pub struct AgentConfig {
pub max_len: usize,
pub head_max_len: usize,
pub head_layers: usize,
pub temperature: Vec<f32>,
pub temperature_by_options: HashMap<String, f32>,
}

#[derive(Debug, Deserialize)]
pub struct EncoderConfig {
pub hidden_size: usize,
pub intermediate_size: usize,
pub num_attention_heads: usize,
pub num_hidden_layers: usize,
pub vocab_size: usize,
pub norm_eps: f32,
pub local_attention: usize,
pub layer_types: Vec<String>,
pub rope_parameters: serde_json::Value,
}

pub struct Config {
pub agent: AgentConfig,
pub encoder: EncoderConfig,
}
impl Config {
pub fn load(dir: &Path) -> Result<Self> {
let read = |name| fs::read(dir.join(name)).with_context(|| format!("read {name}"));
let agent: AgentConfig = serde_json::from_slice(&read("rl_agent_config.json")?)?;
let encoder_bytes = read("encoder/config.json")?;
let raw: serde_json::Value = serde_json::from_slice(&encoder_bytes)?;
ensure!(
raw["model_type"] == "modernbert"
&& raw["hidden_activation"] == "gelu"
&& raw["attention_bias"] == false
&& raw["mlp_bias"] == false
&& raw["norm_bias"] == false,
"unsupported encoder activation, bias or model type"
);
let encoder: EncoderConfig = serde_json::from_slice(&encoder_bytes)?;
ensure!(
agent.max_len == 512 && agent.head_max_len == 192 && agent.head_layers == 2,
"native Laya supports max_len=512, head_max_len=192, head_layers=2"
);
ensure!(
encoder.hidden_size == 1024
&& encoder.intermediate_size == 2624
&& encoder.num_attention_heads == 16
&& encoder.num_hidden_layers == 28
&& encoder.vocab_size == 50368
&& encoder.local_attention == 128
&& encoder.norm_eps == 1e-5,
"unsupported encoder configuration"
);
let expected: Vec<_> = (0..28)
.map(|i| {
if i % 3 == 0 {
"full_attention"
} else {
"sliding_attention"
}
})
.collect();
ensure!(
encoder.layer_types == expected,
"unsupported attention schedule"
);
for (kind, theta) in [("full_attention", 160000.0), ("sliding_attention", 10000.0)] {
let r = &encoder.rope_parameters[kind];
ensure!(
r["rope_type"] == "default" && r["rope_theta"].as_f64() == Some(theta),
"unsupported RoPE configuration"
);
}
ensure!(
agent.temperature.len() == 3
&& agent
.temperature
.iter()
.chain(agent.temperature_by_options.values())
.all(|t| t.is_finite() && *t > 0.0),
"invalid temperatures"
);
Ok(Self { agent, encoder })
}
}
2 changes: 2 additions & 0 deletions src/models/laya/src/lib.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
pub mod config;
pub mod weights;
Loading
Loading