diff --git a/Cargo.lock b/Cargo.lock index 4b9adda..715fbd3 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -949,10 +949,9 @@ dependencies = [ "anyhow", "axum", "half", - "libloading", "memmap2", + "omni-qwen3-5-native", "safetensors 0.8.0", - "serde", "serde_json", "tokenizers", "tokio", @@ -982,6 +981,31 @@ dependencies = [ "tempfile", ] +[[package]] +name = "omni-open-jev-native" +version = "0.1.0" +dependencies = [ + "anyhow", + "axum", + "omni-qwen3-5-native", + "serde_json", + "tokenizers", + "tokio", +] + +[[package]] +name = "omni-qwen3-5-native" +version = "0.1.0" +dependencies = [ + "anyhow", + "half", + "libloading", + "memmap2", + "safetensors 0.8.0", + "serde", + "serde_json", +] + [[package]] name = "once_cell" version = "1.21.4" diff --git a/Cargo.toml b/Cargo.toml index 661c83e..e7f3326 100644 --- a/Cargo.toml +++ b/Cargo.toml @@ -1,3 +1,3 @@ [workspace] -members = ["src/frontend", "src/models/cua_s1/native", "src/models/laya"] +members = ["src/frontend", "src/models/cua_s1/native", "src/models/qwen3_5/native", "src/models/open_jev/native", "src/models/laya"] resolver = "3" diff --git a/README.md b/README.md index dfe39f9..cbf99e6 100644 --- a/README.md +++ b/README.md @@ -27,9 +27,17 @@ System1-Omni models, designed around a Rust frontend, model-owned execution, and high-performance CUDA and Metal backends. The Rust frontend forwards requests to a separately running model worker. The -Cua-S1 4B 0.2 `text` adapter has a native worker with CUDA kernels in this -repository; other in-repository model engines and GPU backends are not -implemented yet. +Cua-S1 4B 0.2 `text` adapter and Open-Jev-27B-v1.1 have native workers using +shared CUDA kernels in this repository. + +## News + +- **2026-10-03:** Added [Open-Jev-27B-v1.1](recipe/open_jev/native.md) + support through a native Rust/CUDA worker: **7.47× faster than raw HF Transformers** + by mean warm HTTP latency, **362.21→48.50 ms** on one H200. Measured over + 74 single-candidate JevBench `noul` requests per pass, with two measured passes + per backend (BF16, concurrency 1). See the + [HF Transformers baseline, results and OpenJev-Fast comparison](recipe/open_jev/validation.md). ## Features @@ -38,9 +46,9 @@ implemented yet. model workers. - **Model-owned execution.** Each model owns its preprocessing, batching, state, execution, and kernel selection; shared utilities stay minimal. -- **Native CUDA worker.** The Cua-S1 4B 0.2 `text` adapter runs as a native - worker with CUDA kernels, or as a Python worker that serves as the - correctness reference. +- **Native CUDA workers.** The Cua-S1 4B 0.2 `text` adapter and + Open-Jev-27B-v1.1 run as native workers with shared CUDA kernels. Cua-S1 + also has a Python worker that serves as the correctness reference. - **LAYA text serving.** LAYA runs as an external Python worker for text requests, with an in-repository CPU checkpoint reader. - **CUDA and Metal backends.** High-performance GPU operations for NVIDIA @@ -80,9 +88,10 @@ repository root. | [`recipe/`](recipe/) | Model setup instructions, launch commands, configuration examples, and example requests. | | [`docs/`](docs/) | Project documentation and architecture assets. | -The frontend, Cua-S1 native worker and Laya checkpoint reader are Cargo -workspace members. The other model and backend directories currently document -planned work; they do not prescribe process boundaries. +The frontend, both native workers, their shared Qwen3.5/3.8 prefill +implementation and the Laya checkpoint reader are Cargo workspace members. +The other model and backend directories currently document planned work; +they do not prescribe process boundaries. ## Getting Started @@ -104,12 +113,14 @@ for a CPU text worker and response checks, or the Cua-S1 recipes for the LAYA can run as an external Python worker for text requests; its in-repository model engine is still planned. The Cua-S1 4B 0.2 `text` adapter -runs as a Python worker or as a native worker on CUDA: +runs as a Python worker or as a native worker on CUDA. Open-Jev-27B-v1.1 +runs as a native Rust/CUDA worker: | Model | Status | | --- | --- | | LAYA | [External worker](recipe/laya/README.md); [Python worker on Apple Silicon (MPS) and CPU](recipe/laya/apple-silicon.md); [CPU checkpoint reader](src/models/laya/README.md); model execution planned | | Cua-S1 4B 0.2 (`text` adapter) | [Python worker](recipe/cua_s1/text.md); [native worker](recipe/cua_s1/native.md), CUDA, run on sm_89 | +| Open-Jev-27B-v1.1 | [Native Rust/CUDA worker](recipe/open_jev/native.md); eager independent text candidates; [H200 validation](recipe/open_jev/validation.md) | [Supported models and hardware](docs/supported-models.md) lists the devices and where each worker has been run. @@ -117,15 +128,16 @@ and where each worker has been run. ## Benchmarks See the [GPU serving benchmark](benchmarks/README.md) for request replay, -output-fidelity checks, and the CUDA comparison protocol. GPU performance -measurements are pending. +output-fidelity checks, and the CUDA comparison protocol. The +[Open-Jev H200 results](recipe/open_jev/validation.md) cover 74 single-candidate +requests and a matched comparison with raw HF Transformers and OpenJev-Fast. ## Roadmap -The current focus is the Cua-S1 native CUDA worker and the serving benchmark -harness. Planned work includes the in-repository LAYA model engine, additional -model engines and GPU backends including Metal, and per-model performance -measurements as implementations are added and validated. +The current focus is the native Cua-S1 and Open-Jev CUDA workers and the serving +benchmark harness. Planned work includes the in-repository LAYA model engine, +additional model engines and GPU backends including Metal, and per-model +performance measurements as implementations are added and validated. diff --git a/recipe/README.md b/recipe/README.md index 6215704..ebdd603 100644 --- a/recipe/README.md +++ b/recipe/README.md @@ -8,6 +8,8 @@ the worker and connect the Rust frontend. - [Cua-S1 4B 0.2 native text worker](cua_s1/native.md): build the CUDA library and the Rust worker, export the merged weights and start the worker. +- [Open-Jev-27B-v1.1 native text worker](open_jev/native.md): export the merged + text backbone and trained decision head, then serve with Rust and CUDA. Recipes contain setup, launch commands and examples. Reusable implementation code belongs under `src/`. diff --git a/recipe/cua_s1/native.md b/recipe/cua_s1/native.md index 018feff..0e35200 100644 --- a/recipe/cua_s1/native.md +++ b/recipe/cua_s1/native.md @@ -28,7 +28,7 @@ each exact prompt length warms the GEMM plans and captures the forward pass; later requests replay it with freshly uploaded token ids. At most eight lengths are cached. Growing the scratch allocation clears the captures before freeing their buffers. Capture adds first-use latency; leave the variable unset to use -the eager control. Rebuild both the worker and CUDA library together (ABI 3). +the eager control. Rebuild both the worker and CUDA library together (ABI 4). If capture fails, the worker returns the completed eager result and disables Graph capture/replay for its remaining lifetime, logging the failure to stderr. @@ -39,5 +39,5 @@ The request tests need no GPU; the kernel tests compare attention and the chunke ```sh cargo test -p omni-cua-s1-native CUA_S1_CUDA_LIB=$PWD/target/release/libqwen3_5_cuda.so \ - cargo test --release -p omni-cua-s1-native --test kernels -- --ignored + cargo test --release -p omni-qwen3-5-native --test kernels -- --ignored ``` diff --git a/recipe/open_jev/example-request.json b/recipe/open_jev/example-request.json new file mode 100644 index 0000000..b1bef5f --- /dev/null +++ b/recipe/open_jev/example-request.json @@ -0,0 +1,27 @@ +{ + "state": "A customer says: I was charged twice for one order. The service is working normally. There is no sign of unauthorized access.", + "questions": { + "route": { + "type": "choice", + "instructions": "Which team should handle this issue?", + "criteria": { + "billing": "Problems with charges, invoices, refunds or payments.", + "security": "Unauthorized access or account compromise.", + "technical": "Service unavailable or a software malfunction." + } + }, + "refund_review": { + "type": "noul", + "instructions": "Does this message describe a duplicate charge?" + }, + "urgency": { + "type": "score", + "instructions": "Assess the urgency using only the given evidence.", + "criteria": [ + "Routine: no service disruption or active security compromise is reported.", + "Urgent: an ongoing service disruption is reported.", + "Critical: active unauthorized access is reported." + ] + } + } +} diff --git a/recipe/open_jev/export_merged.py b/recipe/open_jev/export_merged.py new file mode 100644 index 0000000..d64cb64 --- /dev/null +++ b/recipe/open_jev/export_merged.py @@ -0,0 +1,70 @@ +"""Export the pinned Open-Jev-27B-v1.1 text backbone and scalar head on CPU. + +Use the reference environment documented in native.md. This preparation step +needs about 110 GB of host RAM and 52 GB of output storage, without a GPU. +""" + +import argparse +import json +from pathlib import Path + +import torch +from peft import PeftModel +from transformers import AutoModelForImageTextToText, AutoTokenizer + +BASE_REVISION = "1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0" +CHECKPOINT_REVISION = "28cf73067d5b337860bbef3c85b8b82ba8730956" + + +def main(): + parser = argparse.ArgumentParser(description=__doc__.splitlines()[0]) + parser.add_argument("--base", required=True, type=Path) + parser.add_argument("--checkpoint", required=True, type=Path) + parser.add_argument("--out", required=True, type=Path) + parser.add_argument("--max-length", type=int, default=4096) + args = parser.parse_args() + config = json.loads((args.checkpoint / "model.json").read_text()) + if config["model_id"] != "Qwen/Qwen3.8-27B" or config["revision"] != BASE_REVISION: + raise ValueError("expected Open-Jev-27B-v1.1's pinned base") + if args.out.exists(): + raise ValueError("output already exists; choose a new export directory") + if not 1 <= args.max_length <= 16384: + raise ValueError("max length must be within 1..=16384") + temperature = json.loads((args.checkpoint / "temperature.json").read_text())["temperature"] + tokenizer = AutoTokenizer.from_pretrained(args.base, local_files_only=True) + marker = "\x00OMNI_OPEN_JEV\x00" + chat = tokenizer.apply_chat_template( + [{"role": "user", "content": marker}], tokenize=False, + add_generation_prompt=True, enable_thinking=False, + ) + if chat.count(marker) != 1: + raise ValueError("expected a single-user text chat template") + prefix, suffix = chat.split(marker) + head = torch.load(args.checkpoint / "head.pt", map_location="cpu", weights_only=True) + if head["weight"].shape != (1, 5120) or head["bias"].shape != (1,): + raise ValueError("expected a 5120-wide trained scalar head") + if not all(torch.isfinite(v).all() for v in head.values()): + raise ValueError("non-finite scalar head") + full = AutoModelForImageTextToText.from_pretrained( + args.base, torch_dtype=torch.bfloat16, device_map={"": "cpu"}, + attn_implementation="sdpa", local_files_only=True, + ) + backbone = full.model.language_model + del full + backbone = PeftModel.from_pretrained(backbone, args.checkpoint / "adapter") + backbone = backbone.merge_and_unload(safe_merge=True) + backbone.save_pretrained(args.out, max_shard_size="5GB") + tokenizer.save_pretrained(args.out) + # Written last: the native worker refuses incomplete exports or plain base weights. + (args.out / "open_jev_export.json").write_text(json.dumps({ + "format": "open-jev-text-merged/1", + "model_id": config["model_id"], "base_revision": BASE_REVISION, + "checkpoint_revision": CHECKPOINT_REVISION, "temperature": temperature, + "max_length": args.max_length, "chat_prefix": prefix, "chat_suffix": suffix, + "head_weight": head["weight"].float().reshape(-1).tolist(), + "head_bias": head["bias"].float().item(), + }, allow_nan=False) + "\n") + + +if __name__ == "__main__": + main() diff --git a/recipe/open_jev/native.md b/recipe/open_jev/native.md new file mode 100644 index 0000000..7537d48 --- /dev/null +++ b/recipe/open_jev/native.md @@ -0,0 +1,131 @@ +# Open-Jev-27B-v1.1 native text worker + +The worker owns request compilation, tokenization, candidate scoring and typed +responses in Rust. It uses the native CUDA prefill implementation introduced in +[PR #19](https://github.com/ThinkFlowLab/system1-omni/pull/19), shared with Cua-S1 +under [`src/models/qwen3_5/native/`](../../src/models/qwen3_5/native/). +Python is required only to prepare the merged checkpoint. + +It supports `choice` (1–255 candidates), `score` (2–10 levels), and `noul` +(yes/no). Each candidate has an independent prompt; the last hidden state goes +through Open-Jev's trained FP32 scalar head. Noul uses logits `[0, score]`. +The saved calibration temperature is applied before normalizing each complete +question. There is no autoregressive generation. Structured state and descriptions +use Open-Jev's sorted JSON rendering; question and candidate order is preserved. + +## Prepare the checkpoint + +Use the reference dependencies from +[Open-Jev @ 3308a15](https://github.com/Zefan-Cai/Open-Jev/tree/3308a15ccd7eea1df7a37d6ddc39b023b801ba16): +PyTorch 2.8 or newer, Transformers 5.10.2, PEFT 0.19.1, Accelerate 1.13.0, +and safetensors. An optional `kernels` installation must be compatible with that +Transformers release. Run these commands from the repository root: + +```sh +hf download Qwen/Qwen3.8-27B \ + --revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 \ + --local-dir weights/Qwen3.8-27B +hf download ZefanCai/Open-Jev-27B-v1.1 \ + --revision 28cf73067d5b337860bbef3c85b8b82ba8730956 \ + --include 'package/checkpoint/*' --local-dir weights/Open-Jev-27B-v1.1 +CUDA_VISIBLE_DEVICES='' python recipe/open_jev/export_merged.py \ + --base weights/Qwen3.8-27B \ + --checkpoint weights/Open-Jev-27B-v1.1/package/checkpoint \ + --out weights/open-jev-27b-merged +``` + +CPU export needs roughly 110 GB of RAM and 52 GB of output storage. It merges +LoRA in BF16 and saves the trained head, temperature, and single-user chat +template in `open_jev_export.json`. The worker refuses a plain base checkpoint +or an incomplete export. The saved limit defaults to 4096 tokens per candidate; +`--max-length` may raise it to 16384. Oversize prompts fail before inference. + +## Build and serve + +The CUDA kernels require compute capability 8.0 or newer. The current build +target below is Ada (`89`); pass your GPU's compute capability explicitly. +The CUDA shared library and both Rust workers must be rebuilt together because +the gated-attention entry point updates the library ABI to version 4 alongside +the shared CUDA Graph entry points. + +```sh +src/backends/cuda/qwen3_5/build.sh target/release 89 +cargo build --release --locked -p omni-open-jev-native -p omni-jev +OPEN_JEV_MODEL=weights/open-jev-27b-merged \ + target/release/omni-open-jev-native +``` + +`OPEN_JEV_HOST` and `OPEN_JEV_PORT` default to `127.0.0.1` and `8000`. +`OPEN_JEV_CUDA_LIB` overrides the default library next to the executable. +The worker loads all text weights onto visible CUDA device 0, performs a real +warmup inference, then exposes `/health` and `/v1/systemone`. +Use a reservation before any GPU command on hosts with a GPU scheduler. + +In another terminal, start the existing Rust frontend: + +```sh +OMNI_JEV_BIND=127.0.0.1:8080 OMNI_JEV_BACKEND_URL=http://127.0.0.1:8000 \ + target/release/omni-jev +curl http://127.0.0.1:8080/v1/systemone \ + -H 'Content-Type: application/json' --data-binary @recipe/open_jev/example-request.json +``` + +The worker accepts the model's base name `Qwen/Qwen3.8-27B`, `open-jev`, +`jev-latest`, and `open-jev-27b-v1.1`; the response model is the base name, +matching Open-Jev. Error wording and metadata differ from the reference service. +Requests are bounded to 4 MiB, 4096 questions, and 65536 candidate sequences. + +## Validation and optimization scope + +Tests and fixtures live in the repository-level `tests/` tree: Open-Jev's typed +contract and tokenizer cases are in +[`tests/open_jev/`](../../tests/open_jev/), and shared Qwen JSON, configuration +and CUDA reference tests are in [`tests/qwen3_5/`](../../tests/qwen3_5/). +The default suites below run on CPU without downloading model weights: + +```sh +cargo test --locked -p omni-open-jev-native -p omni-qwen3-5-native +cargo test --locked -p omni-jev --test frontend +``` + +The frontend mock-worker API coverage is tracked in +[issue #46](https://github.com/ThinkFlowLab/system1-omni/issues/46) and +[PR #58](https://github.com/ThinkFlowLab/system1-omni/pull/58). Checkpoint tokenizer +and CUDA kernel tests are opt-in; the latter require a GPU reservation: + +```sh +# Inside a GPU reservation, after building the library: +CUA_S1_CUDA_LIB=$PWD/target/release/libqwen3_5_cuda.so \ + cargo test --release --locked -p omni-qwen3-5-native --test kernels -- --ignored +``` + +CPU golden fixtures come from Open-Jev's request compiler and response formatter +at the revision above. Kernel tests compare attention and Gated DeltaNet with +float64 references and require exact BF16 equality between fused attention gating +and a separate gate pass. Residual RMSNorm and packed SiLU are checked against +rounded references, including odd widths and unaligned pointers. The shared +kernel tests retain PR #19's +`CUA_S1_CUDA_LIB` environment variable. + +The worker reuses PR #19's fused norm, activation, QK/RoPE and chunked Gated +DeltaNet operations. Attention's sigmoid gate is fused into its output epilogue, +preserving both BF16 rounding points and removing one launch and one output +read/write pass per full-attention layer (16 layers for this model). Residual +RMSNorm keeps thread values in registers at widths 2560/5120. MLP SiLU uses +16-byte BF16 loads/stores when width, stride and pointers permit it, retaining +both BF16 rounding points; other layouts use the scalar path. + +This recipe leaves `CUA_S1_GRAPH` unset and runs one eager forward pass per +candidate. Set `CUA_S1_GRAPH=1` on the worker to enable CUDA Graph replay. The +shared backend retains at most 64 graphs, keyed by exact candidate token length; +growing the scratch buffer clears them. Capturing a new length first runs an +eager forward to initialize its plans, then captures and replays the forward. +This adds cost for new lengths, so graph mode remains opt-in. Warm replay is +validated on the 74-case H200 workload: its mean HTTP latency is 2.03% below +eager execution after all workload lengths are warmed. Tokenization, transfers +and the CPU scalar head remain outside the graph. Prefix sharing, GEMM autotuning, +quantization and multimodal inference are not implemented. The +[H200 validation](validation.md) reports full-checkpoint results for 74 +single-candidate requests, including probability differences and timing +variability. It does not establish general accuracy parity or a speedup over +OpenJev-Fast; the author's B300 results use different hardware and workloads. diff --git a/recipe/open_jev/validation.md b/recipe/open_jev/validation.md new file mode 100644 index 0000000..76232a8 --- /dev/null +++ b/recipe/open_jev/validation.md @@ -0,0 +1,226 @@ +# Open-Jev H200 validation + +## Raw HF Transformers comparison, 2026-10-03 + +The native Rust/CUDA worker delivers a **7.47× speedup over raw HF Transformers** +by mean warm HTTP latency: **362.21→48.50 ms (86.61% lower)** on one H200. +This comparison covers 74 real JevBench `noul` requests, one candidate each, +80–3399 tokens, BF16, max length 16384 and concurrency 1. Each backend reuses +one server for an excluded feasibility pass and two measured passes: +148 measured requests per backend. + +| Configuration | Overall mean (ms) | Mean, pass 1 / pass 2 (ms) | P50, pass 1 / pass 2 (ms) | P95, pass 1 / pass 2 (ms) | Correct / 74 | +| --- | ---: | ---: | ---: | ---: | ---: | +| Raw HF Transformers (unmerged LoRA; PyTorch fallback) | 362.209 | 362.238 / 362.180 | 324.822 / 321.321 | 651.999 / 652.911 | 64 | +| Native Rust/CUDA (cached RMSNorm + packed SiLU) | 48.503 | 48.471 / 48.535 | 25.797 / 24.895 | 215.696 / 218.882 | 64 | +| Original OpenJev-Fast | 50.936 | 51.097 / 50.775 | 24.651 / 24.622 | 248.121 / 251.193 | 63 | + +Native and HF agree on all 74 thresholded decisions and both score **64/74**; +their maximum probability difference is **0.020423**. Fast scores 63/74 with +one different decision (`hard-opus-a-temporal_numeric-09`); maximum native/Fast +probability difference is 0.034353. Each backend's probabilities are exactly +unchanged across its feasibility and measured passes. These counts do not +establish statistical accuracy superiority or full numerical parity. + +**Raw HF baseline:** the original Open-Jev server and `DecisionModel`, with +160 unmerged PEFT LoRA modules, the trained scalar head and saved temperature. +Full attention uses stock SDPA; runtime assertions verify all 48 linear-attention +layers use `torch_chunk_gated_delta_rule`, stock convolution and +`Qwen3_5RMSNormGated`. Because the shared environment contains optional kernels, +the baseline makes Transformers' `is_flash_linear_attention_available` and +`is_causal_conv1d_available` checks return false before importing model classes. +It uses no custom Fast model, `torch.compile`, CUDA Graph replay or prefix cache. +Native uses merged LoRA and eager execution (`CUA_S1_GRAPH=0`); original Fast +retains its custom kernels and CUDA Graph stack, reporting 30 retained graphs. +This comparison changes the complete backend; it does not isolate one optimization. + +Timing includes localhost HTTP through the same frozen Rust frontend, +tokenization, worker execution and UTF8 response decoding. Client body +serialization and response JSON parsing are excluded. Downloads, preparation, +process-to-readiness, warmup, first inference after readiness and the complete +feasibility pass are excluded. No Nsight launcher or trace collection was used. +P50 is the median; P95 uses JevBench's `sorted[int(0.95*N)-1]` rank. +The observed mean pass ranges are disjoint; two-pass ranges are not confidence +intervals. Native's mean is 4.78% below Fast in this run, while Fast has a lower +median. Multi-candidate prefix sharing and broader JevBench coverage remain +unmeasured; native CUDA Graph replay is reported separately below. The author's +17.3 ms B300 result uses different hardware and workload. + +### Frozen controls and reproduction + +- Exact H200 GPU 2, UUID `GPU-cbf66259-f4ab-0ede-1811-82037dde5924`, NUMA 0, + CPUs 0–15. The archived driver name is `NVIDIA L20X`, with 143771 MiB and + SM90 / 132 SMs. The same device, affinity, requests, model and prepared + environment are used for all three backends; no shared caches are dropped. +- Frozen native worker/frontend: `202c0e163f868334a99d88407056ebe61dbb2dce`; + accepted packed-SiLU CUDA library SHA256: + `e033315d4c67127809e62f41d991a149ac8ee0ce81827997af063815fd00d435`. + PR source at measurement: `ad1cb818d2ea001bbd8c84643e1a7241d0fc26c8`. +- [Original Open-Jev](https://github.com/Zefan-Cai/Open-Jev/tree/3308a15ccd7eea1df7a37d6ddc39b023b801ba16) + at `3308a15ccd7eea1df7a37d6ddc39b023b801ba16`. Fast, JevBench, request JSON, + base, adapter and temperature use the same pins recorded in the October 2 + controls below and the [native recipe](native.md). +- Actual runtime: Torch 2.13.0+cu130, CUDA 13.0, Transformers 5.10.2, + PEFT 0.19.1. Reuse prepared weights and SM90 extensions offline. +- Use `gpu run --gpu-ids 2 --wait 10m --timeout 45m --note