You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
System1-Omni plans LAYA as its first model (#3, src/models/laya/). Since mid-September, several other open System 1 decision models have appeared on Hugging Face. They answer the same typed questions that /v1/systemone serves (#1), choice, score and noul, but they execute in different ways. Each model engine runs its own execution path. A second model shaped differently from LAYA is the first test of the frontend contract and the model/backend split.
This issue proposes a second model, lists the candidates, and asks the design questions that a second model raises. #3 stays focused on LAYA on Apple Silicon.
Architecture: a frozen Qwen3-8B encoder with last-token pooling, followed by two small projection heads (state and action) trained with a bidirectional InfoNCE loss. A candidate's logit is exp(logit_scale) * cos(state_head(s), action_head(c)), and each answer is the softmax over the question's candidates. A noul question has two candidates (true, false), and score is the expected level index.
Serving shape: the encoder runs as a separate vllm serve Qwen/Qwen3-8B --runner pooling process that exposes /v1/embeddings. clm-serve loads the heads and serves GET /health, POST /v1/systemone, POST /v1/rank and GET /v1/models. The /v1/systemone and /health routes already match the backend contract in [Feature] Add a Rust frontend for omni-modal, prefill-only Jev models #1.
Cross-request state: the encoder embeds states and candidates separately. Candidate vectors can therefore be reused across requests. src/clm/cache.py claims one fixed device arena at startup (2% of device memory by default, CLM_ACTION_CACHE), splits it into pools by vector width, and uses LRU eviction. Keys are namespaced by head generation.
Reported results (from the authors, not reproduced here): zero-shot results on par with Jev on computer use, games and tool calling; up to 9× lower latency than Jev; 13× faster than Jev at about 1,000 candidates when candidate embeddings are cached.
Roadmap: the authors announce a multimodal CLM-35B for early October.
CLM differs from LAYA in three areas the README assigns to model engines: state that persists across requests, the placement of the encoder process, and memory use on Apple Silicon. Implementing CLM second shows which LAYA engine code both models use before any of it moves into shared modules.
Candidate models
Grouped by execution shape. Downloads are Hugging Face's 30-day counts on 2026-09-27.
Extend the README's Supported models table with a row per candidate, giving its execution family and status (Planned, Candidate, Implemented). Each new model gets its own [New Model] issue.
Questions for discussion
Cache ownership. CLM's speedup depends on its candidate-vector arena. The embedder keeps a second host-side cache (cache_size=200_000 texts). Should cross-request caches stay private to each model engine, with a shared rule only for budgets and /health reporting? Or should the repository define a common cache contract now?
Encoder placement. Stage two can keep the encoder as a separate pooling worker, as the reference does, or run it in-process with the heads. A separate worker matches the milestone in [Feature] Add a Rust frontend for omni-modal, prefill-only Jev models #1 and reuses vLLM on CUDA. It adds one local HTTP hop for each cache miss. As far as we know, official vLLM builds on macOS run on CPU only.
Apple Silicon. Qwen3-8B's bf16 weights take about 16 GB (8.2B parameters × 2 bytes). A 16 GB Mac cannot hold them alongside the OS. A Metal path likely needs 8-bit or 4-bit encoder weights. Quantized embeddings need a parity check, because the heads were trained on bf16 embeddings. The reference code also never selects MPS (default_device() returns cuda or cpu). On non-CUDA devices the arena assumes 8 GiB of total memory.
Probability semantics. CLM's probabilities are relative to the candidate set in the request, and score is an expectation over levels. Do these match the semantics the frontend promises for LAYA closely enough to share output-parity tolerances, or does each model declare its own?
Validation for the CLM follow-up
The same structure as #3, with the reference set to CLM's own clm-serve plus vLLM stack on the same hardware and inputs:
Decision agreement and probability differences on a fixed set that covers choice, score and noul, input lengths and option counts, with tolerances declared before comparison.
Warm latency p50/p95 with a cold cache and with a warm cache, reported separately, with and without the frontend.
Load, warmup and readiness time, arena size, and peak memory at a declared concurrency.
On Apple Silicon: the encoder dtype and quantization, plus the device each stage actually ran on.
Acceptance criteria
The README lists candidate models with their execution family and status.
Motivation
System1-Omni plans LAYA as its first model (#3,
src/models/laya/). Since mid-September, several other open System 1 decision models have appeared on Hugging Face. They answer the same typed questions that/v1/systemoneserves (#1),choice,scoreandnoul, but they execute in different ways. Each model engine runs its own execution path. A second model shaped differently from LAYA is the first test of the frontend contract and the model/backend split.This issue proposes a second model, lists the candidates, and asks the design questions that a second model raises. #3 stays focused on LAYA on Apple Silicon.
Highlight: CLM (Contrastive Language Model)
Contrastive-LM/CLM-v0.1-8Bwas released on 2026-09-21 under Apache-2.0, with code at Contrastive-LM/CLM. The notes below come from readingsrc/clm/atbb42c6c.exp(logit_scale) * cos(state_head(s), action_head(c)), and each answer is the softmax over the question's candidates. Anoulquestion has two candidates (true,false), andscoreis the expected level index.vllm serve Qwen/Qwen3-8B --runner poolingprocess that exposes/v1/embeddings.clm-serveloads the heads and servesGET /health,POST /v1/systemone,POST /v1/rankandGET /v1/models. The/v1/systemoneand/healthroutes already match the backend contract in [Feature] Add a Rust frontend for omni-modal, prefill-only Jev models #1.src/clm/cache.pyclaims one fixed device arena at startup (2% of device memory by default,CLM_ACTION_CACHE), splits it into pools by vector width, and uses LRU eviction. Keys are namespaced by head generation.CLM differs from LAYA in three areas the README assigns to model engines: state that persists across requests, the placement of the encoder process, and memory use on Apple Silicon. Implementing CLM second shows which LAYA engine code both models use before any of it moves into shared modules.
Candidate models
Grouped by execution shape. Downloads are Hugging Face's 30-day counts on 2026-09-27.
The agent-side list in system1-agents currently runs Jev (API only), LAYA and Cua-S1 Nano. ThinkFlowLab/system1-agents#14 tracks Cua-S1 4B.
Proposed change
clm-servethroughOMNI_JEV_BACKEND_URL. This stage needs no model code. It tests the [Feature] Add a Rust frontend for omni-modal, prefill-only Jev models #1 acceptance criterion that the public interface has no LAYA-specific assumptions. Stage two addssrc/models/clm/with the heads and cache in the engine and CUDA and Metal paths for the encoder.[New Model]issue.Questions for discussion
cache_size=200_000texts). Should cross-request caches stay private to each model engine, with a shared rule only for budgets and/healthreporting? Or should the repository define a common cache contract now?default_device()returnscudaorcpu). On non-CUDA devices the arena assumes 8 GiB of total memory./v1/systemonecovers CLM's typed questions. CLM's/v1/rankover free-form candidates (best-of-N, tool names) has no equivalent in [Feature] Add a Rust frontend for omni-modal, prefill-only Jev models #1. Should it stay out of scope, or be proposed as a frontend route?scoreis an expectation over levels. Do these match the semantics the frontend promises for LAYA closely enough to share output-parity tolerances, or does each model declare its own?Validation for the CLM follow-up
The same structure as #3, with the reference set to CLM's own
clm-serveplus vLLM stack on the same hardware and inputs:choice,scoreandnoul, input lengths and option counts, with tolerances declared before comparison.Acceptance criteria
choice,scoreandnoulfrom an unmodifiedclm-serve, and the responses match direct backend calls.[New Model]issue for CLM records the encoder placement, cache design and Metal memory plan.Related: #1 (frontend), #3 (LAYA on Apple Silicon), #6 (layout). Agent side: ThinkFlowLab/system1-agents#20 (LAYA), ThinkFlowLab/system1-agents#24 (CLM), ThinkFlowLab/system1-agents#25 (Julia-1).