Skip to content

[RFC] Support emerging System 1 models, with CLM as the second model after LAYA #9

Description

@Yunaik

Motivation

System1-Omni plans LAYA as its first model (#3, src/models/laya/). Since mid-September, several other open System 1 decision models have appeared on Hugging Face. They answer the same typed questions that /v1/systemone serves (#1), choice, score and noul, but they execute in different ways. Each model engine runs its own execution path. A second model shaped differently from LAYA is the first test of the frontend contract and the model/backend split.

This issue proposes a second model, lists the candidates, and asks the design questions that a second model raises. #3 stays focused on LAYA on Apple Silicon.

Highlight: CLM (Contrastive Language Model)

Contrastive-LM/CLM-v0.1-8B was released on 2026-09-21 under Apache-2.0, with code at Contrastive-LM/CLM. The notes below come from reading src/clm/ at bb42c6c.

  • Architecture: a frozen Qwen3-8B encoder with last-token pooling, followed by two small projection heads (state and action) trained with a bidirectional InfoNCE loss. A candidate's logit is exp(logit_scale) * cos(state_head(s), action_head(c)), and each answer is the softmax over the question's candidates. A noul question has two candidates (true, false), and score is the expected level index.
  • Serving shape: the encoder runs as a separate vllm serve Qwen/Qwen3-8B --runner pooling process that exposes /v1/embeddings. clm-serve loads the heads and serves GET /health, POST /v1/systemone, POST /v1/rank and GET /v1/models. The /v1/systemone and /health routes already match the backend contract in [Feature] Add a Rust frontend for omni-modal, prefill-only Jev models #1.
  • Cross-request state: the encoder embeds states and candidates separately. Candidate vectors can therefore be reused across requests. src/clm/cache.py claims one fixed device arena at startup (2% of device memory by default, CLM_ACTION_CACHE), splits it into pools by vector width, and uses LRU eviction. Keys are namespaced by head generation.
  • Reported results (from the authors, not reproduced here): zero-shot results on par with Jev on computer use, games and tool calling; up to 9× lower latency than Jev; 13× faster than Jev at about 1,000 candidates when candidate embeddings are cached.
  • Roadmap: the authors announce a multimodal CLM-35B for early October.

CLM differs from LAYA in three areas the README assigns to model engines: state that persists across requests, the placement of the encoder process, and memory use on Apple Silicon. Implementing CLM second shows which LAYA engine code both models use before any of it moves into shared modules.

Candidate models

Grouped by execution shape. Downloads are Hugging Face's 30-day counts on 2026-09-27.

Family Models Execution path Work beyond LAYA
Encoder, single pass LAYA; AFM-DE (LAYA fine-tune, same checkpoint layout); OpenJev (Verdict) (151M, 20.6k downloads); Julia-1 (144M mmBERT-small, 2 to 20 options, 8,192 tokens) ModernBERT forward pass over state and options AFM-DE is expected to load in the LAYA engine. OpenJev adds an abstain slot and ships ONNX.
Contrastive dual encoder CLM-v0.1-8B Qwen3-8B pooled embeddings, two projection heads, cosine softmax Bounded cross-request vector cache, 8B encoder, separate encoder process
Decoder, prefill only Kev 0.5B to 9B (4B: 8.9k downloads); Cua-S1 4B Qwen3.5 base with LoRA and a classification head, one prefill Decoder prefill kernels, LoRA application, per-model head
Constrained branching LFM2.5-350M-RLCD, LFM2.5-2.6B-RLCD Prefill once, then branch attention and convolution state per candidate State branching for a hybrid convolution/attention model. The authors describe it as experimental and uncalibrated.

The agent-side list in system1-agents currently runs Jev (API only), LAYA and Cua-S1 Nano. ThinkFlowLab/system1-agents#14 tracks Cua-S1 4B.

Proposed change

  1. Keep LAYA as the first model. [Help wanted] Accelerated Laya serving on Apple Silicon #3 continues unchanged.
  2. Pick CLM as the second model.
  3. Integrate CLM in two stages. Stage one points the [Feature] Add a Rust frontend for omni-modal, prefill-only Jev models #1 frontend at an unmodified clm-serve through OMNI_JEV_BACKEND_URL. This stage needs no model code. It tests the [Feature] Add a Rust frontend for omni-modal, prefill-only Jev models #1 acceptance criterion that the public interface has no LAYA-specific assumptions. Stage two adds src/models/clm/ with the heads and cache in the engine and CUDA and Metal paths for the encoder.
  4. Add no shared model abstraction yet, following [Discussion] Minimal repository layout for frontend, models, and GPU backends #6. Extract shared code once LAYA and CLM both exist and show what they have in common.
  5. Extend the README's Supported models table with a row per candidate, giving its execution family and status (Planned, Candidate, Implemented). Each new model gets its own [New Model] issue.

Questions for discussion

  1. Cache ownership. CLM's speedup depends on its candidate-vector arena. The embedder keeps a second host-side cache (cache_size=200_000 texts). Should cross-request caches stay private to each model engine, with a shared rule only for budgets and /health reporting? Or should the repository define a common cache contract now?
  2. Encoder placement. Stage two can keep the encoder as a separate pooling worker, as the reference does, or run it in-process with the heads. A separate worker matches the milestone in [Feature] Add a Rust frontend for omni-modal, prefill-only Jev models #1 and reuses vLLM on CUDA. It adds one local HTTP hop for each cache miss. As far as we know, official vLLM builds on macOS run on CPU only.
  3. Apple Silicon. Qwen3-8B's bf16 weights take about 16 GB (8.2B parameters × 2 bytes). A 16 GB Mac cannot hold them alongside the OS. A Metal path likely needs 8-bit or 4-bit encoder weights. Quantized embeddings need a parity check, because the heads were trained on bf16 embeddings. The reference code also never selects MPS (default_device() returns cuda or cpu). On non-CUDA devices the arena assumes 8 GiB of total memory.
  4. API surface. /v1/systemone covers CLM's typed questions. CLM's /v1/rank over free-form candidates (best-of-N, tool names) has no equivalent in [Feature] Add a Rust frontend for omni-modal, prefill-only Jev models #1. Should it stay out of scope, or be proposed as a frontend route?
  5. Probability semantics. CLM's probabilities are relative to the candidate set in the request, and score is an expectation over levels. Do these match the semantics the frontend promises for LAYA closely enough to share output-parity tolerances, or does each model declare its own?

Validation for the CLM follow-up

The same structure as #3, with the reference set to CLM's own clm-serve plus vLLM stack on the same hardware and inputs:

  • Decision agreement and probability differences on a fixed set that covers choice, score and noul, input lengths and option counts, with tolerances declared before comparison.
  • Warm latency p50/p95 with a cold cache and with a warm cache, reported separately, with and without the frontend.
  • Load, warmup and readiness time, arena size, and peak memory at a declared concurrency.
  • On Apple Silicon: the encoder dtype and quantization, plus the device each stage actually ran on.

Acceptance criteria

  • The README lists candidate models with their execution family and status.
  • Stage one works: the [Feature] Add a Rust frontend for omni-modal, prefill-only Jev models #1 frontend serves choice, score and noul from an unmodified clm-serve, and the responses match direct backend calls.
  • A [New Model] issue for CLM records the encoder placement, cache design and Metal memory plan.
  • Questions 1 to 5 have a recorded decision before stage two starts.

Related: #1 (frontend), #3 (LAYA on Apple Silicon), #6 (layout). Agent side: ThinkFlowLab/system1-agents#20 (LAYA), ThinkFlowLab/system1-agents#24 (CLM), ThinkFlowLab/system1-agents#25 (Julia-1).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions