You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Valen evaluates text, images and video against supplied candidates without autoregressive answer generation. Its Qwen3.5 backbone overlaps with System1-Omni's existing backend work, while its learned decision head and multimodal processing require model-specific integration.
The released 2B preview is Sokoban-trained and requires both the Valen checkpoint and its pinned Qwen3.5-2B base. The separate General results are not evaluation results for this preview.
First integration slice
Serve the reference implementation behind the existing Rust frontend. Target text and single-image choice requests first; document unsupported question types and media explicitly. Broader noul, score, video and native acceleration can follow as separate slices after the reference path is usable.
Done when
Pin the reference code, base model and adapter/checkpoint revisions and record their respective licenses.
Document input construction, candidate ordering, decision-head readout, probability normalization, confidence and usage semantics.
Implement model-owned loading, preprocessing, execution and response construction under src/models/valen/.
Preserve all trained components required by the preview: decision head, language LoRA, visual merger and trained visual blocks. Do not load only the head.
Compare real text and image outputs with the pinned reference, using declared probability tolerances and decision-agreement criteria.
Validate a real first inference after readiness and the Rust-frontend request path; report the tested hardware and any checks that could not run.
Add setup, launch and example requests under recipe/valen/, tests under root tests/, and update the support matrix, docs navigation and [Tracking]: Model support status #83.
Follow CONTRIBUTING.md. Coordinate response-field decisions with #61. Native vision work for Cua-S1 (#56, #59, #63, #64) is related infrastructure, not evidence that Valen already runs natively.
Community help wanted
Help with checkpoint/reference investigation, the worker adapter, CPU contract tests, GPU parity checks, or an independently reproduced recipe. Contributors familiar with Valen or Qwen3.5 vision are especially welcome.
Please comment with the slice you would like to take and any relevant environment. The first goal is a faithful, runnable reference service; performance claims come with separate measured evidence.
Model and motivation
Valen evaluates text, images and video against supplied candidates without autoregressive answer generation. Its Qwen3.5 backbone overlaps with System1-Omni's existing backend work, while its learned decision head and multimodal processing require model-specific integration.
The released 2B preview is Sokoban-trained and requires both the Valen checkpoint and its pinned Qwen3.5-2B base. The separate General results are not evaluation results for this preview.
First integration slice
Serve the reference implementation behind the existing Rust frontend. Target text and single-image
choicerequests first; document unsupported question types and media explicitly. Broadernoul,score, video and native acceleration can follow as separate slices after the reference path is usable.Done when
src/models/valen/.recipe/valen/, tests under roottests/, and update the support matrix, docs navigation and [Tracking]: Model support status #83.Follow CONTRIBUTING.md. Coordinate response-field decisions with #61. Native vision work for Cua-S1 (#56, #59, #63, #64) is related infrastructure, not evidence that Valen already runs natively.
Community help wanted
Help with checkpoint/reference investigation, the worker adapter, CPU contract tests, GPU parity checks, or an independently reproduced recipe. Contributors familiar with Valen or Qwen3.5 vision are especially welcome.
Please comment with the slice you would like to take and any relevant environment. The first goal is a faithful, runnable reference service; performance claims come with separate measured evidence.