Skip to content

Repository files navigation

System1Bench: Benchmarking Jev-Style Decision Models

Verify benchmark artifacts

System 1 决策模型评测基准 · An independent, reproducible evaluation collection for models that turn a supplied state into typed choice, noul and score decisions.

v0.2 includes our real runs of Laya English, Laya Multilingual, Llama-3.1-8B-Instruct and Qwen3-8B. Jev has not been evaluated. The framework accepts a model adapter for future Jev runs; “Jev-style” describes the interface/task category, not a shared architecture.

System1Bench combines 15 public source datasets/collections into 36 suites and controls. It covers semantic classification, bilingual inference, structured workflow decisions, explicit out-of-scope intent detection, safety, long-context retrieval and actual option-order perturbations. Scores stay separate by source and reference quality. There is no blended “decision intelligence” score.

Measured results — all tasks

Every score below comes from our own local model runs in this repository. No third-party model scores are copied. Jev has not been evaluated.

4 checkpoints × 26,450 decisions = 105,800 measured decisions, including controls. Failures: 0; complete inputs: 105,800/105,800.

Laya uses native decision heads. The Llama and Qwen baselines use zero-shot constrained next-token answer selection; Qwen thinking is disabled. See the exact comparison protocol. These are fixed direct-decision baselines, not best-achievable LLM scores.

Values are accuracy against each source reference. Teacher/synthetic and authored/AI-reviewed rows measure reference agreement, not independently human-verified correctness. We publish all suites without an overall blended score. Full confidence intervals, F1, Brier/ECE, ordinal errors, language/length slices and timings are in the report, CSV and JSON.

Main tasks (28 suites)

Task Decisions / model Laya English Laya Multilingual Llama-3.1-8B-Instruct Qwen3-8B Reference
ag_news 1000 94.30% 92.40% 88.90% 86.80% dataset_provided
emotion 1000 59.10% 53.00% 50.50% 54.80% dataset_provided
banking77 1000 55.30% 51.20% 54.10% 66.80% dataset_provided
boolq 1000 84.60% 77.70% 64.80% 83.50% dataset_provided
boolq_choice 1000 83.60% 77.40% 68.70% 83.30% dataset_provided
sst5 1000 34.60% 29.50% 32.00% 45.70% dataset_provided
sst5_choice 1000 49.60% 35.60% 42.20% 43.90% dataset_provided
xnli_en 1000 86.00% 81.70% 46.20% 78.70% dataset_provided
xnli_zh 1000 61.50% 74.30% 40.50% 69.10% dataset_provided
massive_en 1000 54.10% 42.30% 57.20% 66.30% dataset_provided
massive_zh 1000 30.40% 33.20% 53.50% 62.60% dataset_provided
prompt_injections 116 70.69% 57.76% 62.93% 63.79% dataset_provided
typed_decisions 2000 36.35% 34.90% 51.20% 55.50% synthetic_teacher
jevbench_original 72 70.83% 41.67% 75.00% 83.33% authored_or_AI_reviewed
jevbench_easy 48 95.83% 89.58% 100.00% 100.00% authored_or_AI_reviewed
jevbench_hard 111 29.73% 32.43% 34.23% 46.85% authored_or_AI_reviewed
reflexbench_reflex-public-choice-v1 95 58.95% 46.32% 60.00% 84.21% authored_or_AI_reviewed
jev_laya_triage 501 62.48% 55.69% 75.65% 80.04% synthetic_teacher
jev_laya_moderation 426 67.84% 48.36% 88.97% 77.46% synthetic_teacher
jev_laya_routing 411 64.23% 51.09% 61.31% 83.70% synthetic_teacher
jev_laya_claims 300 90.00% 80.00% 61.67% 98.67% synthetic_teacher
jev_laya_reviews 300 70.67% 42.33% 83.00% 87.00% synthetic_teacher
jev_laya_guard 292 60.62% 33.22% 88.01% 83.22% synthetic_teacher
jev_laya_multilingual 256 57.81% 62.50% 94.53% 98.44% synthetic_teacher
jev_laya_needle 900 50.22% 47.44% 77.89% 87.78% programmatic
clinc150_oos 5500 55.73% 64.76% 55.20% 67.78% dataset_provided
turtlebench 1532 42.62% 41.64% 42.17% 44.19% dataset_provided
aegis2_prompt 1928 49.59% 57.05% 66.80% 72.67% human_prompt_annotation

Same-order repeats and reversed-option controls (8 suites)

Task Decisions / model Laya English Laya Multilingual Llama-3.1-8B-Instruct Qwen3-8B Reference
banking77_repeat 100 49.00% 44.00% 62.00% 72.00% dataset_provided
banking77_reversed 100 58.00% 44.00% 37.00% 54.00% dataset_provided
massive_en_repeat 100 54.00% 50.00% 62.00% 74.00% dataset_provided
massive_en_reversed 100 52.00% 43.00% 60.00% 66.00% dataset_provided
jevbench_original_repeat 36 61.11% 58.33% 80.56% 83.33% authored_or_AI_reviewed
jevbench_original_reversed 36 58.33% 55.56% 66.67% 86.11% authored_or_AI_reviewed
reflexbench_reflex-public-choice-v1_repeat 95 58.95% 46.32% 60.00% 84.21% authored_or_AI_reviewed
reflexbench_reflex-public-choice-v1_reversed 95 58.95% 45.26% 57.89% 82.11% authored_or_AI_reviewed

Option-order agreement

Agreement compares predictions on the same inputs; it is not accuracy. The LLM reversal also reassigns answer codes.

Model / task N Original → repeat Repeat → reversed
Laya English / banking77 100 100.00% 52.00%
Laya English / massive_en 100 100.00% 50.00%
Laya English / jevbench_original 36 100.00% 91.67%
Laya English / reflexbench_reflex-public-choice-v1 95 100.00% 89.47%
Laya Multilingual / banking77 100 100.00% 62.00%
Laya Multilingual / massive_en 100 100.00% 54.00%
Laya Multilingual / jevbench_original 36 100.00% 77.78%
Laya Multilingual / reflexbench_reflex-public-choice-v1 95 100.00% 80.00%
Llama-3.1-8B-Instruct / banking77 100 96.00% 35.00%
Llama-3.1-8B-Instruct / massive_en 100 95.00% 49.00%
Llama-3.1-8B-Instruct / jevbench_original 36 100.00% 77.78%
Llama-3.1-8B-Instruct / reflexbench_reflex-public-choice-v1 95 100.00% 82.11%
Qwen3-8B / banking77 100 99.00% 61.00%
Qwen3-8B / massive_en 100 100.00% 68.00%
Qwen3-8B / jevbench_original 36 100.00% 86.11%
Qwen3-8B / reflexbench_reflex-public-choice-v1 95 100.00% 85.26%

Explicit out-of-scope detection (CLINC150 + OOS)

Model In-scope accuracy OOS precision OOS recall OOS F1
Laya English 49.11% 38.44% 85.50% 0.5304
Laya Multilingual 74.42% 86.59% 21.30% 0.3419
Llama-3.1-8B-Instruct 52.51% 41.91% 67.30% 0.5165
Qwen3-8B 65.93% 57.52% 76.10% 0.6552

Ordinal scoring (SST5, fixed 0–4 scale)

Lower MAE is better; within-one agreement allows an error of one scale point.

Model Argmax MAE ↓ Expected-score MAE ↓ Within one ↑
Laya English 0.9270 0.9144 77.80%
Laya Multilingual 1.3000 1.2946 58.70%
Llama-3.1-8B-Instruct 1.0350 0.9999 71.30%
Qwen3-8B 0.6560 0.6482 90.80%

Controlled local decision speed — all workloads

These are our own new measurements on one A100 80GB PCIe, using the existing four adapters. Each cell shows the median over eight rounds and its 95% bootstrap interval. Encoding and structured-answer construction are included; model loading, network/queue time and validation are excluded. The full performance report includes p95, reference agreement, memory, token counts and limits. See the frozen protocol and raw records.

Workload Model Batch 1 request p50 ms Batch 8 requests/s Batch 32 requests/s
ag_news Laya English 10.38 [10.29, 18.35] 441.20 [436.82, 444.37] 573.63 [570.54, 582.44]
ag_news Laya Multilingual 8.51 [8.46, 8.94] 696.15 [690.33, 703.78] 1037.98 [910.16, 1048.95]
ag_news Llama-3.1-8B-Instruct 30.69 [30.39, 30.78] 47.11 [46.67, 47.61] 47.07 [46.89, 48.16]
ag_news Qwen3-8B 30.96 [30.71, 31.02] 46.46 [45.26, 47.51] 47.25 [47.13, 47.79]
boolq Laya English 10.30 [10.23, 10.32] 257.19 [256.58, 261.75] 249.67 [248.99, 255.97]
boolq Laya Multilingual 8.55 [8.50, 8.71] 448.91 [372.81, 457.76] 456.62 [451.39, 482.04]
boolq Llama-3.1-8B-Instruct 36.86 [36.53, 37.19] 29.18 [29.00, 29.47] 25.18 [25.15, 25.53]
boolq Qwen3-8B 33.51 [33.34, 33.71] 27.81 [27.55, 28.07] 23.78 [23.69, 23.89]
sst5 Laya English 10.27 [10.21, 10.41] 516.73 [512.80, 520.98] 805.83 [765.12, 823.88]
sst5 Laya Multilingual 8.47 [8.38, 8.53] 746.61 [715.84, 752.40] 1305.71 [1286.94, 1318.23]
sst5 Llama-3.1-8B-Instruct 29.20 [29.08, 29.31] 54.06 [53.02, 54.85] 57.68 [57.50, 58.33]
sst5 Qwen3-8B 29.27 [29.12, 29.36] 55.05 [54.60, 56.62] 56.30 [56.12, 57.19]
banking77 Laya English 12.66 [12.55, 12.74] 165.13 [134.06, 168.60] 180.98 [158.88, 185.87]
banking77 Laya Multilingual 9.93 [9.81, 10.09] 265.39 [257.80, 267.94] 295.38 [292.77, 300.42]
banking77 Llama-3.1-8B-Instruct 112.93 [111.15, 113.03] 9.00 [8.99, 9.05] 9.02 [9.00, 9.03]
banking77 Qwen3-8B 119.91 [118.15, 120.11] 8.41 [8.41, 8.47] 8.44 [8.43, 8.46]
massive_en Laya English 11.44 [11.35, 20.04] 225.14 [221.28, 229.00] 260.98 [205.78, 264.71]
massive_en Laya Multilingual 9.50 [9.36, 13.45] 357.15 [292.98, 364.11] 420.73 [414.38, 432.53]
massive_en Llama-3.1-8B-Instruct 91.18 [90.88, 91.29] 12.02 [12.02, 12.06] 12.25 [12.23, 12.26]
massive_en Qwen3-8B 96.01 [95.87, 96.06] 11.35 [11.31, 11.44] 11.60 [11.58, 11.60]
massive_zh Laya English 11.38 [11.34, 11.48] 223.72 [218.93, 224.30] 256.27 [249.63, 258.63]
massive_zh Laya Multilingual 9.48 [9.34, 17.08] 360.41 [248.90, 366.44] 424.18 [417.61, 435.98]
massive_zh Llama-3.1-8B-Instruct 91.15 [90.66, 91.19] 12.02 [12.00, 12.03] 12.31 [12.28, 12.32]
massive_zh Qwen3-8B 95.66 [95.44, 95.81] 11.37 [11.35, 11.39] 11.64 [11.63, 11.67]
clinc150_oos Laya English 21.37 [21.29, 21.68] 58.55 [57.44, 58.76] 61.65 [50.61, 62.48]
clinc150_oos Laya Multilingual 13.60 [13.49, 13.77] 101.92 [100.23, 102.75] 109.08 [106.06, 110.01]
clinc150_oos Llama-3.1-8B-Instruct 201.76 [201.59, 201.94] 4.50 [4.50, 4.51] 4.52 [4.50, 4.53]
clinc150_oos Qwen3-8B 214.74 [214.46, 215.18] 4.17 [4.17, 4.18] 4.21 [4.21, 4.23]
jev_laya_triage Laya English 13.89 [13.83, 19.89] 121.58 [121.06, 122.84] 124.05 [113.78, 125.68]
jev_laya_triage Laya Multilingual 9.30 [9.27, 9.43] 225.77 [197.01, 228.87] 244.28 [241.82, 246.21]
jev_laya_triage Llama-3.1-8B-Instruct 100.02 [99.51, 100.78] 12.58 [12.46, 12.64] 12.61 [12.55, 12.67]
jev_laya_triage Qwen3-8B 102.95 [102.66, 103.92] 12.08 [11.99, 12.27] 12.39 [12.35, 12.46]
needle_100 Laya English 11.96 [11.95, 12.00] 234.66 [233.99, 235.23] 269.75 [267.47, 270.78]
needle_100 Laya Multilingual 9.08 [9.00, 12.76] 404.35 [336.13, 405.09] 488.97 [404.88, 493.09]
needle_100 Llama-3.1-8B-Instruct 69.02 [68.65, 69.14] 20.89 [20.86, 20.94] 22.15 [22.14, 22.17]
needle_100 Qwen3-8B 72.63 [72.45, 73.39] 20.26 [20.23, 20.33] 21.75 [21.72, 21.78]
needle_1000 Laya English 34.79 [34.76, 34.83] 39.37 [38.08, 39.49] 41.13 [41.00, 41.22]
needle_1000 Laya Multilingual 18.30 [18.26, 18.37] 75.38 [75.24, 75.66] 79.72 [79.02, 80.03]
needle_1000 Llama-3.1-8B-Instruct 202.67 [202.39, 202.87] 4.96 [4.95, 4.96] 5.02 [5.00, 5.05]
needle_1000 Qwen3-8B 214.47 [213.60, 214.78] 4.65 [4.64, 4.66] 4.69 [4.69, 4.70]
needle_4000 Laya English 171.11 [170.28, 171.40] 6.57 [6.55, 6.59] 6.65 [6.63, 6.66]
needle_4000 Laya Multilingual 92.94 [92.84, 93.04] 12.50 [12.46, 12.53] 12.56 [12.52, 12.57]
needle_4000 Llama-3.1-8B-Instruct 705.78 [705.57, 706.11] 1.15 [1.15, 1.15] 1.16 [1.16, 1.16]
needle_4000 Qwen3-8B 758.71 [757.50, 759.77] 1.06 [1.06, 1.06] 1.07 [1.07, 1.07]

Requests include all questions in a state. Fixed-batch throughput is completed requests divided by summed prediction time; it is not server capacity. No generated-token speed, optimal-serving-engine result, architecture-only speedup, or overall cross-task winner is claimed. Jev was not run.

Run provenance

The original Laya measurements are retained from v0.1. The two LLM runs are new in v0.2; no prior result is relabeled as a new run. Every model has per-suite compressed raw predictions and a completed-run manifest: Laya English, Laya Multilingual, Llama-3.1-8B-Instruct, Qwen3-8B.

Reproduce

Python 3.11+ and a CUDA GPU are required for the Laya runs. The reference run used Python 3.12, PyTorch 2.7.1+cu118, Transformers 5.16.1 and Laya 0.3.20. Laya source hashes and model byte hashes appear in each model's metadata. A shared environment with incompatible PyTorch/Transformers versions may need a fresh virtual environment.

git clone https://github.com/CYMCharming/system1bench.git
cd system1bench
python -m venv .venv
source .venv/bin/activate
pip install -e '.[laya]'
python -m system1bench.download
python -m system1bench.prepare
CUDA_VISIBLE_DEVICES=0 python -m system1bench.run --model english --output my-results
CUDA_VISIBLE_DEVICES=0 python -m system1bench.run --model multilingual --output my-results
python -m system1bench.metrics --results my-results
python -m system1bench.report --results my-results

Downloads use immutable revisions and verify SHA256. Standard source data are stored locally under data/; they are not covered by this repository's license. The checked-in results/ are the published runs. Use a new output directory for a different environment, checkpoint, code or batch size. Interrupted suites are rerun atomically; completed suites must match input/configuration fingerprints.

To verify/recompute the published numbers without downloading data or a model:

pip install -e .
python -m unittest discover -s tests -v
python -m system1bench.metrics
python -m system1bench.report

The compressed per-suite results contain source IDs, gold labels, raw answers, probabilities, request hashes, token audits and per-batch timing. They omit original text and rationale annotations. A reader can reconstruct requests from pinned sources and recompute all published metrics independently.

Llama/Qwen reproduction, prompt conversion and probability semantics are documented in LLM_BASELINES.md. These local baselines use the exact same frozen inputs.

Add Jev or another model

Implement the adapter contract in ADAPTERS.md, then use --adapter your_package.module:Adapter --model your_model_name. The runner passes only state and questions; source references are separate. Keep native raw answers and declare model/version, prompt conversion, probability semantics and input budget limitations. No Jev credentials, paid API calls, mocked leaderboard runs or copied third-party model scores are included.

Interpretation

Classification accuracy is a component measure, not end-to-end agent success. Several workflow references are synthetic/teacher labels; their scores measure agreement. Public data may overlap model training, and AG News/BoolQ are known training-task families for Laya. Chinese states mostly use English instructions.

Laya runs use an 8,192-token maximum and 4,096-token prompt-head budget, but Laya internally caps each candidate description at 48 tokens. Read the actual truncation audits and the shared complete-input subset before comparing checkpoints. A larger configured budget alone does not prove full input retention.

Citation

@misc{system1bench2026,
  author = {CYMCharming},
  title = {System1Bench: Benchmarking Jev-Style Decision Models},
  year = {2026},
  howpublished = {\url{https://github.com/CYMCharming/system1bench}},
  note = {Version 0.2; cite the commit and original datasets for reproducibility}
}

Code: MIT. External data/models retain their original terms; see NOTICE.md.

About

System1Bench: Benchmarking Jev-Style Decision Models. Reproducible Laya baselines, typed decisions, bilingual evaluation, OOS detection, safety and robustness.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages