The leaderboard has no floor. A model that always answers "no" gets 64.9% accuracy on jailbreak (259 of the 399 rows are benign), which makes the real numbers harder to read.
What to build
A provider that answers every question with the most common label in the task's calibration rows, run on the benchmark and added as a verified result, so every table shows what "no skill" scores.
Where to look
src/thinkless/bench/decisions/run.py: build_submission_provider shows how providers are built from a spec.
src/thinkless/bench/decisions/tasks.py: load_rows(task, "calibration") gives the rows to count labels from.
tests/unit/test_decision_benchmark.py: the fake providers there are a good template.
Done when
thinkless bench decisions run --provider majority works for all four question kinds (for extraction, every field is empty).
- The result is in
benchmarks/decisions/results/verified/ and the leaderboard page is regenerated.
- A unit test covers it.
The leaderboard has no floor. A model that always answers "no" gets 64.9% accuracy on
jailbreak(259 of the 399 rows are benign), which makes the real numbers harder to read.What to build
A provider that answers every question with the most common label in the task's calibration rows, run on the benchmark and added as a verified result, so every table shows what "no skill" scores.
Where to look
src/thinkless/bench/decisions/run.py:build_submission_providershows how providers are built from a spec.src/thinkless/bench/decisions/tasks.py:load_rows(task, "calibration")gives the rows to count labels from.tests/unit/test_decision_benchmark.py: the fake providers there are a good template.Done when
thinkless bench decisions run --provider majorityworks for all four question kinds (for extraction, every field is empty).benchmarks/decisions/results/verified/and the leaderboard page is regenerated.