Accuracy on a 77-way or 151-way task hides where a model fails. The leaderboard should show, for each choice task, the classes each provider gets wrong most often.
What to build
A short table under each choice task on the leaderboard page: the five classes with the lowest F1 for each provider, computed from the predictions in the result files (never from stored numbers), like the existing per-language table for massive.
Where to look
src/thinkless/bench/decisions/leaderboard.py: _locale_breakdown is the closest example.
src/thinkless/bench/decisions/score.py: correct_rows and _macro_f1.
Done when
thinkless bench decisions leaderboard benchmarks/decisions/results --out docs/decision-benchmark/leaderboard.md produces the tables, and CI's leaderboard check passes.
- A unit test covers a small synthetic case.
Accuracy on a 77-way or 151-way task hides where a model fails. The leaderboard should show, for each choice task, the classes each provider gets wrong most often.
What to build
A short table under each choice task on the leaderboard page: the five classes with the lowest F1 for each provider, computed from the predictions in the result files (never from stored numbers), like the existing per-language table for
massive.Where to look
src/thinkless/bench/decisions/leaderboard.py:_locale_breakdownis the closest example.src/thinkless/bench/decisions/score.py:correct_rowsand_macro_f1.Done when
thinkless bench decisions leaderboard benchmarks/decisions/results --out docs/decision-benchmark/leaderboard.mdproduces the tables, and CI's leaderboard check passes.