Skip to content

Add a majority-class baseline to the decision benchmark #8

Description

@inboxpraveen

The leaderboard has no floor. A model that always answers "no" gets 64.9% accuracy on jailbreak (259 of the 399 rows are benign), which makes the real numbers harder to read.

What to build

A provider that answers every question with the most common label in the task's calibration rows, run on the benchmark and added as a verified result, so every table shows what "no skill" scores.

Where to look

  • src/thinkless/bench/decisions/run.py: build_submission_provider shows how providers are built from a spec.
  • src/thinkless/bench/decisions/tasks.py: load_rows(task, "calibration") gives the rows to count labels from.
  • tests/unit/test_decision_benchmark.py: the fake providers there are a good template.

Done when

  • thinkless bench decisions run --provider majority works for all four question kinds (for extraction, every field is empty).
  • The result is in benchmarks/decisions/results/verified/ and the leaderboard page is regenerated.
  • A unit test covers it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    benchmarkThe decision benchmark: tasks, metrics, runsgood first issueGood for newcomers

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions