Codex /goal runner for ProgramBench tasks.
GoalBench runs Codex CLI goal mode against ProgramBench cleanroom task images,
packages each generated replacement as submission.tar.gz, evaluates with
ProgramBench, and publishes a static report. It is a Codex /goal scaffold
measurement, not the official ProgramBench mini-SWE-agent baseline.
ProgramBench asks an agent to rebuild a CLI program from only a compiled
executable and bundled documentation. GoalBench keeps that shape, but swaps the
agent scaffold for Codex CLI /goal.
Current reportable tracks:
| Track | Config | Prompt | Compliance label |
|---|---|---|---|
| Mini-SWE-compatible | configs/full-miniswecompat-xhigh.json |
Short mini-SWE-style prompt with /goal prepended. Kept because the current published site result uses it. |
Mini-SWE-compatible no internet |
Paper prompt + /goal |
configs/cpx62-paperprompt-xhigh.json |
ProgramBench paper prompt with /goal prepended, using mini-SWE-style task execution. |
Paper prompt + /goal no internet |
| Paper prompt + Goal contract | configs/cpx62-goalcontract-xhigh.json |
ProgramBench paper prompt plus a stronger Codex Goal completion contract, using mini-SWE-style task execution. | Paper prompt + Goal contract no internet |
All reportable no-internet tracks use strict host egress, black-box target access, mini-SWE-style task workspaces for the paper-prompt tracks, and post-run audits.
GitHub / Pages
^
| publish only from coordinator
|
local laptop --ssh--> coordinator VM -----------------------------+
| |
| runs Codex /goal sessions |
| owns merged report |
| |
+--> target Docker containers |
| one clean ProgramBench image per task |
| |
+--> ProgramBench evaluator |
| package -> audit -> eval -> summarize |
|
+--> optional eval/inference workers
same repo commit, separate shards
no publishing
Codex runs on the VM host, not inside the target container. Docker is used for the reference target during inference and by ProgramBench during evaluation.
target_sets/*.txt
|
v
run-batch.py watch
|
+-- prepare per-task run root
| CODEX_INITIAL_PROMPT.md
| run.json
| start-target.sh
| start-codex-goal.sh
|
+-- start target container
|
+-- start tmux Codex session
| codex --enable goals ...
| send /goal objective, then paste GOAL_PROMPT.md
|
+-- watch transcript
| running -> goal_done
| failed before goal_done -> fresh retry only
|
v
run-batch.py finalize
|
+-- hard gates
| prompt starts with /goal
| transcript shows /goal
| run.json has model/reasoning/mode/strict egress
|
+-- package-submission
+-- audit-run.py
+-- ProgramBench eval
+-- summarize-results.py
|
v
static report
Every benchmark prompt is rendered to GOAL_PROMPT.md. For auditability, the
paper-prompt tracks also write CODEX_INITIAL_PROMPT.md as the literal
/goal-prefixed paper prompt:
/goal <rendered ProgramBench/GoalBench prompt>
start-codex-goal.sh starts an interactive Codex tmux session, creates the
Goal with a compact objective, then submits the rendered benchmark prompt as the
work instruction. This keeps /goal continuation attached to a live TUI
session instead of relying on a one-shot CLI argument:
codex --enable goals --disable plugins --disable apps \
-m gpt-5.5 \
-c model_reasoning_effort=xhigh \
-c trust_level=trusted \
-C solution \
--yolo --no-alt-screen
tmux send-keys "/goal <objective>" Enter
tmux load-buffer GOAL_PROMPT.md
tmux paste-buffer
tmux send-keys EnterFor the paper-prompt tracks, the bytes immediately after /goal are the
ProgramBench paper prompt from Appendix 8.2, followed by a small harness context
block with instance id, target command, package command, and workspace details.
The Goal-contract variant adds the Codex Goal contract before the same paper
rules, making outcome, verification, boundaries, iteration policy, and blocked
stop conditions explicit.
Before publishing, hard gates check both the prompt file and transcript:
CODEX_INITIAL_PROMPT.md begins with /goal
tmux transcript contains /goal
run.json records model, reasoning effort, inference mode, and strict egress
Reportable no-internet runs use layered controls. The prompt alone is not the security boundary.
Codex process
user: codex-runner
|
+-- UID-scoped iptables
| only loopback is allowed for codex-runner
|
+-- local OpenAI allowlist proxy
| Codex talks to http://127.0.0.1:18080
| proxy makes the OpenAI/Codex network calls
|
+-- guard-bin PATH
| blocks source lookup, package installs, binary analysis tools,
| broad host traversal, and direct docker access
| allows ordinary local git commands, but blocks git source acquisition
| allows loopback curl/wget, but blocks external URL fetches
|
+-- target wrapper
sudo -n /usr/local/bin/pb-target-exec <container> <command>
only allows docker exec into pb-goal-* target containers as agent
The coordinator/root account still needs normal network for setup, Docker pulls, Git operations, and publishing. The restricted Codex task user does not.
GoalBench uses a host-side wrapper to transport allowed black-box CLI interactions into the target container. This differs from mini-SWE-agent's in-container execution, but the wrapper is restricted to normal user-interface observations of the target executable and forbids source lookup, binary reading, disassembly, tracing, instrumentation, and evaluator/test access.
Required host checks before a serious run:
scripts/doctor.sh configs/cpx62-paperprompt-xhigh.json
sudo scripts/linux-openai-egress-guard.sh status codex-runner
sudo -H -u codex-runner sudo -n /usr/local/bin/pb-target-exec __pb-wrapper-check trueThe wrapper check should fail with refusing non ProgramBench target container,
not command not found and not sudo: a password is required.
Use Linux amd64 for serious runs. ProgramBench task images are published for
linux/amd64; Mac/ARM is smoke-test only.
git clone git@github.com:Muhtasham/goalbench.git
cd goalbench
scripts/bootstrap-linux-vm.sh --codex-user codex-runner
codex login
scripts/doctor.sh configs/cpx62-paperprompt-xhigh.jsonBootstrap installs Docker, uv, tmux, Codex CLI if missing, a sibling
../ProgramBench checkout, Codex goal/fast defaults, and the target wrapper
used by reportable black-box runs.
Start inference in a named tmux session so SSH disconnects do not stop the run:
RUN_VERSION="$(date -u +%Y%m%dT%H%M%SZ)"
tmux new-session -d -s "goalbench-watch-$RUN_VERSION" \
"cd '$PWD' && uv run python scripts/run-config.py watch configs/cpx62-paperprompt-xhigh.json --run-version '$RUN_VERSION'"Start a second tmux session to finalize completed rows sequentially:
tmux new-session -d -s "goalbench-finalize-$RUN_VERSION" \
"cd '$PWD' && while true; do uv run python scripts/run-config.py finalize configs/cpx62-paperprompt-xhigh.json --run-version '$RUN_VERSION' --programbench-repo ../ProgramBench --allow-partial --limit 1; sleep 60; done"watch can run up to max_parallel Codex sessions. finalize --limit 1
packages, audits, and evaluates one ready task per call, so a simple loop can
keep evaluation moving without running multiple ProgramBench evals on the same
VM.
For a 5-way run, split the 200 task list into five 40-task shard files and use one batch name per shard:
goalbench-vm shard 0/5 cpx62-paperprompt-xhigh-shard-0-of-5
goalbench-eval-1 shard 1/5 cpx62-paperprompt-xhigh-shard-1-of-5
goalbench-eval-2 shard 2/5 cpx62-paperprompt-xhigh-shard-2-of-5
goalbench-eval-3 shard 3/5 cpx62-paperprompt-xhigh-shard-3-of-5
goalbench-eval-4 shard 4/5 cpx62-paperprompt-xhigh-shard-4-of-5
Each VM runs two tmux supervisors:
pb-paperprompt-xhigh-shard-N-<version>
inference watcher, max 10 Codex /goal sessions
pb-paperprompt-xhigh-finalize-shard-N-<version>
sequential package -> audit -> ProgramBench eval loop
Keep every VM on the same Git commit, Codex version, config, and RUN_VERSION.
Only the coordinator should merge artifacts and publish the static site.
GoalBench does not paste continuation prompts into failed Codex sessions.
Retryable:
session_failed_before_goal_done: tmux/Codex ended before/goalcompleted and no valid submission was produced. Retry starts a fresh solution directory and records attempt metadata.
Not retryable:
- valid low-scoring submissions
- audit violations
- evaluator results
- normal no-submission exits unless explicitly reported as a fresh attempt
Command:
uv run python scripts/run-config.py retry \
configs/cpx62-paperprompt-xhigh.json \
--run-version "$RUN_VERSION" \
--failed \
--max-attempts 2| Mode | Meaning |
|---|---|
mini-swe-compatible-nointernet |
Shorter parity prompt with the same no-internet enforcement. Kept for the already-published site result. |
paper-prompt-nointernet |
ProgramBench paper prompt with /goal prepended, strict egress, and mini-SWE-style task execution. |
paper-prompt-goal-contract-nointernet |
ProgramBench paper prompt with a stronger Goal contract plus strict egress and mini-SWE-style task execution. |
The static report mirrors ProgramBench's headline shape:
- resolved rate
- almost-resolved rate at
score >= 0.95 - average behavioral pass rate
- estimated cost
- Codex calls
- wall-clock time
- per-task detail pages
Cost is estimated from local Codex token logs and a refreshed OpenAI pricing snapshot. It is not authoritative billing.
Build locally:
uv run python scripts/build-report.py --output-dir docs
uv run python scripts/privacy-scan.pyPublic output includes sanitized aggregate rows, public eval summaries, and per-task aggregate Codex trace summaries. Raw Codex logs and submission tarballs stay local by default.
scripts/doctor.sh configs/cpx62-paperprompt-xhigh.json
uv run python scripts/run-config.py status configs/cpx62-paperprompt-xhigh.json
uv run python scripts/run-config.py finalize configs/cpx62-paperprompt-xhigh.json --programbench-repo ../ProgramBench --allow-partial --limit 1
uv run python scripts/validate-run-gates.py local_state/batches/<batch>/<version>/results.csv
scripts/backup-run-root.sh --batch-name <batch> --run-version <version>
uv run ruff check .
uv run ty check
uv run pre-commit run --all-files