Single-L20 post-training, verifier-guided inference, and executable benchmark infrastructure for code models.
The completed three-seed SFT/RLVR campaign is a negative result, separate from the historical inference-system results below. The selected RLVR seed improved rStar development from 76/200 to 80/200 (+2.0 points; paired 95% CI [-4.0, +8.0]) and matched MBPP validation at 64/90. It then regressed on both reused EvalPlus guardrails: HumanEval+ 137/164 to 110/164 and MBPP+ 271/378 to 257/378. This does not establish a retained model-quality improvement.
Training used two RTX 4090 GPUs per run, not the original single-L20 setup. The date-held-out LiveCodeBench evaluation is blocked by an unavailable source artifact; it has no replacement score. See the campaign report and receipts for selection rules, seed-level results, and output-format/semantic failure diagnosis. No post-audit tuning is included in this result.
L20-CodeForge is kept as the executable-code benchmark and post-training sandbox in the L20 project family. Its scope is benchmark protocol, candidate generation, repair, verifier-guided inference, trajectory data, and reward signals for code models.
For serving, kernel, and runtime infrastructure work, use single-gpu-inference-lab. For from-scratch pretraining and public checkpoint release artifacts, use l20-edu-135m-pretrain. This repository should stay focused on executable coding benchmarks rather than becoming a second general L20 infrastructure repo.
L20-CodeForge is an eval-first research stack for making a small GPU budget produce measurable coding capability. The project focuses on the pieces that matter in real post-training work: clean benchmark protocol, public/private test separation, candidate generation, repair, verifier experiments, trajectory data, and negative-result audits.
Status: the repository currently demonstrates strong system-level gains on public benchmarks. It does not yet claim a +15 point greedy model-weight improvement over the base model.
Core docs: Reproducibility, Architecture, and paper draft.
| Benchmark | Base model / checkpoint | Protocol | Baseline | L20-CodeForge result | Delta | Artifact |
|---|---|---|---|---|---|---|
LiveCodeBench release_v6 full suite |
Qwen2.5-Coder-7B-Instruct |
temperature=0.8, n=8, public-test selection, hidden replay |
297/1055 (28.15%) greedy |
403/1055 (38.20%) |
+106 tasks, +10.05 points |
benchmarks/livecodebench_full_release_v6_2026_05_22/ |
| EvalPlus HumanEval+ | Qwen2.5-Coder-7B-Instruct |
clean public-signal system, official EvalPlus scoring | 84.8% greedy |
92.7% |
+7.9 points |
benchmarks/evalplus_l20_codeforge_2026_05_22/ |
| EvalPlus MBPP+ | Qwen2.5-Coder-7B-Instruct |
clean public-signal system, official EvalPlus scoring | 72.2% greedy |
81.7% |
+9.5 points |
benchmarks/evalplus_l20_codeforge_2026_05_22/ |
X-Coder medium control12 |
IIGroup/X-Coder-RL-Qwen2.5-7B |
strict code generation, public-only repair, one public-feedback round | 0/12 auto/strict starter-prefix checks |
4/12 |
+4 tasks |
docs/MILESTONE_9_XCODER_L20_PLUS15_PROBE.md |
Cross-benchmark guardrail: benchmarks/generalization_scorecard_2026_05_23/
records a PASS gate over full LiveCodeBench plus EvalPlus HumanEval+/MBPP+.
The 60-task equal-candidate-pool audit
finds public-test and behavior selection both at 21/60, versus first-candidate
19/60 (paired p=0.5). Public selection took 24.008 seconds and behavior
selection 152.367 seconds in the historical runs, with shared generation of
1849.462 seconds. This is equal candidate generation, not equal total compute;
behavior adds selection overhead without a pass-count gain on this shard.
The oracle is only 23/60. This retrospective audit is not a fresh held-out result.
| Artifact | SHA-256 |
|---|---|
benchmarks/generalization_scorecard_2026_05_23/scorecard.json |
1eb0402378ea25732225b29d7ba367b6111ab3351e54cc7c01fa7646a7a12712 |
benchmarks/livecodebench_full_release_v6_2026_05_22/full_n8_public_select_summary.json |
2a0ff919aa15eb9ecdf74824f7bf790a23f6d0197ef74970b6190c60e0e00772 |
benchmarks/evalplus_l20_codeforge_2026_05_22/summary.csv |
08732bbb76450f92ef3c02fa97a163aba01f71028365072c205c5a3af45d5550 |
See REPRODUCIBILITY.md for the full hash-verification command and expected output.
- The LiveCodeBench and EvalPlus numbers above are system-level results unless a row explicitly says greedy model baseline.
- Public tests, public examples, and public prompt metadata may be used for candidate selection and repair. Hidden/private tests are reserved for final measurement and audits.
- This is not an official leaderboard submission. It is a reproducible local checkpoint with saved generations, reports, hashes, and protocol notes.
- Targeted probes, such as first-12 rescues or code-prefix experiments, are not presented as broad benchmark results.
- The current +15 goal is still open: the next milestone is to turn system-level gains into a cleaner greedy or single-sample model-capability improvement without overfitting to visible tests.
L20-CodeForge is organized around one constraint: a single NVIDIA L20 should be enough to run a serious coding post-training loop if the loop is selective, measured, and executable.
The repository contains:
- LiveCodeBench and EvalPlus evaluation harnesses with saved reports.
- Public-test selection, repair, and behavior-test tooling.
- Candidate health audits for syntax, entrypoint, and execution failure modes.
- Failure-driven algorithmic-code verifier audits with labeled false-positive, false-negative, faulty-code kill-rate, and behavior-diversity gates.
- A trajectory schema for repo-repair agents and model training data.
- SFT, DPO, reward-function, and GRPO/RLVR scaffolding.
- A migrated T4 GRPO reasoning experiment under
experiments/grpo_t4/, retained as an experiment artifact rather than a separate public repo. - Single-L20 setup scripts and GPU sanity checks.
Local development:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -e ".[dev,bench]"
python -m pytest -q
python -m l20_codeforge verify-artifacts
python -m l20_codeforge profile
python -m l20_codeforge smoke-loop
python -m l20_codeforge audit-code-verifier \
examples/verifier_audit.square.jsonl \
--output artifacts/verifier/square-audit.json \
--min-reference-solutions 2 \
--min-faulty-kill-rate 1.0 \
--max-false-positive-rate 0.0 \
--max-false-negative-rate 0.0 \
--fail-on-gatesThe test suite should exit successfully; optional dependency tests may skip.
verify-artifacts should return "status": "PASS"; it checks the five pinned
historical benchmark artifacts listed in REPRODUCIBILITY.md. The newer RLVR
receipt and equal-pool replay tests are separate checks in the test suite.
On an L20 host:
bash scripts/bootstrap_remote.sh
source .venv/bin/activate
python scripts/check_gpu.py
python -m pytest -qThe bootstrap script creates an isolated Python environment, installs a CUDA PyTorch stack, and installs this package with training and development dependencies. It does not download large model weights.
Build the cross-benchmark scorecard:
python scripts/build_generalization_scorecard.py \
--output-dir benchmarks/generalization_scorecard_2026_05_23Expected output includes:
{
"status": "PASS",
"checks": [
{
"name": "lcb_overall_improves",
"value": 0.100474,
"threshold": 0.0,
"passed": true
}
]
}Re-run the packaged EvalPlus scoring:
python -m l20_codeforge eval-evalplus humaneval \
benchmarks/evalplus_l20_codeforge_2026_05_22/samples/humaneval.mixed-target.literal-combined.public-consensus-selected.samples.jsonl \
--output /tmp/humaneval_recheck.json \
--parallel 8
python -m l20_codeforge eval-evalplus mbpp \
benchmarks/evalplus_l20_codeforge_2026_05_22/samples/mbpp.temp08.n5-plus-basefallback-n30.public-consensus-shortest-selected.samples.jsonl \
--output /tmp/mbpp_recheck.json \
--parallel 8LiveCodeBench full-suite reproduction requires a local materialized
release_v6 JSONL with private tests. That file is intentionally not committed.
The committed package includes saved generations, compact summaries, hashes,
and evaluator outputs. See
benchmarks/livecodebench_full_release_v6_2026_05_22/README.md for the full
generation and replay commands.
| Path | Purpose |
|---|---|
benchmarks/generalization_scorecard_2026_05_23/scorecard.json |
Machine-readable LCB + EvalPlus gate. |
benchmarks/livecodebench_full_release_v6_2026_05_22/full_n8_public_select_summary.json |
Headline full LCB n=8 public-selection result. |
benchmarks/livecodebench_full_release_v6_2026_05_22/README.md |
Full LCB protocol, hashes, breakdowns, and commands. |
benchmarks/evalplus_l20_codeforge_2026_05_22/summary.csv |
EvalPlus greedy baselines and clean system rows. |
docs/MILESTONE_8_LCB_PLUS15_FOCUS.md |
Research plan for converting system gain into model-capability gain. |
docs/MILESTONE_9_XCODER_L20_PLUS15_PROBE.md |
X-Coder probe, control slices, positive results, and overfitting checks. |
docs/MILESTONE_10_FAILURE_DRIVEN_RLVR.md |
Verifier-first RLVR decision, audit gates, reward ablation, and four-GPU pilot boundary. |
configs/qwen25_coder_7b_failure_driven_rlvr.yaml |
Pinned Phase-A data/verifier gates and proposed 100-step GRPO pilot. |
docs/MILESTONE_11_BASE_SFT_RLVR_CAMPAIGN.md |
Active Base to verified-SFT to RLVR campaign, frozen data, promotion gate, and evidence boundary. |
configs/qwen25_coder_7b_base_sft_rlvr_20260829.yaml |
Exact active campaign data hashes, training settings, topology, and no-regression gates. |
benchmarks/code_rlvr_base_sft_rlvr_2026_08_29/ |
Frozen data, LiveCodeBench overlap, and GPU smoke receipts. |
benchmarks/code_rlvr_retention_v2_2026_08_30/ |
Three-seed retention-aware SFT/RLVR v2 protocol, GPU receipts, development selection, and failed EvalPlus guardrail. |
scripts/evaluate_lcb_generations.py |
Hidden replay, public selection, behavior-input selection, and variable-candidate handling. |
scripts/regenerate_lcb_final_answers.py |
Second-pass code regeneration with optional public-test feedback. |
src/l20_codeforge/ |
Package code for data, envs, evals, rewards, inference, training, and GPU profiling. |
- Full-suite
n=8public-test selection moved Qwen2.5-Coder-7B-Instruct from297/1055to403/1055on LiveCodeBenchrelease_v6. - EvalPlus clean public-signal systems improved HumanEval+ and MBPP+ without using EvalPlus extra tests for selection.
- One public-feedback repair round on the X-Coder medium control slice lifted
the gate from
2/12after public-only repair to4/12. - The evaluation harness caught optimistic small-subset results and forced the project onto full-suite and cross-benchmark scorecards.
These failures are kept in the repository because they are useful research signal, not noise.
- Small LiveCodeBench subsets were optimistic relative to the full 1,055-task suite.
- Automatic starter-prefix prompting did not generalize on the medium
control12slice (0/12). - Multi-source code-only repair did not improve the medium fail10 slice.
- A second public-feedback repair round produced public-test signal but
0/8hidden passes, which is an overfitting warning. - Input-only adaptive differential tests had little leverage when the candidate pool lacked multiple public-passing alternatives.
- The first expected-output verifier pass regressed the targeted replay, so it needs calibration or a stronger verifier before it can affect headline runs.
configs/ L20-first experiment configs
docs/ architecture notes, milestones, and runbooks
scripts/ benchmark, repair, audit, setup, and GPU utilities
src/l20_codeforge/
agents/ mini-SWE-agent trajectory adapter
context/ repo context packing
data/ task, trajectory, SFT, and preference builders
envs/ local repo execution adapters
evals/ EvalPlus, patch, SFT, and eval-card tooling
gpu/ L20 profile and memory policy
inference/ candidate selectors
rewards/ executable and patch-quality reward functions
training/ SFT and TRL-compatible reward helpers
tests/ unit and regression tests
benchmarks/ committed benchmark packages and audit outputs
- Build a 2K-task, source- and license-audited faulty-code pool with exact and near-duplicate LiveCodeBench exclusion receipts.
- Admit verifier tests only after references, known-correct solutions, faulty kill rate, false-positive rate, and false-negative rate pass the Milestone 10 gates.
- Run matched verified-SFT and 100-step dense-vs-binary RLVR ablations; keep the SWE-Gym adapter as a negative control.
- Improve greedy
n=1on held-out algorithmic tasks without regressing EvalPlus; reportpass@4and public selection separately. - Keep frozen/hidden tests out of model selection and run the legacy full LiveCodeBench replay only after the recipe is fixed.
- LiveCodeBench: https://livecodebench.github.io/
- EvalPlus: https://github.com/evalplus/evalplus
- Qwen2.5-Coder: https://qwenlm.github.io/blog/qwen2.5-coder-family/
- X-Coder model card: https://huggingface.co/IIGroup/X-Coder-RL-Qwen2.5-7B
- HardTests: https://arxiv.org/abs/2505.24098
- RobustTests: https://arxiv.org/abs/2608.24135
- TRL GRPO trainer: https://huggingface.co/docs/trl/grpo_trainer
- mini-SWE-agent: https://github.com/SWE-agent/mini-swe-agent
- SWE-bench: https://www.swebench.com/