Skip to content

rl: add the elastic resource benchmark contract, gating, evidence and paired work - #66

Merged
michaellchung merged 1 commit into
mainfrom
feat/rl-elastic-benchmark
Sep 29, 2026
Merged

michaellchung merged 1 commit into
mainfrom
feat/rl-elastic-benchmark

Conversation

@michaellchung

Copy link
Copy Markdown
Contributor

The GPU-free half of the rl-elastic-resource-benchmark study harness: scripts/benchmark_rl_elastic.py plans and reports studies comparing fixed GPU partitions (trainer / rollout / standby) against scheduled and automatic resizing inside one RL island.

What lands

  • Versioned study manifest with a canonical hash, in calibration and formal modes
  • Config / edge / pool validation gated behind runtime capability attestation — an item only becomes runnable once a runner attests it can run it
  • Evidence index with tamper-rejecting resume
  • Dry-run plan CLI, layered results and a report that exits non-zero while the matrix is incomplete
  • Paired splits, schedules and scenarios reusing the legacy benchmark_rl pairing through yeto/rl/elastic_benchmark/legacy.py

scripts/benchmark_rl.py, its CLI and its result fields are unchanged.

No GPU runner ships here. None of the four commands loads a model, imports torch/ray/miles, or creates cloud resources.

Conflicts

None — every file is new (yeto/rl/elastic_benchmark/, the script, the test module, the doc). Nothing existing is touched.

Spec

The planning contract lives in the miles repository under openspec/changes/rl-elastic-resource-benchmark/, as docs/RL_ELASTIC_BENCHMARK.md states; there is deliberately no OpenSpec change for it in this repo.

Tests

19 pass locally (tests/test_rl_elastic_benchmark.py).

🤖 Generated with Claude Code

…red work

Implements the GPU-free half of the rl-elastic-resource-benchmark change:
versioned study manifest with canonical hash and calibration/formal modes,
config/edge/pool validation gated by runtime capability attestation,
evidence index with tamper-rejecting resume, dry-run plan CLI, layered
results and report, and paired splits/schedules/scenarios that reuse the
legacy benchmark_rl pairing without changing its CLI.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@michaellchung
michaellchung merged commit 8653152 into main Sep 29, 2026
1 of 6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant