Outcome and baseline
Provide xbench alongside xtest, with live progress/metrics and durable,
machine-readable benchmark results. Existing serving and report-only runs are
useful inputs. The current test harness owns resource/process management and
test artifacts; it does not already define a benchmark result contract.
This issue covers the benchmark contract, a minimal prototype, and a scoped
target plan. Production CLI and result-retention implementation follow that
plan.
Suggested implementation route
- Define benchmark cases and results: workload, environment, warmup,
measurements, units, repetitions, failures/incomplete runs, raw artifacts,
and retention. Identify which metrics should be live and which become
final only after measurement completes.
- Audit reuse of GPU leases, process supervision, endpoint reservation,
requirements, and artifact ownership. tests.harness is source-owned and
excluded from runtime wheels; decide source-only versus packaged execution
before proposing installed reuse. Separate reusable mechanisms from pytest
collection, JUnit, verdict aggregation, and test-status rendering.
- Prototype one explicit benchmark case with live measurements, stored raw
evidence, and a consistent final result. Exercise cancellation and cleanup.
Propose the case/result schema, packaging boundary, and retention contract.
- After accepting the plan, implement the minimal canonical runner and result
workflow. Share only demonstrated resource/lifecycle mechanisms; benchmark
completion and measurement validity remain independent of test verdicts.
- Validate repeatability, live/stored consistency, failure records, descendant
cleanup before GPU release, active-run artifact protection, and explicit
retention behavior.
Discovery completion evidence
- A reviewed case/result proposal covering measurement and failure semantics.
- A concrete reuse/packaging decision and minimal live-and-stored prototype.
- Reproducible cancellation/cleanup evidence and a scoped implementation
recommendation, or an evidence-backed no-go.
Benchmark measurements do not become performance acceptance gates merely by
using this runner. Timeline observability can improve diagnosis but is not a
prerequisite. The initial issue creates no separate research sub-issues.
References
Outcome and baseline
Provide
xbenchalongsidextest, with live progress/metrics and durable,machine-readable benchmark results. Existing serving and report-only runs are
useful inputs. The current test harness owns resource/process management and
test artifacts; it does not already define a benchmark result contract.
This issue covers the benchmark contract, a minimal prototype, and a scoped
target plan. Production CLI and result-retention implementation follow that
plan.
Suggested implementation route
measurements, units, repetitions, failures/incomplete runs, raw artifacts,
and retention. Identify which metrics should be live and which become
final only after measurement completes.
requirements, and artifact ownership.
tests.harnessis source-owned andexcluded from runtime wheels; decide source-only versus packaged execution
before proposing installed reuse. Separate reusable mechanisms from pytest
collection, JUnit, verdict aggregation, and test-status rendering.
evidence, and a consistent final result. Exercise cancellation and cleanup.
Propose the case/result schema, packaging boundary, and retention contract.
workflow. Share only demonstrated resource/lifecycle mechanisms; benchmark
completion and measurement validity remain independent of test verdicts.
cleanup before GPU release, active-run artifact protection, and explicit
retention behavior.
Discovery completion evidence
recommendation, or an evidence-backed no-go.
Benchmark measurements do not become performance acceptance gates merely by
using this runner. Timeline observability can improve diagnosis but is not a
prerequisite. The initial issue creates no separate research sub-issues.
References