Skip to content

Benchmark Workflows: define and prototype the xbench contract #9

Description

@Coekjan

Outcome and baseline

Provide xbench alongside xtest, with live progress/metrics and durable,
machine-readable benchmark results. Existing serving and report-only runs are
useful inputs. The current test harness owns resource/process management and
test artifacts; it does not already define a benchmark result contract.

This issue covers the benchmark contract, a minimal prototype, and a scoped
target plan. Production CLI and result-retention implementation follow that
plan.

Suggested implementation route

  1. Define benchmark cases and results: workload, environment, warmup,
    measurements, units, repetitions, failures/incomplete runs, raw artifacts,
    and retention. Identify which metrics should be live and which become
    final only after measurement completes.
  2. Audit reuse of GPU leases, process supervision, endpoint reservation,
    requirements, and artifact ownership. tests.harness is source-owned and
    excluded from runtime wheels; decide source-only versus packaged execution
    before proposing installed reuse. Separate reusable mechanisms from pytest
    collection, JUnit, verdict aggregation, and test-status rendering.
  3. Prototype one explicit benchmark case with live measurements, stored raw
    evidence, and a consistent final result. Exercise cancellation and cleanup.
    Propose the case/result schema, packaging boundary, and retention contract.
  4. After accepting the plan, implement the minimal canonical runner and result
    workflow. Share only demonstrated resource/lifecycle mechanisms; benchmark
    completion and measurement validity remain independent of test verdicts.
  5. Validate repeatability, live/stored consistency, failure records, descendant
    cleanup before GPU release, active-run artifact protection, and explicit
    retention behavior.

Discovery completion evidence

  • A reviewed case/result proposal covering measurement and failure semantics.
  • A concrete reuse/packaging decision and minimal live-and-stored prototype.
  • Reproducible cancellation/cleanup evidence and a scoped implementation
    recommendation, or an evidence-backed no-go.

Benchmark measurements do not become performance acceptance gates merely by
using this runner. Timeline observability can improve diagnosis but is not a
prerequisite. The initial issue creates no separate research sub-issues.

References

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:benchmarksBenchmark execution, evaluation methodology, and baselines.researchEvidence gathering, compatibility investigation, or an exploratory prototype.roadmapA direction-level issue with an explicitly bounded initial stage.

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions