Skip to content

Evaluation and Baselines: define methodology and baseline compatibility #10

Description

@Coekjan

Outcome and baseline

Evaluate CrossPool functionality, resource efficiency, throughput, latency,
and tail latency with fair, reproducible methods. Existing qualification owns
numerical, topology, serving, and report-only performance evidence. CrossPool
currently defines no performance bound while the system remains incomplete.

This issue delivers a reviewable evaluation protocol and baseline compatibility
report. Full evaluation runs and published comparative results are subsequent
work, not conditions for closing this methodology issue.

Suggested implementation route

  1. Specify hardware, concrete models, arrival process, workloads, memory
    budgets, warmup, concurrency, SLOs, metrics/units, repetitions, and raw
    artifacts. Define how failed, incomplete, and unavailable runs are reported
    before results are collected. Identify the question answered by each
    experiment and which variables must remain comparable.
  2. Audit native SGLang, potentially vLLM, MuxServe, kvcached, and CrossPool
    ablations against those experiments. Record concrete versions, compatible
    models/topologies, resource-sharing semantics, necessary configuration,
    and reasons a baseline is unavailable or not directly comparable.
  3. Recommend a bounded initial experiment set and review the protocol and
    compatibility report. Specify a small paired sanity run to check the
    measurement method before large runs; make its resource requirements clear.
  4. Once the method and resources are accepted, execute reproducible runs,
    retain raw inputs/results, and use xbench where applicable or an explicit
    reproducible runner. Report excluded or incomplete cases transparently.
  5. Analyze throughput/latency/resource trade-offs and uncertainty, publish the
    qualified comparison scope, and identify evidence that justifies the next
    bounded experiment rather than broad support claims.

Discovery completion evidence

  • A protocol another engineer can execute with explicit resource and workload
    inputs, measurement rules, artifact ownership, and failure treatment.
  • A baseline compatibility matrix and justified initial experiment selection.
  • A reviewed method recommendation or an evidence-backed conclusion that the
    proposed comparisons require a different scope.

Unavailable baselines do not justify widening production interfaces. Creating
performance regression gates requires a separate acceptance decision.
Benchmark Workflows, Timeline observability, and the Product Capability being
evaluated improve this work but are not blanket blocking dependencies.

References

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:benchmarksBenchmark execution, evaluation methodology, and baselines.researchEvidence gathering, compatibility investigation, or an exploratory prototype.roadmapA direction-level issue with an explicitly bounded initial stage.

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions