Outcome and baseline
Evaluate CrossPool functionality, resource efficiency, throughput, latency,
and tail latency with fair, reproducible methods. Existing qualification owns
numerical, topology, serving, and report-only performance evidence. CrossPool
currently defines no performance bound while the system remains incomplete.
This issue delivers a reviewable evaluation protocol and baseline compatibility
report. Full evaluation runs and published comparative results are subsequent
work, not conditions for closing this methodology issue.
Suggested implementation route
- Specify hardware, concrete models, arrival process, workloads, memory
budgets, warmup, concurrency, SLOs, metrics/units, repetitions, and raw
artifacts. Define how failed, incomplete, and unavailable runs are reported
before results are collected. Identify the question answered by each
experiment and which variables must remain comparable.
- Audit native SGLang, potentially vLLM, MuxServe, kvcached, and CrossPool
ablations against those experiments. Record concrete versions, compatible
models/topologies, resource-sharing semantics, necessary configuration,
and reasons a baseline is unavailable or not directly comparable.
- Recommend a bounded initial experiment set and review the protocol and
compatibility report. Specify a small paired sanity run to check the
measurement method before large runs; make its resource requirements clear.
- Once the method and resources are accepted, execute reproducible runs,
retain raw inputs/results, and use xbench where applicable or an explicit
reproducible runner. Report excluded or incomplete cases transparently.
- Analyze throughput/latency/resource trade-offs and uncertainty, publish the
qualified comparison scope, and identify evidence that justifies the next
bounded experiment rather than broad support claims.
Discovery completion evidence
- A protocol another engineer can execute with explicit resource and workload
inputs, measurement rules, artifact ownership, and failure treatment.
- A baseline compatibility matrix and justified initial experiment selection.
- A reviewed method recommendation or an evidence-backed conclusion that the
proposed comparisons require a different scope.
Unavailable baselines do not justify widening production interfaces. Creating
performance regression gates requires a separate acceptance decision.
Benchmark Workflows, Timeline observability, and the Product Capability being
evaluated improve this work but are not blanket blocking dependencies.
References
Outcome and baseline
Evaluate CrossPool functionality, resource efficiency, throughput, latency,
and tail latency with fair, reproducible methods. Existing qualification owns
numerical, topology, serving, and report-only performance evidence. CrossPool
currently defines no performance bound while the system remains incomplete.
This issue delivers a reviewable evaluation protocol and baseline compatibility
report. Full evaluation runs and published comparative results are subsequent
work, not conditions for closing this methodology issue.
Suggested implementation route
budgets, warmup, concurrency, SLOs, metrics/units, repetitions, and raw
artifacts. Define how failed, incomplete, and unavailable runs are reported
before results are collected. Identify the question answered by each
experiment and which variables must remain comparable.
ablations against those experiments. Record concrete versions, compatible
models/topologies, resource-sharing semantics, necessary configuration,
and reasons a baseline is unavailable or not directly comparable.
compatibility report. Specify a small paired sanity run to check the
measurement method before large runs; make its resource requirements clear.
retain raw inputs/results, and use
xbenchwhere applicable or an explicitreproducible runner. Report excluded or incomplete cases transparently.
qualified comparison scope, and identify evidence that justifies the next
bounded experiment rather than broad support claims.
Discovery completion evidence
inputs, measurement rules, artifact ownership, and failure treatment.
proposed comparisons require a different scope.
Unavailable baselines do not justify widening production interfaces. Creating
performance regression gates requires a separate acceptance decision.
Benchmark Workflows, Timeline observability, and the Product Capability being
evaluated improve this work but are not blanket blocking dependencies.
References