Outcome and baseline
Port a selected CrossPool capability set to a non-CUDA accelerator, with Ascend
as the first candidate. Engine-neutral values and model semantics provide a
starting point, but current production depends on CUDA-specific execution,
communication, synchronization, and process-sharing behavior.
This issue covers reuse/gap analysis, the smallest end-to-end feasibility
prototype, and a scoped port recommendation. Production portability requires
an accepted target plan and its own qualification.
Suggested implementation route
- Audit concrete SGLang Ascend and
sgl-kernel-npu versions for reusable
serving/model/operator paths. Map CrossPool requirements for CUDA Graph,
CUDA IPC, NVSHMEM, CCCL/cooperative synchronization, MPS, Device-memory
observations, and timeline sources to demonstrated equivalents or gaps.
- Choose the smallest capability/model/topology combination that can test the
critical gaps and run it on real Ascend hardware. Exercise execution,
communication, process sharing, memory accounting, and lifecycle. Record
source evidence separately from demonstrated runtime behavior.
- Propose the selected platform boundary, reusable components, necessary
control/data-plane changes, and explicit unsupported capabilities. Write a
scoped target plan with model, serving, and lifecycle qualification.
- After accepting the plan, implement the justified platform mechanisms,
reusing the existing NPU stack where suitable and preserving domain owners.
- Qualify the selected port end to end and publish support only for the
model/engine/platform combinations backed by direct evidence.
Discovery completion evidence
- A versioned reuse audit and platform-gap matrix with evidence for each claim.
- Raw results from the minimal end-to-end hardware prototype.
- A reviewed scoped port recommendation or an evidence-backed no-go.
Hardware unavailability does not establish infeasibility or satisfy the
prototype requirement. Empty backend abstractions in the CUDA implementation
are outside scope. Model Coverage and Timeline observability are beneficial
relationships; the three coverage dimensions remain independent.
References
Outcome and baseline
Port a selected CrossPool capability set to a non-CUDA accelerator, with Ascend
as the first candidate. Engine-neutral values and model semantics provide a
starting point, but current production depends on CUDA-specific execution,
communication, synchronization, and process-sharing behavior.
This issue covers reuse/gap analysis, the smallest end-to-end feasibility
prototype, and a scoped port recommendation. Production portability requires
an accepted target plan and its own qualification.
Suggested implementation route
sgl-kernel-npuversions for reusableserving/model/operator paths. Map CrossPool requirements for CUDA Graph,
CUDA IPC, NVSHMEM, CCCL/cooperative synchronization, MPS, Device-memory
observations, and timeline sources to demonstrated equivalents or gaps.
critical gaps and run it on real Ascend hardware. Exercise execution,
communication, process sharing, memory accounting, and lifecycle. Record
source evidence separately from demonstrated runtime behavior.
control/data-plane changes, and explicit unsupported capabilities. Write a
scoped target plan with model, serving, and lifecycle qualification.
reusing the existing NPU stack where suitable and preserving domain owners.
model/engine/platform combinations backed by direct evidence.
Discovery completion evidence
Hardware unavailability does not establish infeasibility or satisfy the
prototype requirement. Empty backend abstractions in the CUDA implementation
are outside scope. Model Coverage and Timeline observability are beneficial
relationships; the three coverage dimensions remain independent.
References