Skip to content

Accelerator Portability: establish Ascend reuse and platform feasibility #8

Description

@Coekjan

Outcome and baseline

Port a selected CrossPool capability set to a non-CUDA accelerator, with Ascend
as the first candidate. Engine-neutral values and model semantics provide a
starting point, but current production depends on CUDA-specific execution,
communication, synchronization, and process-sharing behavior.

This issue covers reuse/gap analysis, the smallest end-to-end feasibility
prototype, and a scoped port recommendation. Production portability requires
an accepted target plan and its own qualification.

Suggested implementation route

  1. Audit concrete SGLang Ascend and sgl-kernel-npu versions for reusable
    serving/model/operator paths. Map CrossPool requirements for CUDA Graph,
    CUDA IPC, NVSHMEM, CCCL/cooperative synchronization, MPS, Device-memory
    observations, and timeline sources to demonstrated equivalents or gaps.
  2. Choose the smallest capability/model/topology combination that can test the
    critical gaps and run it on real Ascend hardware. Exercise execution,
    communication, process sharing, memory accounting, and lifecycle. Record
    source evidence separately from demonstrated runtime behavior.
  3. Propose the selected platform boundary, reusable components, necessary
    control/data-plane changes, and explicit unsupported capabilities. Write a
    scoped target plan with model, serving, and lifecycle qualification.
  4. After accepting the plan, implement the justified platform mechanisms,
    reusing the existing NPU stack where suitable and preserving domain owners.
  5. Qualify the selected port end to end and publish support only for the
    model/engine/platform combinations backed by direct evidence.

Discovery completion evidence

  • A versioned reuse audit and platform-gap matrix with evidence for each claim.
  • Raw results from the minimal end-to-end hardware prototype.
  • A reviewed scoped port recommendation or an evidence-backed no-go.

Hardware unavailability does not establish infeasibility or satisfy the
prototype requirement. Empty backend abstractions in the CUDA implementation
are outside scope. Model Coverage and Timeline observability are beneficial
relationships; the three coverage dimensions remain independent.

References

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:portabilityNon-CUDA accelerator feasibility and portability.researchEvidence gathering, compatibility investigation, or an exploratory prototype.roadmapA direction-level issue with an explicitly bounded initial stage.

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions