Skip to content

Serving-engine Coverage: establish vLLM feasibility and integration scope #6

Description

@Coekjan

Outcome and baseline

Support another serving engine while keeping engine-owned types and state out
of core Plans, Registries, Projections, and the native ABI. vLLM is the first
candidate. The existing SGLang integration and xpool.ops.ffn_shim expose a
useful boundary, while the FfnAgent also has a bounded dependency on SGLang
low-level Expert kernels that needs explicit treatment.

This issue covers compatibility investigation, a representative end-to-end
spike, and the integration target design. It does not establish vLLM support.

Suggested implementation route

  1. Audit a concrete vLLM version's extension surfaces for module replacement,
    weight filtering, TP/DP, eager/graph execution, output ownership, process
    lifecycle, and failure propagation. Identify which existing integration
    behavior is engine-neutral and which remains SGLang-owned.
  2. Prototype a representative qualified model through the FFN shim and current
    Transport/Fabric path. Observe startup, request completion, output
    ownership, eager/captured execution, and failure/shutdown. Document the
    treatment of low-level SGLang operator dependencies.
  3. Propose a scoped integration plan, supported model/topology combinations,
    bootstrap/configuration ownership, and qualification requirements. Decide
    whether public extension seams suffice and justify any narrower dependency.
  4. After accepting the plan, implement the concrete vLLM integration and keep
    engine translation at that integration's ownership boundary.
  5. Qualify numerical behavior, eager/graph serving, topology, lifecycle, and
    failure for the selected scope while retaining valid SGLang evidence.

Discovery completion evidence

  • A compatibility and ownership report tied to a concrete upstream version.
  • Reproducible representative spike evidence, including lifecycle behavior.
  • A reviewed integration plan or an evidence-backed no-go naming insufficient
    extension surfaces or incompatible execution contracts.

Create a shared provider abstraction only after a second actual integration
demonstrates a shared seam. Missing hardware is an unmet prerequisite. Model
Coverage and Timeline observability can help evaluation; neither implies a
blanket cross-engine model-coverage claim.

References

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:servingContext Parallel Serving and additional serving integrations.researchEvidence gathering, compatibility investigation, or an exploratory prototype.roadmapA direction-level issue with an explicitly bounded initial stage.

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions