Skip to content

Context Parallel Serving: establish SGLang feasibility and target design #5

Description

@Coekjan

Outcome and baseline

Support SGLang Prefill and Decode Context Parallelism with degree greater than
one while preserving FFN row ownership and Elastic KV Cache capacity semantics.
CrossPool currently rejects these modes, and its attention topology and FFN
handoff use TP/DP rank geometry. Removing validation alone would leave the
ownership and capacity contracts unresolved.

This issue covers the SGLang support audit, installed-serving prototype, and
target-design recommendation. Production integration is subsequent work.

Suggested implementation route

  1. Audit the pinned SGLang implementation for concrete model and topology
    support. Treat Prefill and Decode Context Parallelism separately. Identify
    runnable configurations, row partition/gather behavior, KV position
    partitioning, and eager/captured execution requirements.
  2. Run the smallest installed-serving experiments that reveal FFN input/output
    row ownership, KV Capacity Group membership, and graph behavior for the
    proposed modes. Retain rank-local shape and protocol evidence rather than
    inferring correctness from successful HTTP completion alone.
  3. Propose the attention topology, FFN row, capacity-group, and generation
    contracts. Specify exactly which model/mode/topology combinations proceed
    and the numerical, graph, serving, and reclamation evidence each requires.
  4. After accepting the target plan, implement the topology, handoff, and
    capacity changes at their owners and admit only the designed modes.
  5. Qualify eager and graph results, row delivery, group capacity transitions,
    serving completion, and failure/shutdown for each selected combination.

Discovery completion evidence

  • A source-backed SGLang support matrix with runnable configurations.
  • Raw installed-serving prototype evidence for proposed supported modes.
  • A reviewed target plan or an evidence-backed no-go identifying the specific
    model, mode, or contract that prevents proceeding.

Missing resources leave required experiments incomplete. Existing
model-qualification and installed-serving harnesses are reusable evidence
surfaces; existing TP support does not establish context-parallel support.

References

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:servingContext Parallel Serving and additional serving integrations.researchEvidence gathering, compatibility investigation, or an exploratory prototype.roadmapA direction-level issue with an explicitly bounded initial stage.

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions