Skip to content

[RFC] PyTorch profiling interface with coordinated PD capture and per-rank artifacts聽#71

Description

@bjf-frz

Implementation status

馃毀 An implementation is available in PR #72, currently open and not yet merged. It provides CLI/Python/HTTP profiling controls, coordinated PD capture, and per-rank background parsing and verified archives. The PR description lists activation methods, parameter defaults, the collection workflow, artifacts, and validation results.

The implementation in commit ea4a4bb simplifies activation to explicit start calls, removes ProfileConfig.enabled, and collects native events without custom phase/request annotations. The design below records the original proposal; its enabled field, custom annotations, and session-ID directory layout should not be treated as the current interface. Artifacts use timestamped rank directories. Profiling, PD, and serving CPU regression tests report 76 passed, 9 skipped. GPU/CUDA Graph and multi-GPU PD/NIXL validation remains outstanding; implementation availability does not mean all RFC acceptance requirements are complete.

Objective

Add a configurable PyTorch Profiler interface for inference in vllm-rlt. Support operator/kernel timelines, shapes, stacks, memory and FLOPs, with bounded collection windows and start/stop control. Collect independently by rank, then parse, compress and clean up exports on a background thread.

Use native PyTorch collection and scheduling semantics. Configuration options are independent; there are no named collection presets. This proposal exposes core options rather than reproducing the entire native profiler API.

Core configuration

Python and CLI use the same collection options. The original proposed configuration is listed below; an implementation is available in #72, with interface differences noted in the implementation status above.

Python field CLI argument Default Meaning
enabled --profile False Enable profiling
output_dir --profile-dir PATH required when enabled Root directory for rank artifacts
activities --profile-activities cpu,cuda CPU for CPU inference; CPU and CUDA for CUDA inference Select collected activities
record_shapes --profile-record-shapes False Record operator input shapes
with_stack --profile-with-stack False Record source stacks
profile_memory --profile-memory False Record tensor allocations and frees
with_flops --profile-with-flops False Estimate FLOPs for supported operators
wait --profile-wait N 0 Wait steps per cycle
warmup --profile-warmup N 1 Profiler warmup steps per cycle
active --profile-active N 10 Recorded steps per cycle
repeat --profile-repeat N 1 Cycle count; zero repeats until stopped

The schedule defaults are vllm-rlt choices for a bounded capture, not PyTorch constructor defaults. Forward these four fields to torch.profiler.schedule. Validate nonnegative wait/warmup/repeat and positive active. Validate requested activities and the output directory before starting collection. Unsupported activities produce an error rather than silently falling back.

Keep native meanings and prerequisites from the PyTorch Profiler reference. For example, FLOPs are available only for supported operators, and detailed memory timeline export requires shapes, stacks and memory collection together. Enabling memory alone still records allocation events; it must not silently enable the other options.

Experimental configuration, module hierarchy, execution-trace observers, custom trace IDs, cross-cycle event accumulation, custom factories/callbacks, post-processing timeouts and advanced schedule modifiers are not exposed by this interface. They can be considered later if a concrete use requires them.

CLI example supported by the implementation PR:

vllm-rlt --model /path/to/model --device cuda \
  --prompt "Explain matrix multiplication." --max-tokens 32 \
  --profile --profile-dir /tmp/rlt-prof \
  --profile-activities cpu,cuda \
  --profile-record-shapes --profile-with-stack --profile-memory \
  --profile-wait 2 --profile-warmup 1 --profile-active 10 --profile-repeat 1

Interface and lifecycle

Use a shared vllm_rlt/profiling.py module for configuration, session control and export handling. The runtime owns the native profiler and its trace callback. Provide start_profile(config), stop_profile(), profile_status() and wait_for_profile_artifacts() through the Python integration. Exact placement on the LLM/engine facade follows existing lifecycle ownership.

Offline CLI collection starts after model loading and runtime initialization. Python callers and serving controls can start collection around the workload of interest. Serving routes start/stop to the execution worker, not merely the HTTP handler. PD dispatches the same configuration and session ID to all workers and gathers their status and artifact locations.

For step-scheduled collection, advance profiler.step() exactly once after each completed LLMEngine.step() in the ordinary engine. In PD, the P worker uses prefill_step() instead; advance after a prefill iteration that submitted work, and after a D worker's engine.step(). Do not count idle polling as work or advance again at the coordinator. Include transfer progression within the worker's profiling context. An engine step is not an output token, request or recurrent loop, and does not imply GPU completion. Each rank advances its own schedule; equal step numbers do not imply globally synchronized work. Do not add per-step device barriers.

Reject a second active session. Stop closes the native profiler and submits remaining completed exports, without waiting for background parsing/compression. Return collection state and artifact-processing state separately. On inference errors, finalize collection without masking the original error. Clear session attachment state so a later start creates a new profiler. Disabled profiling creates no profiler, files or additional synchronization.

Add native record_function ranges where useful for scheduling, prefill, recurrent execution, sampling and KV transfer. Keep phase names consistent across offline and serving paths.

Coordinating a PD collection window

Use explicit session start/stop for a trace covering the same requests across P and D. Construct the native profiler without a step schedule for this window; the existing wait/warmup/active/repeat parameters apply to step-scheduled collection only. The Python start operation accepts an optional schedule: omitting it for an explicitly controlled session means continuous recording until stop. CLI automatic capture uses the configured step schedule. Do not silently ignore supplied schedule parameters or treat matching local step counts as a coordinated PD window.

Reuse the existing coordinator/worker Pipe messages:

  1. PDEngine creates a session ID and sends profile_start with rank identity and collection configuration to every participating worker.
  2. PDWorker.command() starts the local profiler in the worker execution thread and replies profile_started. It processes control messages even when idle. The coordinator handles these acknowledgments before transfer-ID routing in _message(), since profiling commands do not belong to a transfer.
  3. Once every worker has acknowledged, return success to the caller. For a dedicated request window, the caller then submits the target requests. Warm the model before starting this session if initialization is not of interest. This readiness handshake ensures coverage, not simultaneous GPU execution.
  4. Once those requests finish and their KV transfers are released, send profile_stop to all workers. Each worker finalizes its local capture, queues completed files for background processing and replies profile_stopped. This message is separate from the existing worker-shutdown stop command.
  5. Gather per-rank processing status and artifact locations. Waiting for archives is a separate operation and does not hold up inference.

Continue polling existing request and transport messages while waiting for profiling acknowledgments. Use a control timeout; on partial start failure, stop ranks that started and report the missing/failed ranks. Match acknowledgments by session ID so late replies cannot affect a later session.

For an already busy service, start/stop captures a shared overlapping interval; it does not automatically isolate individual requests or drain production traffic. Record local start/stop times and mark requests cut by the boundaries. Retain the existing transfer ID and available request trace ID in annotations for prefill, transfer submission/completion and decode activation. Correlate P/D traces by those IDs and rank roles. Do not infer cross-host timestamp alignment or complete NIXL/RDMA device visibility from these host-side annotations.

Per-rank collection

Each participating rank owns a profiler and exports its own recording cycles. Use distributed global rank when available. For PD workers without a process group, assign session-unique ranks at registration; a single process uses rank 0. Record rank, worker role, hostname, PID and device. Local rank alone does not uniquely identify workers across hosts.

<profile-dir>/<session-id>/
  manifest.json
  rank-00000/
    cycle-00000.tar.gz
  rank-00001/
    cycle-00000.tar.gz

Use separate temporary staging directories for each rank/cycle. The session manifest lists expected ranks, archive locations and incomplete/failed ranks. The coordinator gathers status and paths, not the large raw traces. Paths on non-shared storage include their host identity. No automatic timeline merging or cross-rank clock alignment is required.

Background parsing and packaging

When a native export finishes writing and closes its files, enqueue their paths and immutable rank/cycle metadata to one dedicated background thread per rank. Parsing, compression, verification and cleanup all execute on that thread. The inference thread does not wait for this processing. Never pass a live profiler or mutable tensor/event buffers to the background thread.

Native collection and export retain their own overhead. A background thread also shares CPU, memory and the Python GIL; this design avoids synchronous post-processing on the inference path, but does not guarantee zero contention. Keep one consumer per rank, queue paths rather than trace contents, and parse large files incrementally where practical. Keep pending work discoverable from staging directories so queue pressure does not block inference or drop exports.

For each rank/cycle:

  1. Parse the raw trace into operators.csv and summary.json: operator/kernel names, counts, CPU/device durations, units and available shape/stack/memory details. Preserve thread/stream attribution where present. Label inclusive versus self time and do not equate summed kernel duration with wall time. Missing or uncollected fields remain explicitly unavailable.
  2. Include stack and memory exports when the enabled options meet their native prerequisites. Parse their available details without enabling extra collection options implicitly.
  3. Package raw exports, parsed results and rank/cycle metadata into a temporary .tar.gz archive. Include file sizes and checksums in its inventory.
  4. Reopen the archive, verify its contents and checksums, and atomically publish its final filename. Record the archive path and processing status.
  5. Delete that job's standalone raw files and intermediate parsed files only after verification and publication. Both original and parsed data remain inside the compressed archive.

On parsing or packaging failure, retain the staging files and report the error. Cleanup touches only registered files in that job's staging directory. Report cleanup failure separately so it can be retried without repeating collection.

wait_for_profile_artifacts() waits for queued processing and returns archive paths or errors. Normal CLI exit and graceful worker shutdown drain and join the background thread after inference stops. An interrupted rank retains staged data and is marked incomplete; collection completion alone does not imply archives are ready. Disk exhaustion and missing rank results must be visible in status.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    RFCRequests for comments on major architectural changes or design choices

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions