Implementation status
馃毀 An implementation is available in PR #72, currently open and not yet merged. It provides CLI/Python/HTTP profiling controls, coordinated PD capture, and per-rank background parsing and verified archives. The PR description lists activation methods, parameter defaults, the collection workflow, artifacts, and validation results.
The implementation in commit ea4a4bb simplifies activation to explicit start calls, removes ProfileConfig.enabled, and collects native events without custom phase/request annotations. The design below records the original proposal; its enabled field, custom annotations, and session-ID directory layout should not be treated as the current interface. Artifacts use timestamped rank directories. Profiling, PD, and serving CPU regression tests report 76 passed, 9 skipped. GPU/CUDA Graph and multi-GPU PD/NIXL validation remains outstanding; implementation availability does not mean all RFC acceptance requirements are complete.
Objective
Add a configurable PyTorch Profiler interface for inference in vllm-rlt. Support operator/kernel timelines, shapes, stacks, memory and FLOPs, with bounded collection windows and start/stop control. Collect independently by rank, then parse, compress and clean up exports on a background thread.
Use native PyTorch collection and scheduling semantics. Configuration options are independent; there are no named collection presets. This proposal exposes core options rather than reproducing the entire native profiler API.
Core configuration
Python and CLI use the same collection options. The original proposed configuration is listed below; an implementation is available in #72, with interface differences noted in the implementation status above.
| Python field |
CLI argument |
Default |
Meaning |
enabled |
--profile |
False |
Enable profiling |
output_dir |
--profile-dir PATH |
required when enabled |
Root directory for rank artifacts |
activities |
--profile-activities cpu,cuda |
CPU for CPU inference; CPU and CUDA for CUDA inference |
Select collected activities |
record_shapes |
--profile-record-shapes |
False |
Record operator input shapes |
with_stack |
--profile-with-stack |
False |
Record source stacks |
profile_memory |
--profile-memory |
False |
Record tensor allocations and frees |
with_flops |
--profile-with-flops |
False |
Estimate FLOPs for supported operators |
wait |
--profile-wait N |
0 |
Wait steps per cycle |
warmup |
--profile-warmup N |
1 |
Profiler warmup steps per cycle |
active |
--profile-active N |
10 |
Recorded steps per cycle |
repeat |
--profile-repeat N |
1 |
Cycle count; zero repeats until stopped |
The schedule defaults are vllm-rlt choices for a bounded capture, not PyTorch constructor defaults. Forward these four fields to torch.profiler.schedule. Validate nonnegative wait/warmup/repeat and positive active. Validate requested activities and the output directory before starting collection. Unsupported activities produce an error rather than silently falling back.
Keep native meanings and prerequisites from the PyTorch Profiler reference. For example, FLOPs are available only for supported operators, and detailed memory timeline export requires shapes, stacks and memory collection together. Enabling memory alone still records allocation events; it must not silently enable the other options.
Experimental configuration, module hierarchy, execution-trace observers, custom trace IDs, cross-cycle event accumulation, custom factories/callbacks, post-processing timeouts and advanced schedule modifiers are not exposed by this interface. They can be considered later if a concrete use requires them.
CLI example supported by the implementation PR:
vllm-rlt --model /path/to/model --device cuda \
--prompt "Explain matrix multiplication." --max-tokens 32 \
--profile --profile-dir /tmp/rlt-prof \
--profile-activities cpu,cuda \
--profile-record-shapes --profile-with-stack --profile-memory \
--profile-wait 2 --profile-warmup 1 --profile-active 10 --profile-repeat 1
Interface and lifecycle
Use a shared vllm_rlt/profiling.py module for configuration, session control and export handling. The runtime owns the native profiler and its trace callback. Provide start_profile(config), stop_profile(), profile_status() and wait_for_profile_artifacts() through the Python integration. Exact placement on the LLM/engine facade follows existing lifecycle ownership.
Offline CLI collection starts after model loading and runtime initialization. Python callers and serving controls can start collection around the workload of interest. Serving routes start/stop to the execution worker, not merely the HTTP handler. PD dispatches the same configuration and session ID to all workers and gathers their status and artifact locations.
For step-scheduled collection, advance profiler.step() exactly once after each completed LLMEngine.step() in the ordinary engine. In PD, the P worker uses prefill_step() instead; advance after a prefill iteration that submitted work, and after a D worker's engine.step(). Do not count idle polling as work or advance again at the coordinator. Include transfer progression within the worker's profiling context. An engine step is not an output token, request or recurrent loop, and does not imply GPU completion. Each rank advances its own schedule; equal step numbers do not imply globally synchronized work. Do not add per-step device barriers.
Reject a second active session. Stop closes the native profiler and submits remaining completed exports, without waiting for background parsing/compression. Return collection state and artifact-processing state separately. On inference errors, finalize collection without masking the original error. Clear session attachment state so a later start creates a new profiler. Disabled profiling creates no profiler, files or additional synchronization.
Add native record_function ranges where useful for scheduling, prefill, recurrent execution, sampling and KV transfer. Keep phase names consistent across offline and serving paths.
Coordinating a PD collection window
Use explicit session start/stop for a trace covering the same requests across P and D. Construct the native profiler without a step schedule for this window; the existing wait/warmup/active/repeat parameters apply to step-scheduled collection only. The Python start operation accepts an optional schedule: omitting it for an explicitly controlled session means continuous recording until stop. CLI automatic capture uses the configured step schedule. Do not silently ignore supplied schedule parameters or treat matching local step counts as a coordinated PD window.
Reuse the existing coordinator/worker Pipe messages:
PDEngine creates a session ID and sends profile_start with rank identity and collection configuration to every participating worker.
PDWorker.command() starts the local profiler in the worker execution thread and replies profile_started. It processes control messages even when idle. The coordinator handles these acknowledgments before transfer-ID routing in _message(), since profiling commands do not belong to a transfer.
- Once every worker has acknowledged, return success to the caller. For a dedicated request window, the caller then submits the target requests. Warm the model before starting this session if initialization is not of interest. This readiness handshake ensures coverage, not simultaneous GPU execution.
- Once those requests finish and their KV transfers are released, send
profile_stop to all workers. Each worker finalizes its local capture, queues completed files for background processing and replies profile_stopped. This message is separate from the existing worker-shutdown stop command.
- Gather per-rank processing status and artifact locations. Waiting for archives is a separate operation and does not hold up inference.
Continue polling existing request and transport messages while waiting for profiling acknowledgments. Use a control timeout; on partial start failure, stop ranks that started and report the missing/failed ranks. Match acknowledgments by session ID so late replies cannot affect a later session.
For an already busy service, start/stop captures a shared overlapping interval; it does not automatically isolate individual requests or drain production traffic. Record local start/stop times and mark requests cut by the boundaries. Retain the existing transfer ID and available request trace ID in annotations for prefill, transfer submission/completion and decode activation. Correlate P/D traces by those IDs and rank roles. Do not infer cross-host timestamp alignment or complete NIXL/RDMA device visibility from these host-side annotations.
Per-rank collection
Each participating rank owns a profiler and exports its own recording cycles. Use distributed global rank when available. For PD workers without a process group, assign session-unique ranks at registration; a single process uses rank 0. Record rank, worker role, hostname, PID and device. Local rank alone does not uniquely identify workers across hosts.
<profile-dir>/<session-id>/
manifest.json
rank-00000/
cycle-00000.tar.gz
rank-00001/
cycle-00000.tar.gz
Use separate temporary staging directories for each rank/cycle. The session manifest lists expected ranks, archive locations and incomplete/failed ranks. The coordinator gathers status and paths, not the large raw traces. Paths on non-shared storage include their host identity. No automatic timeline merging or cross-rank clock alignment is required.
Background parsing and packaging
When a native export finishes writing and closes its files, enqueue their paths and immutable rank/cycle metadata to one dedicated background thread per rank. Parsing, compression, verification and cleanup all execute on that thread. The inference thread does not wait for this processing. Never pass a live profiler or mutable tensor/event buffers to the background thread.
Native collection and export retain their own overhead. A background thread also shares CPU, memory and the Python GIL; this design avoids synchronous post-processing on the inference path, but does not guarantee zero contention. Keep one consumer per rank, queue paths rather than trace contents, and parse large files incrementally where practical. Keep pending work discoverable from staging directories so queue pressure does not block inference or drop exports.
For each rank/cycle:
- Parse the raw trace into
operators.csv and summary.json: operator/kernel names, counts, CPU/device durations, units and available shape/stack/memory details. Preserve thread/stream attribution where present. Label inclusive versus self time and do not equate summed kernel duration with wall time. Missing or uncollected fields remain explicitly unavailable.
- Include stack and memory exports when the enabled options meet their native prerequisites. Parse their available details without enabling extra collection options implicitly.
- Package raw exports, parsed results and rank/cycle metadata into a temporary
.tar.gz archive. Include file sizes and checksums in its inventory.
- Reopen the archive, verify its contents and checksums, and atomically publish its final filename. Record the archive path and processing status.
- Delete that job's standalone raw files and intermediate parsed files only after verification and publication. Both original and parsed data remain inside the compressed archive.
On parsing or packaging failure, retain the staging files and report the error. Cleanup touches only registered files in that job's staging directory. Report cleanup failure separately so it can be retried without repeating collection.
wait_for_profile_artifacts() waits for queued processing and returns archive paths or errors. Normal CLI exit and graceful worker shutdown drain and join the background thread after inference stops. An interrupted rank retains staged data and is marked incomplete; collection completion alone does not imply archives are ready. Disk exhaustion and missing rank results must be visible in status.
Implementation status
馃毀 An implementation is available in PR #72, currently open and not yet merged. It provides CLI/Python/HTTP profiling controls, coordinated PD capture, and per-rank background parsing and verified archives. The PR description lists activation methods, parameter defaults, the collection workflow, artifacts, and validation results.
The implementation in commit
ea4a4bbsimplifies activation to explicit start calls, removesProfileConfig.enabled, and collects native events without custom phase/request annotations. The design below records the original proposal; itsenabledfield, custom annotations, and session-ID directory layout should not be treated as the current interface. Artifacts use timestamped rank directories. Profiling, PD, and serving CPU regression tests report 76 passed, 9 skipped. GPU/CUDA Graph and multi-GPU PD/NIXL validation remains outstanding; implementation availability does not mean all RFC acceptance requirements are complete.Objective
Add a configurable PyTorch Profiler interface for inference in vllm-rlt. Support operator/kernel timelines, shapes, stacks, memory and FLOPs, with bounded collection windows and start/stop control. Collect independently by rank, then parse, compress and clean up exports on a background thread.
Use native PyTorch collection and scheduling semantics. Configuration options are independent; there are no named collection presets. This proposal exposes core options rather than reproducing the entire native profiler API.
Core configuration
Python and CLI use the same collection options. The original proposed configuration is listed below; an implementation is available in #72, with interface differences noted in the implementation status above.
enabled--profileFalseoutput_dir--profile-dir PATHactivities--profile-activities cpu,cudarecord_shapes--profile-record-shapesFalsewith_stack--profile-with-stackFalseprofile_memory--profile-memoryFalsewith_flops--profile-with-flopsFalsewait--profile-wait N0warmup--profile-warmup N1active--profile-active N10repeat--profile-repeat N1The schedule defaults are vllm-rlt choices for a bounded capture, not PyTorch constructor defaults. Forward these four fields to
torch.profiler.schedule. Validate nonnegative wait/warmup/repeat and positive active. Validate requested activities and the output directory before starting collection. Unsupported activities produce an error rather than silently falling back.Keep native meanings and prerequisites from the PyTorch Profiler reference. For example, FLOPs are available only for supported operators, and detailed memory timeline export requires shapes, stacks and memory collection together. Enabling memory alone still records allocation events; it must not silently enable the other options.
Experimental configuration, module hierarchy, execution-trace observers, custom trace IDs, cross-cycle event accumulation, custom factories/callbacks, post-processing timeouts and advanced schedule modifiers are not exposed by this interface. They can be considered later if a concrete use requires them.
CLI example supported by the implementation PR:
vllm-rlt --model /path/to/model --device cuda \ --prompt "Explain matrix multiplication." --max-tokens 32 \ --profile --profile-dir /tmp/rlt-prof \ --profile-activities cpu,cuda \ --profile-record-shapes --profile-with-stack --profile-memory \ --profile-wait 2 --profile-warmup 1 --profile-active 10 --profile-repeat 1Interface and lifecycle
Use a shared
vllm_rlt/profiling.pymodule for configuration, session control and export handling. The runtime owns the native profiler and its trace callback. Providestart_profile(config),stop_profile(),profile_status()andwait_for_profile_artifacts()through the Python integration. Exact placement on the LLM/engine facade follows existing lifecycle ownership.Offline CLI collection starts after model loading and runtime initialization. Python callers and serving controls can start collection around the workload of interest. Serving routes start/stop to the execution worker, not merely the HTTP handler. PD dispatches the same configuration and session ID to all workers and gathers their status and artifact locations.
For step-scheduled collection, advance
profiler.step()exactly once after each completedLLMEngine.step()in the ordinary engine. In PD, the P worker usesprefill_step()instead; advance after a prefill iteration that submitted work, and after a D worker'sengine.step(). Do not count idle polling as work or advance again at the coordinator. Include transfer progression within the worker's profiling context. An engine step is not an output token, request or recurrent loop, and does not imply GPU completion. Each rank advances its own schedule; equal step numbers do not imply globally synchronized work. Do not add per-step device barriers.Reject a second active session. Stop closes the native profiler and submits remaining completed exports, without waiting for background parsing/compression. Return collection state and artifact-processing state separately. On inference errors, finalize collection without masking the original error. Clear session attachment state so a later start creates a new profiler. Disabled profiling creates no profiler, files or additional synchronization.
Add native
record_functionranges where useful for scheduling, prefill, recurrent execution, sampling and KV transfer. Keep phase names consistent across offline and serving paths.Coordinating a PD collection window
Use explicit session start/stop for a trace covering the same requests across P and D. Construct the native profiler without a step schedule for this window; the existing wait/warmup/active/repeat parameters apply to step-scheduled collection only. The Python start operation accepts an optional schedule: omitting it for an explicitly controlled session means continuous recording until stop. CLI automatic capture uses the configured step schedule. Do not silently ignore supplied schedule parameters or treat matching local step counts as a coordinated PD window.
Reuse the existing coordinator/worker Pipe messages:
PDEnginecreates a session ID and sendsprofile_startwith rank identity and collection configuration to every participating worker.PDWorker.command()starts the local profiler in the worker execution thread and repliesprofile_started. It processes control messages even when idle. The coordinator handles these acknowledgments before transfer-ID routing in_message(), since profiling commands do not belong to a transfer.profile_stopto all workers. Each worker finalizes its local capture, queues completed files for background processing and repliesprofile_stopped. This message is separate from the existing worker-shutdownstopcommand.Continue polling existing request and transport messages while waiting for profiling acknowledgments. Use a control timeout; on partial start failure, stop ranks that started and report the missing/failed ranks. Match acknowledgments by session ID so late replies cannot affect a later session.
For an already busy service, start/stop captures a shared overlapping interval; it does not automatically isolate individual requests or drain production traffic. Record local start/stop times and mark requests cut by the boundaries. Retain the existing transfer ID and available request trace ID in annotations for prefill, transfer submission/completion and decode activation. Correlate P/D traces by those IDs and rank roles. Do not infer cross-host timestamp alignment or complete NIXL/RDMA device visibility from these host-side annotations.
Per-rank collection
Each participating rank owns a profiler and exports its own recording cycles. Use distributed global rank when available. For PD workers without a process group, assign session-unique ranks at registration; a single process uses rank 0. Record rank, worker role, hostname, PID and device. Local rank alone does not uniquely identify workers across hosts.
Use separate temporary staging directories for each rank/cycle. The session manifest lists expected ranks, archive locations and incomplete/failed ranks. The coordinator gathers status and paths, not the large raw traces. Paths on non-shared storage include their host identity. No automatic timeline merging or cross-rank clock alignment is required.
Background parsing and packaging
When a native export finishes writing and closes its files, enqueue their paths and immutable rank/cycle metadata to one dedicated background thread per rank. Parsing, compression, verification and cleanup all execute on that thread. The inference thread does not wait for this processing. Never pass a live profiler or mutable tensor/event buffers to the background thread.
Native collection and export retain their own overhead. A background thread also shares CPU, memory and the Python GIL; this design avoids synchronous post-processing on the inference path, but does not guarantee zero contention. Keep one consumer per rank, queue paths rather than trace contents, and parse large files incrementally where practical. Keep pending work discoverable from staging directories so queue pressure does not block inference or drop exports.
For each rank/cycle:
operators.csvandsummary.json: operator/kernel names, counts, CPU/device durations, units and available shape/stack/memory details. Preserve thread/stream attribution where present. Label inclusive versus self time and do not equate summed kernel duration with wall time. Missing or uncollected fields remain explicitly unavailable..tar.gzarchive. Include file sizes and checksums in its inventory.On parsing or packaging failure, retain the staging files and report the error. Cleanup touches only registered files in that job's staging directory. Report cleanup failure separately so it can be retried without repeating collection.
wait_for_profile_artifacts()waits for queued processing and returns archive paths or errors. Normal CLI exit and graceful worker shutdown drain and join the background thread after inference stops. An interrupted rank retains staged data and is marked incomplete; collection completion alone does not imply archives are ready. Disk exhaustion and missing rank results must be visible in status.