Outcome and baseline
Extend AtnAgent-to-FfnAgent and FfnAgent-to-FfnAgent Fabric traffic across hosts.
Instance-to-AtnAgent Transport remains a host-local, rank-local path. The
current symmetric NVSHMEM Fabric uses a fixed PE world, Device-side publication,
static placement, and canonical generation failure on one host.
This issue covers actual two-host feasibility evidence and a host-aware target
design. Production cross-host bootstrap and serving remain follow-on work.
Suggested implementation route
- Specify and qualify the two-host GPU/NIC topology, peer access, network,
NVSHMEM/IBGDA stack, bootstrap, and permissions. Record the concrete
environment and compare GPU-initiated and CPU-proxy paths.
- Run a raw two-host prototype that verifies publication ordering, remote
completion, and Graph replay. Measure communication and NIC/QP/completion
costs and exercise participant failure and shutdown. Retain raw evidence
for both the working path and any transport limitation.
- Propose host identity, NIC/GPU topology, placement, bootstrap, lifecycle,
failure, and shutdown contracts. Evaluate whether generation-wide fail-stop
remains appropriate and write a scoped host-aware control-plane plan.
- After accepting the plan, implement host-aware configuration and Fabric
activation while preserving local Transport ownership and explicit process
supervision.
- Qualify multi-host numerical delivery, Graph replay, startup/readiness,
failure propagation, and safe shutdown under the selected deployment.
Discovery completion evidence
- Reproducible qualification details for a real two-host environment.
- Raw ordering, Graph, completion-cost, and failure/shutdown prototype evidence.
- A reviewed host-aware design recommendation or an evidence-backed no-go.
An unavailable second host or NIC leaves these required experiments incomplete;
it is not evidence of transport infeasibility. Existing Fabric contracts and
Timeline observability benefit this work. A speculative transport-backend
registry is outside scope.
References
Outcome and baseline
Extend AtnAgent-to-FfnAgent and FfnAgent-to-FfnAgent Fabric traffic across hosts.
Instance-to-AtnAgent Transport remains a host-local, rank-local path. The
current symmetric NVSHMEM Fabric uses a fixed PE world, Device-side publication,
static placement, and canonical generation failure on one host.
This issue covers actual two-host feasibility evidence and a host-aware target
design. Production cross-host bootstrap and serving remain follow-on work.
Suggested implementation route
NVSHMEM/IBGDA stack, bootstrap, and permissions. Record the concrete
environment and compare GPU-initiated and CPU-proxy paths.
completion, and Graph replay. Measure communication and NIC/QP/completion
costs and exercise participant failure and shutdown. Retain raw evidence
for both the working path and any transport limitation.
failure, and shutdown contracts. Evaluate whether generation-wide fail-stop
remains appropriate and write a scoped host-aware control-plane plan.
activation while preserving local Transport ownership and explicit process
supervision.
failure propagation, and safe shutdown under the selected deployment.
Discovery completion evidence
An unavailable second host or NIC leaves these required experiments incomplete;
it is not evidence of transport infeasibility. Existing Fabric contracts and
Timeline observability benefit this work. A speculative transport-backend
registry is outside scope.
References