Skip to content
#

torch-distributed

Here are 3 public repositories matching this topic...

Language: All
Filter by language

Simulated Multi-GPU inference engine implementing Megatron-style Tensor Parallelism, GPipe Pipeline Parallelism, and KV Cache Sharding from scratch. Features bit-identical fidelity validation on Qwen2-0.5B weights and analytical communication cost modeling.

  • Updated Jul 29, 2026
  • Python

Improve this page

Add a description, image, and links to the torch-distributed topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the torch-distributed topic, visit your repo's landing page and select "manage topics."

Learn more