Documentation · Installation · Examples
TileFoundry is a tile-based, agentic platform for automatic high-performance program generation across hardware.
- 08/2026 🎉: TileFoundry 0.0.1 is on PyPI — the first public release.
- 08/2026 📦: Four worked examples added — Qwen3-1.7B (tilelang), Qwen3.5-35B-A3B (tilelang), MiniCPM3-4B (CuTeDSL) and granite-4.0-h-small (CUDA C) — each one a real agent run kept whole, with the decode throughput it measured.
TileFoundry needs Python 3.12 or newer.
pip install tilefoundryCheck the install — it prints the commands an agent will ask:
tilefoundryRunning a model an agent generates additionally needs one NVIDIA GPU and the checkpoint already on disk.
There is no API to learn first. Give your coding agent this, with a checkpoint directory of your own:
Get real tokens out of Qwen3-1.7B on TileFoundry, and make it fast.
Weights and config: <checkpoint directory>
Backend: tilelang.
Everything about TileFoundry is to be asked of the `tilefoundry` command -- do not
ask a person, do not go looking elsewhere. The model itself is yours to research.
Done when this runs from outside, prints the continuation, and reports a
tokens-per-second number measured over the whole generation:
python run.py --prompt "Write a detailed explanation of how a GPU executes a matrix multiplication." --max-new-tokens 2048
Measure over a long generation -- 2048 new tokens, more than 2000 characters of
text. A 32-token sample is too short for the number to mean anything.
That is the whole input — nothing under it is written by hand.
Claude Opus 5 at xhigh reasoning effort ran this prompt for 2.1 hours with no
interaction, and reached 612 tok/s on one H200. What it wrote is
examples/qwen3_1_7b-tilelang/.
This project is licensed under the MIT License.