Software engineer at Anthropic, passionate about AI infrastructure and high-performance systems.
- Scaling LLM inference systems
- GitHub: @jia-gao
π‘ Open to collaborations on open-source AI infrastructure and ML systems projects!
Software engineer at Anthropic, passionate about AI infrastructure and high-performance systems.
π‘ Open to collaborations on open-source AI infrastructure and ML systems projects!
htop for GPU pods on Kubernetes β per-pod GPU utilization, memory, temperature, power, and waste detection
Forked from vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Python
Forked from vllm-project/semantic-router
Intelligent Mixture-of-Models Router for Efficient LLM Inference
Python
Forked from sgl-project/sglang
SGLang is a fast serving framework for large language models and vision language models.
Python