Building high-performance AI inference systems, distributed LLM serving infrastructure, and GPU-optimized real-time pipelines.
- π₯ Focused on low-latency AI inference and GPU optimization
- π Experienced with TensorRT-LLM, CUDA, vLLM, ONNX Runtime
- π§ Interested in distributed AI systems and runtime performance engineering
- π Passionate about quantization, CUDA Graphs, and scalable AI infrastructure
- π MS in Robotics & Autonomous Systems @ Arizona State University
Inference & GPU Optimization
TensorRT β’ TensorRT-LLM β’ CUDA β’ CUDA Graphs β’ ONNX Runtime β’ vLLM
Machine Learning
PyTorch β’ Transformers β’ FP16 β’ BF16 β’ FP8 β’ INT8 Quantization
Systems & Infrastructure
Docker β’ Linux β’ Async Execution β’ Multi-threading
Profiling & Performance
Nsight Systems β’ Nsight Compute β’ PyTorch Profiler β’ CUDA Events
- High-throughput LLM serving
- CUDA kernel optimization
- Low-latency inference systems
- Distributed AI infrastructure
- GPU runtime performance engineering
