Interested in GPU kernels, distributed training, and inference optimization
lately working on:
- distributed training + inference profiler, profiling which CUDA kernels are blocking NCCL collectives
- agentic triton kernel correctness tester catching bugs that naive checks miss for LLM generated kernels



