Skip to content
View maitimeraki's full-sized avatar
🏠
Working from home
🏠
Working from home

Block or report maitimeraki

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
maitimeraki/README.md

πŸ‘‹ Hey, I'm Anupam Maiti

LLM Systems Engineer | Distributed Inference Specialist | Performance Optimization Enthusiast

I build scalable, high-performance systems that push the boundaries of what's possible with large language models. My focus is on bridging the gap between cutting-edge ML research and production-grade infrastructure.


🎯 What I Do

I design and implement distributed systems for efficient LLM inference, tackling challenges like:

  • Sparse Expert Routing β€” Optimizing mixture-of-experts (MoE) models for inference
  • Distributed Caching Strategies β€” Smart prefetching and cache management for latency reduction
  • Request Scheduling & Batching β€” Maximizing throughput while maintaining response quality
  • Model Quantization β€” Reducing model size without sacrificing accuracy
  • Performance-First Architecture β€” Building systems that scale from edge devices to data centers

Current Focus: Sparse LLM v2

Building a production-grade inference engine that intelligently routes requests to specialized experts, manages distributed storage efficiently, and maintains low latency at scale.

Key Components:

  • 🧠 Expert Routing: Intelligent predictor for selecting optimal experts
  • πŸ’Ύ Multi-tier Storage: Local SSD, network backends, and intelligent prefetching
  • πŸ”„ Request Scheduling: Batch optimization with dynamic queue management
  • ⚑ Performance Optimization: Quantization, caching, and pipeline prefetching

πŸ› οΈ Technical Stack

Languages: Python Β· TypeScript Β· JavaScript Β· Java Β· HTML5

ML & Systems: Python PyTorch FastAPI

Backend & DevOps: Express.js AWS Firebase

Frontend: React TypeScript TailwindCSS Redux


πŸ’‘ Areas of Expertise

Domain Focus
LLM Systems Inference optimization, model serving, distributed execution
System Design Scalability, performance profiling, bottleneck identification
Backend Architecture API design, microservices, request handling pipelines
Frontend Development React, state management, responsive design
DevOps & Cloud AWS, deployment automation, monitoring

πŸš€ Featured Projects

Sparse LLM v2 β€” Efficient Distributed Inference

Intelligent LLM inference engine with expert routing and smart caching

  • Expert router with performance prediction
  • Multi-tier storage (local SSD + network backends)
  • Dynamic request scheduling and batching
  • Quantization support for model compression
  • Prefetching pipeline for latency reduction

Full-Stack Applications

Experience building end-to-end web applications with modern tech stacks, from backend APIs to responsive frontends


πŸ“ˆ My Philosophy

Optimize for impact, not perfection. The best system is one that ships and learns. I focus on identifying real bottlenecks, solving them elegantly, and measuring the results.

  • Performance-First Mindset β€” Every millisecond matters in inference
  • System Thinking β€” Understanding end-to-end data flow and trade-offs
  • Production Readiness β€” Building for scale and reliability from day one
  • Continuous Learning β€” Staying ahead in the rapidly evolving ML infrastructure space

🌐 Let's Connect

LinkedIn GitHub


πŸ“Š GitHub Stats

GitHub Stats Top Languages

Open to collaborations on distributed systems, ML infrastructure, and performance optimization projects.

Pinned Loading

  1. ML-LIFECYCLE-SYSTEM ML-LIFECYCLE-SYSTEM Public

    Production-Grade ML Model Lifecycle Management System.

    Python

  2. MailQuell MailQuell Public

    Production-grade Gmail automation and inbox intelligence platform.

    JavaScript

  3. Code Code Public

    Self-improving agent orchestration system with autonomous loop execution, parallel agent spawning, and multi-LLM support.

    Python

  4. END-TO-END-FINE-TUNING END-TO-END-FINE-TUNING Public

    Jupyter Notebook

  5. SOPHISTICATED-AGENT SOPHISTICATED-AGENT Public

    Jupyter Notebook

  6. QuillSense QuillSense Public

    Personal AI Suggestion System designed to analyze user's writing history and patterns, and provide tailored writing suggestions.