Fine-tuning small language models (SLMs), engineering agentic systems, and optimizing LLM inference.
-
Y-Intelligence
- 🌳 Forest Hills, Queens
- https://y-intelligence.vercel.app/
Pinned Loading
-
tunnel-engine
tunnel-engine PublicA production-grade LLM infra engine that uses vLLM for fast inference, LMCache for smart context caching, and LiteLM to handle multi-model routing and load balancing.
Python 4
-
Tachikoma-Observatory
Tachikoma-Observatory PublicA lightweight evaluation framework for testing Small Language Models (SLMs) on precise tool-calling capabilities to ensure they are production-ready.
Python 2
-
-
Medical-QA-Agent
Medical-QA-Agent PublicThis project demonstrates how to create a conversational agent for medical question-answering using a local LLM Mistral and Haystack's RAG pipeline. Unlike typical setups that rely on cloud-based …
Jupyter Notebook 1
If the problem persists, check the GitHub status page or contact support.




