[SIGCOMM 2023] Lightning: A Reconfigurable Photonic-Electronic SmartNIC for Fast and Energy-Efficient Inference
-
Updated
Nov 17, 2023 - Verilog
[SIGCOMM 2023] Lightning: A Reconfigurable Photonic-Electronic SmartNIC for Fast and Energy-Efficient Inference
[Long Term Support] [SIGCOMM 2023] Lightning: A Reconfigurable Photonic-Electronic SmartNIC for Fast and Energy-Efficient Inference
The official repo for the paper "Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching"
Official repo for the ACSOS 2021 paper on how to manage many deep learning models at the edge!
NexusKV is a model-state intelligence layer for inference systems.
Locus is an engine-neutral inference control plane for compute and model-state placement.
Local inference serving with adaptive batching, benchmark sweeps, and regression gates for batching tradeoffs.
Reactive LLM inference.
Reproducible Apple-Silicon (Apple MLX) deployment and one-token verification kit for the FreeToken edge-native MoE serving API — pinned, loopback-only, schema-validated evidence.
GPU-aware LLM serving runtime with async scheduling, dynamic batching, streaming, admission control, metrics, and mock/llama.cpp/vLLM backends.
Machine-readable companion to the IEEE OJ-CS survey 'Semantic Caching and Response Reuse for Large Language Model Services: A Survey' (Chukkapalli, Mishra, Naik, 2026): 21-work evidence matrix, systematic-search log, proposed benchmark trace schema, stdlib-only contract validator, and CPU pilot. Code MIT; data CC-BY-4.0.
GSM artifact for exact online GNN inference serving with degree-based scheduling and memory-aware batch division.
Your GPU is not 90% busy. truthscale measures what each node can actually deliver, reports the gap, and makes that the signal you scale on.
Put AI compute in the 82 million houses that already have a grid connection. An exact-arithmetic feasibility study of residential AI inference as an alternative to the data center buildout.
cost-aware-inference-cluster: FastAPI inference-serving prototype with Redis queues, worker heartbeats, dynamic batching, autoscaling simulation, and local Docker Compose benchmark evidence.
Inference gateway for batching AI model requests through bounded micro-batching, queue backpressure, per-request result routing, and reproducible performance evaluation.
Add a description, image, and links to the inference-serving topic page so that developers can more easily learn about it.
To associate your repository with the inference-serving topic, visit your repo's landing page and select "manage topics."