Here are
22 public repositories
matching this topic...
Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.
Updated
Aug 23, 2026
Python
A native .NET LLM inference engine for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/Linux with full GPU capability
A PyTorch native library for training speculative decoding models
Updated
Aug 20, 2026
Python
DeepSeek V4 Flash 284B on AMD Strix Halo (gfx1151) — up to 32 tok/s decode & ~250 tok/s prefill via ROCmFPX, DSpark & ROCm 7.2
Optimized two-node DGX Spark deployment recipe for DeepSeek V4 Flash with vLLM, DSpark, and NVFP4 KV cache
Updated
Aug 12, 2026
Shell
DeepSeek-V4-Flash-0731 on 8x RTX 3090 (SM86) with vLLM — verified FP8 serving, benchmarks, build guide, and reproducible release.
Updated
Aug 9, 2026
Python
Anemll + MiaAI-Lab dual DGX Spark DSV4F+DSpark on vLLM 0.25.2.dev0 — 100% abliterated weights, Spec-Kit one-shot agent package. Full credit Anemll/MiaAI.
Updated
Aug 17, 2026
Python
DeepSeek-V4-Flash + DSpark speculative decoding on a pair of NVIDIA DGX Sparks (vLLM TP=2 over RoCE) — tuned recipe, overlays that halve multi-turn TTFT, contamination-guarded benchmarks, ops runbook
Updated
Jul 3, 2026
Python
Lossless inference speedup benchmarking suite for local LLMs using DeepSeek DSpark speculative draft heads.
Updated
Jun 29, 2026
Python
MiniMax M3 NVFP4 on 4x DGX Spark with NVIDIA DSpark, native multi-node vLLM TP=4, reasoning, and tool calling
Updated
Jul 25, 2026
Shell
DSpark-style speculative decoding for OCR VLMs on Apple Silicon (MLX). Faster local OCR for DeepSeek-OCR-2, GLM-OCR, Unlimited-OCR.
Updated
Jul 28, 2026
Python
DeepSeek V4 Flash 0731 on one RTX PRO 6000/WSL2: CUDA 13.3.1, vLLM-Moet, native OpenCode tools, DSpark-5 greedy, 128K FP8 KV, plus clean-MAX fallback.
Updated
Aug 10, 2026
JavaScript
Agent-Level Speculative Orchestration & Formal Dual-Engine Code Generation (Rust, FastMCP, DeepSeek)
Updated
Aug 24, 2026
Rust
Universal Dual-Engine & Speculative Curation Configuration Kit for Antigravity, Cursor, Claude Code, Windsurf & Grok (FastMCP)
Updated
Aug 24, 2026
Python
SGLang patches for RTX 4090/SM89: GLM-5.2 FP8/256K, NVFP4 Marlin BF16, and Qwen3.8 DSpark.
Updated
Aug 15, 2026
Python
Running Cogni-Brain on DGX Spark · Nemotron-3.5-Lightning-30B-A3B-NVFP4 + DSpark (1M context)
Updated
Aug 17, 2026
Python
Exactness-first fast iteration lab for improving DSpark inference speed
Updated
Aug 10, 2026
Python
Verified two-node GB10 DeepSeek-V4-Flash-0731 DSpark Graph-8 deployment and benchmarks
Updated
Aug 4, 2026
Python
Source-only offline deployment, testing, and operations toolkit for DeepSeek-V4-Flash-0731 on 4x/8x NVIDIA A100 GPUs, pinned to a reviewed community vLLM R1 stack.
Updated
Aug 24, 2026
HTML
DeepSeek-V4-Flash-0731 on 4x RTX PRO 6000 Blackwell: DSpark K5, 1M context, uv/no-container; C1/C4/C8 = 212/287/390 output tok/s.
Updated
Aug 24, 2026
Shell
Improve this page
Add a description, image, and links to the
dspark
topic page so that developers can more easily learn about it.
Curate this topic
Add this topic to your repo
To associate your repository with the
dspark
topic, visit your repo's landing page and select "manage topics."
Learn more
You can’t perform that action at this time.