Skip to content

chore: release main - #11

Open
github-actions[bot] wants to merge 1 commit into
mainfrom
release-please--branches--main
Open

github-actions[bot] wants to merge 1 commit into
mainfrom
release-please--branches--main

Conversation

@github-actions

@github-actions github-actions Bot commented Apr 26, 2026 •

Copy link
Copy Markdown

🤖 I have created a release beep boop

higgs-bench: 1.0.3

Dependencies

higgs: 2.0.0

2.0.0 (2026-09-19)

⚠ BREAKING CHANGES

  • higgs start is now config/profile-only, attach is a strict daemon dashboard, shellenv/exec fail fast on invalid or unreachable targets, and exact local model matches now take precedence over regex routes.

Features

  • add --profile CLI flag, TUI routing tab, and config visibility (#45) (95dda05)
  • add higgs exec -- <command> subcommand (#49) (b61dbf5)
  • add higgs run -- <command> subcommand (#47) (72339d2)
  • add name field to ModelConfig (#43) (54147e0)
  • add qwen3.5/qwen3.6 turboquant stack (4f165ee)
  • API routes — thinking mode streaming and tool call validation (5f7e62c)
  • chat: streaming tool-call deltas + Qwen-friendly arg normalisation (#164) (8f01163)
  • cli: higgs ui opens the desktop app (09a5b4c)
  • config: first-class config surface for the MLA latent KV cache (#263) (cdb3c4e)
  • desktop: Tauri dashboard app, request tracing, and richer server metrics (72d8809)
  • desktop: Tauri dashboard app, request tracing, and richer server metrics (487716f)
  • doctor: validate ServerSection fields (0250c42)
  • doctor: validate ServerSection fields (554371f)
  • engine: chunked prefill yield quantum for the batch serving loop (PR8 / ds4 P9) (#259) (247147e)
  • engine: disk-backed KV prefix store for restart-resume (PR6 / ds4 P4) (#257) (0435d11)
  • harden CLI, dashboard, routing, and MLX runtime (4dfc930)
  • harden default security posture (1b41a5c)
  • higgs: migrate from huggingface-cli to hf (#180) (aaecdb4)
  • resize routing sidebar (913e015)
  • server config — TurboQuant, thinking budget, chunked prefill settings (5b63e4c)
  • serve: stream prefill progress (llama.cpp-compatible prompt_progress) (#184) (e7b2c8a)
  • ship mlx.metallib alongside the higgs binary (#39) (deaa322)
  • tune qwen3.6 thinking defaults (6969834)
  • unified AI gateway with proxy routing and format translation (#38) (7c5668b)

Bug Fixes

  • address CI lint and test failures (0d8e259)
  • address coderabbit issues (6f10bc6)
  • address follow-up review issues (5c857ce)
  • address review comments (36c6484)
  • address review follow-ups (d2b312e)
  • backtick TcpListener in doc comment for clippy doc_markdown (94acea7)
  • cargo fmt + restore higgs crate build on PR #74 (050af95)
  • clarify metallib recovery hint (645dbb0)
  • clear lint and review blockers (9c8b878)
  • config: timeout validation and figment model merge (441884d)
  • correct comment about signal handler behavior (81cbd71)
  • default Qwen3.6 to non-thinking mode (#150) (372020e)
  • desktop: address review findings and bundle the CLI in the app (bb9a5c3)
  • doctor: bracket IPv6 host when binding port; add api_key/rate_limit positive-path tests (7177476)
  • higgs: validate config writes, and make metrics see failed requests (#272) (cfe0ab0)
  • improve Anthropic API compatibility (16ee50d)
  • improve Anthropic API compatibility (ff7e96c)
  • log warning instead of silently discarding signal handler error (6946974)
  • make pre-push pass on dust stack (2695165)
  • normalize auto_router.model to a name like routes do (1e741c5)
  • normalize auto_router.model to a name like routes do (3b52db4)
  • remove incorrect default_value on Option<u8> CLI arg — PR #75 (bf87bbd)
  • resolve auto_router model by basename when config uses full path (#41) (7220fa2)
  • restore higgs crate build on PR #75 — chat.rs usage stubs (f92d139)
  • restore pattern matching fallback in force routing mode (c1bb2c3)
  • restore pattern matching fallback in force routing mode (217249f)
  • route first streamed token through incremental detokenizer (722218d)
  • routes: close think tag on length-stopped reasoning in OpenAI path (5c2ea56)
  • satisfy clippy duration lint (20c6aae)
  • split long first doc paragraphs for clippy (c195312)
  • stabilize dust stack CI (8e629e3)
  • strip thinking tags from Anthropic route, fix force routing, add proxy usage tracking (89f8db2)
  • strip thinking tags, fix force routing, add proxy usage metrics (0ffbabb)
  • suffix float literals for new float_literal_f32_fallback lint (#230) (83029e4)
  • support clippy across toolchains (abc316d)
  • support mlx qwen3.6 smoke (7046849)
  • type alias for build_router signature assertion (d084592)
  • use 128+signal convention for signal-killed child exit codes (89b7ae0)

Performance Improvements

  • sse: pre-serialize static chunk fields for streaming routes (056f4b4)
  • sse: pre-serialize static chunk fields for streaming routes (d8e02ba)
higgs-engine: 2.0.0

2.0.0 (2026-09-19)

⚠ BREAKING CHANGES

  • higgs start is now config/profile-only, attach is a strict daemon dashboard, shellenv/exec fail fast on invalid or unreachable targets, and exact local model matches now take precedence over regex routes.

Features

  • add qwen3.5/qwen3.6 turboquant stack (4f165ee)
  • bench: context-frontier benchmark with shared cache rollback helper (PR2 / ds4 P8) (#253) (4510835)
  • bonsai-q1: packed engine scaffold with upstream MLX guard (#142) (fe43aab)
  • bonsai: run Bonsai-Q1 bits=1 on vanilla MLX via JIT qmv_fast kernels (no mlx-rs fork) (#182) (b604754)
  • chat: streaming tool-call deltas + Qwen-friendly arg normalisation (#164) (8f01163)
  • enable MiniCPM5-1B (explicit head_dim + special-token decode + tojson kwarg + tool parser) (#176) (05464e2)
  • engine: chunked prefill yield quantum for the batch serving loop (PR8 / ds4 P9) (#259) (247147e)
  • engine: disk-backed KV prefix store for restart-resume (PR6 / ds4 P4) (#257) (0435d11)
  • harden CLI, dashboard, routing, and MLX runtime (4dfc930)
  • harden default security posture (1b41a5c)
  • inference engine — chunked prefill, prefix cache, thinking budget, MTP (4fa0971)
  • models: add Gemma 3 and Gemma 4 inference support (4f301fc)
  • models: MLA latent KV cache with weight absorption for DeepSeek-V2 (PR3 / ds4 P1) (#254) (482d2ec)
  • mtp: release speculative decoding optimizations (0e5e458)
  • mtp: release speculative decoding optimizations (1d15599)
  • Qwen3.5 model architecture, TurboQuant KV cache, and model-level optimizations (1514737)
  • qwen35: Qwen3.6 MoE MTP speculative decode + checkpoint memory-safety fix (#183) (4989955)
  • serve: stream prefill progress (llama.cpp-compatible prompt_progress) (#184) (e7b2c8a)
  • thinking mode support — reasoning parser and chat template (a3cc52e)
  • tune qwen3.6 thinking defaults (6969834)

Bug Fixes

  • address coderabbit issues (6f10bc6)
  • address follow-up review issues (5c857ce)
  • address remaining draft review issues (4396e88)
  • address review items across models and engine — PR #74 (8863e2d)
  • cargo fmt + restore higgs crate build on PR #74 (050af95)
  • clear lint and review blockers (9c8b878)
  • deps: migrate to sha2 0.11 (#273) (eb817f9)
  • doc_markdown lints in detok test comments (deefbe3)
  • engine,models: cache config propagation, ndim guards, FSM ordering (37d2ddd)
  • engine,models: clippy clean lib targets on PR #74 (51→0 errors) (ebf3192)
  • engine: clippy clean 6 files on PR #74 (85→51 errors) (7dc7370)
  • engine: prefix cache full-hit, scheduler leak, gather bounds check (3e35b3f)
  • make pre-push pass on dust stack (2695165)
  • rename first-token detok bindings to avoid shadowing (8d16696)
  • route first streamed token through incremental detokenizer (722218d)
  • split long first doc paragraphs for clippy (c195312)
  • stabilize dust stack CI (8e629e3)

Performance Improvements

  • incremental detokenization for streaming generation (a949e50)
higgs-models: 2.0.0

2.0.0 (2026-09-19)

⚠ BREAKING CHANGES

  • higgs start is now config/profile-only, attach is a strict daemon dashboard, shellenv/exec fail fast on invalid or unreachable targets, and exact local model matches now take precedence over regex routes.

Features

  • add qwen3.5/qwen3.6 turboquant stack (4f165ee)
  • bench: context-frontier benchmark with shared cache rollback helper (PR2 / ds4 P8) (#253) (4510835)
  • bonsai-q1: packed engine scaffold with upstream MLX guard (#142) (fe43aab)
  • bonsai: run Bonsai-Q1 bits=1 on vanilla MLX via JIT qmv_fast kernels (no mlx-rs fork) (#182) (b604754)
  • cache: AnyCache::trim_by dispatcher for spec-decode rollback (#143) (229c111)
  • config: first-class config surface for the MLA latent KV cache (#263) (cdb3c4e)
  • enable MiniCPM5-1B (explicit head_dim + special-token decode + tojson kwarg + tool parser) (#176) (05464e2)
  • engine: disk-backed KV prefix store for restart-resume (PR6 / ds4 P4) (#257) (0435d11)
  • harden CLI, dashboard, routing, and MLX runtime (4dfc930)
  • models: add Gemma 3 and Gemma 4 inference support (4f301fc)
  • models: MLA latent KV cache with weight absorption for DeepSeek-V2 (PR3 / ds4 P1) (#254) (482d2ec)
  • models: per-tensor quantization settings with dense-mode loading (PR9 / ds4 P2 loader half) (#260) (92277ca)
  • quantize: MoE calibration tooling + loader fixes; asymmetric-quality claim negative (#264) (9c40cce)
  • qwen3_next: mixed-bit Qwen3.5 GDN BA loading fallback (#148) (cc18616)
  • Qwen3.5 model architecture, TurboQuant KV cache, and model-level optimizations (1514737)
  • qwen35: Qwen3.6 MoE MTP speculative decode + checkpoint memory-safety fix (#183) (4989955)
  • serve: stream prefill progress (llama.cpp-compatible prompt_progress) (#184) (e7b2c8a)

Bug Fixes

  • address coderabbit issues (6f10bc6)
  • address review items across models and engine — PR #74 (8863e2d)
  • align turboquant tests with rand 0.10 (d0617c4)
  • bonsai: apply causal mask during prefill (#203) (ecd78e5)
  • cargo fmt + restore higgs crate build on PR #74 (050af95)
  • clear lint and review blockers (9c8b878)
  • deps: update rust crate safetensors to 0.7 (#151) (a177ef6)
  • deps: update rust crate safetensors to 0.8 (#201) (27d20e2)
  • engine,models: cache config propagation, ndim guards, FSM ordering (37d2ddd)
  • engine,models: clippy clean lib targets on PR #74 (51→0 errors) (ebf3192)
  • gate rng import to tests (1161eaa)
  • honor qwen3 next gate quantization (00958e1)
  • make pre-push pass on dust stack (2695165)
  • models: clippy clean for higgs-models on PR #74 (part 2/2) (1b4470b)
  • models: clippy clean in 5/7 files on PR #74 (part 1/2) (af57c32)
  • resolve new clippy lints from stable toolchain update (#286) (8fb011c)
  • restore rng import and checkout pin (8771fa0)
  • satisfy clippy on qwen3.6 tests (ab2920a)
  • stabilize dust stack CI (8e629e3)
  • suffix float literals for new float_literal_f32_fallback lint (#230) (83029e4)
  • support mlx qwen3.6 smoke (7046849)

Performance Improvements

  • dtype: preserve fp16 through scalar multiply in deepseek_v2 + siglip (5dfbd15)
  • dtype: preserve fp16 through scalar multiply in deepseek_v2 + siglip (90aec29)
  • models: opt-in fused MoE gate+up — 3→2 expert matmuls per layer (#141) (60d7cb4)
  • sampling: partial-sort path for small top-k (67a5a06)
  • sampling: partial-sort path for small top-k (bab4519)

This PR was generated with Release Please. See documentation.

@github-actions
github-actions Bot force-pushed the release-please--branches--main branch 2 times, most recently from e097b69 to ffb4186 Compare June 24, 2026 08:14
@github-actions
github-actions Bot force-pushed the release-please--branches--main branch from ffb4186 to 025ec98 Compare August 11, 2026 08:33
@github-actions
github-actions Bot force-pushed the release-please--branches--main branch from 025ec98 to 97853ac Compare September 19, 2026 08:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants