The model to consider.
LiquidAI/LFM2.5-350M, using the constrained-scoring approach from notnotsamuel/LFM2.5-350M-RLCD. Weights and reference code are pinned in the implementation. This follows my feasibility note on #9.
The closest planned or implemented model.
LFM takes a different execution path from LAYA and CLM: prefill a prompt, then score candidate continuations while branching both attention KV and convolution state. The branching method comes from the reference implementation; the weights remain unchanged.
What's your difficulty of supporting the model you want?
The first slice is a Transformers worker for text state and independent choice questions. Each question gets its own prompt and prefill, with a fixed internal field name. Candidate batches start from an unchanged prefix cache. An explicit batch size limits simultaneous branches.
The main checks are state isolation, reference scoring parity and keeping unrelated questions and external question IDs out of a question's model input. The required confidence field is documented as uncalibrated entropy concentration.
Use case and motivation
This would add a small hybrid convolution/attention model to the decision API and provide a measured memory/latency trade-off for larger candidate sets.
The first slice is implemented locally. On one L40S, FP16, 17 cases across six batch settings pass oracle checks; 24 near-tie checks have no selection flips. Direct worker/frontend parity passes 10/10 cases. At 255 candidates and a 3,877-token prefix, batch 32 reduces peak allocated memory from 17.97 to 2.88 GiB, with p50 increasing from 328.84 to 357.25 ms.
These measurements use concurrency 1, eager attention and the PyTorch convolution fallback. The slice includes inference and reproducible checks, without RL/RLCD training or native Rust/CUDA execution. I would like feedback on starting with this HF-first worker before a native follow-up.
Before submitting a new issue...
The model to consider.
LiquidAI/LFM2.5-350M, using the constrained-scoring approach from notnotsamuel/LFM2.5-350M-RLCD. Weights and reference code are pinned in the implementation. This follows my feasibility note on #9.
The closest planned or implemented model.
LFM takes a different execution path from LAYA and CLM: prefill a prompt, then score candidate continuations while branching both attention KV and convolution state. The branching method comes from the reference implementation; the weights remain unchanged.
What's your difficulty of supporting the model you want?
The first slice is a Transformers worker for text state and independent choice questions. Each question gets its own prompt and prefill, with a fixed internal field name. Candidate batches start from an unchanged prefix cache. An explicit batch size limits simultaneous branches.
The main checks are state isolation, reference scoring parity and keeping unrelated questions and external question IDs out of a question's model input. The required confidence field is documented as uncalibrated entropy concentration.
Use case and motivation
This would add a small hybrid convolution/attention model to the decision API and provide a measured memory/latency trade-off for larger candidate sets.
The first slice is implemented locally. On one L40S, FP16, 17 cases across six batch settings pass oracle checks; 24 near-tie checks have no selection flips. Direct worker/frontend parity passes 10/10 cases. At 255 candidates and a 3,877-token prefix, batch 32 reduces peak allocated memory from 17.97 to 2.88 GiB, with p50 increasing from 328.84 to 357.25 ms.
These measurements use concurrency 1, eager attention and the PyTorch convolution fallback. The slice includes inference and reproducible checks, without RL/RLCD training or native Rust/CUDA execution. I would like feedback on starting with this HF-first worker before a native follow-up.
Before submitting a new issue...