Motivation
Agent workloads often ask several questions about the same observation. Independent full prefills can repeatedly compute an expensive image/video or text prefix.
Valen issue #29 documents a concrete example: its Qwen path shares processor encoding but repeats backbone execution per question and per Score level. Fifteen five-level Score questions require 75 explicit backbone calls. This is a source-level call count, not a measured speedup opportunity.
System1-Omni could make shared-observation serving useful across compatible decision models.
Proposed first step
Produce a focused design and benchmark plan before implementing a general cache or scheduler. Select one runnable model and freeze paired request manifests with one observation and increasing numbers of questions/candidates.
Hypothesis: reusing compatible observation computation reduces total multi-question latency while preserving reference decisions and declared numerical tolerances. Memory cost and cache-hit/miss behavior must be measured alongside latency.
Questions for community feedback
- Which reuse boundary is both correct and valuable first: preprocessing, vision features, or the language prefix's execution state?
- For Qwen's hybrid attention/recurrent layers, what state must be retained or copied to execute independent branches correctly?
- What defines compatibility and invalidation: model/adapter revision, media/preprocessing identity, prompt prefix, positions and precision?
- How should memory be bounded and state ownership handled across requests?
- Would candidate/question batching be a simpler first improvement for the selected workload?
Preserve independent question semantics. In particular, Valen Score branches see only their own level description; concatenating all levels/questions into a new prompt changes the computation.
RFC deliverables
Use the existing benchmark protocol and architecture contracts. Valen measurements depend on its reference-worker integration (#84); an already runnable text model can support an initial study.
Community help wanted
We welcome real multi-question agent traces, workload design, CPU request fixtures, profiling, and review of cache/recurrent-state correctness. Please comment with a workload or a boundary you can investigate.
Feedback is open; no implementation deadline or speedup is committed by this RFC.
Motivation
Agent workloads often ask several questions about the same observation. Independent full prefills can repeatedly compute an expensive image/video or text prefix.
Valen issue #29 documents a concrete example: its Qwen path shares processor encoding but repeats backbone execution per question and per Score level. Fifteen five-level Score questions require 75 explicit backbone calls. This is a source-level call count, not a measured speedup opportunity.
System1-Omni could make shared-observation serving useful across compatible decision models.
Proposed first step
Produce a focused design and benchmark plan before implementing a general cache or scheduler. Select one runnable model and freeze paired request manifests with one observation and increasing numbers of questions/candidates.
Hypothesis: reusing compatible observation computation reduces total multi-question latency while preserving reference decisions and declared numerical tolerances. Memory cost and cache-hit/miss behavior must be measured alongside latency.
Questions for community feedback
Preserve independent question semantics. In particular, Valen Score branches see only their own level description; concatenating all levels/questions into a new prompt changes the computation.
RFC deliverables
Use the existing benchmark protocol and architecture contracts. Valen measurements depend on its reference-worker integration (#84); an already runnable text model can support an initial study.
Community help wanted
We welcome real multi-question agent traces, workload design, CPU request fixtures, profiling, and review of cache/recurrent-state correctness. Please comment with a workload or a boundary you can investigate.
Feedback is open; no implementation deadline or speedup is committed by this RFC.